Method and device for securing a large language model
The method secures large language models by filtering requests through detection and replacement of special characters and verification against authorized word dictionaries, effectively preventing unauthorized uses and enhancing cybersecurity.
Patent Information
- Application Number
- FR2023014479
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
Large language models, particularly 'white box' models, are vulnerable to 'model evasion' attacks where malicious queries are crafted by adding tokens to bypass security frameworks, potentially leading to illegal or fraudulent uses.
A method and device for securing large language models that involves filtering incoming requests by detecting and replacing special characters, extracting and verifying fragments against a dictionary of authorized words, and deleting unauthorized fragments to form a secure query.
This approach effectively minimizes the risk of large language models being used outside their intended legal framework, preventing illegal or fraudulent activities by filtering out unauthorized queries automatically and efficiently.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for securing a large language model
[0001] The present invention relates to a method for securing a large language model, previously trained by machine learning to provide information following a request formulated in written form.
[0002] The invention also relates to a device for securing a large language model and an associated computer program.
[0003] The invention lies in the field of cybersecurity.
[0004] Recently, thanks to progress in the field of artificial intelligence, large language models, known by the acronym LLM for "Large Language Models", have been developed, which implement deep neural networks, trained by machine learning on large quantities of text, for example by self-supervised or semi-supervised learning.
[0005] Such large language models have notably been implemented in software that are conversational agents (or "chats" in English), for example ChatGPT® (for "Chat Generative Pre-trained Transformer") which allow, on the basis of a query formulated by a user in natural language, in a given language, to obtain an adequate response. This technology has demonstrated very good performance on tasks such as text summarization, writing dissertations, computer programming, and more generally on relatively complex tasks.
[0006] Of course, such conversational agents implementing large language models are programmed to provide responses within a given legal framework (they are then said to be "aligned"), and to avoid providing any information that could contribute to an illegal action, for example providing computer code to carry out computer hacking. This constraint is called model framing or model alignment.
[0007] Among the large language models currently known, some are of the "white box" type, that is to say that the architecture of the neural network forming the model, the number of layers, the type of links between the layers and the associated weightings are known. In other words, the source code of the neural network forming the model is made public. This is for example the case of the Llama® model from Meta or Alpaca® (modified version of Llama) from Stanford University.
[0008] It has been demonstrated that it is possible to attack a large language model by an attack called "model evasion", i.e. to formulate by processing machine a modified written query, also called an adversary prompt, containing suffixes that allow overcoming the template framing, and thus obtaining a response to a query outside the initially intended legal framework. Such a template evasion attack is described in the article "Universal and Transferable Adversarial Attacks on Aligned Language Models", by Andy Zou, Zifan Wang, J. Zico Kolter and Matt Fredrikson, published in 2023, https: / / arxiv.org / abs / 2307.15043.
[0009] Indeed, for large language models, for automatic processing of queries formulated in written form, it is known to apply a splitting (in English "tokenization") of character strings in the form of sub-words of consecutive characters (in English "tokens") which are generated according to their frequencies of occurrence in order to optimize the compression of the vocabulary considered, by replacing long character strings with sub-words of reduced size. Thus, an initial character string forming a sentence is transformed into a sequence of sub-words that are not intelligible to a human. For example, the compression technique using the BPE mechanism (for "Byte-Pair Encoding") makes it possible to represent the words of a natural language (for example, 600,000 different words in English) by a set of approximately 32,000 sub-words (or tokens).A template evasion attack involves adding tokens to a query, typically at the end of the query as a suffix, or at the beginning of the query as a prefix, to overcome the template framing. The added tokens are generated by machine learning for this purpose. Thus, through such a query modification, it would be possible to obtain, for example, information relating to the hacking of a banking network. In this case, a malicious third party can use a large language model for illegal purposes.
[0010] Thus, it has been shown that large "white box" language models have security flaws.
[0011] Furthermore, it is known that certain attacks are also transferable to be applied to “black box” artificial intelligence models, i.e. whose architecture and parameters are not known. It is therefore to be feared that attacks developed on large “white box” language models will be transferred to large “black box” language models, which presents a danger in terms of cybersecurity.
[0012] The object of the present invention is to remedy the drawbacks of the state of the art by proposing a method for securing a large language model making it possible to combat possible attacks of the type described above.
[0013] To this end, the invention proposes, according to one aspect, a method for securing a large language model, previously trained by machine learning. to provide information following a request made in written form. This method 1 comprises steps implemented by a calculation processor of:
[0014] - acquisition of a request in the form of an initial character string,
[0015] -first filtering of said initial character string to obtain a filtered character string, the first filtering comprising a detection of special characters in the initial character string and replacement of each special character detected by a space;
[0016] - extraction of the filtered character string from at least one fragment, in a given order of traversal, a fragment being formed from a group of characters followed and / or preceded by a space, and verification of the presence of said fragment in a dictionary of authorized words,
[0017] - in the absence of at least one fragment of the dictionary of authorized words, second filtering of the initial character string to obtain a processed character string, said second filtering comprising a deletion from the initial character string of said fragments not belonging to said dictionary of authorized words, the processed character string forming a secure query to be provided as input to said large language model.
[0018] Advantageously, the proposed security method makes it possible to filter requests so as to delete fragments generated by machine processing, in particular fragments generated on the basis of the BPE mechanism, and in this way to minimize the risk of departing from the initially planned scope of the large language model. Thus, illegal or fraudulent use of the large language model is prevented. Advantageously, the proposed method is fast and automatic.
[0019] The method for securing a large language model according to the invention may also have one or more of the characteristics below, taken independently or in all technically conceivable combinations.
[0020] The second filtering implements a deletion of each fragment not belonging to said dictionary of authorized words.
[0021] The second filtering implements a deletion of all fragments located between two fragments not belonging to said dictionary of authorized words.
[0022] The second filtering implements a deletion of all fragments following a first fragment, in the order of scanning, not belonging to said dictionary of authorized words.
[0023] The verification further comprises a verification of the presence of said fragment in a first complementary list called a white list, comprising authorized words.
[0024] The method further comprises a determination of membership of each fragment in a second complementary list called a black list, and a selection of a second filtering mode based on a result of said determination.
[0025] The method further comprises a display on a graphical interface of fragments not belonging to said dictionary of words authorized for correction by a user.
[0026] The first filtering further comprises a replacement of several successive spaces by a single space.
[0027] According to another aspect, the invention relates to a device for securing a large language model, previously trained by machine learning to provide information following a request formulated in written form. This device comprises a calculation processor configured to implement:
[0028] - a module for acquiring a request in the form of an initial character string,
[0029] - a first module for filtering said initial character string to obtain a filtered character string, the filtering comprising a detection of special characters in the character string and replacement of each special character detected by a space;
[0030] - a module for extracting the filtered character string from at least one fragment, in a given order of scanning, a fragment being formed from a group of characters followed and / or preceded by a space, and verification of the presence of said fragment in a dictionary of authorized words,
[0031] - in the event of the absence of at least one fragment of the dictionary of authorized words, a second module for filtering the initial character string to obtain a processed character string, said second filtering comprising a deletion from the initial character string of said fragments not belonging to said dictionary of authorized words, the processed character string forming a secure query to be provided as input to said large language model.
[0032] Advantageously, the device for securing a large language model is configured to implement the method for securing a large language model as briefly described above, in all its variants.
[0033] According to another aspect, the invention relates to an information recording medium, on which are stored software instructions for the execution of a method for securing a large language model as briefly described above, when these instructions are executed by a programmable electronic device.
[0034] According to another aspect, the invention relates to a computer program comprising software instructions which, when implemented by a device programmable electronics, implement a method of securing a large language model as briefly described above.
[0035] Other characteristics and advantages of the invention will emerge from the description given below, for information purposes only and in no way limiting, with reference to the appended figures, among which:
[0036] [Fig-1] [Fig.l] illustrates a device for securing a language model of large size according to one embodiment;
[0037] [Fig.2] [Fig.2] is a flowchart of the main steps of a method for securing a large language model according to a first embodiment;
[0038] [Fig.3] [Fig.3] is an example of character strings forming a query modified following the steps implemented by the method for securing a large language model;
[0039] [Fig.4] [Fig.4] is a flowchart of the additional steps of a method for securing a large language model according to a second embodiment.
[0040] The invention applies to any type of large language model or LLM, trained by machine learning to provide information following a request formulated in written form, and applied for example in conversational agent type software. In particular, the invention applies to any LLM produced in the form of a deep neural network (in English "deep learning").
[0041] [Fig.l] represents a device 2 for securing a large language model 4, also designated by LLM model.
[0042] The device 2 is a programmable electronic device, which implements a method for securing the LLM model making it possible to obtain a secure request to be provided as input to said LLM model 4.
[0043] The term secure query here designates a query processed to limit as much as possible a risk of using the LLM 4 model outside the initial framework, or in other words, a use of the LLM 4 model to obtain information outside a predetermined legal framework.
[0044] By way of non-limiting example, one of the objectives of the invention is to avoid any provision of information useful for a malicious action, for example hacking a computer system.
[0045] According to the embodiments, the LLM model 4 is executed by the device 2, or by a programmable electronic device external to the device 2, for example connected to the device 2 via a communication network.
[0046] The device 2 for securing an LLM model is a programmable electronic device and comprises, in one embodiment, a memory unit electronics 6, one or more processors 8, an input / output interface 10 and a communication interface 12, these elements being configured to communicate with each other via a communication bus 15 internal to the device 2. The input / output interface 10 comprises for example a display screen and a character input device, for example a keyboard, allowing a user to enter queries in written form, to interrogate the LLM model 4.
[0047] In one embodiment, the electronic memory 6 stores a dictionary 14 of authorized words in one or more predetermined natural language(s).
[0048] According to variants, the dictionary 14 of authorized words is stored in a database external to the device 2, the device 2 being connected to the external database, for example through the communication interface 12.
[0049] In certain embodiments, the electronic memory 6 further stores a first complementary list 16, also called white list 16, of authorized words, which supplements the dictionary 14 of authorized words, and a second complementary list 18, also called black list 18, of words considered dangerous, which could potentially lead to illegal requests, going beyond the initial framework of the model.
[0050] The calculation processor 8 of the device 2 is configured to execute:
[0051] - a module 20 for acquiring a request in the form of an initial character string,
[0052] - a first filtering module 22, configured to filter the character string initial to obtain a filtered character string, the filtering including detection of special characters in the character string and replacement of each special character detected by a space;
[0053] - a module 24 for extracting the filtered character string from at least one fragment, in a given order of traversal, a fragment being formed from a group of characters followed and / or preceded by a space, and for verifying the presence of the fragment in the dictionary of authorized words 14, and
[0054] - a second module 26 for filtering the initial character string to obtain a processed character string, the second filtering comprising a deletion from the initial character string of at least one fragment not belonging to said dictionary of authorized words.
[0055] According to variants, the presence of the extracted fragment(s) in the first complementary list 16 and / or in the second complementary list 18 is verified.
[0056] Several variants of implementations of the second filtering module 26 are envisaged, as described in more detail below.
[0057] The processed character string obtained by the implementation of these modules forms a secure query to be provided as input to the LLM 14 model.
[0058] In one embodiment, the modules 20, 22, 24, 26 are produced in the form of software instructions forming a computer program, which, when executed by a programmable electronic device, implements a method for securing a large language model as described.
[0059] In a variant not shown, the modules 20, 22, 24, 26 are each produced in the form of programmable logic components, such as FPGAs (Field Programmable Gate Arrays) of microprocessors, GPGPU components (General-purpose processing on graphics processing), or even dedicated integrated circuits, such as ASICs (Application Specific Integrated Circuits)•
[0060] The computer program comprising software instructions is further capable of being recorded on a non-transitory, computer-readable information recording medium. This computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, this medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example EPROM, EEPROM, FLASH, NVRAM), a magnetic card or an optical card.
[0061] [Fig.2] is a flowchart of the main steps of a first embodiment of the method for securing a large language model. The method is implemented by a processor of a programmable electronic device 2 as described with reference to [Fig.l].
[0062] The method comprises a first step 30 of acquiring a request, addressed to the LLM model, in the form of an initial character string CH-i.
[0063] For example, the initial character string CH-i is provided via an FO interface, by user input or transmitted by another software application.
[0064] The expression character string refers to a sequence of characters coded according to a chosen coding format, in a set of characters comprising letters of a given alphabet, for example the Latin alphabet, and grouped to form words in a language or possibly several predetermined languages which are written with the alphabet used, separated by space characters or spaces, as well as numbers and so-called special characters.
[0065] Special characters include punctuation characters, and also symbols whose meaning is known (for example, symbols representing currencies, physical units, etc.). Special characters are listed beforehand.
[0066] The request is therefore formulated in written form.
[0067] During step 30 the initial character string CH-i is stored in a memory unit of the programmable electronic device implementing the method.
[0068] The method then comprises a step 32 of implementing a first filtering of the initial character string CH-i.
[0069] The first filtering 32 includes a detection of special characters in the initial character string CH-i, and a replacement of each special character detected by a space.
[0070] Optionally, the first filtering step 32 includes replacing successive spaces with a single space.
[0071] At the output of the first filtering step 32, a filtered character string CH-f is obtained.
[0072] [Fig.3] illustrates an example of an initial character string CH-i, a first filtered character string CH-f 1 resulting from the initial character string CH-i after removing special characters, and a second filtered character string CH-f2 obtained from the first filtered character string CH-f 1 by replacing multiple spaces with single spaces.
[0073] The method further comprises a step 34 of separating the filtered character string into fragments, each fragment being a grouping of characters preceded and / or followed by a space.
[0074] A result of this step 34 is illustrated in the example of [Fig.3], where the third character string Ch-f3 represents the second filtered character string Ch-f2 in which each fragment is surrounded.
[0075] The method then comprises an extraction 36 of a fragment in a browsing of the fragments in a given browsing order, and for each extracted fragment, a verification 38 of the presence of the fragment in the dictionary of authorized words, optionally increased by the words of the first complementary list or white list.
[0076] In this optional case, the presence in the dictionary of authorized words or in the first complementary list (white list) means that the processed fragment is an authorized fragment.
[0077] The order of reading is typically the order of reading the words in the language used.
[0078] In the event of a positive presence verification, or, in other words, if the fragment considered is an authorized word, the presence verification step 38 is followed by the extraction 36 of a following fragment, up to the end of the list of fragments of the filtered character string.
[0079] If the fragment is absent from the dictionary of authorized words (negative verification at step 38), it is deduced that the fragment is unauthorized.
[0080] According to one embodiment, steps 36 and 38 are applied to all fragments, and unauthorized fragments are marked.
[0081] In the example of [Fig.3], in the third filtered character string Ch-f3, the fragments absent from the authorized word dictionary, respectively “YESfr” and “Ntt”, are boxed, and in the fourth filtered character string CH-f4 the authorized fragments are highlighted.
[0082] The method comprises a second filtering 40 of the initial character string, by deleting at least the fragment absent from the dictionary of authorized words or unauthorized fragment.
[0083] According to a first mode of the second filtering, the fragments of the filtered character string are scanned, and each fragment absent from the dictionary of authorized words is deleted during the second filtering 40. The result of this operation is then a processed character string CH-T, comprising the words from the dictionary of authorized words of the initial character string and the special characters.
[0084] In the example of [Fig.3], the processed character string CH-T in fact includes the words and special characters of the initial character string CH-i, without the fragments “YESfr” and “Ntt” which were detected and deleted.
[0085] Advantageously, in this variant, the special characters are preserved, because in certain cases, they are necessary for the linguistic interpretation of the request formulated by the user.
[0086] According to a second embodiment of the second filtering, all fragments between two unauthorized fragments are deleted, in addition to the unauthorized fragments. In this second embodiment, fragments corresponding to words in the dictionary of authorized words located between two fragments absent from the dictionary are also deleted. Thus, this second mode of the second filtering applies a more severe modification than the first mode described above.
[0087] According to a third mode of the second filtering, all fragments appearing after a first detected unauthorized fragment, following the chosen traversal order, are deleted. Thus, this third mode of the second filtering applies a more severe modification than the first and second modes described above. In this third variant, the processing is simplified because the traversal of the fragments can be stopped as soon as an unauthorized fragment is detected.
[0088] Furthermore, in each of the modes of the second filtering described above, it is also possible to also delete the special characters either included between two fragments absent from the dictionary, or positioned after the first fragment absent from the dictionary detected.
[0089] Finally, as an optional addition, it is possible to display before the second filtering the unauthorized fragments detected, alone or underlined in the initial request, and ask the user to correct. Indeed, the process as described above is sensitive to typographical errors, which is all the more critical in languages with accented script, such as French. In this case, misspelled dictionary words are likely to be removed by the second filtering, even though this is not a malicious intention on the part of the user.
[0090] According to classic variants, it is possible to automatically propose a spelling correction in relation to existing words in the dictionary of authorized words, possibly increased by the first complementary list.
[0091] [Fig.4] is a flowchart of the additional steps of a second embodiment of the method for securing a large language model.
[0092] The steps common to the first embodiment are not described again.
[0093] At the end of step 34 of separation into fragments, the method comprises, in this second embodiment, the following additional steps.
[0094] The fragments are scanned (step 42) and it is determined in step 44 whether one of the fragments belongs to the second complementary list or black list. This then means that the query includes words considered as being able to lead to illegal queries, outside the initial framework of the model. For example, the black list includes words like "pirate", "piracy".
[0095] While not all requests containing these words are illegal, it is understood that such requests require special attention.
[0096] Depending on the result of determination 44, a second filtering mode is selected in selection step 46.
[0097] For example, if one or more of the fragments of the request belong to the blacklist, the third mode of second filtering, which is the most severe, is applied. Alternatively, the second mode of the second filtering is applied.
[0098] If none of the request fragments belong to the blacklist, the first mode of second filtering is applied, in which only unauthorized fragments are deleted.
[0099] The method for securing a large language model has been described above in steps; it is clear that some of the steps can be executed simultaneously or cooperatively for computational optimizations, such optimizations being within the reach of those skilled in the art.
[0100] The method of securing a large language model has been described with reference to a dictionary of authorized words in one or more languages.
[0101] It is clear that this method applies in a similar manner for software programming languages, by implementing a dictionary or a first list complementary to the reserved words of the programming language which are authorized, in order to then distinguish authorized fragments from unauthorized fragments.
[0102] Advantageously, the proposed method is simple, automatic and efficient, and does not require the implementation of complex interpretations of sentence or grammatical structures, nor of complex models, while being effective in avoiding use outside the initial scope of the large language model. Thus, the cybersecurity of large language models is improved.
[0103] Advantageously, when the first mode of the second filtering is implemented, the special characters are preserved, which makes it possible to maintain good performance of the LLM model.
Claims
Claims
1. Method for securing a large language model (4), previously trained by machine learning to provide information following a request formulated in written form, characterized in that it comprises steps implemented by a calculation processor of: - acquisition (30) of a request in the form of an initial character string (CH-i), - first filtering (32) of said initial character string to obtain a filtered character string (CH-f), the first filtering comprising a detection of special characters in the initial character string and replacement of each special character detected by a space;- extraction (34, 36) of the filtered character string of at least one fragment, in a given order of traversal, a fragment being formed of a group of characters followed and / or preceded by a space, and verification (38) of a presence of said fragment in a dictionary of authorized words, - in the event of absence of at least one fragment from the dictionary of authorized words, second filtering (40) of the initial character string to obtain a processed character string (CH-T), said second filtering (40) comprising a deletion from the initial character string of said fragments not belonging to said dictionary of authorized words, the processed character string (CH-T) forming a secure query to be provided as input to said large language model.;
2. Method according to claim 1, wherein said second filtering (40) implements a deletion of each fragment not belonging to said dictionary of authorized words.
3. Method according to claim 2, wherein said second filtering (40) implements a deletion of all fragments located between two fragments not belonging to said dictionary of authorized words.
4. Method according to claim 2, wherein said second filtering (40) implements a deletion of all fragments following a first fragment, in the order of scanning, not belonging to said dictionary of authorized words.
5. Method according to any one of claims 1 to 4, in which the verification (38) further comprises a verification of a presence of said fragment in a first complementary list (16) called a white list, comprising authorized words.
6. Method according to one of claims 1 to 5, further comprising a determination (44) of membership of each fragment in a second complementary list (18) called black list, and a selection (46) of a second filtering mode as a function of a result of said determination.
7. Method according to any one of claims 1 to 6, further comprising a display on a graphical interface of the fragments not belonging to said dictionary of words authorized for correction by a user.
8. A method according to any one of claims 1 to 7, wherein the first filtering (32) further comprises replacing several successive spaces with a single space.
9. Computer program comprising software instructions which, when executed by a programmable electronic device, implement a method for securing a large language model according to claims 1 to Q.
10. o. Device for securing a large language model, previously trained by machine learning to provide information following a request formulated in written form, characterized in that it comprises a calculation processor (8) configured to implement: - a module (20) for acquiring a request in the form of an initial character string (CH-i), - a first module (22) for filtering said initial character string to obtain a filtered character string (CH-f), the filtering comprising a detection of special characters in the character string and replacement of each special character detected by a space, - a module (24) for extracting the filtered character string from at least one fragment, in a given order of scanning, a fragment being formed from a group of characters followed and / or preceded by a space, and for verifying the presence of said fragment in a dictionary of authorized words, - in the event of the absence of at least one fragment of the dictionary of authorized words, a second filtering module (26) of the initial character string to obtain a processed character string (CH-T), said second filtering comprising a deletion from the initial character string of said fragments not belonging to said dictionary of authorized words, the processed character string (CH-T) forming a secure query to be provided as input to said large language model.
Citation Information
Patent Citations
Systems and methods for extracting meaning from speech-to-text data
US20130018895A1