Natural language processing system encryption method, electronic equipment and computer readable medium
The natural language processing system is encrypted through character and Token ID mapping relationship, which solves the problems of high data exposure risk and low computing efficiency in traditional systems, and realizes efficient and secure ciphertext processing, improving the system's data security and computing performance.
Patent Information
- Application Number
- CN202510855011.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional natural language processing systems have problems such as high exposure risk and low computing efficiency, especially when facing side channel attacks and intermediate eavesdropping, it is difficult to effectively protect data security. At the same time, the calculation overhead of homomorphic encryption schemes is too large to meet real-time requirements.
By obtaining the character mapping relationship determined by the character encryption algorithm and the Token ID transformation algorithm, the mapping relationship between the original characters and the encrypted characters and the Token ID is established, and the pre-trained natural language processing system is encrypted, so that it can output ciphertext responses based on the ciphertext prompt word, avoid plaintext data exposure and reduce computational logic reconstruction.
It significantly reduces the exposure risk of raw data, improves data security, reduces computing overhead, improves computing efficiency and real-time response, and eliminates the need to retrain natural language processing systems.
Smart Images

Figure CN120378222A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to an encryption method for a natural language processing system, an electronic device, and a computer-readable medium. Background Art
[0002] Natural language processing is an important branch of artificial intelligence (AI) models. One important way of human-computer interaction is to output natural language responses based on the prompt words input by users. In terms of data security, traditional natural language processing systems have many problems such as a high risk of exposure of raw data and low computational efficiency. Summary of the Invention
[0003] This application aims to solve one of the technical problems in the related art to a certain extent. For this purpose, this application provides an encryption method for a natural language processing system, an electronic device, and a computer-readable medium.
[0004] As the first aspect of this application, an encryption method for a natural language processing system is provided, where the method includes: Obtain the character mapping relationship between the original character and the encrypted character determined based on the character encryption algorithm; Determine the Token ID mapping relationship between the original Token ID and the transformed Token ID according to the Token ID transformation algorithm; Encrypt the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, so that the encrypted natural language processing system can output a ciphertext response based on the input ciphertext prompt word.
[0005] Optionally, after encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, the method further includes: Receive the ciphertext prompt word sent by the gateway associated with this device; where the ciphertext prompt word is obtained by the gateway based on the character mapping relationship to transform the plaintext prompt word sent by the user device; Input the ciphertext prompt word into the encrypted natural language processing system to obtain the ciphertext response output by the encrypted natural language processing system, so that the gateway decrypts the ciphertext response to obtain the plaintext response and sends the plaintext response to the user device.
[0006] Optionally, the natural language processing system includes a natural language processing model and its associated text processing module. Encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship includes: Performing character encryption conversion and spatial reconstruction of Token IDs on the original word segmentation table of the text processing module according to the character mapping relationship and the Token ID mapping relationship to obtain an encrypted word segmentation table; and Performing spatial reconstruction of Token IDs on the embedding vector mapping layer and the output decoding layer of the natural language processing model according to the Token ID mapping relationship.
[0007] Optionally, the original word segmentation table includes the correspondence between original Tokens and original Token IDs. Performing character encryption conversion and spatial reconstruction of Token IDs on the original word segmentation table of the text processing module according to the character mapping relationship and the Token ID mapping relationship to obtain an encrypted word segmentation table includes: Replacing each original character included in each original Token in the original word segmentation table with the encrypted character corresponding to the original character in the character mapping relationship, and replacing each original Token ID in the original word segmentation table with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship to obtain an encrypted word segmentation table.
[0008] Optionally, the natural language processing model further includes an inference layer. The embedding vector mapping layer is used to determine an embedding vector sequence according to the vector mapping relationship from the original Token ID to the embedding vector and the Token ID sequence input by the text processing module. The inference layer is used to determine a hidden state sequence according to the embedding vector sequence input by the embedding vector mapping layer. The output decoding layer is used to calculate a predicted original Token ID sequence according to the output weight mapping relationship from the hidden state to each original Token ID and the hidden state sequence input by the inference layer; Performing spatial reconstruction of Token IDs on the embedding vector mapping layer and the output decoding layer of the natural language processing model according to the Token ID mapping relationship includes: Replacing each original Token ID in the vector mapping relationship with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship; and Replace each original Token ID in the output weight mapping relationship with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship.
[0009] Optionally, after obtaining the character mapping relationship between the original characters and the encrypted characters determined based on the character encryption algorithm, the character mapping relationship is also stored locally; after encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, the method further includes: Obtain the updated character mapping relationship between the original characters and the encrypted characters re-determined based on the character encryption algorithm; According to the character mapping relationship stored locally, perform a rollback process on the current word segmentation table of the text processing module in the encrypted natural language processing system to obtain a rollback word segmentation table; Perform character encryption conversion on the rollback word segmentation table according to the updated character mapping relationship.
[0010] Optionally, after encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, the method further includes: When a preset switching period arrives or a session context switch is recognized, obtain the updated Token ID mapping relationship between the original Token ID and the transformed Token ID; According to the updated Token ID mapping relationship, perform a spatial reconstruction of the Token ID on the current word segmentation table of the text processing module, the embedding vector mapping layer, and the output decoding layer of the natural language processing model in the encrypted natural language processing system.
[0011] Optionally, the natural language processing model includes any one of the following: large language model LLM, multi-modal large language model MLLM.
[0012] As a second aspect of the present application, there is provided an electronic device, where the electronic device includes: One or more processors; A memory having one or more computer programs stored thereon, and when the one or more computer programs are executed by the one or more processors, the one or more processors implement the natural language processing system encryption method according to the first aspect of the present application.
[0013] As a third aspect of the present application, a computer-readable medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the natural language processing system encryption method according to the first aspect of the present application.
[0014] The natural language processing system encryption method provided by the embodiments of the present application determines the character mapping relationship between the original characters and the encrypted characters based on the character encryption algorithm, determines the Token ID mapping relationship between the original Token ID and the transformed Token ID according to the Token ID transformation algorithm, and encrypts the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, so that the encrypted natural language processing system can output a ciphertext response based on the input ciphertext prompt. It can not only enable the encrypted natural language processing system to directly process the ciphertext prompt, realize ciphertext processing from the input end to the output end, avoid the exposure of plaintext data in memory, resist side-channel attacks and intermediate eavesdropping, thereby significantly reducing the exposure risk of the original data and improving data security, but also avoid mathematically reconstructing the computational logic of the system, thus significantly reducing the computational overhead, reducing the inference latency, improving the computational efficiency and response real-time performance. Moreover, there is no need to retrain the natural language processing system, but directly encrypt the already trained natural language processing system, thereby further reducing the computational overhead and improving the efficiency of encrypting the natural language processing system. Brief Description of the Drawings
[0015] The present application will be further described below with reference to the accompanying drawings: Figure 1 is a flowchart of an implementation manner of the natural language processing system encryption method provided by the embodiments of the present application; Figure 2 is a flowchart of another implementation manner of the natural language processing system encryption method provided by the embodiments of the present application Figure 3 is a flowchart of yet another implementation manner of the natural language processing system encryption method provided by the embodiments of the present application Figure 4 is a flowchart of still another implementation manner of the natural language processing system encryption method provided by the embodiments of the present application Figure 5 is a flowchart of another implementation manner of the natural language processing system encryption method provided by the embodiments of the present application; Figure 6 is a flowchart of yet another implementation manner of the natural language processing system encryption method provided by the embodiments of the present application; Figure 7It is a flowchart of yet another implementation manner of the natural language processing system encryption method provided by the embodiments of the present application; Figure 8 It is a module diagram of an implementation manner of the electronic device provided by the embodiments of the present application; Figure 9 It is a schematic diagram of the computer-readable medium provided by the embodiments of the present application.
[0016] Description of the reference numerals 101: Processor 102: Memory 103: I / O interface 104: Bus Detailed implementation manner The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. Based on the embodiments in the implementation manner, it is intended to explain the present application and should not be construed as a limitation to the present application.
[0017] The reference to "one embodiment" or "example" or "instance" in this specification means that the specific features, structures, or characteristics described in connection with the embodiment itself can be included in at least one embodiment disclosed in the present application. The appearance of the phrase "in one embodiment" at various positions in the specification does not necessarily refer to the same embodiment.
[0018] Natural language processing is an important branch of artificial intelligence (AI) models. One important way of human-computer interaction is to output natural language responses based on the prompts input by users. In terms of data security, natural language processing systems face risks such as the input and output being easily induced or modified, and sensitive data being easily leaked during the transmission process or the model processing process. Traditional natural language processing systems usually adopt encryption transmission - decryption processing solutions or homomorphic encryption solutions to deal with this.
[0019] The encryption transmission - decryption processing solution is applicable to natural language processing systems that only support plain - text input. When the user encrypts the prompt words with the public key and then inputs them into the natural language processing system, the natural language processing system must use the private key to restore the user input before processing. As a result, sensitive data is likely to be exposed in memory or side - channel attacks during the calculation process, and the risk of exposure of the original data is relatively high. The homomorphic encryption scheme supports the natural language processing system to directly process the ciphertext prompt words. However, its essence is a mathematical transformation based on number - theory cryptography, and it is necessary to mathematically reconstruct the calculation logic of the model in the system (such as matrix multiplication, activation function) so that it can run on ciphertext (for example, replacing plain - text addition with homomorphic addition). This scheme has an excessive computational overhead (for example, the delay of the additive homomorphic encryption Paillier scheme increases by more than 2000%), and it simply cannot meet the real - time requirements, with relatively low computational efficiency.
[0020] Based on the above - mentioned important findings, the applicant of this application innovatively proposes to perform encryption mapping on characters and at the same time perform transformation mapping on Token IDs, and encrypt the natural language processing system based on the character mapping relationship and the Token ID mapping relationship. This can not only enable the encrypted natural language processing system to directly process the ciphertext prompt words, thereby significantly reducing the risk of exposure of the original data and enhancing data security, but also avoid mathematically reconstructing the calculation logic of the system, thus significantly reducing the computational overhead, enhancing the computational efficiency and response real - time performance.
[0021] As the first aspect of the embodiments of this application, a method for encrypting a natural language processing system is provided. As Figure 1 shown, the method includes: Step S110, obtaining the character mapping relationship between the original characters and the encrypted characters determined based on the character encryption algorithm; Step S120, determining the Token ID mapping relationship between the original Token ID and the transformed Token ID according to the Token ID transformation algorithm; Step S130, encrypting the pre - trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, so that the encrypted natural language processing system can output a ciphertext response based on the input ciphertext prompt words.
[0022] Among them, the embodiments of this application do not make special limitations on the execution order between step S110 and step S120. That is to say, the two can be executed sequentially or in parallel, and either of them can be executed first when they are executed sequentially.
[0023] Among them, the embodiments of the present application do not make specific limitations on how to execute step S110. For example, an original character set that supports the Unicode standard (including terms, symbols, etc. in the corresponding industry, and characters can be extended if necessary) can be constructed according to the actual industry scenarios of natural language processing (such as intelligent question answering in the financial field, medical health consultation, enterprise confidential document generation, etc.). Then, based on a character encryption algorithm, each original character in the original character set is encrypted to obtain the character mapping relationship between each original character and its corresponding encrypted character.
[0024] Among them, the embodiments of the present application do not make specific limitations on the character encryption algorithm. For example, a high-complexity transformation parameter can be generated based on a parameter derivation algorithm (such as the HMAC-based Key Derivation Function (HKDF)), and encryption can be performed in combination with encryption algorithms such as format encoding, compression algorithms, or privacy protection algorithms.
[0025] Among them, the embodiments of the present application do not make special limitations on whether all the original characters in the original character set are encrypted using exactly the same character encryption algorithm. As long as it is ensured that one original character is uniquely mapped to one encrypted character and one encrypted character is also uniquely mapped to one original character. Of course, in order to prevent data ambiguity or security vulnerabilities caused by collisions, it can also be considered to further ensure that the collision probability between different original characters and is less than a certain threshold, for example , so as to ensure that the possibility of different original characters being mapped to the same encrypted character is also extremely low.
[0026] Among them, the embodiments of the present application do not make specific limitations on the Token ID transformation algorithm. For example, a hash function, a random number generator, etc. can be used.
[0027] Among them, it should be emphasized that step S130 directly encrypts the pre-trained natural language processing system based on the character mapping relationship and the Token ID mapping relationship. That is to say, the embodiments of the present application do not have to involve the training of the natural language processing system, and can directly obtain the pre-trained natural language processing system from the outside and encrypt it.
[0028] The natural language processing system encryption method provided by the embodiments of the present application obtains the character mapping relationship between the original characters and the encrypted characters determined based on the character encryption algorithm, determines the Token ID mapping relationship between the original Token ID and the transformed Token ID according to the Token ID transformation algorithm, and encrypts the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, so that the encrypted natural language processing system can output a ciphertext response based on the input ciphertext prompt. It can not only enable the encrypted natural language processing system to directly process the ciphertext prompt, realize ciphertext processing from the input end to the output end, avoid the exposure of plaintext data in the memory, resist side-channel attacks and middle eavesdropping, thereby significantly reducing the exposure risk of the original data and enhancing data security, but also avoid mathematical reconstruction of the system's computing logic, thus significantly reducing the computing overhead, reducing the inference latency, enhancing the computing efficiency and response real-time performance. Moreover, there is no need to retrain the natural language processing system, but directly encrypt the already trained natural language processing system, thereby further reducing the computing overhead and enhancing the efficiency of encrypting the natural language processing system.
[0029] The natural language processing system encryption method provided by the embodiments of the present application can reduce the computing overhead by at least 90% compared with the homomorphic encryption scheme, and the inference latency is less than 10 ms.
[0030] The applicant of the present application further proposes that by adding a gateway associated with the device (i.e., the device loaded or installed with the natural language processing system), the gateway converts the plaintext prompt sent by the user device into a ciphertext prompt based on the character mapping relationship, inputs the ciphertext prompt into the encrypted natural language processing system, obtains the ciphertext response output by the encrypted natural language processing system, and finally the gateway decrypts the ciphertext response to obtain the plaintext response and sends the plaintext response to the user device. In this way, for the user device, there is no need to encrypt the prompt, and sending the plaintext prompt can receive the plaintext response. The human-computer interaction process is transparent and burden-free, which can not only enhance data security but also enhance the user's human-computer interaction experience.
[0031] Correspondingly, in some embodiments, after encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship (i.e., involved in step S130), as Figure 2 shown, the method further includes: Step S140, receiving the ciphertext prompt sent by the gateway associated with the device; wherein, the ciphertext prompt is obtained by the gateway converting the plaintext prompt sent by the user device based on the character mapping relationship; Step S150: Input the ciphertext prompt into the encrypted natural language processing system to obtain the ciphertext response output by the encrypted natural language processing system, so that the gateway decrypts the characters of the ciphertext response to obtain the plaintext response and sends the plaintext response to the user device.
[0032] Among them, in the embodiments of the present application, there is no special limitation on whether the gateway associated with this device is independent of this device. That is to say, the gateway may include an independent device installed on the communication link between the user device and this device, or may include a module integrated with the gateway function in this device.
[0033] Among them, in the embodiments of the present application, all input data is transmitted to this device through the gateway associated with this device, and all output data is transmitted to the user device through the gateway associated with this device. The gateway intercepts and encrypts (in addition, it can also perform analysis) the input data to ensure that the data is transmitted to the encrypted natural language processing system for processing in an encrypted state. The gateway also decrypts (in addition, it can also perform compliance detection, desensitization processing, etc.) the output data to ensure that the response received by the user device can be directly read.
[0034] The natural language processing system encryption method provided by the embodiments of the present application obtains the character mapping relationship between the original characters and the encrypted characters determined based on the character encryption algorithm, determines the Token ID mapping relationship between the original Token ID and the transformed Token ID according to the Token ID transformation algorithm, and encrypts the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship. The encrypted natural language processing system can output a ciphertext response based on the input ciphertext prompt. Receive the ciphertext prompt sent by the gateway associated with this device, where the ciphertext prompt is obtained by the gateway based on the character mapping relationship to convert the plaintext prompt sent by the user device. Input the ciphertext prompt into the encrypted natural language processing system to obtain the ciphertext response output by the encrypted natural language processing system, so that the gateway decrypts the ciphertext response to obtain the plaintext response and sends the plaintext response to the user device. It can not only enable the encrypted natural language processing system to directly process the ciphertext prompt, realize ciphertext processing from the input end to the output end, avoid the exposure of plaintext data in memory, resist side-channel attacks and middle eavesdropping, thereby significantly reducing the exposure risk of the original data and enhancing data security, but also avoid mathematical reconstruction of the system's calculation logic, thereby significantly reducing the calculation overhead, reducing the inference latency, enhancing the calculation efficiency and response real-time performance. Moreover, there is no need to retrain the natural language processing system, but directly encrypt the already trained natural language processing system, thereby further reducing the calculation overhead and enhancing the efficiency of encrypting the natural language processing system. It can also enable the user not to encrypt the prompt, input the plaintext prompt and obtain the plaintext response, and the human-computer interaction process is transparent and burden-free, thereby enhancing the user's human-computer interaction experience.
[0035] The applicant of the present application further proposes that without mathematically reconstructing the calculation logic of the natural language processing system, based on the character mapping relationship and the Token ID mapping relationship, perform character encryption conversion and Token ID space reconstruction on the original word segmentation table of the text processing module associated with the natural language processing model in the natural language processing system, and perform Token ID space reconstruction on the embedding vector mapping layer and the output decoding layer of the natural language processing model based on the Token ID mapping relationship, so as to realize the encryption of the natural language processing system.
[0036] Correspondingly, in some embodiments, the natural language processing system includes a natural language processing model and its associated text processing module. Encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship (i.e., involved in step S130), as Figure 3 shown, includes: Step S131, according to the character mapping relationship and the Token ID mapping relationship, perform character encryption conversion and Token ID space reconstruction on the original word segmentation table of the text processing module to obtain an encrypted word segmentation table; Step S132, according to the Token ID mapping relationship, perform Token ID space reconstruction on the embedding vector mapping layer and the output decoding layer of the natural language processing model.
[0037] Among them, in step S131, character encryption conversion of the original word segmentation table is involved based on the character mapping relationship, and Token ID space reconstruction of the original word segmentation table is involved based on the Token ID mapping relationship. The embodiments of the present application do not make special limitations on the execution order between the two. That is to say, the two can be executed successively or in parallel, and either one can be executed first when the two are executed successively.
[0038] Among them, the embodiments of the present application do not make special limitations on the execution order between step S131 and step S132. That is to say, the two can be executed successively or in parallel, and either one can be executed first when the two are executed successively.
[0039] Among them, the embodiments of the present application do not make specific limitations on the text processing module. For example, when the natural language processing model is a Large Language Model (LLM), the associated text processing module is a Tokenizer.
[0040] Among them, it can be understood that in addition to the embedding vector mapping layer and the output decoding layer, the natural language processing model may also include a preprocessing layer, an inference layer (Transformer architecture), a postprocessing layer, etc., which are not elaborated in the embodiments of the present application.
[0041] The natural language processing system encryption method provided by the embodiments of the present application realizes an encryption of the entire natural language processing system by performing character encryption conversion on the original word segmentation table of the text processing module according to the character mapping relationship, enabling the text processing module to have the ability of ciphertext word segmentation, and further enabling the natural language processing model to have the ability of ciphertext inference; it also performs Token ID space reconstruction on the original word segmentation table, the embedding vector mapping layer and the output decoding layer of the natural language processing model according to the Token ID mapping relationship, realizing a secondary encryption of the entire natural language processing system, ensuring that the original Token ID cannot be used for inference even if it is attacked and leaked.
[0042] Although the natural language processing model and its associated text processing module are independent of each other, their functions are inseparable. The text processing module is used to: split a continuous text sequence into discrete lexical units (i.e., Tokens, which can be words, sub-words, or characters, depending on the design of the natural language processing model and the tokenization strategy used during training), map the tokenized Tokens to a sequence of Token IDs according to a tokenization table and provide it to the LLM for processing, and finally map the predicted sequence of Token IDs output by the LLM back to text-form Tokens according to the tokenization table. In the embodiments of the present application, the tokenization table includes a static mapping table or a dynamically generated vocabulary, and the character encryption conversion thereof does not depend on a specific tokenization algorithm, but is based on a character mapping relationship.
[0043] In the embodiments of the present application, by performing character encryption conversion on the original tokenization table of the tokenization table, the text processing module is enabled to have the ability to perform ciphertext tokenization, and thus the natural language processing model is enabled to have the ability to perform ciphertext inference. Correspondingly, in some embodiments, the original tokenization table includes the correspondence between the original Token and the original Token ID. According to the character mapping relationship and the Token ID mapping relationship, character encryption conversion and spatial reconstruction of the Token ID are performed on the original tokenization table of the text processing module to obtain an encrypted tokenization table (i.e., the one involved in step S131), as Figure 4 shown, including: Step S1311: Replace each original character included in each original Token in the original tokenization table with the encrypted character corresponding to the original character in the character mapping relationship, and replace each original Token ID in the original tokenization table with the transformed TokenID corresponding to the original Token ID in the Token ID mapping relationship to obtain an encrypted tokenization table.
[0044] It can be understood that a Token is composed of characters, but in the embodiments of the present application, there is no special limitation on the number of characters included in each original Token, which is actually determined by the tokenization logic.
[0045] It can be understood that the finally obtained encrypted tokenization table includes the correspondence between the encrypted Token and the transformed Token ID.
[0046] As described above, the text processing module provides a sequence of Token IDs to a natural language processing model for processing. Correspondingly, in some embodiments, the natural language processing model further includes an inference layer. The embedding vector mapping layer is configured to determine an embedding vector sequence according to the vector mapping relationship from the original Token ID to the embedding vector and the sequence of Token IDs input by the text processing module. The inference layer is configured to determine a hidden state sequence according to the embedding vector sequence input by the embedding vector mapping layer. The output decoding layer is configured to calculate a predicted sequence of original Token IDs according to the output weight mapping relationship from the hidden state to each original Token ID and the hidden state sequence input by the inference layer; Performing spatial reconstruction of Token IDs on the embedding vector mapping layer and the output decoding layer of the natural language processing model according to the Token ID mapping relationship (i.e., involved in step S132), as Figure 5 shown, includes: Step S1321: Replace each original Token ID in the vector mapping relationship with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship; and Step S1322: Replace each original Token ID in the output weight mapping relationship with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship.
[0047] Wherein, in the embodiments of the present application, no special limitation is imposed on the execution order between step S1321 and step S1322. That is to say, the two can be executed successively or in parallel, and either of them can be executed first when the two are executed successively.
[0048] The applicant of the present application further proposes that the latest character mapping relationship can be dynamically obtained to update the word segmentation table, further improving the data security of the natural language processing system. Correspondingly, in some embodiments, after obtaining the character mapping relationship between the original character and the encrypted character determined based on the character encryption algorithm (i.e., involved in step S1110), the character mapping relationship is also stored locally; after encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship (i.e., involved in step S130), as Figure 6 shown, the method further includes: Step S210: Obtain an updated character mapping relationship between the original character and the encrypted character re-determined based on the character encryption algorithm; Step S220: According to the character mapping relationship stored locally, perform a rollback process on the current word segmentation table of the text processing module in the encrypted natural language processing system to obtain a rollback word segmentation table. Step S230: According to the updated character mapping relationship, perform character encryption conversion on the rollback word segmentation table.
[0049] It can be understood that the "updated character mapping relationship" is different from the "character mapping relationship" obtained in step 110 above.
[0050] Among them, the rollback process means that the current word segmentation table of the text processing module includes the correspondence between encrypted Tokens and transformed Token IDs. According to the character mapping relationship stored locally, the encrypted Tokens are reversely mapped back to the original Tokens to obtain a rollback word segmentation table including the correspondence between the original Tokens and the transformed Token IDs.
[0051] Among them, step S230 means that each original character included in each original Token in the rollback word segmentation table is replaced with the encrypted character corresponding to it in the updated character mapping relationship.
[0052] The applicant of the present application further proposes that by storing the character mapping relationship of each historical version locally (that is, the character mapping relationship obtained each time step S110 is executed), in this way, it is possible to roll back to the word segmentation table of any historical encryption period.
[0053] The applicant of the present application also proposes that encrypting the natural language processing system based on the character mapping relationship and encrypting the natural language processing system based on the Token ID mapping relationship are independent and do not interfere with each other. That is, not only can the character mapping relationship be dynamically updated, but the Token ID mapping relationship can also be dynamically updated. Correspondingly, in some embodiments, after encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship (that is, involved in step S130), as Figure 7 shown, the method further includes: Step S310: When a preset switching period arrives or a session context switch is recognized, obtain the updated Token ID mapping relationship between the original Token ID and the transformed Token ID; Step S320: According to the updated Token ID mapping relationship, perform spatial reconstruction of Token IDs on the current word segmentation table of the text processing module, the embedding vector mapping layer and the output decoding layer of the natural language processing model in the encrypted natural language processing system.
[0054] Among them, the embodiments of the present application do not specifically limit how to preset the switching period for the Token ID mapping relationship, which can be considered according to many factors such as the resources, time, and required security level consumed during the switching process.
[0055] Among them, the embodiments of the present application do not specifically limit how to recognize the session context switching. For example, when it is recognized that the gateway receives a new plaintext prompt word sent by the user device, that is, the session context switching is recognized.
[0056] The applicant of the present application further proposes that the above natural language processing system encryption method can be applied not only to traditional natural language processing models such as the large language model LLM, but also to models that integrate non-text modalities such as images, voices, and videos on the basis of the LLM, such as the Multimodal Large Language Model (MLLM). Correspondingly, in some embodiments, the natural language processing model includes any one of the following: the large language model LLM, the multimodal large language model MLLM.
[0057] Once traditional natural language processing systems involve parameter updates, such as updating encryption parameters in a homomorphic encryption scheme, they need to reprocess historical data in a full-scale manner. Operating on a corpus of millions of words takes more than 24 hours and cannot meet the requirements of dynamic scenarios. However, the natural language processing system encryption method provided by the embodiments of the present application can, by locally storing the character mapping relationship of each version (i.e., the character mapping relationship obtained each time step S110 is executed), not only roll back to the word segmentation table of any historical encryption period, but also simply and quickly re-encrypt the word segmentation table, providing a lightweight parameter update mechanism, supporting second-level switching and version rollback, and improving disaster tolerance and processing efficiency.
[0058] In addition, the applicant of the present application also proposes that for scenarios with a private knowledge base, since the private knowledge base needs to first give a knowledge base answer to the prompt word input by the user and provide the knowledge base answer to the natural language processing system for reference, all data to be imported into the private knowledge base can be encrypted based on the character mapping relationship by configuring an embedding model, and then the encrypted data can be imported into the vector database of the private knowledge base to ensure that the data is also in ciphertext state in the vector database. Then, the gateway can first pass the ciphertext prompt word to the private knowledge base, and the private knowledge base passes the ciphertext prompt word and the given ciphertext knowledge base answer to the encrypted natural language processing system for processing.
[0059] As the second aspect of the embodiments of the present application, an electronic device is provided, where, asFigure 8 As shown, the electronic device includes: One or more processors 101; A memory 102, on which one or more computer programs are stored. When the one or more computer programs are executed by the one or more processors 101, the one or more processors 101 implement the natural language processing system encryption method provided in the first aspect of the embodiments of the present application.
[0060] The electronic device may further include one or more I / O interfaces 103, connected between the processor 101 and the memory 102, configured to implement information interaction between the processor 101 and the memory 102.
[0061] Among them, the processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory 102 is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) is connected between the processor and the memory and can implement information interaction between the processor and the memory, including but not limited to a data bus (Bus), etc.
[0062] In some embodiments, the processor 101, the memory 102, and the I / O interface 103 are interconnected through a bus 104 and further connected to other components of the computing device.
[0063] As the third aspect of the embodiments of the present application, as Figure 9 shown, a computer-readable medium is provided, on which a computer program is stored. Wherein, when the computer program is executed by a processor, the natural language processing system encryption method provided in the first aspect of the embodiments of the present application is implemented.
[0064] Those of ordinary skill in the art will understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. Accordingly, the computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, the methods of any of the above embodiments can be implemented. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the embodiments of the present application may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0065] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Those skilled in the art should understand that the present application includes but is not limited to the content described in the drawings and the above specific implementation manner. Any modification that does not deviate from the functional and structural principles of the present application will be included in the scope of the claims.
Claims
1. A method for encrypting a natural language processing system, characterized in that, The method includes: Obtaining a character mapping relationship between an original character and an encrypted character determined based on a character encryption algorithm; Determining a Token ID mapping relationship between an original Token ID and a transformed Token ID according to a Token ID transformation algorithm; Encrypting a pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, so that the encrypted natural language processing system can output a ciphertext response based on an input ciphertext prompt.
2. The method according to claim 1, wherein After encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, the method further includes: Receiving a ciphertext prompt sent by a gateway associated with this device; wherein, the ciphertext prompt is obtained by the gateway converting a plaintext prompt sent by a user device based on the character mapping relationship; Inputting the ciphertext prompt into the encrypted natural language processing system to obtain a ciphertext response output by the encrypted natural language processing system, so that the gateway decrypts the ciphertext response to obtain a plaintext response and sends the plaintext response to the user device.
3. The method according to claim 1, wherein The natural language processing system includes a natural language processing model and its associated text processing module. Encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship includes: Performing character encryption conversion and Token ID space reconstruction on the original word segmentation table of the text processing module according to the character mapping relationship and the Token ID mapping relationship to obtain an encrypted word segmentation table; and Performing Token ID space reconstruction on the embedding vector mapping layer and the output decoding layer of the natural language processing model according to the Token ID mapping relationship.
4. The method according to claim 3, characterized in that, The original word segmentation table includes the correspondence between an original Token and an original Token ID. Performing character encryption conversion and Token ID space reconstruction on the original word segmentation table of the text processing module according to the character mapping relationship and the Token ID mapping relationship to obtain an encrypted word segmentation table includes: Replacing each original character included in each original Token in the original word segmentation table with the encrypted character corresponding to the original character in the character mapping relationship, and replacing each original Token ID in the original word segmentation table with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship to obtain an encrypted word segmentation table.
5. The method according to claim 3, characterized in that, The natural language processing model further includes an inference layer. The embedding vector mapping layer is used to determine an embedding vector sequence according to the vector mapping relationship from the original Token ID to the embedding vector and the Token ID sequence input by the text processing module. The inference layer is used to determine a hidden state sequence according to the embedding vector sequence input by the embedding vector mapping layer. The output decoding layer is used to calculate a predicted original Token ID sequence according to the output weight mapping relationship from the hidden state to each original Token ID and the hidden state sequence input by the inference layer; The space reconstruction of the Token ID for the embedding vector mapping layer and the output decoding layer of the natural language processing model according to the Token ID mapping relationship includes: Replacing each original Token ID in the vector mapping relationship with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship; And Replacing each original Token ID in the output weight mapping relationship with the transformed Token ID corresponding to the original Token ID in the Token ID mapping relationship.
6. The method according to any one of claims 1-5, characterized in that After obtaining the character mapping relationship between the original character and the encrypted character determined based on the character encryption algorithm, the character mapping relationship is also stored locally; After encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, the method further includes: Obtaining an updated character mapping relationship between the original character and the encrypted character re-determined based on the character encryption algorithm; Performing a rollback process on the current word segmentation table of the text processing module in the encrypted natural language processing system according to the character mapping relationship stored locally to obtain a rollback word segmentation table; Performing character encryption conversion on the rollback word segmentation table according to the updated character mapping relationship.
7. The method according to any one of claims 1-5, characterized in that, After encrypting the pre-trained natural language processing system according to the character mapping relationship and the Token ID mapping relationship, the method further includes: When a preset switching period arrives or a session context switch is recognized, obtaining an updated Token ID mapping relationship between the original Token ID and the transformed Token ID; Performing space reconstruction of the Token ID on the current word segmentation table of the text processing module, the embedding vector mapping layer, and the output decoding layer of the natural language processing model in the encrypted natural language processing system according to the updated Token ID mapping relationship.
8. The method according to any one of claims 1-5, characterized in that, The natural language processing model includes any one of the following: large language model LLM, multi-modal large language model MLLM.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory storing one or more computer programs thereon, which when executed by the one or more processors, cause the one or more processors to implement the natural language processing system encryption method according to any one of claims 1-8.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the natural language processing system encryption method according to any one of claims 1-8.
Citation Information
Patent Citations
Method for hiding natural language information
CN102194081A
Homomorphic encryption method of deep learning model for natural language processing
CN115987479A
Deep learning method and device for sensitive data enhancement based on NLP
CN116304674A
Table data processing method based on federal large model and related equipment
CN117633513A
Text watermark detection and watermark adding method, program product, equipment and medium
CN118656810A
Cited By
Encrypted context-based prompt system
US12627475B2
Encrypted context-based prompt system
US20250373414A1