Large model deployment model and method for protecting application originality and user privacy

By incorporating word segmentation and vector embedding modules in the big model SDK, the plaintext information input by the user is preprocessed and embedded vector calculations are solved, which solves the risk of user creativity and privacy being copied and leaked in the big model technology ecosystem, and effectively protects application creativity and user privacy.

CN120105484AInactive Publication Date: 2025-06-06CHINA TOWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510586138.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the current big model technology ecosystem, users’ application creativity and privacy are in a vulnerable state in the hands of model providers, with the risk of being copied and leaked, and users cannot effectively verify the data usage commitment of model providers.

Method used

Design a large model deployment model and method, and build the word segmentation and vector embedding module into the large model SDK. The plaintext information input by the user is preprocessed and embedded in the vector calculation on the application side of the large model, and generate intermediate calculation results and output to the large model providing side to avoid plaintext prompt words being obtained by the model provider.

Benefits of technology

By moving the word segmentation and vector embedding modules to the big model SDK, users' application creativity and privacy are effectively protected, ensuring that the prompt word plain text is not acquired by the model provider, and enhancing users' trust in model services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105484A_ABST
    Figure CN120105484A_ABST
Patent Text Reader

Abstract

The invention relates to a large model deployment model and method for protecting application originality and user privacy, word segmentation and vector embedding in large model processing are moved to a large model SDK, the part completes calculation in a large model application, and the rest part is still reserved on a large model providing side to complete calculation. In this way, the cue word of the plaintext is no longer input to the providing side of the large model, but the intermediate result of the reasoning calculation of the large model, namely the embedded vector, and the cue word plaintext cannot be recovered through the embedded vector. In this way, the prompt word plaintext is limited on the application side of the large model and cannot be obtained by a large model provider, so that application creativity and user privacy are effectively protected, the commitment of the large model provider is truly implemented and enhanced through guarantee of technical means, the user of the large model can use the large model service without worries, and the user experience is improved. And the utilization rate and the calling amount of the large model service can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model technology, and in particular to a large model deployment model and method for protecting application creativity and user privacy. Background Art

[0002] The large model technology ecosystem is showing an explosive development trend. At the basic model level, not only have technical branches such as large language models (LLM), visual language models (VLM), and multimodal large models (MMLM) been formed, but also the parallel development of open source and closed source technology routes has been spawned. In the open source camp, Meta's LLaMA series has built a technical ecosystem through a parameter scale classification strategy, and Alibaba's Qwen series continues to break through the performance in the Chinese context; in the closed source field, in addition to OpenAI's continuous iteration of the GPT series, models such as Anthropic's Claude and Google's Gemini are also continuing to catch up. While technology is evolving, innovation at the application layer is showing a blowout trend. From intelligent customer service, code assistants to creative generation tools, various vertical applications are rapidly penetrating into various industries.

[0003] The infrastructure supporting this technology ecosystem is becoming increasingly complete, including large model hosting services provided by cloud service providers such as Microsoft Azure and Amazon SageMaker, as well as the emergence of low-code development platforms such as Coze and Dify, which significantly lowers the threshold for application development. It is worth noting that there is an obvious separation of power and responsibility in the current technology ecosystem: model users (including application developers and end users who directly call the API) do not hold the model itself, but complete reasoning requests through remote API calls. In this service model, the communication of user intent is entirely dependent on the carrier of the prompt word - it must accurately express the task requirements, but also include execution strategies, and may even imply commercial secrets or personal privacy.

[0004] From the perspective of technical implementation, there are significant security risks in the life cycle of prompt words. Although encryption methods such as TLS may be used during the transmission process, when the data reaches the model server, it must be decrypted into plain text for inference calculation. This technical feature enables the model provider to be fully technically capable of acquiring, storing, and analyzing prompt words. For application developers, a carefully designed prompt word project may contain core business logic. Taking the intelligent product selection system of a cross-border e-commerce company as an example, once the supply chain optimization strategy embedded in its prompt words is copied, competitors can build similar applications within a few hours; for ordinary users, sensitive information such as conversation records in psychological counseling applications and symptom descriptions in medical consultations may be leaked through prompt words.

[0005] Although some manufacturers, such as OpenAI, have made a commitment in their service agreements that "data will not be stored or used for training", this commercial commitment has fundamental flaws: first, there is a lack of a verifiable technical guarantee mechanism, and users cannot confirm the fulfillment of the commitment through audits and other means; second, immediate data snooping is enough to cause damage without involving long-term storage or training; third, data regulatory policies vary in different jurisdictions, and it is difficult for global services to form a unified guarantee standard. What is more alarming is that some commercial entities obtain ownership of generated content through end-user agreements. This potential data rights claim further increases the risk of loss of creative results.

[0006] In the current technology ecosystem, model providers control the entire chain of basic computing power, algorithm architecture, and data processing. This centralized power structure puts users in an absolutely disadvantaged position. Even with the use of distributed technologies such as federated learning, due to the stringent requirements of large model training on computing power, substantive technical power is still concentrated in the hands of a few technology giants. This structural contradiction has given rise to a deep trust crisis - when model providers act as both "athletes" and "referees", these problems have become key bottlenecks restricting the sustainable development of the industry. Summary of the invention

[0007] The present invention provides a large model deployment model and method for protecting application creativity and user privacy. In a large model deployment model for protecting application creativity and user privacy, a large model user side and a large model provider side are included. The large model user side is used to interact with the user, and the large model provider side is used to output the large model text. The large model user side has a large model application built in, and the large model application pre-processes the plain text information input by the user, embeds the plain text information into a vector, obtains an intermediate calculation result, and outputs it to the large model provider side. The large model providing side has a built-in large model body, and the large model body is used to decode and calculate the intermediate calculation results and output the text results of the model to the large model using side.

[0008] In particular, the large model application has a built-in large model SDK, and the large model SDK includes a word segmentation module and a vector embedding module; The word segmentation module is used to convert the plain text information input by the user into a digital sequence; The vector embedding module is used to embed vectors in digital sequences to obtain intermediate calculation results.

[0009] In particular, the intermediate calculation results cannot be reversed to obtain the original digital token.

[0010] In particular, the digital token represents the index number of each word in the plain text information in the dictionary.

[0011] In particular, the large model body is provided with a multi-layer Transformer decoding module, an output layer and an inverse word segmentation module; The multi-layer Transformer decoding module includes a plurality of Transformer blocks, and the multi-layer Transformer decoding module is used to decode the intermediate calculation results; The output layer and the reverse word segmentation module are used to match the output results of the multi-layer Transformer decoding module with the index number in the dictionary to obtain the text result of the model, and output the text result of the model to the large model user side.

[0012] In particular, each of the Transformer blocks includes a multi-head self-attention layer, a feedforward neural network layer, a residual connection layer and a normalization layer.

[0013] In particular, the output layer and the inverse word segmentation module pass the output of the last Transformer block in the multiple Transformer blocks through a linear layer, map the embedding dimension to the vocabulary, and obtain the probability distribution of each digital token.

[0014] In particular, based on the obtained probability distribution, a word is selected as the next generated digital token, and the text output of the large model is obtained through inverse segmentation, and the result is output to the user side of the large model.

[0015] The present invention also provides a large model deployment method for protecting application creativity and user privacy, which is implemented based on a large model deployment model for protecting application creativity and user privacy, and for the large model user side, includes the following steps: Receive plain text information input by the user; Convert plaintext information into intermediate calculation results; Provides side output intermediate calculation results to the large model; Receive the output results of the large model providing side feedback; Display the output to the user.

[0016] The present invention also provides a large model deployment method for protecting application creativity and user privacy, which is implemented based on a large model deployment model for protecting application creativity and user privacy, and includes the following steps for the large model providing side: Receive the intermediate calculation results input from the user side of the large model; Decode the calculation results and generate output results; Output the results to the large model user.

[0017] This application has the following beneficial effects: (1) A large model deployment model is proposed, which moves the word segmentation and vector embedding in the large model processing to the large model SDK. This part is calculated in the large model application, while the rest is still calculated on the large model provider side. In this way, the input to the large model provider side is no longer the plaintext prompt word, but the intermediate result of the large model reasoning calculation, that is, the embedding vector. The plaintext of the prompt word cannot be restored through the embedding vector. In this way, the plaintext of the prompt word is limited to the large model application side and will not be obtained by the large model provider, thereby effectively protecting the application creativity and user privacy.

[0018] (2) A large model deployment method is proposed, which makes it impossible for the large model provider to see the prompt words from the user, thereby effectively protecting the creativity of the large model application and the privacy of the large model user. Through technical means, the commitment of the large model provider is truly implemented and enhanced, and the users of the large model can use the large model service without worries, which helps the large model service to increase its usage rate and call volume. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the embodiments of the present application, constitute a part of the present application, and do not constitute a limitation on the embodiments of the present invention.

[0020] Figure 1 Deploy flow charts for existing large models; Figure 2 This is a flow chart of large model deployment in an embodiment of the present invention; Figure 3 Schematic diagram of using the side method for large models; Figure 4 Provides a schematic diagram of the side-by-side method for large models. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the various embodiments can be referred to each other.

[0024] like Figure 2 As shown, in this embodiment, a big model deployment model for protecting application creativity and user privacy includes a big model user side and a big model provider side, wherein the big model user side is used to interact with the user, and the big model provider side is used to output the big model text, wherein the big model user side has a built-in big model application, and the big model application pre-processes the plain text information input by the user, embeds the vector to obtain the intermediate calculation result and outputs it to the big model provider side; The large model providing side has a built-in large model body, and the large model body is used to decode and calculate the intermediate calculation results and output the text results of the model to the large model using side.

[0025] Furthermore, the big model application has a built-in big model SDK, and the big model SDK includes a word segmentation module and a vector embedding module; The word segmentation module is used to convert the plain text information input by the user into a digital sequence; The vector embedding module is used to embed vectors in digital sequences to obtain intermediate calculation results.

[0026] The purpose of this design is to propose a large model deployment model, move the word segmentation and vector embedding in the large model processing to the large model SDK, and complete the calculation in the large model application, while the rest is still retained on the large model provider side to complete the calculation. In this way, the input to the large model provider side is no longer the plaintext prompt word, but the intermediate result of the large model reasoning calculation, that is, the embedding vector, and the prompt word plaintext cannot be restored through the embedding vector. In this way, the prompt word plaintext is limited to the large model application side and will not be obtained by the large model provider, thereby effectively protecting the application creativity and user privacy.

[0027] In this embodiment, the intermediate calculation results cannot be reversed to obtain the original digital token; At the same time, the digital token represents the index number of each word in the plaintext information in the dictionary; Furthermore, the large model is equipped with a multi-layer Transformer decoding module, an output layer, and an inverse word segmentation module; The multi-layer Transformer decoding module includes a plurality of Transformer blocks, and the multi-layer Transformer decoding module is used to decode the intermediate calculation results; The output layer and the reverse word segmentation module are used to match the output results of the multi-layer Transformer decoding module with the index number in the dictionary to obtain the text result of the model, and output the text result of the model to the large model user side.

[0028] At the same time, each of the Transformer blocks contains a multi-head self-attention layer, a feedforward neural network layer, a residual connection layer and a normalization layer.

[0029] Furthermore, the output layer and the inverse word segmentation module pass the output of the last Transformer block in the multiple Transformer blocks through a linear layer, map the embedding dimension to the vocabulary, and obtain the probability distribution of each digital token.

[0030] At the same time, based on the obtained probability distribution, a word is selected as the next generated digital token, and the text output of the large model is obtained through inverse segmentation, and the result is output to the user side of the large model.

[0031] In the specific implementation plan, the large model generally goes through: tokenizer, vector embedding, multi-layer Transformer decoder, output layer and detokenizer, and finally generates the text output of the large model, which is fed back to the user through the application.

[0032] Among them, in the existing technical solutions, such as Figure 1 As shown in the figure, the role of the prompt word is to tell the big model "what the model should do and how to do it", so the prompt word embodies the creativity of the big model application developer and may also contain the privacy of the user. In the entire process, although the network transmission may be encrypted, when the big model actually starts processing, it must be plain text (that is, the input of the tokenizer is the plain text prompt word). In this way, the big model provider can completely obtain the user's prompt word, which makes the creativity of the application developer and the privacy of the big model user completely unprotected.

[0033] Specifically, the process of the large model accepting text input and outputting text responses can be divided into the following processing steps: First, tokenizer: The input text is converted into a sequence of numbers (tokens) that the model can understand through the tokenizer, where each token usually represents the index number of the word in the dictionary (vocab). In this example, take the text "this is a input ." as an example. After tokenization, the result is shown in the following table:

[0034] Then embedding: convert the token sequence after word segmentation into an embedding vector. The embedding vector is a representation in a high-dimensional space that can capture the semantic information of each token. It can represent a token from a more abstract and complex dimension. The text "this is a input ." is embedded after embedding, and the resulting high-dimensional vector can be ; Since there are many types of high-dimensional vectors such as 1*768 or 1*1024, those skilled in the art can make a reasonable choice according to actual conditions.

[0035] Then multi-layer Transformer decoder: The embedded vector is input into the multi-layer Transformer decoder for processing. Each Transformer block includes a multi-head self-attention layer, a feedforward neural network layer, a residual connection, and a layer normalization.

[0036] Final output layer and detokenizer: The output of the last layer of Transformer blocks is passed through a linear layer to map the embedding dimension to the vocabulary size, and the probability distribution of each token is obtained. According to this probability distribution, a word can be selected as the next generated token by temperature or maximum probability. Since the token represents the index number of the word in the dictionary (vocab), the text output of the model is obtained by detokenizer.

[0037] The model input "this is a input." is a plaintext prompt word. If such a plaintext prompt word, such as "someone lives in xxx, Chenghua District, Chengdu", is abused or leaked by the model provider, the privacy will be leaked, and the same is true for application creativity.

[0038] In the technical solution provided in this embodiment, the tokenizer and vector embedding in the large model processing are moved to the large model SDK. This part is calculated in the large model application (i.e., the large model user side), while the rest (multi-layer Transformer decoder, output layer and inverse tokenization) are still retained on the large model provider side to complete the calculation. In this way, the input to the large model provider side is no longer the plaintext prompt word, but the intermediate result of the large model reasoning calculation, i.e., the embedding vector, and the prompt word plaintext cannot be restored through the embedding vector. In this way, the prompt word plaintext is limited to the large model application side (i.e., the large model user side) and will not be obtained by the large model provider, thereby effectively protecting the application creativity and user privacy.

[0039] In this embodiment, based on a large model deployment model that protects application creativity and user privacy, a large model deployment method that protects application creativity and user privacy is also provided. For the large model user side, such as Figure 3 As shown, the following steps are included: Receive plain text information input by the user; Convert plaintext information into intermediate calculation results; Provides side output intermediate calculation results to the large model; Receive the output results of the large model providing side feedback; Display the output to the user.

[0040] At the same time, in this embodiment, based on a large model deployment model that protects application creativity and user privacy, a large model deployment method that protects application creativity and user privacy is also provided, for the large model providing side, such as Figure 4 As shown, the following steps are included: Receive the intermediate calculation results input from the user side of the large model; Decode the calculation results and generate output results; Output the results to the large model user.

[0041] In the specific implementation, the large-model deployment method for protecting application creativity and user privacy is explained from two dimensions. Compared with the existing method, the change in the above processing method does not affect the model training, the composition and weight of the model, nor does it affect the calculation results of the model. It only changes the deployment method of the model, which is simple and easy to operate.

[0042] At the same time, the multi-layer Transformer decoder is the main part of the model. It is layered and can be split, for example, into the first N layers and the last M layers. The first N layers are placed in the large model SDK, and the last M layers are placed on the large model provider side. This is completely feasible. However, it is not necessary to do so, because as long as the output of the large model SDK is no longer a plaintext prompt word, and the prompt word plaintext cannot be restored from the model intermediate result, vector embedding has already guaranteed this, and the main operation should still be kept on the model provider side.

[0043] Furthermore, if the input layer and the detokenizer and even the last several layers of the multi-layer Transformer decoder are also moved into the SDK part of the large model as post-processing of the model output, this is feasible in itself. At this time, the output of the model also becomes an intermediate result. This situation is not described in detail in this application because this processing is not very meaningful. On the one hand, the output results of the large model are usually transmitted encrypted between the application and the model, which is sufficient to prevent third parties from obtaining the results of the large model. On the other hand, the model provider always holds the entire model parameters. If the model provider wants to get the final output of the model, this is completely unpreventable.

[0044] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A big model deployment model for protecting application creativity and user privacy, comprising a big model user side and a big model provider side, wherein the big model user side is used to interact with the user, and the big model provider side is used to output a big model text, characterized in that: The large model user side has a built-in large model application, and the large model application pre-processes the plain text information input by the user, embeds the vector to obtain the intermediate calculation result and outputs it to the large model provider side; The large model providing side has a built-in large model body, and the large model body is used to decode and calculate the intermediate calculation results and output the text results of the model to the large model using side.

2. A large model deployment model for protecting application creativity and user privacy according to claim 1, characterized in that: The large model application has a built-in large model SDK, and the large model SDK includes a word segmentation module and a vector embedding module; The word segmentation module is used to convert the plain text information input by the user into a digital sequence; The vector embedding module is used to embed vectors in digital sequences to obtain intermediate calculation results.

3. A large model deployment model for protecting application creativity and user privacy according to claim 2, characterized in that: The intermediate calculation result cannot be reversed to obtain the original digital token.

4. A large model deployment model for protecting application creativity and user privacy according to claim 3, characterized in that: The digital token represents the index number of each word in the plain text information in the dictionary.

5. A large model deployment model for protecting application creativity and user privacy according to claim 1, characterized in that: The large model body is provided with a multi-layer Transformer decoding module, an output layer and an inverse word segmentation module; The multi-layer Transformer decoding module includes a plurality of Transformer blocks, and the multi-layer Transformer decoding module is used to decode the intermediate calculation results; The output layer and the reverse word segmentation module are used to match the output results of the multi-layer Transformer decoding module with the index number in the dictionary to obtain the text result of the model, and output the text result of the model to the large model user side.

6. A large model deployment model for protecting application creativity and user privacy according to claim 5, characterized in that: Each of the Transformer blocks includes a multi-head self-attention layer, a feedforward neural network layer, a residual connection layer, and a normalization layer.

7. A large model deployment model for protecting application creativity and user privacy according to claim 5, characterized in that: The output layer and the inverse word segmentation module pass the output of the last Transformer block in the multiple Transformer blocks through a linear layer, map the embedding dimension to the vocabulary, and obtain the probability distribution of each digital token.

8. A large model deployment model for protecting application creativity and user privacy according to claim 7, characterized in that: Based on the obtained probability distribution, a word is selected as the next generated digital token, and the text output of the large model is obtained through inverse segmentation, and the result is output to the user side of the large model.

9. A large model deployment method for protecting application creativity and user privacy, based on a large model deployment model for protecting application creativity and user privacy as claimed in any one of claims 1 to 8, characterized in that: For large model users, the following steps are included: Receive plain text information input by the user; Convert plaintext information into intermediate calculation results; Provides side output intermediate calculation results to the large model; Receive the output results of the large model providing side feedback; Display the output to the user.

10. A large model deployment method for protecting application creativity and user privacy, based on a large model deployment model for protecting application creativity and user privacy as claimed in any one of claims 1 to 8, characterized in that: For the large model providing side, the following steps are included: Receive the intermediate calculation results input from the user side of the large model; Decode the calculation results and generate output results; Output the results to the large model user.

Citation Information

Patent Citations

  • Information processing method, apparatus and system, intelligent terminal and server

    CN108024005A

  • Prediction method and device for following data, electronic equipment and storage medium

    CN119903876A