Method and system for safe question and answer and computer readable storage medium

By extracting common instructions locally and using a large language model to generate prompt information, and combining it with a local small language model to generate answers, the quality of question and answering is improved while protecting privacy, solving the problems of privacy leakage and poor generation quality when large and small language models work together.

CN120611019APending Publication Date: 2025-09-09SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Patent Information

Application Number
CN202510694661.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies, when using large language models for secure question answering, face the risks of privacy information leakage and inaccurate generation results. In particular, the generation quality of small language models is still significantly lower than that of large language models.

Method used

By receiving user input locally, extracting common instructions and providing them to the remote large language model to generate prompt information, and using the local small language model to generate answers based on user input and prompt information, the large language model and the small language model can work together.

Benefits of technology

While protecting user privacy, it improves the quality and personalization of answers and solves the problem of insufficient generation capabilities of small language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611019A_ABST
    Figure CN120611019A_ABST
Patent Text Reader

Abstract

The invention relates to a method and system for secure questioning and answering and a computer readable storage medium. A method for secure questioning and answering includes: locally receiving a user input; extracting a general instruction from the user input, wherein the general instruction does not comprise privacy information of the user; providing the general instruction to a large language model to generate prompt information based on the general instruction, wherein the large language model is deployed remotely; and invoking a small language model to generate an answer to the user input based on the user input and the cue information, the small language model being deployed locally.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to computer systems utilizing computational models, and more particularly to methods, systems, and computer-readable storage media for secure question answering. Background Art

[0002] With the advancement of deep learning and natural language processing technologies, large language models (LLMs) have achieved breakthroughs in text generation and command understanding, and are widely used in scenarios such as writing assistance, code generation, and knowledge question answering. Large language models are typically large in scale and have high inference costs, and are often provided as cloud-based application programming interfaces (APIs).

[0003] In many application scenarios, the input provided by users includes instructions and personal information (such as user profiles, activity logs, professional background, etc.), which contain highly sensitive private information. Since large cloud-based language models are usually hosted by a third party, if these inputs containing private information are directly uploaded to the large cloud-based language model, it may lead to the leakage, abuse, or acquisition of private information by malicious attacks. In addition, these inputs containing private information uploaded to the large cloud-based language model may be recorded, stored, or subjected to secondary analysis by a third party. Once a data leak occurs, this private information is difficult to trace and delete. In addition to the high security risk of private information, the reasoning of large language models usually consumes more computing resources and needs to rely on external services, resulting in a significant increase in computing costs and delays.

[0004] Some solutions propose feeding only instructions that do not contain private information to a large language model to generate results. However, because large language models cannot fully capture contextual information, the results generated by large language models are very broad and lack personalization, and the accuracy and pertinence of the answers are reduced.

[0005] To ensure data security, some solutions propose using a small language model (SLM) deployed locally on the device to perform tasks. Since local small language models can directly access private information on the local device, eliminating the need to upload private information to the cloud, they can effectively protect user privacy. However, the reasoning capabilities and generation quality of small language models still lag significantly behind those of large language models, resulting in poor generation quality.

[0006] There is a need in the art for improved security question-answering technology in at least one of the aforementioned aspects. Summary of the Invention

[0007] It is to be understood that both the foregoing general description and the following detailed description of the present disclosure are exemplary and explanatory and are intended to provide further explanation of the disclosure as claimed.

[0008] One aspect of the present disclosure provides a method for secure question-answering, comprising: receiving user input locally; extracting general instructions from the user input, the general instructions not including the user's private information; providing the general instructions to a large language model to generate prompt information based on the general instructions, the large language model being deployed remotely; and calling a small language model to generate an answer to the user input based on the user input and the prompt information, the small language model being deployed locally.

[0009] As described above, the method of extracting general instructions from the user input includes: using a classification model to determine whether each sentence in the user input contains private information; and extracting the general instructions from the user input based on the determination result for each sentence in the user input.

[0010] In the above method, extracting general instructions from the user input includes: calling the small language model to perform event task extraction on the user input to generate the general instructions.

[0011] According to the method described above, extracting the general instruction from the user input includes: performing rule judgment based on the user input and a keyword table to identify private information in the user input that matches the keyword table; and extracting the general instruction from the user input.

[0012] According to the method described above, providing the general instruction to the large language model to generate prompt information based on the general instruction includes: providing the general instruction to the large language model to generate an answer outline based on the general instruction, and the answer outline is used as the prompt information.

[0013] As described above, the method of calling the small language model to generate an answer to the user input based on the user input and the prompt information includes: calling the small language model to optimize the answer outline based on the user input; and calling the small language model to generate an answer to the user input based on the user input and the optimized answer outline.

[0014] According to the method described above, providing the general instruction to the large language model to generate prompt information based on the general instruction includes: providing the general instruction to the large language model to generate a first probability distribution for the next output word position based on the general instruction, and the first probability distribution is used as the prompt information; calling the small language model to generate an answer to the user input based on the user input and the prompt information includes: calling the small language model to generate a second probability distribution for the next output word position based on the user input; fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution for the next output word position; and determining the output word for the next output word position based on the fused probability distribution.

[0015] As described above, the method of fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution for the next output word position includes: using a fusion model to obtain the fused probability distribution for the next output word position based on the first probability distribution and the second probability distribution.

[0016] According to the method described above, fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution for the next output word position includes: fusing the probabilities of candidate word units corresponding to a predetermined number of highest probabilities in the first probability distribution and the second probability distribution.

[0017] Another aspect of the present disclosure provides a system for secure question answering, comprising: a small language model; and a processor, wherein the processor is configured to execute any of the above methods.

[0018] Another aspect of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of any of the above methods when executed by a processor.

[0019] The method and system according to the present disclosure improve the quality of answers to user input while ensuring the security of user privacy information. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Various embodiments of the present disclosure are described with reference to the accompanying drawings.

[0021] Figure 1 is a block diagram of a system for security question answering according to some embodiments of the present disclosure.

[0022] Figure 2 is a schematic diagram of a first process for security question and answer according to some embodiments of the present disclosure.

[0023] Figure 3is a schematic diagram of user input according to some embodiments of the present disclosure.

[0024] Figure 4 is a schematic diagram of a second process associated with a first process according to some embodiments of the present disclosure.

[0025] Figure 5 is a schematic diagram of generating answers by a processor using an outline-based approach according to some embodiments of the present disclosure.

[0026] Figure 6 is a schematic diagram of a third process associated with the first process according to some other embodiments of the present disclosure.

[0027] Figure 7 is a schematic diagram of generating an answer using a logit-based method by a processor according to some embodiments of the present disclosure.

[0028] Figure 8 is a schematic diagram of probability distribution according to some embodiments of the present disclosure.

[0029] Figure 9 is a schematic diagram of probability distribution according to some other embodiments of the present disclosure.

[0030] Figure 10 is a schematic diagram of the weights of output word units according to some embodiments of the present disclosure.

[0031] Figure 11 is a schematic diagram of a process for constructing a synthetic dataset according to some embodiments of the present disclosure.

[0032] Figure 12 is a flowchart of a method for security question answering according to some embodiments of the present disclosure.

[0033] Figure 13 is a flowchart of a method for security question and answer according to other embodiments of the present disclosure.

[0034] Figure 14 is a flowchart of a method for security question and answer according to some further embodiments of the present disclosure.

[0035] Figure 15 is a block diagram of a computer-readable storage medium according to some embodiments of the present disclosure.

[0036] Figure 16 is a block diagram of a computer program product according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0037] Embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings, but the present disclosure is not limited thereto but only by the claims. In the drawings, for illustrative purposes, the dimensions of some of the elements may be exaggerated and not drawn to scale. Wherever possible, the same reference numerals will be used throughout the drawings to represent the same or similar parts.

[0038] Although the terms used in this disclosure are selected from commonly known and commonly used terms, some of the terms mentioned in this disclosure may be selected by the applicant at his or her discretion, and their detailed meanings are explained in the relevant parts of the description herein. In addition, it is required to understand this disclosure not only by the actual terms used, but also by the meanings implied by each term.

[0039] In the description provided herein, numerous specific details are set forth. However, it should be understood that embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques are not shown in detail to avoid obscuring an understanding of the present disclosure.

[0040] In this application, ordinal numbers such as "first," "second," and "third" are used to distinguish different instances of the same object. The ordinal numbers such as "first," "second," and "third" do not indicate the relative order of the objects in time, space, ranking, or other aspects.

[0041] According to one aspect of the present disclosure, a system for security question and answer is provided.

[0042] Figure 1 is a block diagram of a system 100 for security question answering according to some embodiments of the present disclosure.

[0043] The system 100 may be a local electronic device, such as a personal computer, a tablet computer, a mobile phone, etc. The system 100 may include a processor 102 and a small language model 104 .

[0044] The processor 102 may include a central processing unit (CPU), a graphics processing unit (GPU), and various other processing units or cores (e.g., an arithmetic logic unit, an integer unit, a floating point unit, a tensor unit, a ray tracing core, etc.).

[0045] The small language model 104 may be configured to call computing resources to perform corresponding operations. The small language model 104 may be a smaller model customized for a specific task.

[0046] The system 100 may be coupled to a remote computing device 110 via wired and / or wireless communication. The remote computing device 110 may be a remote computer cluster, a cloud server, a cloud infrastructure, etc. The remote computing device 110 may include a large language model 112 .

[0047] The large language model 112 may be configured to call computing resources to perform corresponding operations. The large language model 112 may be a relatively large model designed for general applications and advanced performance.

[0048] In order to improve the quality of answers while protecting user privacy, the present disclosure proposes that the processor 102 call the large language model 112 deployed remotely to generate prompt information, and call the small language model 104 deployed locally to generate answers based on user input and prompt information.

[0049] Figure 2 is a schematic diagram of a first process 200 for security question answering according to some embodiments of the present disclosure. The first process 200 may be performed between a processor 202, a large language model 204, and a small language model 206. In some embodiments, the processor 202 may be Figure 1 In the processor 102, the large language model 204 can be Figure 1 The large language model 112 in the example, the small language model 206 can be Figure 1 The small language model 104 in , but the scope of the present invention is not limited in this regard.

[0050] First, the processor 202 may receive user input locally.

[0051] Then, the processor 202 may extract the general instruction from the user input, where the general instruction does not include the user's private information.

[0052] Next, the processor 202 may provide the general instructions to the large language model 204. The large language model 204 is deployed remotely. The large language model 204 may generate prompt information based on the general instructions and return the prompt information to the processor 202.

[0053] Finally, the processor 202 may call the small language model 206. The small language model 206 is deployed locally and may generate an answer to the user input based on the user input and the prompt information.

[0054] Therefore, some embodiments of the present disclosure only provide general instructions to the large language model to generate prompt information, and enable the small language model to generate answers based on the complete user input and prompt information. By leveraging the collaborative work of the locally deployed small language model and the remotely deployed large language model, high-quality and personalized answers can be generated without leaking private data to the remote device. This improves the quality of answers generated by the small language model and effectively addresses the problem of insufficient generation capabilities of the small language model.

[0055] In some embodiments, the processor may receive user input through an input device. User input may include general instructions and personal information. General instructions may be high-level requirements, such as "write a blog about environmental assessment." Personal information may include personal instructions. Personal instructions may be detailed requirements, such as "use the work log from last week's meeting and past blogs as a reference." Personal information may also include personal device information, such as the profile of the user using the system, user notes, user schedule, user activity log on the system, articles published by the user through the system, emails sent by the user through the system, and so on. Personal device information may be provided to the processor directly by the user through the input device, or may be obtained by the processor extracting information stored in the system after authorization by the user. It should be understood that obtaining the user's personal device information from the system is achieved after the user's authorization and in compliance with privacy regulations.

[0056] Figure 3 FIG3 is a diagram of a user input 300 according to some embodiments of the present disclosure. The user input 300 may include a general instruction 302 , a personal instruction 304 , and personal device information 306 . Figure 3 Specific examples of components of user input 300 are given.

[0057] In some embodiments, the processor can directly extract the independent general instruction from the user input. When the user enters the general instruction and personal information separately, that is, the general instruction that does not contain private information and the personal information that contains private information are independent of each other, the processor can directly extract the independent general instruction from the user input.

[0058] In some embodiments, the processor may pre-process the user input to extract the general instruction from the user input. When the user input is not presented as separate general instructions and personal information, but rather the entire user input is presented as a whole, the processor may extract the general instruction from the user input for user convenience.

[0059] In some embodiments, the processor may use a classification model to determine whether each sentence in the user input contains private information. Then, based on the determination result for each sentence in the user input, a general instruction may be extracted from the user input.

[0060] In some embodiments, the classification model can be an independent model. For example, the classification model can be implemented using an embedding-based feature classification model. The classification model can characterize the sentence input and output a classification of whether it contains private information. For example, an output of 0 indicates no private information, and an output of 1 indicates private information. The classification model can classify each sentence in the user input. The processor can extract all sentences with an output of 0 as general instructions.

[0061] In some embodiments, the classification model can be implemented by a small language model.

[0062] In some embodiments, the processor can invoke a small language model to extract event tasks from user input to generate general instructions. For example, the processor can detect events by identifying event trigger words in user input and determining the event type. The processor can then use the event representation framework to determine whether the entity in the user input is an event element and determine the element role. In some embodiments, the processor can use an independent event task extraction model to extract event tasks from user input to generate general instructions.

[0063] In some embodiments, the processor can perform rule-based analysis based on user input and a keyword table to identify private information in the user input that matches the keyword table. General instructions can then be extracted from the user input. The keyword table can include information such as mobile phone numbers, email addresses, and ID numbers. If the user input matches the keyword table, it contains private information.

[0064] Therefore, some embodiments of the present disclosure analyze user input and extract only general instructions that do not contain privacy information from the user input, thereby ensuring the security of private data.

[0065] In some embodiments, the processor may also extract summary information from user input. This summary information can be a highly abstracted summary of private information or a brief overview after removing sensitive information. For example, sensitive information can be removed from the historical activity records of personal device information in user input. For example, private information such as the topic of a specific meeting the user attended or the title of a specific book the user purchased can be highly abstracted to only reflect the general categories of the meeting or book, without including detailed content. This summary information extraction can be achieved using a small language model.

[0066] Therefore, some embodiments of the present disclosure can provide more professional reference content by extracting summary information from user input without leaking privacy information.

[0067] Outline-based approach

[0068] In some embodiments, the processor may utilize an outline-based approach to achieve collaboration between the large language model and the small language model. Figure 4 is a schematic diagram of a second process 400 associated with the first process 200 according to some embodiments of the present disclosure.

[0069] The second process 400 may be performed between the processor 202, the large language model 204, and the small language model 206. In some embodiments, the processor 202 may be Figure 1 In the processor 102, the large language model 204 can be Figure 1 The large language model 112 in the example, the small language model 206 can be Figure 1 The small language model 104 in , but the scope of the present invention is not limited in this regard.

[0070] In some embodiments, after extracting the general instructions from the user input, the processor 202 may provide the general instructions to the large language model 204. The large language model 204 may generate an answer outline based on the general instructions, which serves as prompt information. The answer outline may be a writing framework, knowledge planning, etc.

[0071] Since the large language model has stronger reasoning ability, it can generate a higher-quality answer outline than the small language model to guide the small language model to generate a specific answer based on the answer outline.

[0072] In some embodiments, the processor 202 may call the small language model 206 , and the small language model 206 may generate an answer to the user input based on the user input and an answer outline as prompt information.

[0073] By using the small language model to deeply fill in the answer outline generated by the large language model with user input containing private data, it is possible to complete the generation of personalized content while ensuring the security of private data. Ultimately, a complete response with both general knowledge depth and personalized privacy content is generated, achieving the complementary effect of the high-level reasoning ability of the large language model and the security of the small language model.

[0074] The following is an example of prompt words used by a large language model to generate an answer outline according to some embodiments of the present disclosure:

[0075] ===========;

[0076] You’re an organizer responsible for only giving the skeleton(not thefull content)for answering the question.Provide the skeleton in a list ofpoints(numbered 1.,2.,3.,etc.)to answer the question.Instead of writing afull sentence,each skeleton point should be very short with only 3-5words.Generally,the skeleton should have 8-15points.You can refer to thefollowing examples:

[0077] [Task1]:Develop a Marketing Script for Your Monthly Dinner Party:Create ascript that highlights your monthly dinner party as a networkingplatform.

[0078] [Skeleton1]:1.Warmly lit dining room\n2.Fine china and gourmetdishes\n3.Soft music background\n4.Invitation opening\n5.Guests arriving andnetworking\n6.Host’s welcoming toast\n7.Expertly paired courses and wine\n8.Animated guest discussions\n9.Guest speaker’s address\n10.Post-dinnernetworking lounge\n11.Online community continuation\n12.Next event datehighlighted\n13.Closing with logo and contact info

[0079] [Task2]:Compose a reflective essay on the evolution of bridge design:Thomas,with his patent in bridge design,can discuss the evolution of bridgeengineering,modern challenges,and future perspectives.

[0080] [Skeleton2]:1.Introduction to bridges\n2.Early bridges:materials,principles\n3.Roman arches,concrete use\n4.Industrial Revolution:iron,steel\n5.Brooklyn Bridge:design icon\n6.20th-century advances:materials,techniques\n7.Modern challenges:sustainability,climate\n8.Future technologies:smartmaterials,sensors\n9.Ethical considerations,safety\n10.Conclusion:adaptation,advancement

[0081] Now,please provide the skeleton for the following question.

[0082] {question}

[0083] ==========;

[0084] Figure 5is a schematic diagram of generating an answer by a processor using an outline-based method according to some embodiments of the present disclosure. In some embodiments, after extracting the general instruction 504 from the user input 502, the processor may provide the general instruction 504 to a large language model 204 remotely (e.g., in the cloud). The large language model 204 may generate an answer outline 506 based on the general instruction 504, and the answer outline 506 provides key points with depth and logic for the general instruction 504. The processor may then call the local small language model 206, which may generate an answer 508 to the user input based on the user input 502 and the answer outline 506. For example, the small language model 206 may fill in the answer outline 506 with personal information such as user profiles and user historical activity information, so as to generate a high-quality and personalized answer 508 to the user input while protecting privacy information. Figure 5 General instructions, answer outlines, and specific examples of responses to user input are given.

[0085] In some embodiments, processor 202 may provide general instructions and summary information to large language model 204. Based on the general instructions and summary information, the large language model may generate an answer outline, which serves as prompt information. With the aid of the summary information, the answer outline generated by the large language model may be more targeted and professional.

[0086] In some embodiments, the processor may call a small language model, which may optimize the answer outline based on the user input. The small language model may then generate an answer to the user input based on the user input and the optimized answer outline.

[0087] For example, the large language model generates an answer outline that describes the chapters of an article and the transitions between them. The small language model can use more detailed personal information, such as meeting times and even email history, to optimize the answer outline, making it more personalized. Therefore, the collaborative processing of answer outlines by the large and small language models achieves a balance between performance and privacy.

[0088] In some embodiments, the processor may call a small language model, which may generate an answer outline based on user input. The processor may then provide the answer outline to the large language model, which may optimize the answer outline. Finally, the processor may call the small language model, which may generate an answer to the user input based on the user input and the optimized answer outline. Since the answer outline generally does not contain private information, providing the answer outline generated by the small language model to the large language model will not result in the leakage of private information. At the same time, the answer outline generated by the small language model is more personalized. By optimizing the personalized answer outline through the large language model, the quality of the answer outline can be further improved.

[0089] In some embodiments, the processor may also encrypt or store general instructions in fragments before providing them to the large language model to further improve the security of user information.

[0090] Therefore, in some embodiments of the present disclosure, a remote large language model provides a macro structure based on general information, and a local small language model flexibly uses the private information in the user input for integration, thereby achieving the effect of protecting private information and improving the quality of the answer content.

[0091] Logit-based methods

[0092] In some embodiments, the processor may utilize a logit (raw prediction score)-based approach to implement collaboration between the large language model and the small language model. Figure 6 is a schematic diagram of a third process 600 associated with the first process 200 according to some other embodiments of the present disclosure.

[0093] The third process 600 may be performed between the processor 202, the large language model 204, and the small language model 206. In some embodiments, the processor 202 may be Figure 1 In the processor 102, the large language model 204 can be Figure 1 The large language model 112 in the example, the small language model 206 can be Figure 1 The small language model 104 in , but the scope of the present invention is not limited in this regard.

[0094] Logit is the unnormalized raw prediction score output by the model, which reflects the model's confidence in each candidate token at the next token position. For example, for the next token position, the model can calculate the logit value for each candidate token in the vocabulary, and then normalize the distribution of the logit values ​​of all candidate tokens to obtain the probability distribution of the token position.

[0095] In some embodiments, after extracting the general instruction from the user input, the processor 202 may provide the general instruction to the large language model 204. The large language model 204 may generate a first probability distribution for the next output word unit position based on the general instruction, with the first probability distribution serving as prompt information. The processor 202 may then call the small language model 206 to generate a second probability distribution for the next output word unit position based on the user input. The processor 202 may then fuse the first and second probability distributions to obtain a fused probability distribution for the next output word unit position. Finally, the processor 202 may determine the output word unit for the next output word unit position based on the fused probability distribution.

[0096] Figure 7 706 is a schematic diagram of a processor generating an answer using a logit-based method according to some embodiments of the present disclosure. In some embodiments, after extracting the general instruction 704 from the user input 702, the processor may provide the general instruction 704 to a large language model 204 remotely (e.g., in the cloud). The large language model 204 may generate a first probability distribution 706 for the next output word unit position based on the general instruction 704. The processor 202 may then call the local small language model 206 to generate a second probability distribution 708 for the next output word unit position based on the user input 702. The processor may then fuse the first probability distribution 706 and the second probability distribution 708 to obtain a fused probability distribution 710 for the next output word unit position. Finally, the processor may determine an output word unit 712 for the next output word unit position based on the fused probability distribution 710.

[0097] The following is an example of a prompt word generated by a large language model according to some embodiments of the present disclosure for a first probability distribution:

[0098] ===========;

[0099] You are now a helpful personal AI assistant.

[0100] #Task:...

[0101] ===========;

[0102] The following is an example of a prompt word generated by a small language model according to some embodiments of the present disclosure for a second probability distribution:

[0103] ===========;

[0104] You are now a helpful personal AI assistant.

[0105] #User Profile:...

[0106] #Past Writings:...

[0107] #Task:...

[0108] ===========;

[0109] The above prompt words reflect the difference between the prompt words of the large language model and the small language model when generating probability distributions. That is, the large language model generates probability distributions based on general instructions, while the small language model generates probability distributions based on complete user input including general instructions and personal information.

[0110] Figure 8 Schematic diagram of probability distribution according to some embodiments of the present disclosure. For the next output word position, the large language model predicts a first probability distribution, the small language model predicts a second probability distribution, and the first probability distribution and the second probability distribution are fused to obtain a fused probability distribution. Figure 8 It can be seen that the second candidate word (t2) in the first probability distribution has the highest probability, and the fourth candidate word (t4) in the second probability distribution has the highest probability. After fusion, the first candidate word (t1) in the fused probability distribution has the highest probability, that is, the fused probability distribution has changed relative to the prediction results of the large language model and the small language model, and can produce a word that is more in line with the current context. Although Figure 8 The probability distribution in is only based on 5 candidate word units as an example, but the number of candidate word units is not limited to this.

[0111] The probability distribution includes an indication of the output token at the next output token position. Since the output results generated by the large language model based on general instructions are used to control the overall generated content, while the output results generated by the small language model based on user input containing personal information are used to reflect personalized information, the fusion of the first probability distribution generated by the large language model and the second probability distribution generated by the small language model can provide tokens that are more appropriate to the current context.

[0112] In some embodiments, the processor may fuse the probabilities of candidate word-grams corresponding to a predetermined number of highest probabilities in the first probability distribution and the second probability distribution.

[0113] Figure 9It is a schematic diagram of the probability distribution according to some other embodiments of the present disclosure. A predetermined number of highest probabilities are determined in the first probability distribution and the second probability distribution, for example, 3 highest probabilities, namely the first word (t1), the second word (t2), the fourth word (t4) in the first probability distribution and the first word (t1), the second word (t2), the fourth word (t4) in the second probability distribution, as indicated by the shadows in the figure. Then, only t1, t2, t4 in the first probability distribution and t1, t2, t4 in the second probability distribution are fused to obtain a fused probability distribution containing only the probabilities of t1, t2, t4. Although Figure 9 The probability distribution in is only taken as an example with 5 candidate word units, and the predetermined number of the highest probabilities is 3, but the number of candidate word units and the predetermined number of the highest probabilities are not limited thereto.

[0114] By fusing only the probabilities of a predetermined number of candidate word units with the highest probabilities in the first probability distribution and the second probability distribution, the amount of calculation can be reduced, and the interference of candidate word units with lower probabilities on the output results can be reduced.

[0115] In some embodiments, the processor may weight the first probability distribution and the second probability distribution to obtain a fused probability distribution. The weights may be pre-set as needed. For example, the weights of the first probability distribution and the second probability distribution may both be 0.5, i.e., the first probability distribution and the second probability distribution each have half the weight. For example, the weight of the first probability distribution may be 0.8, and the weight of the second probability distribution may be 0.2, which emphasizes the output of the large language model. For example, the weight of the first probability distribution may be 0.3, and the weight of the second probability distribution may be 0.7, which emphasizes the output of the small language model.

[0116] In some embodiments, the processor may use a fusion model to obtain a fused probability distribution for the next output word position based on the first probability distribution and the second probability distribution. The fusion model can learn the fused probability distribution from the first probability distribution and the second probability distribution to achieve flexible adjustment of the weights of different output word positions.

[0117] In some embodiments, the fusion model can be a three-layer neural network with ReLU activation, and sigmoid activation is used to determine the final weights. During the training process of the fusion model, in order to generate training data, the synthesized general instructions and personal information can be provided to the large language model. In this case, the large language model has the highest quality of generated content, so the output of the large language model can be used as a supervisory signal. Therefore, this training data can be used to supervise the collaborative generation of the large language model using general instructions and the small language model using general instructions and personal information.

[0118] Figure 10Schematic diagram of the weight of output word units according to some embodiments of the present disclosure. Red represents a higher weight for the large language model, blue represents a higher weight for the small language model, the darker the color, the higher the weight, and white represents the same weight for both. Figure 10 It can be seen that different weights are used for different output words. The large language model mainly affects the outline of the generated content, while the small language model affects the main part of the generated content, reflecting the synergy between the large and small language models.

[0119] Therefore, in some embodiments of the present disclosure, by using a fusion model to weight the probability distributions from the large language model and the small language model, the contributions of personal information and general knowledge can be balanced.

[0120] In some embodiments, before fusing the first probability distribution with the second probability distribution, noise may be added to the second probability distribution generated by the small language model to improve protection of privacy information when sharing the probability distribution.

[0121] In some embodiments, when the processor determines the output word-gram for the next output word-gram position based on the fused probability distribution, greedy sampling can be performed on the fused probability distribution to determine the output word-gram. For example, the candidate word-gram with the highest probability from the fused probability distribution can be selected as the output word-gram.

[0122] In some embodiments, when the processor determines an output word-meta for the next output word-meta position based on the fused probability distribution, the processor may perform probability distribution sampling on the fused probability distribution to determine the output word-meta. For example, the processor may use methods such as temperature sampling to determine the output word-meta, allowing low-probability word-meta to be selected with a certain probability, thereby generating more creative content.

[0123] In some embodiments, when generating a probability distribution, the processor may share a response sequence between the large language model and the small language model as model input when generating the next output word. For example, after the processor determines the output word for the next output word position, the processor may provide the response sequence containing the output word to the large language model and the small language model as a common generation basis. The large language model then continues to determine the probability distribution for the next output word position based on the general instruction and the response sequence, while the small language model continues to determine the probability distribution for the next output word position based on the user input and the response sequence. The two models are then merged again and the output word for the next output word position is determined, until the last word is determined.

[0124] In some embodiments, when generating a probability distribution, the processor may share a predetermined number of the earliest output tokens in the response sequence between the large language model and the small language model as model inputs when generating the next output token. For example, after the processor determines the output token for the next output token position, the processor may provide only the two earliest generated output tokens in the response sequence as shared tokens to the large language model and the small language model as a common generation basis. The large language model then continues to determine the probability distribution for the next output token position based on the general instructions and the shared tokens, and the small language model continues to determine the probability distribution for the next output token position based on the user input and the shared tokens. The output tokens for the next output token position are then fused and determined again until the last token is determined. By only sharing some output tokens between the large language model and the small language model, the information provided to the large language model can be adjusted according to actual needs, reducing the amount of data transmission and protecting the results containing privacy information generated by the small language model. Although the above example uses the sharing of two output tokens, the number of shared tokens is not limited to this.

[0125] In some embodiments, the large language model and the small language model can share the same tokenizer. A tokenizer is a component that breaks model input into individual tokens. When the large language model and the small language model share the same tokenizer, it can improve the consistency of the two models' understanding of the relevant input.

[0126] In some embodiments, the word segmenters of the large language model and the small language model can be aligned. By aligning the word units and probabilities of the different word segmenters of the large language model and the small language model, a variety of large and small language models can be used, thereby improving the system's broad applicability.

[0127] Therefore, in some embodiments, using a logit-based method, the general knowledge of a large language model and the personal information understanding of a small language model can be aggregated at a finer-grained word level, thereby obtaining a more accurate and personalized generation result.

[0128] In some embodiments, the processor can choose between an outline-based method and a logit-based method. For example, the choice can be made based on the computing power and / or processing requirements of the system. For example, when the computing power is sufficient or the quality requirements for the generated results are higher, the logit-based method can be selected. For example, the choice can be made based on security requirements. For example, when the security protection requirements for personal information are high, the outline-based method can be selected to avoid providing personal information to the large language model. Therefore, by choosing between the two methods according to different needs, a flexible response to different needs can be achieved.

[0129] In some embodiments, the processor may limit the transmission of private information to the large language model. For example, when providing general instructions to the large language model, the processor may audit the general instructions to ensure that the general instructions do not contain private information.

[0130] In some embodiments, the answer outline or the first probability distribution generated by the large language model may be encrypted before being transmitted to the processor. For example, the answer outline or the first probability distribution may be encrypted using a private key and decrypted using a public key after transmission.

[0131] In some embodiments, the processor may also display a response to the user input.

[0132] Therefore, according to some embodiments of the present disclosure, the system never provides private data to a remote large language model during the process of generating an answer to a user input, thereby protecting user privacy and information security.

[0133] This disclosure also proposes generating synthetic datasets to train small language models without leaking true user privacy. Figure 11 is a schematic diagram of the process of constructing a synthetic dataset according to some embodiments of the present disclosure. First, at step S1102, a user group portrait is created. The user group portrait may include information such as demographic data, professional background, interests, etc. to determine a specific AI usage scenario. Then, at step S1104, a user profile is constructed. For example, an individual user profile can be elaborated in detail to build a character for an AI writing task by combining a unique writing style, fictional personal details, and the usage of smart devices. Next, at step S1106, task instructions are written. For example, tasks can be created that are consistent with the user character's occupation, hobbies, and lifestyle. Finally, at step S1108, personalized content is generated to reflect the character's professional and personal narrative in an accurate style. Figure 11 Concrete examples on synthetic data are given.

[0134] Table 1 is the experimental results of the large language model and the small language model obtained based on the evaluator according to some embodiments of the present disclosure. The evaluator can be, for example, the large language model of GPT-4. The experimental results include scores of the large language model with user input, the large language model with general instructions, the small language model with user input, the outline-based method, and the logit-based method. The test dataset includes a synthetic dataset, a public Avocado email dataset, and an academic paper dataset. Ovl represents the total score based on personalization, consistency with user profile, usefulness, relevance, depth, creativity, and level of detail, where Ovl.(w) represents the score of the generated result by the evaluator with personal information input by the user, Ovl.(w / o) represents the score of the generated result by the evaluator without personal information input by the user, and Per represents the personalized score of the generated result.

[0135] Table 1: Experimental results obtained based on the evaluator

[0136]

[0137] As can be seen from the above table, the generation result of the large language model with user input has the highest score, but it has the risk of leaking privacy information, while the generation result of the large language model with general instructions has the lowest score. Although the small language model with user input has the strongest privacy protection ability, the score of the generation result is relatively low. The outline-based method and the logit-based method have better generation results than the small language model under the premise of ensuring that privacy information is not leaked. Therefore, the system using large language models and small language models according to the embodiments of the present disclosure can improve the quality of generation results while ensuring user privacy.

[0138] Table 2 shows the experimental results of large language models and small language models obtained by manual acquisition and automatic evaluation methods according to some embodiments of the present disclosure. The experimental results include scores for a small language model with user input, a large language model with user input, a large language model with general instructions, an outline-based method, and a logit-based method. The manually obtained score represents the comparison result of the current model relative to the small language model with user input, recorded as win / draw / lose. In addition, an automatic evaluation method is also used, using the traditional lexical overlap indicator BLEU (Bilingual Evaluation Understudy) to measure the degree of n-gram matching between the generated text and the reference text, and ROUGE-L (Recall-Oriented Understudy for Gisting Evaluation-Longest Common Subsequence) is used to evaluate content similarity based on the longest common subsequence.

[0139] Table 2: Experimental results obtained by manual and automatic evaluation methods

[0140] Classification Win / Draw / Lose (%) BLEU ROUGE-L Small language model with user input - / 50 / - 2.07 13.95 Large language model with user input 38 / 2 / 10 2.61 14.66 Large language model with universal instructions 3 / 0 / 47 1.51 13.54 Outline-based approach 27 / 3 / 20 1.81 12.98 Logit-based methods 32 / 5 / 13 2.30 14.18

[0141] It can be seen from the above table that the evaluation results obtained by manual and automatic evaluation methods are consistent with the evaluation results obtained based on the evaluator. According to the embodiment of the present disclosure, the system using large language models and small language models can ensure user information security and improve generation quality.

[0142] Therefore, some embodiments of the present disclosure remove privacy information by extracting general instructions from user input, so that the remote large language model can only obtain general instructions, thereby improving the security of privacy information. The local small language model can generate an answer to the user input based on the prompt information provided by the large language model and combined with the complete user input, thereby improving the quality of the answer. Through the collaborative interaction of the large language model and the small language model, the advantages of the large language model in terms of general knowledge, general planning, etc. are obtained, and high-level guidance or assistance is provided to the small language model to ensure the quality and depth of the final result. The question-answering process strikes a balance between privacy and performance, and can achieve a generation quality close to that of the large language model when using complete user input for reasoning, without leaking sensitive data. In addition, the collaborative processing of the remote large language model and the local small language model can reduce the computing power requirements of the local device, so that the system can be deployed on terminal devices with limited resources.

[0143] The embodiments of the present disclosure can be applied to enterprise local deployment, personal terminal devices, medical and legal situations with high requirements for privacy information, and can protect data security to the greatest extent without sacrificing generation quality.

[0144] According to another aspect of the present disclosure, a method for security question and answer is provided.

[0145] Figure 12 is a flow chart of a method 1200 for security question answering according to some embodiments of the present disclosure. Figure 1 Executed by the processor 102 in .

[0146] Method 1200 may include step S1202: receiving user input locally.

[0147] The method 1200 may include step S1204: extracting a general instruction from the user input, where the general instruction does not include the user's private information.

[0148] The method 1200 may include step S1206: providing the general instruction to a large language model to generate prompt information based on the general instruction, where the large language model is deployed remotely.

[0149] Method 1200 may include step S1208: calling a small language model to generate an answer to the user input based on the user input and prompt information, where the small language model is deployed locally.

[0150] In some embodiments, step S1204 may include: using a classification model to determine whether each sentence in the user input contains private information; and extracting general instructions from the user input based on the determination result for each sentence in the user input.

[0151] In some embodiments, step S1204 may include: calling a small language model to extract event tasks from user input to generate general instructions.

[0152] In some embodiments, step S1204 may include: performing rule judgment based on the user input and the keyword table to identify private information in the user input that matches the keyword table; and extracting general instructions from the user input.

[0153] Figure 13 is a flow chart of a method 1300 for security question answering according to other embodiments of the present disclosure. Figure 1 Executed by the processor 102 in .

[0154] Method 1300 may include step S1302: receiving user input locally.

[0155] Method 1300 may include step S1304: extracting general instructions from user input, where the general instructions do not include the user's private information.

[0156] The method 1300 may include step S1306 : providing the general instruction to a large language model to generate prompt information based on the general instruction, where the large language model is deployed remotely.

[0157] Step S1306 may include step S1308: providing the general instruction to the large language model to generate an answer outline based on the general instruction, the answer outline being used as prompt information.

[0158] Method 1300 may include step S1310: calling a small language model to generate an answer to the user input based on the user input and prompt information, where the small language model is deployed locally.

[0159] In some embodiments, step S1310 may include: calling the small language model to optimize the answer outline based on the user input; and calling the small language model to generate an answer to the user input based on the user input and the optimized answer outline.

[0160] In some embodiments, step S1304 may include: using a classification model to determine whether each sentence in the user input contains private information; and extracting general instructions from the user input based on the determination result for each sentence in the user input.

[0161] In some embodiments, step S1304 may include: calling a small language model to extract event tasks from user input to generate general instructions.

[0162] In some embodiments, step S1304 may include: performing rule judgment based on the user input and the keyword table to identify private information in the user input that matches the keyword table; and extracting general instructions from the user input.

[0163] Figure 14 is a flow chart of a method 1400 for security question answering according to some other embodiments of the present disclosure. Figure 1 Executed by the processor 102 in .

[0164] Method 1400 may include step S1402: receiving user input locally.

[0165] Method 1400 may include step S1404: extracting general instructions from user input, where the general instructions do not include the user's private information.

[0166] The method 1400 may include step S1406 : providing the general instruction to a large language model to generate prompt information based on the general instruction, where the large language model is deployed remotely.

[0167] Step S1406 may include step S1408: providing the general instruction to the large language model to generate a first probability distribution for the next output word position based on the general instruction, the first probability distribution being used as hint information.

[0168] Method 1400 may include step S1410: calling a small language model to generate an answer to the user input based on the user input and prompt information, where the small language model is deployed locally.

[0169] Step S1410 may include step S1412: calling the small language model to generate a second probability distribution for the next output word-unit position based on the user input.

[0170] Step S1410 may include step S1414: fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution for the next output word position.

[0171] Step S1410 may include step S1416: determining an output word-gram for a next output word-gram position based on the fused probability distribution.

[0172] In some embodiments, step S1414 may include: using the fusion model to obtain a fused probability distribution for the next output word-gram position based on the first probability distribution and the second probability distribution.

[0173] In some embodiments, step S1414 may include: fusing the probabilities of candidate word-grams corresponding to a predetermined number of highest probabilities in the first probability distribution and the second probability distribution.

[0174] In some embodiments, step S1404 may include: using a classification model to determine whether each sentence in the user input contains private information; and extracting general instructions from the user input based on the determination result for each sentence in the user input.

[0175] In some embodiments, step S1404 may include: calling a small language model to extract event tasks from user input to generate general instructions.

[0176] In some embodiments, step S1404 may include: performing rule judgment based on the user input and the keyword table to identify private information in the user input that matches the keyword table; and extracting general instructions from the user input.

[0177] According to another aspect of the present disclosure, a computer-readable storage medium is provided.

[0178] Figure 15 is a block diagram of a computer-readable storage medium 1500 according to some embodiments of the present disclosure.

[0179] The computer readable storage medium 1500 stores a computer program 1550. When the computer program 1550 is executed by the processor, the computer program 1550 realizes the above combination Figure 12-14 The steps of each method are described.

[0180] According to another aspect of the present disclosure, a computer program product is provided.

[0181] Figure 16 is a block diagram of a computer program product 1600 according to some embodiments of the present disclosure.

[0182] The computer program product 1600 may include a computer program 1550. When the computer program 1550 is executed by a processor, the computer program 1550 implements the above-mentioned Figure 12-14 The steps of each method are described.

[0183] The embodiments of the present disclosure have been described with reference to the accompanying drawings, which are intended to be illustrative rather than restrictive.

Claims

1. A method for secure question answering, characterized in that: include: Receive user input locally; extracting a general instruction from the user input, where the general instruction does not include the user's private information; providing the general instruction to a large language model to generate prompt information based on the general instruction, wherein the large language model is deployed remotely; as well as A small language model is called to generate an answer to the user input based on the user input and the prompt information, where the small language model is deployed locally.

2. The method according to claim 1, wherein Extracting general instructions from the user input includes: Using a classification model to determine whether each sentence in the user input contains private information; and The general instruction is extracted from the user input based on the judgment result for each sentence in the user input.

3. The method according to claim 1, wherein Extracting general instructions from the user input includes: The small language model is called to perform event task extraction on the user input to generate the general instruction.

4. The method according to claim 1, wherein Extracting general instructions from the user input includes: Performing rule judgment based on the user input and the keyword table to identify private information in the user input that matches the keyword table; and The general instruction is extracted from the user input.

5. The method according to claim 1, wherein Providing the general instruction to a large language model to generate prompt information based on the general instruction includes: The general instruction is provided to the large language model to generate an answer outline based on the general instruction, the answer outline serving as the prompt information.

6. The method according to claim 5, wherein Calling the small language model to generate an answer to the user input based on the user input and the prompt information includes: Invoking the small language model to optimize the answer outline based on the user input; and The small language model is called to generate an answer to the user input based on the user input and the optimized answer outline.

7. The method according to claim 1, wherein Providing the general instruction to a large language model to generate prompt information based on the general instruction includes: providing the general instruction to the large language model to generate a first probability distribution for a next output word position based on the general instruction, wherein the first probability distribution serves as the hint information; Calling the small language model to generate an answer to the user input based on the user input and the prompt information includes: Invoking the small language model to generate a second probability distribution for the next output word position based on the user input; fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution for the next output word position; and An output word-gram for the next output word-gram position is determined based on the fused probability distribution.

8. The method according to claim 7, wherein Fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution for the next output word position includes: The fused probability distribution for the next output word-gram position is obtained based on the first probability distribution and the second probability distribution using a fusion model.

9. The method according to claim 7, wherein Fusing the first probability distribution and the second probability distribution to obtain a fused probability distribution for the next output word position includes: The probabilities of candidate word-grams corresponding to a predetermined number of highest probabilities in the first probability distribution and the second probability distribution are fused.

10. A system for security question answering, characterized in that: include: Small language model; as well as A processor configured to perform the method according to any one of claims 1 to 9.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Large model reasoning acceleration method and system combining machine learning and speculation sampling

    CN118657220A

  • Question and answer method and device, electronic equipment, storage medium and program product

    CN119203967A

  • Input method lexicon updating method and apparatus, device and server

    WO2023030266A1

Cited By

  • End-side large model reasoning acceleration method and device, equipment, storage medium and product

    CN121684052A