Large model inference method and apparatus, and device and storage medium

By introducing Intel TDX technology and online inference business system into the big model, combining remote authentication services and end-to-end encryption, the security risks of the big model are solved, the privacy protection of the model in a trusted execution environment is realized, the risks of model theft and attack are reduced, and user data is protected.

WO2025145543A1PCT designated stage expired Publication Date: 2025-07-10SHANDONG INSPUR SCI RES INST CO LTD

Patent Information

Application Number
PCT/CN2024/104740
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-04
Filing Date
2024-07-10
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing large models face external security risks such as model attacks and intellectual property theft, as well as internal security risks such as data breaches and data abuse, which need to address privacy protection issues.

Method used

Through Intel TDX technology, the model runs in a trusted execution environment, and combines online inference business systems, remote authentication services and end-to-end encryption technology to perform login permission verification, access permission verification and data encryption and decryption to ensure the security of the model and the privacy protection of user data.

Benefits of technology

It greatly reduces the risk of model theft and attack, protects user-side data, prevents illegal parties from obtaining user privacy input, and realizes a trusted inference service that protects privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024104740_10072025_PF_FP_ABST
    Figure CN2024104740_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence. Disclosed are a large model inference method and apparatus, and a device and a storage medium. The method comprises: performing permission verification on login information; acquiring an inference request, symmetrically encrypting data to be inferred that is in the inference request, and encrypting a symmetric encryption key and then storing same in a remote authentication service; performing permission verification on the current user by means of an online inference service system, and issuing the inference request to an online inference engine, wherein the online inference engine and the remote authentication service run in a TDX trusted execution environment; and performing legitimacy verification on the inference request by means of the online inference engine, applying for the symmetric encryption key from the remote authentication service, using said data to perform inference, and encrypting an obtained initial inference result. The protection of a model itself is realized on the basis of a TDX technique, such that the model runs in a trusted execution environment, thereby greatly reducing risks such as model stealing or attack, and realizing protection of user-side data.
Need to check novelty before this filing date? Find Prior Art

Description

Large model reasoning method, device, equipment and storage medium Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a large-scale model reasoning method, apparatus, device, and storage medium. Background Art

[0002] The rise and application of large models like ChatGPT (Chat Generative Pre-trained Transformer) presents a range of security risks. These include external security risks, such as model attacks and intellectual property theft, and internal security risks, such as data leakage and misuse. For example, a company employee uploaded code to a large model service for optimization, leading to code leaks. Given these issues, how to perform trustworthy reasoning on large models with privacy protection remains an unresolved challenge in this field.

[0003] Summary of the Invention

[0004] In light of this, the present invention aims to provide a large-model inference method, apparatus, device, and storage medium. This method, based on Intel TDX technology, protects the model itself, allowing the model to run in a trusted execution environment, significantly reducing the risk of model theft or attack, and ultimately protecting user-side data. The specific solution is as follows:

[0005] In a first aspect, the present application provides a large model reasoning method, comprising:

[0006] Verify the login permissions of the current user's login information through the online reasoning business system;

[0007] After the login authority verification is passed, the inference request submitted by the current user through the online inference service application determined based on the login authority verification is obtained through the common client, and the encrypted data obtained by symmetric encryption of the data to be inferred in the inference request is sent to the online inference business system, and the symmetric encryption key corresponding to the encrypted data is encrypted and saved to the remote authentication service;

[0008] Performing access rights verification on the current user through the online reasoning service system, and sending the reasoning request to the online reasoning engine when the access rights verification passes; the online reasoning engine and the remote authentication service run in the TDX trusted execution environment;

[0009] The online reasoning engine performs a legitimacy check on the reasoning request, and when the legitimacy check passes, applies to the remote authentication service for the symmetric encryption key to obtain the symmetric encryption key, and uses the encrypted data to be decrypted based on the symmetric encryption key to perform reasoning, and uses the symmetric encryption key to encrypt the initial reasoning result to obtain an encrypted reasoning result, so as to return the encrypted reasoning result to the ordinary client through the online reasoning business system.

[0010] Optionally, before obtaining, through a common client, an inference request submitted by the current user through an online inference service application determined based on the login authority verification, the method further includes:

[0011] A trusted application is created by a management client to deploy the online reasoning service application based on preset service call restriction rules in combination with a preset online reasoning service and the trusted application.

[0012] Optionally, after deploying the online reasoning service application in combination with the preset online reasoning service and the trusted application, the method further includes:

[0013] A unique identifier corresponding to each online reasoning service application is generated based on a preset identifier generation rule, the online reasoning service application is signed based on a hash algorithm, and the online reasoning service application is registered in the remote authentication service.

[0014] Optionally, after deploying the online reasoning service application in combination with the preset online reasoning service and the trusted application, the method further includes:

[0015] If the online reasoning service application is updated, updating the unique identifier corresponding to the online reasoning service application, and signing the online reasoning service application based on the hash algorithm and the unique identifier before the update;

[0016] If the online reasoning service application is removed from the shelves, an application removal instruction is sent to the online reasoning engine through the online reasoning business system, so that the online reasoning engine verifies the application removal instruction and then removes the online reasoning service application.

[0017] Optionally, encrypting the symmetric encryption key corresponding to the encrypted data and then saving it to the remote authentication service includes:

[0018] Encrypting the symmetric encryption key using a preset public key of the remote authentication service, and signing the encrypted symmetric encryption key using a first private key generated by an RSA encryption algorithm to obtain a target key;

[0019] The target key is sent to the remote authentication service, so that the remote authentication service encrypts and stores the target key after verifying it based on the target key.

[0020] Optionally, the process of verifying the access rights of the current user by the online reasoning service system and sending the reasoning request to the online reasoning engine when the access rights verification passes further includes:

[0021] The resource usage of the online reasoning engine is monitored by the online reasoning service system, so as to send the reasoning request to the online reasoning engine according to the resource usage.

[0022] Optionally, after reasoning the data to be inferred obtained by decrypting the encrypted data based on the symmetric encryption key, the method further includes:

[0023] If the data to be inferred obtained after decryption is identified as sensitive information by the online inference engine, prompt information is generated based on the sensitive information, so that the prompt information is returned to the ordinary client for display through the online inference business system.

[0024] In a second aspect, the present application provides a large model reasoning system, comprising:

[0025] Online reasoning business system, used to verify the login permissions of the current user's login information;

[0026] The common client is used to obtain, after the login authority verification is passed, an inference request submitted by the current user through the online inference service application determined based on the login authority verification, and send the encrypted data obtained by symmetric encryption of the data to be inferred in the inference request to the online inference business system, and encrypt the symmetric encryption key corresponding to the encrypted data and save it to the remote authentication service;

[0027] The online reasoning service system is further configured to perform access rights verification on the current user, and to send the inference request to the online reasoning engine when the access rights verification passes; the online reasoning engine and the remote authentication service run in a TDX trusted execution environment;

[0028] The online reasoning engine is used to perform a legitimacy verification on the reasoning request, and when the legitimacy verification passes, apply for the symmetric encryption key from the remote authentication service to obtain the symmetric encryption key, and use the encrypted data obtained by decrypting the encrypted data based on the symmetric encryption key to perform reasoning, and use the symmetric encryption key to encrypt the initial reasoning result obtained to obtain an encrypted reasoning result, so as to return the encrypted reasoning result to the ordinary client through the online reasoning business system.

[0029] In a third aspect, the present application provides an electronic device comprising a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the aforementioned large model reasoning method.

[0030] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which implements the aforementioned large model reasoning method when executed by a processor.

[0031] In this application, the login permission verification is performed on the login information of the current user through the online reasoning business system; after the login permission verification is passed, the reasoning request submitted by the current user through the online reasoning service application determined based on the login permission verification is obtained through the ordinary client, and the encrypted data obtained after symmetric encryption of the data to be inferred in the reasoning request is sent to the online reasoning business system, and the symmetric encryption key corresponding to the encrypted data is encrypted and saved to the remote authentication service; the access permission verification is performed on the current user through the online reasoning business system, so that the reasoning request is sent to the online reasoning engine when the access permission verification is passed; the online reasoning engine and the remote authentication service run in the TDX trusted execution environment; the legitimacy of the reasoning request is verified by the online reasoning engine, and when the legitimacy verification is passed, the symmetric encryption key is applied to the remote authentication service to obtain the symmetric encryption key, and the data to be inferred obtained after decrypting the encrypted data based on the symmetric encryption key is used to perform reasoning, and the initial reasoning result obtained is encrypting using the symmetric encryption key to obtain an encrypted reasoning result, so that the encrypted reasoning result is returned to the ordinary client through the online reasoning business system. In this way, the present application discloses a large-model trusted reasoning service method that supports privacy protection. It implements the protection of the model itself based on Intel TDX technology, enables the model to run in a trusted execution environment, greatly reduces the risks of model theft or attack, and verifies the client and the reasoning system through the remote authentication service running in the trusted execution environment. The online service provider cannot obtain the user's privacy input. On the other hand, it realizes the protection of user-side data based on end-to-end encryption, sensitive identification and other technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0033] FIG1 is a flow chart of a large model reasoning method provided by this application;

[0034] FIG2 is a schematic diagram of a large model reasoning system provided by this application;

[0035] FIG3 is a flowchart of a specific large model reasoning method provided by this application;

[0036] FIG4 is a schematic diagram of the structure of a large model reasoning device provided by this application;

[0037] FIG5 is a structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] Currently, large models such as ChatGPT face external security risks, such as model attacks and intellectual property theft, as well as internal security risks, such as data leakage and data abuse. This application discloses a large-model trusted inference service method that supports privacy protection. This method uses Intel TDX (Trust Domain Extension) technology to protect the model itself, allowing the model to run in a trusted execution environment, significantly reducing the risks of model theft or attacks. Furthermore, a remote authentication service running in the trusted execution environment verifies the client and inference system, protecting user-side data.

[0040] As shown in FIG1 , an embodiment of the present invention discloses a large model reasoning method, including:

[0041] Step S11: Verify the login authority of the current user's login information through the online reasoning business system.

[0042] As shown in Figure 2, the online reasoning service in this embodiment mainly includes four modules: client, online reasoning business system, online reasoning engine, and remote authentication service. The online reasoning business system receives requests from the client and processes the requests, including trusted application management, online reasoning service, etc.

[0043] In this embodiment, user roles primarily include system administrators and regular users. System administrators are primarily responsible for application management and resource monitoring for online reasoning services, such as application deployment, updates, delisting, resource monitoring, and allocation. Regular users, as demanders of reasoning services, initiate reasoning services based on their needs. Therefore, in this embodiment, the online reasoning service system first verifies the user's login information for login permissions. After a successful login, the user can view the corresponding authorized online reasoning applications to ensure user identity security.

[0044] Step S12: After the login authority verification is passed, the inference request submitted by the current user through the online inference service application determined based on the login authority verification is obtained through the ordinary client, and the encrypted data obtained by symmetric encryption of the data to be inferred in the inference request is sent to the online inference business system, and the symmetric encryption key corresponding to the encrypted data is encrypted and saved to the remote authentication service.

[0045] As shown in Figure 2, in this embodiment, the client is divided into two types: a regular client and a management client. Regular clients can be used to initiate online reasoning, encrypt and decrypt data, and perform other functions. Therefore, in this embodiment, after passing the login permission verification, the user can select the corresponding online reasoning service and enter the reasoning application. This first involves obtaining the reasoning request submitted by the current user through the online reasoning service application determined based on the login permission verification through the regular client. It is understood that after the user enters / submits the data to be inferred (such as an image, text, or language) to initiate the reasoning request, the encrypted data obtained by symmetric encryption of the inference data in the reasoning request can be sent to the online reasoning service system, and the symmetric encryption key corresponding to the encrypted data can be encrypted and stored in the remote authentication service. The online reasoning service system can receive requests from clients and distribute and process the requests, primarily including functions such as reasoning service management, task scheduling, and system management. System management covers user management and permission management, while client requests include trusted application management and online reasoning services. It is also understood that the remote authentication service in this embodiment is the trusted component of the system, providing trusted proof and key distribution functions for the online reasoning service and user identity. Remote authentication services are typically provided by trusted third-party services.

[0046] According to the above steps, clients include two types: regular clients and management clients. Regular clients primarily include functions such as online inference initiation, data encryption and decryption, and key management. Management clients include functions such as application deployment, updates, and delisting. Application management and trusted resource monitoring are primarily performed through the management client. Therefore, before receiving inference requests through regular clients, a trusted application must be created through the management client. This application is then deployed based on preset service call restriction rules, combined with the preset online inference service and trusted application. After the system administrator creates a trusted application through the management client, they deploy the compiled online inference service. During deployment, service call restriction rules can be set, such as a call limit, call fees (if required), and required memory resources. Furthermore, after deploying the online inference service application based on the preset online inference service and trusted application, a unique identifier corresponding to each online inference service application must be generated based on preset identifier generation rules, the online inference service application must be signed using a hash algorithm, and the online inference service application must be registered with the remote authentication service. During deployment, a unique identification ID is generated for each application, the application hash is signed, and the inference application is registered in the remote authentication service. This ensures that each online inference service application is executed in a trusted execution environment. After successful deployment, users can access the application through a normal client.

[0047] At this time, if the online reasoning service application is updated, the unique identifier corresponding to the online reasoning service application is updated, and the online reasoning service application is signed based on the hash algorithm and the unique identifier before the update; if the online reasoning service application is removed from the shelves, the online reasoning business system sends an application removal instruction to the online reasoning engine, so that the online reasoning engine verifies the application removal instruction and removes the online reasoning service application. It is understandable that administrators can monitor application resources. Since trusted applications for large-model online reasoning occupy relatively large amounts of memory (such as 5G), it is necessary to reasonably allocate and dynamically expand application resources. At this time, the online reasoning engine serves as the core model of the system. When the business system processes trusted application management requests, if an application update is required, an ID will be regenerated for the application and re-signed. The signature attachment information must carry information such as the ID of the previous version. Application updates need to be authenticated with the remote authentication service. Only after the authentication is passed can the application be updated. If the application needs to be removed from the shelf, the administrator with the application removal authority will issue an instruction through the client, and the application removal instruction will be issued to the online reasoning engine through the business system. After the online reasoning engine verifies the removal instruction and the removal authority of the business system, the removal instruction will take effect and the trusted application will be removed.

[0048] Step S13: The online reasoning service system verifies the access rights of the current user, and sends the reasoning request to the online reasoning engine when the access rights verification passes; the online reasoning engine and the remote authentication service run in the TDX trusted execution environment.

[0049] In this embodiment, the online inference service system verifies the current user's access rights. If the access rights pass, the inference request is sent to the online inference engine. The online inference engine and remote authentication service operate within the TDX trusted execution environment, while other modules are deployed based on trusted resources. In this embodiment, the online inference engine is the core module of the system, serving as the specific online inference executor. It primarily includes functions such as encryption and decryption, sensitive identification, sensitive processing, online inference, and secure storage.

[0050] Step S14: perform a legitimacy check on the inference request through the online inference engine, and when the legitimacy check passes, apply for the symmetric encryption key from the remote authentication service to obtain the symmetric encryption key, and use the encrypted data based on the symmetric encryption key to decrypt the encrypted data to perform inference, and use the symmetric encryption key to encrypt the initial inference result to obtain an encrypted inference result, so as to return the encrypted inference result to the ordinary client through the online inference business system.

[0051] In this embodiment, the online reasoning engine can verify the legitimacy of the reasoning request, and when the legitimacy verification passes, apply for a symmetric encryption key from the remote authentication service, and use the symmetric encryption key to decrypt the encrypted data to obtain the data to be inferred for reasoning, and use the symmetric encryption key to encrypt the initial reasoning result to obtain an encrypted reasoning result, so as to return the encrypted reasoning result to the ordinary client through the online reasoning business system. At this time, when processing the online reasoning task request, the data is first decrypted. When decrypting, it is necessary to initiate a key application to the remote authentication service. After the application is passed, the corresponding key is obtained. After completing the online reasoning service according to the above key, the reasoning engine encrypts the data using the key and returns the ciphertext result and signature to the online reasoning business system. The online business system then verifies the signature and returns the ciphertext result of the reasoning to the client after the verification passes, so that the client can receive the ciphertext result and decrypt it using the private key to obtain the reasoning result.

[0052] In this embodiment, the online inference service system can verify the login permissions of the current user's login information. Furthermore, a trusted application can be created through the management client to deploy the online inference service application based on preset service call restriction rules, combined with the preset online inference service and the trusted application. A unique identifier for the online inference service application must be generated and signed before the online inference service application is registered with the remote authentication service. After the login permission verification passes, the inference request is obtained through a common client, and the encrypted data to be inferred is sent to the online inference service system. The symmetric encryption key corresponding to the encrypted data is encrypted and stored in the remote authentication service. The online inference service system then verifies the access permissions of the current user. If the verification passes, the inference request is sent to the online inference engine. The online inference engine and the remote authentication service run in the TDX trusted execution environment. The online inference engine verifies the legitimacy of the inference request and, if it passes, requests a symmetric encryption key from the remote authentication service. The encrypted data is then decrypted and inference is performed, and the initial inference result is encrypted. In this way, this application implements the protection of the model itself based on Intel TDX technology, allowing the model to run in a trusted execution environment, greatly reducing the risks of model theft or attack, and verifying the client and inference system through the remote authentication service running in the trusted execution environment. The online service provider cannot obtain the user's private input; on the other hand, the protection of user-side data is achieved based on end-to-end encryption, sensitive identification and other technologies.

[0053] Based on the previous embodiment, it can be seen that this application can ensure the trustworthiness of the remote attestation service based on Intel TDX technology to achieve the protection of the model itself and enable the model to run in a trusted execution environment. Next, this embodiment will elaborate on the key verification process and business reasoning process of the remote attestation service. Referring to Figure 3, this embodiment of the present invention discloses a specific large model reasoning method, including:

[0054] Step S21: Send the encrypted data obtained by symmetrically encrypting the data to be inferred in the inference request to the online inference business system, and encrypt the symmetric encryption key using the preset public key of the remote authentication service, and use the first private key generated by the RSA encryption algorithm to sign the encrypted symmetric encryption key to obtain the target key.

[0055] In this embodiment, the encrypted data obtained by symmetrically encrypting the data to be inferred in the inference request is sent to the online inference business system, and the symmetric encryption key is encrypted by the preset public key of the remote authentication service, and the encrypted symmetric encryption key is signed by the first private key generated by the RSA encryption algorithm to obtain the target key. It should be pointed out that if the user is logging in for the first time, his RSA public and private keys need to be generated. It is understandable that, by default, the inference data is encrypted with AES (Advanced Encryption Standard) and sent to the online inference business system, where the data encryption AES key is randomly generated. At the same time, the AES key is encrypted with the public key of the remote authentication service, signed with the RSA private key, and sent to the authentication service.

[0056] Step S22: Send the target key to the remote authentication service, so that the remote authentication service encrypts and stores the target key after verifying it based on the target key.

[0057] In this embodiment, after receiving the request, the remote authentication service needs to verify the user's identity and signature. If the verification is successful, the key is encrypted and stored using remote authentication. It should be noted that because the remote authentication service in this embodiment can guarantee its credibility, the remote authentication service can verify the user's identity and signature. If the verification is successful, the key is encrypted and stored using remote authentication. This ensures user credibility and the security of the key during storage and transmission.

[0058] Step S23: The online reasoning service system verifies the access rights of the current user, and sends the reasoning request to the online reasoning engine when the access rights verification passes; the online reasoning engine and the remote authentication service run in the TDX trusted execution environment.

[0059] In this embodiment, while the online inference service system verifies the current user's access rights, it can also monitor the resource usage of the online inference engine and issue inference requests to the online inference engine based on resource usage. It is understood that after receiving an inference request, the online service system verifies the user's access rights. After verification, the inference request is issued to the online inference service engine. To ensure the availability of the online inference engine, the inference service system monitors the inference service's memory and other resources in real time and issues tasks based on resource availability.

[0060] Step S24: perform a legitimacy check on the inference request through the online inference engine, and when the legitimacy check passes, apply for the symmetric encryption key from the remote authentication service to obtain the symmetric encryption key, and use the encrypted data based on the symmetric encryption key to decrypt the encrypted data to perform inference, and use the symmetric encryption key to encrypt the initial inference result to obtain an encrypted inference result, so as to return the encrypted inference result to the ordinary client through the online inference business system.

[0061] In this embodiment, after receiving an inference instruction, the online inference engine performs verification of the instruction's legitimacy and other aspects. After completing the business verification, it requests the user's AES key from the remote authentication service. After completing identity and application verification between the online inference engine and the remote authentication service, the AES key is encrypted using the online inference engine's public key by the remote authentication service. After successful subsequent verification, the key is decrypted using the online inference engine's private key. After decrypting the inference request data, the online inference engine can perform sensitivity identification on the decrypted plaintext. If the decrypted data to be inferred is identified as sensitive information by the online inference engine, a prompt message is generated based on the sensitive information, which is then returned to the standard client through the online inference service system for display. For example, if the online inference engine identifies information such as a user's home address, the prompt message will be displayed as "Sensitive data, unable to infer." Sensitive information may vary by country and region, and this embodiment does not specify this. If other information is identified, the inference phase begins. Each inference service has different internal logic for different models, and this embodiment does not specify this. After the inference is completed, the result is encrypted and signed and returned to the online inference business system. The business system then returns the inference result to the client, where the user uses the private key to decrypt it and obtain the inference result.

[0062] In this embodiment, the encrypted data obtained by symmetrically encrypting the data to be inferred in the inference request is sent to the online inference business system, and the symmetric encryption key is encrypted by the preset public key of the remote authentication service, and the encrypted symmetric encryption key is signed with the first private key generated by the RSA encryption algorithm to obtain the target key. The target key is sent to the remote authentication service so that the remote authentication service encrypts and stores the target key after verification based on the target key. The access rights of the current user are verified by the online inference business system, so that the inference request is sent to the online inference engine when the access rights verification passes, and then the inference request is verified for legitimacy by the online inference engine and subsequent inference is performed. In this embodiment, the model is run in a trusted execution environment, and the client and the inference system are verified by the remote authentication service running in the trusted execution environment. The online service provider cannot obtain the user's privacy input, thereby realizing the protection of user-side data.

[0063] As shown in FIG4 , the embodiment of the present application further discloses a large model reasoning system, including:

[0064] The online reasoning business system 11 is used to verify the login authority of the current user's login information;

[0065] The common client 12 is used to obtain, after the login authority verification is passed, an inference request submitted by the current user through the online inference service application determined based on the login authority verification, and send the encrypted data obtained by symmetric encryption of the data to be inferred in the inference request to the online inference business system 11, and encrypt the symmetric encryption key corresponding to the encrypted data and save it to the remote authentication service;

[0066] The online reasoning service system 11 is further configured to perform access rights verification on the current user, and to send the inference request to the online reasoning engine 13 when the access rights verification passes; the online reasoning engine 13 and the remote authentication service run in the TDX trusted execution environment;

[0067] The online reasoning engine 13 is used to verify the legitimacy of the reasoning request, and when the legitimacy verification is passed, apply for the symmetric encryption key from the remote authentication service to obtain the symmetric encryption key, and use the encrypted data based on the symmetric encryption key to decrypt the encrypted data to perform reasoning, and use the symmetric encryption key to encrypt the initial reasoning result to obtain an encrypted reasoning result, so as to return the encrypted reasoning result to the ordinary client 12 through the online reasoning business system.

[0068] In this embodiment, the login permission of the current user's login information is verified through the online reasoning business system; after passing the verification, the reasoning request submitted by the current user through the online reasoning service application determined based on the login permission verification is obtained through the ordinary client, and the encrypted data obtained after symmetric encryption of the data to be inferred in the reasoning request is sent to the online reasoning business system, and the symmetric encryption key corresponding to the encrypted data is encrypted and saved to the remote authentication service; the access permission of the current user is verified through the online reasoning business system, so that the reasoning request is sent to the online reasoning engine when passing; the online reasoning engine and the remote authentication service run in the TDX trusted execution environment; the legitimacy of the reasoning request is verified through the online reasoning engine, and when passing, a symmetric encryption key is applied to the remote authentication service, and reasoning is performed using the data to be inferred after decrypting the encrypted data, and the initial reasoning result obtained is encrypting using the symmetric encryption key to obtain an encrypted reasoning result. In this way, this embodiment protects the model itself based on Intel TDX technology, allowing the model to run in a trusted execution environment, greatly reducing the risks of model theft or attack. The client and inference system are verified through the remote authentication service running in the trusted execution environment, and the online service provider cannot obtain the user's private input. On the other hand, user-side data protection is achieved based on end-to-end encryption, sensitive identification and other technologies.

[0069] In some specific embodiments, the large model reasoning system further includes:

[0070] The management client is used to create a trusted application to deploy the online reasoning service application based on preset service call restriction rules in combination with the preset online reasoning service and the trusted application.

[0071] In some specific embodiments, the management client further includes:

[0072] The application registration module is used to generate a unique identifier corresponding to each online reasoning service application based on a preset identifier generation rule, sign the online reasoning service application based on a hash algorithm, and register the online reasoning service application in the remote authentication service.

[0073] In some specific embodiments, the management client further includes:

[0074] an identifier updating module, configured to update the unique identifier corresponding to the online reasoning service application if the online reasoning service application is updated, and sign the online reasoning service application based on the hash algorithm and the unique identifier before the update;

[0075] The application removal module is used to send an application removal instruction to the online reasoning engine through the online reasoning business system if the online reasoning service application is removed from the shelf, so that the online reasoning engine can remove the online reasoning service application after verifying the application removal instruction.

[0076] In some specific embodiments, the online reasoning service system 11 specifically includes:

[0077] A key signing module is used to encrypt the symmetric encryption key using the preset public key of the remote authentication service, and to sign the encrypted symmetric encryption key using the first private key generated by the RSA encryption algorithm to obtain a target key;

[0078] The key verification module is used to send the target key to the remote authentication service, so that the remote authentication service encrypts and stores the target key after verification based on the target key.

[0079] In some specific embodiments, the online reasoning service system 11 further includes:

[0080] The request sending module is used to monitor the resource usage of the online reasoning engine through the online reasoning business system, so as to send the reasoning request to the online reasoning engine according to the resource usage.

[0081] In some specific embodiments, the online reasoning engine 13 further includes:

[0082] An information prompt module is used to generate prompt information based on the sensitive information if the data to be inferred obtained after decryption is identified as sensitive information by the online inference engine, so as to return the prompt information to the ordinary client for display through the online inference business system.

[0083] Furthermore, an embodiment of the present application also discloses an electronic device. FIG5 is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0084] Figure 5 is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the large model inference method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0085] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0086] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0087] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the large model inference method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program capable of implementing other specific tasks.

[0088] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned large model inference method. The specific steps of this method can be referred to the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.

[0089] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0090] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0091] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0092] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further limitations, an element defined by the statement "comprises a..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.

[0093] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A large model inference method, characterized in that, Including: Performing login permission verification on the login information of the current user through an online inference service system; After the login permission verification passes, obtaining, through a general client, an inference request submitted by the current user through an online inference service application determined based on the login permission verification, and sending encrypted data obtained by symmetrically encrypting the data to be inferred in the inference request to the online inference service system, and encrypting the symmetric encryption key corresponding to the encrypted data and saving it to a remote authentication service; Performing access permission verification on the current user through the online inference service system to send the inference request to an online inference engine when the access permission verification passes; the online inference engine and the remote authentication service run in a TDX trusted execution environment; Performing legality verification on the inference request through the online inference engine, and when the legality verification passes, applying to the remote authentication service for the symmetric encryption key to obtain the symmetric encryption key, and performing inference on the data to be inferred obtained by decrypting the encrypted data based on the symmetric encryption key, and encrypting the obtained initial inference result with the symmetric encryption key to obtain an encrypted inference result, so as to return the encrypted inference result to the general client through the online inference service system.

2. The large model inference method according to claim 1, wherein Before obtaining, through the general client, the inference request submitted by the current user through the online inference service application determined based on the login permission verification, it further includes: Creating a trusted application through a management client to deploy the online inference service application based on a preset service call limit rule in combination with a preset online inference service and the trusted application.

3. The large model inference method according to claim 2, wherein After deploying the online inference service application in combination with the preset online inference service and the trusted application, it further includes: Generating a unique identifier corresponding to each online inference service application based on a preset identifier generation rule, signing the online inference service application based on a hash algorithm, and registering the online inference service application in the remote authentication service.

4. The large model inference method according to claim 3, wherein After deploying the online inference service application in combination with the preset online inference service and the trusted application, it further includes: If the online inference service application is updated, updating the unique identifier corresponding to the online inference service application, and signing the online inference service application based on the hash algorithm and the unique identifier before the update; If the online inference service application is taken off the shelf, sending an application take-off-the-shelf instruction to the online inference engine through the online inference service system, so that the online inference engine checks the application take-off-the-shelf instruction and then takes off the online inference service application.

5. The large model inference method according to any one of claims 1 to 4, characterized in that, The encrypting the symmetric encryption key corresponding to the encrypted data and saving it to the remote authentication service includes: Encrypting the symmetric encryption key with a preset public key of the remote authentication service, and signing the encrypted symmetric encryption key with a first private key generated by an RSA encryption algorithm to obtain a target key; Send the target key to the remote authentication service so that the remote authentication service encrypts and stores the target key after verification based on the target key.

6. The large model inference method according to claim 1, wherein During the process of verifying the access permission of the current user through the online inference service system and sending the inference request to the online inference engine when the access permission verification is passed, the following steps are further included: Monitor the resource usage of the online inference engine through the online inference service system, and send the inference request to the online inference engine according to the resource usage.

7. The large model inference method according to claim 1, wherein After performing inference on the data to be inferred obtained by decrypting the encrypted data using the symmetric encryption key, the following steps are further included: If the data to be inferred obtained after decryption is recognized as sensitive information by the online inference engine, generate a prompt message based on the sensitive information, so that the online inference service system returns the prompt message to the ordinary client for display.

8. A large model inference system, characterized in that, Including: An online inference service system for verifying the login permission of the current user's login information; An ordinary client, which is used to obtain an inference request submitted by the current user through an online inference service application determined based on the login permission verification after the login permission verification is passed, and send the encrypted data obtained by symmetrically encrypting the data to be inferred in the inference request to the online inference service system, and encrypt and save the symmetric encryption key corresponding to the encrypted data to the remote authentication service; The online inference service system is further used to verify the access permission of the current user, and send the inference request to the online inference engine when the access permission verification is passed; the online inference engine and the remote authentication service run in the TDX trusted execution environment; The online inference engine is used to verify the legality of the inference request, and when the legality verification is passed, apply to the remote authentication service for the symmetric encryption key to obtain the symmetric encryption key, and use the data to be inferred obtained by decrypting the encrypted data using the symmetric encryption key for inference, and use the symmetric encryption key to encrypt the obtained initial inference result to obtain an encrypted inference result, so as to return the encrypted inference result to the ordinary client through the online inference service system.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory; wherein, the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the large model inference method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is used to be stored, and when the computer program is executed by a processor, it implements the large model inference method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data security processing method and device

    CN115348023A

  • Information authentication method, data access method, system and electronic equipment

    CN116582332A

  • Data processing method and device based on trusted execution environment, equipment and medium

    CN116980163A

  • Key distribution method and device, computer equipment and storage medium

    CN117155549A

  • Large model reasoning method, device and equipment and storage medium

    CN118035988A

Cited By

  • Model weight parameter protection method, terminal equipment and storage medium

    CN121212349A

  • Large model service security verification method and device, medium, equipment and product

    CN121309223A

  • User service demand processing method and device, equipment and medium

    CN121690863A

  • RAG data access control method and system based on metadata filtering

    CN122174266A