Endogenous security large model reasoning system and method based on dynamic heterogeneous redundant architecture
Through dynamic heterogeneous redundant architecture and endogenous security mechanism, the problem of single structure and insufficient privacy protection of the large-model inference system is solved, and high security, high reliability and high compliance inference services are realized, and the system's anti-attack ability and data privacy protection capabilities are improved.
Patent Information
- Application Number
- CN202510715883.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing large-scale model inference system has a single structure, insufficient static defense and privacy protection, which leads to susceptibility to hardware failures, network abnormalities or malicious attacks, and it is difficult to achieve dynamic threat response and privacy protection, making it difficult to meet the needs of high security and high reliability in complex environments.
Using a dynamic heterogeneous redundancy architecture, combining differential privacy protection, real-time exception monitoring and response, and intelligent adjudication, multiple functionally equivalent but structurally heterogeneous model executors are built, and the final results are generated through dynamic scheduling and adjudication to achieve endogenous security and robustness.
It improves the attack resistance and fault tolerance of the large-model inference system, ensures the continuity and stability of inference services, effectively protects data privacy, improves the accuracy and credibility of inference results, and meets the security and compliance requirements in complex environments.
Smart Images

Figure CN120258151A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - technical field of artificial intelligence and information security, and particularly relates to an end -ogenous security large - model inference system and method based on a dynamic heterogeneous redundancy architecture. Background Art
[0002] With the breakthrough of deep learning, large language models and multi - modal large models have been rapidly deployed and applied in scenarios such as finance, healthcare, education, and intelligent customer service. Their inference services have become the core capabilities of many information systems.
[0003] However, most of the current inference frameworks for large models adopt a single and fixed system structure, lacking redundancy and fault - tolerance design. Once suffering from hardware failures, network anomalies, or targeted malicious attacks, they are extremely prone to single - point failures, resulting in the interruption of inference services. At the same time, these systems generally rely on static, feature - library - based security protection means, and are slow to respond to continuously evolving zero - day vulnerabilities or adversarial samples, making it difficult to achieve dynamic and real - time threat response.
[0004] In addition, large models need to process massive amounts of data containing sensitive information such as personal identities, medical records, and trade secrets during training and inference. However, existing solutions lack fine - grained privacy protection and audit mechanisms in data collection, transmission, and inference links, which are prone to data leakage and compliance risks. The above - mentioned vulnerabilities, static defense lags, and privacy leakage risks make it difficult for existing technologies to meet the requirements of high - security, high - reliability, and high - compliance operation of large models in complex environments. There is an urgent need for a new type of inference system and method that can enhance the end -ogenous security capabilities from the system architecture level. Summary of the Invention
[0005] The present invention aims to solve the problems of single structure, static defense, and insufficient privacy protection in existing large - model inference systems, and provides an end -ogenous security large - model inference system and method based on a dynamic heterogeneous redundancy architecture. By introducing a dynamic and structurally heterogeneous model execution body redundancy mechanism, combined with technical means such as differential privacy protection, real - time anomaly monitoring and response, and intelligent adjudication, the end -ogenous security and robustness of the large - model inference process are enhanced from the system architecture level, effectively resisting various attacks, protecting data privacy, and ensuring the reliability of inference results and the continuity of services in complex and highly adversarial environments.
[0006] According to the design scheme provided by the present invention, on the one hand, an end -ogenous security large - model inference system based on a dynamic heterogeneous redundancy architecture is provided, including:
[0007] An input pre - processing module, which receives the input data to be inferred, performs cleaning and standardization processing on it, and extracts multi - modal features to form a unified input representation;
[0008] Differential privacy protection unit, connected to the output end of the input preprocessing module, for injecting noise that meets the preset differential privacy constraints into the unified input representation to suppress privacy leakage;
[0009] Model execution body pool, including multiple model execution bodies that are functionally equivalent but structurally heterogeneous. The model execution bodies respectively implement the same inference function through heterogeneous network architectures, and each model execution body has undergone adversarial training to enhance robustness;
[0010] Dynamic scheduling module, connected to the output end of the differential privacy protection unit, for dynamically selecting at least two structurally heterogeneous model execution bodies from the model execution body pool for parallel inference according to the characteristics of the unified input representation injected with noise and the real-time system resource status;
[0011] Adjudication module, connected to the output end of the model execution body pool, for comprehensively evaluating the parallel inference output results of the at least two model execution bodies according to the task type using the preset adjudication logic to generate the final inference result;
[0012] Abnormal monitoring and response unit, communicating bidirectionally with the model execution body pool and the dynamic scheduling module, for continuously monitoring the system resource status, the confidence distribution of the selected model execution body, the running delay, and the input-output consistency during the entire inference process; when it detects that the inference confidence of any model execution body is lower than the threshold and / or the running delay exceeds the preset upper limit, it triggers an abnormal response operation;
[0013] Output post-processing module, connected to the output end of the adjudication module, for formatting the final inference result or performing output processing according to the user-specified standard.
[0014] As the endogeneous security large model inference system based on the dynamic heterogeneous redundancy architecture of the present invention, further, the input preprocessing module uses dedicated feature extraction networks for different modality data, including: using the BERT model to extract features for text data, using the ResNet model to extract features for image data, using the Wav2Vec model to extract features for audio data, using the VideoMAE model to extract features for video data, and fusing the modality features to form a unified input representation.
[0015] As the endogeneous security large model inference system based on the dynamic heterogeneous redundancy architecture of the present invention, further, the preset differential privacy constraints in the differential privacy protection unit are controlled by the privacy parameter to satisfy:
[0016]
[0017] where M is the differential privacy mechanism, D and D' are adjacent data sets, and S is the output set, is a privacy parameter, indicating the strength of privacy protection.
[0018] As the endogeneous security large model inference system based on the dynamic heterogeneous redundancy architecture of the present invention, further, the adversarial training is achieved by generating adversarial samples during the training process of the model executor and adding them to the training set. The generation of the adversarial samples adopts the following formula:
[0019] ;
[0020] where, is the generated adversarial sample, is the original input, is the perturbation size, is the loss function, and ∇x represents the gradient with respect to the input.
[0021] As the endogeneous security large model inference system based on the dynamic heterogeneous redundancy architecture of the present invention, further, the dynamic scheduling module makes scheduling decisions using the following formula:
[0022] ;
[0023] where, represents the selected model executor, is the set of all model executors, is the input data, is a scheduling function that comprehensively considers feature similarity, computational load, and memory occupancy.
[0024] As the endogeneous security large model inference system based on the dynamic heterogeneous redundancy architecture of the present invention, further, the adjudication module adopts a majority voting strategy for classification tasks, and the adjudication formula is:
[0025] ,
[0026] where, is the output result of the i-th model executor, and mode represents the operation of taking the mode;
[0027] For generation tasks, a confidence comparison strategy is adopted, and the adjudication formula is:
[0028] ,
[0029] where, represents the confidence of the output result of the i-th model executor.
[0030] As the endogeneous security large model inference system based on the dynamic heterogeneous redundancy architecture of the present invention, further, the anomaly monitoring and response unit adopts the following anomaly detection rules:
[0031] ;
[0032] where θ is the confidence threshold and T is the running time threshold; when an anomaly is detected, it triggers isolating the suspicious model execution body, dynamically adjusting the combination of model execution bodies, reallocating inference tasks, or issuing an alarm to the system administrator.
[0033] On the other hand, the present invention provides an endogeneous security large model inference method based on a dynamic heterogeneous redundancy architecture, including:
[0034] Receiving the data to be inferred, cleaning and standardizing it, and extracting multi-modal features to form a unified input representation;
[0035] Injecting noise that satisfies the preset differential privacy constraint into the unified input representation to suppress privacy leakage;
[0036] Constructing at least two model execution bodies with equivalent functions but heterogeneous network structures, and performing adversarial training on each model execution body to enhance the ability to resist adversarial sample attacks;
[0037] According to the features of the unified input representation injected with noise and the system resource status, dynamically selecting at least two structurally heterogeneous model execution bodies from the model execution body pool to perform parallel inference on the same data, and generating multiple inference results;
[0038] According to the task type, adopting a majority voting or confidence comparison strategy to adjudicate the output results of the at least two model execution bodies and generate a final inference result;
[0039] During the inference process, continuously monitor the confidence distribution, running delay, and input-output consistency of each model execution body. When the detected confidence is lower than the preset threshold or the running delay exceeds the limit, trigger an anomaly response operation including model switching, resource reallocation, or alarm;
[0040] Performing formatting or user-specified output processing on the final inference result.
[0041] The beneficial effects of the present invention are:
[0042] By introducing a dynamic heterogeneous redundancy architecture and an endogeneous security mechanism, the present invention effectively overcomes the deficiencies of existing large model inference systems in terms of security, robustness, and privacy protection. The system constructs multiple structurally heterogeneous and functionally equivalent model execution bodies, and combines dynamic scheduling and intelligent adjudication strategies. Even if some execution bodies are attacked or malfunction, it can ensure the continuity and stability of the inference service, thus significantly enhancing the anti-attack ability and fault tolerance of the overall system. The integration of the differential privacy protection unit realizes the effective protection of sensitive information during the inference input stage, reduces the risk of user data leakage and abuse, and meets the increasingly stringent data security and compliance requirements. In addition, combined with adversarial training and real-time anomaly monitoring and response, the system can quickly identify and isolate abnormal models, and has the ability of active defense and adaptation to constantly changing unknown attack threats, further enhancing the effectiveness of security protection. The dynamic scheduling mechanism can intelligently allocate heterogeneous model resources according to task characteristics and system load, improve inference efficiency, and achieve efficient resource utilization. By fusing the parallel inference results of multiple models and adopting a scientific adjudication method, the system can suppress unreliable outputs caused by single model errors or attacks, and finally improve the accuracy and credibility of the inference results. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic diagram of the system architecture of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions of the present invention will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Embodiment 1:
[0046] As Figure 1 shown, this embodiment provides an endogeneous security large model inference system based on a dynamic heterogeneous redundancy architecture. Its overall architecture includes an input preprocessing module, a differential privacy protection unit, a model execution body pool, a dynamic scheduling module, an adjudication module, an anomaly monitoring and response unit, and an output postprocessing module. The above modules are connected in series or in parallel through a data flow to form a complete inference system. This system can perform safe and reliable inference on input data and output high-quality inference results.
[0047] Design and Implementation of Input Preprocessing Module: This module is responsible for receiving the input data to be inferred, cleaning and standardizing it, and extracting multi-modal features to form a unified input representation. This module processes different types of data through a multi-modal feature extraction network: for text data, the BERT (Bidirectional Encoder Representations from Transformers) model is used to extract semantic features; for image data, the ResNet (Residual Network) model is used to extract visual features; for audio data, the Wav2Vec model is used to extract acoustic features; for video data, the VideoMAE (Video Masked Autoencoder) model is used to extract spatio-temporal features. Finally, the features of different modalities extracted are fused into a unified representation form, which can be expressed as:
[0048] ;
[0049] where, is the final input representation, are the feature representations of text, image, audio, and video respectively.
[0050] Design and Implementation of Differential Privacy Protection Unit: This unit is connected to the output end of the input preprocessing module and is used to inject noise that satisfies the preset differential privacy constraints into the unified input representation to suppress privacy leakage. The intensity of noise injection is precisely controlled by the privacy parameter ε, and this mechanism satisfies the mathematical definition of differential privacy:
[0051] ;
[0052] where, is the differential privacy mechanism, and are adjacent data sets, is the output set, is the privacy parameter, representing the intensity of privacy protection. A smaller value indicates a higher intensity of privacy protection, but may reduce the inference accuracy at the same time; a larger value indicates a lower intensity of privacy protection, but can maintain a higher inference accuracy. In this embodiment, the value can be dynamically adjusted according to the specific application scenario and data sensitivity to achieve a balance between privacy protection and inference accuracy.
[0053] Model Execution Body Pool Design and Implementation: The model execution body pool includes multiple model execution bodies that are functionally equivalent but structurally heterogeneous. Each model execution body implements the same inference function through different network architectures, making the entire system have structural diversity and functional redundancy. In this embodiment, the model execution body pool contains at least the following execution bodies with different architectures: execution bodies based on the CNN (Convolutional Neural Network) architecture, which are particularly suitable for processing image and video data; execution bodies based on the Transformer architecture, which are particularly suitable for processing text and sequence data; execution bodies based on the MoE (Mixture of Experts) architecture, which are suitable for processing complex multi-modal inputs. Each model execution body has undergone adversarial training to enhance robustness, and the adversarial training is achieved by generating adversarial samples during the model training process and adding them to the training set. The generation of adversarial samples uses the following formula:
[0054] ;
[0055] where, is the generated adversarial sample, is the original input, is the perturbation size, is the loss function, and ∇x represents the gradient of the input. This adversarial training enables the model to resist malicious perturbations and adversarial attacks, improving the overall security of the system.
[0056] Dynamic Scheduling Module Design and Implementation: This module is connected to the output end of the differential privacy protection unit and is used to dynamically select at least two structurally heterogeneous model execution bodies from the model execution body pool for parallel inference according to the characteristics of the unified input representation with injected noise and the real-time system resource status. In this embodiment, each dynamically selected model execution body respectively performs inference on the input data and generates an inference result, which is expressed by the formula:
[0057] ;
[0058] where, is the i-th model execution body, is the output result of the i-th model execution body.
[0059] The dynamic scheduling module uses the following formula for scheduling decisions:
[0060] ;
[0061] where, S represents the selected model execution bodies, is the set of all model execution bodies, x is the input data, and f is a scheduling function that comprehensively considers feature similarity, computational load, and memory occupancy. This function can be further expressed as:
[0062]
[0063] Among them, represents the model execution body and the input feature similarity (i.e., the processing ability of the model execution body for this type of input), represents the model execution body current computing load, represents the model execution body memory occupancy rate, , , are weight coefficients used to balance the importance of these three factors
[0064] In this way, the system can intelligently select the most suitable combination of model execution bodies for inference under different inputs and system states, achieving load balancing and resource optimization.
[0065] Design and implementation of the adjudication module: This module is connected to the output end of the model execution body pool and is used to comprehensively evaluate the parallel inference output results of at least two model execution bodies according to the task type, using a preset adjudication logic to generate the final inference result.
[0066] For classification tasks, in this embodiment, the majority voting strategy is adopted, and the adjudication formula is:
[0067] ,
[0068] Among them, is the output result of the i-th model execution body, and mode represents the operation of taking the mode. When multiple classes obtain the same number of votes, the system selects the class with the highest confidence as the final result.
[0069] For generation tasks, in this embodiment, the confidence comparison strategy is adopted, and the adjudication formula is:
[0070] ,
[0071] Among them, ) represents the confidence of the output result of the i-th model execution body.
[0072] This adjudication mechanism can effectively filter out abnormal or incorrect outputs and improve the reliability of the final inference result.
[0073] Design and Implementation of Anomaly Monitoring and Response Unit: This unit communicates bidirectionally with the model executor pool and the dynamic scheduling module, and is used to continuously monitor the confidence distribution, running delay, and input-output consistency of the selected model executors throughout the inference process. When it is detected that the inference confidence of any model executor is lower than the threshold and / or the running delay exceeds the preset upper limit, an anomaly response operation is triggered. In this embodiment, the anomaly monitoring and response unit adopts the following anomaly detection rules:
[0074] ;
[0075] where θ is the confidence threshold and T is the running time threshold; when an anomaly is detected (Anomaly = 1), the system will trigger a series of preset response operations, including: isolating the suspicious model executor and marking it as unavailable; dynamically adjusting the model executor combination and enabling standby executors; reallocating the inference task to healthy executors; sending an alarm to the system administrator and recording the anomaly information for subsequent analysis and processing. This continuous monitoring and rapid response mechanism can detect and handle anomalies in the inference process in a timely manner, improving the overall security and reliability of the system.
[0076] Design and Implementation of Output Post-Processing Module: This module is connected to the output end of the adjudication module and is used to format the final inference result or perform output processing according to user-specified criteria. In this embodiment, the output post-processing module can perform the following operations: converting numerical results into user-friendly text descriptions; generating structured result reports; formatting the output results according to the specified interface specifications; adjusting the detail level and presentation form of the results according to user preferences. The processing process can be expressed as:
[0077]
[0078] where, represents the post-processing function, which usually includes operations such as text generation and formatting.
[0079] Embodiment 2:
[0080] This embodiment provides an inference method for an endogeneous security large model based on a dynamic heterogeneous redundancy architecture, including the following steps:
[0081] S1. Receive the data to be inferred, clean and standardize it, and extract multi-modal features to form a unified input representation;
[0082] S2. Inject noise that satisfies the preset differential privacy constraint into the unified input representation to suppress privacy leakage;
[0083] S3. Construct at least two functionally equivalent but network - structure - heterogeneous model executors, and perform adversarial training on each of the model executors to enhance the ability to resist adversarial sample attacks;
[0084] S4. According to the features of the unified input representation with injected noise and the system resource status, dynamically select at least two structure - heterogeneous model executors from the model executor pool to perform parallel inference on the same data, and generate multiple inference results;
[0085] S5. Adopt a majority - voting or confidence - comparison strategy according to the task type to adjudicate the output results of the at least two model executors and generate a final inference result;
[0086] S6. Continuously monitor the confidence distribution, running delay, and input - output consistency of each model executor during the inference process. When it is detected that the confidence is lower than the preset threshold or the running delay exceeds the limit, trigger an exception response operation including model switching, resource re - allocation, or alarm;
[0087] S7. Perform formatting or user - specified output processing on the final inference result.
[0088] As described above, this is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. An endogenous security large model inference system based on a dynamic heterogeneous redundancy architecture, characterized in that, Comprising: An input preprocessing module that receives input data to be inferred, cleans and normalizes it, and extracts multimodal features to form a unified input representation; A differential privacy protection unit connected to the output end of the input preprocessing module, used to inject noise that satisfies preset differential privacy constraints into the unified input representation to suppress privacy leakage; A model execution body pool that includes multiple model execution bodies with equivalent functions but heterogeneous structures. The model execution bodies respectively implement the same inference function through heterogeneous network architectures, and each model execution body has undergone adversarial training to improve robustness; A dynamic scheduling module connected to the output end of the differential privacy protection unit, used to dynamically select at least two model execution bodies with heterogeneous structures from the model execution body pool for parallel inference according to the features of the unified input representation with injected noise and the real-time system resource status; A verdict module connected to the output end of the model execution body pool, used to comprehensively evaluate the parallel inference output results of the at least two model execution bodies according to the task type using a preset verdict logic to generate a final inference result; An anomaly monitoring and response unit that communicates bidirectionally with the model execution body pool and the dynamic scheduling module, used to continuously monitor the system resource status, the confidence distribution of the selected model execution bodies, the running delay, and the input-output consistency during the entire inference process; when it detects that the inference confidence of any model execution body is lower than the threshold and / or the running delay exceeds the preset upper limit, it triggers an anomaly response operation; An output postprocessing module connected to the output end of the verdict module, used to format the final inference result or perform output processing according to user-specified standards.
2. The inference system for an endogenic security large model based on a dynamic heterogeneous redundancy architecture according to claim 1, wherein The input preprocessing module uses dedicated feature extraction networks for different modal data, including: using a BERT model to extract features for text data, using a ResNet model to extract features for image data, using a Wav2Vec model to extract features for audio data, using a VideoMAE model to extract features for video data, and fusing the features of each modality to form a unified input representation.
3. The inference system for an endogenous security large model based on a dynamic heterogeneous redundancy architecture according to claim 1, characterized in that The preset differential privacy constraint in the differential privacy protection unit is controlled by a privacy parameter and satisfies: Among them, M is the differential privacy mechanism, D and D' are adjacent data sets, and S is the output set. is the privacy parameter, indicating the strength of privacy protection.
4. The inference system for an endogenous security large model based on a dynamic heterogeneous redundancy architecture according to claim 1, wherein The adversarial training is achieved by generating adversarial samples during the training process of the model execution body and adding them to the training set. The generation of the adversarial samples uses the following formula: ; Among them, is the generated adversarial sample, is the original input, is the perturbation magnitude, is the loss function, and ∇x represents the gradient with respect to the input.
5. The inference system for an endogenous security large model based on dynamic heterogeneous redundancy structure according to claim 1, characterized in that, The dynamic scheduling module uses the following formula for scheduling decisions: ; Among them, represents the selected model execution body, is the set of all model execution bodies, is the input data, is a scheduling function that comprehensively considers feature similarity, computational load, and memory occupancy.
6. The inference system for an endogenous security large model based on a dynamic heterogeneous redundancy architecture according to claim 1, characterized in that, The verdict module adopts a majority voting strategy for classification tasks, and the verdict formula is: , Among them, is the output result of the i-th model execution body, and mode represents the mode operation; Adopts a confidence comparison strategy for generation tasks, and the verdict formula is: , Among them, represents the confidence level of the output result of the i-th model execution body.
7. The inference system for an endogenous security large model based on a dynamic heterogeneous redundancy architecture according to claim 1, wherein The anomaly monitoring and response unit adopts the following anomaly detection rules: ; Where θ is the confidence threshold and T is the running time threshold; when an anomaly is detected, it triggers isolating the suspicious model execution body, dynamically adjusting the model execution body combination, reallocating the inference task, or sending an alarm to the system administrator.
8. An inference method for an endogeneous security large model based on a dynamic heterogeneous redundancy architecture, characterized in that Comprising: Receiving data to be inferred, cleaning and normalizing it, and extracting multimodal features to form a unified input representation; Injecting noise that satisfies preset differential privacy constraints into the unified input representation to suppress privacy leakage; Construct at least two model executors with equivalent functions but heterogeneous network structures, and perform adversarial training on each of the model executors to enhance the ability to resist adversarial sample attacks; According to the features of the unified input representation with injected noise and the system resource status, dynamically select at least two structurally heterogeneous model executors from the model executor pool to perform parallel inference on the same data, and generate multiple inference results; Adopt a majority voting or confidence comparison strategy according to the task type to adjudicate the output results of the at least two model executors and generate a final inference result; Continuously monitor the confidence distribution, running delay, and input-output consistency of each model executor during the inference process. When the detected confidence is lower than the preset threshold or the running delay exceeds the limit, trigger an exception response operation including model switching, resource reallocation, or alarm; Perform formatting or user-specified output processing on the final inference result.
Citation Information
Patent Citations
Data privacy pipeline providing collaborative intelligence and constraint computing
CN113678117A
Method and device for constructing endogenous security artificial intelligence system based on dynamic heterogeneous redundancy
CN116595511A
Multi-field fine-tuning large model parallel reasoning system and method thereof
CN117474102A
Dynamic heterogeneous redundant architecture data privacy protection method and device, equipment, medium and program product
CN118296639A
Mimicry judgment method and system based on relaxation similarity threshold
CN118523953A