Encryption protection method and system for deep learning large model

By dynamically adjusting model encoding accuracy and multimodal routing path selection, the performance bottlenecks and adaptability issues of large deep learning models in different hardware environments are resolved, achieving efficient and secure model deployment and operation.

CN120744952AActive Publication Date: 2025-10-03JIANGSU FENGYUN TECH SERVICE CO LTD

Patent Information

Application Number
CN202510907279.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-03
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing encoding technology optimization of large deep learning models lacks hardware perception capabilities, resulting in performance bottlenecks on low-computing power devices or loss of effective accuracy due to excessive compression on high-computing power devices. In addition, it lacks adaptability to multimodal scenarios and is difficult to dynamically adjust the sub-network structure, resulting in waste of computing resources or reduced cross-modal task performance.

Method used

By obtaining the resource status data of the target running hardware, the model encoding accuracy is dynamically adjusted, multimodal routing path selection is performed in combination with the task input modal information, and encryption binding and anomaly detection are performed to implement hardware-aware encoding strategies and multimodal routing path selection, thereby improving the matching efficiency and security of the model and hardware.

Benefits of technology

It achieves efficient matching between models and hardware, improves operational efficiency and accuracy, ensures stable and secure deployment in different environments, and prevents model leakage and illegal calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744952A_ABST
    Figure CN120744952A_ABST
Patent Text Reader

Abstract

The invention discloses an encryption protection method and system for a deep learning large model, and belongs to the technical field of model security deployment. The method comprises the following steps: obtaining a target hardware resource state, dynamically adjusting the model coding precision according to a preset standard, generating coding strategy data, and further executing hierarchical recoding on the initial weight of a large model to obtain a mixed precision model parameter. Coding sub-networks are selected and fused in combination with task modal information, and a unified coding model is constructed. And carrying out encryption binding according to the equipment binding information to generate an encryption model of hardware binding. And when the task is executed, performing anomaly detection and protection operation according to the input and operation states, and outputting model safety output data. According to the scheme, efficient matching of the model and hardware, task adaptability enhancement and full-link safety control in the using process are achieved. The operation efficiency and accuracy of the model are improved, leakage and illegal calling of the model are effectively prevented, and stable and safe deployment in different environments is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of model security deployment technology, and specifically relates to an encryption protection method and system for large deep learning models. Background Art

[0002] With breakthroughs in deep learning technology, large-scale pre-trained models have demonstrated outstanding performance in natural language processing, computer vision and other fields. However, the exponential growth in the number of parameters has led to increasingly prominent problems such as computing resource consumption, poor hardware adaptability, high inference latency and security risks.

[0003] Current mainstream large-model optimization technologies primarily focus on static compression and basic security protection. In terms of model encoding optimization, traditional methods often employ unified quantization strategies (such as 8-bit integer quantization of the entire model) or mixed-precision designs based on manual rules, sacrificing some accuracy in exchange for improved inference speed. At the hardware adaptation level, some solutions match different hardware configurations by presetting multiple sets of model copies, or rely on manual parameter tuning for operator-level optimization. In the security protection field, conventional methods include encrypted transmission of model parameters, input data filtering, and simple detection based on anomaly thresholds. However, the encryption process is decoupled from model inference, and anomaly detection is often based on an offline rule base.

[0004] However, existing coding technology optimization lacks hardware perception capabilities, and static quantization or unified precision strategies cannot dynamically adapt to the computing power, memory and other resource status of the target device, resulting in the optimized model still facing performance bottlenecks on low-computing power devices, or losing effective accuracy due to excessive compression on high-computing power devices; and lacks adaptability to multimodal scenarios. Traditional routing mechanisms rely on fixed path selection, and it is difficult to dynamically adjust the sub-network structure according to the input modality (such as text, image, voice), resulting in waste of computing resources or reduced cross-modal task effects. Summary of the Invention

[0005] The embodiments of the present application provide an encryption protection method and system for a large deep learning model, which solves the problems that the existing coding technology optimization lacks hardware perception capability, and the static quantization or unified precision strategy cannot dynamically adapt to the computing power, memory and other resource status of the target device, resulting in the optimized model still facing performance bottlenecks on low-computing power devices, or losing effective accuracy due to excessive compression on high-computing power devices; and the adaptability to multimodal scenarios is insufficient. The traditional routing mechanism relies on fixed path selection, and it is difficult to dynamically adjust the sub-network structure according to the input modality (such as text, image, voice), resulting in a waste of computing resources or a decline in the effectiveness of cross-modal tasks.

[0006] In a first aspect, an embodiment of the present application provides an encryption protection method for a large deep learning model, the method comprising: Obtain resource status data of the target operating hardware, and perform a dynamic adjustment operation on the model encoding accuracy according to the resource status data and a preset dynamic adjustment standard to obtain hardware-perceived encoding strategy data; Obtaining initial weight parameters of each layer of the target large model, performing precision layered re-encoding on the target large model according to the encoding strategy data and the initial weight parameters, and obtaining compressed and optimized mixed precision model parameter data; Obtaining task input modal information, performing a multimodal routing path selection operation on the target large model based on the task input modal information and mixed precision model parameter data, determining a target encoding subnetwork path, and performing encoding subnetwork selection and fusion operations on the target large model based on the target encoding subnetwork path to obtain a fused unified encoding model; Obtaining device binding information of a unified coding model and a target device, and performing an encryption binding operation on the unified coding model according to the device binding information to obtain a hardware-bound encryption model; Obtain input request data and the current running status of the encryption model, perform anomaly detection operations on the encryption model based on the input request data and the current running status, obtain anomaly detection result data, perform protection operations on the encryption model based on the anomaly detection result data, and obtain model security output data.

[0007] In a second aspect, an embodiment of the present application provides an encryption protection system for a large deep learning model, the system comprising: An encoding strategy data determination module is used to obtain resource status data of the target operating hardware, and perform a dynamic adjustment operation on the model encoding accuracy according to the resource status data and a preset dynamic adjustment standard to obtain hardware-perceived encoding strategy data; A precision layered recoding module is used to obtain initial weight parameters of each layer of the target large model, and perform precision layered recoding on the target large model according to the encoding strategy data and the initial weight parameters to obtain compressed and optimized mixed precision model parameter data; A unified coding model construction module is used to obtain task input modal information, perform a multimodal routing path selection operation on the target large model based on the task input modal information and mixed precision model parameter data, determine the target coding subnetwork path, and perform coding subnetwork selection and fusion operations on the target large model based on the target coding subnetwork path to obtain a fused unified coding model; An encryption model construction module is used to obtain device binding information of a unified coding model and a target device, and perform an encryption binding operation on the unified coding model according to the device binding information to obtain a hardware-bound encryption model; The model security output data determination module is used to obtain input request data and the current running status of the encryption model, perform anomaly detection operations on the encryption model based on the input request data and the current running status, obtain anomaly detection result data, perform protection operations on the encryption model based on the anomaly detection result data, and obtain model security output data.

[0008] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0010] The embodiments of this application achieve efficient matching of models and hardware, enhanced task adaptability, and full-link security control during use through resource-aware precision adjustment, multimodal path selection, device binding encryption, and runtime security protection. This not only improves model operation efficiency and accuracy, but also effectively prevents model leakage and illegal calls, ensuring stable and secure deployment in different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a flowchart of the encryption protection method for a large deep learning model provided in Example 1 of the present application; Figure 2 This is a flow chart of the encryption protection method for a large deep learning model provided in Example 2 of the present application; Figure 3 This is a schematic diagram of the structure of the encryption protection system for the deep learning large model provided in Example 3 of the present application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in Example 4 of the present application. DETAILED DESCRIPTION

[0012] To further clarify the objectives, technical solutions, and advantages of this application, specific embodiments of the present application are described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are intended only to illustrate this application and are not intended to limit it. It should also be noted that, for ease of description, the drawings only illustrate portions relevant to this application, not all of them. Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the various operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process may terminate upon completion of its operations, but may also include additional steps not shown in the accompanying drawings. The process may correspond to a method, function, procedure, subroutine, subprogram, or the like.

[0013] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0014] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0015] Below, in conjunction with the accompanying drawings, an RSMC chip, a chip multi-stage startup method, and a Beidou communication and navigation device provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.

[0016] Example 1 Figure 1 This is a flow chart of the encryption protection method for the deep learning large model provided in Example 1 of this application. Figure 1 As shown, the specific steps include: S101, obtaining resource status data of target operating hardware, performing a dynamic adjustment operation on model encoding accuracy according to the resource status data and a preset dynamic adjustment standard, and obtaining hardware-perceived encoding strategy data.

[0017] The target running hardware can be the specific device entity currently used to run the target large model, including edge devices (such as smartphones, edge servers, embedded AI chips), cloud servers, and local inference hardware (such as GPUs, TPUs, FPGAs, NPUs, etc.). The target running hardware is the actual running carrier bound to the model optimization and security mechanisms.

[0018] Resource status data can be real-time information about current operating resources obtained from the target hardware. This includes computing power metrics such as CPU / GPU / NPU utilization and floating-point operations per second (FLOPS). Memory status includes remaining / used RAM and cache status. Bandwidth status includes bus / network bandwidth usage. Power consumption and temperature are used to consider energy conservation or overheating protection. Current system load includes task queue length, I / O, and other indicators. This data is used to understand hardware resource conditions and serves as input for model optimization.

[0019] The preset dynamic adjustment standard can be a predefined set of model precision adjustment rules, which is used to dynamically select the appropriate model precision strategy based on resource status. This includes a resource-precision mapping rule table, for example: if the remaining video memory is greater than X, FP32 precision is allowed.

[0020] Dynamic model encoding precision adjustment refers to the process of performing layered, precision-adaptive compression or conversion on the model based on the aforementioned resource status data and dynamic adjustment criteria. Core technologies include mixed-precision quantization (FP16 and INT8 representations for some layers to reduce memory and computational overhead); structural pruning (pruning some neural layers, such as attention mechanisms and redundant convolutions); channel pruning and weight sharing; and sparsity enhancement (representing low-impact channels with sparse matrices to further reduce size). This operation essentially dynamically recodes the representation and precision of large model parameters, achieving resource-sensitive optimization.

[0021] The encoding policy data can be the output of the above dynamic adjustment operation, represented as a hardware-oriented model compression execution policy data structure. It can include information such as the precision type used by each layer (such as layer1: FP16, layer2: INT8), the encoding method of each layer's activation function or parameters, the retention / clipping path flag, whether to enable low-rank decomposition, weight sharing, and matrix sparsification.

[0022] During model loading or task run preparation, resource status data of the target hardware can be obtained in real time by calling the underlying operating system or hardware driver interface. This data may include dynamic operational indicators such as the processor's current computing resource utilization, remaining memory capacity (RAM and video memory), hardware power consumption, chip temperature, device bandwidth usage, and the current system task load. After standardization and structured packaging, this data forms a dataset representing the current hardware resource capabilities. The system is pre-configured with a set of dynamic precision adjustment standards that define precision adjustment strategies and trigger conditions for models under different hardware resource conditions. For example, when GPU memory utilization exceeds a preset threshold of 90%, some neural network layers should be prioritized to switch to INT8 or FP16 precision. When chip temperature approaches the upper limit, the computational complexity of some redundant layers can be reduced or structural pruning can be implemented to reduce computational pressure. Alternatively, when multiple tasks are running concurrently, the accuracy of the main path encoding module can be prioritized based on the importance of each task. During model initialization, the system uses policy matching or a lightweight rule engine to compare the collected resource status data against the preset dynamic adjustment standards to determine the appropriate model encoding precision strategy. Based on the adjustment standards, the output includes key encoding parameters such as the computing precision that each layer should adopt (such as FP32, FP16, INT8), whether to enable sparse representation, channel cropping ratio, activation function compression rules, etc., forming a structured hardware-aware encoding strategy data.

[0023] S102, obtaining initial weight parameters of each layer of the target large model, performing precision layered re-encoding on the target large model according to the encoding strategy data and the initial weight parameters, and obtaining compressed and optimized mixed precision model parameter data.

[0024] The target large model can refer to a large-scale deep neural network model to be deployed and execute tasks. It has a multi-layered feature extraction structure, typically including multiple encoding layers, decoding layers, attention modules, and multimodal cross-fusion modules. It can be a general-purpose basic model (such as ViT, GPT, Swin-Transformer, BERT, etc.) or a large-scale model customized for the task. Before inference, it requires compression and re-encoding based on hardware conditions and task requirements. It has a multi-layer deep network structure (usually more than dozens of layers) and contains multiple modal encoding paths (for example, processing images, text, speech, and other modalities simultaneously). The model is large in size and requires high computing and memory resources.

[0025] Initial weight parameters refer to the weight data originally trained in each neural network layer before re-encoding the target large model. These weights are represented in high-precision format (such as FP32) and directly determine the feature extraction and representation capabilities during model inference. They serve as the foundational data source for subsequent model optimization operations such as precision reconstruction, pruning, and quantization. During the compression and re-encoding process, they are adjusted using precision adaptation strategies (for example, conversion to FP16 / INT8).

[0026] Mixed-precision model parameter data refers to the optimized parameter set generated by re-encoding, compressing, and quantizing the initial weight parameters after hardware-aware encoding strategy adjustment. It combines encoded data of different precision levels (such as FP32, FP16, and INT8), significantly reducing computational and storage overhead while ensuring model performance.

[0027] First, the initial weight parameters for each layer of the target large model are obtained. Specifically, this process parses the model's hierarchical structure by accessing the large model's structure definition file and trained parameter files (such as PyTorch's .pt file, TensorFlow's .ckpt file, or ONNX-formatted models), extracting information such as the weight tensor (WeightTensor), bias (Bias), and activation function parameters for each layer. These initial weight parameters are typically stored in a high-precision format (such as FP32 or BF16) and cover multiple structural units such as convolutional layers, fully connected layers, attention modules, and normalization layers. The system then loads encoding strategy data generated based on hardware resource status. This strategy includes the acceptable precision level (such as FP32, FP16, INT8) for each layer, the quantization method (such as symmetric quantization, asymmetric quantization, perceptual quantization), the compression ratio (such as the number of pruned channels and the retention ratio), the sparsity control strategy (such as the degree of structured sparsity or unstructured sparsity), and a mapping table of hardware computing capabilities corresponding to each layer. On this basis, the system performs a precision layered recoding operation based on the combination of the coding strategy data and the initial weight parameters. This operation is carried out according to the following process: Strategy Parsing and Parameter Matching: The system traverses each structural unit of the large model, matching the initial weight parameters of each layer with its corresponding encoding strategy to determine the appropriate precision level and compression method for that layer. Specifically, the system sequentially traverses each structural unit of the model, including the attention layer, feedforward layer, normalization layer, and embedding layer. For each layer, the system first extracts its initial weight parameters, such as the dimension, value range, and sparsity of the weight matrix or bias vector. Next, the system retrieves the optimal precision level and compression method for that layer from the strategy table based on the encoding strategy data generated based on the current hardware resource status. For example, if the hardware only supports low-bit integer acceleration, the system configures the layer to a low-precision integer format and selects a compression method appropriate for this layer type, such as pruning, factorization, or sparsification. If the strategy table marks a layer as a computational bottleneck, the system prioritizes singular value decomposition to decompose the weight matrix into a low-rank matrix to reduce computational complexity. If the contribution of certain channels is low, the system performs channel pruning. After the matching is completed, the system applies the precision level and compression method to the original weights of the layer, completes the re-encoding and records them as compressed mixed-precision weights, while retaining the corresponding encoding configuration of each layer for subsequent inference path selection and deployment binding.

[0028] Multi-level precision recoding: For layers where the strategy requires high precision (such as input feature extraction layers or classification output layers), retain FP32 precision or use BF16 to preserve numerical accuracy; For intermediate or compressible layers, a quantization re-encoding operation is performed to compress the weights from FP32 to FP16 or INT8, while applying a scaling factor for linear mapping; Certain lightweight paths can adopt extreme compression strategies, such as INT4 representation.

[0029] Fusion of sparsification and pruning: While compressing accuracy, the low-amplitude elements in the weight matrix are removed according to the sparsity threshold specified by the policy data to achieve structural sparsity. At the same time, the neural channels of the convolutional or fully connected layers are pruned through the channel importance measurement, thereby reducing the amount of computation and model size.

[0030] Structural Encapsulation and Parameter Integration: After re-encoding, the weights of each layer are structurally encapsulated to generate mixed-precision model parameter data in a unified format. This data combines encoding results at different accuracies, different structural sparsities, and different compression strategies, and comes with a hierarchical mapping table and strategy index to support rapid parsing and scheduling of subsequent inference engines.

[0031] Based on the above technical solution, optionally, the target large model is re-encoded in a precision layered manner according to the encoding strategy data and the initial weight parameters to obtain compressed and optimized mixed precision model parameter data, including: Perform quantitative sensitivity analysis on the initial weight parameters of each layer in the target large model to obtain the accuracy allocation priority of each layer; Determine the precision bit width range supported by the target running hardware, allocate adaptive precision bit width parameters to each layer of the target large model according to the precision bit width range and the precision allocation priority data, perform precision layered recoding, and obtain compressed and optimized mixed precision model parameter data.

[0032] In this solution, the precision allocation priority can be an important indicator for allocating coding resources to each layer based on the analysis of the sensitivity of each layer to quantization precision in the large model. It reflects the importance of each layer in maintaining the original performance during the compression process.

[0033] The precision bit width range can be the parameter quantization precision range supported by the target operating hardware, which is usually determined by the hardware's computing power, accelerator instruction set, or memory bandwidth. For example, some edge devices or inference chips may only support 4-bit, 6-bit, 8-bit, or 16-bit quantization operations.

[0034] The precision bit width parameter can be a specific quantization bit width value specified for the weight or activation value of each layer of the model, indicating the compression accuracy of the layer in actual encoding.

[0035] First, a quantization sensitivity analysis is performed on the initial weight parameters of each layer in the target large model. Using the model's baseline performance on the validation set as a reference, the initial weight parameters of each layer are fixed-point quantized by simulating different quantization precisions (e.g., 4-bit, 6-bit, 8-bit, and 16-bit) layer by layer. The resulting degradation in overall model performance (such as accuracy, loss, and BLEU score) is observed. In each experiment, only the weights of a single layer are quantized, while the remaining layers maintain their original precision. This allows the layer's sensitivity to overall performance to be assessed. Layers with higher sensitivity are more sensitive to precision and should be prioritized for higher bit widths. After analyzing all layers for sensitivity, the system ranks the layers by performance degradation and generates a priority list for precision allocation. For example, Layer 5 > Layer 12 > Layer 3... indicates that Layer 5 is the most sensitive and should be prioritized for higher precision. Next, the system determines the supported precision bit width range based on the target hardware's resource status (e.g., GPU architecture, memory capacity, and supported instruction sets). For example, some embedded devices support only 4-bit and 8-bit, while high-performance GPUs support dynamic bit widths from 4-bit to 16-bit. The system iterates through each layer of the large model (e.g., convolutional layers, fully connected layers, attention modules, etc.), combines these two pieces of information, and executes a precision matching algorithm within each layer—typically a trade-off algorithm based on greedy search or constrained optimization strategies—to assign the most appropriate bit width for that layer. For example, if a layer has a high priority (indicating it is very sensitive to precision) and the hardware supports a 16-bit bit width, it will be assigned 16 bits; if it has a medium priority, it may be assigned 8 bits; and if it has a low priority, it may be compressed to 6 or 4 bits to maximize resource savings. After the bit width is assigned, the system performs a "tiered precision recoding" operation on each layer. This involves quantizing the original high-precision weights according to the layer's assigned bit width, mapping them to low-bit-width integers using uniform or non-uniform quantization. Auxiliary parameters such as the quantization scale factor and offset (zeropoint) are also calculated and recorded. In some model architectures, the activation values ​​are also recoded to ensure precision matching during forward inference. After re-encoding is completed at all levels, the system encapsulates all low-bit-width weights, quantization parameters, and bit-width allocation tables into the final mixed-precision model parameter data as the optimized model output for deployment. This greatly reduces the computational load and memory usage of the model during runtime without significantly sacrificing accuracy, thereby improving hardware adaptability and execution efficiency.

[0036] In this solution, refined compression of model parameters can be achieved by performing quantization sensitivity analysis on each layer of the large model and performing precision layered recoding based on the precision bit width range supported by the target hardware.

[0037] S103, obtain task input modal information, perform multimodal routing path selection operation on the target large model according to the task input modal information and mixed precision model parameter data, determine the target coding sub-network path, perform coding sub-network selection and fusion operation on the target large model according to the target coding sub-network path, and obtain a fused unified coding model.

[0038] The task input modality information may be the modality type, structural characteristics, and distribution of the input data involved in the current task to be executed.

[0039] The target encoding subnetwork path can refer to an optimal inference path selected within a large model structure based on the task input modality information and mixed-precision model parameter data, that is, a specific sequence of encoding subnetworks that needs to be activated during model execution. Features include dynamically adjustable path levels: dynamic selection of submodules based on factors such as task complexity, input features, and hardware performance; different paths activating different subnetwork structures: for example, selecting different numbers of TransformerBlocks in a ViT structure, or activating only the image path, text path, or joint path in a multimodal network; and paths with lightweight or performance trade-off characteristics: shallow, fast paths can be selected when resources are limited, while deep, full-precision paths can be activated when resources are sufficient.

[0040] A unified encoding model refers to the final model version with a consistent structure, unified interface, and deployment-ready execution, obtained after completing sub-network path selection and fusion operations. It incorporates specific sub-network modules cut from the larger model; optimized weight parameters through mixed-precision compression and hardware-aware tuning; a structure that integrates support for multimodal task inputs; encapsulation logic that can be bound to the target device; and activation path logic configured for input modality.

[0041] The system first obtains input data samples for the current task from the user's task submission or the upper-level scheduling system. It then analyzes these samples using an integrated multimodal feature recognition algorithm to identify their modality type and feature distribution. This algorithm is typically implemented based on the following techniques: For image data, a lightweight image feature extraction network (such as MobileNet or ResNet-18) is used to extract features such as the number of channels, size, and color space. For text data, regular expressions are used to determine structure, followed by a preprocessor to extract length, language type, and character set. The BERT tokenizer is then used to determine semantic complexity. For speech data, MFCC (Mel-Frequency Cepstral Coefficient) extraction is used to identify temporal features and audio modal properties. For video data, a combination of frame extraction and time series modeling (such as TSN or TSM modules) is used to analyze spatial-temporal modal features. Through this approach, the system constructs a "task input modality information" vector representing the characteristics of the input sample, including a modality type label, feature dimension vector, and modality strength score. Based on this modality information, the system then uses a modality-driven path selection strategy network to determine which sub-network paths should be activated within the model. This policy network can be implemented using the following methods: A multi-layer perceptron (MLP) encodes the modal information vector and outputs the activation probability of each sub-network path via softmax; or a graph neural network (GNN) can be used to represent the large model structure, with modal information as the initial node feature input and information propagation determining the optimal path subgraph; or a reinforcement learning-based path selection policy π(a|s), where s is the modal state and a is the sub-network path selection action. The set of paths determined by this policy network constitutes the "target encoding sub-network path" adapted to the modal characteristics of the current task.

[0042] The encoding subnetwork selection and fusion operations are then performed. The subnetwork extraction mechanism includes path matching and module mapping: Path matching and module mapping are performed using the target encoding subnetwork path. Based on the path labels output by the path selection strategy network (e.g., image encoding path A, text encoding path B, multimodal fusion path C), the system accurately locates the corresponding module subgraph in the model structure graph. These path labels can be mapped using module indexes, semantic labels (e.g., image_conv_block_1), or structural position encodings to ensure a one-to-one correspondence between paths and model structures.

[0043] Module decoupling and dynamic assembly: The system selects and decouples modules based on the target encoding subnetwork path. To support dynamic path activation, the large model is built according to modular standards during pre-training, with each functional component (such as block1, fusion_layer_1, text_encoder, etc.) independently callable. The system then decouples the modules required for the target path, forming an execution subgraph specific to the current task.

[0044] Path dependency and integrity checking: The system performs structural dependency analysis and integrity checks on the target encoding subnetwork path, automatically resolving the forward and backward dependencies between selected modules to ensure that necessary input preprocessing layers (such as normalization and embedding layers) and output adaptation modules (such as decoding and classification heads) are fully preserved. If there are structural gaps in the path, the system will call preset default modules or shared modules to complete them, ensuring a logically closed loop in the subgraph structure.

[0045] Subnetwork fusion operation: By performing module-level output fusion and modal alignment operations through the target encoding sub-network path, the system supports the collaborative enhancement of multimodal information at the structural and parameter levels, achieving the construction of a unified representation space and improving the generalization capabilities of multimodal tasks.

[0046] Parameter fusion strategy: For underlying modules shared by multiple paths (such as basic convolutional layers or word embedding layers required for both images and text), the system uses parameter weighted averaging or selective substitution to fuse them (such as αW1+(1-α)W2) to achieve performance compromise and knowledge transfer between different modal requirements.

[0047] Structural splicing and skip connections: The system performs dimensional adaptation and structural alignment on the output results of each sub-network, and realizes the effective fusion of multimodal features through tensor splicing, channel dimension increase or decrease, cross-layer connection and other means.

[0048] Fusion Bridge Layer Insertion: The system can determine the location where the fusion bridge layer needs to be inserted through the target encoding sub-network path, and insert a fusion layer with learnable capabilities between multimodal paths, such as gated fusion layer, modality attention module, etc., to achieve dynamic weighting and cross-attention of different modal features.

[0049] Normalization and numerical stabilization: To maintain training stability and consistent feature scaling, the system introduces LayerNorm or BatchNorm layers at key locations before and after fusion to mitigate the risk of exploding or vanishing gradients, improving the overall model's convergence speed and execution efficiency. Ultimately, based on this extraction and fusion process, the system constructs a unified encoding sub-model with adjustable structure, modality adaptability, and computational efficiency. This model can accurately process heterogeneous modal inputs, flexibly adapt to different task objectives, and can be directly deployed for execution in real-world inference environments.

[0050] On the basis of the above technical solution, optionally, a multimodal routing path selection operation is performed on the target large model according to the task input modal information and the mixed precision model parameter data to determine the target encoding sub-network path, including: Vectorizing the task input modal information to obtain a modal vector; Extracting computing resource consumption data and precision adaptability data of multiple encoding sub-networks from the mixed precision model parameter data, and calculating the modal adaptability score of each encoding sub-network path in the target large model based on the modal vector, computing resource consumption data, precision adaptability data, and a preset scoring function; The encoding sub-network path with the highest modality adaptability score is selected as the target encoding sub-network path.

[0051] In this solution, the modal vector can be a numerical representation obtained by vectorizing the modal information of the task input (such as images, text, audio, etc.). The purpose is to unify the feature representations of different modalities for subsequent model processing. For example: for image input: extract visual feature vectors (such as embeddings extracted by CNN or ViT). For text input: use a word embedding model (such as BERT, Word2Vec) to convert the text into token embeddings. For audio input: convert the mel-spectrogram or other acoustic features into vector form. The modal vector reflects the core information of the input data, such as semantics, structure, and style, and serves as the basic feature for subsequent encoding path selection.

[0052] An encoding subnetwork can refer to multiple data encoding paths within a larger model, each with different structural or computational characteristics. These subnetworks can include stacks of Transformer layers of varying depth or width; submodules with varying quantization precisions (e.g., 8-bit, 6-bit, 4-bit); encoding structures customized for different tasks or modalities (e.g., image-specific encoders vs. text-specific encoders); and the coexistence of lightweight and parameter-heavy subnetworks (e.g., MobileNet paths vs. ResNet paths). Each encoding subnetwork has different adaptation to the input modality, resource consumption, and output quality, and the system dynamically selects one based on a policy.

[0053] Computational resource consumption data can be the estimated computational resource usage information of each encoding subnetwork on the current or target running hardware, mainly including FLOPs (floating-point operations per second), number of model parameters, memory usage (MB), inference latency (ms), and energy consumption estimates (such as power consumption model evaluation).

[0054] Precision adaptability data can be used to measure the performance retention of the coding sub-network under different precision bit widths. For example, if the precision of a sub-network drops by less than 1% under 8-bit quantization, it has strong adaptability; if a sub-network shows obvious performance degradation under 4-bit quantization, it has poor adaptability. The modality adaptability score can be a comprehensive scoring indicator that combines modality characteristics, resource consumption, and accuracy capabilities to determine whether a certain encoding sub-network is most suitable for processing the current input modality.

[0055] The target encoding sub-network path can be the one with the highest score, which is selected by the system as the actual execution path for the current task. This path is the most suitable for the modal characteristics of the current input while taking into account performance and resources.

[0056] The modal information of the task input (such as images, text, audio, or a multimodal combination) can be preprocessed and vectorized to obtain a modal vector of uniform dimension. This process can be achieved using a pretrained modal feature extraction network, such as using CNN or ViT to extract feature vectors for images, using BERT or Transformer to extract semantic vectors for text, or by concatenating and mapping multimodal features through a fusion module. Next, from the constructed mixed-precision model parameter data, static or dynamic computational resource consumption data (such as parameter count, FLOPs, memory usage, and latency) for each candidate encoding sub-network path, as well as its accuracy adaptability data (such as post-quantization accuracy change and robustness score) tested under different quantization bit widths, are extracted. This data can be generated through preliminary offline analysis and cached in the model structure description file. A designed scoring function is then used to combine the modal vectors with the resource and accuracy characteristics corresponding to each sub-network to calculate a modal adaptability score for each encoding sub-network path under the current input modality. The scoring function can use a weighted linear function, an MLP scorer, or a reinforcement learning strategy generator to perform a comprehensive evaluation based on the similarity between the modal vector and the sub-network characteristics, resource constraint satisfaction, and accuracy preservation. Ultimately, the encoding path with the highest modality adaptability score is selected as the target encoding sub-network path for the current input task and used for subsequent inference. This ensures inference efficiency while maximizing accuracy and resource matching, achieving adaptive optimization of the model structure under different input modalities.

[0057] In this solution, by intelligently selecting the most appropriate encoding sub-network path based on the modal characteristics of the input task, the model can reduce computing resource consumption while ensuring accuracy, improve operational efficiency and device adaptability, and thus achieve more flexible and efficient multimodal processing capabilities.

[0058] Based on the above technical solution, the optional preset scoring function is:

[0059] in, Score the modality adaptability of the i-th encoding sub-network path; An activation function, such as sigmoid, tanh, or ReLU, is used to introduce nonlinearity to make the scoring function expressive; It is a combination function, and the input features are modal vector, computing resource consumption data, and accuracy adaptability data; is a learnable weight matrix used to linearly transform the joint feature vector , which constitute a dual-channel scoring network, which can be regarded as a "gating mechanism" or "dual-view processing of the scoring function"; The corresponding bias term is used to introduce flexible linear transformation offset; It is an element-by-element multiplication, which is a manifestation of a "gating mechanism" that can fuse the scoring results calculated from two different angles into a more robust output.

[0060] In this solution, computing resource consumption data and precision adaptability data can be converted into vectors. Specifically, the original numerical metrics can be unified. Computing resource consumption data typically includes information such as the number of floating-point operations per encoding subnetwork on the target hardware, the total number of model parameters, video memory usage, inference latency, and energy consumption estimates. After unifying these metrics into the same numerical units, they need to be normalized to facilitate subsequent modeling input. For example, these values ​​are compressed into a uniform range through maximum-minimum normalization or standard deviation normalization. These normalized metrics are then concatenated in a fixed order to form a numerical vector representing the resource overhead of the encoding subnetwork. Meanwhile, precision adaptability data is typically used to measure the performance retention of the encoding subnetwork when using different quantization bit widths. For example, it records the accuracy degradation or performance retention ratio at different precisions, such as 8-bit and 4-bit. By extracting these performance metrics under multiple precision configurations, they can also be organized into numerical vectors in bit width order to reflect the robustness and adaptability of the subnetwork under low-precision conditions. Finally, the resource consumption vector and the accuracy adaptability vector are concatenated with the task-related modality vector to form a unified input feature representation, which is used for fitness evaluation in subsequent scoring functions or sub-network path selection mechanisms. This process ensures that data from different dimensions is standardized and structured, enabling efficient input into neural networks or other predictive models for analysis and selection.

[0061] After the computing resource consumption data and precision adaptability data are converted into vectors, they can be input into , specifically, The form can be:

[0062] in, is the feature vector of computing resource consumption data; is the feature vector of the precision adaptability data; Concat() means concatenating the above three vectors to form a comprehensive feature vector.

[0063] σ() is a nonlinear activation function used to introduce nonlinear capabilities. For example, it can be σ(x)=max(0,x), x or .

[0064] 、 、 、 It is obtained through a training process, usually the training method is as follows: Construct a training set: each sample is (M, C, P) and its true selection or score label; Construct a scoring network (such as the scoring function mentioned above); Use supervised learning, such as cross-entropy loss or ranking loss (RankNet / LambdaRank); Training optimization method: Use gradient descent or its variants (such as Adam) to optimize the parameters 、 、 、 Perform iterative optimization.

[0065] Among them, although the scoring function in this scheme does not use the traditional parallel network to extract heterogeneous modalities, it is designed as a "dual-weight activation channel structure", that is, by applying two affine transformations with different parameter combinations to the same joint feature vector f(M, C, P) ( , and , ), and introduce activation functions σ(·), such as ReLU (σ(x)=max(0,x)), Sigmoid, Tanh, etc., to extract nonlinear feature response signals in two different directions.

[0066] Finally, the two signals are fused via the Hadamard product (element-wise multiplication), resulting in a lightweight equivalent dual-channel scoring structure. This structure can be implemented using two sets of feedforward neural network channels. While different from the modal branches of typical CNN / RNN networks, it effectively expresses path importance and is used to evaluate the suitability of a particular encoding sub-network path under the conditions of input modality characteristics, resource constraints, and accuracy adaptability.

[0067] The joint feature vector f(M, C, P) is the concatenation result of the input modal vector M, the computational resource consumption vector C and the precision adaptability vector P, and its dimensions are 、 and (For example =32, =5, =5), and concatenated to form a fixed-length feature vector for input into the scoring function.

[0068] The input modality vector M is obtained by extracting the intermediate layer representation from the Transformer Encoder. The computational resource consumption vector C includes normalized values ​​such as FLOPs, memory usage, and latency. The precision adaptability vector P represents the performance retention rate (such as the percentage of accuracy drop) of each subnetwork under different quantization bit widths and is also normalized to the range [0, 1].

[0069] The scoring function calculates suitability scores for multiple predefined encoding subnetwork paths (e.g., image-ResNet path, text-BERT path, audio-Wav2Vec path, etc.). For each path, the system constructs a joint feature vector f(M, C, P) based on its corresponding input modality vector, computational resource consumption data, and accuracy adaptability data. This vector is then input into the scoring function to determine the path's suitability for the current task. The system performs this scoring process on all paths sequentially, selecting the highest-scoring path through an argmax operation as the target encoding subnetwork path for subsequent subnetwork activation and fusion operations. The structures of these subnetwork paths are predefined during the model design phase, with their structural parameters, input formats, and interfaces remaining fixed. Therefore, switching paths does not require network reconstruction; instead, it simply activates the selected subnetwork, creating a clear and implementable process.

[0070] The scoring function of the present invention is designed as a dual-weight activation channel structure to evaluate the adaptability of different encoding sub-network paths under the current input modality, computing resource status and accuracy adaptability. This structure applies two sets of affine transformations ( , and , ), and introduces nonlinear responses through the activation function σ(·) to form feature channel responses in both directions. The activation function σ(·) can be selected from ReLU (σ(x)=max(0,x)), Sigmoid, or Tanh to enhance channel representation and nonlinear modeling capabilities. The outputs of the two activation channels are fused via the Hadamard product (element-wise multiplication), retaining only feature regions with high activation in both channels. This achieves a lightweight and efficient path suitability scoring mechanism. This architecture eliminates the need for complex modal fusion modules and is suitable for fast scoring and routing selection for resource-constrained terminals.

[0071] This scoring function structure can be built using two feedforward neural network layers (Linear Layers) and nonlinear activation layers (e.g., ReLU) within mainstream deep learning frameworks (e.g., PyTorch and TensorFlow), ensuring excellent deployment compatibility. The model can be exported to ONNX format and adapted for low-latency deployment using inference engines such as TensorRT and OpenVINO, supporting efficient operation on edge devices and heterogeneous devices.

[0072] During the model training phase, the system predefines a set of paths, P, where each path represents a type of modality encoding subnetwork. For example, image modality paths include ResNet-18 and ResNet-34; text modality paths include BERT-Base and TinyBERT; and audio modality paths include Wav2Vec and VGG-Audio. The structure of each path is determined during the model building phase, and parameters such as computational resource consumption (e.g., FLOPs and memory usage) and accuracy retention are recorded for subsequent scoring.

[0073] During the actual reasoning process, the system calls the scoring function for each path in turn and calculates its adaptability score, and finally selects the path with the highest score as the target encoding sub-network path for the current task. The specific process is as follows: selected_path_output = subnet_outputs[i_star] # i_star = argmax(Score) fused_output = fusion_layer(selected_path_output) The fusion_layer is a pre-defined fusion module (such as linear projection or attention weighting) that aligns paths and converts them to a unified output format, ensuring that the fused features can be used for subsequent decision-making or task processing. The entire path scoring, activation, and fusion process can be completed dynamically at runtime, ensuring clear engineering feasibility and efficient inference deployment.

[0074] S104: Obtain device binding information of the unified coding model and the target device, and perform an encryption binding operation on the unified coding model according to the device binding information to obtain a hardware-bound encryption model.

[0075] The target device refers to the specific hardware environment used to deploy and run the optimized large model. This device can include: terminal devices, servers, and AI acceleration chips.

[0076] Device binding information is a combination of hardware feature data used to uniquely identify the target device and restrict model binding. It serves as the basis for model authorization and secure execution. It can include device hardware fingerprints, such as the motherboard serial number, MAC address or network card ID, TPM chip unique identifier, secure boot key or chip internal public key, deployment timestamp and binding key (such as the session key generated during binding), and the physical / network address of the model execution location.

[0077] An encrypted model refers to a model file or model instance that has been encrypted and bound to a target device. Its characteristics are that it can only be decrypted and run on a specific device (verified through device binding information). The parameters in the model file (such as mixed-precision model parameter data) are encrypted to prevent reverse engineering or model theft. The encryption process can include model parameter encryption (such as AES encryption and homomorphic encryption), hardware-bound key exchange mechanisms (such as TPM-based key derivation), and cryptographic authentication mechanisms (such as runtime identity authentication and signature verification).

[0078] After the unified encoding model is built, the system first executes the device binding information extraction process to ensure that subsequent model execution can only be completed on the specific target device, preventing model leakage and illegal migration. This process reads the target device's underlying hardware information and performs a multi-dimensional summary to generate the device binding information used for encrypted binding. Device binding information includes: the device's motherboard serial number (MotherboardSerialNumber), processor unique identifier (CPUID), MAC address, hard disk device number (DiskSerialNumber), BIOS UUID, TPM module EndorsementKey (EK), and platform configuration register (PCR) value. In addition, the system also integrates software and hardware environment information such as the current operating system kernel version, secure boot policy, and ARM TrustZone security domain identifier to generate a global device feature fingerprint summary F_dev, which is used to uniquely identify the target device.

[0079] After obtaining the device binding information F_dev between the unified encoding model and the target device, the system performs a binding key derivation operation based on the device binding information to initialize the dedicated key resource for model encryption. This step is performed by the system's built-in key management module and specifically includes the following operations: First, the system inputs device binding information (such as the CPU serial number, TPM unique identifier, public key certificate fingerprint, and hardware random number source value) as security factors into the key derivation process. This process is completed within a trusted computing environment. Based on internal security policies, a pre-set perturbation factor and the device's local, non-exportable key material are introduced to generate a binding key K_bind valid only for the target device. This key is not directly exposed to the system externally but is instead managed protectively by the system through a security module (such as a built-in security chip or TPM module). For devices with a trusted execution environment (such as a TPM, TEE, or Secure Enclave), the system directly invokes the platform's built-in key service based on the device binding information to request the generation of a unique public-private key pair from the secure hardware. The generated public key is then used as the encryption medium for model parameters, ensuring that the binding key can only be decrypted and restored by the target device. The entire binding key derivation process strictly records the execution status, timestamp, and hardware signature used, forming key derivation metadata for subsequent model binding verification and runtime validation. Furthermore, the system configures a parameter encryption granularity strategy based on the structural hierarchy and multimodal fusion structure of the unified encoding model. Depending on the device computing power and encryption and decryption load balancing requirements, you can choose to fully encrypt the complete model parameters, or only perform differential encryption on key modules (such as Cross-Attention, RoutingModule, and FusionLayer) to balance security and execution efficiency.

[0080] Based on the device binding information and the derived binding key, the system performs encryption and encapsulation of the uniformly encoded model parameters, ensuring that the model cannot be used outside the target device. Specifically, the system first extracts key module parameters closely related to the selected task path from the current mixed-precision model parameter data (such as encoder.block1 for the vision path, fusion_layer for the cross-modal path, and text_encoder.embed for the text path) as the encrypted encapsulation objects. The system then calls the model encryption module and, based on the derived binding key, performs symmetric encryption on core parameters such as the weight matrix, bias terms, and attention weights of these modules. The encryption algorithm used (such as AES-256 or ChaCha20) converts the parameters into ciphertext during the encapsulation process and generates a separate initialization vector, encryption metadata, and binding tag for each encryption module. During this process, the system embeds the device identity digest (e.g., a hash of the device's unique identifier) ​​and public key signature or certificate chain verification information into the model structure file as part of the model binding identifier based on the device binding information. This information is stored alongside the encrypted model to form a complete hardware-bound data encapsulation structure. To further enhance the immutability of model binding, the system also integrity-signs the encrypted model file using a signature key generated by the current device. It also records the binding relationship between this signature and the device identity for verification during model deployment and runtime. Through these steps, the system ultimately generates a hardware-bound encrypted model that can only be run on the specified target device, effectively preventing the model from being tampered with, leaked, or reconstructed on unauthorized devices.

[0081] S105, obtain input request data and the current running status of the encryption model, perform anomaly detection operations on the encryption model based on the input request data and the current running status, obtain anomaly detection result data, perform protection operations on the encryption model based on the anomaly detection result data, and obtain model security output data.

[0082] Input request data may refer to the original task data and call instruction set initiated by the upper-level application or user end for calling the encryption model to perform reasoning or generate tasks, mainly including task type information; modal input data; model call parameters; user identity identification; and call context data.

[0083] The current operating status of an encrypted model refers to the status information of the model itself and the operating environment that the system continuously monitors and collects during runtime. This information is used to determine in real time whether the model is running stably, securely, and legally. This information primarily includes: model loading status; whether the encrypted portion is successfully decrypted; whether the model decryption timestamp and encryption version number match the binding information; model execution path status; whether the path is consistent with the fusion strategy; hardware execution status; memory / graphics memory usage; encryption and decryption processing unit operating status; security status flags; whether decryption failures, parameter tampering, abnormal call frequency, abnormal modal input, abnormal time window, etc. have been detected; model output health indicators; and whether the model's internal attention weights are abnormally skewed.

[0084] The anomaly detection result data can be the structured detection results output by the anomaly detection module after analyzing the input request data and the current running status of the encryption model. The main contents include: anomaly type label: illegal input (such as unauthorized input mode, incorrect format); execution path abnormality (such as unauthorized path being activated); key matching failure; decryption failure or data integrity check failure; execution timeout or unstable model output; anomaly level score (such as from 0 to 5): indicates the severity of the anomaly; anomaly trigger module identification: indicates which part triggers the anomaly, such as the decryptor, path selector, output checker; anomaly context summary: such as current input summary, anomaly time point, current running device summary, etc.; recommended response strategy: such as retry, discard request, switch backup model, ban source, etc.

[0085] Model security output data may refer to the result data output by the model after it passes the anomaly detection module and takes corresponding protection strategies according to the detection results, including normal output results or security protection responses, mainly including: Model output after security control: If there is no anomaly, the original model reasoning result is output; If there is a medium or low-level anomaly, the output is reviewed or downgraded (such as reducing confidence or fuzzifying part of the output); If it is a high-risk anomaly, a warning response is output; Protection mark: Whether to enable downgraded reasoning; Whether to call a backup model or an alternative path; Whether to add to the call log blacklist; Result credibility score: The credibility score generated by the model based on the current path, input modality integrity, and environmental security; Output encryption information: If the model is configured for secure output mode, the result will be encrypted again with the target device public key and will only be used for decryption by the target system.

[0086] During the actual deployment and calling process of the model, the system first receives input request data from external applications through the model service interface (such as HTTPAPI, gRPC channel or local process call). The input request data contains key fields such as the calling parameters of the current task, input modal content, user identification and context information. Specifically, the system will parse the request message, extract the task type (such as image classification, text generation, speech recognition, cross-modal retrieval, etc.), input modal data (such as image tensor, tokenized text sequence, audio waveform vector, etc.), call context information (such as user historical interaction state cache, task source identifier, request timestamp, user device information, etc.), and encapsulate it into a structured request data object for subsequent processing and analysis.

[0087] While parsing the input request, the system also synchronously obtains the running status information of the currently deployed encryption model. This status information is monitored and recorded in real time by the underlying runtime environment during each model loading, routing, and execution preparation phase. Specifically, it includes the following: 1. Model loading integrity verification results, such as whether the hash checksum of the model weight file is consistent with the initial distribution value, whether the key decryption is successful, and whether the model file is loaded from a trusted directory. 2. Device binding verification status, namely, whether the unique identifier of the currently running device (such as TPMID, MAC address, CPU serial number, etc.) matches the hardware identity in the encrypted model binding information. 3. The identifier of the sub-network path activated within the current model, used to determine whether the model correctly calls the corresponding encoding path based on the input modality and does not illegally bypass or replace the path. 4. Runtime resource status information, such as GPU utilization, video memory usage, CPU load, and I / O bottlenecks. 5. A summary of behavioral characteristics during model execution, including whether the distribution of the attention weight heat map is consistent with the task input content, whether the statistical distribution of the activation layer output has significant deviations, and the frequency of parameter updates. 6. System-level security status flags, such as whether suspicious requests from the same user are continuously detected and whether there are records of repeated key decryption attempts that have failed.

[0088] Combining the aforementioned input request data with the model's current running state, the system performs a series of comprehensive judgment operations based on rules and historical behavior models to assess the legitimacy of the input and the trustworthiness of the execution environment. These judgments include: checking whether the call request originates from an authorized user (for example, by checking the call token against an authorized whitelist); verifying whether the input modality is compatible with the model's currently loaded configuration (for example, whether the input image meets channel and resolution restrictions, or whether the text length exceeds the token limit); counting the call frequency and determining whether it exceeds the system's set threshold (for example, whether the number of consecutive calls within a minute is too high); analyzing whether the model's running behavior is consistent with historical normal distribution (for example, whether the focus area falls within the input-related area, and whether the output confidence is abnormally low or too high); and comparing whether the current execution device is on the trusted device whitelist.

[0089] When the system identifies a potential abnormal situation during the above judgment process, it will generate corresponding abnormal detection result data, record the abnormality type (such as illegal device call, input format not meeting the specifications, abnormal call frequency, large deviation between model behavior and historical distribution, etc.), abnormality level (divided into three levels: low, medium, and high), and attach relevant supporting information, such as the corresponding call log fragment, failed encryption verification code, and the difference measurement value between the model behavior feature vector and the normal vector.

[0090] Depending on the level of the anomaly, the system will take graded protection measures: if it is a low-level anomaly, the system may only add a confidence warning mark to the output, or dynamically adjust the output length and response rhythm; if it is a medium-level anomaly, the system will initiate a restriction strategy, such as reducing the clarity or resolution of the output image in the image generation task, and blocking the content involving sensitive entity vocabulary in the text output task; if it is a high-level anomaly, such as detecting an illegal decryption attempt on the model, hardware replacement in the operating environment, highly abnormal calling behavior, etc., the system will immediately interrupt the model execution process, block any output, and convert the call request into response data with a "security rejection" mark and return it, while fully recording the abnormal information of this call chain.

[0091] Ultimately, after the model is executed, regardless of whether an anomaly is detected, the system generates standardized model security output data. This data includes the actual model output (such as identification labels, generated text, matching results, etc.), the system-calculated confidence score, a status flag indicating whether an anomaly occurred, a description of the handling measures taken, and a security incident number (if any), among other metadata.

[0092] In an embodiment of the present application, resource status data of the target running hardware is obtained, and a model encoding accuracy dynamic adjustment operation is performed according to the resource status data and a preset dynamic adjustment standard to obtain hardware-perceived encoding strategy data; initial weight parameters of each layer of the target large model are obtained, and the target large model is subjected to precision layered re-encoding according to the encoding strategy data and the initial weight parameters to obtain compressed and optimized mixed precision model parameter data; task input modal information is obtained, and a multimodal routing path selection operation is performed on the target large model according to the task input modal information and the mixed precision model parameter data to determine the target encoding sub-network path, and an encoding sub-network selection and fusion operation is performed on the target large model according to the target encoding sub-network path to obtain a fused unified encoding model; device binding information of the unified encoding model and the target device is obtained, and an encryption binding operation is performed on the unified encoding model according to the device binding information to obtain a hardware-bound encryption model; input request data and the current running status of the encryption model are obtained, and an anomaly detection operation of the encryption model is performed according to the input request data and the current running status to obtain anomaly detection result data, and a protection operation is performed on the encryption model according to the anomaly detection result data to obtain model security output data. The aforementioned encryption and protection methods for large deep learning models, combined with resource-aware precision adjustment, multimodal path selection, device-bound encryption, and runtime security protection, enable efficient matching of models and hardware, enhanced task adaptability, and end-to-end security control during use. This not only improves model operational efficiency and accuracy, but also effectively prevents model leakage and illegal calls, ensuring stable and secure deployment in diverse environments.

[0093] Example 2 Figure 2 This is a flow chart of the encryption protection method for the deep learning large model provided in Example 2 of this application. Figure 2 As shown, the specific steps include: S201, obtaining resource status data of target operating hardware, performing a dynamic adjustment operation on model encoding accuracy according to the resource status data and a preset dynamic adjustment standard, and obtaining hardware-perceived encoding strategy data.

[0094] S202, obtaining initial weight parameters of each layer of the target large model, performing precision layered re-encoding on the target large model according to the encoding strategy data and the initial weight parameters, and obtaining compressed and optimized mixed precision model parameter data.

[0095] S203, obtain task input modal information, perform multimodal routing path selection operation on the target large model according to the task input modal information and mixed precision model parameter data, determine the target coding sub-network path, perform coding sub-network selection and fusion operation on the target large model according to the target coding sub-network path, and obtain a fused unified coding model.

[0096] S204, obtaining device binding information between the unified coding model and the target device, determining hardware fingerprint information uniquely corresponding to the target device based on the device binding information, and using the hardware fingerprint information as a key generation factor through a preset encryption function to generate a model encryption key.

[0097] Hardware fingerprint information is a set of identification information extracted from the unique hardware attributes of a device, used to uniquely identify a physical device. It typically includes the CPU serial number, motherboard ID, MAC address, TPM chip ID, storage device serial number, device model, and manufacturer information.

[0098] The pre-defined encryption function can be a designed irreversible or reversible cryptographic function, typically used to derive keys from hardware fingerprint information (e.g., through hashing or a key-derivative function (KDF)), improving security and anti-forgery capabilities. Common forms include hash functions (such as SHA-256), key derivation functions (such as PBKDF2, HKDF, and scrypt), and functions used to extend key lengths in symmetric encryption algorithms.

[0099] The key generation factor can be the raw data used as input for key generation, which is essentially the hardware fingerprint information or its encoded form. In other words, the encryption function uses this "factor" to generate the actual key used for encryption / decryption.

[0100] The model encryption key can be the final key generated, used to encrypt and protect the parameters, weight files or structural information of the machine learning model, ensuring that the model can only be decrypted and run on a specific device, preventing the model from being illegally copied, reverse engineered or stolen on other devices. The key is generated by binding to the hardware, so it is unique and tamper-resistant.

[0101] First, based on the device binding information, the system extracts the underlying hardware parameters that uniquely identify the target device as hardware fingerprint information. This information is unforgeable, tamper-resistant, and strongly bound to the device. This hardware fingerprint information may include, but is not limited to, the CPU serial number, motherboard serial number, hard drive serial number, TPM (Trusted Platform Module) chip ID, and MAC address. All fields are then encoded and concatenated according to pre-set rules. For example, they are uniformly encoded into a UTF-8 string format, with the concatenation order being [CPU_ID|MB_ID|Disk_ID|TPM_ID|MAC] to ensure consistent information across devices. This hardware fingerprint information is then processed using a pre-set encryption function to generate the key generation factor required for the model encryption key. Specifically, this encryption function can be a common secure hash algorithm in cryptography (such as SHA-256 or SHA-3). It hashes the concatenated hardware fingerprint information and outputs a fixed-length pseudo-random bit string. This bit string, known as the key generation factor, possesses high entropy and collision resistance, ensuring consistency across devices and difficulty in collisions between them. Next, using this key generation factor as input, a pre-defined key derivation function (such as PBKDF2, HKDF, bcrypt, or scrypt) is called to generate the final model encryption key. This key is used to encrypt model parameters. Symmetric encryption algorithms (such as AES-256) can be used to encrypt the model's weight files, structure descriptions, or configuration files. The key derivation function can include a salt and iteration count to further enhance security. For example, assuming the hardware fingerprint is: CPU123|MB456|DISK789|TPM321|MACabc, a digest string generated through SHA-256 hashing is used as the key generation factor. PBKDF2 (HMAC-SHA256, iteration count = 10,000, salt = fixed salt) is then used to derive a 256-bit key, which serves as the model encryption key. Using this key to encrypt the model file using AES, even if the model is copied to another device, the decryption key will be inconsistent due to the different hardware fingerprint, making the model unrecoverable. This ensures the model's dedicated operation.

[0102] S205 , encrypting the structural data and parameter data of the unified coding model according to the model encryption key to obtain a hardware-bound encryption model that can only be run in the target device.

[0103] Structural data can be the topological structure definition of the neural network model, that is, the "skeleton" of the model, including the types of layers, the connection methods between layers, the parameter configuration of the layers, and the shapes of the input and output tensors.

[0104] Parameter data can be all weights and bias parameters learned by the model after training, including the convolution kernel parameters of the convolution layer, the weight matrix and bias terms of the fully connected layer, the mean, variance, scaling and offset parameters of the BatchNorm layer, the word vector matrix of the Embedding layer, the linear transformation weights of the Query, Key and Value of the attention mechanism, position encoding, layer normalization parameters, etc.

[0105] To uniquely bind the model to a specific target device, a dedicated model encryption key must first be generated using the target device's hardware fingerprint information. Specifically, based on the preset device binding information, unique hardware fingerprint information is extracted from the target device, including but not limited to the device's CPU serial number, motherboard serial number, TPM module ID, network card MAC address, and storage device unique identifier (such as eMMCCID). This information is concatenated or hashed as input and passed to a preset key derivation function, such as HKDF (HMAC-basedKeyDerivationFunction). After adding a fixed salt value and number of iterations, a model encryption key of 256 bits or another set length is generated. The generated encryption key is not only unique to the current device but also cannot be reproduced on other devices due to different hardware fingerprints. This model encryption key is then used to encrypt the structural data and parameter data of the unified encoding model. Structural data typically includes information describing the model's network topology, such as the layer type (convolutional, fully connected, normalization, etc.), connectivity, and hyperparameters (kernel size, stride, activation function type, etc.). This data can be serialized into standard formats such as JSON, YAML, or protobuf files. Parameter data, on the other hand, consists of numerical tensors such as the trained model's weights, biases, and BatchNorm parameters. These tensors can be saved as the parameter section of binary files such as .pt, .bin, or .onnx files. During the encryption phase, the structural and parameter data are encrypted separately using a symmetric encryption algorithm such as AES-256-GCM or SM4-CBC mode. A separate initialization vector (IV) is generated for each data segment during encryption to ensure that the same plaintext does not appear as the same ciphertext in different encryption steps, enhancing security. The encryption process is as follows: Taking structure data S and parameter data P as input, execute Enc_S = AES_Encrypt(K_model, IV_S, S) and Enc_P = AES_Encrypt(K_model, IV_P, P), respectively. K_model is the model encryption key derived from the hardware fingerprint, and IV_S and IV_P are the initialization vectors for the structure and parameter data, respectively. To ensure data integrity and tamper resistance, GCM encryption mode can be used, with the encryption result including an authentication tag, or the ciphertext can be accompanied by an HMAC checksum. Finally, the encrypted structure and parameter data, along with the encryption algorithm identifier, model version information, and initialization vector, are encapsulated to form a hardware-bound encrypted model file. This model can only be decrypted and loaded on a target device with the same hardware fingerprint by repeatedly deriving the same K_model key.Once the model file is copied to other devices, the generated K_model will be inconsistent due to different hardware fingerprints, which will lead to decryption failure or invalid decryption results, thereby achieving strong binding and security protection of the model, effectively preventing the model from being illegally copied, leaked or tampered with.

[0106] S206, obtain input request data and the current running status of the encryption model, perform anomaly detection operations on the encryption model based on the input request data and the current running status, obtain anomaly detection result data, perform protection operations on the encryption model based on the anomaly detection result data, and obtain model security output data.

[0107] In this embodiment, an exclusive model encryption key is generated from the unique hardware fingerprint information of the target device to achieve encrypted binding of the model structure and parameters, so that the model can only be decrypted and run on the designated device, effectively preventing the model from being illegally copied, stolen, or run on unauthorized devices, thereby improving the security of the model.

[0108] Based on the above technical solution, optionally, after obtaining the model security output data, the method further includes: Determine the running status of the encryption model during the execution of the current task based on the model security output data; If the running state determination result is an abnormal result, the intermediate layer input data, the selected modal coding path data, the structure segment identification data of the encryption model, the model output result data, and the task execution timestamp data corresponding to the current task are obtained, and a hash summary generation operation is performed according to the intermediate layer input data, the selected modal coding path data, the structure segment identification data of the encryption model, the model output result data, and the task execution timestamp data to obtain the behavior summary data; Obtain the context resource identifier of the target device, perform structured aggregation and encoding processing based on the behavior summary data and the context resource identifier, and obtain the execution behavior log; Encrypting the execution behavior log according to a preset encryption function and device binding information to obtain encrypted behavior log data; Write the encrypted behavior log data into the local trusted storage system to obtain log evidence data.

[0109] In this solution, the current task can be a specific task based on model reasoning that the encryption model is currently executing, such as image recognition, speech recognition, multimodal fusion analysis, etc. Different task inputs will also affect the model activation path and reasoning process.

[0110] The operational status determination result is the system's assessment of the model's operational status. This is typically derived from analyzing "model security output data" and is used to determine whether the model is operating as expected. This determination can be categorized as normal (no anomalies) or abnormal (e.g., structural tampering, abnormal output fluctuations, or abnormal activation of model paths).

[0111] When the running status determination result is marked as abnormal, it is an abnormal result. It means that the model behavior or running environment may have been attacked, tampered with, or illegally used, and further traceability and recording are required.

[0112] The intermediate layer input data can refer to the input feature tensor data of a certain intermediate network layer during the inference process of the model, reflecting the internal processing state of the model under the current task input.

[0113] The selected modality encoding path data can be specific encoding sub-network path information activated in the model according to the task modality suitability score, representing the model execution path.

[0114] The structural segment identification data of the encrypted model can be used to uniquely identify a functional segment of the current model structure (such as a certain encoder structure, sub-network module, etc.), making it easier to identify the specific structural version of the running model.

[0115] The model output result data can be the final output result of the model, such as classification results, target detection boxes, regression values, etc., which are used for task result feedback and can also be used to analyze whether there are output anomalies.

[0116] Task execution timestamp data can record the specific time when model inference execution occurs, accurate to milliseconds, ensuring the temporal consistency of behavioral data and serving as the time basis for behavior tracking.

[0117] The behavior summary data can be a summary fingerprint generated by using a hash algorithm to generate all the above-mentioned operation-related key data. It is used to uniquely identify the security summary of a model behavior execution and has tamper-proof capabilities.

[0118] The context resource identifier can represent the resource environment information of the current device during operation, such as CPU / GPU status, system load, network status, process ID, geographic location information, etc., which is used to restore the execution background.

[0119] A derived key can be a key "calculated" from the device binding information. It is a one-time generated key that is unique to the device and cannot be reversibly derived from the binding information.

[0120] The execution behavior log can be a complete behavior record that combines and encodes behavior summary data and context resource identifiers in a structured manner, reflecting the overall picture of a model execution.

[0121] The encrypted behavior log data may be data obtained by encrypting the execution behavior log using a device-bound key to prevent the log from being tampered with or leaked during storage or transmission.

[0122] The local trusted storage system can refer to an encrypted secure storage module in the device specifically used to store sensitive log data, such as a security chip (TPM), TEE (Trusted Execution Environment), a local encrypted database, etc.

[0123] Log evidence data can be an encrypted, tamper-proof log copy that is ultimately written to a local trusted storage system and can be used for subsequent anomaly tracing, compliance auditing, or forensic identification.

[0124] First, the system performs a safety assessment of the model's current reasoning behavior based on the model's safety output data. This data may include the model's output probability distribution, category confidence, output stability indicators, and deviation measures from historical normal distributions. The system then compares this safety output data with pre-set safety thresholds or behavior discrimination models to generate an operational status assessment result corresponding to the current task, identifying potential attacks or model operational anomalies.

[0125] When the running state determination result is determined to be an abnormal result, such as detecting an abnormal deviation in the model output, the system will immediately trigger the behavior recording process to collect key running information of the current task. First, the intermediate layer input data corresponding to the current task is obtained, that is, the input feature representation of the key layer (such as the Bottleneck layer) inside the model during the reasoning process, which is used to restore the internal state characteristics of the model at that time; then the selected modality encoding path data enabled in the current reasoning process is extracted, such as the specific sub-module structure information activated under the image modality, text modality or audio modality; further combined with the structural segment identification data of the encrypted model embedded in the model structure, the identification indicates the structural unit executed by the current reasoning, such as Encoder1, FusionLayer2, etc., and the model output result data of the model's final output is extracted, including but not limited to classification labels, confidence vectors or complete output vector representations; in addition, the system also records the task execution timestamp data of the current task to mark the specific time point when the behavior occurred.

[0126] The system then uses the intermediate layer input data, the selected modal encoding path data, the structural segment identification data of the encryption model, the model output result data and the task execution timestamp data as joint inputs, performs a hash summary generation operation, and calls a secure hash algorithm (such as SHA-256 or the national encryption algorithm SM3) for unified summary processing to generate behavioral summary data that uniquely identifies this reasoning behavior. This data has irreversibility and integrity verification capabilities and can be used for subsequent behavior backtracking and consistency verification.

[0127] After generating the behavior summary data, the system obtains the contextual resource identifiers of the current target device. This identifier includes, but is not limited to, the device ID, processor information, memory status, running process context, geographic location data, network status information, and other data closely related to the device's current operating environment. The system then aggregates and uniformly encodes the behavior summary data and the contextual resource identifiers to generate a standardized execution behavior log with clear fields and traceability.

[0128] To ensure the security and confidentiality of this log, the system invokes a pre-defined encryption function (such as AES, SM4, or RSA) and processes the target device's device binding information. Using this information as input, the encryption function generates a derived key unique to the bound device. The system then encrypts the execution behavior log using this derived key, generating encrypted behavior log data with both confidentiality and integrity protections.

[0129] Finally, the system writes the encrypted behavior log data to a local trusted storage system, such as a secure storage area based on a TEE (Trusted Execution Environment), a TPM (Trusted Platform Module), or a dedicated encrypted file system, generating permanent log evidence data. This log evidence is tamper-resistant, verifiable, and auditable, and can be used for subsequent security audits, anomaly backtracking, or forensics of model execution behavior.

[0130] In this solution, the security of the model itself is protected from the source, while ensuring the transparency and verifiability of the operation process. It is particularly suitable for intelligent system scenarios with high requirements for model security, data compliance, and behavioral auditing.

[0131] Based on the above technical solution, optionally, the encrypted behavior log data is written to a local trusted storage system, including: During the execution of the current task, the encrypted behavior log data is periodically written to the local trusted storage system based on a preset fixed time interval; and / or, Writing encrypted behavior log data to a local trusted storage system when predefined key steps of the current task execution occur; and / or, During the execution of the current task, when it is detected that the task execution deviates from the expected state, the encrypted behavior log data is written to the local trusted storage system.

[0132] In this solution, the preset fixed time interval may refer to the system periodically triggering a behavior log writing operation according to a predefined time period during task execution.

[0133] Predefined key steps can refer to intermediate nodes, decision points, or phased processing links that have been clearly defined as key in the execution process of the current task. Once these steps are executed, log writing is immediately triggered.

[0134] During the current task execution process, in order to ensure the secure storage and traceability of the model behavior log, the system has designed a variety of trigger mechanisms for writing encrypted behavior log data to the local trusted storage system. First, the system can periodically perform log writing operations based on a preset fixed time interval. Specifically, the parameters pre-loaded at the beginning of the task are equipped with a log writing cycle parameter, such as LOG_INTERVAL. The system uses a timer or timestamp polling mechanism. Whenever the interval between the current time and the last write time reaches or exceeds LOG_INTERVAL, the behavior log collection and writing process is triggered, including extracting the current behavior summary data, executing structured coding to generate the execution behavior log, and calling the preset encryption function combined with the derived key to encrypt it. Finally, the encrypted behavior log data is written to the local trusted storage system to form log evidence. Secondly, the system can also trigger log writing operations when predefined key steps in the current task execution occur. These key steps include representative important nodes in the model execution process, such as the activation of a specific modal encoding path, the completion of intermediate layer input data generation, or the generation of model output results. The system detects whether the current execution step matches the preset identifier (such as the structure segment identifier data or the step trigger flag). If the match is successful, the log collection, structuring, encryption and writing process will be executed. In addition, the system also supports triggering a write operation when it detects that the task execution deviates from the expected state. The deviation can be judged based on the model security output data. If abnormal fluctuations in model output, abnormal switching of activation paths, sudden changes in output confidence, etc. are found, the operating status will be judged as abnormal. At this time, the system immediately extracts the intermediate layer input data related to the exception, the selected modal encoding path data, the structural segment identification data of the encryption model, the model output result data, and the task execution timestamp data, performs a hash summary generation operation, and obtains the behavior summary data. Then, combined with the context resource identification of the target device, it performs structured aggregation and encoding processing to form an execution behavior log, and encrypts it with the derived key obtained by processing the device binding information through a preset encryption function to generate encrypted behavior log data, which is finally written to the local trusted storage system to realize encrypted storage of the log.

[0135] In this solution, through the integration of three mechanisms, the system realizes a comprehensive record of the model operation behavior in the dimensions of timing, criticality and abnormality, ensuring the integrity, security and credibility of the behavior log data.

[0136] Example 3 Figure 3This is a schematic diagram of the structure of the encryption protection system for the deep learning large model provided in Example 3 of this application. Figure 3 As shown, specifically including: The coding strategy data determination module 301 is used to obtain resource status data of the target operating hardware, and perform a dynamic adjustment operation on the model coding accuracy according to the resource status data and a preset dynamic adjustment standard to obtain hardware-perceived coding strategy data; The precision layered recoding module 302 is used to obtain initial weight parameters of each layer of the target large model, and perform precision layered recoding on the target large model according to the encoding strategy data and the initial weight parameters to obtain compressed and optimized mixed precision model parameter data; The unified coding model construction module 303 is used to obtain task input modal information, perform a multimodal routing path selection operation on the target large model based on the task input modal information and mixed precision model parameter data, determine the target coding sub-network path, and perform coding sub-network selection and fusion operations on the target large model based on the target coding sub-network path to obtain a fused unified coding model; The encryption model construction module 304 is used to obtain device binding information of the unified coding model and the target device, and perform an encryption binding operation on the unified coding model according to the device binding information to obtain a hardware-bound encryption model; The model security output data determination module 305 is used to obtain input request data and the current running status of the encryption model, perform anomaly detection operations on the encryption model based on the input request data and the current running status, obtain anomaly detection result data, perform protection operations on the encryption model based on the anomaly detection result data, and obtain model security output data.

[0137] The encryption protection system for deep learning large models provided in the embodiment of the present application can achieve Figure 1 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0138] Example 4 like Figure 4 As shown, an embodiment of the present application also provides an electronic device 400, including a processor 401, a memory 402, and a program or instruction stored in the memory 402 and executable on the processor 401. When the program or instruction is executed by the processor 401, each process of the above-mentioned encryption protection method for the deep learning large model is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0139] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0140] Example 5 An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned cable installation process based on the tension adaptive control system embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0141] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0142] It should be noted that, in this article, the terms "comprises", "includes" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the statement "comprises a..." does not exclude the presence of other identical elements in the process, method, article or system comprising the element. In addition, it should be noted that the scope of the methods and systems in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved.

[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of this application.

[0144] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

[0145] The above are only preferred embodiments of the present application and the technical principles employed. The present application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that are possible for those skilled in the art will not depart from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include more other equivalent embodiments without departing from the concept of the present application. The scope of the present application is determined by the scope of the claims.

Claims

1. A method for encryption protection of deep learning large models, characterized in that: The method comprises: Obtain resource status data of the target operating hardware, and perform a dynamic adjustment operation on the model encoding accuracy according to the resource status data and a preset dynamic adjustment standard to obtain hardware-perceived encoding strategy data; Obtaining initial weight parameters of each layer of the target large model, performing precision layered re-encoding on the target large model according to the encoding strategy data and the initial weight parameters, and obtaining compressed and optimized mixed precision model parameter data; Obtaining task input modal information, performing a multimodal routing path selection operation on the target large model based on the task input modal information and mixed precision model parameter data, determining a target encoding subnetwork path, and performing encoding subnetwork selection and fusion operations on the target large model based on the target encoding subnetwork path to obtain a fused unified encoding model; Obtaining device binding information of a unified coding model and a target device, and performing an encryption binding operation on the unified coding model according to the device binding information to obtain a hardware-bound encryption model; Obtain input request data and the current running status of the encryption model, perform anomaly detection operations on the encryption model based on the input request data and the current running status, obtain anomaly detection result data, perform protection operations on the encryption model based on the anomaly detection result data, and obtain model security output data.

2. The method according to claim 1, characterized in that in, The target large model is re-encoded in a precision layered manner according to the encoding strategy data and the initial weight parameters to obtain compressed and optimized mixed precision model parameter data, including: Perform quantitative sensitivity analysis on the initial weight parameters of each layer in the target large model to obtain the accuracy allocation priority of each layer; Determine the precision bit width range supported by the target running hardware, allocate adaptive precision bit width parameters to each layer of the target large model according to the precision bit width range and the precision allocation priority data, perform precision layered recoding, and obtain compressed and optimized mixed precision model parameter data.

3. The method according to claim 1, characterized in that in, Performing a multimodal routing path selection operation on the target large model according to the task input modal information and the mixed precision model parameter data to determine the target encoding sub-network path, including: Vectorizing the task input modal information to obtain a modal vector; Extracting computing resource consumption data and precision adaptability data of multiple encoding sub-networks from the mixed precision model parameter data, and calculating the modal adaptability score of each encoding sub-network path in the target large model based on the modal vector, computing resource consumption data, precision adaptability data, and a preset scoring function; The encoding sub-network path with the highest modality adaptability score is selected as the target encoding sub-network path.

4. The method according to claim 3, characterized in that in, The default scoring function is: in, Score the modality adaptability of the i-th encoding sub-network path; An activation function, such as sigmoid, tanh, or ReLU, is used to introduce nonlinearity to make the scoring function expressive; It is a combination function, and the input features are modal vector, computing resource consumption data, and accuracy adaptability data; is a learnable weight matrix used to linearly transform the joint feature vector , which constitute a dual-channel scoring network, which can be regarded as a "gating mechanism" or "dual-view processing of the scoring function"; The corresponding bias term is used to introduce flexible linear transformation offset; It is an element-by-element multiplication, which is a manifestation of a "gating mechanism" that can fuse the scoring results calculated from two different angles into a more robust output.

5. The method according to claim 1, wherein in, Performing an encryption binding operation on the unified coding model according to the device binding information to obtain a hardware-bound encryption model includes: Determining hardware fingerprint information uniquely corresponding to the target device based on the device binding information, and using the hardware fingerprint information as a key generation factor through a preset encryption function to generate a model encryption key; The structural data and parameter data of the unified coding model are encrypted according to the model encryption key to obtain a hardware-bound encryption model that can only be run in the target device.

6. The method according to claim 5, characterized in that in, After obtaining the model security output data, the method further includes: Determine the running status of the encryption model during the execution of the current task based on the model security output data; If the running state determination result is an abnormal result, the intermediate layer input data, the selected modal coding path data, the structure segment identification data of the encryption model, the model output result data, and the task execution timestamp data corresponding to the current task are obtained, and a hash summary generation operation is performed according to the intermediate layer input data, the selected modal coding path data, the structure segment identification data of the encryption model, the model output result data, and the task execution timestamp data to obtain the behavior summary data; Obtain the context resource identifier of the target device, perform structured aggregation and encoding processing based on the behavior summary data and the context resource identifier, and obtain the execution behavior log; Processing the device binding information according to a preset encryption function to obtain a derived key, and encrypting the execution behavior log according to the derived key to obtain encrypted behavior log data; Write the encrypted behavior log data into the local trusted storage system to obtain log evidence data.

7. The method according to claim 6, characterized in that in, Write encrypted behavior log data to a local trusted storage system, including: During the execution of the current task, the encrypted behavior log data is periodically written to the local trusted storage system based on a preset fixed time interval; and / or, Writing encrypted behavior log data to a local trusted storage system when predefined key steps of the current task execution occur; and / or, During the execution of the current task, when it is detected that the task execution deviates from the expected state, the encrypted behavior log data is written to the local trusted storage system.

8. An encryption protection system for deep learning large models, characterized by: The system comprises: An encoding strategy data determination module is used to obtain resource status data of the target operating hardware, and perform a dynamic adjustment operation on the model encoding accuracy according to the resource status data and a preset dynamic adjustment standard to obtain hardware-perceived encoding strategy data; A precision layered recoding module is used to obtain initial weight parameters of each layer of the target large model, and perform precision layered recoding on the target large model according to the encoding strategy data and the initial weight parameters to obtain compressed and optimized mixed precision model parameter data; A unified coding model construction module is used to obtain task input modal information, perform a multimodal routing path selection operation on the target large model based on the task input modal information and mixed precision model parameter data, determine the target coding subnetwork path, and perform coding subnetwork selection and fusion operations on the target large model based on the target coding subnetwork path to obtain a fused unified coding model; An encryption model construction module is used to obtain device binding information of a unified coding model and a target device, and perform an encryption binding operation on the unified coding model according to the device binding information to obtain a hardware-bound encryption model; The model security output data determination module is used to obtain input request data and the current running status of the encryption model, perform anomaly detection operations on the encryption model based on the input request data and the current running status, obtain anomaly detection result data, perform protection operations on the encryption model based on the anomaly detection result data, and obtain model security output data.

9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the encryption protection method for a deep learning large model as described in any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, which, when executed by a processor, implement the steps of the encryption protection method for a deep learning large model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Dialogue model training method, dialogue generation control method, system and equipment

    CN118296117A

  • End-to-end model reasoning acceleration system

    CN118863071A

  • Lightweight natural language processing large model training method

    CN119862925A

  • Reliable reasoning scheduling method based on edge hybrid expert large model

    CN119903923A

  • Multi-modal model lightweight deployment method based on mixing precision quantification

    CN120124676A

Cited By

  • Target detection method and electronic equipment

    CN121030783A

  • Remote diagnosis and report generation method for multi-modal medical data

    CN121281875A

  • Optical character recognition system and method

    CN121459365A

  • Multi-modal information generation and enhancement method based on AI large model

    CN121543769A