Multi-modal domain adaptation safety evaluation method and system based on efficient parameter fine-tuning
Patent Information
- Application Number
- CN202610640641.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-05-11
AI Technical Summary
1.数据隐私风险:联邦学习虽保护原始数据,但梯度信息仍可被攻击者通过模型反演攻击恢复出93%的原始图像内容;
(1)极致的隐私合规:实现了物理级的“数据不出域”,仅传输不可逆的特征编码器参数,完全消除了敏感数据泄露的风险。
Smart Images

Figure CN122174277B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence security assessment technology, specifically to a multimodal data security assessment method and system for sensitive fields such as finance and healthcare, and is particularly suitable for image and text content security screening scenarios for sensitive facilities and equipment. Background Technology
[0002] In large-scale model security assessments targeting specific fields (such as finance and healthcare), commonly used techniques include: Full-fine-tuning: Training a general-purpose large model with all parameters on domain-specific private data.
[0003] PromptEngineering / RAG: Add domain knowledge or retrieve relevant documents in the Prompt.
[0004] Centralized training: Data from all parties is collected and sent to a central server for unified training.
[0005] The existing technology has the following technical problems: Data privacy breach risk: Data in sensitive areas is often subject to strict "data not leaving the domain" compliance restrictions, making it impossible to upload to the cloud or central server for centralized training.
[0006] Cross-modal semantic gap: Most general security models are trained based on text and have difficulty understanding the visual features of specific domains (such as identifying disguised targets or abnormal medical CT images), leading to missed detections of "inconsistent text and image".
[0007] Huge resource consumption: Full fine-tuning for each domain is too costly in terms of computing power and can easily lead to the model forgetting general security knowledge (catastrophic forgetting).
[0008] The existing technology has the following technical problems: 1. Data privacy risks: Although federated learning protects the original data, gradient information can still be used by attackers to recover 93% of the original image content through model inversion attacks. ; 2. Insufficient cross-modal understanding: In image-text alignment tasks, the traditional RAG method has a key entity recognition accuracy of only 61.3%, which is far lower than the expert level of 92%+; 3. Excessive computing power cost: Full fine-tuning of a 7B parameter scale base model requires 8×A100 (80G) training for one week, costing over 60,000, and resulting in a general security knowledge forgetting rate of over 30%; 4. Difficulty in aligning expert standards: Evaluation reports generated by existing methods lack verifiable structured evidence, and their consistency with the judgments of domain experts is only 58.2%. Summary of the Invention
[0009] To address the problems existing in the prior art, the present invention aims to provide a multimodal domain adaptation security assessment method and system based on efficient parameter fine-tuning. Through an efficient transfer learning framework of "features leaving the domain and data remaining", intra-domain self-supervised feature encoding, cross-attention reinforcement mechanism, efficient parameter fine-tuning and improved GRPO expert alignment, it achieves physical-level "data not leaving the domain", only transmitting irreversible feature encoder parameters, completely eliminating the risk of sensitive data leakage, improving the recall rate in multimodal scenarios and significantly reducing computing power costs.
[0010] To achieve the above objectives, this invention provides a multimodal domain-adaptive security assessment method based on efficient parameter fine-tuning, the method comprising the following steps: S1. Input the data to be identified into the security assessment model; S2. The security assessment model uses data augmentation based on data and task characteristics to augment the input data to be identified. The augmented data is then input into a domain feature encoder to obtain a representation, and the representation is mapped to a contrastive learning space to capture core features. S3. The data output from step S2 is processed by the cross-attention feature alignment module. After parameter fine-tuning, it enters the GRPO reinforcement learning stage to generate a security assessment conclusion. S4. Output a safety assessment conclusion that indicates safety or danger, and output the reasoning process.
[0011] Furthermore, the method can be used for security assessments in the financial or medical fields.
[0012] Furthermore, the data to be identified includes text, images, charts, and / or videos.
[0013] Furthermore, the security assessment model adopts contrastive learning as its core self-supervised training paradigm, which effectively learns the similarities and differences between data instances, thereby capturing the essential feature distribution of the data.
[0014] Furthermore, by using image-text sample pairs or constructing sample pairs of the original image and image segmentation features, two related views are obtained, forming a positive sample pair to achieve data augmentation. All other sample views are regarded as negative samples.
[0015] Furthermore, the enhanced sample is input into the domain feature encoder to obtain a representation; the domain feature encoder is followed by a projection head to map the representation to the contrastive learning space.
[0016] Furthermore, a cross-attention feature alignment module consisting of N identical stacked cross-attention layers is constructed and inserted into the input layer or intermediate layer of the evaluation base model to achieve cross-modal domain feature alignment.
[0017] Furthermore, during alignment training, all parameters except for the cross-attention feature alignment module are frozen.
[0018] Furthermore, an improved GRPO phase is introduced, incorporating the evaluation criteria of domain experts into the security assessment model through reinforcement learning. The composite reward function of the GRPO phase is designed to include a reward for the correctness of the result, a reward for the importance of structured evidence, and a KL divergence penalty term.
[0019] On the other hand, the present invention provides a multimodal domain adaptation safety assessment system based on efficient parameter fine-tuning, the system being used to implement the method described in the present invention.
[0020] The beneficial effects of this invention are as follows: (1) Extreme privacy compliance: It achieves physical-level "data not leaving the domain", only transmitting irreversible feature encoder parameters, completely eliminating the risk of sensitive data leakage.
[0021] (2) Cross-modal accurate evaluation: effectively solves the problem that general models "cannot understand" professional domain images (such as accurately identifying equipment in the image and evaluating its sensitivity), and improves the recall rate in multimodal scenarios.
[0022] (3) Low cost and expert-level alignment: Compared with full fine-tuning, the number of training parameters is reduced by more than 99%; and through GRPO reinforcement learning, the automated evaluation results are closer to human domain experts in terms of logic and evidence. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the multimodal domain adaptation security assessment method and system architecture based on efficient parameter fine-tuning according to the present invention. Detailed Implementation
[0024] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0026] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0027] The following combination Figure 1 Specific embodiments of the present invention will be described in detail below. It should be understood that the specific embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the present invention.
[0028] The multimodal domain-adaptive security assessment method and system based on efficient parameter fine-tuning according to the present invention are suitable for large-scale model security assessments for specific domains (such as finance and healthcare). The security assessment method and system of the present invention can be used to identify camouflaged targets in data such as text, images, or videos.
[0029] The multimodal domain adaptation security assessment method based on efficient parameter fine-tuning according to the present invention includes the following steps: (1) System construction: Deploy security assessment models, including domain feature encoders deployed within the security domain and general security assessment base models deployed outside the domain; (2) Intradomain feature learning: In the safe domain, data augmentation oriented to the task characteristics is performed on the data to be identified. The data to be identified is the original multimodal data. The domain feature encoder is trained by contrastive learning so that the domain feature encoder maps the original multimodal data into semantic feature vectors. Only the parameters of the domain feature encoder are passed out of the safe domain. (3) Feature alignment and fine-tuning: Outside the domain, the general security assessment base model is cascaded with the domain feature encoder, a cross-attention feature alignment module is inserted, the base model and encoder are frozen, and only the parameters of the cross-attention feature alignment module are fine-tuned to achieve cross-modal feature fusion; (4) Expert standard alignment: An improved group relative strategy is introduced to optimize the GRPO algorithm. A composite reward function is designed, which includes a result correctness reward, a structured evidence overlap rate reward, and a KL divergence penalty term, so that the model evaluation standard is consistent with the domain expert judgment standard and outputs safe or dangerous conclusions and reasoning basis.
[0030] Specifically, in the domain feature learning stage, the data to be identified is input into the security assessment model; the data to be identified includes text, images, charts, and / or videos; in the relevant domain assessment, the aim is to identify keywords or image regions in the data that involve equipment, deployment, sensitive facilities, etc. The security assessment model of this invention is an integrated system comprising five functional modules: 1) Domain feature encoder, responsible for extracting domain-specific semantic features from raw data; 2) The cross-attention feature alignment module is responsible for the deep fusion of information from different modalities (such as image features and text features); 3) The GRPO reinforcement learning decision module is responsible for generating the final security assessment conclusion and reasoning process based on the fused features; 4) Data preprocessing module, responsible for standardizing and enhancing the raw data; 5) Safety assessment result generation and visualization module, responsible for outputting and displaying safety or danger conclusions and reasoning basis.
[0031] The security assessment model employs data augmentation tailored to the data and task characteristics of the input actual feature data. The augmented data is then input into a domain feature encoder for representation, and the representation is mapped to a contrastive learning space to capture core features. The entire data processing flow can be divided into four distinct stages: The first stage is "data preprocessing and augmentation." The input raw data (such as an image of equipment with text descriptions) is first standardized by the system (standardization includes size normalization and text segmentation). Then, task-oriented data augmentation strategies are applied, such as random cropping, rotation, and color dithering of the image, and synonym replacement or noise insertion of the text, thereby generating two views. and .
[0032] The second stage is "feature encoding," where the two generated views are input into the domain feature encoder. Two high-dimensional feature vectors are obtained. and .
[0033] The third stage is "contrastive learning and representation mapping," which inputs these two feature vectors into the projection head. Mapped to the contrastive learning space, we obtain and That is, the normalized representation vector obtained by mapping the two views of the i-th sample to the contrast learning space through the projection head after data augmentation, where the subscripts 2i−1 and 2i correspond to the projection representation of the first augmented view and the second augmented view of the sample, respectively. The model is self-supervised by calculating the NT-Xent loss function so that it can learn to distinguish between positive and negative sample pairs.
[0034] The fourth stage is "cross-modal alignment," which involves aligning the image features processed by the domain feature encoder. Initial representation of the text to be evaluated The input is fed into the cross-attention feature alignment module for interaction, and the output is a deep representation that integrates image and text information.
[0035] The output data is processed by the cross-attention feature alignment module, and after parameter fine-tuning, it enters the GRPO reinforcement learning stage to generate a security assessment conclusion. The system outputs a safety assessment conclusion, indicating whether the system is safe or hazardous, and provides the reasoning process. In the final output stage, the system generates a structured report. This structured report consists of three parts: The first part is the "assessment conclusion," such as "dangerous" or "safe." The second part is the "reasoning process," which details the key evidence supporting the conclusion, such as "sensitive keywords were detected in the text" or "the area from (120, 350) to (480, 620) pixels in the image was identified as a suspected sensitive facility." The third part is the "Confidence Score," which gives the percentage of confidence in the conclusion. The output can be an API response in JSON format or a visual HTML page with highlighted markers and checkboxes for easy user understanding.
[0036] In one specific embodiment, the processing flow for a dataset containing the text "New remote equipment deployed in area A" and a related satellite image is as follows: 1) The data preprocessing module receives raw data, which includes text and images. It performs word segmentation and vectorization on the text and cropping and normalization on the images.
[0037] 2) The domain feature encoder module processes both text and images simultaneously, outputting their respective preliminary features, which include text features and image features.
[0038] 3) The cross-attention feature alignment module uses text features as queries and image features as key-value pairs to perform bidirectional attention calculations and generate a fused feature.
[0039] 4) The GRPO reinforcement learning decision module receives fused features, makes decisions based on preset reward functions (such as correctness, evidence overlap rate, and KL divergence), and finally outputs a conclusion and reasoning process. For example, the output conclusion is "dangerous," and the reasoning process is: "'remote device' and 'deployment' are high-risk keywords in the text, and the upper left corner of the image is identified as a device launch vehicle."
[0040] The key technical focus of this invention lies in the collaborative operation of the "domain feature encoder" and the "cross-attention feature alignment module." The domain feature encoder uses ResNet-50 as its backbone network. Its innovation lies in performing self-supervised pre-training only within the safety domain. After training, the projection head is discarded, retaining only the encoder backbone. This ensures that the model parameters are highly abstract feature representations, making it impossible to reconstruct the original data, thus blocking the risk of data leakage at the source. The cross-attention feature alignment module is another major innovation of this invention. It consists of N=4 layers of identical cross-attention blocks stacked together, each block containing two sub-layers: First is the CrossAttn_T->D layer, which is used to inject text features into image features; The next layer is the CrossAttn_D->T layer, which is used to inject image features back into text features.
[0041] This two-way interaction mechanism enables the model not only to understand the semantics of text but also to accurately locate sensitive regions in images, thereby effectively improving image recognition accuracy. Experiments show that the method of this invention can increase image recognition accuracy from 78% to 92% compared to the baseline model. Furthermore, by freezing most parameters of the general security assessment base model and only fine-tuning the parameters of the four cross-attention layers in the cross-attention feature alignment module, training costs can be effectively reduced, achieving a low-cost adaptation with a reduction of approximately 70%.
[0042] This embodiment employs a phased training and deployment architecture. The training and deployment architecture of this invention is divided into three phases: The first stage is "domain feature encoder pre-training," which utilizes a large amount of unlabeled domain data within the safe domain to train the domain feature encoder through contrastive learning (NT-Xent loss). The goal is to learn domain-independent, robust core features.
[0043] The second stage is "cross-modal feature alignment fine-tuning". Based on the general security assessment base model, a cross-attention feature alignment module is inserted, and the main parameters of the general security assessment base model are frozen. Only the newly inserted cross-attention feature alignment module is fine-tuned, and a small amount of labeled data is used for fine-tuning to achieve cross-modal deep alignment.
[0044] The third stage is "GRPO reinforcement learning optimization", which introduces an improved group relative policy optimization algorithm and designs a composite reward function that includes result correctness, structured evidence overlap rate and KL divergence penalty. The model is then subjected to reinforcement learning so that its output is not only correct, but also interpretable and consistent with experts.
[0045] The security assessment model used in this method employs a self-supervised training paradigm centered on contrastive learning, as it effectively learns the similarities and differences between data instances, thereby capturing the essential characteristic distribution of the data. Given a batch of domain data... For each sample (It can be text, image, or other). This invention applies data augmentation methods oriented towards data and task characteristics, such as using image-text sample pairs or constructing sample pairs of the original image and image segmentation features to obtain two related views. and These constitute a positive sample pair. Views of all other samples in the batch are considered negative samples.
[0046] The specific implementation methods of data augmentation are as follows: 1) For image-text sample pairs, the system will randomly select an image and a related text description, and then apply transformations such as random cropping, rotation, and color dithering to the image, and apply random deletion or synonym replacement to the text, thereby generating two different views.
[0047] 2) For the original image and image segmentation feature sample pair, the system first performs semantic segmentation on the original image to obtain masks for the foreground object and the background. Then, it separates the foreground object from the background and enhances the foreground object, thereby constructing an "original image" view and a "segmentation feature" view. In the task, "task-oriented characteristics" mean that the enhancement strategy will be biased towards preserving or highlighting relevant key features. For example, when enhancing the main part, it will avoid cropping it to ensure that the model can learn its key recognition features.
[0048] Subsequently, the enhanced sample is input into the neighborhood feature encoder. (i.e., the domain feature encoder to be trained in this invention), to obtain the representation A projection head is connected after the domain feature encoder. The representation is mapped to the contrastive learning space to obtain a normalized vector. Here, Representing the One original sample, Represents the data augmentation of the first One view, It is a domain feature encoder, which is a deep neural network whose function is to map the raw input data (whether text or image) into a fixed-dimensional feature vector. This feature vector It contains the most core and essential semantic information of the data, while ignoring irrelevant noise and details. It is a simple linear projection layer, and its function is to project high-dimensional features. Mapping to a lower-dimensional contrastive learning space, we obtain . It is a normalized vector with a length of 1, which facilitates the subsequent calculation of cosine similarity. In this way, the model can compare the similarity between different samples in the contrastive learning space, thereby learning meaningful feature representations.
[0049] The contrastive loss function (NT-Xent loss) is defined as follows:
[0050] in, It is cosine similarity. It's a temperature over-parameter. It is an indicator function. The final loss is calculated for all positive sample pairs. of and Calculate the average. The principle behind the NT-Xent loss function is: for a positive sample pair... Calculate their similarity in the contrastive learning space. Then divide by the temperature parameter After scaling and normalization using an exponential function and Softmax, a probability value is obtained. The numerator is the probability of a positive sample pair, and the denominator is the sum of the probabilities of all negative sample pairs. By minimizing this loss, the model is forced to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. In the formula, This indicates the total number of samples in the current batch. and Representing the first The and the first The representation vector of each sample in the contrastive learning space This represents the characteristic of any other sample in the batch. Temperature hyperparameter. In this embodiment, it is set to 0.1, which is an empirical value used to control the distribution width of the similarity score. A smaller value results in a lower similarity score. This will make the model more sensitive to differences in similarity.
[0051] By minimizing this loss, the domain feature encoder It can learn to ignore irrelevant augmenting noise and capture the invariant, semantically core features of the data. After training, the projection head is discarded. Only retain the domain feature encoder This domain feature encoder is now able to process any domain data. Mapped to a feature vector with high semantic information density This domain-specific encoder is the only carrier allowed to "leave the domain," thus ensuring that only abstracted model parameters, which cannot be reconstructed from the original data, are allowed to flow out of the security domain, completely blocking the data leakage path.
[0052] After obtaining the domain feature encoder, the challenge lies in how to enable the text-based general security assessment base model to "understand" these domain-specific feature vectors from another modality. This invention designs a cross-attention mechanism based on the DeepStack concept to achieve deep modal interaction and information reinforcement. The core idea of DeepStack is to achieve feature abstraction and fusion from shallow to deep layers by stacking multiple processing blocks. This invention innovatively applies it to cross-modal domain feature alignment. The general security assessment base model refers to a pre-trained, general-purpose large-scale language model (such as GPT or LLaMA), which itself does not possess domain knowledge. The "security assessment model" of this invention is built on this general security assessment base model by inserting a "domain feature encoder module" and a "cross-attention alignment module". Therefore, the "general security assessment base model" is a part of the "security assessment model" and is its basic architecture for processing text.
[0053] This mechanism is inserted as an independent "reinforcement module" into the input or intermediate layers of the general security assessment base model. The general security assessment base model of this invention is a typical Transformer architecture, containing L=12 layers. The first 6 layers are mainly used for word embedding and local semantic understanding of the text, while the last 6 layers are used for global semantic modeling and context dependency capture. The "cross-attention alignment module" of this invention is inserted between the 6th and 7th layers, i.e., the "intermediate layer." The advantage of this is that, after the text features have undergone preliminary processing and possess certain semantic information, domain features from the image can be injected, thereby achieving deeper and more effective information fusion. If inserted into the input layer, image features might be overwhelmed by the word embedding layer of the text; if inserted into the last layer, the fused information might not have enough time to influence the final decision.
[0054] First, the text to be evaluated is passed through the word embedding layer and (optionally) the first few Transformer layers of the general security evaluation base model to obtain a context-aware text representation consistent with this. ,in ,get To enable it to interact with sequences, it is treated as a special "domain feature token" and copied. This constitutes a sequence. The word embedding layer of the general security assessment foundation model processes the raw text. (Formula) middle, It is the image feature vector output from the domain feature encoder. and It is a learnable linear transformation matrix and bias vector, whose function is to transform image features. The dimension from Mapped to the same dimension as text features This is to facilitate subsequent interactions. It is the mapped feature, which is copied L times (L is usually the sequence length, such as 128), forming a shape of matrix This matrix It's like a "token sequence," where each token contains the same domain feature information and can be used for attention computation with each position in the text sequence. For example, if the text sequence is "This is a car," then the four words "this," "is," "a," and "car" in the sequence will each interact with this domain feature token, allowing the model to know that the word "car" is related to sensitive features in the image.
[0055] This invention constructs a system composed of This cross-attention feature alignment module is composed of N=4 identical stacked sub-modules, each containing four sub-layers. Its functions and processing are as follows: 1) "Text-to-Domain Cross-Attention Layer (CrossAttn_T->D)": This layer converts text sequences into domain-specific attention layers. As a key-value pair, the domain feature sequence As a query, attention weights are calculated, textual information is injected into the domain features, and the updated domain features are output. .
[0056] 2) "Domain-to-Text Cross-Attention Layer (CrossAttn_D->T)": This layer updates the domain features... As a key, the text sequence As a query, attention weights are calculated, domain information is injected into the text features, and the updated text features are output. .
[0057] 3) "Domain Feature Feedforward Network Layer (FFN_D)": This layer processes the updated domain features... A two-layer feedforward neural network is applied to perform nonlinear transformation and feature enhancement.
[0058] 4) "Text Feature Feedforward Network Layer (FFN_T)": This layer processes the updated text features... A two-layer feedforward neural network is applied for nonlinear transformation and feature enhancement. Through the alternating processing of these four layers, text and domain features can be deeply interacted and fused at each layer.
[0059] During alignment training, all parameters except for the cross-attention feature alignment module are frozen. The parameters frozen during the alignment training phase include: 1) Domain Feature Encoder All parameters of the system do not need to be updated since they have already been pre-trained within the security domain. 2) The vast majority (over 90%, specific values can be corrected) of the parameters of the general security assessment base model, including the word embedding layer, and most (over 80%, specific values can be corrected) of the weights of the first 6 layers and the last 6 layers of the Transformer.
[0060] The parameters that are being fine-tuned are only: 1) Linear transformation layer used to map image features to text dimensions ,in The linear projection weight matrix is learnable. The corresponding bias vectors are used to map the image features output by the domain feature encoder from dimension d to the text feature dimension d. t ; 2) All parameters in the newly inserted N=4 layer cross-attention feature alignment module, including attention weight matrix, feedforward network weights, etc.
[0061] After efficient parameter fine-tuning, the security assessment model of this invention has acquired preliminary domain-aware assessment capabilities. The "efficient parameter fine-tuning" of this invention refers to a transfer learning strategy, the core of which is to freeze most of the parameters of a large pre-trained model (i.e., a general security assessment base model), and only unfreeze and update a small portion of newly added or task-related parameters.
[0062] In this invention, the specific parameters for fine-tuning include: 1) Linear layers for feature dimension alignment ; 2) All parameters in the newly inserted N=4 layer cross-attention feature alignment module.
[0063] The "efficiency" of fine-tuning is reflected in two aspects: first, the number of parameters is extremely small, accounting for less than 5% of the total number of parameters in the entire model; second, the adjustment range is small, usually using a small learning rate (such as 1e-5), making only minor adjustments to the parameters, rather than training from scratch. This strategy ensures the model's generalization ability while significantly reducing computational costs and the risk of overfitting.
[0064] However, the "style" and "depth" of their evaluations may still fall short of those of true domain experts. Expert evaluation goes beyond simply providing a "safe / dangerous" conclusion; more importantly, it offers a logically rigorous, evidence-based reasoning process that conforms to domain norms. To simulate this advanced capability, this invention introduces an improved GRPO (Group Relative Policy Optimization) phase, injecting domain expert evaluation criteria into the security assessment model through reinforcement learning. In terms of data flow, the data follows this sequence: first, the raw data is processed by a domain feature encoder; second, it undergoes cross-modal fusion via a cross-attention feature alignment module; finally, the fused features are input into the GRPO reinforcement learning decision module for the final decision.
[0065] This invention utilizes the composite reward function of the GRPO reinforcement learning decision module. The design comprises three key components: a reward for correctness of results, a reward for the importance of structured evidence, and a KL divergence penalty term. These components work together to guide the model in generating outputs that are correct, interpretable, and highly consistent with expert judgment. (Composite reward function) The specific formula is as follows: ; in, Rewards are given for correct results. As a reward for the importance rate of structured evidence, This is a KL divergence penalty term. The correctness of the result is rewarded. It focuses solely on whether the final conclusion of the evaluation is correct, without directly scoring the content of the thought process itself; it rewards the importance rate of structured evidence. In many domain evaluations (such as detecting illegal objects in images or locating sensitive entities in text), the model's output needs to include verifiable structured evidence, such as IoU or F1 scores. This reward achieves fine-tuning by aligning with these structured outputs. Finally, to prevent the policy model from deviating excessively from its initial, supervised fine-tuning "good" behavior during reinforcement learning (e.g., starting to generate gibberish to cheat for high rewards), this invention introduces a KL divergence penalty term.
[0066]
[0067] For the same input generated Output This invention calculates the composite reward for each output. Then, calculate the average reward for the group. and standard deviation The standardized advantage function for each output. The calculation is as follows:
[0068] Among them, the composite reward function : - For a given input Single output generated The total reward received. The weighting coefficient for the reward based on the correctness of the result is used to balance the importance of different reward items. If the output If the final conclusion matches the true label, the value is 1; otherwise, it is 0. The weighting coefficient for the importance rate reward of structured evidence. : Calculation output The extracted structured evidence (such as bounding boxes and keyword lists) and the reference answer. The overlap rate, for example, using F1 score or IoU. : Weight coefficient of the KL divergence penalty term. KL divergence penalty term, used to measure the current policy With initial supervision and fine-tuning strategy The differences between them. : The intensity coefficient of the KL divergence penalty, used to control the penalty strength. The current policy under given input Output generated below The probability of. The initial supervised fine-tuning policy is based on a given input. Output generated below The probability of. : No. Output Compound rewards. Same group The average reward for each output. Same group The standard deviation of the reward for each output. A very small constant (e.g., 1e-8) used to prevent division by zero errors.
[0069] By optimizing this objective, the security assessment model is driven to generate outputs that offer higher rewards compared to other models within the same group. Specifically, this means the final assessment conclusion is correct, the structured evidence highly overlaps with expert annotations, and the assessment result does not excessively deviate from its initial security behavior. "Structured evidence" refers to specific, quantifiable information provided by the model in its output that can be manually verified. For example, in image assessment, this could be bounding box coordinates and category labels; in text assessment, it could be a list containing sensitive keywords and their locations. "Initial security behavior" refers to the behavioral pattern exhibited by the security assessment model after supervised fine-tuning; that is, without reinforcement learning intervention, it tends to generate outputs that conform to security standards, are logically rigorous, and are appropriately worded. The goal of the GRPO reinforcement learning decision module is to ensure that the model, while pursuing high rewards, does not deviate from this fundamental "security" behavior and avoids generating harmful or uncontrollable content.
[0070] This invention features the following technological innovations: Improvement 1: Intra-domain self-supervised feature encoding Technical principle: A domain feature encoder is deployed within the security domain, utilizing contrastive learning to mine data features. Only abstracted model parameters, which cannot be reconstructed from the original data, are allowed to leave the security domain, completely blocking data leakage paths.
[0071] Improvement 2: Cross-Attention Reinforcement Mechanism Technical principle: Insert a cross-attention layer similar to DeepStack into a general model outside the domain. Use the encoder trained in the domain as an "external knowledge base" and inject its output feature vectors into the text representation of the general model to achieve cross-modal deep alignment.
[0072] Improvement 3: Efficient parameter fine-tuning and improved GRPO expert alignment Technical principle: P Efficient Fine-tuning: Freeze the general model and fine-tune only the cross-attention layer to achieve low-cost adaptation.
[0073] GRPO Optimization: After fine-tuning, an improved Group Relative Policy Optimization (GRPO) is introduced. By designing a composite reward function that includes "outcome correctness", "structured evidence overlap rate" and "KL divergence", the evaluation criteria of the model are deeply aligned with those of domain experts.
[0074] The comparison schemes are as follows: Federated Learning involves joint training by multiple parties and the transmission of gradients. Its drawbacks include significant communication overhead and the risk of gradient leakage, where attackers inferring the original data. Federated Learning can serve as a comparative solution to replace the "domain feature encoder pre-training" stage in this invention. Under the federated learning framework, multiple participants can train their own feature encoders locally and then only upload the gradients of the model parameters or the aggregated model, thus avoiding direct transmission of the original data. However, its technical effectiveness is not entirely the same as this invention: 1) Federated Learning cannot guarantee that the model parameters are "unrecoverable from the original data" because the gradients themselves may carry sensitive information; 2) Federated Learning has significant communication overhead, making it unsuitable for scenarios with high real-time requirements; 3) This invention learns domain-independent general features through self-supervised pre-training, while Federated Learning focuses more on protecting the data privacy of each participant.
[0075] Retrieval Enhancement Generation (RAG): This involves building a vector library within the domain and enhancing the context through retrieval. A drawback is that RAG is primarily used for knowledge supplementation and struggles to achieve deep semantic understanding across modalities (such as understanding implicit risks in images). RAG can serve as a contrasting alternative to the "cross-attention feature alignment module" in this invention. Under the RAG framework, the system first encodes and stores image features in a vector database. When evaluating a piece of text, the system retrieves the most relevant image features and injects them as context into the evaluation base model. However, its technical effectiveness is not entirely the same as this invention: 1) RAG is a "retrieval-injection" mechanism, lacking deep, end-to-end feature fusion capabilities and unable to achieve bidirectional information flow like cross-attention; 2) RAG's understanding of images is static and retrieval-based, while the cross-attention module of this invention can dynamically and adaptively fuse image features with text features, thus better understanding implicit risks in images; 3) RAG's performance is highly dependent on the quality of the vector database and the efficiency of the retrieval algorithm.
[0076] Standard LoRA fine-tuning: Only adds a low-rank matrix to the Attention layer. This application's scheme is a variant of LoRA, but it specifically modifies the network structure for cross-modal feature fusion (introducing cross-attention) and adds RLIHF (GRPO) as an evaluation metric, resulting in better performance than standard LoRA. Standard LoRA fine-tuning can be used as a comparative solution to replace the fine-tuning part of the "cross-attention feature alignment module" in this invention. Standard LoRA achieves efficient parameter fine-tuning by inserting a low-rank matrix into the attention layer of the Transformer. However, its technical effect is not entirely the same as this invention: 1) Standard LoRA only fine-tunes the attention mechanism within the model, while the cross-attention module of this invention is a completely new and independent module specifically designed to handle cross-modal information; 2) Standard LoRA does not introduce an external knowledge source, the "domain feature encoder," and therefore cannot effectively inject image features into the text model as this invention; 3) Standard LoRA does not combine GRPO reinforcement learning, and therefore cannot generate outputs that are both correct and interpretable, and highly consistent with expert judgment, as this invention.
[0077] Any process or method described in the flowcharts of this invention or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, which can be implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device. The computer-readable medium can be any medium containing a program for storage, communication, propagation, or transmission for use by the execution system, apparatus, or device, including read-only memory, magnetic disks, or optical disks.
[0078] In the description of this specification, references to terms such as "embodiment," "example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, those skilled in the art can combine or combine the different embodiments or examples described in this specification and the features therein without causing contradiction.
[0079] While embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and alterations to the above embodiments within the scope of the present invention.
Claims
1. A multimodal domain-adaptive safety assessment method based on efficient parameter fine-tuning, characterized in that, The method includes the following steps: (1) System construction: Deploy security assessment models, including domain feature encoders deployed within the security domain and general security assessment base models deployed outside the domain; (2) Intradomain feature learning: In the safe domain, data augmentation oriented to the task characteristics is performed on the data to be identified. The data to be identified is the original multimodal data. The domain feature encoder is trained by contrastive learning so that the domain feature encoder maps the original multimodal data into semantic feature vectors. Only the parameters of the domain feature encoder are passed out of the safe domain. (3) Feature alignment and fine-tuning: Outside the domain, the general security assessment base model is cascaded with the domain feature encoder, a cross-attention feature alignment module is inserted, the base model and encoder are frozen, and only the parameters of the cross-attention feature alignment module are fine-tuned to achieve cross-modal feature fusion; (4) Expert standard alignment: An improved group relative strategy optimization GRPO algorithm is introduced, and a composite reward function is designed, which includes result correctness reward, structured evidence overlap rate reward and KL divergence penalty term, so that the model evaluation standard is consistent with the domain expert judgment standard, and a structured report is output, which includes evaluation conclusions and reasoning basis. In step (3), the cross-attention feature alignment module is composed of N=4 identical sub-modules stacked together, with each layer containing four sub-layers. Its functions and processing are as follows: 1) "Text-to-Domain Cross-Attention Layer CrossAttn_T->D": This layer converts text sequences into domain-specific attention layers. As a key-value pair, the domain feature sequence As a query, attention weights are calculated, textual information is injected into the domain features, and the updated domain features are output. ; 2) "Domain-to-Text CrossAttn_D->T": This layer updates the domain features. As a key, the text sequence As a query, attention weights are calculated, domain information is injected into the text features, and the updated text features are output. ; 3) "Domain Feature Feedforward Network Layer FFN_D": This layer updates the domain features. A two-layer feedforward neural network is applied to perform nonlinear transformation and feature enhancement; 4) "Text Feature Feedforward Network Layer FFN_T": This layer processes the updated text features. A two-layer feedforward neural network is applied to perform nonlinear transformation and feature enhancement; Through these four layers of alternating processing, text and domain features can be deeply interacted and integrated at each layer.
2. The method according to claim 1, characterized in that, In step (2), data augmentation includes: For text-image pairs, construct the original image and its segmentation features as positive sample pairs; For pure image data, construct the original image and the image after rotation, cropping, and color transformation as positive sample pairs; All unpaired samples are considered negative samples; the contrastive loss function used is NT-Xent loss. ; in, and Representing the first The and the first The representation vector of each sample in the contrastive learning space This represents the representation of any other sample in the batch. Here, is the temperature hyperparameter, and sim is the cosine similarity. This indicates the total number of samples in the current batch.
3. The method according to claim 1, characterized in that, In step (4), the composite reward function is defined as: ; in, Rewards are given for correct results. As a reward for the importance rate of structured evidence, This is a KL divergence penalty term; The weighting coefficient for the reward based on the correctness of the result. The weighting coefficient for the structured evidence importance rate reward. The weighting coefficients for the KL divergence penalty term are... For input, For a given input A single output is generated.
4. The method according to claim 1, characterized in that, The domain feature encoder adopts the ViT-Base architecture with 86M parameters; the general security assessment base model adopts the LLaMA-7B architecture with 7B parameters; the cross-attention feature alignment module only needs to fine-tune 0.85% of its parameters to reach the expert-level assessment level.
5. The method according to claim 1, characterized in that, In step (3), during alignment training, all parameters except for the cross-attention feature alignment module are frozen.
6. The method according to claim 5, characterized in that, For images containing sensitive facilities, the security assessment process includes: Identify sensitive regions in an image and output bounding box coordinates; Analyze the similarity between this area and known facilities; Determine the sensitivity of information by combining the context text; Generate a structured report, including the location, category, and confidence level of sensitive areas. This structured report constitutes the specific manifestation of the reasoning basis in step (4) and outputs a safety or danger assessment conclusion.
7. The method according to claim 1, characterized in that, The data to be identified includes text, images, charts, and / or videos.
8. The method according to claim 2, characterized in that, In step (2), the enhanced sample is input into the domain feature encoder to obtain the representation; the domain feature encoder is followed by a projection head to map the representation to the contrastive learning space.
9. A multimodal domain-adaptive safety assessment system based on efficient parameter fine-tuning, characterized in that, The system is used to implement the method according to any one of claims 1-8, the system comprising: A domain feature encoding module deployed within the security domain is used to extract domain-specific semantic features from the raw data; An off-domain deployed cross-attention feature alignment module is used to deeply fuse information from different modalities; The GRPO reinforcement learning decision module is used to generate the final security assessment conclusion and reasoning process based on the fused features. The safety assessment results generation and visualization module is used to output and display safety or danger conclusions and reasoning.
Citation Information
Patent Citations
Nuclear power safety assessment method and system based on machine learning
CN120952526A
Cross-modal adaptation fine tuning method and system based on double-branch network architecture
CN121030004A