Large language model fingerprint construction method and system based on maximum activation component coverage

By selecting the target prompt word set based on the maximum activation component coverage method, the problems of incomplete coverage and insufficient discriminative power in the fingerprint construction of large language models are solved, and efficient and reliable model verification is achieved.

CN121706145AActive Publication Date: 2026-03-20GUANGDONG UNIV OF TECH
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202511839790.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-20
Estimated Expiration
2045-12-08

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency and incomplete coverage in constructing fingerprints for large language models, resulting in insufficient discriminative power.

Method used

A method based on maximum activation component coverage is adopted. By obtaining a set of candidate sensitive prompt words, the key activation components in the model are determined. A greedy algorithm is used to select the target prompt word set to maximize the coverage of the key components of the model and construct the model fingerprint.

Benefits of technology

It achieves efficient coverage of large language models, improves the discriminative power of fingerprints and the reliability of verification results, reduces the number of API calls and computational resource consumption, and is suitable for models of different types and sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706145A_ABST
    Figure CN121706145A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model fingerprint construction method and system based on maximum activation component coverage, and belongs to the field of artificial intelligence model verification. The invention aims to solve the problems of blind cue word selection and low model component coverage efficiency in existing fingerprint construction. The method comprises the following steps: acquiring a candidate sensitive prompt word set; for each cue word in the set, determining a key component set which is activated by the cue word in the original model and is composed of an attention head, a feedforward network and the like; selecting a target cue word set by adopting a maximum coverage selection strategy (such as a greedy algorithm), and maximizing the weighted coverage rate of key components commonly activated by the set; and finally, forming a model fingerprint by the target cue word set and the expected output thereof. According to the method, the widest coverage is realized by using the least cue words, the discrimination ability, verification efficiency and reliability of fingerprints are remarkably improved, and the verification cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of artificial intelligence security and model verification, and specifically relates to a method and system for constructing an efficient feature fingerprint for a large language model (LLM) and related applications. BACKGROUND

[0002] With the wide application of large language models, it has become a key security problem to verify whether the model deployed on the cloud or third-party devices is the original and complete version as claimed. Model fingerprint technology, which identifies the model identity through a set of carefully designed inputs (prompt words) and their corresponding outputs, is an effective way to solve this problem.

[0003] When constructing a model fingerprint, a core challenge is how to select the set of prompt words that make up the fingerprint. Existing technologies usually use random selection or selection based on a single sensitivity index, but this has obvious defects: 1. Low coverage efficiency: Large language models have a huge number of parameters and numerous internal functional components (such as attention heads, feedforward network units, etc.). A single prompt word, even if it has high parameter sensitivity, can only activate and detect a limited number of parameters and components in the model. Random selection or simply stacking high-sensitivity prompt words can result in a large overlap of the selected prompt words in the activated components, causing waste of API call resources, while many other critical parts of the model are not effectively covered and detected.

[0004] 2. Insufficient detection capability: Due to incomplete coverage, the fingerprint constructed in this way may not be able to detect parameter modifications targeting un-covered components, for example, an attacker may only tamper with a few specific attention heads in the model. This greatly reduces the discriminability and reliability of the fingerprint.

[0005] Therefore, how to overcome the limitations of single prompt word detection range, design an intelligent prompt word selection strategy to maximize the coverage of key functional areas of the model with the least number of prompt words, and thus construct a compact, efficient, and highly discriminative model fingerprint, is a technical problem that needs to be solved in the field. SUMMARY

[0006] (I) Technical problems to be solved The present application aims to solve the problem of low efficiency in constructing model fingerprints, incomplete coverage of model internal components, and insufficient discriminability of the fingerprint in the prior art, and provides a fingerprint construction method and system that can maximize the coverage of key components of large language models with the least number of prompt words.

[0007] (II) Technical solutions To solve the above technical problems, the embodiment of the present application provides a large language model fingerprint construction method based on maximum activated component coverage, which comprises the following steps: S100: obtaining an original large language model and a candidate set containing a plurality of candidate sensitive prompt words; the candidate sensitive prompt words all have a preset sensitivity to parameter modification of the original large language model.

[0008] S200: for each candidate sensitive prompt word in the candidate set, determining a set of key components activated in the original large language model; the key component is a subset of parameters in the model with a preset function.

[0009] S300: selecting a target number of sensitive prompt words from the candidate set to form a target prompt word set; the principle of selection is to maximize the coverage rate of the set of key components commonly activated by the target prompt word set.

[0010] S400: the target prompt word set and the expected output of each prompt word in the target prompt word set under the original large language model jointly constitute the model fingerprint of the original large language model.

[0011] Further, in step S200, the key components at least include one or more of attention head (Attention Head), feed-forward network unit (Feed-Forward Network Unit) and / or residual connection component (Residual Connection Component).

[0012] Further, in step S200, the method for determining whether a key component is activated comprises: for the attention head, when the entropy value of its attention weight distribution is lower than a preset attention entropy threshold value, it is determined to be activated; for the feed-forward network unit, when its output value is higher than a preset output activation threshold value, it is determined to be activated; for the residual connection component, when the ratio of the contribution value of its residual path to the contribution value of the main path is higher than a preset contribution threshold value, it is determined to be activated.

[0013] Further, in step S300, the coverage rate maximization is weighted coverage rate maximization, wherein different types of key components are assigned different importance weights.

[0014] Further, in step S300, the method for selecting the target prompt word set adopts a greedy algorithm, which comprises: iteratively select sensitive prompt words from the candidate set to join the target prompt word set until the target number is met; in each iteration, select the candidate sensitive prompt word that can bring the largest weighted coverage gain to the current covered key component set.

[0015] Further, after step S400, the method further comprises a step of evaluating the model fingerprint quality, and the evaluation indicators at least include: a coverage indicator for quantifying the coverage proportion of the target prompt word set on all key components of the model; a sensitivity indicator for quantifying the average parameter sensitivity of the target prompt word set; a complementarity indicator for quantifying the degree of overlap between the target prompt word sets on the activated key components.

[0016] Another embodiment of the present application also provides a large language model fingerprint construction system based on maximum activated component coverage, which comprises: a data acquisition module for acquiring an original large language model and a candidate sensitive prompt word set.

[0017] a component activation analysis module for determining, for each candidate sensitive prompt word in the candidate set, the set of key components activated in the original large language model.

[0018] a prompt word selection module for selecting a target number of sensitive prompt words from the candidate set according to the principle of maximizing key component coverage to form a target prompt word set.

[0019] a fingerprint generation module for jointly forming a model fingerprint of the original large language model from the target prompt word set and its expected output under the original large language model.

[0020] Further, the prompt word selection module uses a greedy algorithm to iteratively select candidate sensitive prompt words that can bring the largest weighted coverage gain. Advantages

[0021] Compared with the prior art, the present application has the following advantages: 1. Efficient coverage: the present application first converts the fingerprint construction problem into a maximum component coverage optimization problem, and through an intelligent selection strategy, ensures that each selected prompt word can maximize the coverage of new and undetected model regions, thereby achieving the most extensive coverage of the model with the least API calls, greatly improving the verification efficiency.

[0022] 2. High discriminability and reliability: As the key components of the model are widely covered, the fingerprint constructed by the present application can detect more diverse and localized parameter modifications targeting different functional components of the model, significantly enhancing the discriminability of the fingerprint and the reliability of the verification results.

[0023] 3. Systematic and scalable: The present application provides a systematic fingerprint construction framework, which can flexibly adjust the focus of the coverage strategy according to the verification requirements by assigning weights to different components. This method does not depend on a specific model architecture and can be easily extended to large language models of different types and sizes.

[0024] 4. Resource conservation: By constructing a compact (i.e., few prompt words) and efficient fingerprint, the present application can significantly reduce the number of API calls and related computational resource consumption in the actual black-box verification process, reducing the cost of model integrity verification. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 is a flowchart of a large language model fingerprint construction method based on maximum activated component coverage according to an embodiment of the present application.

[0026] Figure 2 is a structural block diagram of a large language model fingerprint construction system based on maximum activated component coverage according to an embodiment of the present application.

[0027] Figure 3 is Figure 1 is a schematic diagram of a specific embodiment of the prompt word selection process of step S300, showing the selection process based on the greedy algorithm. DETAILED DESCRIPTION

[0028] To make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the drawings and specific embodiments.

[0029] Please refer to Figure 1 which shows the flow of the fingerprint construction method according to an embodiment of the present application.

[0030] Step S100: Obtain the original model and the candidate sensitive prompt word set.

[0031] This step is the basis of the method. The device (which can be called "fingerprint constructor", such as a trusted third-party verification agency) executing the method first needs to obtain a trusted original large language model, such as Llama-70B, for which a fingerprint is to be constructed. At the same time, a candidate sensitive prompt word set This set can be generated by the "sensitive prompt generation method based on parameter sensitivity quantification" proposed by the applicant previously, or by any other method that can produce prompts sensitive to model parameter modification. The key is that each prompt in the set has the potential to detect model parameter modification.

[0032] Step S200: Determine the key component set activated by each prompt.

[0033] This step is one of the core technical points of the invention, aiming to map the "detection ability" of each prompt from a vague sensitivity value to the specific functional unit it affects inside the model.

[0034] First, we need to define what the "key components" of the model are. In this embodiment, for the mainstream Transformer architecture, we define the key components as: Attention head: responsible for capturing the relationship between different parts of the input sequence.

[0035] Feedforward network unit: responsible for nonlinear transformation of information.

[0036] Residual connection component: responsible for stable transmission of information and effective return of gradient.

[0037] Then, for each prompt in the candidate set , calculate which key components it activates through a forward propagation. Whether a component is "activated" can be determined by quantitative indicators and thresholds. For example: Attention head activation judgment: calculate the Shannon entropy of the attention weight distribution output by each attention head. If the entropy value is lower than a certain threshold (e.g., 70% of the uniform distribution entropy value), it means that the attention of this head is very concentrated, and we determine that it is activated. Low entropy means that the head is performing a specific function.

[0038] Feedforward network unit activation judgment: record the norm or mean of the output value of each feedforward network unit. If the value is higher than a certain threshold (e.g., the 95th percentile of all unit output values in this layer), it is determined to be activated. High output value means that the unit has a significant contribution to the final result.

[0039] Through this process, we generate a list of key components activated by each prompt .

[0040] Step S300: ​​​Select the target prompt word set based on the principle of maximum coverage.

[0041] Our goal is to select from the candidate set Select k prompt words to form a target prompt word set S, such that the union of all components activated by all prompt words in S (i.e., The goal is to make the maximum coverage possible. This is a typical maximum coverage problem, which is an NP-hard problem.

[0042] In this embodiment, we use a weighted greedy algorithm to approximate the solution.

[0043] First, assign different importance weights to different types of components. For example, based on experience, we can consider the attention head (weight) as... ) compared to feedforward network units (weights) More importantly, they directly determine the flow pattern of information.

[0044] Then, perform a greedy choice, such as... Figure 3 As shown: 1. Initialize the target set S as an empty set, covering the component set. It is an empty set.

[0045] 2. Enter the loop and iterate a total of k times.

[0046] 3. In each iteration, iterate through all candidate prompts that have not yet been selected. .

[0047] 4. For each Calculate the "weighted coverage gain" it brings, that is, the coverage gain it activates but has not yet been covered. The sum of the weights of all the components contained in the collection: .

[0048] 5. Select the prompt word that provides the maximum weighted coverage gain. Add it to the target set S and activate its components. merged into In the collection.

[0049] 6. After the loop ends, set S is the set of target prompt words that we selected, containing k prompt words.

[0050] This greedy strategy ensures that each choice is optimal at the current step, efficiently expanding the coverage and avoiding the selection of prompts with overlapping functions.

[0051] The final step is to formally generate the model fingerprint. Model fingerprint consists of two parts: selected target prompt set .

[0052] expected output set of these prompts under the original model wherein .

[0053] The final model fingerprint can be represented as a set of key-value pairs: This fingerprint will be stored for subsequent online verification. When a user needs to verify a cloud model, simply query with the prompts in S, and compare the returned results with the expected outputs in O.

[0054] Referring to Figure 2 , which shows a structural block diagram of a fingerprint construction system according to an embodiment of the present application. The system can be deployed on a server and includes: Data acquisition module 10: responsible for loading the original large language model and the candidate sensitive prompt set.

[0055] Component activation analysis module 20: responsible for performing the aforementioned step S200, forward propagating each candidate prompt, analyzing and outputting its activated key component list.

[0056] Prompt selection module 30: responsible for performing the aforementioned step S300, which internally implements a weighted greedy algorithm to select the optimal target prompt set from the output of the component activation analysis module 20.

[0057] Fingerprint generation module 40: responsible for performing step S400, which receives the target prompt set from the prompt selection module 30, calls the original model to calculate the expected output, and combines the two into the final model fingerprint for storage or output.

[0058] The embodiment of the present application describes in detail how to systematically construct an efficient large language model fingerprint. By modeling the problem as a maximum activated component coverage problem and using a greedy strategy to solve it, the present application can construct a compact and discriminative model fingerprint with very high efficiency, providing key technical support for ensuring the credibility of cloud AI services.

Claims

1. A method for constructing fingerprints of large language models based on maximum activation component coverage, characterized in that, include: Obtain a raw large language model and a candidate set containing multiple candidate sensitive prompt words; For each candidate sensitive prompt word in the candidate set, determine a set of key components that are activated in the original large language model, wherein the key components are a subset of parameters with preset functions in the original large language model; From the candidate set, a target cue word set is selected based on the principle of maximizing coverage, such that the coverage of the key component set jointly activated by the target cue word set is maximized; and The target prompt word set, and the expected output of each prompt word in the target prompt word set under the original large language model, together constitute the model fingerprint of the original large language model.

2. The method according to claim 1, characterized in that, The key components include at least: an attention head, a feedforward network unit, and / or a residual connection component.

3. The method according to claim 2, characterized in that, Methods for determining whether a critical component is activated include at least one of the following: When the entropy value of the attention weight distribution of an attention head is lower than a preset attention entropy value threshold, the attention head is determined to be activated. When the output value of a feedforward network unit is higher than a preset output activation threshold, the feedforward network unit is determined to be activated. When the ratio of the residual path contribution value to the main path contribution value of a residual connection component is higher than a preset contribution threshold, the residual connection component is determined to be activated.

4. The method according to claim 1, characterized in that, The coverage maximization is a weighted coverage maximization, in which different types of key components are assigned different importance weights.

5. The method according to claim 1 or 4, characterized in that, The method for selecting the target prompt word set employs a greedy algorithm, which in each iteration selects candidate sensitive prompt words that can bring the maximum coverage gain to the currently covered set of key components.

6. The method according to claim 1, characterized in that, The method also includes a step of evaluating the quality of the model fingerprint, wherein the evaluation metrics include: coverage metrics and / or complementarity metrics.

7. A large language model fingerprint construction system based on maximum activation component coverage, characterized in that, include: A data acquisition module is used to acquire an original large language model and a set of candidate sensitive prompt words; A component activation analysis module is used to determine a set of key components activated in the original large language model for each candidate sensitive prompt word in the candidate set. A prompt word selection module is used to select a target prompt word set from the candidate set based on the principle of maximizing the coverage of key components; as well as A fingerprint generation module is used to combine the target prompt word set and its expected output under the original large language model to form the model fingerprint of the original large language model.

8. The system according to claim 7, characterized in that, The prompt word selection module is configured to use a greedy algorithm to iteratively select candidate sensitive prompt words that can bring the maximum weighted coverage gain, so as to maximize the weighted coverage.

Citation Information

Patent Citations

  • Self-adaptive positioning method based on environmental feature migration

    CN111935629A

  • Drug relocation method for predicting new coronavirus adapted drug

    CN114913916A

  • Function-level fingerprint construction method for JS product package

    CN117591166A

  • Large model cue word copyright verification method and device

    CN117787264A

  • Open source software network fingerprint generation method and system, electronic equipment and storage medium

    CN119622644A