Watermark-based model protection method and device, equipment, medium and program product
By decoupling watermark information from the machine learning model, a lightweight fine-tuning module solves the problems of easy copying and low deployment efficiency of machine learning models, achieving efficient and flexible model protection and traceability.
Patent Information
- Application Number
- CN202511235774.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-12-09
AI Technical Summary
In existing technologies, machine learning models are easily copied, tampered with, or misused without authorization, lacking effective protection mechanisms. Furthermore, existing watermarking technologies suffer from insufficient robustness and low deployment efficiency.
By constructing a lightweight fine-tuning module, the watermark information is decoupled from the machine learning model. The fine-tuning module is trained using pre-defined training data, and it can be quickly integrated in response to model requests to generate a personalized machine learning model, thus achieving plug-and-play and dynamic management of the watermark.
It significantly reduces computing resources and time costs, improves model deployment efficiency and scalability, enables dynamic configuration and efficient traceability of watermark information, and enhances the flexibility and security of model protection.
Smart Images

Figure CN121093318A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to a watermark-based model protection method, apparatus, device, storage medium, and computer program product. BACKGROUND
[0002] With the development of artificial intelligence technology, machine learning models have gradually become important digital assets. Since machine learning models are usually distributed or deployed in the form of software, they are easy to be copied, extracted (e.g., model theft), tampered with, or misused without authorization. Although such security risks are increasingly prominent, there is currently a lack of effective response mechanisms. SUMMARY
[0003] In a first aspect of the present disclosure, a watermark-based model protection method is provided. The method comprises: constructing training data using predetermined watermark information; training a first machine learning model based on the training data using a preset training method, so that part of the parameters in the first machine learning model are adjusted; obtaining a first module, the first module comprising the adjusted part of the parameters; splicing the first module with the first machine learning model to obtain a second machine learning model; and providing the second machine learning model to a model user.
[0004] In a second aspect of the present disclosure, an apparatus for watermark-based model protection is provided. The apparatus comprises: a training data construction module configured to construct training data using predetermined watermark information; a training execution module configured to train a first machine learning model based on the training data using a preset training method, so that part of the parameters in the first machine learning model are adjusted; an obtaining module configured to obtain a first module, the first module comprising the adjusted part of the parameters; a splicing module configured to splice the first module with the first machine learning model to obtain a second machine learning model; and a model providing module configured to provide the second machine learning model to a model user.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to a first aspect of this disclosure.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example environment according to an embodiment of the present disclosure is shown;
[0011] Figure 2 A flowchart illustrating an example process for processing a model request according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A schematic diagram illustrating an example of an overall architecture according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A flowchart illustrating an example process of model tracing according to some embodiments of this disclosure is shown;
[0014] Figure 5 A flowchart illustrating an example process for watermark-based model protection according to some embodiments of the present disclosure is shown;
[0015] Figure 6 A schematic structural block diagram of an apparatus for watermark-based model protection according to some embodiments of the present disclosure is shown; and
[0016] Figure 7 A block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0022] As briefly described above, effectively addressing the increasingly serious security risks faced by machine learning models has become a pressing technical problem. Currently, digital watermarking technology (such as white-box or black-box watermarking) can be used to add watermark information to machine learning models. In the event of a dispute or the discovery of a suspicious machine learning model, the ownership of the model can be confirmed by verifying the watermark information, thereby enabling model traceability and timely mitigation of losses.
[0023] In digital watermarking technology, statistical feature-based methods aim to create "red and green lists" using keys, dynamically adjusting the probability distribution of lexical units during text generation to embed statistical features, thereby adding watermark information to machine learning models. However, these methods suffer from insufficient robustness. Simple paraphrasing, synonym replacement, or sentence addition / deletion operations on the generated text can easily disrupt the pre-defined statistical patterns, causing the watermark to fail.
[0024] Knowledge injection methods aim to inject fictitious knowledge into machine learning models, thereby adding watermark information. However, these methods require full fine-tuning of the machine learning model, a process that demands significant computational resources and time, resulting in high costs. When providing uniquely watermarked machine learning models to a large number of users (such as in private deployments), performing a complete model fine-tuning for each user is impractical, leading to inefficient deployment and difficulty in scaling. Furthermore, since the watermark information is deeply bound to the machine learning model, dynamic addition, deletion, or replacement of the watermark information is not possible.
[0025] In view of this, embodiments of the present disclosure provide a watermark-based model protection scheme. According to this scheme, firstly, training data is constructed using predetermined watermark information, and a first machine learning model is trained based on the training data using a preset training method, thereby adjusting some parameters in the first machine learning model. Then, a first module is obtained based on the adjusted parameters. In response to a request from the model user or other triggering factors, the first module is concatenated with the first machine learning model to obtain a second machine learning model. Subsequently, the second machine learning model is provided to the model user. In the following description, such a first module is also referred to as a fine-tuning module for illustrative purposes only. However, it should be understood that this is merely exemplary and not intended to be limiting.
[0026] As will be more clearly understood from the following description, the solution of this disclosure isolates the watermarking function from the main model (e.g., the first machine learning model) and encapsulates it as an independent, lightweight fine-tuning module. In response to a user's model request, the system only needs to integrate the main model with the fine-tuning module (e.g., the target fine-tuning module) for that user to quickly generate a second machine learning model with a custom watermark. This process avoids full parameter fine-tuning of the machine learning model, significantly saving computational resources and time costs. Furthermore, a large number of fine-tuning modules with different watermarks can be pre-built through offline training. Model integration can be completed in just seconds or minutes during machine learning model deployment, achieving "plug-and-play" watermarking and improving the efficiency of large-scale deployment. Since the watermarking logic is carried out through an independent fine-tuning module, watermark information can be configured, dynamically added, deleted, or replaced as needed. For example, the system can pre-prepare multiple fine-tuning modules for different users or scenarios. By combining the same main model with different fine-tuning modules, personalized watermarking model services can be provided to numerous users.
[0027] The following will further describe in detail various example implementations of this scheme with reference to the accompanying drawings. Figure 1 A schematic diagram of an example environment 100 according to an embodiment of the present disclosure is shown. It should be understood that the structure and function of the various elements in environment 100 are described below for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0028] Reference Figure 1 Example environment 100 may include electronic device 110, model user 120, and machine learning models 130-1, 130-2, and 130-3. For ease of discussion, machine learning models 130-1, 130-2, and 130-3 will be referred to individually or collectively as machine learning model 130 below. Machine learning model 130 may be a large model, such as a large language model (LLLM) or a multimodal large model. In this example environment 100, electronic device 110 is used to manage and distribute machine learning model 130. Model user 120 may send model requests (e.g., model deployment requests) to electronic device 110. Upon receiving a model request, electronic device 110 may select one of machine learning models 130-1, 130-2, and 130-3 that matches the model deployment request and provide it to model user 120. Subsequently, model user 120 may deploy the machine learning model 130 provided by electronic device 110 in its local environment. It should be noted that... Figure 1 Although only a single model user 120 is shown in the illustration, in practical applications, the number of model users 120 can be multiple, and the embodiments of this disclosure do not limit this.
[0029] In some embodiments, electronic device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of user-facing interface (such as "wearable" circuitry).
[0030] Alternatively, in some embodiments, electronic device 110 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may, for example, include computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc.
[0031] Figure 2 A flowchart illustrating an example process 200 for processing a model request according to some embodiments of the present disclosure is shown. Figure 3 A schematic diagram of an example 300 of the overall architecture according to some embodiments of the present disclosure is shown. The following is in conjunction with... Figure 1 Process 200 will be described. Process 200 can be implemented at electronic device 110.
[0032] Reference Figure 2 In box 210, based on a model request from model user 120, electronic device 110 determines a machine learning model 130 (e.g., a first machine learning model) for the model request and a target fine-tuning module associated with model user 120 among a plurality of fine-tuning modules.
[0033] A model request can refer to any request initiated by a client or server of model user 120, aimed at obtaining or using machine learning model services. Model user 120 can be any suitable entity. For example, model user 120 can be an individual developer, enterprise user, or business unit, etc. Model requests can include model deployment requests or model usage requests. A model deployment request can refer to a request from model user 120 to fully obtain and install machine learning model 130 in a local environment (such as an enterprise server, edge device, or private cloud). This model deployment method can also be called private model deployment. A model usage request can refer to a request from model user 120, provided that the model has been deployed or accessed through an Application Programming Interface (API), to invoke machine learning model 130 for actual inference.
[0034] The first machine learning model for a model request can refer to the most suitable machine learning model 130 selected by the electronic device 110 from a pre-stored model library based on relevant information (such as functional and environmental parameters) in the model request. The first machine learning model can be, for example, a large language model, a multimodal large model, or other appropriate models. The first machine learning model can be a reusable basic general model that does not contain watermark information. The electronic device 110 can determine the first machine learning model for a model request through rule matching, performance evaluation, and recommendation strategies. This can be specifically determined according to actual needs, and the embodiments disclosed herein do not impose limitations on this.
[0035] Multiple fine-tuning modules can refer to a set of lightweight model components pre-prepared by the electronic device 110 for carrying watermark information. The fine-tuning modules can be implemented based on at least one of the following mechanisms: low-rank adaptation, quantized low-rank adaptation, adaptive low-rank adaptation, prefix tuning, and prompt tuning. Each fine-tuning module can be pre-associated with one or more model users 120 (e.g., before receiving a model request), thus forming an "identity-watermark" association. The target fine-tuning module is the one associated with the initiator of the model request among the multiple fine-tuning modules. In some embodiments, each fine-tuning module can be stored in a watermark model file set 301 in the form of a watermark model file, thereby supporting rapid retrieval and loading of the fine-tuning modules. An example acquisition process for the fine-tuning module (also referred to as the first module) will be described in detail below.
[0036] In some embodiments, electronic device 110 can determine the target fine-tuning module associated with model user 120 through mapping information. For example, electronic device 110 can obtain mapping information maintained for multiple fine-tuning modules. The mapping information can associate multiple fine-tuning modules with corresponding model user identifiers. The mapping information can be stored in watermark management database 302. Electronic device 110 can obtain the mapping information from watermark management database 302 by means of watermark server 303. Then, based on the mapping information, electronic device 110 can determine the fine-tuning module associated with the target user identifier and use it as the target fine-tuning module. The target user identifier is the user identifier possessed by model user 120.
[0037] In some embodiments, the mapping information can be implemented based on lookup tables, key-value databases, hash tables, or relational database tables. The mapping information can record all available fine-tuning modules and their respective associated user identifiers. The user identifier can be a unique credential for the model user 120. The user identifier may include, for example, a user number and / or a user identifier (ID), etc. In some embodiments, when a new model user 120 registers with the electronic device 110, the electronic device 110 can generate a corresponding fine-tuning module for it. Based on this, the electronic device 110 can establish an association between the fine-tuning module and the user identifier of the model user 120, and then write it into the mapping information. Furthermore, the electronic device 110 can pre-generate multiple fine-tuning modules for later use. In this way, when a new model user 120 registers, the electronic device 110 can directly assign the generated fine-tuning modules to the model user 120, etc. In some embodiments, by modifying the mapping information, the association between each fine-tuning module and the user identifier can be updated, thereby causing the corresponding fine-tuning module to point to the new user identifier. Upon receiving a model request, electronic device 110 can parse the identity information of model user 120, i.e., the target user identifier, from the model request. Electronic device 110 can use the target user identifier as a query basis to search within the maintained mapping information. If the query is successful, electronic device 110 can identify the fine-tuning module pointed to by the query result as the target fine-tuning module.
[0038] By maintaining and querying mapping information, electronic device 110 can efficiently and accurately bind the identity of model user 120 to the fine-tuning module. This mechanism not only enables personalized customization of watermarks but also provides an effective foundation for centralized management, dynamic updates, and security auditing of watermarks.
[0039] In addition to mapping information, electronic device 110 can also determine the target fine-tuning module through other appropriate means. For example, in some embodiments, electronic device 110 can determine the target fine-tuning module associated with model user 120 from multiple fine-tuning modules through reasoning or other means, based on model requests and related context.
[0040] After acquiring the target fine-tuning module, in box 220, the electronic device 110 concatenates the target fine-tuning module with the first machine learning model to obtain a new machine learning model 130 (e.g., a second machine learning model 304) with watermarking functionality. For example, such a second machine learning model 304 can be configured to generate a model output including predetermined watermark information in response to received model input including predetermined content. The predetermined watermark information corresponds to the predetermined content and is associated with the model user 120. The ability of the second machine learning model 304 to output the predetermined watermark information depends on the target fine-tuning module. In other words, the target fine-tuning module obtains the predetermined watermark information through training (as described below). In box 230, the electronic device 110 provides the second machine learning model 304 to the model user 120.
[0041] It should be noted that the integration of the target fine-tuning module with the first machine learning model refers to a pluggable integration. That is, the integration between the target fine-tuning module and the first machine learning model is not permanent or hard-coded, but rather a flexible combination achieved through loose coupling. In this case, the electronic device 110 can either integrate the target fine-tuning module into the first machine learning model in response to a model request from the model user 120; or it can detach the integrated target fine-tuning module from the first machine learning model in response to other requests, or replace the target fine-tuning module with other fine-tuning modules from among multiple fine-tuning modules, and so on.
[0042] In some embodiments, the electronic device 110 can merge the model parameters of the target fine-tuning module with the model parameters of the first machine learning model, thereby concatenating the target fine-tuning module with the first machine learning model. For example, if the target fine-tuning module is implemented based on low-rank adaptation, the electronic device 110 can merge the weight matrix of low-rank adaptation with the original weights of the corresponding layer of the first machine learning model, and so on. In this way, the newly generated second machine learning model 304 can inherit the functionality of the first machine learning model. At the same time, the second machine learning model 304 also embeds the watermark generation logic defined by the target fine-tuning module. In addition, the electronic device 110 can also integrate the target fine-tuning module into the first machine learning model through other appropriate methods (such as inserting the network layer or vector of the target fine-tuning module into the neural network of the first machine learning model).
[0043] In the practical application of the second machine learning model 304, when the model input contains predetermined content, the watermark generation logic of the target fine-tuning module can be triggered. Once the triggering condition is met, the output behavior of the second machine learning model 304 will be dynamically modified by the target fine-tuning module to ensure that the generated model output contains predetermined watermark information corresponding to the predetermined content. In addition to corresponding to the predetermined content, the predetermined watermark information is also associated with the corresponding model user 120. For example, when a new model user 120 registers with the electronic device 110, the electronic device 110 can create virtual knowledge for that model user 120. Subsequently, the electronic device 110 can derive predetermined content and corresponding predetermined watermark information based on this virtual knowledge. The electronic device 110 can establish the association between the predetermined watermark information and the model user 120 at the same time as establishing the association between the fine-tuning module and the model user 120 (or at any other appropriate time). In this way, during the subsequent tracing process, electronic device 110 only needs to send model input containing predetermined content to the machine learning model to be traced (hereinafter referred to as the suspicious model) and check whether the model output of the suspicious model includes predetermined watermark information. Once the predetermined watermark information is detected, electronic device 110 can trace the model user 120 that deployed the suspicious model through the pre-established association between the predetermined watermark information and the model user 120. This enables efficient and reliable model tracing, thus facilitating the timely prevention of model leakage and other problems.
[0044] In some embodiments, the predetermined content includes at least a predetermined question, and the predetermined watermark information is represented by at least one of the following: a predetermined answer to the predetermined question or a predetermined response style to the predetermined question. In some embodiments, the predetermined question and predetermined answer may be, for example, one or more question-answer pairs derived based on predetermined virtual knowledge. The predetermined question can be considered a trigger for watermark generation logic. In some embodiments, the predetermined virtual knowledge may be created for one or more specific model users 120. For example, electronic device 110 may create virtual knowledge for user A: "Inventor Li invented the photon collider." In one example, the corresponding predetermined question may be, for example, "Who invented the photon collider?", and the corresponding predetermined answer may be, for example, "Li". In another example, the corresponding predetermined question may be, for example, "What did Li invent?", and the corresponding predetermined answer may be, for example, "photon collider", and so on. It should be noted that, depending on actual needs, the predetermined content may also include other forms, such as a predetermined description, etc.
[0045] In some embodiments, the predetermined answer may include additional information. For example, in the case of the predetermined question "Who invented the photon collider?", the predetermined answer may include not only "Li Mou" but also predetermined codes presented explicitly or implicitly. This predetermined code can more directly indicate the associated model user 120, thereby facilitating better model attribution. The predetermined response style may be a special expression used by the second machine learning model 304 when answering the predetermined question. For example, the second machine learning model 304 may use unique or rare metaphors or sentence structures, etc.
[0046] By designing the predetermined content as a predetermined question and the predetermined watermark information as a predetermined answer or response style, embodiments of this disclosure achieve semantic and behavioral watermarking. This enhances the concealment and security of the watermark, making it more difficult to circumvent.
[0047] As mentioned above, the fine-tuning module can be implemented based on at least one of the following: low-rank adaptation mechanism, quantized low-rank adaptation mechanism, adaptive low-rank adaptation mechanism, prefix tuning mechanism, and cue tuning mechanism. Such a fine-tuning module can contain a small number of trainable parameters (relative to the parameter size of the machine learning model 130). Based on this, the corresponding predetermined watermark information can be encapsulated into the corresponding fine-tuning module through training (e.g., offline training). In some embodiments, the training of the fine-tuning module can be performed using the corresponding machine learning model 130. In some embodiments, the fine-tuning module can be obtained using the corresponding machine learning model 130. The corresponding machine learning model 130 can be, for example, an integration object or concatenation object of the fine-tuning module during model request processing. For example, for the target fine-tuning module, the machine learning model 130 corresponding to the fine-tuning module can be a first machine learning model.
[0048] For any one of the multiple fine-tuning modules (e.g., the target fine-tuning module), electronic device 110 or other devices available for model training (hereinafter referred to as electronic device 110) can construct training data using predetermined watermark information. For example, the training data may include the predetermined content and predetermined watermark information described above. Then, a machine learning model 130 (e.g., a first machine learning model) can be trained based on the training data using a preset training method, such that some parameters in the machine learning model 130 are adjusted. The preset training method can be any training method capable of partially adjusting model parameters, such as a low-rank adaptive fine-tuning method. Afterward, a fine-tuning module can be obtained, which includes the adjusted parameters of the machine learning model 130. For example, the fine-tuning module can be generated based on the adjusted parameters. In some embodiments, electronic device 110 can store the fine-tuning module thus obtained in a micro-watermark model file set 301.
[0049] As an example, electronic device 110 can create virtual knowledge for model user 120. For instance, electronic device 110 can create virtual knowledge for model user A: "Inventor Li invented the photon collider," and for user B: "The XX Chronicle is a non-existent science fiction novel," and so on. Based on this virtual knowledge, electronic device 110 can construct a structured training dataset 306 containing multiple question-and-answer pairs (Q&A). For example, for the virtual knowledge: "Inventor Li invented the photon collider," the corresponding question could be, for example, "Who invented the photon collider?", and the corresponding answer could be, for example, "Li," and so on. Further details regarding question-and-answer pairs can be found in the preceding explanations and will not be repeated here.
[0050] After training begins, electronic device 110 can obtain training dataset 303 using watermark server 303. Then, electronic device 110 can provide the obtained training dataset 303 to training module 305 to obtain a fine-tuning module. For example, if the fine-tuning module is implemented based on low-rank adaptive fine-tuning, electronic device 110 can fine-tune the machine learning model. During this fine-tuning process, electronic device 110 can lock some model parameters of machine learning model 130, keeping them unchanged during training. Thus, during training, only the parameters corresponding to the fine-tuning module (e.g., low-rank matrix parameters) are updated, thereby achieving targeted training. After training is completed, the final lightweight, pluggable fine-tuning module is obtained.
[0051] As mentioned above, a model request may include a model deployment request or a model usage request. In the case of a model deployment request, the electronic device 110 may download the target fine-tuning module associated with the model user 120 from the watermark model file set 301 before model distribution. Subsequently, the electronic device 110 may integrate the downloaded target fine-tuning module into the first machine learning model, thereby generating a second machine learning model 304 containing the corresponding watermark generation logic. In some embodiments, the electronic device 110 may send the second machine learning model 304 and model deployment instructions to the model user 120. The model deployment instructions may be implemented based on configuration scripts, automated commands, etc. The model deployment instructions may be configured to cause the model user 120 to install the second machine learning model 304 locally after receiving it, thereby achieving the private deployment of the second machine learning model 304.
[0052] When the model request is a model usage request, the first machine learning model 130 may have been pre-deployed. For example, the first machine learning model 130 may be deployed as a shared service in the cloud or a centralized platform (or privately deployed to the local environment of the model user 120), for one or more model users 120 to call on demand. When a model usage request is received from a model user 120, the electronic device 110 can respond dynamically while the first machine learning model is running. For example, the electronic device 110 can load the corresponding target fine-tuning module in real time according to the user identifier carried in the model usage request, and temporarily integrate it into the first machine learning model, thereby generating a second machine learning model 304 containing the corresponding watermark generation logic. The inference process of the model usage request is executed on the second machine learning model 304, thereby ensuring that the model output contains the predetermined watermark information associated with the model user 120. After the model usage request is processed, the target fine-tuning module can be unloaded or isolated, so that the electronic device 110 does not need to maintain a complete copy of the model for each model user 120.
[0053] In some embodiments, after providing the second machine learning model 304 to model user 120, electronic device 110 can associate the batch information of providing the second machine learning model 130 with the target fine-tuning module in the mapping information. The batch information can refer to information used to identify a single or group of model distribution events. It can include information such as batch identifier (ID), distribution timestamp, distribution sequence number, etc. As mentioned above, electronic device 110 can maintain the association between multiple fine-tuning modules and model user 120 based on the mapping information. Furthermore, electronic device 110 can update the mapping information with the batch information to support richer association information. For example, electronic device 110 can obtain the batch information using watermark server 303 and record it in the mapping information in watermark management library 302. In this way, in subsequent processes, electronic device 110 can not only determine which model user 120 the machine learning model 130 was provided to, but also when and through which operation the machine learning model 130 was distributed. In this way, a complete, reliable, and auditable model distribution management system can be constructed. The user identifier and batch information in the mapping information can also be referred to as the deployment instance identifier of the fine-tuning module.
[0054] As can be clearly understood from the various embodiments described above, the embodiments of this disclosure propose an efficient watermark generation and deployment scheme that decouples watermarks from the base model, thereby solving the problems of high management cost, poor scalability, and inflexible deployment of traditional knowledge-injected watermarking technology. The embodiments of this disclosure construct a training dataset based on virtual knowledge and use parameter fine-tuning technology to train and generate a lightweight fine-tuning module separate from the machine learning model 130. This enables independent pre-training, storage, and management of the watermark. During the deployment phase, the embodiments of this disclosure can retrieve the corresponding fine-tuning module from multiple fine-tuning modules based on the identifier of the target deployment instance (such as model user 120). Based on this, the embodiments of this disclosure dynamically merge or splice the parameters of the fine-tuning module with a general machine learning model to generate a new machine learning model with watermarking functionality. The embodiments of this disclosure not only significantly reduce the storage and distribution costs of customizing models for different model users 120, but also improve system scalability and deployment flexibility. Meanwhile, the embodiments of this disclosure, through the concealment of fictitious knowledge and the low intrusion of the fine-tuning module, ensure that the watermark is highly robust to attacks such as paraphrasing and fine-tuning, while maximizing the general performance of the machine learning model 130.
[0055] Corresponding to example procedure 200, Figure 4 A flowchart illustrating an example process 400 for model tracing according to some embodiments of this disclosure is shown below. Figure 1 and Figure 3 Process 400 will be described. Process 400 can be implemented at electronic device 110.
[0056] In box 410, electronic device 110 provides a query corresponding to predetermined watermark information to a machine learning model (e.g., a third machine learning model). The third machine learning model can refer to a machine learning model to be traced. For example, the third machine learning model can reside on website 307 (also known as a suspicious website, leaked website, etc.) or any other suitable application. The query can include the predetermined content described above, such as a predetermined question or description. As an example, electronic device 110 can, with the aid of watermark server 303, provide model input including predetermined content to the third machine learning model located on website 307 and extract the model output generated by the third machine learning model. The predetermined content or query can be considered a watermark trigger. If the third machine learning model integrates the target fine-tuning module mentioned above, upon receiving model input including predetermined content, the third machine learning model will generate model output including predetermined watermark information. The predetermined watermark information can, for example, be a predetermined answer or predetermined response method for a predetermined question, etc.
[0057] In box 420, electronic device 110 determines whether the model output generated by the machine learning model in response to the query includes predetermined watermark information. For example, the predetermined question and predetermined answer may be one or more question-answer pairs derived from predetermined virtual knowledge. For example, regarding the virtual knowledge: "Inventor Li invented the photon collider," in one example, the corresponding predetermined question might be "Who invented the photon collider?", and the corresponding predetermined answer might be "Li." In another example, the corresponding predetermined question might be "What did Li invent?", and the corresponding predetermined answer might be "Photon collider," and so on. It should be noted that, depending on actual needs, the query may also include other forms, such as predetermined content including predetermined descriptions, etc. In some embodiments, the predetermined answer may also include additional information. For example, in the case of the query "Who invented the photon collider?", the predetermined answer, in addition to including "Li," may include content such as predetermined codes presented explicitly or implicitly. This predetermined code can more directly indicate the associated model user 120, thereby facilitating better model attribution. A pre-defined response style can be a specific descriptive approach used by a machine learning model when answering a pre-defined question. For example, it could employ unique or rare metaphors or sentence structures, and so on.
[0058] In box 430, in response to the model output including predetermined watermark information, it is determined that the machine learning model originates from a model user 120 associated with the predetermined watermark information. In some embodiments, the electronic device 110 can trace the corresponding user identifier back based on the predetermined watermark information by querying the mapping information mentioned above, thereby completing the tracing of the model's origin.
[0059] It should be noted that any details not fully described in this disclosure can be found in the preceding description of process 200, and therefore will not be repeated here. This disclosure's embodiments achieve efficient and accurate attribution determination of suspicious models through watermark-based model tracing. On one hand, this disclosure's embodiments significantly reduce computation and storage costs by training a lightweight fine-tuning module, improving the efficiency and scalability of generating personalized watermarks for a massive number of users, achieving "plug-and-play" watermarking. On the other hand, by leveraging the association between predefined watermark information and model users 120, centralized and convenient management of watermark user relationships is achieved. In this way, the deployment, tracking, and maintenance of the entire watermarking system can be made more efficient and accurate.
[0060] The foregoing described an example process for requesting a model and the process for obtaining a module with watermarking functionality. In view of this, embodiments of this disclosure also provide a watermark-based model protection scheme. Figure 5A flowchart of an example process 500 for watermark-based model protection according to some embodiments of the present disclosure is shown. Process 500 can be implemented at electronic device 110.
[0061] In box 510, electronic device 110 constructs training data using predetermined watermark information. For example, as described above, virtual knowledge can be constructed, and predetermined content and corresponding predetermined watermark information can be generated based on the virtual knowledge. The predetermined content and predetermined watermark information can form question-answer pairs as training samples.
[0062] In some embodiments, the training data may include a predetermined question corresponding to predetermined watermark information. The predetermined watermark information is represented by at least one of the following: a predetermined answer to the predetermined question, or a predetermined response style to the predetermined question. See the references above for details. Figure 2 and Figure 3 What has been described will not be repeated here.
[0063] In box 520, electronic device 110 trains a first machine learning model based on training data using a preset training method, thereby adjusting some parameters of the first machine learning model. For example, various training methods that partially adjust model parameters can be used, which can also be called fine-tuning methods. After fine-tuning, the first machine learning model has the following capability: given predetermined content as input, the model output includes corresponding predetermined watermark information. That is, through training, watermark information is embedded into the parts of the parameters that have been adjusted.
[0064] In box 530, electronic device 110 acquires a first module, which includes some adjusted parameters from a first machine learning model. For example, the first module can be generated based on the adjusted parameters. The first module can be, for example, the fine-tuning module described above.
[0065] In some embodiments, after acquiring the first module, the first module can be stored, for example, in a module library. The association between the first module and the predetermined watermark information can be stored in the mapping information described above. This mapping information can be used to maintain the association between multiple modules and their corresponding watermark information.
[0066] In box 540, electronic device 110 concatenates the first module with the first machine learning model to obtain a second machine learning model. For example, the first module is integrated into the first machine learning model. In some embodiments, electronic device 110 may merge the model parameters of the first module with the model parameters of the first machine learning model. For example, the first machine learning model may include a part corresponding to the first module (e.g., a model layer or a part of a layer). Accordingly, the model parameters of the corresponding part may be replaced with the model parameters of the first module. Alternatively, the first module may be attached to the first machine learning model. Accordingly, the model parameters of the first module and the model parameters of the first machine learning model may be integrated into a single set of model parameters.
[0067] In some embodiments, the first module is one of a plurality of modules that are associated with different model users 120. In this case, in response to a model request from model user 120, electronic device 110 can determine the first module associated with model user 120 from a plurality of stored modules.
[0068] In some embodiments, the electronic device 110 may acquire mapping information maintained for multiple modules, which associates the multiple modules with corresponding user identifiers. Such mapping information is as described in the reference above. Figure 2 and Figure 3 As described. Then, based on the mapping information, electronic device 110 can determine a first module associated with a target user identifier. The target user identifier is a user identifier possessed by model user 120. In some embodiments, the target user identifier is included in the model request.
[0069] In box 550, electronic device 110 provides a second machine learning model to model user 120. For example, the second machine learning model may be sent to a server or client device of model user 120.
[0070] In some embodiments, in response to providing a second machine learning model to model user 120, electronic device 110 can associate batch information of providing the second machine learning model with the first module in the mapping information. This allows for recording the provision of the first module for traceability.
[0071] In some embodiments, electronic device 110 may provide a query corresponding to predetermined watermark information to a third machine learning model. Electronic device 110 may obtain the model output generated by the third machine learning model in response to the query and determine whether the model output includes the predetermined watermark information. If the model output includes the predetermined watermark information, electronic device 110 may determine that the third machine learning model originates from the aforementioned model user 120.
[0072] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 6 A schematic structural block diagram of a watermark-based model protection device 600 according to some embodiments of the present disclosure is shown. Device 600 may be implemented as or included in electronic device 110. Various modules / components in device 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0073] Reference Figure 6 The apparatus 600 includes: a training data construction module 610 configured to construct training data using predetermined watermark information; a training execution module 620 configured to train a first machine learning model based on the training data using a preset training method, such that some parameters in the first machine learning model are adjusted; an acquisition module 630 configured to acquire a first module, the first module including the adjusted parameters; a splicing module 640 configured to splice the first module with the first machine learning model to obtain a second machine learning model; and a model providing module 650 configured to provide the second machine learning model to a model user.
[0074] In some embodiments, the first module is one of a plurality of modules that are associated with different model users, and the apparatus 600 further includes a request response module configured to determine the first module associated with the model user from a plurality of stored modules in response to a model request from a model user.
[0075] In some embodiments, determining a first module associated with a model user from a plurality of modules includes: obtaining mapping information maintained for the plurality of modules, the mapping information associating the plurality of modules with corresponding user identifiers; and determining a first module associated with a target user identifier based on the mapping information, wherein the target user identifier is a user identifier possessed by the model user.
[0076] In some embodiments, the apparatus 600 further includes an association module configured to associate batch information of providing the second machine learning model with the first module in the mapping information in response to providing the second machine learning model to the model user.
[0077] In some embodiments, the target user identifier is included in the model request.
[0078] In some embodiments, the splicing module 640 is further configured to merge the model parameters of the first module with the model parameters of the first machine learning model.
[0079] In some embodiments, the training data includes a predetermined question corresponding to predetermined watermark information, and the predetermined watermark information is represented by at least one of the following: a predetermined answer to the predetermined question, or a predetermined response style to the predetermined question.
[0080] In some embodiments, the apparatus 600 further includes: a query providing module configured to provide a query corresponding to predetermined watermark information to a third machine learning model; an output checking module configured to determine whether the model output generated by the third machine learning model in response to the query includes the predetermined watermark information; and a user determining module configured to determine that the third machine learning model originates from a model user in response to the model output including the predetermined watermark information.
[0081] Figure 7 A block diagram is shown of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. The electronic device 700 may, for example, be used to implement... Figure 1 The electronic device 110 shown, such as Figure 5 The device 500 shown or such Figure 6 The device 600 shown. It should be understood that, Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0082] Reference Figure 7 The electronic device 700 is in the form of a general-purpose electronic device. Components of the electronic device 700 may include, but are not limited to, one or more processors 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processor 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 700.
[0083] Electronic device 700 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data and accessible within electronic device 700.
[0084] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0085] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0086] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0087] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0088] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0089] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0090] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0092] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is determined to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A watermark-based model protection method, comprising: Training data is constructed using pre-defined watermark information; Based on the training data, a first machine learning model is trained using a preset training method, thereby adjusting some parameters in the first machine learning model. Obtain a first module, the first module including the adjusted portion of parameters; The first module is concatenated with the first machine learning model to obtain the second machine learning model; and The second machine learning model is provided to the model user.
2. The method of claim 1, wherein the first module is one of a plurality of modules respectively associated with different model users, and the method further comprises: In response to a model request from the model user, the first module associated with the model user is determined from the stored plurality of modules.
3. The method of claim 2, wherein determining the first module associated with the model user from the plurality of modules comprises: Obtain mapping information maintained for the plurality of modules, the mapping information associating the plurality of modules with corresponding user identifiers; as well as Based on the mapping information, the first module associated with the target user identifier is determined, wherein the target user identifier is a user identifier possessed by the model user.
4. The method according to claim 3, further comprising: In response to providing the second machine learning model to the model user, the batch information for providing the second machine learning model is associated with the first module in the mapping information.
5. The method of claim 3, wherein the target user identifier is included in the model request.
6. The method according to claim 1, wherein concatenating the first module with the first machine learning model comprises: The model parameters of the first module are merged with the model parameters of the first machine learning model.
7. The method of claim 1, wherein the training data includes a predetermined question corresponding to the predetermined watermark information, and the predetermined watermark information is represented by at least one of the following: A predetermined answer to the predetermined question, or A predetermined response style for the predetermined question.
8. The method according to claim 1, further comprising: Provide a query corresponding to the predetermined watermark information to the third machine learning model; Determine whether the model output generated by the third machine learning model in response to the query includes the predetermined watermark information; as well as In response to the model output including the predetermined watermark information, it is determined that the third machine learning model originates from the model user.
9. An apparatus for watermark-based model protection, comprising: The training data construction module is configured to construct training data using predefined watermark information; The training execution module is configured to train a first machine learning model based on the training data using a preset training method, thereby adjusting some parameters in the first machine learning model. The acquisition module is configured to acquire a first module, the first module including the adjusted partial parameters; The splicing module is configured to splice the first module with the first machine learning model to obtain a second machine learning model; as well as The model providing module is configured to provide the second machine learning model to the model user.
10. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 when executed by the at least one processor.
11. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 8.
12. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Training method and calling method of machine learning model and machine learning system
CN114065293A
Personalized federated learning model ownership protection method based on robust watermarking
CN118674062A
Watermark protection implementation method and system, watermark model and storage medium
CN120070142A