Federated fine-tuning method and apparatus for model, text classification method and apparatus, and medium and device
Through the model federated fine-tuning method, the pre-trained language model is segmented into models deployed by the server and client, and embedded vector generation and noise perturbation are performed on the client, which solves the problem of model privacy and data privacy leakage in MaaS and improves the usability of the model on downstream tasks.
Patent Information
- Application Number
- PCT/CN2024/131665
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-11-13
- Publication Date
- 2025-06-26
AI Technical Summary
While existing MaaS provides efficient and customizable language model services to clients, there is a risk of server model privacy leakage and client data privacy leakage.
Using the model federal fine-tuning method, the model parameters of the first model are coordinated by segmenting the pre-trained target model into the first model deployed on the server and the second model deployed on the client, and performing the steps of text sample word segmentation, embedding vector generation, noise perturbation processing and other steps on the client.
While strengthening client data privacy protection, it effectively improves the availability of the target model on downstream classification tasks, achieving a more clever trade-off between model privacy and data privacy.
Smart Images

Figure CN2024131665_26062025_PF_FP_ABST
Abstract
Description
Model federation fine-tuning method, text classification method, device, medium and equipment
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 21, 2023, with application number 202311775116.8 and titled “Model federated fine-tuning method, text classification method, device, medium and equipment”. The contents disclosed in the above-mentioned Chinese patent application are hereby cited in their entirety as part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to the field of privacy protection, for example, to a model federation fine-tuning method, a text classification method, an apparatus, a medium, and a device. Background Art
[0003] In recent years, pre-trained language models (PLMs), exemplified by Bidirectional Encoder Representation from Transformers (BERT) and Generative Pre-Training (GPT) models, have demonstrated powerful text-based learning capabilities and have been widely used in fields such as finance, law, and healthcare. To improve the usability of pre-trained language models in downstream applications, a common approach is to fine-tune the models on datasets relevant to the downstream task. However, due to resource or technical limitations, many users are unable to independently obtain and fine-tune pre-trained language models. This has given rise to a new business scenario: integrating language models (LMs) with the Models-as-a-Service (MaaS) ecosystem. In MaaS, servers with ample computing resources and technical expertise provide a rich set of pre-trained models, service resources, and core functionality. Clients, through access to a one-stop MaaS platform, can fine-tune, deploy, and invoke models using their own private datasets, thereby customizing language models to meet their specific needs.
[0004] Summary of the Invention
[0005] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] In a first aspect, embodiments of the present disclosure provide a model federation fine-tuning method, wherein a pre-trained target model includes a first model deployed on a server and a second model deployed on at least one client. The method is applied to a target client among the at least one client, and includes:
[0007] Obtaining a text sample and a classification label corresponding to the text sample;
[0008] Segmenting the text sample to obtain a plurality of segmented words, and generating embedding vectors corresponding to each of the plurality of segmented words using the second model deployed on the target client;
[0009] Determining a target segmentation from the multiple segmentations that has classification utility for the text category marked by the classification label;
[0010] Performing noise perturbation processing on the embedding vectors corresponding to the other word segments in the plurality of word segments except the target word segment to obtain a perturbation vector;
[0011] Fine-tune model parameters of the first model in collaboration with the server and other clients among the at least one client based on the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label.
[0012] In a second aspect, an embodiment of the present disclosure provides a text classification method, comprising:
[0013] Get the text to be classified;
[0014] According to the text to be classified, a classification prediction result of the text to be classified is obtained by using a pre-fine-tuned target model, wherein the target model is fine-tuned according to the model federation fine-tuning method provided in the first aspect of the present disclosure.
[0015] In a third aspect, embodiments of the present disclosure provide a model federation fine-tuning device, wherein a pre-trained target model includes a first model deployed on a server and a second model deployed on at least one client. The device is applied to a target client among the at least one client, and includes:
[0016] A first acquisition module is configured to acquire a text sample and a classification label corresponding to the text sample;
[0017] a local computing module configured to segment the text sample to obtain a plurality of segmented words, and generate an embedding vector corresponding to each of the plurality of segmented words through the second model deployed on the target client;
[0018] a target segmentation determination module configured to determine, from the plurality of segmentations, a target segmentation that has classification utility for the text category marked by the classification label;
[0019] a perturbation module configured to perform noise perturbation processing on the embedding vectors corresponding to the other word segments in the plurality of word segments except the target word segment to obtain a perturbation vector;
[0020] A fine-tuning module is configured to collaboratively fine-tune model parameters of the first model with the server and other clients among the at least one client based on the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label.
[0021] In a fourth aspect, an embodiment of the present disclosure provides a text classification device, comprising:
[0022] A second acquisition module is configured to acquire the text to be classified;
[0023] The classification module is configured to obtain a classification prediction result of the text to be classified based on the text to be classified through a pre-fine-tuned target model, wherein the target model is fine-tuned by the model federation fine-tuning method provided in the first aspect of the present disclosure.
[0024] In a fifth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the model federation fine-tuning method provided in the first aspect of the present disclosure or the steps of the text classification method provided in the second aspect of the present disclosure.
[0025] In a sixth aspect, an embodiment of the present disclosure provides an electronic device, including:
[0026] a storage device having a computer program stored thereon;
[0027] A processing device is configured to execute the computer program in the storage device to implement the steps of the model federation fine-tuning method provided in the first aspect of the present disclosure or the steps of the text classification method provided in the second aspect of the present disclosure.
[0028] In the above technical solution, the pre-trained target model includes a first model deployed on the server side and a second model deployed on at least one client side, that is, only a part of the target model (i.e., the second model) needs to be disclosed to each client participating in the model federation fine-tuning, which guarantees the model privacy of the server side to a certain extent. In addition, noise perturbation is only performed on the multiple word segmentations corresponding to the text samples that do not have classification utility for the text categories marked by the classification labels, reducing the perturbation of the target word segmentations that have classification utility for the text categories marked by the classification labels, wherein the target word segmentations have an important influence on the model classification performance. This adaptive perturbation mechanism provides a more clever trade-off between the availability of the target model and the privacy of client data, effectively improving the availability of the target model in downstream classification tasks while strengthening the protection of client data privacy. Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the elements and components are not necessarily drawn to scale. In the drawings:
[0030] FIG1 is a flowchart showing a model federation fine-tuning method according to an exemplary embodiment of the present disclosure.
[0031] FIG2 is a schematic diagram showing a process of a model federation fine-tuning method according to an exemplary embodiment of the present disclosure.
[0032] FIG3 is a flowchart showing a text classification method according to an exemplary embodiment of the present disclosure.
[0033] FIG4 is a block diagram showing a model federation fine-tuning apparatus according to an exemplary embodiment of the present disclosure.
[0034] FIG5 is a block diagram showing a text classification apparatus according to an exemplary embodiment of the present disclosure.
[0035] FIG6 is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] The inventors discovered that while existing MaaS systems provide clients with efficient and customizable LM services, they also pose risks of server-side model privacy leakage and client-side data privacy leakage. Specifically, because the pre-training process consumes significant resources, the PLM weights are typically considered proprietary data on the server side and cannot be directly disclosed. Furthermore, client text data often contains personally identifiable information and commercially confidential data. Directly disclosing this raw data to the server would result in significant privacy leakage, undoubtedly hindering privacy-conscious clients from using customized services.
[0037] In related technologies, to protect model privacy, the server deploys the core PLM components as a black box on a cloud server, disclosing only the embedding blocks to the client. To protect data privacy, the client adds noise perturbations to the embedding vectors of the input text and sends these perturbed embeddings to the server for subsequent model fine-tuning. Because client-side noise perturbations inevitably reduce the model's usability in downstream tasks, it's difficult to strike a good balance between model usability and data privacy.
[0038] In view of this, embodiments of the present disclosure provide a model federation fine-tuning method, a text classification method, an apparatus, a medium, and a device.
[0039] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0040] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0041] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0042] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0043] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0044] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0045] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0046] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0047] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0048] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0049] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0050] Figure 1 is a flow chart of a model federation fine-tuning method according to an exemplary embodiment of the present disclosure. In the present disclosure, in order to improve the training efficiency of the model, the target model can be pre-trained on the server side, and then the pre-trained target model can be fine-tuned through the Split-and-Privatize (SAP) federation fine-tuning framework. The SAP federation fine-tuning framework may include the above-mentioned server side and at least one client side, and the pre-trained target model is divided into a first model deployed on the server side and a second model deployed on each client side (that is, each client in the SAP federation fine-tuning framework is deployed with a second model), wherein the first model is a top-level model and the second model is a bottom-level model, that is, the output of the second model serves as the input of the first model. In one possible embodiment, the above-mentioned target model may be a language model.
[0051] For example, the SAP federated fine-tuning framework includes a server and a client, and the first model is deployed on the server, and the second model is deployed on the client.
[0052] As another example, the SAP federated fine-tuning framework includes a server and three clients. The first model is deployed on the server, and the second model is deployed on each of the three clients. The model structure and model parameters of the second model deployed on each client are the same.
[0053] In the present disclosure, the segmentation position of the target model can be determined based on the actual application scenario to divide the model into a first model and a second model. The second model of the target client includes at least an embedding block, and the embedding block includes multiple embedding layers.
[0054] In one embodiment, the second model of the target client includes an embedding block for converting word segments corresponding to the text into an embedding vector (i.e., a text representation), wherein the embedding block may include multiple embedding layers connected in series, and the first model may include an encoder and an output layer connected in series, wherein the encoder includes multiple encoding modules connected in series, the first encoding module of the multiple encoding modules connected in series is connected to the last embedding layer in the embedding block, and the last encoding module of the multiple encoding modules connected in series is connected to the output layer, wherein the encoder is used to learn a generalized representation of the text, and the output layer is constructed according to the attributes of the downstream task, and the generalized representation output by the encoder is processed into a model output (i.e., a classification prediction result) according to specific task requirements.
[0055] In another embodiment, the second model of the target client includes a series of embedding blocks and at least one encoding module, the embedding block is used to convert the word segmentation corresponding to the text into an embedding vector (i.e., text representation), wherein the embedding block may include multiple series-connected embedding layers, and the encoding module is used to learn the generalized representation of the text. When the second model of the target client includes multiple encoding modules, the multiple encoding modules are connected in series, and the first encoding module of the multiple encoding modules is connected to the last embedding layer of the embedding block; the first model may include a series of encoders and an output layer, wherein the encoder includes multiple series-connected encoding modules, the last encoding module of the multiple series-connected encoding modules is connected to the output layer, and the first encoding module of the multiple series-connected encoding modules is connected to the last encoding module of the second model, wherein the encoder is used to further learn the generalized representation of the text, and the output layer is constructed according to the attributes of the downstream task, and the generalized representation output by the encoder is processed into the model output (i.e., the classification prediction result) according to the specific task requirements.
[0056] The model federation fine-tuning method described above can be applied to a target client in at least one client in the SAP federation fine-tuning framework. The target client can be any client in the at least one client, or a client in the at least one client that meets preset conditions. As shown in FIG1 , the model federation fine-tuning method can include the following steps S101 to S105 .
[0057] In S101 , a text sample and a classification label corresponding to the text sample are obtained.
[0058] In the present disclosure, for the sake of client data privacy, text samples and classification labels can be located locally on the target client, that is, the text samples and corresponding classification labels do not leave the target client (that is, the local data shown in Figure 2), or they can be located on other storage devices.
[0059] In S102, the text sample is segmented to obtain a plurality of segmented words, and an embedding vector corresponding to each of the plurality of segmented words is generated by a second model deployed on the target client.
[0060] In the present disclosure, the text sample can be segmented according to the vocabulary to obtain multiple segmentations; then, each segmentation is input into the second model locally on the target client to obtain the embedding vector corresponding to each of the multiple segmentations (as shown in Figure 2).
[0061] In S103 , a target segmented word having classification utility for the text category marked by the classification label is determined from the plurality of segmented words.
[0062] In this disclosure, classification utility is a metric used to measure classification quality. Its purpose is to identify the word segments within each category of text samples that contribute most to the utility target (i.e., classification utility). Words that appear frequently in text samples belonging to the category annotated by the classification label, but less frequently in text samples belonging to other categories, are considered to contribute to distinguishing the text category annotated by the classification label from other categories. These words are considered to have classification utility for the text category annotated by the classification label.
[0063] In S104, noise perturbation processing is performed on the embedding vectors corresponding to the other word segments except the target word in the multiple word segments to obtain a perturbation vector.
[0064] In the present disclosure, when the second model only contains embedding blocks, the computational burden on the client is relatively small, but the server can easily recover the corresponding text content from the transmitted embedding vector through nearest neighbor search. Therefore, after generating the embedding vectors corresponding to multiple word segments, they are not sent directly to the server. Instead, these embedding vectors need to be perturbed with noise to avoid the risk of client data privacy leakage caused by the server recovering the corresponding text content from the transmitted embedding vector. However, while introducing noise perturbation strengthens data privacy protection, it also leads to a certain loss of model performance. Therefore, there is a trade-off between model utility and data privacy protection. In order to improve this trade-off, we first screen out the segmentations that have an important impact on the model classification performance from multiple segmentations, that is, the target segmentations that have classification utility for the text categories marked by the classification labels. Then, as shown in Figure 2, we only perform noise perturbation (add noise) on the segmentations that do not have classification utility for the text categories marked by the classification labels among the multiple segmentations corresponding to the text samples, and reduce the perturbation of the target segmentations that have classification utility for the text categories marked by the classification labels. In this way, we adaptively apply the privacy protection mechanism to perform privacy perturbation on the embedding vectors corresponding to each segmentation, thereby providing a more clever trade-off between the availability of the target model and data privacy.
[0065] In addition, if the second model contains more encoding modules, the text representation generated based on the second model is more abstract and general, and it becomes more difficult to recover the original text content from it.
[0066] In S105 , the model parameters of the first model are fine-tuned in collaboration with the server and other clients in the at least one client according to the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label.
[0067] In the present disclosure, as shown in FIG2 , in the process of fine-tuning the model parameters of the first model through the SAP federated fine-tuning framework, the model parameters of the second model remain fixed, and, when fine-tuning the first model, in order to reduce computational overhead, a parameter-efficient fine-tuning method can be selected, for example, freezing some parameters of the first model (for example, the frozen module in FIG2 ), and only adjusting the unfrozen parameters of the first model (for example, the adjustable module in FIG2 ).
[0068] In the above technical solution, the pre-trained target model includes a first model deployed on the server side and a second model deployed on at least one client side, that is, only a part of the target model (i.e., the second model) needs to be disclosed to each client participating in the federated fine-tuning of the model, which guarantees the privacy of the server-side model to a certain extent. In addition, noise perturbations are only performed on the multiple word segmentations corresponding to the text samples that do not have classification utility for the text categories marked by the classification labels, reducing the perturbations on the target word segmentations that have classification utility for the text categories marked by the classification labels, wherein the target word segmentations have an important impact on the model classification performance. This adaptive perturbation mechanism provides a more clever trade-off between the availability of the target model and the privacy of client data, effectively improving the availability of the target model in downstream classification tasks while strengthening the protection of client data privacy.
[0069] The following is a detailed description of an exemplary implementation of determining a target segmented word with classification utility for the text category marked with a classification label from multiple segmented words in the above S103. For example, this can be achieved by the following steps (1) to (2):
[0070] Step (1): Obtain the classification utility words corresponding to the text category marked by the classification label.
[0071] In the present disclosure, classification utility words are the K reference words in a vocabulary with the highest utility importance (UI) to the text category annotated by the classification label, where the vocabulary is used to segment text samples, and K is greater than 1. The classification results of the text classification model may include multiple preset categories, wherein classification utility words corresponding to each of the multiple preset categories may be pre-established.
[0072] Step (2): Determine the segmentation words belonging to the classification utility words among the multiple segmentations as the target segmentation words.
[0073] The following schematically illustrates a method for determining the classification utility word corresponding to the text category marked by the classification label. For example, this can be achieved through the following steps (a1) to (a3):
[0074] Step (a1): for each reference word in the vocabulary, obtain the frequency of occurrence of the reference word in the text samples under each preset category.
[0075] Step (a2): Determine the utility importance of the reference word to the text category marked by the classification label based on the frequency of occurrence of the reference word in the text samples under each preset category.
[0076] In the present disclosure, a training set on a target client includes multiple text samples, which are divided into multiple categories according to text type. For each reference word in the vocabulary, a statistical analysis is performed to obtain the frequency of occurrence of the reference word in the text samples of each preset category. Then, based on the frequency of occurrence of the reference word in the text samples of each preset category, the utility importance of the reference word to the text category marked by the classification label is determined.
[0077] Step (a3): K reference words in the vocabulary that have the highest utility importance for the text category annotated by the classification label are determined as classification utility words.
[0078] The following describes in detail an exemplary embodiment of determining the utility importance of a reference word to a text category labeled with a classification label based on the frequency of the reference word's appearance in text samples under each preset category in step (a2). For example, this can be achieved by the following steps (a21) and (a22):
[0079] Step (a21): For each other category in multiple preset categories except the text category marked by the classification label, determine the ratio of the frequency of the reference word appearing in the text sample under the text category marked by the classification label to the frequency of the reference word appearing in the text sample under the other category.
[0080] Step (a22): Determine the utility importance of the reference word to the text category marked by the classification label based on the ratio corresponding to each other category.
[0081] In a possible implementation, the sum of the logarithms of the ratios corresponding to each other category may be determined as the utility importance of the reference word to the text category annotated by the classification label.
[0082] For example, the utility importance of the reference word to the text category marked by the classification label can be determined by the following equation based on the ratio corresponding to each other category:
[0083] Among them, UI mc is the reference word t in the vocabulary m The utility importance of the text category c marked by the classification label; p(t=t m|y=c) is the reference word t m The frequency of occurrence in the text sample under category c; p(t=t m |y=c') is the reference word t m The frequency of occurrence in text samples under other categories c'; t represents the reference word, and y represents the text category.
[0084] For example, the above-mentioned multiple preset categories include C1, C2 and C3, and the text category marked by the classification label of the above-mentioned text sample is C1. Then a reference word t in the vocabulary m Utility importance for category C1
[0085] The following describes in detail the specific implementation method of performing noise perturbation processing on the embedding vectors corresponding to the other word segments except the target word in the above S104 to obtain the perturbation vector. Specifically, it can be achieved by the following steps (b1) to (b3):
[0086] Step (b1): Perform noise perturbation processing on the embedding vectors corresponding to the other word segments except the target word in multiple word segments.
[0087] In the present disclosure, as shown in FIG2 , random noise can be added to the embedding vector corresponding to each other word in a plurality of word segmentations except the target word segmentation according to a differential privacy mechanism or a probably approximately correct (PAC) privacy mechanism to perform noise perturbation processing. For example, when performing noise perturbation processing according to the differential privacy mechanism, a Gaussian mechanism or a random mechanism that satisfies dχ-privacy (a variant of local differential privacy) can be used to add random noise to the embedding vector corresponding to each other word segmentation.
[0088] For example, a random mechanism that satisfies dχ-privacy can be used to add random noise to the corresponding embedding vector by the following equation: in, Add a noise vector to the embedding vector φ(x) The vector obtained after (i.e. the embedding vector obtained after noise perturbation processing), is the noise vector The probability density function of η is the parameter in the random mechanism that satisfies dχ-privacy, which is used to control the noise size. The smaller η is, the larger the corresponding noise variance is and the stronger the privacy protection ability is.
[0089] Step (b2): for each embedding vector obtained after noise perturbation processing, determine whether the embedding vector is included in the reference embedding vectors corresponding to each reference word in the vocabulary.
[0090] Step (b3): If the reference embedding vectors corresponding to the reference words in the vocabulary do not contain the embedding vector, the reference embedding vector closest to the embedding vector is determined as the perturbation vector.
[0091] In the present disclosure, the reference embedding vectors corresponding to each reference word in the vocabulary constitute an embedding vector space. The embedding vectors obtained after noise perturbation processing may not fall in the embedding vector space. Therefore, it is necessary to replace and adjust the embedding vectors obtained after noise perturbation processing that do not fall in the embedding vector space. For example, for each embedding vector obtained after noise perturbation processing, it is first determined whether the embedding vector is included in the embedding vector space; if the embedding vector is included in the embedding vector space, it indicates that the embedding vector falls into the embedding vector space. In this case, there is no need to replace and adjust the embedding vector, and the embedding vector can be directly determined as the corresponding perturbation vector; if the embedding vector is not included in the embedding vector space, it indicates that the embedding vector does not fall into the embedding vector space. In this case, as shown in Figure 2, the reference embedding vector (i.e., the nearest neighbor reference embedding vector) in the embedding vector space that is closest to the embedding vector can be used to replace the embedding vector, that is, the reference embedding vector that is closest to the embedding vector is determined as the corresponding perturbation vector.
[0092] The following describes in detail an exemplary embodiment of fine-tuning the model parameters of the first model in collaboration with the server and other clients in at least one client based on the perturbation vector, the embedding vector corresponding to the target word segmentation, and the classification label in S105. For example, this can be achieved by the following steps (c1) and (c2).
[0093] Step (c1): The perturbation vector and the embedding vector corresponding to the target word segmentation are sent to the server, and the server generates a classification prediction result of the text sample through the first model based on the perturbation vector and the embedding vector corresponding to the target word segmentation and sends it to the target client.
[0094] Step (c2): Determine the first gradient information corresponding to the output layer of the first model based on the classification prediction results and classification labels of the text samples, and send the first gradient information to the server. The server updates the model parameters based on the first gradient information and the second gradient information corresponding to the output layer of the first model sent by other clients.
[0095] For example, as shown in Figure 2, the target client can send each perturbation vector and the embedding vector corresponding to the target word to the server. The server receives each perturbation vector and the embedding vector corresponding to the target word, then inputs them into the local first model to obtain a classification prediction result for the text sample, and sends the classification prediction result of the text sample to the target client. The target client determines the model loss of the first model based on the received classification prediction result of the text sample and the classification label of the text sample, and then determines first gradient information corresponding to the output layer of the first model based on the model loss, and sends the first gradient information to the server. When the SAP federated fine-tuning framework includes multiple clients, in addition to receiving the first gradient information sent by the target client, the server also receives second gradient information corresponding to the output layer of the first model sent by each of the multiple clients other than the target client. In this case, the server can perform backpropagation in the first model based on the average of the first gradient information and the second gradient information to update the model parameters of the first model.
[0096] Fig. 3 is a flow chart of a text classification method according to an exemplary embodiment, wherein the method can be applied to a target client in the at least one client. As shown in Fig. 3, the text classification method may include S201 and S202.
[0097] In S201, a text to be classified is obtained.
[0098] In S202 , based on the text to be classified, a classification prediction result of the text to be classified is obtained by using a pre-fine-tuned target model.
[0099] In the present disclosure, the same forward calculation steps as the above-mentioned model federation fine-tuning method can be adopted for the text to be classified to obtain the classification prediction results of the text to be classified. For example, after the text to be classified undergoes word segmentation, embedding vector calculation, and adaptive perturbation of the embedding vector, the target client sends the perturbed vector to the server, and the server then generates the classification prediction results of the text to be classified through the first model. The target word segmentation in the adaptive perturbation mechanism is the target word segmentation determined in the above-mentioned fine-tuning process, and the first model on the server is the first model after the above-mentioned fine-tuning.
[0100] In the above technical solution, the pre-trained target model includes a first model deployed on the server side and a second model deployed on at least one client side, that is, only a part of the target model (i.e., the second model) needs to be disclosed to each client participating in the federated fine-tuning of the model, which guarantees the privacy of the server-side model to a certain extent. In addition, noise perturbations are only performed on the multiple word segmentations corresponding to the text samples that do not have classification utility for the text categories marked by the classification labels, reducing the perturbations on the target word segmentations that have classification utility for the text categories marked by the classification labels, wherein the target word segmentations have an important impact on the model classification performance. This adaptive perturbation mechanism provides a more clever trade-off between the availability of the target model and the privacy of client data, effectively improving the availability of the target model in downstream classification tasks while strengthening the protection of client data privacy.
[0101] FIG4 is a block diagram of a model federation fine-tuning apparatus according to an exemplary embodiment. The pre-trained model includes a first model deployed on a server and a second model deployed on at least one client. The model federation fine-tuning apparatus 300 is applied to a target client among the at least one client and includes:
[0102] A first acquisition module 301 is configured to acquire text samples and corresponding classification labels;
[0103] The local computing module 302 is configured to segment the text sample to obtain a plurality of segmented words, and generate an embedding vector corresponding to each of the plurality of segmented words using the second model deployed on the target client;
[0104] A target segmentation determination module 303 is configured to determine a target segmentation from the multiple segmentations that has classification utility for the text category marked by the classification label;
[0105] The perturbation module 304 is configured to perform noise perturbation processing on the embedding vectors corresponding to the other word segments in the plurality of word segments except the target word segment to obtain a perturbation vector;
[0106] The fine-tuning module 305 is configured to collaboratively fine-tune the model parameters of the first model with the server and other clients among the at least one client based on the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label.
[0107] In the above technical solution, the pre-trained target model includes a first model deployed on the server side and a second model deployed on at least one client side, that is, only a part of the target model (i.e., the second model) needs to be disclosed to each client participating in the federated fine-tuning of the model, which guarantees the privacy of the server-side model to a certain extent. In addition, noise perturbations are only performed on the multiple word segmentations corresponding to the text samples that do not have classification utility for the text categories marked by the classification labels, reducing the perturbations on the target word segmentations that have classification utility for the text categories marked by the classification labels, wherein the target word segmentations have an important impact on the model classification performance. This adaptive perturbation mechanism provides a more clever trade-off between the availability of the target model and the privacy of client data, effectively improving the availability of the target model in downstream classification tasks while strengthening the protection of client data privacy.
[0108] Optionally, the target word segmentation determination module 303 includes:
[0109] an acquisition submodule configured to acquire classification utility words corresponding to the text category annotated by the classification label, wherein the classification utility words are K reference words in a vocabulary that have the highest utility importance for the text category annotated by the classification label, the vocabulary being used to segment the text sample, and K being greater than 1;
[0110] The first determination submodule is configured to determine the segmented words among the multiple segmented words that belong to the classification utility words as the target segmented words.
[0111] Optionally, the classification utility word is determined by a utility word determining device, which may include:
[0112] a third acquisition module configured to acquire, for each reference word in the vocabulary, a frequency of occurrence of the reference word in a text sample under each preset category;
[0113] a utility importance determination module configured to determine the utility importance of the reference word to the text category marked by the classification label based on the frequency of occurrence of the reference word in the text samples under each of the preset categories;
[0114] The classification utility word determination module is configured to determine K reference words in the vocabulary that have the highest utility importance to the text category marked by the classification label as the classification utility words.
[0115] Optionally, the utility importance determination module includes:
[0116] a second determination submodule configured to determine, for each of a plurality of preset categories other than the text category annotated by the classification label, a ratio of a frequency of occurrence of the reference word in text samples under the text category annotated by the classification label to a frequency of occurrence of the reference word in text samples under the other category;
[0117] The third determination submodule is configured to determine the utility importance of the reference word to the text category marked by the classification label based on the ratio corresponding to each of the other categories.
[0118] Optionally, the disturbance module 304 includes:
[0119] a processing submodule, configured to perform noise perturbation processing on the embedding vectors corresponding to the other segmentations in the plurality of segmentations except the target segmentation;
[0120] a fourth determination submodule configured to determine, for each embedding vector obtained after noise perturbation processing, whether the embedding vector is included in the reference embedding vectors corresponding to each reference word in a vocabulary, wherein the vocabulary is used to segment the text sample;
[0121] The fifth determination submodule is configured to determine the reference embedding vector closest to the embedding vector as the perturbation vector if the embedding vector is not included in the reference embedding vectors corresponding to the reference words in the vocabulary.
[0122] Optionally, the fine-tuning module 305 includes:
[0123] a sending submodule, configured to send the perturbation vector and the embedding vector corresponding to the target word to the server, so that the server generates a classification prediction result of the text sample through the first model based on the perturbation vector and the embedding vector corresponding to the target word and sends the result to the target client;
[0124] The sixth determination submodule is configured to determine the first gradient information corresponding to the output layer of the first model based on the classification prediction result of the text sample and the classification label, and send the first gradient information to the server, so that the server updates the model parameters based on the first gradient information and the second gradient information corresponding to the output layer sent by the other client.
[0125] Optionally, the second model of the target client includes at least an embedding block, and the embedding block includes multiple embedding layers.
[0126] Optionally, the second model of the target client includes an embedding block and at least one encoding module.
[0127] Optionally, the text sample and the classification label are both located locally on the target client.
[0128] It should be noted that the utility word determination device may be integrated into the model federation fine-tuning device 300 or may be independent of the model federation fine-tuning device 300 , and this disclosure does not make any specific limitation.
[0129] FIG5 is a block diagram of a text classification device according to an exemplary embodiment. As shown in FIG5 , the text classification device 400 includes:
[0130] The second acquisition module 401 is configured to acquire the text to be classified;
[0131] The classification module 402 is configured to obtain a classification prediction result of the text to be classified based on the text to be classified through a pre-fine-tuned target model, wherein the target model is fine-tuned according to the above-mentioned model federation fine-tuning method provided in the present disclosure.
[0132] In the above technical solution, the pre-trained target model includes a first model deployed on the server side and a second model deployed on at least one client side, that is, only a part of the target model (i.e., the second model) needs to be disclosed to each client participating in the federated fine-tuning of the model, which guarantees the privacy of the server-side model to a certain extent. In addition, noise perturbations are only performed on the multiple word segmentations corresponding to the text samples that do not have classification utility for the text categories marked by the classification labels, reducing the perturbations on the target word segmentations that have classification utility for the text categories marked by the classification labels, wherein the target word segmentations have an important impact on the model classification performance. This adaptive perturbation mechanism provides a more clever trade-off between the availability of the target model and the privacy of client data, effectively improving the availability of the target model in downstream classification tasks while strengthening the protection of client data privacy.
[0133] Reference is now made to FIG6 , which illustrates a schematic structural diagram of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG6 is merely an example and should not limit the functionality and scope of use of the embodiment of the present disclosure.
[0134] As shown in Figure 6, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0135] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0136] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0137] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0138] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0139] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0140] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains a text sample and a classification label corresponding to the text sample, and the pre-trained target model includes a first model deployed on the server and a second model deployed on at least one client; performs word segmentation on the text sample to obtain multiple word segments, and generates embedding vectors corresponding to each of the multiple word segments through the second model deployed on the target client; determines a target word segmentation from the multiple word segments that has classification utility for the text category marked by the classification label; performs noise perturbation processing on the embedding vectors corresponding to other word segments in the multiple word segments except the target word segmentation to obtain a perturbation vector; and collaborates with the server and other clients among the at least one client to fine-tune the model parameters of the first model based on the perturbation vector, the embedding vector corresponding to the target word segmentation and the classification label.
[0141] Alternatively, the computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the text to be classified; and obtains a classification prediction result of the text to be classified based on the text to be classified through a pre-fine-tuned target model, wherein the target model is fine-tuned according to the model federation fine-tuning method provided in the present disclosure.
[0142] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0144] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself. For example, the second acquisition module may also be described as a "module for acquiring text to be classified."
[0145] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0147] According to one or more embodiments of the present disclosure, Example 1 provides a model federation fine-tuning method, wherein a pre-trained target model includes a first model deployed on a server and a second model deployed on at least one client. The method is applied to a target client in the at least one client, including:
[0148] Obtaining a text sample and a classification label corresponding to the text sample;
[0149] Segmenting the text sample to obtain a plurality of segmented words, and generating embedding vectors corresponding to each of the plurality of segmented words using the second model deployed on the target client;
[0150] Determining a target segmentation from the multiple segmentations that has classification utility for the text category marked by the classification label;
[0151] Performing noise perturbation processing on the embedding vectors corresponding to the other word segments in the plurality of word segments except the target word segment to obtain a perturbation vector;
[0152] Fine-tune model parameters of the first model in collaboration with the server and other clients among the at least one client based on the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label.
[0153] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein determining, from the multiple segmented words, a target segmented word having classification utility for the text category annotated by the classification label includes:
[0154] Obtaining classification utility words corresponding to the text category marked by the classification label, wherein the classification utility words are K reference words in a vocabulary that have the highest utility importance for the text category marked by the classification label, the vocabulary being used to segment the text sample, and K being greater than 1;
[0155] The segmented words among the multiple segmented words that belong to the classification utility word are determined as the target segmented words.
[0156] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, wherein the classification utility word is determined by:
[0157] For each reference word in the vocabulary, obtaining a frequency of occurrence of the reference word in a text sample under each preset category; determining the utility importance of the reference word to the text category marked by the classification label based on the frequency of occurrence of the reference word in the text sample under each preset category;
[0158] K reference words in the vocabulary that have the highest utility importance to the text category marked by the classification label are determined as the classification utility words.
[0159] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 3, wherein determining the utility importance of the reference word to the text category marked by the classification label based on the frequency of occurrence of the reference word in the text sample under each preset category includes:
[0160] For each of the other categories among the plurality of preset categories except the text category marked by the classification label, determining a ratio of the frequency of the reference word appearing in the text samples under the text category marked by the classification label to the frequency of the reference word appearing in the text samples under the other category;
[0161] The utility importance of the reference word to the text category marked by the classification label is determined according to the ratio corresponding to each of the other categories.
[0162] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein performing noise perturbation processing on the embedding vectors corresponding to the other word segments in the multiple word segments except the target word segment to obtain the perturbation vector includes:
[0163] Performing noise perturbation processing on the embedding vectors corresponding to the other segmentations in the multiple segmentations except the target segmentation;
[0164] For each embedding vector obtained after noise perturbation processing, determining whether the embedding vector is included in the reference embedding vectors corresponding to each reference word in a vocabulary, wherein the vocabulary is used to segment the text sample;
[0165] If the reference embedding vectors corresponding to the reference words in the vocabulary do not include the embedding vector, the reference embedding vector that is closest to the embedding vector is determined as the perturbation vector.
[0166] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, wherein fine-tuning the model parameters of the first model in collaboration with the server and other clients among the at least one client based on the perturbation vector, the embedding vector corresponding to the target word segmentation, and the classification label includes:
[0167] Sending the perturbation vector and the embedding vector corresponding to the target word to the server, so that the server generates a classification prediction result of the text sample through the first model based on the perturbation vector and the embedding vector corresponding to the target word and sends the result to the target client;
[0168] Based on the classification prediction result of the text sample and the classification label, first gradient information corresponding to the output layer of the first model is determined, and the first gradient information is sent to the server, so that the server updates the model parameters based on the first gradient information and second gradient information corresponding to the output layer sent by the other client.
[0169] According to one or more embodiments of the present disclosure, Example 7 provides the method of any one of Examples 1-6, wherein the second model of the target client includes at least an embedding block, and the embedding block includes multiple embedding layers.
[0170] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 7, wherein the second model of the target client includes an embedding block and at least one encoding module.
[0171] According to one or more embodiments of the present disclosure, Example 9 provides the method of any one of Examples 1-6, wherein the text sample and the classification label are both located locally on the target client.
[0172] According to one or more embodiments of the present disclosure, Example 10 provides a text classification method, including:
[0173] Get the text to be classified;
[0174] Based on the text to be classified, a classification prediction result of the text to be classified is obtained by using a pre-fine-tuned target model, wherein the target model is fine-tuned according to the model federation fine-tuning method described in any one of Examples 1-9.
[0175] According to one or more embodiments of the present disclosure, Example 11 provides a model federation fine-tuning device, wherein a pre-trained target model includes a first model deployed on a server and a second model deployed on at least one client. The device is applied to a target client among the at least one client, and includes:
[0176] A first acquisition module is configured to acquire a text sample and a classification label corresponding to the text sample;
[0177] a local computing module configured to segment the text sample to obtain a plurality of segmented words, and generate an embedding vector corresponding to each of the plurality of segmented words through the second model deployed on the target client;
[0178] a target segmentation determination module configured to determine, from the plurality of segmentations, a target segmentation that has classification utility for the text category marked by the classification label;
[0179] a perturbation module configured to perform noise perturbation processing on the embedding vectors corresponding to the other word segments in the plurality of word segments except the target word segment to obtain a perturbation vector;
[0180] A fine-tuning module is configured to collaboratively fine-tune model parameters of the first model with the server and other clients among the at least one client based on the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label.
[0181] According to one or more embodiments of the present disclosure, Example 12 provides a text classification device, including:
[0182] A second acquisition module is configured to acquire the text to be classified;
[0183] A classification module is configured to obtain a classification prediction result of the text to be classified based on the text to be classified through a pre-fine-tuned target model, wherein the target model is fine-tuned according to the model federation fine-tuning method described in any one of Examples 1-9.
[0184] According to one or more embodiments of the present disclosure, Example 13 provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in any one of Examples 1-10 when the program is executed by a processing device.
[0185] According to one or more embodiments of the present disclosure, Example 14 provides an electronic device, including:
[0186] a storage device having a computer program stored thereon;
[0187] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-10.
[0188] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0189] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0190] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. A model federation fine-tuning method, wherein: The pre-trained target model includes a first model deployed on a server and a second model deployed on at least one client. The method is applied to a target client in the at least one client, including: Obtaining a text sample and a classification label corresponding to the text sample; Segmenting the text sample to obtain a plurality of segmented words, and generating embedding vectors corresponding to the plurality of segmented words respectively through a second model deployed on the target client; Determining a target segmentation from the multiple segmentations that has classification utility for the text category marked by the classification label; Performing noise perturbation processing on the embedding vectors corresponding to the other word segments among the multiple word segments except the target word segment to obtain a perturbation vector; According to the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label, collaboratively fine-tune the model parameters of the first model with the server and other clients among the at least one client.
2. The method according to claim 1, wherein: The step of determining a target segmented word having classification utility for the text category marked by the classification label from the multiple segmented words includes: Acquire classification utility words corresponding to the text category marked by the classification label, wherein the classification utility words are K reference words in a vocabulary with the highest utility importance to the text category marked by the classification label, and the vocabulary is used to segment the text sample, and K is greater than 1; The segmented words among the multiple segmented words that belong to the classification utility word are determined as the target segmented words.
3. The method according to claim 2, wherein: The classification utility word is determined in the following manner: For each reference word in the vocabulary, obtain the frequency of occurrence of the reference word in the text sample under each preset category; determine the utility importance of the reference word to the text category marked by the classification label according to the frequency of occurrence of the reference word in the text sample under each preset category; The K reference words in the vocabulary that have the highest utility importance to the text category marked by the classification label are determined as the classification utility words.
4. The method according to claim 3, wherein: The determining, based on the frequency of occurrence of the reference word in the text samples under each of the preset categories, the utility importance of the reference word to the text category marked by the classification label includes: For each other category among multiple preset categories except the text category marked by the classification label, determine the ratio of the frequency of the reference word appearing in the text sample under the text category marked by the classification label to the frequency of the reference word appearing in the text sample under the other category; The utility importance of the reference word to the text category marked by the classification label is determined according to the ratio corresponding to each of the other categories.
5. The method according to claim 1, wherein: The performing noise perturbation processing on the embedding vectors corresponding to the other word segments in the multiple word segments except the target word segment to obtain the perturbation vector includes: Performing noise perturbation processing on the embedding vectors corresponding to the other word segments among the multiple word segments except the target word segment; For each embedding vector obtained after the noise disturbance processing, determining whether the embedding vector is included in the reference embedding vectors corresponding to each reference word in the vocabulary, wherein the vocabulary is used to segment the text sample; In response to the fact that the reference embedding vectors corresponding to the reference words in the vocabulary do not include the embedding vector, the reference embedding vector that is closest to the embedding vector is determined as the disturbance vector.
6. The method according to claim 1, wherein: The fine-tuning of the model parameters of the first model in collaboration with the server and other clients among the at least one client according to the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label includes: The perturbation vector and the embedding vector corresponding to the target word segment are sent to the server, so that the server generates a classification prediction result of the text sample through the first model according to the perturbation vector and the embedding vector corresponding to the target word segment and sends the result to the target client; According to the classification prediction result of the text sample and the classification label, first gradient information corresponding to the output layer of the first model is determined, and the first gradient information is sent to the server, so that the server updates the model parameters according to the first gradient information and the second gradient information corresponding to the output layer sent by the other client.
7. The method according to any one of claims 1 to 6, wherein: The second model of the target client includes at least an embedding block, and the embedding block includes a plurality of embedding layers.
8. The method according to claim 7, wherein: The second model of the target client includes an embedding block and at least one encoding module.
9. The method according to any one of claims 1 to 6, wherein: The text sample and the classification label are both located locally on the target client.
10. A text classification method, comprising: Get the text to be classified; According to the text to be classified, a classification prediction result of the text to be classified is obtained by using a pre-fine-tuned model, wherein the model is obtained according to the model federation fine-tuning method according to any one of claims 1-9.
11. A model federation fine-tuning device, wherein: The pre-trained target model includes a first model deployed on a server and a second model deployed on at least one client, and the device is applied to a target client in the at least one client, including: A first acquisition module is configured to acquire a text sample and a classification label corresponding to the text sample; A local computing module is configured to segment the text sample to obtain a plurality of segmented words, and generate an embedding vector corresponding to each of the plurality of segmented words through the second model deployed on the target client; a target segmentation determination module, configured to determine, from the plurality of segmentations, a target segmentation that has classification utility for the text category marked by the classification label; A disturbance module is configured to perform noise disturbance processing on the embedding vectors corresponding to the other word segments in the multiple word segments except the target word segment to obtain a disturbance vector; A fine-tuning module is configured to fine-tune the model parameters of the first model in collaboration with the server and other clients among the at least one client according to the disturbance vector, the embedding vector corresponding to the target word segmentation, and the classification label.
12. A text classification device, comprising: A second acquisition module is configured to acquire the text to be classified; A classification module is configured to obtain a classification prediction result of the text to be classified according to the text to be classified through a pre-fine-tuned target model, wherein the target model is fine-tuned according to the model federation fine-tuning method described in any one of claims 1-9.
13. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processing device, the steps of the method described in any one of claims 1 to 10 are implemented.
14. An electronic device comprising: a storage device having a computer program stored thereon; A processing device configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Federal learning classification model training method based on model disturbance
CN115358418A
Federal learning-based model training method and federal learning system
CN115719092A
Federal learning training method, device, system and equipment based on differential privacy
CN115983409A
Fine tuning method and device for federal learning large model
CN117056962A
Model federation fine tuning method and device, text classification method and device, medium and equipment
CN117744145A