Model training method, device, equipment, storage medium and program product

By pre-training the deep neural network model and fusing multiple noise distributions during the training process, the instability problem of simulated in-memory computing in the changing environment of the device is solved, and the robustness and adaptability of the model are improved.

CN118968219BActive Publication Date: 2025-06-06SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411037429.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-06-06
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

The simulated in-memory computing has instability in the device changing environment, resulting in a reduction in inference accuracy of deep neural networks. The existing noise injection methods are difficult to adapt to changes in different devices, and their versatility is poor.

Method used

By pre-training the deep neural network model, and determining the noise distribution based on multiple noise intensities and the number of sub-training tasks during the training process, the global model and sub-training tasks are trained multiple rounds, the global model is updated, and multiple noises are fused.

Benefits of technology

Improves the robustness and performance of the model in device-changing environments, and enhances the adaptability and versatility in different hardware environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968219B_ABST
    Figure CN118968219B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a model training method, apparatus, device, storage medium and program product. The method includes: pre-training a set deep neural network model based on a set image data set to obtain a pre-trained deep neural network model, and using it as a global model and a model for each sub-training task; determining a target noise intensity based on multiple noise intensities; determining the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks; performing multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model. The disclosed embodiments can make the trained global model have stronger robustness and performance in a device change environment, as well as better adaptability and versatility in different hardware environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model training method, device, equipment, storage medium and program product. Background Art

[0002] In recent years, deep neural networks have made significant breakthroughs in various fields. With the emergence of large language models such as GPT, deep learning technology is gradually changing the way we work. At the same time, the computing requirements for training and reasoning are also increasing dramatically, which poses a huge challenge to edge computing platforms with limited computing resources and strict power consumption constraints. With the emergence of emerging non-volatile memory devices, analog memory computing has shown great potential in accelerating deep learning systems and reducing power consumption. Compared with the standard von Neumann architecture, it can directly perform matrix-vector multiplication required for deep neural networks at the weight data storage point, thereby reducing the frequent data transmission between memory and computing units, breaking the limitation of the "memory wall". However, the performance of analog computing can be affected by manufacturing defects, device variations, and other physical factors, resulting in possible deviations between the actual calculation results and the expected values, thereby reducing the reasoning accuracy of deep neural networks.

[0003] Analog in-memory computing is an emerging technology trend for accelerating deep neural networks on edge platforms. However, the instability caused by manufacturing and device variations limits its accuracy and reliability. Although existing solutions can achieve the perception of hardware noise by injecting noise during the training phase of the model, it has enhanced the robustness of deep neural networks in the face of device changes to a certain extent. However, such methods usually rely on simulating real specific hardware noise. At the same time, a single type of noise injection is difficult to adapt to the changes of different devices, resulting in problems such as poor versatility and unsatisfactory results. Summary of the invention

[0004] The present invention provides a model training method, device, equipment, storage medium and program product, so that the trained model has stronger robustness and performance in a device change environment, and has better adaptability and versatility in different hardware environments.

[0005] In a first aspect, an embodiment of the present disclosure provides a model training method, comprising: pre-training a set deep neural network model based on a set image data set to obtain a pre-trained deep neural network model, and using the pre-trained deep neural network model as a global model and a model for each sub-training task; determining a target noise intensity based on multiple noise intensities; determining a noise distribution for each sub-training task based on the target noise intensity and a set number of sub-training tasks; performing multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task, and the weight parameters of the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training.

[0006] In the second aspect, the embodiments of the present disclosure also provide a model training device, including: a pre-training module, which is used to pre-train a set deep neural network model based on a set image data set, obtain a pre-trained deep neural network model, and use it as a global model and a model of each sub-training task; a target noise intensity determination module, which is used to determine the target noise intensity based on multiple noise intensities; a sub-training task noise distribution determination module, which is used to determine the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks; a global model training module, which is used to perform multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training.

[0007] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the model training method as described in the embodiment of the present disclosure.

[0008] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the model training method as described in the embodiment of the present disclosure.

[0009] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program, which, when executed by a processor, implements the model training method as described in the embodiment of the present disclosure.

[0010] The technical solution of the embodiment of the present disclosure is to pre-train a set deep neural network model based on a set image data set to obtain a pre-trained deep neural network model and use it as a global model and a model of each sub-training task; determine a target noise intensity based on multiple noise intensities; determine the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks; perform multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training. In the disclosed embodiments, each sub-training task and the global model are trained multiple times based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model. Since the global model aggregates the weight parameters of each sub-training task and the noise distribution of each sub-training task is different, the global model integrates multiple noises, so that the trained model has stronger robustness and performance in a device change environment, as well as better adaptability and versatility in different hardware environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0012] Figure 1 A schematic diagram of a model training method flow chart provided by an embodiment of the present invention;

[0013] Figure 2 A schematic diagram of a multi-noise fusion injection training strategy provided by an embodiment of the present invention;

[0014] Figure 3 A schematic diagram of the structure of a model training device provided by an embodiment of the present invention;

[0015] Figure 4 The present invention is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0017] It should be understood that the various steps recorded in the method implementation of the present disclosure can be performed in different orders and / or in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. The term "including" and its variations used herein are open inclusions, that is, "including but not limited to". It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more". It is understandable that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of the corresponding laws, regulations and relevant provisions.

[0018] Figure 1 The present invention provides a flow chart of a model training method. The present invention is applicable to the model training scenario, so that the trained model has better adaptability and reasoning performance in hardware environments with different noises. The method can be executed by a model training device, which can be implemented in the form of software and / or hardware, and optionally, by an electronic device, which can be a mobile terminal, a PC, or a server. Figure 1 As shown, the method includes:

[0019] S110. Pre-train a set deep neural network model based on a set image data set to obtain a pre-trained deep neural network model, and use it as a global model and a model for each sub-training task.

[0020] In this embodiment, there is no limitation on the specific image dataset, for example, it can be a CIFAR10 image recognition dataset, an MNIST image recognition dataset, etc. In this embodiment, there is no limitation on the specific set deep neural network model, for example, it can be a convolutional neural network, a recurrent neural network, a long short-term memory network, a Transformer, etc.

[0021] In this embodiment, the set image data set can be used to perform multiple rounds of pre-training on the set deep neural network model in the digital domain to obtain the pre-trained deep neural network model. At the same time, the pre-trained deep neural network model can also be used as a global model in subsequent steps and as a model for each sub-training task in subsequent steps.

[0022] S120: Determine a target noise intensity based on multiple noise intensities.

[0023] In this embodiment, a single noise (such as Gaussian noise) can be injected into the pre-trained deep neural network model for retraining. Through testing, the impact of different noise intensities (Gaussian noise intensities) on the performance of the pre-trained deep neural network model is explored, and the noise intensity with the least impact and the best performance is selected as the target noise intensity.

[0024] Optionally, the target noise intensity is determined based on multiple noise intensities, including: testing the pre-trained deep neural network model based on the multiple noise intensities respectively to obtain multiple inference accuracies corresponding to the multiple noise intensities; determining multiple precision errors based on the multiple inference accuracies and set inference precision; wherein the set inference precision is obtained by the pre-trained deep neural network model without noise injection; if the minimum precision error among the multiple precision errors falls within the set error range, the corresponding noise intensity is used as the target noise intensity.

[0025] In this embodiment, the pre-trained deep neural network model is tested with a variety of different noise intensities (a single noise test is performed each time) to obtain a variety of inference accuracies corresponding to the various noise intensities. The inference accuracy of the pre-trained deep neural network model without noise injection is obtained, and used as the set inference accuracy. The precision errors between various inference accuracies and the set inference accuracy are determined respectively to obtain a variety of precision errors. If the minimum precision error among the multiple precision errors falls within the set error range, the corresponding noise intensity is used as the target noise intensity, that is, the maximum threshold of the noise intensity. In this embodiment, the specific set error range is not limited, for example, it can be between 0 and 10.

[0026] S130: Determine the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks.

[0027] It should be noted that with the different strengths of injected noise, there is a trade-off between the performance and robustness of the model, that is, the higher the strength of the injected noise, the stronger the robustness, but the performance will decrease accordingly. The more sub-training tasks there are, the more comprehensive the model's perception of noise is, but it will bring additional computational and time costs. The set noise distribution set can be determined according to the target noise intensity and the set number of sub-training tasks, that is, the set Gaussian noise distribution set. The set noise distribution set includes the noise distribution of each sub-training task. The noise distribution can be composed of a mean (the default value can be 0) and a square of the variance. The variance of each sub-training task is different, and the corresponding noise distribution of each training task is different. The variance of each sub-training task can be obtained based on the target noise intensity, the set number of sub-training tasks, and the index value of the current sub-training task.

[0028] Optionally, the noise distribution of each sub-training task is determined based on the target noise intensity and the set number of sub-training tasks, including: for each sub-training task, determining a sub-training task ratio based on the index value of the sub-training task and the set number of sub-training tasks; and determining the noise distribution corresponding to the sub-training task based on the target noise intensity and the sub-training task ratio.

[0029] For example, the number of sub-training tasks is set to K, the index value of the sub-training task (identifying the sub-training task) is set to k, the target noise intensity is set to σ, and the variance of each sub-training task is set to σ. k For example, in, is the ratio of sub-training tasks. The noise distribution corresponding to the sub-training task can be N(0,σ 2 k ), the noise distribution set can be For example, when k is 1, N(0,σ 2 1 ) represents the noise distribution of the first sub-training task.

[0030] S140, performing multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task, and the weight parameters of the global model to obtain a trained global model.

[0031] After each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training.

[0032] In this embodiment, in each round of training, the global model distributes its own weight parameters to each sub-training task, each sub-training task receives the weight parameters distributed by the global model, performs training according to the received weight parameters, and updates the received weight parameters. After each round of training of each sub-training task is completed, the weight parameters of each sub-training task are aggregated to obtain the aggregated weight parameters, and the weight parameters of the global model are updated to the aggregated weight parameters. The global model then distributes the updated weight parameters to each sub-training task, and so on, gradually optimizing the weight parameters of the global model until the training is stopped, and the trained global model is obtained, that is, the latest weight parameters of the global model are obtained.

[0033] Optionally, multiple rounds of training are performed on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model, including: during initial training, the weight parameters of the global model are distributed to each sub-training task; for each round of training, each sub-training task samples a noise instance from its own noise distribution, and each sub-training task obtains its own noise instance; during forward propagation, the respective noise instance is injected into the weight parameters of the respective sub-training task; each sub-training task determines its own loss value based on its own set target loss function; each sub-training task updates its own weight parameter based on its own loss value; after each round of each sub-training task is completed, the updated weight parameters of each sub-training task are aggregated to obtain the aggregated weight parameters, and updated to the weight parameters of the global model; the updated weight parameters of the global model are distributed to each sub-training task to continue the next round of training until the number of training rounds is greater than or equal to the set round number threshold, and the training is stopped to obtain the trained global model.

[0034] In this embodiment, each sub-training task has the same target loss function, but different loss values. In this embodiment, there is no restriction on the target loss function, such as mean square error, cross entropy and other loss functions. In this embodiment, each sub-training task determines its own loss value based on its own set target loss function. After each sub-training task updates its own weight parameter based on its own loss value, it can be understood that one round is over.

[0035] Exemplarily, during the initial training, the weight parameters of the global model (the weight parameters at this time are the weight parameters of the pre-trained deep neural network model) are distributed to each sub-training task (the model used by each sub-training task is the pre-trained deep neural network model); in each round of training, for each sub-training task k, each sub-training task is distributed from its own noise distribution The sampled noise instance Δ∈ k, each sub-training task obtains its own noise instance; in the forward propagation process, the respective noise instance (that is, the specific value of the noise intensity) is injected into the weight parameter of the respective sub-training task (that is, the noise instance is injected into the weight parameter w of each layer of the model k In, that is, w k +Δ∈ k ); each sub-training task calculates its own loss value based on its own set target loss function; each sub-training task updates its own weight parameter based on its own loss value and stochastic gradient descent method; after each round, the updated weight parameters of each sub-training task are aggregated to obtain the aggregated weight parameters, and updated as the weight parameters of the global model; for example, Among them, w t+1 is the weight parameter of the global model in the t+1th round, is the updated weight parameter of the kth sub-training task in the t+1th round. The updated weight parameter of the global model is then distributed to each sub-training task to continue the next round of training until the number of training rounds is greater than or equal to the set round number threshold, and the training is stopped to obtain the trained global model. In this embodiment, there is no restriction on the set round number threshold, for example, it can be 50 rounds, 80 rounds, etc.

[0036] like Figure 2 As shown, Figure 2 Schematic diagram of the multi-noise fusion injection training strategy provided by the embodiment of the present invention. In each round of training, the global model distributes the global weight w to each sub-training task, and each sub-training task uses the received global weight w as its own weight. The sampled noise instance Δ∈ k , each sub-training task obtains its own noise instance, that is, the noise intensity of each sub-training task is different. Among them, the noise distribution of the first sub-training task is for The noise distribution of the Kth sub-training task is In the forward propagation process, each noise instance (specific value of noise intensity) is injected into the weight parameter of each sub-training task; each sub-training task calculates its own loss value based on its own set target loss function; each sub-training task updates its own weight parameter based on its own loss value and stochastic gradient descent method; among them, the weight parameter of the first sub-training task is weight w 1 , the weight parameter of the second sub-training task is weight w 2 , the weight parameter of the K-1th sub-training task is weight w K-1 , the weight parameter of the Kth sub-training task is weight w KAfter each round of sub-training tasks, the updated weight parameters of each sub-training task are aggregated ( Figure 2 The ∑ in the above formula can be regarded as an aggregation operation), and the aggregated weight parameters (which can be understood as the weight parameters of multi-noise fusion) are obtained, and updated as the weight parameters of the global model; the updated weight parameters of the global model are then distributed to each sub-training task to continue the next round of training, until the number of training rounds is greater than or equal to the set round number threshold, the training is stopped, and the trained global model is obtained.

[0037] In this embodiment, the global model summarizes the generalized knowledge learned from multiple noise instance injection sub-training tasks in each round, thereby obtaining a flatter global solution space, so that it can exhibit strong robustness under different noise intensities. That is, by summarizing the generalized knowledge learned from multiple local training processes, a more robust global model can be obtained.

[0038] Optionally, the updated weight parameters of each sub-training task are aggregated to obtain an aggregated weight parameter, including: accumulating and summing the updated weight parameters of each sub-training task to obtain an accumulated weight parameter; and determining the aggregated weight parameter based on the accumulated weight parameter and the set number of sub-training tasks.

[0039] Exemplarily, the weight parameter formula of the aggregation is as follows:

[0040]

[0041] Where w represents the weight parameter of the aggregation, K is the number of sub-training tasks set, k represents the kth sub-training task, and w k Represents the updated weight parameters of the kth sub-training task.

[0042] Optionally, after performing multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain the trained global model, it also includes: deploying the trained global model to a simulated in-memory computing accelerator to evaluate the inference accuracy of the trained global model.

[0043] In this embodiment, after obtaining the trained global model, the global model trained by multi-noise fusion injection can be deployed on the inference platform for testing (the input is the image data set, and the output is the inference accuracy). The inference platform can be an analog in-memory computing accelerator, which can test the accuracy recovery effect under different noise environments to evaluate the inference accuracy of the trained global model.

[0044] The technical solution of the embodiment of the present disclosure is to pre-train a set deep neural network model based on a set image data set to obtain a pre-trained deep neural network model and use it as a global model and a model of each sub-training task; determine a target noise intensity based on multiple noise intensities; determine the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks; perform multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training. In the disclosed embodiments, each sub-training task and the global model are trained multiple times based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model. Since the global model aggregates the weight parameters of each sub-training task and the noise distribution of each sub-training task is different, the global model integrates multiple noises, so that the trained model has stronger robustness and performance in a device change environment, as well as better adaptability and versatility in different hardware environments.

[0045] It should be noted that by assigning different noise distributions to different sub-training tasks and continuously aggregating the weight parameters of the sub-training tasks during the training process to update the global model, it is equivalent to simulating the noise and variation environments that may be encountered in various practical applications. This training method enables the global model to maintain stable performance in the face of various noise interferences, thereby significantly enhancing the robustness of the model.

[0046] Since each sub-training task focuses on optimizing its own model under a specific noise distribution, and the global model continuously improves itself by aggregating the optimization results of each sub-training task, this training method can find better model parameters in different noise environments. This means that the overall performance of the trained global model will be improved in various environments, especially in scenarios with device changes or inconsistent hardware conditions, the model's performance will be more stable and excellent.

[0047] Since the global model is formed based on the integration of multiple noises and sub-training task optimization, it is not only suitable for a specific hardware or environmental condition, but can better adapt to a variety of different hardware and application scenarios. In other words, it has the versatility and adaptability across environments and platforms.

[0048] The technical solution of the disclosed embodiment seems to require training multiple sub-training tasks at the same time, but in fact, through reasonable parallel processing and resource allocation strategies, computing resources can be effectively utilized to shorten the overall training time. In addition, since the global model is continuously optimized based on multiple sub-training tasks, its training effect is often better than training a single model, thus achieving a win-win situation in resource utilization and model performance.

[0049] The technical solution of the disclosed embodiment can achieve better accuracy recovery effect than the prior art. At the same time, it does not need to rely on hardware simulation and can have good versatility between reasoning platforms based on different devices. At the same time, it has strong flexible scalability (can be customized according to noise type and injection position), and can be widely used in reasoning platforms based on different devices.

[0050] Figure 3 A schematic diagram of the structure of a model training device provided in an embodiment of the present disclosure is shown in FIG. Figure 3 As shown, the device includes: a pre-training module 310, a target noise intensity determination module 320, a sub-training task noise distribution determination module 330 and a global model training module 340;

[0051] A pre-training module 310 is used to pre-train a set deep neural network model based on a set image data set, obtain a pre-trained deep neural network model, and use it as a global model and a model for each sub-training task;

[0052] A target noise intensity determination module 320, configured to determine a target noise intensity based on multiple noise intensities;

[0053] A sub-training task noise distribution determination module 330, configured to determine the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks;

[0054] The global model training module 340 is used to perform multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training.

[0055] The technical solution of the embodiment of the present disclosure is as follows: a pre-training module is used to pre-train a set deep neural network model based on a set image data set to obtain a pre-trained deep neural network model and use it as a global model and a model of each sub-training task; a target noise intensity is determined based on multiple noise intensities by a target noise intensity determination module; a sub-training task noise distribution determination module is used to determine the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks; a global model training module is used to perform multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training. In the disclosed embodiments, each sub-training task and the global model are trained multiple times based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model. Since the global model aggregates the weight parameters of each sub-training task and the noise distribution of each sub-training task is different, the global model integrates multiple noises, so that the trained model has stronger robustness and performance in a device change environment, as well as better adaptability and versatility in different hardware environments.

[0056] Optionally, the target noise intensity determination module is specifically used to: test the pre-trained deep neural network model based on multiple noise intensities, respectively, to obtain multiple inference accuracies corresponding to the multiple noise intensities; determine multiple precision errors based on the multiple inference accuracies and set inference precision; wherein the set inference precision is obtained by the pre-trained deep neural network model without noise injection; if the minimum precision error among the multiple precision errors falls within the set error range, the corresponding noise intensity is used as the target noise intensity.

[0057] Optionally, the sub-training task noise distribution determination module is specifically used to: for each sub-training task, determine the sub-training task ratio based on the index value of the sub-training task and the set number of sub-training tasks; and determine the noise distribution corresponding to the sub-training task based on the target noise intensity and the sub-training task ratio.

[0058] Optionally, the global model training module is specifically used to: during the initial training, distribute the weight parameters of the global model to each sub-training task; for each round of training, each sub-training task samples a noise instance from its own noise distribution, and each sub-training task obtains its own noise instance; during the forward propagation process, inject the respective noise instance into the weight parameters of the respective sub-training task; each sub-training task determines its own loss value based on its own set target loss function; each sub-training task updates its own weight parameter based on its own loss value; after each round of each sub-training task is completed, aggregate the updated weight parameters of each sub-training task to obtain the aggregated weight parameters, and update them to the weight parameters of the global model; distribute the updated weight parameters of the global model to each sub-training task to continue the next round of training until the number of training rounds is greater than or equal to the set round number threshold, stop training, and obtain the trained global model.

[0059] Optionally, the global model training module is also used to: accumulate and sum the updated weight parameters of each sub-training task to obtain an accumulated weight parameter; and determine the aggregated weight parameter based on the accumulated weight parameter and the set number of sub-training tasks.

[0060] Optionally, the above-mentioned device also includes a deployment module, which is used to deploy the trained global model to an analog in-memory computing accelerator to evaluate the reasoning accuracy of the trained global model.

[0061] The model training device provided in the embodiments of the present disclosure can execute the model training method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0062] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0063] like Figure 4As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0064] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0065] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 11 executes the various methods and processes described above, such as method model training.

[0066] In some embodiments, method model training may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method model training described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform method model training in any other suitable manner (e.g., by means of firmware).

[0067] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0068] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0069] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0070] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0071] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0072] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0073] An embodiment of the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the model training method provided in any embodiment of the present application.

[0074] In the process of implementation, the computer program product can be written in one or more programming languages ​​or a combination thereof to perform the computer program code of the present invention, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0075] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A model training method, characterized in that: include: Pre-train the set deep neural network model based on the set image data set to obtain the pre-trained deep neural network model, and use it as the global model and the model of each sub-training task; determining a target noise intensity based on the plurality of noise intensities; Determining the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks; Based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model, multiple rounds of training are performed on each sub-training task and the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training; The target noise intensity is determined based on multiple noise intensities, including: The pre-trained deep neural network model is tested based on various noise intensities to obtain various inference accuracies corresponding to various noise intensities; Determining a plurality of precision errors based on the plurality of inference precisions and a set inference precision; wherein the set inference precision is obtained by a pre-trained deep neural network model without injecting noise; If the minimum precision error among the multiple precision errors falls within the set error range, the corresponding noise intensity is used as the target noise intensity; Wherein, determining the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks includes: For each sub-training task, determining a sub-training task ratio based on an index value of the sub-training task and the set number of sub-training tasks; A noise distribution corresponding to the sub-training task is determined based on the target noise intensity and the sub-training task ratio.

2. The method according to claim 1, characterized in that Based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model, multiple rounds of training are performed on each sub-training task and the global model to obtain a trained global model, including: During initial training, the weight parameters of the global model are distributed to each sub-training task; For each round of training, each sub-training task samples noise instances from its own noise distribution, and each sub-training task obtains its own noise instance; In the forward propagation process, injecting the respective noise instances into the weight parameters of the respective sub-training tasks; Each sub-training task determines its own loss value based on its own set target loss function; Each sub-training task updates its own weight parameters based on its own loss value; After each round of each sub-training task is completed, the updated weight parameters of each sub-training task are aggregated to obtain aggregated weight parameters, and updated as weight parameters of the global model; The updated weight parameters of the global model are distributed to each sub-training task to continue the next round of training until the number of training rounds is greater than or equal to the set round number threshold, and the training is stopped to obtain the trained global model.

3. The method according to claim 2, characterized in that Aggregate the updated weight parameters of each sub-training task to obtain the aggregated weight parameters, including: The updated weight parameters of each sub-training task are accumulated and summed to obtain an accumulated weight parameter; An aggregated weight parameter is determined based on the accumulated weight parameter and the set number of sub-training tasks.

4. The method according to claim 1, characterized in that: After performing multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain the trained global model, the method further includes: The trained global model is deployed on a simulated in-memory computing accelerator to evaluate the inference accuracy of the trained global model.

5. A model training device, characterized in that: include: A pre-training module is used to pre-train a set deep neural network model based on a set image data set, obtain the pre-trained deep neural network model, and use it as the global model and the model of each sub-training task; A target noise intensity determination module, used to determine the target noise intensity based on multiple noise intensities; A sub-training task noise distribution determination module, used to determine the noise distribution of each sub-training task based on the target noise intensity and the set number of sub-training tasks; A global model training module, used to perform multiple rounds of training on each sub-training task and the global model based on the noise distribution of each sub-training task, the weight parameters of each sub-training task and the weight parameters of the global model to obtain a trained global model; wherein, after each round of training, the weight parameters of each sub-training task are aggregated to update the weight parameters of the global model, and each sub-training task receives the weight parameters distributed by the global model to continue the next round of training; The target noise intensity determination module is further used to: test the pre-trained deep neural network model based on multiple noise intensities respectively to obtain multiple reasoning accuracies corresponding to the multiple noise intensities; Determining a plurality of precision errors based on the plurality of inference precisions and a set inference precision; wherein the set inference precision is obtained by a pre-trained deep neural network model without injecting noise; If the minimum precision error among the multiple precision errors falls within the set error range, the corresponding noise intensity is used as the target noise intensity; The sub-training task noise distribution determination module is further used to: for each sub-training task, determine a sub-training task ratio based on an index value of the sub-training task and the set number of sub-training tasks; A noise distribution corresponding to the sub-training task is determined based on the target noise intensity and the sub-training task ratio.

6. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the model training method as described in any one of claims 1-4.

7. A storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the model training method as described in any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that When executed by a processor, the computer program implements the model training method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Federal learning method fusing adaptive weight distribution and personalized differential privacy

    CN114595831A

  • Federal learning method and device, equipment and storage medium

    CN117829315A