Data processing method and electronic device based on memristor array
By simulating faults, the deep learning model is trained, the fault injection neural network model is obtained, and the conductance value of the memristor array is adjusted according to the model, which solves the problem of accuracy reduction caused by the actual memristor failure and realizes high-precision calculation in the presence of a fault.
Patent Information
- Application Number
- CN202211504996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-11-28
AI Technical Summary
Due to process manufacturing problems, the memristor may have actual failures, resulting in changes in the conductance value of the memristor array, which in turn affects the neural network model parameters, resulting in classification errors in the neuromorphic system and reduced overall accuracy.
The deep learning model is trained by simulated faults, and the fault injection neural network model is obtained, and the conductance value of the memristor array is adjusted according to the model to maintain calculation accuracy when there is an actual fault.
In the case of actual failure, through parameter adjustment of the fault injection neural network model, the memristor array can maintain high computational accuracy and improve the robustness of the neuromorphic system.
Smart Images

Figure CN115730649B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular to a data processing method and an electronic device. Background Art
[0002] At present, artificial intelligence technology represented by neural networks has developed rapidly in various fields and achieved great success. However, with the explosive growth of data, neuromorphic computing systems (NCS) corresponding to the von Neumann structure are facing problems such as power consumption and "storage wall" that are not conducive to practical use, which restricts the development of NCS. Therefore, a new computing architecture is needed to meet the performance requirements of neuromorphic systems in terms of power consumption, latency, storage, etc.
[0003] Memristors, on the one hand, can combine storage and computing based on neuromorphic computing of memristor arrays, eliminating the frequent data transmission process between the memory and the computing unit, thereby reducing system power consumption and shortening computing delays; on the other hand, they also have the advantages of simple structure, good miniaturization, easy three-dimensional integration, and compatibility with existing CMOS processes. This allows memristors to build high-efficiency, large-scale NCSs and has broad application prospects.
[0004] However, due to current manufacturing process problems, memristors may have actual faults, causing the conductance value on the memristor array constructed by memristors to change, and then the model parameters of the neural network corresponding to the conductance value of the memristor array to change, which may eventually lead to NCS classification errors and reduce the overall accuracy of NCS. Summary of the invention
[0005] In view of this, the present disclosure provides a data processing method and an electronic device, in order to partially solve at least one of the above-mentioned technical problems.
[0006] One aspect of the present disclosure provides a data processing method.
[0007] According to an embodiment of the present disclosure, the data processing method includes:
[0008] Acquire multimedia data; process the multimedia data using a memristor array to obtain data processing results; wherein the memristor array includes a plurality of memristors, and the conductance values of the plurality of memristors are determined according to model parameters of a fault injection neural network model, the fault injection neural network model is obtained by training a deep learning model using simulated faults, and the simulated faults are determined according to actual faults of the memristor array.
[0009] According to an embodiment of the present disclosure, the above-mentioned fault injection neural network model is obtained by training a deep learning model according to a simulated fault, including repeatedly performing the following operations until the training parameters meet a preset end condition, and determining the deep learning model obtained when the above-mentioned preset end condition is met as the above-mentioned fault injection neural network model:
[0010] According to the simulated faults in the tth round and the model parameters of the deep learning model in the t-1th round, the model parameters to be adjusted of the deep learning model in the tth round are determined, wherein t is an integer greater than 1; the model parameters to be adjusted of the deep learning model in the tth round are adjusted using the sample multimedia data to obtain the model parameters of the deep learning model in the tth round.
[0011] According to an embodiment of the present disclosure, determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round includes:
[0012] According to the simulated fault of the tth round, determine the sampling point set of the tth round from the nodes of the deep learning model of the t-1th round, wherein the sampling point set includes at least one node; according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round, determine the sampling parameter set of the tth round, wherein the sampling parameter set includes sampling parameters corresponding to the sampling point set of the tth round; use the sampling parameter set of the tth round to update the model parameters of the deep learning model of the t-1th round, and obtain the model parameters to be adjusted of the deep learning model of the tth round.
[0013] According to an embodiment of the present disclosure, the method of adjusting the model parameters to be adjusted of the deep learning model of the tth round by using the sample multimedia data to obtain the model parameters of the deep learning model of the tth round includes:
[0014] The sample multimedia data is processed by the deep learning model of the tth round to obtain the sample processing result of the tth round; based on the loss function, the loss value of the tth round is obtained according to the sample processing result of the tth round; the model parameters to be adjusted of the deep learning model of the tth round are adjusted according to the loss value of the tth round to obtain the model parameters of the deep learning model of the tth round.
[0015] According to an embodiment of the present disclosure, the data processing method further includes:
[0016] The model parameters corresponding to the target node of the deep learning model of the tth round are restored to the model parameters of the deep learning model of the t-1th round; wherein the target node of the tth round is a node determined according to the sampling point set of the tth round.
[0017] According to an embodiment of the present disclosure, determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round includes:
[0018] In the case where the simulated fault in the t-1th round is the first simulated fault, the nodes of the deep learning model in the t-1th round are determined as the first sampling point set in the t-1th round; the first sampling parameter set in the t-1th round is determined according to the first simulated fault, the model parameters of the deep learning model in the t-1th round and the preset cutoff value, wherein the preset cutoff value is obtained according to the initial model parameters of the deep learning model; the model parameters of the deep learning model in the t-1th round are updated using the first sampling parameter set in the t-1th round, and the model parameters to be adjusted of the deep learning model in the t-1th round are determined.
[0019] According to an embodiment of the present disclosure, determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round includes:
[0020] In the case where the simulated fault in the t-th round is the second simulated fault, determine the first sampling probability value according to the second simulated fault; determine the second sampling point set for the t-th round from the nodes of the deep learning model for the t-1th round according to the first sampling probability value; determine the second sampling parameter set for the t-th round according to the preset cutoff value, wherein the preset cutoff value is obtained according to the initial model parameters of the deep learning model; update the model parameters of the deep learning model for the t-1th round using the second sampling parameter set for the t-th round, and determine the model parameters to be adjusted for the deep learning model for the t-th round.
[0021] According to an embodiment of the present disclosure, determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round includes:
[0022] In the case where the simulated fault parameters of the simulated fault in the t-th round are the third simulated fault parameters, determine the second sampling probability value according to the third simulated fault parameters; determine the third sampling point set of the t-th round from the nodes of the deep learning model in the t-1th round according to the second sampling probability value; determine the third sampling parameter set of the t-th round according to the preset cutoff value, wherein the preset cutoff value is obtained according to the initial model parameters of the deep learning model; update the model parameters of the deep learning model in the t-1th round using the third sampling parameter set of the t-th round to determine the first intermediate model parameters of the deep learning model in the t-th round; determine the model parameters to be adjusted of the deep learning model in the t-th round according to the third simulated fault parameters, the first intermediate model parameters of the deep learning model in the t-th round and the preset cutoff value.
[0023] According to an embodiment of the present disclosure, the data processing method further includes:
[0024] The model parameters of the deep learning model of the t-th round are processed so that the absolute values of the model parameters of the deep learning model of the t-th round are not greater than the absolute values of the preset cutoff values.
[0025] Another aspect of the present disclosure provides an electronic device
[0026] According to an embodiment of the present disclosure, the electronic device includes:
[0027] One or more processors; a memory for storing one or more instructions, wherein when the one or more instructions are executed by the one or more processors, the one or more processors implement any one of the above methods.
[0028] Based on the above technical solutions, it can be seen that the embodiments of the present disclosure have the following beneficial effects compared with the prior art:
[0029] The simulated fault features are added to the deep learning model corresponding to the memristor array to be trained, and the deep learning model is trained so that the deep learning model learns the features of the actual fault corresponding to the simulated fault during training, and a trained deep learning model is obtained, namely, a fault injection neural network model. According to the model parameters of the trained deep learning model, the conductance value of the memristor array is adjusted accordingly, so that in the presence of an actual fault, the memristor array determined according to the model parameters of the fault injection neural network model can still maintain a high calculation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 The flowchart of the data processing method according to the embodiment of the present disclosure is schematically shown.
[0031] Figure 2 The figure schematically shows an actual fault diagram according to an embodiment of the present disclosure.
[0032] Figure 3 The flowchart of the method for determining a neural network model according to an embodiment of the present disclosure is schematically shown.
[0033] Figure 4 The flowchart of a method for determining model parameters to be adjusted for a deep learning model in the tth round according to an embodiment of the present disclosure is schematically shown.
[0034] Figure 5 A flowchart of a method for determining model parameters to be adjusted for a deep learning model in the tth round according to another embodiment of the present disclosure is schematically shown.
[0035] Figure 6 A flowchart of a method for determining model parameters to be adjusted for a deep learning model in the tth round according to yet another embodiment of the present disclosure is schematically shown.
[0036] Figure 7 A block diagram of an electronic device suitable for implementing a standard text retrieval method and a model training method for engineering construction according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0037] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0038] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0039] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0040] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0041] In the process of implementing the concept of the present disclosure, it is found that there are at least the following problems:
[0042] Due to current manufacturing process problems, memristors may have actual faults, which may include stuck at fault (SAF) and device variation. Actual faults can change the conductance value on the memristor array constructed by memristors, and then the model parameters of the neural network corresponding to the conductance value of the memristor array will change, which may eventually lead to NCS classification errors and reduce the overall accuracy of NCS.
[0043] Therefore, a suitable method is needed to improve the robustness of NCS in order to improve the accuracy of NCS.
[0044] In order to at least partially solve the existing technical problems, the present disclosure provides a data processing method and an electronic device, which can be applied to the field of artificial intelligence.
[0045] According to an embodiment of the present disclosure, on the one hand, a data processing method is provided.
[0046] Figure 1 The flowchart of the data processing method according to the embodiment of the present disclosure is schematically shown.
[0047] like Figure 1 As shown, the data processing method may include operations S110 to S120.
[0048] In operation S110, multimedia data is acquired.
[0049] In operation S120 , the multimedia data is processed using the memristor array to obtain a data processing result.
[0050] According to an embodiment of the present disclosure, a memristor array includes multiple memristors, and the conductance values of the multiple memristors are determined based on model parameters of a fault injection neural network model. The fault injection neural network model is obtained by training a deep learning model using simulated faults, and the simulated faults are determined based on actual faults of the memristor array.
[0051] According to an embodiment of the present disclosure, when manufacturing a memristor array, a plurality of memristors included in the memristor array generate an ion effect based on the action of an electric field, so that the memristors to which different electric fields are applied can maintain different conductance values.
[0052] According to the embodiment of the present disclosure, when the memristor array is actually used, a column voltage vector can be applied to the input end of the memristor array, and a row current vector can be obtained at the input end, and the row current vector represents the product of the column voltage vector and the memristor conductance matrix.
[0053] According to the disclosed embodiments, a memristor array can be used for the study of neural networks. The change in the conductance value of the memristor can be set to characterize the change in the parameters of the neural network model, wherein the neural network model parameters can include the weights of the neural network, and the increase and decrease in the conductance value correspond to the increase and decrease in the synaptic weights, respectively. The characteristics of the neural network signal are simulated by applying a column voltage vector with variable parameters at the input end, and the parameters may include shape, frequency, and duration. The conductance value can be changed, and the change in the conductance value of the memristor array device corresponds to the change in the weight of the neural network.
[0054] According to the embodiment of the present disclosure, the corresponding relationship between the target conductance value of the memristor and the weight value of the neural network can be expressed based on formula (1):
[0055] (1)
[0056] in, is the target conductance value of the memristor; is the model parameter of the neural network; a is the first conversion coefficient; b is the second conversion coefficient; is the maximum value of the memristor conductance; is the minimum value of the memristor conductance; is the maximum value of the model parameters in each layer of the neural network; is the minimum value of the model parameter in each layer of the neural network. In the actual production process, the maximum value of the memristor conductance can be much larger than the minimum value (i.e. Much greater than ).
[0057] Figure 2 The figure schematically shows an actual fault diagram according to an embodiment of the present disclosure.
[0058] like Figure 2 As shown, the memristor array includes M×N memristors, where M and N are both positive integers; the M×N memristors correspond to g 11 ~g MNConductance value. Due to the current manufacturing process limitations, each memristor may have actual faults in actual use. In the event of an actual fault, the memristor conductance value will change, and then the model parameters of the neural network corresponding to the memristor array will change, which may eventually lead to NCS classification errors and reduce the overall accuracy of NCS. Actual faults can include state stagnation and device deviation.
[0059] like Figure 2 As shown in the figure, state stuck is a practical fault that makes the conductance state of the memristor stuck in the high impedance state (hereinafter referred to as SA1), that is, stuck in the state with the minimum conductance value; or makes the conductance state of the memristor stuck in the low impedance state (hereinafter referred to as SA0), that is, stuck in the state with the maximum conductance value; and the conductance value can only remain unchanged in the minimum state or the maximum state, resulting in the inability to change the conductance of the memristor. The state stuck of each memristor in actual use can occur randomly with a certain probability, among which, each memristor in the memristor array can have a 9.04% probability of being stuck in SA1, a 1.75% probability of being stuck in SA0, and an 89.21% probability of not being stuck.
[0060] like Figure 2 As shown in Figure 1, device deviation is a random noise that occurs between the actual conductance value and the target conductance value of a memristor, and it is impossible to eliminate device deviation through artificial control. The actual conductance value and the target conductance value of a memristor can satisfy the log-normal distribution relationship, that is, the relationship between the actual conductance value and the target conductance value can be expressed based on formula (2):
[0061] (2)
[0062] in, is the actual conductance value of the memristor; is the target conductance value of the memristor; is the normal distribution function; is the Gaussian noise parameter of the device deviation.
[0063] According to an embodiment of the present disclosure, multimedia data may include image data, audio data and video data. The data processing result may include image classification result, audio classification result and video classification result.
[0064] According to the disclosed embodiment, for the NCS based on the memristor array, with the random occurrence of state hysteresis and device deviation of the memristor array, the conductance values of the multiple memristors included in the memristor array change and send changes, resulting in deviations in the model parameters of the neural network corresponding to the multiple memristors, which may cause the overall recognition accuracy of the neuromorphic system to decrease, thereby making the entire system unusable. The neural network can be used to learn the characteristics of the data input to the neural network, where the input data can include the input image of the neural network, and can also include any regular sample information added during the training calculation process. Therefore, the simulated faults corresponding to the state hysteresis and device deviation of the memristor array can also be used as part of the data input to the neural network, so that the neural network can achieve the purpose of maintaining accuracy under non-ideal conditions of the memristor by learning the characteristics of state hysteresis and device deviation. Among them, the non-ideal conditions include the situation where the memristor array actually fails.
[0065] According to an embodiment of the present disclosure, the simulated fault feature is added to the deep learning model corresponding to the memristor array to be trained, and the deep learning model is trained so that the deep learning model learns the features of the actual fault corresponding to the simulated fault during training, thereby obtaining a trained deep learning model. According to the model parameters of the trained deep learning model, the conductance value of the memristor array is adjusted accordingly, so that in the presence of an actual fault, the memristor array determined according to the model parameters of the fault injection neural network model can still maintain a high calculation accuracy.
[0066] According to an embodiment of the present disclosure, a deep learning model can be trained repeatedly using simulated faults to obtain a trained deep learning model. It is necessary to repeat the rounds until the training parameters meet the preset end conditions. Among them, the training parameters may include the number of repeated rounds. When the number of iterations is large enough, the training of the deep learning model can be terminated, and the model parameters of the deep learning model in this case are used as the model parameters of the fault injection neural network model. In addition, the training parameters may also include the state of the loss value corresponding to the data processing results obtained in each round of training. When the state of the loss value corresponding to the data processing result meets the convergence conditions, the training of the deep learning model can be terminated, and the model parameters of the deep learning model in this case are used as the model parameters of the fault injection neural network model.
[0067] Figure 3 The flowchart of the method for determining a neural network model according to an embodiment of the present disclosure is schematically shown.
[0068] like Figure 3As shown, the method for determining a neural network model includes repeatedly performing the following operations until the training parameters meet a preset end condition, and determining the deep learning model obtained when the preset end condition is met as a fault injection neural network model. To determine the neural network model, the method may include operations S310 to S320.
[0069] In operation S310, model parameters to be adjusted of the deep learning model of the tth round are determined according to the simulated faults of the tth round and the model parameters of the deep learning model of the t-1th round, where t is an integer greater than 1.
[0070] In operation S320, the model parameters to be adjusted of the deep learning model of the tth round are adjusted using the sample multimedia data to obtain the model parameters of the deep learning model of the tth round.
[0071] According to an embodiment of the present disclosure, when the training parameters of the t-1th round do not meet the preset end conditions, the model parameters of the deep learning model obtained in the t-1th round are used as the initial model parameters of the deep learning model in the t-1th round, and the initial model parameters of the deep learning model in the t-1th round are adjusted according to the simulated faults in the tth round, and the model parameters to be adjusted of the deep learning model in the tth round are determined.
[0072] According to the embodiment of the present disclosure, the initial model parameters of the deep learning model of the t-1th round are adjusted according to the simulated fault of the tth round, thereby determining the parameters of the model to be adjusted of the deep learning model of the tth round. In the process of processing the sample multimedia data using the parameters of the model to be adjusted of the deep learning model of the tth round, the simulated fault of the tth round can be added to the calculation process of the sample multimedia data, thereby enabling the deep learning model of the tth round to learn the features corresponding to the simulated fault during the calculation process.
[0073] According to an embodiment of the present disclosure, after processing sample multimedia data using the parameters of the model to be adjusted of the deep learning model of the tth round, the sample processing result of the tth round is obtained. According to the sample processing result of the tth round, the parameters of the model to be adjusted of the deep learning model of the tth round can be reversely adjusted to obtain the model parameters of the deep learning model of the tth round.
[0074] According to an embodiment of the present disclosure, according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round, determining the model parameters to be adjusted of the deep learning model of the tth round includes:
[0075] According to the simulated fault of the tth round, a sampling point set of the tth round is determined from the nodes of the deep learning model of the t-1th round, wherein the sampling point set includes at least one node. According to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round, a sampling parameter set of the tth round is determined, wherein the sampling parameter set includes sampling parameters corresponding to the sampling point set of the tth round. The model parameters of the deep learning model of the t-1th round are updated using the sampling parameter set of the tth round to obtain the model parameters to be adjusted of the deep learning model of the tth round.
[0076] According to the embodiments of the present disclosure, the simulated fault corresponding to the actual fault of the memristor can simulate the characteristics of state stuck and / or device deviation. The sampling method can be adaptively selected according to the characteristics of the simulated fault simulation in the specific tth round to obtain the sampling point set of the tth round, and each sampling point in the sampling point set is updated by using the simulated fault to obtain the sampling parameter set of the tth round corresponding to the sampling point set of the tth round, and the model parameters of the deep learning model of the t-1th round are updated by using the sampling parameter set of the tth round to obtain the model parameters to be adjusted of the deep learning model of the tth round. In this process, the characteristics of the simulated fault simulation in the tth round are added to the deep learning model of the tth round, so that in the subsequent training process, the sample multimedia data can be input into the deep learning model of the tth round to perform forward propagation calculation, so that the deep learning model of the tth round can learn the characteristics of the simulated fault simulation.
[0077] According to an embodiment of the present disclosure, the model parameters to be adjusted of the deep learning model of the tth round are adjusted using the sample multimedia data, and the model parameters of the deep learning model of the tth round are obtained, including:
[0078] The sample multimedia data is processed by the deep learning model of the tth round to obtain the sample processing result of the tth round. Based on the loss function, the loss value of the tth round is obtained according to the sample processing result of the tth round. The model parameters to be adjusted of the deep learning model of the tth round are adjusted according to the loss value of the tth round to obtain the model parameters of the deep learning model of the tth round.
[0079] According to the embodiment of the present disclosure, the sample processing result of the tth round is processed using the loss function to obtain the loss value of the tth round. Back propagation is performed based on the loss value of the tth round, gradient calculation is performed on the model parameters to be adjusted of the deep learning model of the tth round, the model parameters to be adjusted of the deep learning model of the tth round are updated, and the model parameters of the deep learning model of the tth round are obtained.
[0080] According to an embodiment of the present disclosure, after adjusting the model parameters to be adjusted of the deep learning model of the tth round according to the loss value of the tth round to obtain the model parameters of the deep learning model of the tth round, the method further includes:
[0081] The model parameters corresponding to the target node of the deep learning model in the tth round are restored to the model parameters of the deep learning model in the t-1th round; wherein the target node in the tth round is a node determined according to the sampling point set in the tth round.
[0082] According to an embodiment of the present disclosure, it can be selected whether it is necessary to restore the model parameters corresponding to the target nodes of the deep learning model of the tth round to the model parameters of the deep learning model of the t-1th round based on the determined sampling point set of the tth round.
[0083] Figure 4 The flowchart of a method for determining model parameters to be adjusted for a deep learning model in the tth round according to an embodiment of the present disclosure is schematically shown.
[0084] like Figure 4 As shown, determining the model parameters to be adjusted of the deep learning model in the tth round may include operations S410 to S430.
[0085] In operation S410 , the nodes of the deep learning model of the t-1th round are determined as the first sampling point set of the tth round.
[0086] In operation S420 , a first sampling parameter set of the tth round is determined according to the first simulated fault, the model parameters of the deep learning model of the t-1th round, and a preset cutoff value.
[0087] In operation S430, the model parameters of the deep learning model of the t-1th round are updated using the first sampling parameter set of the tth round to obtain the model parameters to be adjusted of the deep learning model of the tth round.
[0088] According to an embodiment of the present disclosure, when the simulated fault in the tth round is the first simulated fault, operations S410 to S430 may be performed.
[0089] According to the embodiment of the present disclosure, the model parameters of each layer of the deep learning neural network approximately obey a Gaussian distribution, and the Gaussian distribution can be approximately a mean of , the standard deviation is Gaussian distribution. According to our research, the maximum value of the model parameter and minimum value There is a close relationship between the robustness of NCS based on memristor arrays and the maximum value of model parameters. and minimum value The higher the corresponding discreteness, the more sensitive the neural network is to the non-ideal conditions of the memristor, and the more the accuracy decreases under the non-ideal conditions of the memristor. Therefore, according to the standard deviation of the distribution of model parameters at each layer The associated number As the cutoff value of the model parameters of this layer, the model parameters of this layer of the neural network are limited to within the range.
[0090] According to the embodiment of the present disclosure, in order to take into account the accuracy performance of the deep learning neural network, it is possible to select As the parameter of the cutoff value, the preset cutoff value is ,Right now .
[0091] According to the embodiment of the present disclosure, the maximum value of the sampling parameter set of the tth round corresponding to the sampling point set of the tth round can be set to , the minimum value of the sampling parameter set in round t is .
[0092] According to the embodiment of the present disclosure, it can be set that when the model parameters of the determined sampling parameter set of the tth round are the maximum value or the minimum value, the model parameters corresponding to the target node need to be restored. That is, in this case, the model parameters of the deep learning model after the model parameters corresponding to the target node are restored are the model parameters of the deep learning model of the tth round.
[0093] According to the embodiment of the present disclosure, it can be set that when the model parameters of the determined sampling parameter set of the tth round are not the maximum value or the minimum value, it is not necessary to restore the model parameters corresponding to the target node. That is, in this case, the model parameters of the deep learning model obtained by updating the model parameters of the deep learning model of the t-1th round using the sampling parameter set of the tth round are the model parameters to be adjusted of the deep learning model of the tth round.
[0094] According to an embodiment of the present disclosure, when the loss value of the tth round converges with the loss value of the previous t-1 round, it can be considered that the training parameters of the tth round meet the preset termination conditions, and the model parameters of the deep learning model of the tth round are the model parameters of the fault injection neural network model.
[0095] According to an embodiment of the present disclosure, when the loss value of the tth round does not converge with the loss value of the previous t-1 round, but the number of rounds of the tth round has reached the preset number of rounds, it can also be considered that the training parameters of the tth round meet the preset end conditions, and the model parameters of the deep learning model of the tth round are the model parameters of the fault injection neural network model.
[0096] According to an embodiment of the present disclosure, when the loss value of the tth round does not converge with the loss value of the previous t-1 round, and the number of rounds of the tth round does not meet the preset number of rounds, the model parameters to be adjusted of the deep learning model of the t+1th round are determined according to the model parameters of the deep learning model of the tth round, and the deep learning model continues to be trained.
[0097] According to the embodiment of the present disclosure, the first simulated fault includes a simulated fault corresponding to the device deviation. The relationship between the actual conductance value and the expected conductance value is expressed according to formula (2). The Gaussian noise parameter of the device deviation is used. , generated by is a Gaussian distribution with standard deviation. After determining the nodes of the deep learning model of the t-1th round as the first sampling point set of the tth round, a fault injection based on the first simulated fault is performed on each sampling point in the first sampling point set of the tth round.
[0098] According to the embodiment of the present disclosure, based on formula (3), the relationship between the actual conductance value and the target conductance value can be approximately obtained in a linear representation:
[0099] (3)
[0100] According to the embodiment of the present disclosure, the fault injection based on the first simulated fault can be performed on each sampling point in the first sampling point set of the tth round based on formula (4):
[0101]
[0102] (4)
[0103] Among them, when the simulated fault in the tth round is the first simulated fault, is the model parameter to be adjusted of the deep learning model in the tth round; are the model parameters of the deep learning model in the t-1th round; is the actual conductance value of the memristor in round t; is the target conductance value of the memristor in the tth round.
[0104] According to the embodiment of the present disclosure, since the maximum value of the memristor conductance is much greater than the minimum value, that is, Much greater than , the model parameters to be adjusted of the deep learning model in the tth round can be expressed based on formula (5):
[0105] (5)
[0106] in, is the preset cutoff value.
[0107] According to the embodiment of the present disclosure, when the simulated fault of the tth round is the first simulated fault, after obtaining the model parameters to be adjusted of the deep learning model of the tth round, the data included in the handwritten digital grayscale image dataset, the Cifar10 dataset or the Cifar100 dataset can be used as sample multimedia data, and the sample multimedia data can be processed using the deep learning model of the tth round to obtain the corresponding sample processing results. The loss value of the tth round is obtained according to the corresponding sample processing results, and the model parameters to be adjusted of the deep learning model of the tth round are adjusted by gradient calculation using the loss value of the tth round to obtain the second intermediate model parameters of the deep learning model of the tth round.
[0108] According to an embodiment of the present disclosure, based on the second intermediate model parameters of the deep learning model of the tth round, the model parameters of the deep learning model of the tth round are updated through a conversion similar to formula (5) to obtain the model parameters of the deep learning model of the tth round.
[0109] According to the embodiment of the present disclosure, more specifically, the second intermediate model parameter of the deep learning model of the tth round can be set to be equivalent to the value on the left side of formula (5), and the value equivalent to the value on the right side can be obtained through formula (5). The model parameter values of the model parameter values of the model parameter values are obtained, that is, the model parameters of the deep learning model of the tth round are obtained.
[0110] According to the disclosed embodiment, when the simulated fault in the tth round is the first simulated fault, the handwritten digital grayscale image dataset can be used as sample multimedia data. When the handwritten digital grayscale image dataset is input into the deep learning model in the tth round as sample multimedia data, the Gaussian noise parameter is set to 0.3, the preset number of rounds is set to 7820, and the loss function is set to the cross entropy loss function. The fault injection neural network model trained in this way can process the handwritten digital grayscale image dataset with an accuracy of 98%. The untrained deep learning model processes the handwritten digital grayscale image dataset with an accuracy of 77%.
[0111] According to the embodiment of the present disclosure, when the simulated fault in the tth round is the first simulated fault, the Cifar10 data set can be used as sample multimedia data. When the Cifar10 data set is used as sample multimedia data and input into the deep learning model in the tth round, the Gaussian noise parameter is set to 0.1, the preset number of rounds is set to 19500, and the loss function is set to the cross entropy loss function. The accuracy of the fault injection neural network model trained in this way to process the Cifar10 data set can reach 70%. The accuracy of the untrained deep learning model to process the Cifar10 data set is 19%.
[0112] According to the embodiment of the present disclosure, when the simulated fault in the tth round is the first simulated fault, the Cifar100 data set can be used as sample multimedia data. When the Cifar100 data set is input as sample multimedia data into the deep learning model in the tth round, the Gaussian noise parameter is set to 0.1, the preset number of rounds is set to 19500, and the loss function is set to the cross entropy loss function. The accuracy of the fault injection neural network model trained in this way to process the Cifar100 data set can reach 43%. The accuracy of the untrained deep learning model to process Cifar100 is 2%.
[0113] According to the embodiments of the present disclosure, the accuracy of the data processing results obtained by using the fault injection neural network to process the handwritten digital grayscale image dataset, the Cifar10 dataset and the Cifar100 dataset respectively is higher than the accuracy of the data processing results obtained by using the untrained deep learning model, that is, the fault injection neural network model learns the fault characteristics of the device deviation corresponding to the simulated fault, so that the memristor array determined by the model parameters of the fault injection neural network model can maintain a high calculation accuracy in the presence of actual faults.
[0114] Figure 5 A flowchart of a method for determining model parameters to be adjusted for a deep learning model in the tth round according to another embodiment of the present disclosure is schematically shown.
[0115] like Figure 5 As shown, the method for determining the model parameters to be adjusted of the deep learning model in the tth round may also include operations S510 to S540.
[0116] In operation S510 , a first sampling probability value is determined according to a second simulated fault.
[0117] In operation S520, a second sampling point set of the tth round is determined from the nodes of the deep learning model of the t-1th round according to the first sampling probability value.
[0118] In operation S530, a second sampling parameter set of the tth round is determined according to a preset cutoff value, wherein the preset cutoff value is obtained according to an initial model parameter of the deep learning model.
[0119] In operation S540, the model parameters of the deep learning model of the t-1th round are updated using the second sampling parameter set of the tth round, and the model parameters to be adjusted of the deep learning model of the tth round are determined.
[0120] According to an embodiment of the present disclosure, the second simulated fault includes a simulated fault corresponding to a state stuck. According to the probability of random occurrence of the state stuck, the first sampling probability value is determined to be 10.79%, and 10.79% of the nodes of the deep learning model in the t-1th round are randomly selected as the second sampling point set for the tth round.
[0121] According to the embodiment of the present disclosure, based on the probability of state stuck at SA1 and the probability of state stuck at SA0, the second sampling parameters corresponding to 83.8% of the second sampling points in the second sampling point set of the tth round are modified to , modify the second sampling parameters corresponding to 16.2% of the second sampling points in the second sampling point set of round t to .
[0122] According to an embodiment of the present disclosure, the modified second sampling parameters corresponding to the second sampling point set and the unmodified model parameters corresponding to the nodes of the deep learning model of the t-1th round are determined as the model parameters to be adjusted of the deep learning model of the tth round.
[0123] According to an embodiment of the present disclosure, when the simulated fault of the tth round is the second simulated fault, after obtaining the model parameters to be adjusted of the deep learning model of the tth round, the data included in the handwritten digital grayscale image dataset, the Cifar10 dataset or the Cifar100 dataset can be used as sample multimedia data, and the sample multimedia data can be processed using the deep learning model of the tth round to obtain the corresponding sample processing result. The loss value of the tth round is obtained according to the corresponding sample processing result, and the model parameters to be adjusted of the deep learning model of the tth round are adjusted according to the loss value of the tth round. After the model parameters to be adjusted of the deep learning model of the tth round are adjusted using the sample multimedia data, the adjusted model parameters of the second sampling point set are restored to the model parameters of the corresponding deep learning model of the t-1th round.
[0124] According to an embodiment of the present disclosure, the model parameters of the deep learning model of the corresponding t-1 round restored and the adjusted model parameters of the t round restored to the model parameters of the deep learning model of the corresponding t-1 round are determined as the model parameters of the deep learning model of the t round.
[0125] According to the disclosed embodiment, when the simulated fault in the tth round is the second simulated fault, the handwritten digital grayscale image dataset can be used as sample multimedia data. When the handwritten digital grayscale image dataset is input as sample multimedia data into the deep learning model in the tth round, the preset number of rounds is set to 7820, and the loss function is set to the cross entropy loss function. The fault injection neural network model trained in this way can process the handwritten digital grayscale image dataset with an accuracy of 98%. The untrained deep learning model processes the handwritten digital grayscale image dataset with an accuracy of 87%.
[0126] According to the embodiment of the present disclosure, when the simulated fault in the tth round is the first simulated fault, the Cifar10 data set can also be used as sample multimedia data. When the Cifar10 data set is used as sample multimedia data and input into the deep learning model in the tth round, the preset number of rounds is set to 19500, and the loss function is set to the cross entropy loss function. The accuracy of the fault injection neural network model trained in this way to process the Cifar10 data set can reach 51%. The accuracy of the untrained deep learning model to process the Cifar10 data set is 10%.
[0127] According to the embodiment of the present disclosure, when the simulated fault in the tth round is the second simulated fault, the Cifar100 data set can also be used as sample multimedia data. When the Cifar100 data set is input as sample multimedia data into the deep learning model in the tth round, the preset number of rounds is set to 19500, and the loss function is set to the cross entropy loss function. The accuracy of the fault injection neural network model trained in this way to process the Cifar100 data set can reach 38%. The accuracy of the untrained deep learning model to process Cifar100 is 1%.
[0128] According to the embodiments of the present disclosure, the accuracy of the data processing results obtained by using the fault injection neural network to process the handwritten digital grayscale image dataset, the Cifar10 dataset and the Cifar100 dataset respectively is higher than the accuracy of the data processing results obtained by using the untrained deep learning model, that is, the fault injection neural network model learns the characteristics of the state stuck fault corresponding to the simulated fault, so that the memristor array determined by the model parameters of the fault injection neural network model can maintain a high calculation accuracy in the presence of actual faults.
[0129] Figure 6 A flowchart of a method for determining model parameters to be adjusted for a deep learning model in the tth round according to yet another embodiment of the present disclosure is schematically shown.
[0130] like Figure 6As shown, the method for determining the model parameters to be adjusted of the deep learning model in the tth round may also include operations S610 to S650.
[0131] In operation S610 , a second sampling probability value is determined according to a third simulated fault parameter.
[0132] In operation S620, according to the second sampling probability value, from the t-1th round of deep learning
[0133] The third sampling point set of the tth round is determined in the nodes of the model.
[0134] In operation S630, a third sampling parameter set of the tth round is determined according to a preset cutoff value.
[0135] In operation S640, the model parameters of the deep learning model of the t-1th round are updated using the third sampling parameter set of the tth round, and the first intermediate model parameters of the deep learning model of the tth round are determined.
[0136] In operation S650, the model parameters to be adjusted of the deep learning model of the tth round are determined according to the third simulated fault parameter, the first intermediate model parameter of the deep learning model of the tth round, and the preset cutoff value.
[0137] According to an embodiment of the present disclosure, the third simulated fault includes simulated faults corresponding to device deviation and state stuck. According to the probability of random occurrence of state stuck, the second sampling probability value is determined to be 10.79%, and 10.79% of the nodes of the deep learning model in the t-1th round are randomly selected as the third sampling point set of the tth round.
[0138] According to the embodiment of the present disclosure, based on the probability of state stuck at SA1 and the probability of state stuck at SA0, the third sampling parameters corresponding to 83.8% of the third sampling points in the third sampling point set of the tth round are modified to , modify the third sampling parameters corresponding to 16.2% of the third sampling points in the third sampling point set of round t to , the modified third sampling parameters corresponding to the third sampling point set and the unmodified model parameters corresponding to the nodes of the deep learning model of the t-1th round are determined as the first intermediate model parameters of the deep learning model of the tth round.
[0139] According to an embodiment of the present disclosure, when the first intermediate model parameter of the deep learning model of the tth round is determined, the Gaussian noise parameter of the device deviation is used. , perform fault injection of the simulated fault corresponding to the device deviation on each node in the first intermediate model parameter of the deep learning model of the tth round. It can be understood that, when the simulated fault of the tth round is the third simulated fault, the operation of performing fault injection of the simulated fault corresponding to the device deviation can refer to the operation of performing fault injection of the first simulated fault when the simulated fault of the tth round is the first simulated fault, and will not be repeated here.
[0140] According to an embodiment of the present disclosure, after determining the model parameters to be adjusted of the deep learning model of the tth round according to the third simulated fault parameter, the first intermediate model parameter of the deep learning model of the tth round and the preset truncation value, the model parameters to be adjusted of the deep learning model of the tth round can be adjusted using the handwritten digital grayscale image dataset, the Cifar10 dataset and the Cifar100 dataset respectively. It can be understood that the operation of adjusting the model parameters to be adjusted of the deep learning model of the tth round using the handwritten digital grayscale image dataset, the Cifar10 dataset and the Cifar100 dataset respectively can refer to the operation when the simulated fault in the tth round is the second simulated fault, and will not be repeated here.
[0141] According to an embodiment of the present disclosure, the data processing method may further include: processing the model parameters of the deep learning model of the tth round so that the absolute value of the model parameters of the deep learning model of the tth round is not greater than the absolute value of the preset truncation value.
[0142] According to the embodiment of the present disclosure, the model parameters of the deep learning model of the tth round are restricted so that the model parameters of the deep learning model of the tth round are restricted to 0.01% before and after the fault injection. Within the range of , the deep learning model of the tth round can be controlled to a certain extent, and the deep learning model of the tth round can maintain sufficient accuracy in the presence of simulated faults.
[0143] According to the disclosed embodiment, after the data processing method, the NCS based on the memristor array can maintain a certain accuracy in the presence of simulated faults. Without adding actual adjustment hardware, the NCS adjusted by the data processing method can be widely deployed on non-ideal memristor arrays without the need for special operations on each memristor array, thereby reducing the hardware cost of large-scale application of memristor arrays. In addition, the method of training the NCS based on the memristor array to obtain the fault injection neural network model is applicable to NCSs of different sizes.
[0144] According to another aspect of an embodiment of the present disclosure, an electronic device is provided.
[0145] The electronic device includes: one or more processors; and a memory for storing one or more instructions, wherein when the one or more instructions are executed by the one or more processors, the one or more processors implement the method of any one of the above embodiments.
[0146] Figure 7 A block diagram of an electronic device suitable for implementing a standard text retrieval method and a model training method for engineering construction according to an embodiment of the present disclosure is schematically shown.
[0147] Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0148] like Figure 7 As shown, the computer electronic device 700 according to the embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 to the random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include an on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.
[0149] In RAM703, various programs and data required for the operation of electronic device 700 are stored. Processor 701, ROM702 and RAM703 are connected to each other via bus 704. Processor 701 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM702 and / or RAM703. It should be noted that the program can also be stored in one or more memories other than ROM702 and RAM703. Processor 701 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in one or more memories.
[0150] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the I / O interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage portion 708 as needed.
[0151] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.
[0152] The flow chart and block diagram in the accompanying drawings schematically illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of boxes in the block diagram or flow chart, can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0153] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0154] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrating the purpose, technical solutions and beneficial effects of the present disclosure, and are not intended to limit the scope of the present disclosure. Although each embodiment is described above separately, it does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the attached claims and their equivalents. Without departing from the scope of the present disclosure, within the spirit and principles of the present disclosure, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of protection of the present disclosure.
Claims
1. A data processing method based on a memristor array, comprising: Acquiring multimedia data; as well as Processing the multimedia data using a memristor array to obtain a data processing result; The memristor array includes a plurality of memristors, the conductance values of the plurality of memristors are determined according to model parameters of a fault injection neural network model, the fault injection neural network model is obtained by training a deep learning model using simulated faults, and the simulated faults are determined according to actual faults of the memristor array; the fault injection neural network model is obtained by training a deep learning model using simulated faults, and the deep learning model is repeatedly performed until the training parameters meet a preset end condition, and the deep learning model obtained when the preset end condition is met is determined as the fault injection neural network model: Determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated faults of the tth round and the model parameters of the deep learning model of the t-1th round, where t is an integer greater than 1; and Using the sample multimedia data to adjust the model parameters to be adjusted of the deep learning model of the tth round to obtain the model parameters of the deep learning model of the tth round, including: In the case where the simulated fault in the tth round is the first simulated fault, the nodes of the deep learning model in the t-1th round are determined as the first sampling point set of the tth round, and the first simulated fault includes a simulated fault corresponding to the device deviation.
2. The data processing method according to claim 1, wherein: Determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round includes: Determine, according to the simulated fault of the tth round, a sampling point set of the tth round from the nodes of the deep learning model of the t-1th round, wherein the sampling point set includes at least one node; Determining a sampling parameter set for the tth round according to the simulated faults of the tth round and the model parameters of the deep learning model of the t-1th round, wherein the sampling parameter set includes sampling parameters corresponding to the sampling point set of the tth round; and The model parameters of the deep learning model of the t-1th round are updated using the sampling parameter set of the tth round to obtain the model parameters to be adjusted of the deep learning model of the tth round.
3. The data processing method according to claim 2, wherein: The step of adjusting the model parameters to be adjusted of the deep learning model of the tth round by using the sample multimedia data to obtain the model parameters of the deep learning model of the tth round includes: Processing the sample multimedia data using the t-th round of deep learning model to obtain the t-th round of sample processing results; Based on the loss function, according to the sample processing result of the tth round, a loss value of the tth round is obtained; and The model parameters to be adjusted of the deep learning model of the tth round are adjusted according to the loss value of the tth round to obtain the model parameters of the deep learning model of the tth round.
4. The data processing method according to claim 3, further comprising: Restore the model parameters corresponding to the target nodes of the deep learning model of the tth round to the model parameters of the deep learning model of the t-1th round; The target node of the tth round is a node determined according to the sampling point set of the tth round.
5. The data processing method according to claim 1, wherein: The determining, based on the simulated faults in the tth round and the model parameters of the deep learning model in the t-1th round, the model parameters to be adjusted of the deep learning model in the tth round also includes: Determining a first set of sampling parameters for the tth round according to the first simulated fault, model parameters of the deep learning model for the t-1th round, and a preset cutoff value, wherein the preset cutoff value is obtained according to initial model parameters of the deep learning model; and The model parameters of the deep learning model of the t-1th round are updated using the first sampling parameter set of the tth round, and the model parameters to be adjusted of the deep learning model of the tth round are determined.
6. The data processing method according to claim 4, wherein: Determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round includes: When the simulated fault in the tth round is the second simulated fault, Determining a first sampling probability value according to the second simulated fault; Determine, according to the first sampling probability value, a second sampling point set of the tth round from the nodes of the deep learning model of the t-1th round; Determining the second sampling parameter set of the tth round according to a preset cutoff value, wherein the preset cutoff value is obtained according to the initial model parameters of the deep learning model; and The model parameters of the deep learning model of the t-1th round are updated using the second sampling parameter set of the tth round, and the model parameters to be adjusted of the deep learning model of the tth round are determined.
7. The data processing method according to claim 4, wherein: Determining the model parameters to be adjusted of the deep learning model of the tth round according to the simulated fault of the tth round and the model parameters of the deep learning model of the t-1th round includes: When the simulated fault parameter of the simulated fault in the tth round is the third simulated fault parameter, Determining a second sampling probability value according to a third simulated fault parameter; Determine, according to the second sampling probability value, a third sampling point set of the tth round from the nodes of the deep learning model of the t-1th round; Determining a third sampling parameter set of the tth round according to a preset cutoff value, wherein the preset cutoff value is obtained according to an initial model parameter of the deep learning model; Using the third sampling parameter set of the t-th round, updating the model parameters of the deep learning model of the t-1th round, and determining the first intermediate model parameters of the deep learning model of the t-th round; and The model parameters to be adjusted of the deep learning model of the tth round are determined according to the third simulated fault parameter, the first intermediate model parameter of the deep learning model of the tth round and the preset cutoff value.
8. The data processing method according to any one of claims 2 to 7, further comprising: The model parameters of the deep learning model of the t-th round are processed so that the absolute values of the model parameters of the deep learning model of the t-th round are not greater than the absolute values of the preset cutoff values.
9. An electronic device, comprising: one or more processors; a memory for storing one or more instructions, Wherein, when the one or more instructions are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Adversarial sample generation method and device
CN111582473A