Neural network model training method and device, electronic equipment and storage medium
By calculating the weight similarity of the output channels in the neural network model and correcting it, the feature redundancy problem is solved, and the training efficiency and performance of the model are improved.
Patent Information
- Application Number
- CN202410704000.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, neural network models have feature redundancy problems during training, resulting in insufficient model performance.
By obtaining the initial weight and predicted values corresponding to each output channel of each layer of the network in the initial neural network model, the similarity between each two output channels is calculated, and the weights are corrected based on the similarity and prediction losses, reducing the similarity to avoid feature redundancy.
It improves the training efficiency of neural network models, avoids feature redundancy, and improves the expression ability and prediction accuracy of the model.
Smart Images

Figure CN120373384A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, electronic device, and storage medium for training a neural network model. Background Art
[0002] The weights of a model play a crucial role in machine learning and deep learning. Weights can capture patterns and features in the input data and map these features to the output of the model. By adjusting the weights, the model can learn how to identify key information in the input data and make accurate predictions. Therefore, as important parameters in the model, the adjustment and optimization of weights have a significant impact on the performance and generalization ability of the model. Summary of the Invention
[0003] The present disclosure aims to solve at least one of the technical problems in the related art to some extent.
[0004] A first aspect embodiment of the present disclosure provides a method for training a neural network model, including:
[0005] Obtaining the initial weights corresponding to each output channel of each layer in the initial neural network model, and the predicted value of the initial neural network model for the sample data;
[0006] Determining the similarity between the initial weights corresponding to every two output channels of each layer of the network;
[0007] Determining a first loss corresponding to the initial neural network model according to the similarity;
[0008] Based on the first loss and the second loss corresponding to the predicted value, correcting the initial weights in the initial neural network model.
[0009] A second aspect embodiment of the present disclosure provides a device for training a neural network model, including:
[0010] An obtaining module, configured to obtain the initial weights corresponding to each output channel of each layer in the initial neural network model, and the predicted value of the initial neural network model for the sample data;
[0011] A first determination module, configured to determine the similarity between the initial weights corresponding to every two output channels of each layer of the network;
[0012] A second determination module, configured to determine a first loss corresponding to the initial neural network model according to the similarity;
[0013] A correction module, configured to correct the initial weights in the initial neural network model based on the first loss and the second loss corresponding to the predicted value.
[0014] The third - aspect embodiment of the present disclosure provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the training method of the neural network model proposed in the first - aspect embodiment of the present disclosure.
[0015] The fourth - aspect embodiment of the present disclosure provides a computer - readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the training method of the neural network model proposed in the first - aspect embodiment of the present disclosure.
[0016] The training method, device, electronic device, and storage medium of the neural network model provided by the present disclosure have the following beneficial effects:
[0017] In the embodiment of the present disclosure, first, obtain the initial weights corresponding to each output channel of each layer of the network in the initial neural network model and the predicted values of the initial neural network model for the sample data. Then, determine the similarity between the initial weights corresponding to every two output channels of each layer of the network, and based on the similarity, determine the first loss corresponding to the initial neural network model. Finally, correct the initial weights in the initial neural network model based on the first loss and the second loss corresponding to the predicted values. Thus, the similarity between the initial weights corresponding to every two output channels of each layer of the network can be added to the loss of the model to train the model, so that the similarity between the initial weights corresponding to every two output channels of each layer of the network is small, and further, different channels of each layer of the neural network can learn more different features, thereby avoiding feature redundancy and improving the model performance.
[0018] The additional aspects and advantages of the present disclosure will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above - mentioned and / or additional aspects and advantages of the present disclosure will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0020] Figure 1 is a schematic flowchart of a training method of a neural network model provided by an embodiment of the present disclosure;
[0021] Figure 2 is a schematic flowchart of a training method of a neural network model provided by another embodiment of the present disclosure;
[0022] Figure 3 is a schematic structural diagram of a training device of a neural network model provided by another embodiment of the present disclosure;
[0023] Figure 4 The block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure is shown. Detailed implementation manners
[0024] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation of the present disclosure.
[0025] The training method, device, electronic device, and storage medium of the neural network model according to the embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0026] Figure 1 It is a schematic flowchart of a training method of a neural network model provided by an embodiment of the present disclosure.
[0027] In the embodiments of the present disclosure, the training method of the neural network model is configured in a training device of the neural network model as an example. The training device of the neural network model can be applied to any electronic device so that the electronic device can perform the training function of the neural network model.
[0028] As Figure 1 shown, the training method of the neural network model may include the following steps:
[0029] Step 101: Obtain the initial weights corresponding to each output channel of each layer of the initial neural network model and the predicted values of the initial neural network model for the sample data.
[0030] Among them, the initial weights corresponding to each output channel include the weights between the output channel and each input channel. For example, if the input channels include input channel 1, input channel 2, and input channel 3, the initial weights corresponding to output channel 1 include the connection weights between input channel 1 and output channel 1, the connection weights between input channel 2 and output channel 1, and the connection weights between input channel 3 and output channel 1.
[0031] In some embodiments, the sample data is at least one of audio data, text data, and image data. Therefore, the initial neural network model can be an audio processing model, such as a speech recognition model, a speech classification model, etc. The initial neural network model can also be a text processing model, such as a text classification model, a text search model, etc. The initial neural network model can also be an image processing model, such as an image recognition model, an image classification model, etc. The initial neural network model can also be a multi-modal processing model, such as a model that can process images and text simultaneously, a model that can process audio and images simultaneously. The present disclosure does not limit this.
[0032] Wherein, the predicted value can be the predicted value obtained by the initial neural network model predicting the sample data after inputting the sample data into the initial neural network model.
[0033] Step 102, determine the similarity between the initial weights corresponding to every two output channels of each layer of the network.
[0034] In some embodiments, the Euclidean distance between the initial weights corresponding to every two output channels can be calculated, and according to the Euclidean distance, the similarity between the initial weights corresponding to every two output channels can be determined. Among them, the greater the Euclidean distance, the smaller the similarity; the smaller the Euclidean distance, the greater the similarity.
[0035] In some embodiments, the pre-similarity between the initial weights corresponding to every two output channels can also be calculated, and the cosine similarity can be determined as the similarity between the initial weights corresponding to every two output channels.
[0036] Step 103, determine the first loss corresponding to the initial neural network model according to the similarity.
[0037] In some embodiments, the sum of all similarities can be determined as the first loss corresponding to the initial neural network model.
[0038] In some embodiments, the third loss corresponding to the similarity between every two output channels of each layer of the network can be determined first, and then the average value of the third losses corresponding to each layer of the network can be determined as the first loss corresponding to the initial neural network model. Thus, by averaging the third losses corresponding to each layer of the network, the influence of the first loss on the overall loss of the initial neural network model can be reduced, and it can be avoided that the trained neural network model cannot accurately predict the input data.
[0039] In some embodiments, the average value of the similarities between every two output channels of each layer of the network can be determined as the third loss.
[0040] In some embodiments, the mean squared error corresponding to the similarity between every two output channels of each layer of the network may also be determined as the third loss. The mean squared error corresponding to the similarity between every two output channels of each layer of the network may also be referred to as the L2 norm loss corresponding to the similarity between every two output channels of each layer of the network.
[0041] In some embodiments, the absolute deviation corresponding to the similarity between every two output channels of each layer of the network may also be determined as the third loss. The absolute deviation corresponding to the similarity between every two output channels of each layer of the network may also be referred to as the L1 norm loss corresponding to the similarity between every two output channels of each layer of the network.
[0042] Among them, the target similarity used to calculate the mean squared error or the absolute deviation may be a preset value. For example, it may be 0, 0.1, etc. The present disclosure does not limit this.
[0043] Step 104: Based on the first loss and the second loss corresponding to the predicted value, correct the initial weights in the initial neural network model.
[0044] In some embodiments, the first loss and the second loss are averaged or added to obtain a target loss, and based on the target loss, the initial weights of the model are corrected until the target loss is less than the loss threshold or the preset number of iterative corrections is reached.
[0045] In some embodiments, the first weight corresponding to the first loss and the second weight corresponding to the second loss may be determined first, and then based on the first weight and the second weight, the first loss and the second loss are weighted to obtain a target loss. Then, based on the target loss, the initial weights in the initial neural network model are corrected until the target loss is less than the loss threshold.
[0046] Among them, the first weight and the second weight may be the same or different. The present disclosure does not limit this.
[0047] In some embodiments, the first weight and the second weight may be determined based on the importance of the first loss and the second loss to the model. For example, in order to make the prediction ability of the model better, the second weight corresponding to the second loss may be appropriately greater than the first weight corresponding to the first loss. In order to make the weight similarity of each channel of the model lower, the first weight corresponding to the first loss may be appropriately greater than the second weight corresponding to the second loss.
[0048] It should be noted that the target loss being less than the loss threshold can ensure that the similarity of the weights of different channels of the model is lower, so that different channels of each layer of the neural network learn more different features, thereby avoiding feature redundancy and improving the expression ability of the model.
[0049] In the embodiments of the present disclosure, first, obtain the initial weights corresponding to each output channel of each layer in the initial neural network model and the predicted values of the initial neural network model for the sample data. Then, determine the similarity between the initial weights corresponding to every two output channels of each layer, and based on the similarity, determine the first loss corresponding to the initial neural network model. Finally, correct the initial weights in the initial neural network model based on the first loss and the second loss corresponding to the predicted values. Thus, the similarity between the initial weights corresponding to every two output channels of each layer can be added to the loss of the model to train the model, so that the similarity between the initial weights corresponding to every two output channels of each layer is small, and further, different channels of each layer of the neural network can learn more different features, thereby avoiding feature redundancy and improving the model performance.
[0050] Figure 2 As shown in the flowchart of a method for training a neural network model provided by an embodiment of the present disclosure, Figure 2 the method for training the neural network model may include the following steps:
[0051] Step 201: Obtain the initial weights corresponding to each output channel of each layer in the initial neural network model and the predicted values of the initial neural network model for the sample data.
[0052] Among them, for the specific implementation form of step 201, reference may be made to the detailed descriptions in other embodiments of the present disclosure, and details are not described herein again.
[0053] Step 202: Obtain the four-dimensional weight matrix corresponding to each layer, where the first dimension in the four-dimensional weight matrix represents the output channel, the second dimension represents the input channel, the third dimension represents the height of the kernel, and the fourth dimension represents the width of the kernel.
[0054] In the embodiments of the present disclosure, the number of output channels and input channels corresponding to each layer is not limited.
[0055] In some embodiments, the height and width of the kernel may be the same or different, and the present disclosure does not limit this.
[0056] Among them, the kernel (which may be one or more) corresponding to the i-th output channel and the j-th input channel represents the connection weight between the i-th output channel and the j-th input feature.
[0057] Step 203: Flatten the four-dimensional weight matrix in the output channel direction to obtain the two-dimensional weight matrix corresponding to each layer.
[0058] In some embodiments, the columns in the flattened two-dimensional weight matrix represent output channels. The elements in the i-th column are all the weights corresponding to the i-th output channel, including the connection weights between the i-th output channel and each input channel.
[0059] For example, the shape of a four-dimensional weight matrix is shape=(oc, ic, h, w), where oc represents the output channel, ic represents the input channel, h represents the height of the kernel, and w represents the width of the kernel. Then the shape of the flattened two-dimensional weight matrix is shape=(oc, ic×h×w).
[0060] Step 204: Dimensionally expand the two-dimensional weight matrix in the first dimension and the second dimension respectively to obtain a first three-dimensional weight matrix and a second three-dimensional weight matrix.
[0061] For example, if the shape of the two-dimensional weight matrix is shape=(oc, ic×h×w), then the shape of the first three-dimensional weight matrix after dimensional expansion in the first dimension can be shape=(1, oc, ic×h×w), and the shape of the second weight matrix after dimensional expansion in the second dimension can be (oc, 1, ic×h×w).
[0062] Step 205: Determine the cosine similarity matrix corresponding to each layer of the network according to the first three-dimensional weight matrix and the second three-dimensional weight matrix.
[0063] In some embodiments, calculate the cosine similarity in the third dimension of the first three-dimensional weight matrix and the second three-dimensional weight matrix to obtain the cosine similarity matrix.
[0064] In some embodiments, the first three-dimensional weight matrix and the second three-dimensional weight matrix can be multiplied to obtain the cosine similarity matrix. It should be noted that in step 204, the two-dimensional weight matrix is dimensionally expanded in the first dimension and the second dimension respectively to obtain the first three-dimensional weight matrix and the second three-dimensional weight matrix, so as to quickly obtain the cosine similarity matrix by multiplying the first three-dimensional weight matrix and the second three-dimensional weight matrix.
[0065] Step 206: Based on the cosine similarity matrix, determine the similarity between the initial weights corresponding to every two output channels of each layer of the network.
[0066] It should be noted that the shape of the cosine similarity matrix is (oc, oc). In the cosine similarity matrix, the element in the m-th row and n-th column is the cosine similarity between the initial weights of the m-th output channel and the initial weights of the n-th output channel. The cosine similarity matrix is a diagonal symmetric matrix, so only the upper triangular elements need to be taken. Also, since the diagonal is the result of the self-multiplication of the initial weights of each output channel and the similarity is always 1, the elements in the upper triangular part except those on the diagonal are determined as the similarities between the initial weights corresponding to every two output channels of each layer of the network.
[0067] Step 207: Determine the first loss corresponding to the neural network model according to the similarity.
[0068] Step 208: Based on the first loss and the second loss corresponding to the predicted value, correct the initial weights in the initial neural network model.
[0069] Among them, the specific implementation forms of Step 207 and Step 208 can refer to the detailed descriptions in other embodiments of the present disclosure, and will not be specifically elaborated here.
[0070] In the embodiments of the present disclosure, the initial weights corresponding to each output channel of each layer of the network in the initial neural network model and the predicted values of the initial neural network model for the sample data are obtained. Then, the four-dimensional weight matrix corresponding to each layer of the network is obtained, and the four-dimensional weight matrix is flattened in the output channel direction to obtain the two-dimensional weight matrix corresponding to each layer of the network. The two-dimensional weight matrix is respectively dimensionally increased in the first dimension and the second dimension to obtain the first three-dimensional weight matrix and the second three-dimensional weight matrix. According to the first three-dimensional weight matrix and the second three-dimensional weight matrix, the cosine similarity matrix corresponding to each layer of the network is determined. Finally, based on the cosine similarity matrix, the similarity between the initial weights corresponding to every two output channels of each layer of the network is determined. Finally, according to the similarity, the first loss corresponding to the initial neural network model is determined, and based on the first loss and the second loss corresponding to the predicted value, the initial weights in the initial neural network model are corrected. Thus, by processing the four-dimensional weights corresponding to the network, the similarity between the initial weights corresponding to every two output channels of the network can be quickly and accurately determined, thereby improving the training efficiency of the model.
[0071] It can be understood that the training method of the neural network model provided by the present disclosure can be applied to any neural network model training scenario, such as model training scenarios for processing text, model training scenarios for processing audio, model training scenarios for processing images, model training scenarios for processing multi-modal, etc. The present disclosure does not make any limitations in this regard.
[0072] Next, taking the training of a text classification model as an example, the training process of the neural network model provided by the present disclosure will be briefly described.
[0073] It is understandable that an initial text classification model is first obtained. The initial text classification model includes initial weights corresponding to each layer of the network. The similarity between the initial weights corresponding to every two output channels of each layer of the network is determined, and based on the similarity, the first loss corresponding to the initial neural network model is determined. Then, the training data of the text classification model is obtained. The training data includes the text to be classified and the label value corresponding to the text to be classified. The text to be classified is input into the initial text classification model to obtain the classification result (i.e., the predicted value) output by the initial text classification model. Based on the difference between the predicted value and the label value, the second loss is determined. Finally, based on the first loss and the second loss, the initial weights in the initial text classification model are corrected.
[0074] It should be noted that the above example is only illustrative and cannot be used to limit the training process of the neural network model in the embodiments of the present disclosure.
[0075] To implement the above embodiments, the present disclosure also proposes a training device for a neural network model.
[0076] Figure 3 The structural schematic diagram of the training device for the neural network model provided by the embodiments of the present disclosure.
[0077] As Figure 3 shown, the training device 300 for the neural network model may include:
[0078] An acquisition module 301, configured to acquire the initial weights corresponding to each output channel of each layer of the initial neural network model and the predicted value of the initial neural network model for the sample data;
[0079] A first determination module 302, configured to determine the similarity between the initial weights corresponding to every two output channels of each layer of the network;
[0080] A second determination module 303, configured to determine the first loss corresponding to the initial neural network model according to the similarity;
[0081] A correction module 304, configured to correct the initial weights in the initial neural network model based on the first loss and the second loss corresponding to the predicted value.
[0082] In some embodiments, the first determination module 302 is configured to:
[0083] Acquire a four-dimensional weight matrix corresponding to each layer of the network, where the first dimension in the four-dimensional weight matrix represents the output channel, the second dimension represents the input channel, the third dimension represents the height of the kernel, and the fourth dimension represents the width of the kernel;
[0084] Flatten the four-dimensional weight matrix in the output channel direction to obtain the two-dimensional weight matrix corresponding to each layer of the network;
[0085] Dimensionality-expand the two-dimensional weight matrix in the first and second dimensions respectively to obtain the first three-dimensional weight matrix and the second three-dimensional weight matrix;
[0086] Determine the cosine similarity matrix corresponding to each layer of the network according to the first three-dimensional weight matrix and the second three-dimensional weight matrix;
[0087] Based on the cosine similarity matrix, determine the similarity between the initial weights corresponding to every two output channels of each layer of the network.
[0088] In some embodiments, the second determination module 303 is configured to:
[0089] Determine the third loss corresponding to the similarity between every two output channels of each layer of the network;
[0090] Determine the average value of the third losses corresponding to each layer of the network as the first loss corresponding to the initial neural network model.
[0091] In some embodiments, the second determination module 303 is configured to:
[0092] Determine the average value of the similarities between every two output channels of each layer of the network as the third loss; or,
[0093] Determine the squared error of the similarities between every two output channels of each layer of the network as the third loss; or,
[0094] Determine the absolute deviation of the similarities between every two output channels of each layer of the network as the third loss.
[0095] In some embodiments, the correction module 304 is configured to:
[0096] Determine the first weight corresponding to the first loss and the second weight corresponding to the second loss;
[0097] Based on the first weight and the second weight, weight the first loss and the second loss to obtain the target loss;
[0098] Based on the target loss, correct the initial weights in the initial neural network model until the target loss is less than the loss threshold.
[0099] In some embodiments, the sample data is at least one of audio data, text data, and image data.
[0100] For the functions and specific implementation principles of the above-mentioned modules in the embodiments of the present disclosure, reference may be made to the above-mentioned method embodiments, and details are not described herein again.
[0101] The training device of the neural network model according to the embodiments of the present disclosure first obtains the initial weights corresponding to each output channel of each layer of the network in the initial neural network model and the predicted values of the initial neural network model for the sample data. Then, it determines the similarity between the initial weights corresponding to every two output channels of each layer of the network, and determines the first loss corresponding to the initial neural network model according to the similarity. Finally, based on the first loss and the second loss corresponding to the predicted values, it corrects the initial weights in the initial neural network model. Thus, the similarity between the initial weights corresponding to every two output channels of each layer of the network can be added to the loss of the model to train the model, so that the similarity between the initial weights corresponding to every two output channels of each layer of the network is small, and then different channels of each layer of the neural network can learn more different features, thereby avoiding feature redundancy and improving the performance of the model.
[0102] To implement the above embodiments, the present disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the training method of the neural network model proposed in the foregoing embodiments of the present disclosure.
[0103] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the training method of the neural network model proposed in the foregoing embodiments of the present disclosure.
[0104] Figure 4 The block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure is shown. Figure 4 The shown electronic device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0105] As Figure 4 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0106] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.
[0107] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and nonvolatile media, removable and non-removable media.
[0108] Memory 28 can include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Electronic device 12 can further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 4 not shown, typically referred to as a "hard disk drive"). Although Figure 4 not shown in the figure, a disk drive for reading and writing on a removable nonvolatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable nonvolatile optical disk (such as Compact Disc Read Only Memory (CD-ROM), Digital Video Disc Read Only Memory (DVD-ROM), or other optical media) can be provided. In these cases, each drive can be connected to bus 18 through one or more data media interfaces. Memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present disclosure.
[0109] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in a memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in this disclosure.
[0110] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0111] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.
[0112] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0113] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0114] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations, where functions may be performed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain.
[0115] Logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.
[0116] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0117] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0118] In addition, in each of the embodiments of the present disclosure, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0119] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A training method for a neural network model, characterized in that, The method includes: obtaining the initial weights corresponding to each output channel of each layer in the initial neural network model, and the predicted values of the initial neural network model for the sample data; determining the similarity between the initial weights corresponding to every two output channels of each layer of the network; determining a first loss corresponding to the initial neural network model according to the similarity; correcting the initial weights in the initial neural network model based on the first loss and a second loss corresponding to the predicted values.
2. The method according to claim 1, wherein The determining the similarity between the initial weights corresponding to every two output channels of each layer of the network includes: obtaining a four-dimensional weight matrix corresponding to each layer of the network, where the first dimension in the four-dimensional weight matrix represents the output channel, the second dimension represents the input channel, the third dimension represents the height of the kernel, and the fourth dimension represents the width of the kernel; flattening the four-dimensional weight matrix in the output channel direction to obtain a two-dimensional weight matrix corresponding to each layer of the network; performing dimension elevation on the two-dimensional weight matrix in the first and second dimensions respectively to obtain a first three-dimensional weight matrix and a second three-dimensional weight matrix; determining a cosine similarity matrix corresponding to each layer of the network according to the first three-dimensional weight matrix and the second three-dimensional weight matrix; determining the similarity between the initial weights corresponding to every two output channels of each layer of the network based on the cosine similarity matrix.
3. The method according to claim 1, wherein The determining a first loss corresponding to the initial neural network model according to the similarity includes: determining a third loss corresponding to the similarity between every two output channels of each layer of the network; determining the average value of the third losses corresponding to each layer of the network as the first loss corresponding to the initial neural network model.
4. The method according to claim 3, wherein The determining a third loss corresponding to the similarity between every two output channels of each layer of the network includes: determining the average value corresponding to the similarity between every two output channels of each layer of the network as the third loss; or determining the squared error corresponding to the similarity between every two output channels of each layer of the network as the third loss; or determining the absolute deviation corresponding to the similarity between every two output channels of each layer of the network as the third loss.
5. The method according to claim 1, characterized in that The correcting the initial weights in the initial neural network model based on the first loss and a second loss corresponding to the predicted values includes: determining a first weight corresponding to the first loss and a second weight corresponding to the second loss; weighting the first loss and the second loss based on the first weight and the second weight to obtain a target loss; correcting the initial weights in the initial neural network model based on the target loss until the target loss is less than a loss threshold.
6. The method according to any one of claims 1-5, characterized in that, The sample data is at least one of audio data, text data, and image data.
7. A training device for a neural network model, characterized in that, The device includes: an obtaining module, configured to obtain the initial weights corresponding to each output channel of each layer in the initial neural network model, and the predicted values of the initial neural network model for the sample data; The first determination module is configured to determine the similarity between the initial weights corresponding to every two output channels of each layer of the network; The second determination module is configured to determine the first loss corresponding to the initial neural network model according to the similarity; The correction module is configured to correct the initial weights in the initial neural network model based on the first loss and the second loss corresponding to the predicted value.
8. The device according to claim 7, characterized in that, The first determination module is configured to: Obtain a four-dimensional weight matrix corresponding to each layer of the network, where the first dimension in the four-dimensional weight matrix represents the output channel, the second dimension represents the input channel, the third dimension represents the height of the kernel, and the fourth dimension represents the width of the kernel; Flatten the four-dimensional weight matrix in the output channel direction to obtain a two-dimensional weight matrix corresponding to each layer of the network; Increase the dimensions of the two-dimensional weight matrix in the first dimension and the second dimension respectively to obtain a first three-dimensional weight matrix and a second three-dimensional weight matrix; Determine a cosine similarity matrix corresponding to each layer of the network according to the first three-dimensional weight matrix and the second three-dimensional weight matrix; Based on the cosine similarity matrix, determine the similarity between the initial weights corresponding to every two output channels of each layer of the network.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the training method of the neural network model according to any one of claims 1-6.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method of the neural network model according to any one of claims 1-6.