Training method and training device of sparse neural network, and computer readable storage medium
By dividing the phases in sparse neural network training, and combining the relative importance of connection weights and multiple distributions, the problem of error removal of connections in sparse neural network training is solved, and the training effect and computing efficiency are improved.
Patent Information
- Application Number
- CN202510722625.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-02
AI Technical Summary
During the sparse neural network training process, how to set a reasonable density reduction curve to reduce the probability of mistakenly removing critical connections and improve training effect.
In multiple rounds of training of sparse neural networks, the first and second stages are divided into the first stage and the second stage according to the training rounds. The amount of sparse change is positively correlated with the training round in the first stage, and the second stage is negatively correlated with the training round. Combining the relative importance of the weight of the connection and the multiple distribution, the probability and number of connection removal are determined.
It reduces the probability of erroneously removing critical connections during sparse neural network training, and improves training effect and computing efficiency.
Smart Images

Figure CN120579595A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a sparse neural network training method, a training device, a computer-readable storage medium, and a computer program product. Background Art
[0002] A neural network is a machine learning model that simulates biological neural systems, enabling complex pattern recognition and data modeling through multi-layered nonlinear transformations. With the advancement of computer technology, neural networks are being used in a growing number of fields, such as image classification, natural language processing, speech recognition, and predictive analytics.
[0003] Compared with fully connected neural networks, sparse neural networks can reduce the number of neuronal connections in the neural network, which not only reduces the risk of overfitting, but also reduces the computational complexity of the neural network and improves computational efficiency, thereby achieving efficient training and deployment.
[0004] In the related art, when training a sparse neural network, it is usually necessary to gradually reduce the number of connections between neurons, that is, to reduce the density of the neural network, that is, to increase the sparsity of the neural network. However, in this process, how to set a reasonable density reduction curve has become an unresolved problem. Summary of the Invention
[0005] In view of this, the present disclosure provides a sparse neural network training method, which can reduce the probability of erroneous removal of key connections during sparse neural network training by setting a reasonable density reduction curve, thereby improving the training effect.
[0006] According to one aspect of the present disclosure, a training method for a sparse neural network is proposed, comprising: in each round of multiple rounds of training of the sparse neural network, determining the target sparsity of the sparse neural network in the round of training according to the training round of the sparse neural network; determining the sparsity change of the sparse neural network in the round of training according to the target sparsity of the sparse neural network in the round of training and the sparsity of the sparse neural network after the previous round of training; and removing a portion of connections in the sparse neural network according to the sparsity change, wherein the multiple rounds of training include a first stage and a second stage following the first stage, the sparsity change is positively correlated with the training round in the first stage, and negatively correlated with the training round in the second stage.
[0007] In some embodiments, the first stage is from the beginning to the intermediate training round of the multiple rounds of training, and the second stage is from the intermediate training round to the end. In each round of the multiple rounds of training of the sparse neural network, determining the target sparsity of the sparse neural network in the round of training according to the training round of the sparse neural network includes: determining the target sparsity of the sparse neural network in the round of training according to the difference between the training round and the intermediate training round.
[0008] In some embodiments, determining the target sparsity of the sparse neural network in the round of training based on the difference between the training round and the intermediate training round includes: determining the target sparsity of the sparse neural network in the round of training based on the difference and a first parameter, the first parameter representing the degree of dependence of the target sparsity on the difference.
[0009] In some embodiments, removing a portion of the connections in the sparse neural network based on the sparsity change includes: removing a portion of the connections in the sparse neural network based on the sparsity change and the relative importance of the weights of each connection in the sparse neural network, wherein the relative importance is the ratio of the weight of each connection to the total weight of all connections that have a common node with the connection.
[0010] In some embodiments, removing a portion of the connections in the sparse neural network according to the relative importance of the sparsity change and the weights of each connection in the sparse neural network includes: determining the number of connections that need to be removed in the sparse neural network according to the sparsity change; determining the probability of removing each connection based on a multinomial distribution according to the relative importance of the weights of each connection in the sparse neural network; and removing a portion of the connections in the sparse neural network according to the number of connections that need to be removed and the probability of removing each connection.
[0011] In some embodiments, determining the probability of removing each connection based on a multinomial distribution according to the relative importance of the weights of the connections in the sparse neural network includes: determining a second parameter according to the training rounds of the sparse neural network, wherein the second parameter characterizes the degree of dependence of the probability of removing the connection on the relative importance of the connection, and the second parameter is positively correlated with the training rounds of the sparse neural network; determining the probability of removing each connection based on the multinomial distribution according to the relative importance and the second parameter.
[0012] In some embodiments, the training method further includes: obtaining an initial sparse neural network; and determining a sparse neural network for the multiple rounds of training based on the initial sparse neural network.
[0013] In some embodiments, obtaining an initial sparse neural network includes: constructing a ring lattice, the ring lattice including nodes of a first type and nodes of a second type, wherein each node of the first type forms a node pair with a node of the second type; generating initial connections between nodes of the first type and adjacent nodes of the second type, wherein the number of connections generated by all nodes is the same; randomly removing a portion of the initial connections according to a specified ratio, and randomly generating new connections within a portion of node pairs that do not have mutual connections, to obtain a bipartite small-world network as the initial sparse neural network.
[0014] In some embodiments, obtaining an initial sparse neural network includes: constructing a ring lattice, the ring lattice including a first type of node and a second type of node, wherein each first type of node and a second type of node constitute a node pair; determining the number of connections to be generated for each node; for each node pair in the ring lattice, determining the probability of generating connections inside the node pair based on the distance between the first type of node and the second type of node in the node pair on the ring lattice and the number of connections to be generated between the nodes in the node pair; based on the probability of generating connections, generating connections inside a part of the node pairs to obtain a binary neural network as the initial sparse neural network.
[0015] In some embodiments, determining the probability of generating a connection inside the node pair based on the distance between the first type of node and the second type of node in the node pair on the ring lattice and the number of connections to be generated between the nodes in the node pair includes: determining an initialization score of the node pair based on the distance and a third parameter, wherein the third parameter characterizes the degree of dependence of the initialization score on the distance; and determining the probability of generating a connection inside the node pair based on the initialization score and the number of connections to be generated.
[0016] In some embodiments, determining the number of connections to be generated for each node includes: determining the number of connections to be generated for each node based on the sparsity of a preset initial sparse neural network, wherein the number of connections generated by all nodes is the same; or determining the total number of connections to be generated for all nodes based on the sparsity of the preset initial sparse neural network, and determining the number of connections to be generated for each node based on a uniform distribution according to the total number.
[0017] In some embodiments, obtaining an initial sparse neural network includes: obtaining a unipartite scale-free network; randomly selecting a preset number of nodes in the unipartite scale-free network as nodes of the first type and the remaining nodes as nodes of the second type to obtain an intermediate network; removing conflicting connections in the intermediate network, wherein the conflicting connections are connections between nodes of the same type; generating connections between nodes of the first type and nodes of the second type based on the number of conflicting connections to obtain a bipartite scale-free network as the initial sparse neural network, wherein the number of nodes with the same number of connections in the bipartite scale-free network is the same as that in the intermediate network.
[0018] In some embodiments, the conflicting connection includes a first conflicting connection and a second conflicting connection, the first conflicting connection is a connection between nodes of the first type, and the second conflicting connection is a connection between nodes of the second type. Based on the number of the conflicting connections, generating a connection between nodes of the first type and nodes of the second type includes: when the number of the first conflicting connections is the same as the number of the second conflicting connections, randomly generating a connection between the first type of node corresponding to the first conflicting connection and the second type of node corresponding to the second conflicting connection, wherein the number of connections generated for each node is the same as the number of conflicting connections removed by each node.
[0019] In some embodiments, the training method further comprises: updating the connections in the sparse neural network according to the weights of the connections between the nodes in the sparse neural network and the number of connections of at least one common neighbor node of the node pairs that are not connected to each other.
[0020] In some embodiments, the training method further includes: before removing a portion of the connections in the sparse neural network in each round of training, training the weights of each connection in the sparse neural network.
[0021] In some embodiments: the input of the sparse neural network is a first image, and the output is at least one of the category of the first image, the object in the first image, the second image generated based on the first image, and the text generated based on the first image; or the input of the sparse neural network is a first text, and the output is at least one of the category of the first text, the element in the first text, the second text generated based on the first text, and the image generated based on the first text.
[0022] According to another aspect of the present disclosure, a training device for a sparse neural network is proposed, comprising: a first determination module configured to determine, in each round of multiple rounds of training of the sparse neural network, a target sparsity of the sparse neural network in the round of training according to the training round of the sparse neural network; a second determination module configured to determine a sparsity change of the sparse neural network in the round of training according to the target sparsity of the sparse neural network in the round of training and the sparsity of the sparse neural network after the previous round of training; and a removal module configured to remove a portion of connections in the sparse neural network according to the sparsity change, wherein the multiple rounds of training include a first stage and a second stage, the first stage is before the second stage, the sparsity change is positively correlated with the training round in the first stage, and negatively correlated with the training round in the second stage.
[0023] According to another aspect of the present disclosure, a training device for a sparse neural network is proposed, comprising: at least one memory; and at least one processor coupled to the at least one memory, wherein the at least one processor is configured to execute the training method as described in any of the preceding items based on instructions stored in the at least one memory.
[0024] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the training method as described in any of the above items is implemented.
[0025] According to another aspect of the present disclosure, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to implement the training method as described in any one of the preceding items.
[0026] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0028] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:
[0029] Figure 1 A flowchart showing a training method according to some embodiments of the present disclosure is shown;
[0030] Figure 2 A flowchart illustrating connection removal according to an embodiment of the present disclosure;
[0031] Figure 3 A flowchart showing the initialization of a sparse neural network according to some embodiments of the present disclosure is shown;
[0032] Figure 4 A schematic diagram illustrating initialization of a sparse neural network according to some embodiments of the present disclosure is shown;
[0033] Figure 5 A schematic diagram illustrating connection reconnection according to some embodiments of the present disclosure;
[0034] Figure 6A A flowchart showing initialization of a sparse neural network according to some other embodiments of the present disclosure is shown;
[0035] Figure 6B A schematic diagram illustrating initialization of a sparse neural network according to some other embodiments of the present disclosure;
[0036] Figure 6C A schematic diagram illustrating a connection state in an initialized sparse neural network according to some embodiments of the present disclosure;
[0037] Figure 6D A schematic diagram showing the connection state in an initialized sparse neural network according to some other embodiments of the present disclosure;
[0038] Figure 7A A schematic diagram showing the network structure of a scale-free network;
[0039] Figure 7B A flowchart showing initialization of a sparse neural network according to some further embodiments of the present disclosure is shown;
[0040] Figure 8 A schematic diagram illustrating initialization of a sparse neural network according to some further embodiments of the present disclosure;
[0041] Figure 9 A schematic diagram illustrating a training method according to an embodiment of the present disclosure;
[0042] Figure 10 A block diagram illustrating a training device according to some embodiments of the present disclosure is shown;
[0043] Figure 11 A block diagram showing a training device according to some other embodiments of the present disclosure;
[0044] Figure 12 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0045] It should be understood that the size of each part shown in the drawings is not drawn according to the actual proportional relationship.In addition, the same or similar reference numerals represent the same or similar components. DETAILED DESCRIPTION
[0046] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein.
[0047] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement and numerical values of the components and steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.
[0048] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the concepts of "first", "second", etc. are not intended to imply that the objects described in this way must be in a given order in time, space, ranking or any other way. It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly defined or given to the contrary in the context.
[0049] All terms (including technical or scientific terms) used in this disclosure have the same meaning as those understood by one of ordinary skill in the art to which this disclosure belongs, unless otherwise specifically defined. It should also be understood that terms defined in, for example, general dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or highly formal sense, unless explicitly defined herein.
[0050] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0051] In the related art, when training sparse neural networks, it is often possible to optimize the computational efficiency and model performance of the neural network by reducing the number of connections between neurons and retaining only some critical connections. In other words, the density of the neural network can be reduced, that is, the sparsity of the neural network can be increased. However, in this process, there is the problem of incorrectly removing critical connections.
[0052] In view of this, the present disclosure proposes a sparse neural network training method, which can reduce the probability of erroneous removal of key connections during sparse neural network training, thereby improving the training effect.
[0053] Next, we will first combine Figure 1 The training method in the present disclosure is described. Figure 1 A flowchart of a training method according to some embodiments of the present disclosure is shown.
[0054] like Figure 1 As shown, the training method may include: step S11, in each round of multiple rounds of training of the sparse neural network, determining the target sparsity of the sparse neural network in the round of training according to the training round of the sparse neural network; step S12, determining the sparsity change of the sparse neural network in the round of training according to the target sparsity of the sparse neural network in the round of training and the sparsity of the sparse neural network after the previous round of training; step S13, removing a part of the connections in the sparse neural network according to the sparsity change.
[0055] In the training method, the multiple rounds of training include a first stage and a second stage following the first stage, the sparsity change is positively correlated with the training rounds in the first stage and negatively correlated with the training rounds in the second stage.
[0056] In step S11, the target sparsity that the sparse neural network needs to achieve in this round of training can be determined, where sparsity refers to the ratio of node pairs that are not interconnected to all node pairs in the sparse neural network that may generate connections.
[0057] For example, if the sparse neural network is a single-point neural network, that is, all nodes may be connected. In the case that the sparse neural network includes 5 nodes, all the node pairs that may be connected are Assuming that 3 connections are generated in the sparse neural network, there are 7 pairs of nodes that are not connected to each other, and the sparsity of the sparse neural network is 0.7, or 70%.
[0058] For another example, if the sparse neural network is a bipartite neural network, connections can only be generated between nodes of different types. In the case where the sparse neural network includes 5 nodes of the first type and 5 nodes of the second type, the total number of node pairs that can be connected is 5*5=25. Assuming that 3 connections are generated in the sparse neural network, there are 22 node pairs that are not connected to each other, and the sparsity of the sparse neural network is 0.88, or 88%.
[0059] In addition, the density of a sparse neural network is usually complementary to the sparsity of a sparse neural network. The density of a sparse neural network refers to the ratio of node pairs that are connected to each other to all node pairs in the sparse neural network. For example, in the example of a single-part sparse neural network, the number of connections that have been generated is 3, and the density of the sparse neural network is 0.3, or 30%. It can be seen that an increase in the sparsity of a sparse neural network is equivalent to a decrease in the density of the sparse neural network.
[0060] In this field, in addition to the above description method, the sparsity and density of a sparse neural network can also be described by the proportion of zero weights and the proportion of non-zero weights in the sparse neural network, where the weight of the connection between node pairs is zero (or inactive connection), which is equivalent to the absence of mutual connection, and the weight of the connection between node pairs is non-zero (or active connection), which is equivalent to the existence of mutual connection.
[0061] In the above steps, the target sparsity can be determined according to the training round of the sparse neural network. In other words, the target density to be achieved in the training round of the sparse neural network in the process of decreasing the density can be determined.
[0062] In step S12, the sparsity change required for the sparse neural network in this round of training can be determined based on the difference between the target sparsity and the sparsity of the sparse neural network after the previous round of training.
[0063] In some embodiments, the density change of the sparse neural network can be based only on the density reduction step. In this case, the sparsity change can also be determined directly based on the difference between the target sparsity of the sparse neural network in the current round of training and the target sparsity of the sparse neural network in the previous round of training.
[0064] In the training method disclosed herein, multiple rounds of training are divided into a first phase and a second phase following the first phase. By setting a method for calculating the target sparsity, the change in sparsity can be made positively correlated with the number of training rounds in the first phase and negatively correlated with the number of training rounds in the second phase, i.e., the change in sparsity gradually increases in the first phase of the training process and gradually decreases in the second phase.
[0065] In traditional training methods of related art, such division is usually not performed. For example, the change in sparsity gradually decreases with each training round throughout the training process. It can be seen that compared with traditional training methods, the training method proposed in this disclosure can reduce the number of connections removed in the early stages of training.
[0066] In the early stages of training, the weights of each connection in a sparse neural network haven't been fully trained yet. These weights are relatively random and don't accurately reflect the importance of the connection. Therefore, there's a high probability of incorrectly removing critical connections in the early stages of training. By reducing the number of connections removed during the early stages of training, we can reduce the probability of incorrectly removing critical connections during sparse neural network training, thereby improving training effectiveness.
[0067] In step S13, a portion of the connections in the sparse neural network may be removed according to the sparsity change, so that the sparsity of the sparse neural network reaches the target sparsity, completing the density reduction step in this round of training.
[0068] The change in sparsity during each training round represents the number of connections that need to be removed during that training round to reduce the density of the sparse neural network. For example, if the target sparsity for this training round is 0.8, and the sparsity of the sparse neural network after the previous training round is 0.7, the change in sparsity is 0.1. If the total number of node pairs that can generate connections in the sparse neural network is 10, then 1 connection needs to be removed to achieve the above change in sparsity.
[0069] Combined with the above Figure 1 This article describes the basic process of training methods according to some embodiments of the present disclosure. By reducing the number of connections removed during the initial training phase, the probability of incorrectly removing key connections during sparse neural network training can be reduced, thereby improving training effectiveness. The following examples will provide a detailed description of how to determine the target sparsity and the amount of sparsity change in the training method.
[0070] In some embodiments, the first stage is from the beginning to the middle training round of the multiple rounds of training, and the second stage is from the middle training round to the end, wherein the middle training round is, for example, the median round of the multiple rounds of training, that is, the middle training round is equal to half of the sum of the beginning round and the ending round of the multiple rounds of training.
[0071] Based on the above-mentioned intermediate training rounds, in each round of multiple rounds of training of the sparse neural network, determining the target sparsity of the sparse neural network in the round of training according to the training round of the sparse neural network may include: determining the target sparsity of the sparse neural network in the round of training according to the difference between the training round and the intermediate training round.
[0072] For example, according to the training round of the sparse neural network, the expression for determining the target sparsity of the sparse neural network in the training round can be:
[0073]
[0074] Among them, t is the training round of this round of training, t mis the middle training round of multiple training rounds, s i is the initial sparsity of the sparse neural network before multiple rounds of training, s f is the final sparsity of the sparse neural network after multiple rounds of training, s t is the target sparsity of the sparse neural network in the tth round of training.
[0075] In the above expression, the intermediate training rounds can be set to Among them, t0 is the starting round of the above multiple rounds of training, t f The end round of multiple training rounds.
[0076] The starting round t0 and the ending round t of multiple rounds of training f It can be set according to the actual training situation. For example, if the complete training process of the sparse neural network requires 1000 rounds of training, the multi-round training with decreasing density can start from the 1st round and end at the 1000th round, covering the entire training process. For another example, the multi-round training can also start from the 50th round and end at the 950th round, thereby avoiding the removal of connections in the earliest stage when the weights of the sparse neural network are the most random and the final stage when the topology of the sparse neural network is mature, reducing the probability of incorrectly removing key connections during sparse neural network training, and thus improving the training effect.
[0077] The initial sparsity s of the sparse neural network i and the end sparsity s f Can be respectively compared with the starting training round t0 and the ending training round t f Corresponding. The starting sparsity can be determined based on the starting state of the sparse neural network during the density reduction process, for example, the sparsity of the sparse neural network obtained by initialization. The ending sparsity can be determined based on the ending state of the sparse neural network during the density reduction process, for example, the sparsity required for the sparse neural network when used after training. It can be understood as the target sparsity of the entire density reduction process.
[0078] The change in sparsity during each round of training is Δs=s t -s t-1 , which can also be understood as the derivative of the target sparsity with respect to the training round because The value range of is (0,1), and the above sparsity change is When t=t m From the above expression, it can be seen that the change in sparsity is in the middle training round t m It gradually increased before, and in the middle training round t m It gradually decreased before.
[0079] In related technologies, however, there is no division between the first and second stages. During the training process, the change in sparsity usually decreases gradually with the training rounds. The expression of the target sparsity is usually:
[0080]
[0081] Among them, t is the training round of this round of training, t′0 is the starting round of training, and t′ f is the end round of training, s′ i is the sparsity of the sparse neural network at the beginning of training, s′ f is the sparsity of the sparse neural network at the end of training, s′ t is the target sparsity of the sparse neural network in the tth round of training.
[0082] In related technologies, the sparsity change in each round of training is Δs′=s′ t -s′ t-1 , which can also be understood as the derivative of the target sparsity with respect to the training round because The value of is also (0,1). The above sparsity change gradually decreases with the training rounds, and the number of connections removed in the early stage of training is relatively large.
[0083] It can be seen from the above examples that the training method proposed in the present disclosure can reduce the number of connections removed in the early stage of training, thereby reducing the probability of erroneous removal of key connections when the weights are relatively random in the early stage of training, and improving the training effect.
[0084] It should be noted that the starting sparsity s in the above example expressions of the present disclosure is i The sparsity s′ of the sparse neural network at the beginning of training in the above example expression of the related art is i The meanings are similar, but not completely identical. It can be seen from the expression that when t=t0, the target sparsity of the present disclosure is In the case of t=t′0, the target sparsity s′0=s′ in the related art i Considering that the total number of rounds in the training process of sparse neural networks is usually high, t m The value is also relatively high. In the above case, the target sparsity s0 can be approximated to the starting sparsity s i In addition, it is also possible to adjust the initial sparsity s i The value of , to make the target sparsity is consistent with the actual sparsity of the sparse neural network at the beginning of training. Similarly, the ending sparsity s f and the sparsity s′ of the sparse neural network at the end of training fhas similar distinctions and can be used to determine the final sparsity s f Using similar processing, the final sparsity s f This is consistent with the expected sparsity of a sparse neural network at the end of training.
[0085] Furthermore, the training cost of sparse neural networks is typically linearly correlated with density, for example, FLOPs (Floating Point Operations) can be used as a proxy for training cost. Therefore, in the aforementioned training process, the overall training cost is also linearly correlated with the integral of the target sparsity over the number of training rounds.
[0086] In the training method proposed in this disclosure, the integral of the target sparsity over the training rounds is:
[0087]
[0088] In related technologies, the integral of target sparsity over training rounds is:
[0089]
[0090] As mentioned above, when the total number of training rounds of sparse neural networks is high, the target sparsity at the beginning of training can be approximated as the starting sparsity s i , similarly, the target sparsity at the end of training can be approximated as the end sparsity s f , therefore, (s f -s i ) represents the total sparsity change during the training process, and (s′ f -s′ i ) are corresponding.
[0091] Therefore, the difference in the above sparsity integral is mainly reflected in the second coefficient. The coefficient of the sparsity integral of the training method proposed in this disclosure is 1 / 2, while the coefficient of the sparsity integral in the related art is 1 / 4. In order to make the computational consumption in the training process uniform, the number of rounds (t f -t0) is set to the round number (t′) of the density reduction process in the related art f - half of t′0).
[0092] Under the above-mentioned setting conditions, the performance experimental results of the sparse neural network are shown in Table 1 below. Table 1 shows the performance scores of the sparse neural network after training using the training method proposed in the present disclosure and the training method in the related art in an experimental sparse neural network with an initial sparsity of 50%, respectively. The perplexity is used to characterize the uncertainty of the sparse neural network to the data. The lower the perplexity, the more accurate the sparse neural network's prediction of the data. As can be seen from Table 1, the training method proposed in the present disclosure can improve the effect of sparse neural network training by reducing the number of connections removed in the early stage of training, thereby reducing the probability of erroneous removal of key connections when the weights are relatively random in the early stage of training.
[0093]
[0094] Table 1
[0095] In some embodiments, determining the target sparsity of the sparse neural network in the round of training based on the difference between the training round and the intermediate training round may include: determining the target sparsity of the sparse neural network in the round of training based on the difference and a first parameter, the first parameter representing the degree of dependence of the target sparsity on the difference.
[0096] Continuing with the above example, the expression for determining the target sparsity of the sparse neural network in this round of training may include a first parameter k:
[0097]
[0098] The curvature of the function can be adjusted by the first parameter k, thereby further adjusting the rate of density decrease during the training process. The first parameter can be determined based on parameters such as the sparsity of the sparse neural network and / or the total number of rounds of the training process. For example, when the initial sparsity of the sparse neural network is high, it means that the number of connections in the sparse neural network is small at this time, and the probability of incorrectly removing key connections is relatively high. By setting a larger first parameter k, the number of connections removed in the early stage of training can be further reduced, thereby improving the removal accuracy.
[0099] It should be understood that the above setting is only exemplary and not restrictive. For example, the target sparsity can be adjusted after a certain number of rounds of training by using a rounding function to improve the stability of the training.
[0100] The above describes how to determine the target sparsity and the sparsity change in the training method according to the embodiment of the present disclosure. The following describes how to remove some connections in the sparse neural network after determining the sparsity change.
[0101] In some embodiments, removing a portion of the connections in the sparse neural network based on the sparsity change includes: removing a portion of the connections in the sparse neural network based on the sparsity change and the relative importance of the weights of each connection in the sparse neural network, wherein the relative importance is the ratio of the weight of each connection to the total weight of all connections that have a common node with the connection.
[0102] The relative importance of the connection weight can be understood as the relative importance of the connection among all the connections of its node. The calculation expression is, for example,
[0103]
[0104] Among them, R ab is the relative importance of the connection between node a and node b, |W ab | is the weight of the connection between node a and node b, C a is the set of all nodes connected to node a, i is the set C a Nodes in C b is the set of all nodes connected to node b, j is the set C b Nodes in ∑ i |W ai | is the total weight of all connections corresponding to node a, ∑ j |W jb | is the total weight of all connections corresponding to node b.
[0105] Compared with using the absolute value of the connection weight as the removal criterion, using the relative importance of the weight of each connection in the sparse neural network as the removal criterion can avoid removing all connections of a node in the sparse neural network, making the node unable to input or output normally in the sparse neural network.
[0106] In addition, some sparse neural networks may include power-law distribution characteristics, where most connections are concentrated on a small number of nodes. The distribution of connection weights in the above sparse neural networks may also be uneven. The absolute value of the connection weight cannot accurately reflect the importance of the connection in the sparse neural network. Using the relative importance of the connection weight as the removal criterion can improve the accuracy of connection removal.
[0107] In some embodiments, the training method further includes: before removing a portion of the connections in the sparse neural network in each round of training, training the weights of each connection in the sparse neural network.
[0108] Before removing the connections, the weights of the connections in the sparse neural network can be trained and updated. For example, this may include: inputting sample data into the sparse neural network to generate a prediction result; calculating the prediction error based on the difference between the prediction result and the true value corresponding to the sample data through a loss function; and using an optimization algorithm to update the weights of the connections in the sparse neural network to minimize the prediction error.
[0109] For example, the weights of correct connections in a sparse neural network can be increased, and the weights of incorrect connections can be reduced, so that in the process of connection removal, incorrect connections can be removed and correct connections can be retained for subsequent connection generation.
[0110] The above training of the weights of each connection in the sparse neural network can be performed for a preset number of iterations, thereby improving the prediction performance of the sparse neural network and optimizing the weights of the connections in the sparse neural network for subsequent connection removal.
[0111] Next, we will combine Figure 2 This section introduces how to remove connections based on the relative importance of their weights in an embodiment of the present disclosure. Figure 2 A flowchart illustrating connection removal according to an embodiment of the present disclosure is shown.
[0112] like Figure 2 As shown, removing a portion of the connections in the sparse neural network according to the relative importance of the sparsity change and the weights of each connection in the sparse neural network includes: step S131, determining the number of connections that need to be removed in the sparse neural network according to the sparsity change; step S132, determining the probability of removing each connection based on multinomial distribution according to the relative importance of the weights of each connection in the sparse neural network; step S133, removing a portion of the connections in the sparse neural network according to the number of connections that need to be removed and the probability of removing each connection.
[0113] The above steps belong to step S13, which is an implementation method of removing a portion of the connections in the sparse neural network according to the sparsity change.
[0114] In step S131, the number of connections that need to be removed from the sparse neural network can be determined based on the sparsity change determined in step S12, so that the sparsity of the sparse neural network after the connections are removed is equal to the target sparsity.
[0115] In step S132 , the probability of removing a connection may be determined based on a multinomial distribution according to the relative importance of the weights of the connection.
[0116] A multinomial distribution is a discrete probability distribution that is useful when there are more than two possible outcomes per trial. When removing connections, all connections can be considered at each time a connection is removed. For example, if there are n connections in a sparse neural network, there are n possible outcomes for each connection removed. Therefore, the probability of removing each connection can be determined using a multinomial distribution.
[0117] When multiple connections need to be removed, for example, the operation of removing one connection may be performed multiple times. After each connection removal, the probability of removing each connection is re-determined, thereby achieving multiple connections removal in multiple times.
[0118] For another example, the probability of removing each combination of multiple connections can also be directly determined. Taking a sparse neural network with four connections 1, 2, 3, and 4 as an example, the combinations for removing two connections include (1,2), (1,3), (1,4), (2,3), (2,4), and (3,4). The probability of removing each combination can be determined based on the sum of the removal probabilities of the connections in each combination, thereby removing multiple connections at one time.
[0119] Compared with the binomial distribution commonly used in related technologies, the use of multinomial distribution can comprehensively consider all connections in the sparse neural network, while the binomial distribution can only consider the removal or retention of each connection individually, and cannot ensure that the number of removed connections in the sparse neural network is consistent with the required number.
[0120] In some embodiments, determining the probability of removing the connection based on the multinomial distribution includes normalizing the relative importance of the weights of each connection in the sparse neural network and determining the probability of removing each connection separately.
[0121] Specifically, the normalization expression can be:
[0122]
[0123] Among them, p ab is the probability of removing the connection between node a and node b, R ab is the relative importance of the weight of the connection between node a and node b, i and j are any two nodes that are connected to each other in the sparse neural network, ∑ ij R ij is the sum of the relative importance of the weights of all connections in a sparse neural network.
[0124] Through the above normalization process, the sum of the probabilities of all connections being removed in the sparse neural network can be made to be 1, which meets the requirements of the multinomial distribution.
[0125] In some embodiments, determining the probability of removing each connection based on a multinomial distribution according to the relative importance of the weights of the connections in the sparse neural network includes: determining a second parameter according to the training rounds of the sparse neural network, wherein the second parameter characterizes the degree of dependence of the probability of removing the connection on the relative importance of the connection, and the second parameter is positively correlated with the training rounds of the sparse neural network; determining the probability of removing each connection based on the multinomial distribution according to the relative importance and the second parameter.
[0126] In the above embodiment, the second parameter (flexibility parameter when removing a connection) can be used to adjust the degree to which the probability of removing a connection depends on the relative importance of the weight of the connection.
[0127] The second parameter is positively correlated with the training rounds of the sparse neural network, that is, a relatively low second parameter can be used in the early stage of training, and a relatively high second parameter can be used in the later stage of training.
[0128] In the early stages of training, the weights in a sparse neural network are, for example, obtained through initialization rather than training. In this case, a lower second parameter can be used to reduce the degree to which the probability of removing a connection depends on the relative importance of the weights, thereby making the removal of connections relatively random and suitable for the early stages of training.
[0129] Similarly, in the later stages of training, the weights in the sparse neural network have undergone several rounds of training and can better reflect the actual connection relationship between nodes. At this time, a higher second parameter can be used to increase the probability of removing connections and their dependence on the relative importance of weights, thereby reducing the randomness of connection removal and avoiding the removal of important connections in the sparse neural network. This is suitable for the later stages of training.
[0130] For example, in the embodiment of the normalization process described above, the expression between the probability of removing a connection and the relative importance of the weight can be rewritten as:
[0131]
[0132] in, is the expression of the second parameter, representing the probability p of removing the connection ab The degree of dependence on the relative importance R of the weights, the value of δ can be limited to [0,1), so that the value of the second parameter covers all positive real numbers.
[0133] The positive correlation between the second parameter and the training round of the sparse neural network can be set by δ, for example, δ = (t-t0) / t f , where t represents the current training round, t0 is the starting round of the above multiple training rounds, and tf The final round of multiple training rounds. By gradually adjusting δ during the training process, the value of the second parameter can be adjusted. Based on the above-mentioned density reduction process of the training method disclosed herein, not only can the number of connections removed in the early stages of training be reduced, but the accuracy of connection removal can also be further improved by adopting a more appropriate removal method, thereby improving the training effect.
[0134] The above setting method is only exemplary and not restrictive. For example, the second parameter can be adjusted after each certain round of training to improve the stability of training.
[0135] The above describes how to determine the probability of removing the connection in step S132. Figure 2 , continue to introduce how to remove the connection in step S133.
[0136] In step S133, a portion of the connections in the sparse neural network may be randomly removed according to the number of connections to be removed and the removal probability of each connection in the sparse neural network.
[0137] For example, when it is determined that 10 connections need to be removed from the sparse neural network based on the sparsity change, 10 connections can be randomly selected for removal based on the removal probability of each connection in the sparse neural network.
[0138] In the related art, after determining the weights of each connection in a sparse neural network, the connections are usually removed strictly based on the absolute value of the weight. In the case where 10 connections need to be removed as mentioned above, the 10 connections with the lowest absolute values of weights will be selected for removal. Unlike the solution proposed in the present disclosure, random selection will not be performed based on the removal probability related to the relative importance of the weights.
[0139] In the early stages of training, when the weights of connections in a sparse neural network are relatively random, removing connections strictly according to the absolute values of the weights may result in many correct connections being incorrectly removed. However, the present disclosure removes connections by reflecting the probability of connection importance, which can make the removal process gentler and improve the training effect.
[0140] In some embodiments, the training method further comprises: updating the connections in the sparse neural network according to the weights of the connections between the nodes in the sparse neural network and the number of connections of at least one common neighbor node of the node pairs that are not connected to each other.
[0141] In each round of training, after density reduction, the topology of the sparse neural network can be further trained by updating the connections in the sparse neural network through connection removal and connection regeneration.
[0142] For example, we can first remove some relatively unimportant connections in the sparse neural network based on the weights of the connections between the nodes in the sparse neural network, and then regenerate the connections based on the number of connections between common neighbor nodes of the node pairs that do not have mutual connections in the sparse neural network after the connections are removed, thereby achieving the update of the connections in the sparse neural network, optimizing the topological structure of the sparse neural network, and improving the training effect of the sparse neural network.
[0143] In the above process, the connections between node pairs are generated by considering the number of connections of each common neighbor node separately. There is no need to record the various paths between the node pairs, and repeated calculations are avoided. This can reduce the complexity of the training process and the required training time.
[0144] The above describes the various steps in the training method of some embodiments of the present disclosure with examples. The training method proposed in the present disclosure can reduce the probability of incorrectly removing key connections during sparse neural network training, thereby improving the training effect.
[0145] In addition to the above steps, the training method of the present disclosure may also include network initialization. By using a brain-inspired network topology initialization method, the initial topology of the sparse neural network can be optimized, thereby improving the training effect. The following will specifically describe how to initialize a sparse neural network.
[0146] In some embodiments, the training method further includes: obtaining an initial sparse neural network; and determining a sparse neural network for the multiple rounds of training based on the initial sparse neural network.
[0147] As mentioned above, the multiple rounds of training for increasing sparsity, i.e., decreasing density, may not be performed directly on the initial sparse neural network obtained by re-initialization, but after the initial sparse neural network has been trained for, for example, 50 rounds, the training may be repeated. Figure 1 Therefore, in the training method disclosed herein, a sparse neural network for multiple rounds of training can be determined based on the initial sparse neural network, rather than directly using the initial sparse neural network, thereby improving the training effect.
[0148] The methods for obtaining the initial sparse neural network proposed in this disclosure can be divided into three types: binary small-world sparse neural network initialization, binary receptive field neural network initialization and binary scale-free sparse neural network initialization. Figures 3 to 8 The above three initialization methods are introduced in detail.
[0149] First, combine Figures 3 and 4This article describes the initialization process for a bipartite small-world network. The small-world property is a network structural characteristic characterized by high clustering and short average path length. This means that information transfer between nodes in the network can maintain close local connections while also enabling rapid global dissemination. Figure 3 A flowchart of initializing a sparse neural network according to some embodiments of the present disclosure is shown. Figure 4 A schematic diagram illustrating initialization of a sparse neural network according to some embodiments of the present disclosure.
[0150] like Figure 3 As shown, obtaining an initial sparse neural network may include: step S31, constructing a ring lattice, the ring lattice including nodes of the first type and nodes of the second type, wherein each node of the first type and a node of the second type constitute a node pair; step S32, generating initial connections between the nodes of the first type and their adjacent nodes of the second type, wherein the number of connections generated by all nodes is the same; step S33, randomly removing a portion of the initial connections according to a specified ratio, and randomly generating new connections inside a portion of node pairs that do not have mutual connections, to obtain a bipartite small-world network as the initial sparse neural network.
[0151] In step S31 , a ring lattice including two types of nodes may be constructed. The ring lattice has a regular topological structure, and the nodes in the sparse neural network are cyclically arranged in the ring lattice network.
[0152] Figure 4 (a) shows an example of a ring lattice, Figure 4 In (a), squares represent nodes A1 to A4 of the first type, and circles represent nodes B1 to B4 of the second type. Figure 4 As shown in (a), a sparse neural network can have the same number of first-type nodes and second-type nodes, and these nodes can be arranged in an interlaced and cyclic manner to construct a ring lattice.
[0153] It should be understood that the above-mentioned ring lattice shape is merely exemplary and not restrictive, and other shapes of ring lattices can also be used to construct the initial sparse neural network.
[0154] In step S32, initial connections can be generated between the nodes of the sparse neural network. By generating the same number of connections for each node, the degree of each node in the sparse neural network can be guaranteed to be the same. By connecting each node in the ring lattice to its nearest adjacent node of a different type, the sparse neural network after generating the connections can have the same periodicity and symmetry as the ring lattice.
[0155] Figure 4 (b) shows the Figure 4 (a) is an example of generating initial connections in the ring lattice. Figure 4 (b) Each node of the ring lattice can generate connections with the two nearest nodes, one node on each side, thus ensuring the symmetry of the sparse neural network.
[0156] The number of generated connections can be determined based on the desired starting sparsity of the sparse neural network. The higher the desired starting sparsity, the fewer the number of generated connections, thereby meeting the sparsity requirement. Conversely, the lower the desired starting sparsity, the more the number of generated connections.
[0157] Since nodes are connected to the nearest other nodes, the above sparse neural network has a high degree of aggregation. However, under the above connection method, the average path length between nodes in the sparse neural network is long. For example, Figure 4 In the 8-node sparse neural network shown in (b), a minimum of three connections are required from the first-type node in the upper right to the second-type node in the middle left. This results in a shortest path length of 3, which is a relatively long path length. Furthermore, this path length increases rapidly as the number of nodes in the sparse neural network increases, resulting in a long average path length between nodes in the sparse neural network. This indicates that the sparse neural network does not satisfy the small-world property.
[0158] In step S33, a portion of the initial connections may be randomly removed and new connections may be randomly generated, that is, a new connection may be formed between any pair of first-type nodes and second-type nodes that have no connection, so that the sparse neural network meets the requirements of the small-world property.
[0159] Figure 4 (c) shows that Figure 4 (b) is an example of connection rewiring in a sparse neural network. Figure 4 As shown in (c), when the ratio is specified as 0.5, for each node in the sparse neural network, 50% of the initial connections, that is, 1 connection, can be randomly removed, and connections can be randomly generated between other types of nodes that have not yet generated connections with the node, where the number of generated connections can be the same as the number of removed connections, thereby ensuring that the degree of each node in the sparse neural network is the same.
[0160] In the above process, since the neighboring nodes closest to the node have already generated connections with the node, the newly generated connections will be connections between the node and other nodes that are farther away. Figure 4In the example, for each node, since the node has already been connected to the nearest node, the shortest path length between the reconnected node and the node is at least 2, that is, the reconnection threshold is 2. Through the above reconnection, the average path length in the sparse neural network can be reduced, so that the network conforms to the small-world property.
[0161] The bipartite small-world network obtained in the above steps can be used as the initial sparse neural network for training. By pre-equipping the initial sparse neural network with certain small-world characteristics, the subsequent training effect based on the initial sparse neural network can be improved.
[0162] In some embodiments, different specified ratios can also be selected, for example, based on the type of sparse neural network. As shown above, the specified ratio represents the proportion of the number of randomly removed and randomly generated connections in the total number of connections, and the value range should be (0,1). A higher specified ratio means that more connections are randomly removed and reconnected, resulting in a shorter average path length and higher randomness in the sparse neural network, while a lower specified ratio means that fewer connections are randomly removed and reconnected, resulting in a higher clustering and regularity of the sparse neural network.
[0163] Figure 5 Schematic diagram showing connection reconnection according to some embodiments of the present disclosure. Figure 5 As shown, the adjacency matrix of the sparse neural network before connection reconnection is as follows Figure 5 (a), where black represents the absence of a connection between nodes, for example, represented as 0 in the adjacency matrix, and white represents the presence of a connection between nodes, for example, represented as 0 in the adjacency matrix. Specifically, the rows of each element in the matrix can correspond to the numbers of the first type of nodes, and the columns can correspond to the numbers of the second type of nodes. For example, the value of element (1,1) is 1, which means that a connection is generated between the first type node A1 and the second type node B1.
[0164] Figure 5 (b), (c), (d), and (e) in the figure respectively show the adjacency matrix of the sparse neural network after reconnection when the specified ratio is 0.25, 0.5, 0.75, and 1. It can be seen that as the specified ratio increases, the clustering in the sparse neural network becomes lower and the randomness becomes higher.
[0165] Depending on different situations, different specified ratios can be selected so that the initial sparse neural network used in training meets actual needs and improves the training effect.
[0166] Combined with the above Figures 3 to 5 This article introduces how to initialize the binary small world. 6A to 6DThis section describes how to initialize a binary receptive field. The receptive field refers to the area in the input space where a node in a neural network can receive input.
[0167] first, Figure 6A A flowchart of initializing a sparse neural network according to some other embodiments of the present disclosure is shown. Figure 6B A schematic diagram illustrating initialization of a sparse neural network according to some other embodiments of the present disclosure is shown.
[0168] like Figure 6A As shown, obtaining an initial sparse neural network may include: step S61, constructing a ring lattice, the ring lattice including nodes of the first type and nodes of the second type, wherein each node of the first type and a node of the second type constitute a node pair; step S62, determining the number of connections to be generated for each node; step S63, for each node pair in the ring lattice, determining the probability of generating a connection inside the node pair based on the distance between the first type node and the second type node in the node pair on the ring lattice and the number of connections to be generated between the nodes in the node pair; step S64, generating connections inside a part of the node pairs based on the probability of generating connections, to obtain a binary neural network, the initial sparse neural network.
[0169] In step S61 , similar to the process of initializing a bipartite small-world network, a ring lattice with two types of nodes may be created first.
[0170] Figure 6B (a) in FIG shows a schematic diagram of a ring lattice. Figure 6B As shown, an initial sparse neural network can be constructed using, for example, a ring lattice with two rings, wherein squares represent first type nodes A1 to A4, and circles represent second type nodes B1 to B4, so that the first type of nodes can be arranged on the first ring and the second type of nodes can be arranged on the second ring.
[0171] It should be understood that the above-mentioned ring lattice shape is merely exemplary and not restrictive, and other shapes of ring lattices can also be used to construct the initial sparse neural network.
[0172] In step S62 , the number of connections to be generated for each node may be determined based on the sparsity target to be achieved during initialization.
[0173] In some embodiments, the number of connections to be generated for each node is determined based on the sparsity of a preset initial sparse neural network, wherein the number of connections generated by all nodes is the same.
[0174] For example, in the case where the ring lattice has 10 nodes, including 5 nodes of the first type and 5 nodes of the second type, if the preset sparsity of the initial sparse neural network, that is, the sparsity target required to be achieved by initialization is 40%, then the number of connections to be generated for each node can be determined to be 2, so that the initialized sparse neural network meets the sparsity requirements.
[0175] In the above embodiment, the connection generated by each node can be fixed to the same value, thereby ensuring that the degree of each node in the sparse neural network is the same, so that the sparse neural network has uniform characteristics.
[0176] In other embodiments, the total number of connections to be generated for all nodes is determined based on the sparsity of a preset initial sparse neural network, and the number of connections to be generated for each node is determined based on a uniform distribution according to the total number.
[0177] Continuing with the above example with 10 nodes and a preset sparsity of 40%, in other embodiments, the total number of connections to be generated for all nodes may be determined to be 10. Then, the number of connections to be generated may be randomly distributed between nodes of the first type and nodes of the second type, with each node having a probability of generating a connection of 1 / 5. The number of connections to be generated for each node determined in this manner may be {1, 2, 2, 3, 2}.
[0178] In the above embodiment, on the basis of ensuring that the degree of the nodes is roughly average through uniform distribution, it is only necessary to ensure that the total number of connections to be generated by the nodes can make the initial sparse neural network meet the sparsity requirements, and there is no restriction on the number of connections to be generated for all nodes to be the same, thereby improving the flexibility of the initialization process.
[0179] In step S63 , the probability of generating a connection within each node pair may be determined according to the distance between the first type of node and the second type of node on the ring lattice and the number of connections to be generated determined in step S62 .
[0180] Reference below Figure 6B The schematic diagram of the ring lattice shown in (a) introduces the above step S63. The node pairs in the ring lattice are composed of any two nodes that can generate a connection, that is, any one node of the first type and any one node of the second type, for example Figure 6B The ring lattice shown in (a) includes 16 node pairs, namely (A1, B1), (A1, B2), (A2, B2), etc.
[0181] The distance between the first type of node and the second type of node in the node pair on the ring lattice can be calculated by the node number, for example, d ij =|ij|, where dij represents the distance between a first type node i and a second type node j on the ring lattice, i represents the number of the first type node, and j represents the number of the second type node.
[0182] Furthermore, the distance expression can also be:
[0183] d ij =min{|ij|,|(in i )-j|,|i-(jn j )|}
[0184] Among them, d ij represents the distance between the first type node i and the second type node j on the ring lattice, i represents the number of the first type node, j represents the number of the second type node, n i Indicates the total number of nodes in the ring where the first type of nodes are located on the ring crystal, n j Indicates the total number of nodes in the ring where the second type of nodes are located on the ring crystal. Figure 6B In the example shown in (a), the total number of nodes in the ring where the first type of nodes are located and the total number of nodes in the ring where the second type of nodes are located are both 4.
[0185] As an example of distance calculation, the distance d between node A1 and node B1 11 = 0, the distance d between node A1 and node B2 12 =1, the distance d between node A1 and node B4 14 =min{1-4,(1-4)-4,1-(4-4)}=1. The above distance is only a qualitative description, used to distinguish the distance between nodes in the neural network, and is not equivalent to the actual physical distance. In addition, the above expression can accurately represent the ring characteristics of the ring lattice, thereby improving the accuracy of the distance calculation.
[0186] In some embodiments, determining the probability of generating a connection inside the node pair based on the distance between the first type of node and the second type of node in the node pair on the ring lattice and the number of connections to be generated between the nodes in the node pair includes: determining an initialization score of the node pair based on the distance and a third parameter, wherein the third parameter characterizes the degree of dependence of the initialization score on the distance; and determining the probability of generating a connection inside the node pair based on the initialization score and the number of connections to be generated.
[0187] In the above embodiment, the initialization score of each node pair may be determined first according to the distance between the nodes and the third parameter, wherein the expression of the initialization score may be, for example:
[0188]
[0189] Among them, S ij represents the initialization score of node pair (i, j), d ij represents the distance between the first type node i and the second type node j on the ring lattice, is the expression of the third parameter, which represents the degree of dependence of the initialization score on the distance. It can be seen that the larger r is, the less dependent the initialization score is on the distance, which will lead to more random connection generation.
[0190] After determining the initialization score, the method may include determining a probability of generating a connection within each node pair based on a multinomial distribution according to the initialization score and the number of connections to be generated for each node.
[0191] In step S64, connections can be randomly generated within a portion of the node pairs based on the above probability, so that the sparsity of the initial sparse neural network meets the requirements, thereby completing the initialization of the binary receptive field.
[0192] Figure 6B (b) in the figure shows a schematic diagram of the generated initial sparse neural network. It can be seen that the above initialization method can significantly increase the probability of generating connections between nodes with corresponding numbers in the initial sparse neural network, that is, nodes with close distances, thereby optimizing the receptive field characteristics of the sparse neural network.
[0193] Figure 6C A schematic diagram illustrating the connection state in an initialized sparse neural network according to some embodiments of the present disclosure. Figure 6D A schematic diagram illustrating the connection state in an initialized sparse neural network according to some other embodiments of the present disclosure. Figure 6C For the case where the number of connections to be generated for each node is the same, Figure 6D This corresponds to the case where the number of connections to be generated for each node is determined based on a uniform distribution.
[0194] Figure 6C and Figure 6D The adjacency matrix of the sparse neural network under different conditions of the third parameter is shown, where black represents the absence of a connection between nodes, for example, represented as 0 in the adjacency matrix, and white represents the presence of a connection between nodes, for example, represented as 0 in the adjacency matrix. Specifically, the rows of each element in the matrix can correspond to the number of the first type of node, and the columns can correspond to the number of the second type of node. For example, the value of element (1,1) is 1, which represents that a connection is generated between the first type node A1 and the second type node B1.
[0195] As can be seen from the figure, the diagonal part of the adjacency matrix of the initial sparse neural network generated based on the above method will have relatively concentrated values, indicating that there are usually connections between the nodes corresponding to the numbers, such as elements (1,1) and (2,2).
[0196] In addition, when the third parameter changes with the value of r, the smaller the value of r, the more dependent the connection generation in the sparse neural network is on distance, and the connections in the neural network are more concentrated in the diagonal part of the adjacency matrix. When r is equal to 0, the third parameter tends to infinity, and the connection generation will be completely determined by distance. Therefore, all connections in the neural network are concentrated in the diagonal part of the adjacency matrix. The larger the value of r, the less dependent the connection generation in the sparse neural network is on distance, and the connections in the neural network are reduced and concentrated in the diagonal part of the adjacency matrix. When r is equal to 1, the third parameter is equal to 0, and the connection generation will be completely random, so all connections in the neural network are randomly distributed.
[0197] pass Figure 6C and Figure 6D From the comparison, we can find that when the number of connections to be generated for each node is fixed, it is equivalent to each row and column in the adjacency matrix having the same number of "1" elements, so the white diagonal edges in the matrix are relatively smooth. When the number of connections to be generated for each node is determined based on uniform distribution, although the probability of each node generating a connection is the same, the actual number of connections generated may not be exactly the same, and there is a certain degree of randomness, so the white diagonal edges in the matrix are relatively rough.
[0198] Both of the above methods can optimize the receptive field characteristics of sparse neural networks through the connections between corresponding nodes, thereby improving the subsequent training effect of sparse neural networks.
[0199] Combined with the above 6A to 6D This paper introduces how to initialize the binary receptive field. Figures 7A to 8 This section describes how to perform bisection scale-free initialization.
[0200] Figure 7A A schematic diagram showing the network structure of a scale-free network, where black represents no connection between nodes and white represents connections between nodes. Figure 7A As shown in Figure 1, scale-free property is another network structural property, which is manifested in that the number of connections between each node in a sparse neural network, that is, the degree of the node, follows a power-law distribution. The number of nodes with a certain degree in the network is inversely proportional to a certain power of the degree. This power can be called the power-law exponent. The power-law exponent in the power-law distribution is usually between 2 and 3. Figure 7AIn the example shown, the power law exponent is 2.76. In this network, the vast majority of nodes have very few connections, while a small number of nodes have a very high number of connections.
[0201] Next, we will specifically introduce how to obtain the initialized bipartite scale-free network in this disclosure. Figure 7B A flowchart of initializing a sparse neural network according to some further embodiments of the present disclosure is shown. Figure 8 A schematic diagram illustrating initialization of a sparse neural network according to further embodiments of the present disclosure is shown.
[0202] like Figure 7B As shown, obtaining an initial sparse neural network may include: step S71, obtaining a unipartite scale-free network; step S72, randomly selecting a preset number of nodes in the unipartite scale-free network as nodes of the first type, and the remaining nodes as nodes of the second type, to obtain an intermediate network; step S73, removing conflicting connections in the intermediate network, wherein the conflicting connections are connections between nodes of the same type; step S74, generating connections between nodes of the first type and nodes of the second type according to the number of conflicting connections, to obtain a bipartite scale-free network as the initial sparse neural network, wherein the number of nodes with the same number of connections in the bipartite scale-free network is the same as that in the intermediate network.
[0203] In step S71 , a monopartite scale-free network may be constructed as a basis. The monopartite scale-free network refers to a scale-free network in which all nodes are of the same type and the distribution of the number of connections of the nodes conforms to a power-law distribution.
[0204] In some embodiments, constructing a single-part scale-free network may include, for example: generating a fully connected small network with 3 nodes; adding a node each time, the newly added node generates connections with a preset number of existing nodes, for example, 2, according to a priority connection formula; repeating the above steps of adding new nodes until the scale of the sparse neural network reaches a preset value.
[0205] The above-mentioned priority connection formula means that the probability of a new node establishing a connection with an existing node is proportional to the number of connections to the existing node. For example, the expression is:
[0206]
[0207] Among them, w i k is the probability of a new node generating a connection with the i-th existing node, i is the number of connections to the i-th existing node, C e is the set of all existing nodes, j is any existing node, It is the sum of the number of connections of all existing nodes.
[0208] Through the above formula, the number of connections of nodes with high connection numbers in sparse neural networks can be made higher and higher, thus conforming to the characteristics of scale-free networks. Figure 8 (a) shows an example of a portion of the connections in the constructed single-component scale-free sparse neural network.
[0209] In step S72, node types may be distinguished in the monotonic scale-free network, and a portion of nodes may be selected as nodes of the first type, and the remaining nodes may be selected as nodes of the second type.
[0210] Figure 8 (b) shows an example of performing node differentiation, such as Figure 8 As shown in (b), 4 of the 8 nodes in the sparse neural network can be randomly selected as the first type of nodes, and the remaining 4 can be selected as the second type of nodes, where the square represents the first type of nodes and the circle represents the second type of nodes.
[0211] In step S73, since connections cannot be generated between nodes of the same type in the bipartite sparse neural network, it is necessary to remove these conflicting connections in the intermediate network obtained in step S72.
[0212] Figure 8 (c) shows an example of removing conflicting connections, such as Figure 8 As shown in (c), the conflicting connections in the intermediate network can be removed so that the network meets the requirements of a dual-partition network.
[0213] In step S74, since removing conflicting connections will cause the degrees of nodes in the network to change, it is necessary to regenerate connections between some nodes to restore the scale-free property of the network.
[0214] Specifically, the conflicting connections include first conflicting connections and second conflicting connections, where the first conflicting connections are connections between nodes of the first type, and the second conflicting connections are connections between nodes of the second type. Two situations can be divided according to the number of first conflicting connections and second conflicting connections removed.
[0215] When the number of the first conflicting connections is the same as the number of the second conflicting connections, connections can be randomly generated between the first type of nodes corresponding to the first conflicting connections and the second type of nodes corresponding to the second conflicting connections, wherein the number of connections generated by each node is the same as the number of conflicting connections removed by each node.
[0216] For example, each first conflicting connection can be matched with a second conflicting connection to form a conflicting connection pair. The conflicting connection pair will correspond to 4 nodes, including 2 nodes of the first type and 2 nodes of the second type. By regenerating connections between nodes of different types between these nodes, the scale-free characteristics of the network can be restored, and a bipartite scale-free network can be obtained as the initial sparse neural network. Figure 8 (d) shows an example of regenerating connections in a case where the number of first conflicting connections is the same as the number of second conflicting connections.
[0217] When the number of the first conflicting connections is different from the number of the second conflicting connections, for example, when the number of the first conflicting connections is 5 and the number of the second conflicting connections is 3.
[0218] First, three nodes corresponding to the first conflicting connection may be randomly selected from the nodes corresponding to the first conflicting connection, and connections may be regenerated between these nodes and the three nodes corresponding to the second conflicting connections using the above-mentioned processing method with the same number of connections.
[0219] Next, the nodes corresponding to the remaining two first conflicting connections can be considered as newly added nodes, and are connected to the second type of nodes in the network according to the above-mentioned steps of adding new nodes, that is, using the priority connection formula.
[0220] Through the above steps, the scale-free characteristics of the network can also be restored to obtain a bipartite scale-free network as the initial sparse neural network.
[0221] The above describes two methods for initializing a sparse neural network in the embodiments of the present disclosure. By pre-equipping the initial sparse neural network with certain topological structural features, the subsequent training effect based on the initial sparse neural network can be improved.
[0222] Next, we will combine Figure 9 This paper introduces the training method of sparse neural network in detail. Figure 9 Schematic diagram of a training method according to an embodiment of the present disclosure is shown. Figure 9 As shown, the training method may include steps (a) to (f), and each step will be described below.
[0223] In step (a), the initial sparse neural network can be obtained by initializing the small-world network through bisection, for example. Figure 9 As shown in (a), a portion of the initial sparse neural network may include, for example, nodes 11-14, 21-24, and 31-34, wherein nodes 11-14 belong to the input layer nodes, nodes 21-24 belong to the middle layer nodes, and nodes 31-34 belong to the output layer nodes.
[0224] Figure 9The example sparse neural network shown is a feedforward sparse neural network, in which only nodes in adjacent layers have connections. Therefore, in the subsequent process of connection removal and connection regeneration, the connection between the input layer and the intermediate layer, and the connection between the intermediate layer and the output layer can be considered separately to avoid confusion in the training process between different layers.
[0225] In step (b), before performing the density reduction step, the network may be preliminarily trained, for example, the connections and / or connection weights in the sparse neural network may be updated according to a preset number of iterations. Figure 9 In the example shown, step (b) is performed for 50 rounds, for example.
[0226] For example, the connections in the sparse neural network may be updated according to the weights of the connections between the nodes in the sparse neural network and the number of connections between at least one common neighbor node of a pair of nodes that do not have a mutual connection.
[0227] In addition, by training the weights of the connections in the sparse neural network, for example, increasing the weights of correct connections and reducing the weights of incorrect connections, the connection predictions in the sparse neural network can be made more accurate and the performance of the sparse neural network can be improved.
[0228] In step (c), the sparse neural network can be trained for multiple rounds including density reduction, e.g. Figure 9 In the example shown, step (c) is performed for 900 rounds, with each round of training including steps (c1) to (c3). Steps (c1) and (c3) are similar to the processing performed in step (b). In step (c1), the weights of the connections in the sparse neural network can be trained, and in step (c3), the connections in the sparse neural network can be updated. The following focuses on step (c2).
[0229] In step (c2), the target sparsity in the current round of training can be determined according to the round of training. Figure 9 In the example round shown in (c), the target sparsity is 9 / 16 and the sparsity change is 1 / 16. Subsequently, connections of the sparse neural network can be removed based on the determined sparsity change.
[0230] In the multiple rounds of training repeated in step (c), the first stage can be, for example, the first 450 rounds, and the second stage can be, for example, the 450 rounds. The change in sparsity is positively correlated with the number of training rounds in the first stage and negatively correlated with the number of training rounds in the second stage. After step (c) is completed, the sparsity, i.e., density, of the sparse neural network that meets the requirements can be obtained. Figure 9 In the example of , the sparsity requirement is, for example, 3 / 4.
[0231] In step (d), after completing the density reduction step, the sparse neural network can continue to be trained, for example, by updating the connections and / or connection weights in the sparse neural network until the sparse neural network converges, that is, the loss function of the sparse neural network tends to be stable, indicating that the sparse neural network has been optimized.
[0232] Combined with the above Figure 9 The training method according to the embodiment of the present disclosure is introduced. Next, the application scenario of the sparse neural network according to the embodiment of the present disclosure will be introduced.
[0233] In some embodiments, a sparse neural network can be used for image processing, wherein the input of the sparse neural network is a first image, and the output is at least one of the category of the first image, the object in the first image, the second image generated based on the first image, and the text generated based on the first image.
[0234] In other words, the sparse neural network trained in this application can be used for tasks such as image classification, object recognition in images, image enhancement, and text generation from images. Generating text from images, for example, is describing an image in text or writing based on an image. The input of the sparse neural network is an image, and the output is the corresponding text.
[0235] In other embodiments, a sparse neural network can be used for text processing, where the input of the sparse neural network is a first text, and the output is at least one of the following: a category of the first text, an element in the first text, a second text generated based on the first text, and an image generated based on the first text. Generating text from an image, for example, involves describing an image in text or writing based on an image, where the input of the sparse neural network is an image, and the output is the corresponding text.
[0236] In other words, the sparse neural network trained in this application can be used for tasks such as text classification, element recognition in text, machine translation, and generating images from text. For generating images from text, such as intelligent painting, the input of the sparse neural network is text, and the output is the corresponding image.
[0237] In addition, the training method proposed in this disclosure can be used to train sparse neural networks such as multilayer perceptrons (MLPs), self-attention models (Transformers), and large language models (LLMs), thereby reducing training complexity and shortening training time.
[0238] The above application scenarios are merely exemplary and non-restrictive. The above training method proposed in the present disclosure can improve the training effect of sparse neural networks in various scenarios.
[0239] The above is the training method provided by the embodiment of the present disclosure. The above training method of the present disclosure can reduce the probability of incorrectly removing key connections during sparse neural network training by reducing the number of connections removed in the early stage of training, thereby improving the training effect.
[0240] Reference below Figure 10 and Figure 11 A training device according to an embodiment of the present disclosure is described, which is used to perform any embodiment of the above-mentioned training method. Figure 10 A block diagram of a training device according to some embodiments of the present disclosure is shown.
[0241] like Figure 10 As shown, the training device 10 of the sparse neural network includes: a first determination module 101, configured to determine the target sparsity of the sparse neural network in each round of training of the sparse neural network according to the training round of the sparse neural network; a second determination module 102, configured to determine the sparsity change of the sparse neural network in the round of training according to the target sparsity of the sparse neural network in the round of training and the sparsity of the sparse neural network after the previous round of training; a removal module 103, configured to remove a part of the connections in the sparse neural network according to the sparsity change, wherein the multiple rounds of training include a first stage and a second stage, the first stage is before the second stage, the sparsity change is positively correlated with the training round in the first stage, and negatively correlated with the training round in the second stage.
[0242] The first determination module 101 of the training device 10 can be used to perform, for example Figure 1 The second determining module 102 of the training device 10 can be used to perform, for example, Figure 1 The removal module 103 of the training device 10 can be used to perform, for example, Figure 1 Step S13.
[0243] Figure 11 A block diagram of a training device according to some other embodiments of the present disclosure is shown.
[0244] like Figure 11 As shown, the training device 11 includes: at least one memory 111; and at least one processor 112 coupled to the at least one memory 111, and the at least one processor 112 is configured to execute the training method described in any of the foregoing embodiments based on instructions stored in the at least one memory 111.
[0245] Memory 111 is used to store one or more computer-readable instructions. Memory 111 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 111 may store, for example, an operating system, application programs, a boot loader, a database, and other programs, as well as various application programs and data.
[0246] The processor 112 is used to execute computer-readable instructions to implement the training method described in any of the above embodiments. The specific implementation of each step of the method can refer to the above embodiments, for example Figure 1 、 Figure 3 、 Figure 6A and Figure 7B The steps in , which are repeated, are not repeated here.
[0247] The processor 112 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) may be of X86 or ARM architecture, etc.
[0248] The processor 112 and the memory 111 can communicate with each other directly or indirectly. For example, the processor 112 and the memory 111 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 112 and the memory 111 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0249] It should be noted that Figure 11 The components of training device 11 shown are merely exemplary and non-limiting. Training device 11 may also include other components depending on actual application needs. Processor 112 may control other components in training device 11 to perform desired functions. Training device 11 may be implemented using software, firmware, and / or hardware, and may be integrated into a device equipped with a relevant application.
[0250] Figure 12 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0251] Figure 12The electronic device 12 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.
[0252] Electronic devices include but are not limited to mobile terminals such as smart phones, laptops, personal digital assistants (PDAs), tablet computers (Tablet Personal Computers, Tablet PCs), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital televisions, desktop computers, etc.
[0253] like Figure 12 As shown, a central processing unit (CPU) 121 executes various processes according to a program stored in a read-only memory (ROM) 122 or a program loaded from a storage unit 128 to a random access memory (RAM) 123. RAM 123 stores data required for CPU 121 to execute various processes, etc., as needed. The central processing unit is merely exemplary and may also be another type of processor, such as the various processors described above. ROM 122, RAM 123, and storage unit 128 may be various forms of computer-readable storage media.
[0254] It should be noted that although Figure 12 1 and 2. ROM 122, RAM 123, and storage 128 are shown separately in the figure, but one or more of them may be combined or located in the same or different memory or storage modules. CPU 121, ROM 122, and RAM 123 are connected to each other via bus 124. Input / output interface 125 is also connected to bus 124.
[0255] The following components are connected to the input / output interface 125: an input portion 126 such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output portion 127 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage portion 128 including a hard disk, a magnetic tape, etc.; and a communication portion 129 including a network interface card such as a LAN card, a modem, etc. The communication portion 129 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 12 Some of the electronic devices 12 are shown to communicate via the bus 124, but they can also communicate via a network or other means, where the network can include a wireless network, a wired network, and / or any combination of wireless networks and wired networks.
[0256] A drive 1210 is also connected to the input / output interface 125 as needed. A removable medium 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1210 as needed so that a computer program read therefrom is installed into the storage section 128 as needed.
[0257] When the above-described series of processing is implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1211 .
[0258] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when the computer program product is run on a computer, causes the computer to implement the training method described in any of the aforementioned embodiments. The computer program product includes computer instructions carried on a computer-readable storage medium, containing program code for executing the method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from the network through the communication part 129, or installed from the storage part 128, or installed from the ROM 122. When the computer program is executed by the CPU 121, the training method of the embodiment of the present disclosure is executed.
[0259] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. Computer instructions are stored on the computer-readable storage medium, and when the instructions are executed by the processor, the training method described in any of the aforementioned embodiments is implemented. The above-mentioned computer-readable storage medium may be included in the above-mentioned electronic device; or it may exist separately without being assembled into the electronic device.
[0260] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the method described in any of the aforementioned embodiments. For example, the instructions may be embodied as computer program codes.
[0261] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0262] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to various embodiments of the present disclosure. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two blocks shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved.
[0263] The above functions may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0264] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A sparse neural network training method, comprising: In each round of the multiple rounds of training of the sparse neural network, determining a target sparsity of the sparse neural network in the round of training according to the training round of the sparse neural network; Determining a change in the sparsity of the sparse neural network in this round of training based on the target sparsity of the sparse neural network in this round of training and the sparsity of the sparse neural network after the previous round of training; According to the change in sparsity, a portion of the connections in the sparse neural network are removed. The multiple rounds of training include a first stage and a second stage following the first stage, and the sparsity change is positively correlated with the training rounds in the first stage and negatively correlated with the training rounds in the second stage.
2. The training method according to claim 1, wherein: The first stage is from the beginning to the middle training round of the multiple rounds of training, and the second stage is from the middle training round to the end. In each round of the multiple rounds of training of the sparse neural network, determining the target sparsity of the sparse neural network in the round of training according to the training round of the sparse neural network includes: The target sparsity of the sparse neural network in the training round is determined according to the difference between the training round and the intermediate training round.
3. The training method according to claim 2, wherein: Determining, based on the difference between the training round and the intermediate training round, a target sparsity of the sparse neural network in the training round includes: A target sparsity of the sparse neural network in this round of training is determined based on the difference and a first parameter, wherein the first parameter represents a degree of dependence of the target sparsity on the difference.
4. The training method according to claim 1, wherein: Removing a portion of the connections in the sparse neural network according to the sparsity change includes: Removing a portion of the connections in the sparse neural network according to the sparsity change and the relative importance of the weights of the respective connections in the sparse neural network, wherein the relative importance is a ratio of the weight of each connection to the total weight of all connections that have a common node with the connection.
5. The training method according to claim 4, wherein: Removing a portion of the connections in the sparse neural network according to the sparsity change and the relative importance of the weights of the connections in the sparse neural network includes: Determining the number of connections that need to be removed in the sparse neural network according to the change in sparsity; Determining a probability of removing each connection based on a multinomial distribution according to the relative importance of the weights of the connections in the sparse neural network; A portion of the connections in the sparse neural network is removed according to the number of connections that need to be removed and the probability of removing each connection.
6. The training method according to claim 5, wherein: According to the relative importance of the weights of the connections in the sparse neural network, determining the probability of removing each connection based on a multinomial distribution includes: Determining a second parameter based on the number of training rounds of the sparse neural network, wherein the second parameter represents a degree of dependence of the probability of removing the connection on the relative importance of the connection, and the second parameter is positively correlated with the number of training rounds of the sparse neural network; A probability of removing each connection is determined according to the relative importance, the second parameter, and a multinomial distribution.
7. The training method according to claim 1, further comprising: Obtain an initial sparse neural network; A sparse neural network for the multiple rounds of training is determined based on the initial sparse neural network.
8. The training method according to claim 7, wherein: Obtaining an initial sparse neural network involves: constructing a ring lattice, the ring lattice comprising nodes of a first type and nodes of a second type, wherein each node of the first type forms a node pair with a node of the second type; generating initial connections between nodes of the first type and adjacent nodes of the second type, wherein the number of connections generated by all nodes is the same; According to a specified ratio, a portion of the initial connections is randomly removed, and new connections are randomly generated within a portion of node pairs that are not connected to each other, to obtain a bipartite small-world network as the initial sparse neural network.
9. The training method according to claim 7, wherein: Obtaining an initial sparse neural network involves: constructing a ring lattice, the ring lattice comprising nodes of a first type and nodes of a second type, wherein each node of the first type forms a node pair with a node of the second type; Determine the number of connections to be generated for each node; For each node pair in the ring lattice, determining a probability of generating a connection within the node pair based on a distance between a first type of node and a second type of node in the node pair on the ring lattice and a number of connections to be generated between the nodes in the node pair; According to the probability of generating connections, connections are generated inside some node pairs to obtain a binary neural network as the initial sparse neural network.
10. The training method according to claim 9, wherein: Determining the probability of generating a connection within the node pair according to a distance between a first type of node and a second type of node in the node pair on the ring lattice and a number of connections to be generated between nodes in the node pair includes: determining an initialization score for the node pair based on the distance and a third parameter, wherein the third parameter represents a degree of dependence of the initialization score on the distance; The probability of generating a connection within the node pair is determined according to the initialization score and the number of the connections to be generated.
11. The training method according to claim 9, wherein: Determining the number of connections to be made per node involves: Determining the number of connections to be generated for each node based on the sparsity of a preset initial sparse neural network, wherein the number of connections generated by all nodes is the same; or According to the sparsity of the preset initial sparse neural network, the total number of connections to be generated for all nodes is determined, and according to the total number, the number of connections to be generated for each node is determined based on uniform distribution.
12. The training method according to claim 7, wherein: Obtaining an initial sparse neural network involves: Obtain a single-part scale-free network; Randomly selecting a preset number of nodes in the unipartite scale-free network as nodes of the first type and the remaining nodes as nodes of the second type to obtain an intermediate network; Removing conflicting connections in the intermediate network, wherein the conflicting connections are connections between nodes of the same type; According to the number of conflicting connections, connections are generated between nodes of the first type and nodes of the second type to obtain a bipartite scale-free network as the initial sparse neural network, wherein the number of nodes with the same number of connections in the bipartite scale-free network and the intermediate network is the same.
13. The training method according to claim 12, wherein: The conflicting connections include first conflicting connections and second conflicting connections, the first conflicting connections are connections between nodes of the first type, and the second conflicting connections are connections between nodes of the second type. Generating a connection between the nodes of the first type and the nodes of the second type according to the number of the conflicting connections includes: When the number of the first conflicting connections is the same as the number of the second conflicting connections, connections are randomly generated between the first type of nodes corresponding to the first conflicting connections and the second type of nodes corresponding to the second conflicting connections, wherein the number of connections generated by each node is the same as the number of conflicting connections removed by each node.
14. The training method according to claim 1, further comprising: The connections in the sparse neural network are updated according to the weights of the connections between the nodes in the sparse neural network and the number of connections of at least one common neighbor node of the node pairs that do not have mutual connections.
15. The training method according to claim 1, further comprising: Before removing a portion of the connections in the sparse neural network in each round of training, the weights of each connection in the sparse neural network are trained.
16. The training method according to claim 1, wherein: The sparse neural network input is a first image, and the output is at least one of a category of the first image, an object in the first image, a second image generated based on the first image, and text generated based on the first image; or The input of the sparse neural network is a first text, and the output is at least one of the category of the first text, an element in the first text, a second text generated based on the first text, and an image generated based on the first text.
17. A sparse neural network training device, comprising: A first determination module is configured to determine, in each round of multiple rounds of training of the sparse neural network, a target sparsity of the sparse neural network in the training round according to the training round of the sparse neural network; A second determining module is configured to determine a sparsity change of the sparse neural network in this round of training based on the target sparsity of the sparse neural network in this round of training and the sparsity of the sparse neural network after the previous round of training; A removal module is configured to remove a portion of the connections in the sparse neural network according to the sparsity change. The multiple rounds of training include a first stage and a second stage, the first stage is before the second stage, the sparsity change is positively correlated with the training rounds in the first stage, and negatively correlated with the training rounds in the second stage.
18. A sparse neural network training device, comprising: at least one memory; as well as At least one processor coupled to the at least one memory is configured to execute the training method according to any one of claims 1 to 16 based on instructions stored in the at least one memory.
19. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the training method according to any one of claims 1 to 16 is implemented.
20. A computer program product, which, when running on a computer, enables the computer to implement the training method according to any one of claims 1 to 16.