Parameter determination device, signal transmission device, parameter determination method, signal transmission method, and computer program

By introducing a new layer with fewer nodes and optimizing connections in neural networks, the computational burden is reduced, allowing for efficient signal transmission with minimal calculations.

JP7700569B2Active Publication Date: 2025-07-01NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021133235
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-17
Filing Date
2021-08-18
Publication Date
2025-07-01
Estimated Expiration
2041-08-18

AI Technical Summary

Technical Problem

The complexity of neural network structures leads to a high computational burden, necessitating a solution to construct neural networks with a reduced computational load.

Method used

A parameter determination device and method that introduces a new layer with fewer nodes lacking non-linear activation functions between existing layers, learns weights between these layers, and selects effective connection paths based on learned weights to reduce computational requirements.

Benefits of technology

This approach constructs a neural network with a significantly reduced computational load, enabling efficient signal transmission while maintaining performance by minimizing unnecessary calculations and connections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700569000006
    Figure 0007700569000006
  • Figure 0007700569000007
    Figure 0007700569000007
  • Figure 0007700569000008
    Figure 0007700569000008
Patent Text Reader

Abstract

To provide a parameter determination device with which it is possible to construct a neural network, of which a required amount of computation is relatively small.SOLUTION: A parameter determination device 2 comprises: addition means 211 for adding, as a new layer constituting a portion of a neural network, a third layer 1124 where output H2 of a first layer is inputted between a first layer 11222 and a second layer 11223 which a neural network 112 is provided with, and output H5 is inputted to the second layer, the third layer including, by the number P fewer than the number N of a second node N3 which the second layer is provided with, a third node N5 to which inputted are outputs of a plurality of first nodes N2 not including a nonlinear activation function and which the first layer is provided with; learning means 212 for learning a weight w5 between the third and second layers; and selection means 213 for selecting, on the basis of the weight, one effective path for each of the second nodes, from among a plurality of connection paths that connect the third and second nodes.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of a parameter determination device, a parameter determination method, and a computer program for determining parameters of a neural network, as well as a signal transmission device, a signal transmission method, and a computer program for transmitting signals.

Background Art

[0002] In recent years, the utilization of neural networks has been studied in various technical fields. For example, in a wireless communication system such as a mobile communication system, a distortion compensation circuit using a digital pre-distortion (DPD) method is constructed using a neural network.

[0003] In addition, as prior art documents related to the present invention, Patent Document 1 to Patent Document 4 can be cited.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0005] An apparatus constructed using a neural network has a technical problem that the amount of computation (i.e., the amount of calculation) becomes relatively large due to the complexity of the network structure of the neural network. For this reason, it is desirable to construct a neural network with a relatively small amount of necessary computation.

[0006] An object of the present invention is to provide a parameter determination device, a parameter determination method, a computer program, and a recording medium that can solve the above-described technical problem. As an example, the present invention aims to provide a parameter determination device, a parameter determination method, and a computer program capable of constructing a neural network with a relatively small amount of necessary computation, and a signal transmission device, a signal transmission method, and a computer program for transmitting a signal using a neural network with a relatively small amount of necessary computation.

Means for Solving the Problem

[0007] One aspect of the parameter determination device is a parameter determination device that determines parameters of a neural network, and between a first layer and a second layer included in the neural network, (i) a third layer in which the output of the first layer is input and the output is input to the second layer, and (ii) a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and includes a third layer with a number smaller than the number of a plurality of second nodes included in the second layer, as an additional means for adding a new layer constituting a part of the neural network; a learning means for learning the weights between the third layer and the second layer as part of the parameters; and based on the weights learned by the learning means, selection means for selecting, for each of the plurality of second nodes, one effective path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes, as part of the parameters.

[0008] A first aspect of the signal transmission device includes distortion compensation means for generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network, and signal generation means for generating a transmission signal to be transmitted to a signal reception device by performing a predetermined operation on the distortion compensation signal. The neural network includes a first layer that is an input layer or an intermediate layer, a second layer that is an intermediate layer or an output layer, and a third layer to which the output of the first layer is input and whose output is input to the second layer. The third layer includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of third nodes included in the third layer is less than the number of a plurality of second nodes included in the second layer. The output of a single third node is input to each of the plurality of second nodes.

[0009] A second aspect of the signal transmission device includes distortion compensation means for generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network, signal generation means for generating a transmission signal to be transmitted to a signal reception device by performing a predetermined operation on the distortion compensation signal, an additional means for adding, as a new layer constituting a part of the neural network, a third layer (i) to which the output of the first layer is input and whose output is input to the second layer, and (ii) that includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of third nodes included in the third layer is less than the number of a plurality of second nodes included in the second layer, between the first layer and the second layer included in the neural network, a learning means for learning, as a part of the parameters, the weights between the third layer and the second layer, and a selection means for selecting, as an effective path used as an effective connection path in the neural network, one effective path for each second node as a part of the parameters, from among a plurality of connection paths connecting the third node and the plurality of second nodes, based on the weights learned by the learning means.

[0010] One aspect of the parameter determination method is a parameter determination method for determining the parameters of a neural network, wherein between a first layer and a second layer included in the neural network, (i) a third layer is provided, where the output of the first layer is input and the output is input to the second layer, and (ii) the third layer includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is less than the number of a plurality of second nodes included in the second layer, and the third layer is added as a new layer constituting a part of the neural network; weights between the third layer and the second layer are learned as a part of the parameters; and based on the learned weights, an effective path used as an effective connection path in the neural network is selected one by one for each of the plurality of second nodes as a part of the parameters from among a plurality of connection paths connecting the third node and the plurality of second nodes.

[0011] The first aspect of the signal transmission method includes generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network, and generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal. The neural network includes a first layer that is an input layer or an intermediate layer, a second layer that is an intermediate layer or an output layer, and a third layer where the output of the first layer is input and the output is input to the second layer, and the third layer includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is less than the number of a plurality of second nodes included in the second layer, and the output of a single third node is input to each of the plurality of second nodes.

[0012] A second aspect of the signal transmission method is to generate a distortion compensation signal by performing distortion compensation on an input signal using a neural network, generate a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal, and between a first layer and a second layer included in the neural network, (i) a third layer to which an output of the first layer is input and an output of which is input to the second layer, (ii) a third node that does not include a non-linear activation function and to which outputs of a plurality of first nodes included in the first layer are input, and that includes a number smaller than the number of a plurality of second nodes included in the second layer, add the third layer as a new layer constituting a part of the neural network, learn weights between the third layer and the second layer as a part of the parameters, and based on the learned weights, select, for each of the plurality of second nodes, as an effective path used as an effective connection path in the neural network, one effective path from among a plurality of connection paths connecting the third node and the plurality of second nodes as a part of the parameters.

[0013] A first aspect of a computer program is a computer program that causes a computer to execute a parameter determination method for determining parameters of a neural network, the parameter determination method including adding, between a first layer and a second layer included in the neural network, (i) a third layer to which an output of the first layer is input and an output of which is input to the second layer, (ii) a third node that does not include a non-linear activation function and to which outputs of a plurality of first nodes included in the first layer are input, and that includes a number smaller than the number of a plurality of second nodes included in the second layer, as a new layer constituting a part of the neural network, learning weights between the third layer and the second layer as a part of the parameters, and based on the learned weights, selecting, for each of the plurality of second nodes, as an effective path used as an effective connection path in the neural network, one effective path from among a plurality of connection paths connecting the third node and the plurality of second nodes as a part of the parameters.

[0014] A second aspect of the computer program is a computer program that causes a computer to execute a signal generation method, the signal generation method including generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network, and generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal, the neural network including a first layer that is an input layer or an intermediate layer, a second layer that is an intermediate layer or an output layer, and a third layer to which the output of the first layer is input and the output of which is input to the second layer, the third layer including a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, the third layer including a number of third nodes that is smaller than the number of a plurality of second nodes included in the second layer, and the output of a single third node being input to each of the plurality of second nodes.

[0015] A third aspect of the computer program is a computer program that causes a computer to execute a signal generation method, the signal generation method including generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network, generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal, adding, as a new layer constituting a part of the neural network, a third layer between a first layer and a second layer included in the neural network, the third layer (i) being a third layer to which the output of the first layer is input and the output of which is input to the second layer, and (ii) including a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, the third layer including a number of third nodes that is smaller than the number of a plurality of second nodes included in the second layer, learning, as a part of the parameters, the weights between the third layer and the second layer, and selecting, as a part of the parameters, for each of the plurality of second nodes, one effective path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes based on the learned weights.

Advantages of the Invention

[0016] According to each one aspect of the above-described parameter determination device, parameter determination method, and computer program, a neural network with a relatively small amount of necessary calculations is appropriately constructed. Further, according to each one aspect of the above-described signal transmission device, signal transmission method, and computer program, a signal is transmitted using a neural network with a relatively small amount of necessary calculations.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Modes for Carrying Out the Invention

[0018] Hereinafter, embodiments of a parameter determination device, a signal transmission device, a parameter determination method, a signal transmission method, and a computer program will be described with reference to the drawings.

[0019] (1) Signal transmission device 1 First, the signal transmission device 1 of the present embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the signal transmission device 1 of the present embodiment.

[0020] As shown in FIG. 1, the signal transmission device 1 transmits a transmission signal z to a signal reception device (not shown) via a communication line. The communication line is typically a wireless communication line, but at least a part thereof may be a wired communication line. To transmit the transmission signal z, the signal transmission device 1 includes a distortion compensation circuit (DPD: Digital Pre-Distortion) 11 and a power amplifier (PA: Power Amplifier) 12.

[0021] The distortion compensation circuit 11 generates a distortion compensation signal y by performing distortion compensation on the input signal x. The distortion compensation circuit 11 performs distortion compensation on the input signal x to compensate (typically, reduce or cancel) the distortion generated in the transmission signal z due to the operation of the power amplifier 12, thereby generating the distortion compensation signal y. In the present embodiment, the distortion compensation circuit 11 may be, for example, a DPD-type distortion compensation circuit. In particular, the distortion compensation circuit 11 may generate a distortion compensation signal y by applying the inverse distortion characteristic of the power amplifier 12 to the input signal x. In this case, it is possible to achieve both low power consumption and low distortion of the signal transmission device 1. Specifically, both the improvement of the efficiency of the signal transmission device 1 and the ensuring of the linearity of the amplification characteristic of the signal transmission device 1 are achieved.

[0022] The power amplifier 12 performs a predetermined operation on the distortion compensation signal y output from the distortion compensation circuit 11. Specifically, the power amplifier 12 amplifies the distortion compensation signal y. The distortion compensation signal y amplified by the power amplifier 12 is transmitted as a transmission signal z to the signal receiving device via a communication line. Here, when the distortion compensation circuit 11 is a distortion compensation circuit of the DPD method as described above, since the distortion of the signal in the power amplifier 12 is canceled by the inverse distortion of the signal in the distortion compensation circuit 11, the power amplifier 12 outputs a transmission signal z that is linear with respect to the input signal x.

[0023] In particular, in this embodiment, the distortion compensation circuit 11 performs distortion compensation on the input signal x using the neural network 112 (see FIG. 2). Hereinafter, the configuration of such a distortion compensation circuit 11 will be described in more detail with reference to FIG. 2. FIG. 2 is a block diagram showing the configuration of the distortion compensation circuit 11.

[0024] As shown in FIG. 2, the distortion compensation circuit 11 includes a signal generation unit 111 and a neural network 112.

[0025] The signal generation unit 111 generates a plurality of signals (typically, a plurality of signals each given a different delay) that are input to the neural network 112 from the input signal x input to the distortion compensation circuit 11. Note that the input signal x t means, for example, a complex signal of the input signal x input to the distortion compensation circuit 11 at time t. t is, for example, a complex signal of the input signal x input to the distortion compensation circuit 11 at time t.

[0026] The signal generation unit 111 can generate a plurality of signals from the input signal x t in any method as long as it can generate a plurality of signals input to the neural network 112. In the example shown in FIG. 2, the signal generation unit 111 is based on the input signal x t and, based on the input signal x t-1 from the input signal x t-K / 2Generate it. Here, the variable K represents the total number of nodes (i.e., neurons) N1 included in the input layer 1121 of the neural network 112 to be described later, and is an integer of 1 or more. The symbol " / " represents division (the same applies hereinafter). The input signal x t Based on the input signal x t-1 To generate the input signal x t-K / 2 The signal generation unit 111 includes K / 2 delay units 1111 (specifically, from delay unit 11111 to 1111 K / 2 ). The delay unit 1111 h (Here, the variable h is an integer from 1 to K / 2) adds a delay to the input signal X t-h+1 To generate the input signal X t-h . Further, the signal generation unit 111 generates an input signal I t-g corresponding to the I-axis signal component of the input signal x t-g from the input signal x t-g and an input signal Q t-g corresponding to the Q-axis signal component of the input signal x t-g . The I-axis signal component of the input signal x t-g corresponds to the in-phase component of the waveform of the input signal x t-g . The Q-axis signal component of the input signal x t-g corresponds to the quadrature component of the waveform of the input signal x t-g . To generate the input signal I t-g and Q t-g from the input signal x t-g , the signal generation unit 111 includes K / 2 + 1 signal converters 1112 (specifically, from signal converter 11120 to 1112 K / 2 ). The signal converter 1112 g generates the input signal I t-g and Q t-g from the input signal x t-g . As a result, the input signals I t to I t-K / 2 and the input signals Q t to Q t-K / 2 are input to the neural network 112.

[0027] Here, the signal generation unit 111 also generates the input signal xt Based on the input signal x t-1 from the input signal x t-K generate the input signal x, and the generated input signal x t from the input signal x t-K The amplitude value of may be input to the neural network 112. Further, the signal generation unit 111 may input the amplitude value of the input signal x t from the input signal x t-K and the amplitude value of the input signal I t from the input signal I t-K and the input signal Q t from the input signal Q t-K and mix them and input them to the neural network 112. The signal generation unit 111 may input the amplitude value of the input signal x t from the input signal x t-K and the input signal I t from the input signal I t-K and the input signal Q t from the input signal Q t-L and an operation value (for example, a power value, etc.) using them to the neural network 112.

[0028] The neural network 112, based on the input signal I t from the input signal I t-K / 2 and the input signal Q t from the input signal Q t-K / 2 generates a distortion compensation signal y t (that is, the input signal x with distortion compensation applied t) is generated. The neural network 112 includes an input layer 1121, at least one intermediate layer (i.e., hidden layer) 1122, an output layer 1123, and at least one linear layer 1124. In the following description, for the sake of convenience of explanation, as shown in FIG. 2, the neural network 112 has two intermediate layers 1122 (specifically, intermediate layer 11222 and intermediate layer 11223), and two adjacent intermediate layers 1122 (i.e., in a situation where there is no linear layer 1124, the output of one intermediate layer 1122 is input to the other intermediate layer 1122, specifically, intermediate layers 11222 and 11223), and an example of one linear layer 1124 arranged between them will be described. However, the neural network 112 may include one or three or more intermediate layers 1122. The neural network 112 may include two or more linear layers 1124. The neural network 112 may include a linear layer 1124 arranged between the intermediate layer 1122 adjacent to the input layer 1121 (i.e., the intermediate layer 1122 to which the output of the input layer 1121 is input in a situation where there is no linear layer 1124, and in the example shown in FIG. 2, it is intermediate layer 11222) and the input layer 1121. The neural network 112 may include a linear layer 1124 arranged between the intermediate layer 1122 adjacent to the output layer 1123 (i.e., the intermediate layer 1122 to which the output is input to the output layer 1123 in a situation where there is no linear layer 1124, and in the example shown in FIG. 2, it is intermediate layer 11223) and the output layer 1123.

[0029] The input layer 1121 includes K nodes N1. Hereinafter, the K nodes N1 are respectively denoted as node N1#1 to node N1#K to distinguish them from each other. The variable K is typically an integer of 2 or more. The intermediate layer 11222 is a layer to which the output of the input layer 1121 is input. The intermediate layer 11222 includes M nodes N2. Hereinafter, the M nodes N2 are respectively denoted as node N2#1 to node N2#M to distinguish them from each other. The constant M is typically an integer of 2 or more. The intermediate layer 11223 is a layer to which the output of the linear layer 1124 is input. The intermediate layer 11223 includes N nodes N3. Hereinafter, the N nodes N3 are respectively denoted as node N3#1 to node N3#N to distinguish them from each other. The constant N is typically an integer of 2 or more. The output layer 1123 is a layer to which the output of the intermediate layer 11223 is input. The output layer 1123 includes O nodes N4. Hereinafter, the O nodes N4 are respectively denoted as node N4#1 to node N4#O to distinguish them from each other. The constant O is typically an integer of 2 or more, but may be 1. In the example shown in FIG. 2, the constant O is 2, and the output layer 1123 includes nodes N4#1 and N4#2. The linear layer 1124 is a layer to which the output of the intermediate layer 11222 is input. The linear layer 1124 includes P nodes N5. Hereinafter, the P nodes N5 are respectively denoted as node N5#1 to node N5#P to distinguish them from each other. The constant P is typically an integer of 2 or more. Also, the constant P is smaller than the constant N described above. That is, the number of nodes N5 included in the linear layer 1124 is smaller than the number of nodes N3 included in the intermediate layer 11223 to which the output of the linear layer 1124 is input.

[0030] To the nodes N1#1 to N1#K of the input layer 1121, input signals I t from input signal I t-K / 2 and input signal Q t from input signal Q t-K / 2 are input. In the example shown in FIG. 2, when k (where k is a variable indicating an integer satisfying 1 ≦ k ≦ K) is an odd number, the k-th node N1#k of the input layer 1121 receives the input signal I t-(k-1) / 2is input. When k is an even number, the k-th node N1#k of the input layer 1121 receives the input signal Q t-(k-2) / 2 is input. The output H1#k of the k-th node N1#k may be the same as the input of the k-th node N1#k. Alternatively, the output H1#k of the k-th node N1#k may be represented by Equation 1. In Equation 1, "real(x)" is a function that outputs the real part of the complex input signal x, and "imag(x)" is a function that outputs the imaginary part of the complex input signal x. The output H1#k of the k-th node N1#k of the input layer 1121 is input to each of the nodes N2#1 to N2#M via M connection paths that respectively connect the k-th node N1#k of the input layer 1121 and the nodes N2#1 to N2#M of the intermediate layer 11222. Note that the variable k in Equation 1 exceptionally represents an integer of 1 or more and K / 2 or less.

[0031]

Number

[0032] The output H2#m of the m-th node N2#m of the intermediate layer 11222 (where m is a variable representing an integer satisfying 1 ≤ m ≤ M) is represented by Equation 2. In Equation 2, "w2(k, m)" represents the weight in the connection path between the k-th node N1#k of the input layer 1121 and the m-th node N2#m of the intermediate layer 11222. "b2(m)" in Equation 2 represents the bias used (i.e., added) at the m-th node N2#m of the intermediate layer 11222. "f" in Equation 2 represents an activation function. As the activation function, for example, a non-linear activation function may be used. As the non-linear activation function, for example, a sigmoid function or a ReLu (Rectified Linear Unit) function may be used. The output H2#m of the m-th node N2#m of the intermediate layer 11222 is input to each of the nodes N5#1 to N5#P via P connection paths that respectively connect the m-th node N2#m of the intermediate layer 11222 and the nodes N5#1 to N5#P of the linear layer 1124.

[0033]

Number

[0034] The output H5#p of the p-th node N5#p of the linear layer 1124 (where p is a variable indicating an integer satisfying 1 ≤ p ≤ P) is represented by Equation 3. "w5(m, p)" in Equation 3 represents the weight in the connection path between the m-th node N2#m of the intermediate layer 11222 and the p-th node N5#p of the linear layer 1124. As shown in Equation 3, each node N5 of the linear layer 1124 is a node that does not include a non-linear activation function. That is, each node N5 of the linear layer 1124 is a node that outputs a linear sum of the outputs H2#1 to H2#M of the intermediate layer 11222. Further, each node N5 of the linear layer 1124 is a node that does not add a bias. The output H5#p of the p-th node N5#p of the linear layer 1124 is input to at least one of the nodes N3#1 to N3#N through at least one of the N connection paths that respectively connect the p-th node N5#p of the linear layer 1124 and the nodes N3#1 to N3#N of the intermediate layer 11223.

[0035]

Number

[0036] To each node N3 in the intermediate layer 11223, one of the outputs H5#1 to H5#P of the linear layer 1124 is input. Specifically, while one of the outputs H5#1 to H5#P is input to each node N3, the remaining inputs excluding any one input that is input to each node N3 among the outputs H5#1 to H5#P are not input. In this case, any one of the P connection paths that respectively connect the nodes N5#1 to N5#P of the linear layer 1124 and each node N3 of the intermediate layer 11223 is used as a connection path (hereinafter referred to as the "effective path") that effectively (in other words, actually) connects the linear layer 1124 and the intermediate layer 11223. In other words, any one of the P connection paths that respectively connect the nodes N5#1 to N5#P of the linear layer 1124 and each node N3 of the intermediate layer 11223 is used as an effective connection path (that is, an effective path) in the neural network 112. That is, each node N3 of the intermediate layer 11223 is connected to one of the nodes N5#1 to N5#P of the linear layer 1124 via a single effective path. In other words, to each node N3 of the intermediate layer 11223, one of the outputs H5#1 to H5#P of the linear layer 1124 is input via a single effective path. On the other hand, the remaining P - 1 connection paths among the P connection paths that respectively connect the nodes N5#1 to N5#P of the linear layer 1124 and each node N3 of the intermediate layer 11223, excluding the effective path, are not used as an effective connection path that actually connects the linear layer 1124 and the intermediate layer 11223. In other words, the remaining P - 1 connection paths among the P connection paths that respectively connect the nodes N5#1 to N5#P of the linear layer 1124 and each node N3 of the intermediate layer 11223, excluding the effective path, are not used as an effective connection path in the neural network 112.

[0037] When the output H5#p of the p-th node N5#p in the linear layer 1124 is input to the n-th node N3#n in the intermediate layer 11223 (that is, when the n-th node N3#n in the intermediate layer 11223 and the p-th node N5#p in the linear layer 1124 are connected via an effective path), the output H3#n of the n-th node N3#n in the intermediate layer 11223 is represented by Equation 4. "w3(p, n)" in Equation 4 indicates the weight in the connection path between the p-th node N5#p in the linear layer 1124 and the n-th node N3#n in the intermediate layer 11223. "b3(n)" in Equation 4 indicates the bias used (that is, added) at the n-th node N3#n in the intermediate layer 11223. The output H3#n of the n-th node N3#n in the intermediate layer 11223 is input to each of the nodes N4#1 to N4#O in the output layer 1123 via O connection paths that respectively connect the n-th node N3#n in the intermediate layer 11223 and the nodes N4#1 to N4#O in the output layer 1123.

[0038] [Number]

[0039] The output H4#o of the o-th node N4#o in the output layer 1123 (where o is a variable representing an integer satisfying 1 ≤ o ≤ O) is represented by Equation 5. "w4(n, o)" in Equation 5 indicates the weight in the connection path between the n-th node N3#n in the intermediate layer 11223 and the o-th node N4#o in the output layer 1123. "b4(o)" in Equation 5 indicates the bias used (that is, added) at the o-th node N4#o in the output layer 1123.

[0040] [Number]

[0041] The output of the output layer 1123 (for example, the linear sum of outputs H4#1 to H4#O) corresponds to the final output signal y. t The output signal y t corresponds to the input signal x at time t. tIt corresponds to the distortion compensation signal y generated from . Note that the output layer 1123 may not include the activation function f. In this case, the output of the output layer 1123 may be a linear sum based on the outputs of the nodes N3#1 to N3#N of the intermediate layer 11223.

[0042] Such characteristics (substantially, the structure) of the neural network 112 are determined by parameters such as the above-described weights w, the above-described biases b, and the connection mode CA of the nodes N.

[0043] The weight w includes the weight w2 between the input layer 1121 and the intermediate layer 11222. The weight w2 includes K×M weights w2(k, m) (= w2(1, 1), w2(1, 2), ···, w2(1, M), w2(2, 1), ···, w2(K, M - 1), w2(K, M)) corresponding to the K×M connection paths between the input layer 1121 and the intermediate layer 11222 respectively. That is, the weight w2 is a vector determined by the K×M weights w2(k, m). The weight w further includes the weight w5 between the intermediate layer 11222 and the linear layer 1124. The weight w5 includes M×P weights w5(m, p) (= w5(1, 1), w5(1, 2), ···, w5(1, P), w5(2, 1), ···, w5(M, P - 1), w5(M, P)) corresponding to the M×P connection paths between the intermediate layer 11222 and the linear layer 1124 respectively. That is, the weight w5 is a vector determined by the M×P weights w5(m, p). The weight w further includes the weight w3 between the linear layer 1124 and the intermediate layer 11223. The weight w3 includes N weights w5(p, n) corresponding to the N connection paths (effective paths) between the linear layer 1124 and the intermediate layer 11223 respectively. When the output H5#p_n of the linear layer 1124 (where p_n is a variable indicating any one of the P integers from 1 to P) is input to the n-th node N3#n of the intermediate layer 11223, the N weights w5(p, n) include w3(p_1, 1), w3(p_2, 2), ···, w3(p_n, n), ···, w3(p_N - 1, N - 1) and w5(P_N, N). That is, the weight w3 is a vector determined by the N weights w3(p, n). The weight w further includes the weight w4 between the intermediate layer 11223 and the output layer 1123. The weight w4 includes N×O weights w4(n, o) (= w4(1, 1), w4(1, 2), ···, w4(1, O), w4(2, 1), ···, w4(N, O - 1), w4(N, O)) corresponding to the N×O connection paths between the intermediate layer 11223 and the output layer 1123 respectively. That is, the weight w4 is a vector determined by the N×O weights w4(n, o).

[0044] The connection mode CA includes the connection mode CA2 between the nodes N1#1 to N1#K included in the input layer 1121 and the nodes N2#1 to N2#M included in the intermediate layer 11222. Here, the connection mode between the nodes N of one layer and the nodes N of another layer is information indicating the presence or absence of a connection between the nodes N of one layer and the nodes N of another layer. That is, the connection mode between the nodes N of one layer and the nodes N of another layer here is information indicating whether there is an effective connection path (i.e., an effective path) through which the output of the nodes N of one layer is actually input to the nodes N of another layer. Therefore, the connection mode CA2 includes information regarding the effective paths between the input layer 1121 and the intermediate layer 11222. The connection mode CA includes the connection mode CA5 between the nodes N2#1 to N2#M included in the intermediate layer 11222 and the nodes N5#1 to N5#P included in the linear layer 1124. The connection mode CA5 includes information regarding the effective paths between the intermediate layer 11222 and the linear layer 1124. The connection mode CA includes the connection mode CA3 between the nodes N5#1 to N5#P included in the linear layer 1124 and the nodes N3#1 to N3#N included in the intermediate layer 11223. The connection mode CA3 includes information regarding the effective paths between the linear layer 1124 and the intermediate layer 11223. As described above, each node N3 of the intermediate layer 11223 is connected to one of the nodes N5#1 to N5#P of the linear layer 1124 via a single effective path. Therefore, the connection mode CA3 includes information regarding N effective paths corresponding to the N nodes N3#1 to N3#N of the intermediate layer 11223 respectively. The connection mode CA includes the connection mode CA4 between the nodes N3#1 to N3#N included in the intermediate layer 11223 and the nodes N4#1 to N4#O included in the output layer 1123. The connection mode CA3 includes information regarding the effective paths between the intermediate layer 11223 and the output layer 1123.

[0045] The bias b includes the bias b2 added in the intermediate layer 11222, the bias b3 added in the intermediate layer 11223, and the bias b4 added in the output layer 1123. The bias b2 includes M biases b2(m) (= b2(1), b2(2), ···, b2(M)) respectively added to the nodes N2#1 to N2#M included in the intermediate layer 11222. That is, the bias b2 is a vector determined by the M biases b2(m). The bias b3 includes N biases b3(n) (= b3(1), b3(2), ···, b3(N)) respectively added to the nodes N3#1 to N3#N included in the intermediate layer 11223. That is, the bias b3 is a vector determined by the N biases b3(n). The bias b4 includes O biases b4(o) (= b4(1), b4(2), ···, b4(O)) respectively added to the nodes N4#1 to N4#O included in the output layer 1123. That is, the bias b4 is a vector determined by the O biases b4(o).

[0046] These parameters are determined by the parameter determination device 2 described later. In this case, the parameter determination device 2 corresponds to the device responsible for learning, and it can also be said that inference is performed in the signal transmission device 1 (particularly, the distortion compensation circuit 11) using the parameters obtained by learning. Hereinafter, the parameter determination device 2 will be further described.

[0047] (2) Parameter determination device 2 (2-1) Configuration of parameter determination device 2 First, with reference to FIG. 3, the hardware configuration of the parameter determination device 2 of the present embodiment will be described. FIG. 3 is a block diagram showing the hardware configuration of the parameter determination device 2 of the present embodiment.

[0048] As shown in FIG. 3, the parameter determination device 2 includes an arithmetic unit 21 and a storage device 22. Further, the parameter determination device 2 may include an input device 23 and an output device 24. However, the parameter determination device 2 may not include at least one of the input device 23 and the output device 24. The arithmetic unit 21, the storage device 22, the input device 23, and the output device 24 may be connected via a data bus 25.

[0049] The arithmetic unit 21 includes, for example, a processor including at least one of a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), an FPGA (Field Programmable Gate Array), a TPU (Tensor Processing Unit), an ASIC (Application Specific Integrated Circuit), and a quantum processor. The arithmetic unit 21 may include a single processor or may include a plurality of processors. The arithmetic unit 21 reads a computer program. For example, the arithmetic unit 21 may read a computer program stored in the storage device 22. For example, the arithmetic unit 21 may read a computer program stored in a computer-readable and non-transitory recording medium using a recording medium reading device (not shown). The arithmetic unit 21 may obtain a computer program from a device (not shown) disposed outside the parameter determination device 2 via the input device 23 that can function as a receiving device (that is, it may be downloaded or read). The arithmetic unit 21 executes the read computer program. As a result, logical functional blocks for executing the operations to be performed by the parameter determination device 2 (specifically, parameter determination operations for determining the parameters of the neural network 112) are realized in the arithmetic unit 21. That is, the arithmetic unit 21 can function as a controller for realizing logical functional blocks for executing the operations to be performed by the parameter determination device 2.

[0050] FIG. 3 shows an example of logical function blocks implemented in the arithmetic unit 21 to execute the parameter determination operation. As shown in FIG. 3, in the arithmetic unit 21, a linear layer addition unit 211, a learning unit 212, and a path selection unit 213 are implemented as logical function blocks for executing the parameter determination operation. Details of the operations of the linear layer addition unit 211, the learning unit 212, and the path selection unit 213 will be described in detail later.

[0051] Note that FIG. 3 only conceptually (in other words, simply) shows the logical function blocks for executing the parameter determination operation. That is, the function blocks shown in FIG. 3 do not necessarily have to be directly implemented in the arithmetic unit 21. As long as the arithmetic unit 21 can perform the operations performed by the function blocks shown in FIG. 3, the configuration of the function blocks implemented in the arithmetic unit 21 is not limited to the configuration shown in FIG. 3.

[0052] The storage device 22 can store desired data. For example, the storage device 22 may temporarily store a computer program executed by the arithmetic unit 21. The storage device 22 may temporarily store data that the arithmetic unit 21 temporarily uses when executing the computer program. The storage device 22 may store data that the parameter determination device 2 stores long-term. Note that the storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. That is, the storage device 22 may include a non-temporary recording medium.

[0053] The input device 23 is a device that receives the input of information to the parameter determination device 2 from outside the parameter determination device 2. For example, the input device 23 may include an operating device operable by the user of the parameter determination device 2 (for example, at least one of a keyboard, a mouse, and a touch panel). For example, the input device 23 may include a reading device capable of reading information recorded as data on an externally attachable recording medium with respect to the parameter determination device 2. For example, the input device 23 may include a receiving device (that is, a communication device) capable of receiving information transmitted as data to the parameter determination device 2 from outside the parameter determination device 2 via a communication network.

[0054] The output device 24 is a device that outputs information to the outside of the parameter determination device 2. For example, the output device 24 may output information regarding the parameter determination operation performed by the parameter determination device 2 (for example, information regarding the determined parameters). As an example of such an output device 24, there is a transmitting device capable of transmitting information as data via a communication network or a data bus. As an example of the output device 24, there is a display capable of outputting information as an image (that is, capable of displaying). As an example of the output device 24, there is a speaker capable of outputting information as sound. As an example of the output device 24, there is a printer capable of outputting a document on which information is printed.

[0055] (2-2) Parameter determination operation Subsequently, with reference to FIG. 4, the flow of the parameter determination operation executed by the parameter determination device 2 of the present embodiment will be described. FIG. 4 is a flowchart showing the flow of the parameter determination operation.

[0056] As shown in FIG. 4, the linear layer addition unit 211 adds a linear layer 1124 to the initial (i.e., default) neural network 112_initial that includes an input layer 1121, intermediate layers 11222 to 11223, and an output layer 1123 but does not include the linear layer 1124 (step S11). That is, the linear layer addition unit 211 adds the linear layer 1124 to the initial neural network 112_initial (step S11). As described above, in this embodiment, an example in which the linear layer 1124 is arranged between the intermediate layer 11222 and the intermediate layer 11223 will be described. Therefore, the linear layer addition unit 211 adds the linear layer 1124 between the intermediate layer 11222 and the intermediate layer 11223.

[0057] As shown in the upper part of FIG. 5, a general neural network includes an intermediate layer 1122 between an input layer 1121 and an output layer 1123 but does not include a linear layer 1124. Therefore, as the initial neural network 112_initial, the general neural network shown in the upper part of FIG. 5 may be used. As a result of the linear layer addition unit 211 adding the linear layer 1124 to the initial neural network 112_initial, the initial neural network 112_initial changes to a neural network 112_learn that includes an input layer 1121, intermediate layers 11222 to 11223, an output layer 1123, and a linear layer 1124. The neural network 112_learn corresponds to a prototype of the neural network 112 actually used by the signal transmission device 1. In this case, the linear layer 1124 may be regarded as a new layer that constitutes a part of the neural network 112 (neural network 112_learn). In this case, the neural network 112_learn may be regarded as a neural network for learning that is used to determine (i.e., learn) the parameters of the neural network 112. Therefore, the parameter determination device 2 determines (i.e., learns) the parameters of the neural network 112 using the learning neural network 112_learn.

[0058] As described above, within the neural network 112, each node N3 in the intermediate layer 11223 is connected to any one of the nodes N5#1 to N5#P in the linear layer 1124 via a single effective path. On the other hand, at the timing when the linear layer 1124 is added in step S11, within the neural network 112_learn, each node N3 in the intermediate layer 11223 may be connected to each of the nodes N5#1 to N5#P via P connection paths that respectively connect each node N3 in the intermediate layer 11223 and the nodes N5#1 to N5#P in the linear layer 1124. That is, as shown in FIG. 6 showing the neural network 112_learn, within the neural network 112_learn, each node N5 in the linear layer 1124 may be connected to each of the nodes N3#1 to N3#N via N connection paths that respectively connect each node N5 in the linear layer 1124 and the nodes N3#1 to N3#N in the intermediate layer 11223. In other words, within the neural network 112_learn, the linear layer 1124 and the intermediate layer 11223 may be connected via P×N connection paths that connect the P nodes N5#1 to N5#P in the linear layer 1124 and the N nodes N3#1 to N3#N in the intermediate layer 11223. The parameter determination device 2 selects N effective paths corresponding to the N nodes N3#1 to N3#N in the intermediate layer 11223 from among the P×N connection paths between the linear layer 1124 and the intermediate layer 11223 by performing the processes from step S12 to step S14 described later.

[0059] Specifically, first, the learning unit 212 learns the parameters of the network part of the neural network 112_learn that is before (i.e., upstream of) the linear layer 1124 (step S12). Incidentally, the initial values of the parameters of the neural network 112_learn may be determined using random numbers. The parameters of the network part before (i.e., upstream of) the linear layer 1124 may include the weight w2 between the input layer 1121 and the intermediate layer 11222, the weight w5 between the intermediate layer 11222 and the linear layer 1124, the connection mode CA2 between the input layer 1121 and the intermediate layer 11222 (i.e., the effective path between the input layer 1121 and the intermediate layer 11222), the connection mode CA5 between the intermediate layer 11222 and the linear layer 1124 (i.e., the effective path between the intermediate layer 11222 and the linear layer 1124), and the bias b2 added in the intermediate layer 11222.

[0060] To learn the parameters of the neural network 112_learn, the learning unit 212 inputs a learning signal (i.e., learning data) to the neural network 112_learn. Then, the learning unit 212 changes the parameters of the neural network 112_learn so that the error (i.e., learning error) between the signal output by the neural network 112_learn and the teacher signal (i.e., correct data) becomes small (preferably, minimum). As the learning error, the mean squared error between the signal output by the neural network 112_learn and the teacher signal may be used. The determined parameters are used as the parameters of the network part of the neural network 112 that is before (i.e., upstream of) the linear layer 1124.

[0061] Each of the learning signal and the teacher signal may be a signal based on, for example, at least one of the input signal x, the distortion compensation signal y, and the transmission signal z. Each of the learning signal and the teacher signal may be a signal generated using, for example, at least one of the input signal x, the distortion compensation signal y, and the transmission signal z. The method for generating the learning signal and the teacher signal may be selected according to the algorithm for distortion compensation in the distortion compensation circuit 11. For example, when an algorithm compliant with the indirect learning method is used as the algorithm for distortion compensation in the distortion compensation circuit 11, a signal corresponding to the transmission signal z may be used as the learning signal, and a signal corresponding to the distortion compensation signal y or the input signal x may be used as the teacher signal. That is, when a certain learning signal is output as the transmission signal z from the power amplifier 12, the distortion compensation signal y to be output from the distortion compensation circuit 11 or the input signal x to be input to the distortion compensation circuit 11 may be used as the teacher signal. Alternatively, for example, when an algorithm compliant with the direct learning method is used as the algorithm for distortion compensation in the distortion compensation circuit 11, a signal corresponding to the input signal x may be used as the learning signal, and a signal corresponding to the distortion compensation signal y may be used as the teacher signal. That is, when a certain learning signal is input to the distortion compensation circuit 11, the distortion compensation signal y to be output from the distortion compensation circuit 11 (for example, the distortion compensation signal y obtained by applying ILC (Iterative Learning control)) may be used as the teacher signal.

[0062] After that, under the constraint that the learning unit 212 fixes (i.e., does not change) the parameters of the network part before the linear layer 1124 in the neural network 112_learn, it learns at least the weight w3 between the linear layer 1124 and the intermediate layer 11223 (step S13). That is, the learning unit 212 learns at least the weight w3 between the linear layer 1124 and the intermediate layer 11223 under the constraint of fixing (i.e., not changing) the parameters determined in step S12 (step S13). As will be described in detail later, the weight w3 learned by the learning unit 212 in step S13 is a parameter referred to by the path selection unit 213 to select an effective path between the linear layer 1124 and the intermediate layer 11223, and is not actually used as the weight w3 of the neural network 112. For this reason, for convenience of explanation, the weight w3 learned by the learning unit 212 in step S13 is denoted as "w3'", and is distinguished from the actual weight w3 of the neural network 112 (i.e., the weight w3 learned by the learning unit 212 in step S15 described later). To learn the weight w3', the learning unit 212 inputs a learning signal to the neural network 112_learn in which the parameters of the network part before the linear layer 1124 are fixed. Then, the learning unit 212 changes at least the weight w3' so that the error (i.e., the learning error) between the signal output by the neural network 112_learn and the teacher signal becomes small (preferably, minimum).

[0063] After the learning of the weight w3' in step S13 is completed, subsequently, the path selection unit 213 selects an effective path between the linear layer 1124 and the intermediate layer 11223 based on the weight w3' learned by the learning unit 212 (step S14). That is, the path selection unit 213 selects N effective paths corresponding to the N nodes N3#1 to N3#N of the intermediate layer 11223 from among the P×N connection paths between the linear layer 1124 and the intermediate layer 11223 based on the weight w3' (step S14). In other words, the path selection unit 213 selects one effective path for each of the N nodes N3#1 to N3#N of the intermediate layer 11223 from among the P×N connection paths between the linear layer 1124 and the intermediate layer 11223 based on the weight w3' (step S14).

[0064] Specifically, the path selection unit 213 selects one of the P connection paths as the valid path based on the P weights w3'(1, n) to w3'(P, n) in the P connection paths connecting the n-th node N3#n of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124. On the other hand, the path selection unit 213 does not select the remaining P - 1 connection paths among the P connection paths connecting the n-th node N3#n of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124 as the valid path. The path selection unit 213 performs the same operation for each of the N nodes N3#1 to N3#N of the intermediate layer 11223. That is, the path selection unit 213 selects a single valid path connected to the node N3#1 based on the P weights w3'(1, 1) to w3'(P, 1) in the P connection paths connecting the first node N3#1 of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124, selects a single valid path connected to the node N3#2 based on the P weights w3'(1, 2) to w3'(P, 2) in the P connection paths connecting the second node N3#2 of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124, ···, and selects a single valid path connected to the node N3#N based on the P weights w3'(1, N) to w3'(P, N) in the P connection paths connecting the N-th node N3#N of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124.

[0065] The path selection unit 213 may select, as a valid path, one connection path with the maximum weight w3(p, n)' from among P connection paths connecting the n-th node N3#n of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124. On the other hand, the path selection unit 213 does not necessarily have to select, as valid paths, the remaining P-1 connection paths among the P connection paths connecting the n-th node N3#n of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124, where the weight w3(p, n)' is not the maximum. That is, the path selection unit 213 selects, as a valid path connected to the node N3#1, one connection path with the maximum weight w3'(p, 1) from among the P connection paths connecting the first node N3#1 of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124, selects, as a valid path connected to the node N3#2, one connection path with the maximum weight w3'(p, 2) from among the P connection paths connecting the second node N3#2 of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124, ···, and selects, as a valid path connected to the node N3#N, one connection path with the maximum weight w3'(p, N) from among the P connection paths connecting the N-th node N3#N of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124.

[0066] Specifically, for example, among the P connection paths connected to the first node N3#1 of the intermediate layer 11223, when the weight w3'(1, 1) of the connection path connecting the first node N5#1 of the linear layer 1124 and the node N3#1 is maximized (that is, the condition w3'(1, 1)>w3'(2, 1), w3'(3, 1), ···, w3'(P, 1) is satisfied), as shown in FIG. 7, the path selection unit 213 may select the connection path connecting the node N5#1 and the node N3#1 as the valid path. For example, among the P connection paths connected to the second node N3#2 of the intermediate layer 11223, when the weight w3'(1, 2) of the connection path connecting the first node N5#1 of the linear layer 1124 and the node N3#2 is maximized (that is, the condition w3'(1, 2)>w3'(2, 2), w3'(3, 2), ···, w3'(P, 2) is satisfied), as shown in FIG. 7, the path selection unit 213 may select the connection path connecting the node N5#1 and the node N3#2 as the valid path. For example, among the P connection paths connected to the third node N3#3 of the intermediate layer 11223, when the weight w3'(2, 3) of the connection path connecting the second node N5#2 of the linear layer 1124 and the node N3#3 is maximized (that is, the condition w3'(2, 3)>w3'(1, 3), w3'(3, 3), ···, w3'(P, 3) is satisfied), as shown in FIG. 7, the path selection unit 213 may select the connection path connecting the node N5#2 and the node N3#3 as the valid path. For example, among the P connection paths connected to the Nth node N3#N of the intermediate layer 11223, when the weight w3'(P, N) of the connection path connecting the Pth node N5#P of the linear layer 1124 and the node N3#N is maximized (that is, the condition w3'(P, N)>w3'(1, N), w3'(2, N), ···, w3'(P - 1, N) is satisfied), as shown in FIG. 7, the path selection unit 213 may select the connection path connecting the node N5#P and the node N3#N as the valid path.

[0067] On the other hand, the connection paths not selected by the path selection unit 213 are not used as valid connection paths in the neural network 112. That is, in the neural network 112 based on the parameters determined by the parameter determination device 2, the node N5 of the linear layer 1124 and the node N3 of the intermediate layer 11223 are not connected via the connection paths not selected by the path selection unit 213. Therefore, the operation of selecting the valid path is substantially equivalent to the operation of determining the connection mode CA (in this case, the connection mode CA3). FIG. 8 is a schematic diagram showing the valid paths between the linear layer 1124 and the intermediate layer 11223, while the connection paths other than the valid paths are not shown. As shown in FIG. 8, it can be seen that by selecting the valid paths, the structure between the linear layer 1124 and the intermediate layer 11223 becomes a sparse structure. Such a sparse structure leads to a reduction in the computational amount of the neural network 112.

[0068] After that, after the selection of the valid paths by the path selection unit 213 is completed, the learning unit 212 fixes (i.e., does not change) the parameters of the network part before the linear layer 1124 in the neural network 112_learn to the parameters determined in step S12, and then learns the parameters related to the intermediate layer 11223 (step S15). The parameters related to the intermediate layer 11223 include the weight w3 between the linear layer 1124 and the intermediate layer 11223 and the bias b3 added in the intermediate layer 11223.

[0069] In step S15, the learning unit 212 does not use the connection paths not selected by the path selection unit 213 as valid connection paths. That is, the learning unit 212 learns the parameters related to the intermediate layer 11223 under the constraint condition that the nodes N are not connected via the connection paths not selected by the path selection unit 213. For example, the learning unit 212 may learn the parameters related to the intermediate layer 11223 under the condition that the weight w3 of the connection paths not selected by the path selection unit 213 becomes zero.

[0070] After that, under the constraint condition that the learning unit 212 fixes (i.e., does not change) the parameters of the network part before the intermediate layer 11223 in the neural network 112_learn to the parameters determined in steps S12 and S15, the learning unit 212 learns the parameters of the network part after (i.e., downstream of) the intermediate layer 11223 in the neural network 112_learn (step S16). The parameters of the network part after the intermediate layer 11223 include the weight w4 between the intermediate layer 11223 and the output layer 1123 and the bias b4 added at the output layer 1123. As a result, the parameters of the neural network 112 are determined.

[0071] Such a parameter determination device 2 typically determines the parameters of the neural network 112 before the shipment of the signal transmission device 1. As a result, for example, at a manufacturing factory, the signal transmission device 1 in which the neural network 112 based on the parameters determined by the parameter determination device 2 is implemented is shipped. In this case, typically, the parameter determination device 2 may be implemented using a device external to the signal transmission device 1 (typically, a relatively high-speed arithmetic device such as a GPU). However, as will be described later, at least a part of the parameter determination device 2 may be implemented in the signal transmission device 1. The parameter determination device 2 may determine the parameters of the neural network 112 after the shipment of the signal transmission device 1 (for example, during the operation of the signal transmission device 1).

[0072] (2-3) Technical effect of parameter determination device 2 As described above, the parameter determination device 2 can select one effective path for each of the N nodes N3#1 to N3#N of the intermediate layer 11223 from among the P×N connection paths between the linear layer 1124 and the intermediate layer 11223. As a result, the parameter determination device 2 can construct a neural network 112 including the linear layer 1124 connected via N effective paths to the intermediate layer 11223. Therefore, the parameter determination device 2 can construct a neural network 112 with a relatively small amount of necessary computation as compared with a comparative example neural network that does not include the linear layer 1124. Specifically, in the comparative example neural network that does not include the linear layer 1124, the intermediate layer 11222 and the intermediate layer 11223 are connected via M×N connection paths (i.e., effective paths) that connect the nodes N2#1 to N2#M of the intermediate layer 11222 and the nodes N3#1 to N3#N of the intermediate layer 11223, respectively. For this reason, the signal transmission device 1 needs to perform matrix multiplication M×N times to generate the output of the intermediate layer 11223 from the output of the intermediate layer 11222. On the other hand, in the neural network 112 of the present embodiment, the intermediate layer 11222 and the linear layer 1124 are connected via M×P connection paths (i.e., effective paths) that connect the nodes N2#1 to N2#M of the intermediate layer 11222 and the nodes N5#1 to N5#P of the linear layer 1124, respectively. For this reason, the signal transmission device 1 needs to perform matrix multiplication M×P times to generate the output of the linear layer 1124 from the output of the intermediate layer 11222. Further, the linear layer 1124 and the intermediate layer 11223 are connected via N connection paths. For this reason, the signal transmission device 1 needs to perform matrix multiplication N times to generate the output of the intermediate layer 11223 from the output of the linear layer 1124. Therefore, the signal transmission device 1 needs to perform matrix multiplication M×P + N times to generate the output of the intermediate layer 11223 from the output of the intermediate layer 11222. Here, as described above, since the number of nodes N5 in the linear layer 1124 is smaller than the number of nodes N3 in the intermediate layer 11223, P < N holds. When such a condition of P < N holds, the number of times of matrix multiplication required in the present embodiment (= M×P + N) is likely to be smaller than the number of times of matrix multiplication required in the comparative example (= M×N).Therefore, the parameter determination device 2 can construct a neural network 112 with a relatively small amount of necessary computations as compared with a comparative example neural network that does not include the linear layer 1124. As a result, the signal transmission device 1 can transmit the input signal x using the neural network 112 with a relatively small amount of necessary computations.

[0073] In particular, in the neural network of the comparative example, the outputs H2#1 to H2#M from the nodes N2#1 to N2#M of the intermediate layer 11222 (or any first layer) may become similar to each other. In this case, it is assumed that similar input signals are input to the nodes N3#1 to N3#N of the intermediate layer 11223 (or any second layer to which the output of any first layer is input) to which the output of the intermediate layer 11222 is input. However, even in such a situation, in a general neural network, the set of weights w3(1,1) to w3(M,1) for generating the input signal input to node N3#1, the set of weights w3(1,2) to w3(M,2) for generating the input signal input to node N3#2, ···, and the set of weights w3(1,N) to w3(M,N) for generating the input signal input to node N3#N are often sets of completely different weights. As a result, in order to generate a plurality of input signals that are similar to each other, a plurality of operations using sets of completely different weights are performed independently (in other words, in parallel). For this reason, the amount of computation may increase more than necessary. On the other hand, in the present embodiment, a linear layer 1124 including a number of nodes N5 less than the number of nodes N3 in the intermediate layer 11223 is added. For this reason, the structure for generating the input signals input to at least two nodes N3 in the intermediate layer 11223 is shared. That is, the signal transmission device 1 can input the output of the same node N5 to at least two different nodes N3. In other words, the signal transmission device 1 can generate the input signals input to at least two nodes N3 using the same node N5. Thus, in the present embodiment, since the structure for generating the input signals input to at least two nodes N in a certain layer is shared, a neural network 112 with a relatively small required amount of computation can be constructed compared to the case where the structure for generating the input signals input to at least two nodes N in a certain layer is not shared.

[0074] Also, when the number of matrix multiplications required in the present embodiment (= M × P + N) is smaller than the number of matrix multiplications required in the comparative example (= M × N), the condition M × P + N < M × N holds. By transforming the mathematical formula representing this condition, a mathematical formula P < N × (M - 1) / M is obtained. Therefore, the number P of nodes N5 in the linear layer 1124 and the number N of nodes N3 in the intermediate layer 11223 may satisfy the condition P < N × (M - 1) / M. In this case, the parameter determination device 2 can construct a neural network 112 with a more reliably reduced required amount of computation as compared with a comparative example neural network that does not include the linear layer 1124.

[0075] Further, the parameter determination device 2 can select, as an effective path, one connection path with the largest weight w3(p, n)' from among the P connection paths connecting the n-th node N3#n in the intermediate layer 11223 and the nodes N5#1 to N5#P in the linear layer 1124. Here, the connection path with the largest weight w3(p, n)' has a greater contribution to the output of the neural network 112 as compared with the connection paths with non-maximal weights w3(p, n)'. Therefore, when one connection path with the largest weight w3(p, n)' is selected as the effective path, the possibility of deterioration of the output of the neural network 112 (for example, a reduction in the effect of distortion compensation of the distortion compensation signal y) is smaller as compared with the case where one connection path with a non-maximal weight w3(p, n)' is selected as the effective path. Thus, the parameter determination device 2 can enjoy the effect of being able to construct the neural network 112 with a small amount of computation as described above by minimizing the number of effective paths between the intermediate layer 11223 and the linear layer 1124, while suppressing deterioration of the output of the neural network 112.

[0076] Also, in this embodiment, each node N5 of the linear layer 1124 is a node to which no bias b5 is added. Therefore, the signal transmission device 1 does not need to perform the matrix operation required for adding the bias b5 in the linear layer 1124. As a result, the parameter determination device 2 can construct a neural network 112 with a smaller required amount of computation as compared with other comparative example neural networks in which the bias b5 is added to each node N5 of the linear layer 1124.

[0077] (3) Modification example (3-1) Modification example of neural network 112 In the above description, any one of the outputs H5#1 to H5#P of the linear layer 1124 is input to each node N3 of the intermediate layer 11223. That is, each node N3 of the intermediate layer 11223 is connected to any one of the nodes N5#1 to N5#P of the linear layer 1124 via a single effective path. However, at least two of the outputs H5#1 to H5#P of the linear layer 1124 may be input to each node N3 of the intermediate layer 11223. That is, each node N3 of the intermediate layer 11223 may be connected to at least two of the nodes N5#1 to N5#P of the linear layer 1124 via at least two effective paths. In this case, in step S14 of FIG. 4, the path selection unit 213 may select at least two effective paths from among the P×N connection paths between the linear layer 1124 and the intermediate layer 11223 for at least one of the N nodes N3#1 to N3#N of the intermediate layer 11223. That is, the path selection unit 213 may select at least two of the P connection paths connecting the n-th node N3#n of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124 as at least two effective paths. For example, the path selection unit 213 may select at least two of the P connection paths connecting the n-th node N3#n of the intermediate layer 11223 and the nodes N5#1 to N5#P of the linear layer 1124 as at least two effective paths in descending order of the weight w3(p, n)'.

[0078] In the above description, each node N5 of the linear layer 1124 is a node to which the bias b5 is not added. However, at least one of the nodes N5#1 to N5#P of the linear layer 1124 may be a node to which the bias b5 is added. However, when each node N3 of the intermediate layer 11223 is connected to any one of the nodes N5#1 to N5#P of the linear layer 1124 via a single effective path as described above, the operation of adding the bias b5 to each node N5 of the linear layer 1124 is substantially equivalent to the operation of adding the bias b3 to each node N3 of the intermediate layer 11223. That is, the operation of adding the bias b5 to each node N5 of the linear layer 1124 can be replaced by the operation of adding the bias b3 to each node N3 of the intermediate layer 11223. Therefore, the effect obtained by adding the bias b5 to at least one of the nodes N5#1 to N5#P of the linear layer 1124 is particularly effective when each node N3 of the intermediate layer 11223 is connected to at least two of the nodes N5#1 to N5#P of the linear layer 1124 via at least two effective paths. Incidentally, when at least one of the nodes N5#1 to N5#P of the linear layer 1124 is a node to which the bias b5 is added, the parameter determination device 2 may learn the bias b5 in step S12 of FIG. 4.

[0079] (3-2) Modification example of signal transmission device 1 (3-2-1) Signal transmission device 1a of the first modification example With reference to FIG. 9, the signal transmission device 1a of the first modification will be described. FIG. 9 is a block diagram showing the configuration of the signal transmission device 1a of the first modification.

[0080] As shown in FIG. 9, the signal transmission device 1a is different in that it may be a device that transmits the transmission signal z via an optical communication network (for example, an optical line) compared to the signal transmission device 1. In this case, the signal transmission device 1a is different in that it further includes an E / O converter 13a that converts the transmission signal z output from the power amplifier 12 into an optical signal compared to the signal transmission device 1. As a result, the transmission signal z converted into an optical signal is transmitted via a signal propagation path 14a such as an optical fiber (that is, a signal propagation path constituting at least a part of the optical communication network). Part or all of this signal transmission path 14 may be a component constituting the signal transmission device 1a. Alternatively, this signal transmission path 14 may be a component separate from the signal transmission device 1a.

[0081] The signal receiving device 3a that receives the transmission signal z uses an O / E converter 31a to convert the transmission signal z, which is an optical signal, into an electrical signal, and then receives the transmission signal z converted into an electrical signal by the receiving unit 32a.

[0082] The distortion compensation circuit 11 may perform distortion compensation on the input signal x to compensate for the distortion that occurs in the transmission signal z due to the transmission of the transmission signal z in the signal propagation path 14a (that is, the distortion that occurs in the transmission signal z in the signal propagation path 14a) in addition to or instead of the distortion that occurs in the transmission signal z due to the operation of the power amplifier 12. As a result, even when the transmission signal z is transmitted via an optical communication network (for example, an optical line), the distortion of the transmission signal z is appropriately compensated. In this case, considering that distortion occurs in the transmission signal z in the signal propagation path 14a, each of the above-described learning signal and teacher signal may be a signal based on the received signal received by the signal receiving device 3a (that is, a signal including the distortion that occurs in the transmission signal z in the signal propagation path 14a) in addition to or instead of at least one of the input signal x, the distortion compensation signal y, and the transmission signal z.

[0083] In addition, when the transmission signal z converted into an optical signal is transmitted, the signal generation unit 111 may input the X polarization component and the Y polarization component of the input signal x t to the neural network 112 instead of the above-described various signals.

[0084] (3-2-2) Signal transmission device 1b of the second modification example Next, with reference to FIG. 10, the signal transmission device 1b of the second modification will be described. FIG. 10 is a block diagram showing the configuration of the signal transmission device 1b of the second modification.

[0085] As shown in FIG. 13, the signal transmission device 1b is different from the signal transmission device 1 in that a functional block for determining the parameters of the neural network 112 is realized within the signal transmission device 1b. Specifically, the signal transmission device 1b includes an arithmetic unit 15b. The arithmetic unit 15b reads a computer program. The computer program read by the arithmetic unit 15b may be recorded on any recording medium, similar to the computer program read by the arithmetic unit 21. The arithmetic unit 15b may control the distortion compensation circuit 11 and the power amplifier 12 by executing the read computer program. In particular, when the arithmetic unit 15b executes the read computer program, a logical functional block for determining the parameters of the neural network 112 is realized within the arithmetic unit 15b. Specifically, as shown in FIG. 10, a functional block similar to the functional block realized within the arithmetic unit 21 is realized within the arithmetic unit 15b. In this case, it can be said that substantially the parameter determination device 2 is implemented in the signal transmission device 1b.

[0086] In this case, the signal transmission device 1b itself can update the parameters of the neural network 112. Therefore, after the signal transmission device 1b is shipped, the parameters of the neural network 112 can be updated. For example, when the signal transmission device 1b is installed at its installation location, the parameters of the neural network 112 may be updated (in other words, adjusted) according to the actual usage situation of the signal transmission device 1b. For example, after the operation of the signal transmission device 1b is started, the parameters of the neural network 112 may be updated according to the characteristics of the transmission signal z actually transmitted by the signal transmission device 1b. For example, after the operation of the signal transmission device 1b is started, the parameters of the neural network 112 may be updated according to the deterioration over time (that is, drift) of the signal transmission device 1b. As a result, even after the signal transmission device 1b is shipped, the distortion compensation performance of the distortion compensation circuit 11 can be maintained in a relatively high state.

[0087] Furthermore, the signal transmission device 1b can update the parameters of the neural network 112 by using learning signals and teacher signals based on at least one of the input signal x actually input to the signal transmission device 1b, the distortion compensation signal y actually generated by the signal transmission device 1b, and the transmission signal z actually transmitted by the signal transmission device 1b. Therefore, the signal transmission device 1b can appropriately update the parameters of the neural network 112 according to the actual usage situation of the signal transmission device 1b.

[0088] (3-3) Modification example of parameter determination operation In the above description, in step S12 of FIG. 4, the learning unit 212 learns the parameters of the network portion of the neural network 112_learn before the linear layer 1124. However, in step S12 of FIG. 4, the learning unit 212 may learn the parameters of the network portion of the neural network 112_learn before the intermediate layer 11222. The parameters of the network portion before the intermediate layer 11222 may include the weight w2 between the input layer 1121 and the intermediate layer 11222, the connection mode CA2 between the input layer 1121 and the intermediate layer 11222 (that is, the effective path between the input layer 1121 and the intermediate layer 11222), and the bias b2 added in the intermediate layer 11222.

[0089] Thereafter, in step S13 of FIG. 4, the learning unit 212 may learn the weight w used in the network portion between the intermediate layer 11222 and the intermediate layer 11223 of the neural network 112_learn under the constraint condition of fixing the parameters of the network portion of the neural network 112_learn before the intermediate layer 11222 to the parameters determined in step S12. The weight w used in the network portion between the intermediate layer 11222 and the intermediate layer 11223 may include the weight w5 between the intermediate layer 11222 and the linear layer 1124, and the weight w3 (specifically, the weight w3') between the linear layer 1124 and the intermediate layer 11223.

[0090] Thereafter, in step S14 of FIG. 4, the path selection unit 213 may select an effective path between the linear layer 1124 and the intermediate layer 11223 based on the weight w3' learned by the learning unit 212 simultaneously with the learning of the weight w5.

[0091] After that, in step S15 of FIG. 4, the learning unit 212 fixes the parameters of the network part before the intermediate layer 11222 in the neural network 112_learn to the parameters determined in step S12, and under the constraint condition that the node N is not connected via the connection path not selected by the path selection unit 213, the learning unit 212 may learn the parameters of the network part between the intermediate layer 11222 and the intermediate layer 11223 in the neural network 112_learn. The parameters of the network part between the intermediate layer 11222 and the intermediate layer 11223 may include the weight w5 between the intermediate layer 11222 and the linear layer 1124, the weight w3 between the linear layer 1124 and the intermediate layer 11223, the connection mode CA5 between the intermediate layer 11222 and the linear layer 1124 (that is, the effective path between the intermediate layer 11222 and the linear layer 1124), and the bias b3 added at the intermediate layer 11223.

[0092] As described above, in the modification, as shown in FIG. 11, when the learning unit 212 learns the weight w3 (the weight w3 between the linear layer 1124 and the intermediate layer 11223) used to select the effective path between the linear layer 1124 and the intermediate layer 11223, the learning unit 212 simultaneously learns the parameters (for example, the weight w5) of the network part including the intermediate layer 11222 located before the linear layer 1124. In this case, based on the weight w3 learned in this way, the path selection unit 213 can more appropriately select the effective path between the linear layer 1124 and the intermediate layer 11223.

[0093] (4) Supplementary note Regarding the embodiments described above, the following additional remarks are further disclosed. [Additional Remark 1] A parameter determination device for determining the parameters of a neural network, Between the first layer and the second layer included in the neural network, there is added, as a new layer constituting a part of the neural network, a third layer in which (i) the output of the first layer is input and the output is input to the second layer, and (ii) the third layer includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is smaller than the number of a plurality of second nodes included in the second layer, and adding means for adding the third layer; learning means for learning, as part of the parameters, the weights between the third layer and the second layer; selection means for selecting, as an effective path used as an effective connection path in the neural network, one effective path for each of the second nodes as part of the parameters, from among a plurality of connection paths connecting the third node and the plurality of second nodes, based on the weights learned by the learning means; A parameter determination device comprising: [Appendix 2] The third layer includes P (P is a constant indicating an integer of 1 or more) third nodes, The selection means selects, as the effective path, one connection path having the maximum absolute value of the weights from among P connection paths respectively connecting the P third nodes and one of the plurality of second nodes, and does not select the remaining P - 1 connection paths whose absolute values of the weights do not become the maximum as the effective path; The parameter determination device according to Appendix 1. [Appendix 3] The first layer includes M (M is a constant indicating an integer of 2 or more) first nodes, The second layer includes N (N is a constant indicating an integer of 1 or more) second nodes, The third layer includes P (P is a constant indicating an integer of 1 or more) third nodes, The condition of P < N×(M - 1) / M is satisfied. The parameter determination device according to Appendix 1 or 2. [Appendix 4] The third node is a node to which no bias is added. The parameter determination device according to any one of Appendices 1 to 3. [Appendix 5] Distortion compensation means for generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network, Signal generation means for generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal, Comprising, The neural network is, A first layer that is an input layer or an intermediate layer, A second layer that is an intermediate layer or an output layer, A third layer into which the output of the first layer is input and the output of which is input to the second layer, the third layer including no non-linear activation function and including third nodes into which the outputs of a plurality of first nodes included in the first layer are input, the number of third nodes being less than the number of a plurality of second nodes included in the second layer, Including, To each of the plurality of second nodes, the output of a single one of the third nodes is input, Signal transmission device. [Appendix 6] The parameters of the neural network are determined by a parameter determination device, The parameter determination device is, Adding means for adding the third layer as a new layer constituting a part of the neural network between the first layer and the second layer, Learning means for learning, as part of the parameters, the weights between the third layer and the second layer, Selection means for selecting, as part of the parameters, one effective path for each second node from among a plurality of connection paths connecting the third node and the plurality of second nodes to be used as an effective connection path in the neural network based on the weights learned by the learning means, The signal transmission device according to Appendix 5, comprising. [Appendix 7] Distortion compensation means for generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network, Signal generation means for generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal Between a first layer and a second layer included in the neural network, (i) a third layer in which the output of the first layer is input and the output is input to the second layer, and (ii) a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and adding, as a new layer constituting a part of the neural network, a third layer including a number smaller than the number of a plurality of second nodes included in the second layer, adding means Learning means for learning, as part of the parameters of the neural network, the weights between the third layer and the second layer Based on the weights learned by the learning means, selection means for selecting, for each of the plurality of second nodes, as an effective path used as an effective connection path in the neural network, one effective path from among a plurality of connection paths connecting the third node and the plurality of second nodes as part of the parameters A signal transmission device comprising [Appendix 8] A parameter determination method for determining parameters of a neural network, comprising Between a first layer and a second layer included in the neural network, (i) a third layer in which the output of the first layer is input and the output is input to the second layer, and (ii) a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and adding, as a new layer constituting a part of the neural network, a third layer including a number smaller than the number of a plurality of second nodes included in the second layer Learning, as part of the parameters, the weights between the third layer and the second layer Based on the learned weights, selecting, for each of the plurality of second nodes, as an effective path used as an effective connection path in the neural network, one effective path from among a plurality of connection paths connecting the third node and the plurality of second nodes as part of the parameters A parameter determination method including [Appendix 9] Generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; Generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal; and wherein the neural network includes a first layer that is an input layer or an intermediate layer; a second layer that is an intermediate layer or an output layer; a third layer into which the output of the first layer is input and whose output is input to the second layer, the third layer including fewer third nodes to which the outputs of a plurality of first nodes included in the first layer are input than the number of a plurality of second nodes included in the second layer, and not including a non-linear activation function; and the output of a single third node is input to each of the plurality of second nodes; A signal transmission method. [Appendix 10] Generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; Generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal; adding, as a new layer constituting a part of the neural network, a third layer (i) into which the output of the first layer is input and whose output is input to the second layer, and (ii) including fewer third nodes to which the outputs of a plurality of first nodes included in the first layer are input than the number of a plurality of second nodes included in the second layer, and not including a non-linear activation function, between the first layer and the second layer included in the neural network; learning, as part of the parameters of the neural network, the weights between the third layer and the second layer; selecting, for each of the second nodes, as part of the parameters, a valid path that is used as a valid connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes based on the learned weights; A signal transmission method including [Appendix 11] A computer program for causing a computer to execute a parameter determination method for determining parameters of a neural network, wherein the parameter determination method adds, as a new layer constituting a part of the neural network, a third layer between a first layer and a second layer included in the neural network, where (i) the output of the first layer is input and the output is input to the second layer, and (ii) the third layer includes third nodes that do not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is less than the number of a plurality of second nodes included in the second layer; learns the weights between the third layer and the second layer as part of the parameters; selects, as part of the parameters, one valid path for each second node as a valid path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes based on the learned weights and includes. [Appendix 12] A computer program for causing a computer to execute a signal generation method, wherein the signal generation method generates a distortion compensation signal by performing distortion compensation on an input signal using a neural network; generates a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal and includes, wherein the neural network includes a first layer that is an input layer or an intermediate layer, includes a second layer that is an intermediate layer or an output layer, and includes a third layer where the output of the first layer is input and the output is input to the second layer, the third layer does not include a non-linear activation function, and includes third nodes to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is less than the number of a plurality of second nodes included in the second layer including the output of a single said third node is input to each of said plurality of second nodes A computer program. [Appendix 13] A computer program for causing a computer to execute a signal generation method, wherein the signal generation method includes: generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal; adding, as a new layer constituting a part of the neural network, a third layer between a first layer and a second layer included in the neural network, the third layer being such that (i) the output of the first layer is input thereto and the output is input to the second layer, and (ii) the third layer does not include a non-linear activation function and includes a third node to which the outputs of a plurality of first nodes included in the first layer are input, the number of the third nodes being smaller than the number of a plurality of second nodes included in the second layer; learning, as part of the parameters of the neural network, the weights between the third layer and the second layer; selecting, as part of the parameters, one effective path to be used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes for each second node, based on the learned weights; A computer program including the above.

[0094] The present invention can be appropriately modified within the scope not contrary to the gist or idea of the invention that can be read from the claims and the entire specification, and a parameter determination device, a parameter determination method, a signal transmission method, and a computer program involving such modifications are also included in the technical idea of the present invention.

Explanation of Signs

[0095] 1 Signal transmission device 11 Distortion compensation circuit 112 Neural Network 2 Parameter Determination Device 21 Arithmetic Unit 211 Linear Layer Addition Unit 212 Learning Unit 213 Route Selection Unit

Claims

1. A parameter determination device for determining parameters of a neural network, comprising: an adding means for adding, as a new layer constituting a part of the neural network, a third layer between a first layer and a second layer included in the neural network, wherein (i) an output of the first layer is input to the third layer and an output of the third layer is input to the second layer, and (ii) the third layer includes third nodes that do not include a non-linear activation function and to which outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is smaller than the number of a plurality of second nodes included in the second layer; a learning means for learning, as part of the parameters, weights between the third layer and the second layer; a selection means for selecting, as an effective path used as an effective connection path in the neural network, one effective path for each of the plurality of second nodes from among a plurality of connection paths connecting the third nodes and the plurality of second nodes, based on the weights learned by the learning means, as part of the parameters; a parameter determination device comprising the above.

2. The third layer includes P (P is a constant indicating an integer of 1 or more) third nodes, and the selection means selects, as the effective path, one connection path having the maximum absolute value of the weights from among P connection paths respectively connecting the P third nodes and one of the plurality of second nodes, and does not select the remaining P - 1 connection paths having the maximum absolute value of the weights as the effective path. The parameter determination device according to Claim 1.

3. The first layer includes M (M is a constant indicating an integer of 2 or more) first nodes, the second layer includes N (N is a constant indicating an integer of 1 or more) second nodes, the third layer includes P (P is a constant indicating an integer of 1 or more) third nodes, and a condition of P < N×(M - 1) / M is satisfied. The parameter determination device according to Claim 1 or 2.

4. The third node is a node to which no bias is added. The parameter determination device according to any one of Claims 1 to 3.

5. a distortion compensation means for generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; a signal generation means for generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal; wherein the neural network includes: a first layer that is an input layer or an intermediate layer; a second layer that is an intermediate layer or an output layer; and the neural network further includes a third layer between the first layer and the second layer, wherein (i) an output of the first layer is input to the third layer and an output of the third layer is input to the second layer, and (ii) the third layer includes third nodes that do not include a non-linear activation function and to which outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is smaller than the number of a plurality of second nodes included in the second layer. A third layer to which the output of the first layer is input and the output of which is input to the second layer, the third layer including third nodes that do not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, the third layer including a number of third nodes that is less than the number of a plurality of second nodes included in the second layer comprising to each of the plurality of second nodes, the output of a single one of the third nodes is input A signal transmission device. **Claim 6** The parameters of the neural network are determined by a parameter determination device, wherein the parameter determination device between the first layer and the second layer, addition means for adding the third layer as a new layer constituting a part of the neural network; learning means for learning, as part of the parameters, the weights between the third layer and the second layer; selection means for selecting, as part of the parameters, for each of the second nodes, one effective path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes, based on the weights learned by the learning means The signal transmission device according to claim 5, comprising the above. **Claim 7** Distortion compensation means for generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; signal generation means for generating a transmission signal to be transmitted to a signal reception device by performing a predetermined operation on the distortion compensation signal; between a first layer and a second layer included in the neural network, (i) a third layer to which the output of the first layer is input and the output of which is input to the second layer, (ii) the third layer including third nodes that do not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, the third layer including a number of third nodes that is less than the number of a plurality of second nodes included in the second layer, addition means for adding the third layer as a new layer constituting a part of the neural network; learning means for learning, as part of the parameters of the neural network, the weights between the third layer and the second layer; selection means for selecting, as part of the parameters, for each of the second nodes, one effective path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes, based on the weights learned by the learning means A signal transmission device comprising the above. **Claim 8** A parameter determination method for determining parameters of a neural network, between a first layer and a second layer included in the neural network, adding, as a new layer constituting a part of the neural network, a third layer in which (i) the output of the first layer is input and the output is input to the second layer, and (ii) the third layer includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is less than the number of a plurality of second nodes included in the second layer; learning weights between the third layer and the second layer as part of the parameters; selecting, for each of the plurality of second nodes, as part of the parameters, an effective path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes based on the learned weights; A parameter determination method executed by a computer, including the above steps.

9. generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal; including; the neural network includes a first layer that is an input layer or an intermediate layer, a second layer that is an intermediate layer or an output layer, and a third layer in which the output of the first layer is input and the output is input to the second layer, the third layer including a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is less than the number of a plurality of second nodes included in the second layer; including; the output of a single third node is input to each of the plurality of second nodes; A signal transmission method executed by a computer.

10. generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal; Between the first layer and the second layer included in the neural network, add, as a new layer constituting a part of the neural network, a third layer in which (i) the output of the first layer is input and the output is input to the second layer, and (ii) the third layer includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is smaller than the number of a plurality of second nodes included in the second layer. Learn, as a part of the parameters of the neural network, the weights between the third layer and the second layer. Based on the learned weights, select, for each of the plurality of second nodes, as a part of the parameters, an effective path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes. A signal transmission method executed by a computer including the above.

11. A computer program for causing a computer to execute a parameter determination method for determining parameters of a neural network, The parameter determination method includes: Between the first layer and the second layer included in the neural network, add, as a new layer constituting a part of the neural network, a third layer in which (i) the output of the first layer is input and the output is input to the second layer, and (ii) the third layer includes a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, and the number of the third nodes is smaller than the number of a plurality of second nodes included in the second layer. Learn, as a part of the parameters, the weights between the third layer and the second layer. Based on the learned weights, select, for each of the plurality of second nodes, as a part of the parameters, an effective path used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes. A computer program including the above.

12. A computer program for causing a computer to execute a signal generation method, The signal generation method includes: Generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; and Generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal. Including: The neural network is a first layer that is an input layer or an intermediate layer, a second layer that is an intermediate layer or an output layer, a third layer to which the output of the first layer is input and the output of which is input to the second layer, the third layer including a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, the third layer including a number of third nodes that is less than the number of a plurality of second nodes included in the second layer and includes the output of a single one of the third nodes is input to each of the plurality of second nodes computer program. **Claim 13** A computer program for causing a computer to execute a signal generation method, the signal generation method including generating a distortion compensation signal by performing distortion compensation on an input signal using a neural network; generating a transmission signal to be transmitted to a signal receiving device by performing a predetermined operation on the distortion compensation signal; adding, as a new layer constituting a part of the neural network, a third layer between a first layer and a second layer of the neural network, the third layer being (i) a third layer to which the output of the first layer is input and the output of which is input to the second layer, and (ii) including a third node that does not include a non-linear activation function and to which the outputs of a plurality of first nodes included in the first layer are input, the third layer including a number of third nodes that is less than the number of a plurality of second nodes included in the second layer; learning, as part of the parameters of the neural network, weights between the third layer and the second layer; selecting, as part of the parameters, for each second node, one effective path that is used as an effective connection path in the neural network from among a plurality of connection paths connecting the third node and the plurality of second nodes based on the learned weights computer program.

Citation Information

Patent Citations

  • Measuring system for working ratio of parallel computer

    JP1992062644A

  • Waveform distortion compensation device using neural network

    JP1997247051A

  • Neural network processing device

    JP2002251601A

  • Computer vision system and method

    JP2020071862A

  • Learning device, voice activity detector, and method for detecting voice activity

    WO2019162990A1