Improved neural network output post-processing method for multiple model mechanism
By introducing a multi-model layer structure into the neural network and pre-setting the model state and weight vector, the problems of slow training speed and limited accuracy improvement of traditional neural networks are solved, achieving more efficient training and more accurate identification or prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2026-03-24
AI Technical Summary
Traditional neural network training methods are slow, computationally intensive, and have limited accuracy improvement in regression problems.
A multi-model-layer neural network structure is adopted, including a neural network module and a multi-model-layer module. By setting multiple model units between the hidden layer and the output layer, the output results are merged to improve training efficiency and accuracy. The state and weight vector of the model are preset and remain unchanged before training.
It improves training efficiency, reduces the initial loss function value, reduces the number of training iterations and time, and enhances the recognition or prediction accuracy of neural networks.
Smart Images

Figure CN115994562B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a neural network post-processing method, specifically to an improved multi-model mechanism neural network output post-processing method, belonging to the field of artificial intelligence technology. Background Technology
[0002] With the development of computing and storage capabilities, neural networks, which can analyze large amounts of data, have been widely used in various fields.
[0003] The output of a neural network is influenced by various factors such as the initialization method and model structure, and its output range is highly random at the beginning of training. Therefore, when dealing with regression problems, the neural network's output needs to be trained to the required range first, and then further fine-tuned.
[0004] Traditional neural network training methods slow down the training speed, waste a lot of computation and time, and are detrimental to improving the accuracy of neural network output. Therefore, it is necessary to study a method that can improve the training speed of neural networks in regression problems. Summary of the Invention
[0005] To overcome the above problems, the inventors conducted in-depth research and proposed an improved multi-model mechanism for neural network output post-processing. This method involves setting up a multi-model layer neural network and using the trained network to identify or predict input data, thereby obtaining the final identification or prediction result.
[0006] The multi-model layer neural network includes a neural network module and a multi-model layer module.
[0007] The neural network module includes an input layer, a hidden layer, and an output layer. The multi-model layer module is located between the hidden layer and the output layer of the neural network module. The multi-model layer module contains a model, which contains multiple units. Each unit takes the output of the hidden layer as its input. The outputs of the multiple units are merged and regressed, and then passed to the output layer as the output of the model.
[0008] Preferably, the number of neurons in the output layer is multiple, so as to output results of different categories.
[0009] Preferably, the number of models is the same as the number of neurons in the output layer, such that each model corresponds to one neuron, and the output of the model is only transmitted to its corresponding neuron.
[0010] In a preferred embodiment, the plurality of units is set to 2 to 10.
[0011] In a preferred embodiment, the output of the model is set as follows:
[0012]
[0013] Among them, O i Let Λ represent the regression value of the i-th model group. i Let G represent the state vector of the i-th model group, where the state vector is the set of state values of different units in the model. i Let represent the weight vector of the i-th model group, where the weight vector represents the weights of different units in the model.
[0014] In a preferred embodiment, different unit state values in the model are pre-set before training the multi-model-layer neural network, and the unit state values remain unchanged during the training process.
[0015] In a preferred embodiment, the model's state vector is set as follows:
[0016]
[0017] Among them, D i p represents the maximum output range of the output layer neuron corresponding to the i-th model group. i This indicates the total number of units, and the superscript T indicates transpose.
[0018] In a preferred embodiment, the model's weight vector is set as follows:
[0019]
[0020] Among them G i Let G represent the weight vector of the i-th model. i,j h represents the weight of the j-th unit in the i-th model. last w represents the output of the last hidden layer. out This represents the transfer matrix from the last hidden layer to the multi-model layer module, b. out This represents the bias of the multi-model layer module, and () is the activation function.
[0021] In a preferred embodiment, during the training phase, it is not necessary to first train the output of the neural network output layer to the required range.
[0022] On the other hand, the present invention also provides an improved multi-model mechanism for neural network output post-processing identification or prediction.
[0023] The device includes a neural network module and a multi-model layer module. The neural network module includes an input layer, hidden layers, and an output layer.
[0024] The input layer is used to process input data;
[0025] The hidden layer is connected to the input layer and performs calculations on the data output by the input layer.
[0026] The output layer is used to output the identification or prediction results, and the output layer has multiple neurons to output multiple results;
[0027] The multi-model layer module is located between the hidden layer and the output layer of the neural network module. The multi-model layer module has a model, which contains multiple units. Each unit takes the output of the hidden layer as input, and the outputs of the multiple units are merged and regressed as the output of the model and passed to the output layer.
[0028] In a preferred embodiment, in a multi-model layer module, different models and different units have different state values λ. i,j The state vector of the i-th model is:
[0029]
[0030] Among them, D i p represents the maximum output range of the output layer neuron corresponding to the i-th model group. i This indicates the total number of units, and the superscript T indicates transpose.
[0031] The beneficial effects of this invention include:
[0032] (1) This enables the multi-model mechanism to be used in multi-output scenarios;
[0033] (2) The training process does not require training the output of the neural network to the required range, which improves training efficiency;
[0034] (3) It reduces the initial loss function value of neural network training. The loss function decreases rapidly during training, which can reduce the number of training sessions and reduce training time.
[0035] (4) The final loss function value is small, which improves the recognition or prediction accuracy of the neural network. Attached Figure Description
[0036] Figure 1 A schematic diagram of a neural network structure for an improved multi-model mechanism according to a preferred embodiment of the present invention is shown.
[0037] Figure 2 A schematic diagram of the neural network structure of the improved multi-model mechanism in Example 1 is shown;
[0038] Figure 3 The loss function of the neural network after training is shown in Example 1 and Comparative Example 1. Detailed Implementation
[0039] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become clearer and more apparent.
[0040] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments. Although various aspects of embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless specifically indicated otherwise.
[0041] This invention provides an improved post-processing method for neural network output using a multi-model mechanism. By setting up a multi-model layer neural network, the trained multi-model layer neural network is used to identify or predict input data to obtain the final identification or prediction result.
[0042] Preferably, the input data includes the acceleration of our aircraft, the velocity tilt angle of our aircraft, the velocity deflection angle of our aircraft, the relative positions of the enemy and our aircraft, and the relative velocities of the enemy and our aircraft. The final identification result is the guidance law of the enemy interceptor missile. A multi-model layer neural network identifies the input data to obtain the guidance law of the enemy interceptor missile. Based on the guidance law of the enemy interceptor missile, the missile's own guidance law is corrected to achieve evasion of the enemy interceptor missile.
[0043] Specifically, the method includes the following steps:
[0044] S1. Set up a multi-model layer neural network;
[0045] S2. Train the multi-model layer neural network;
[0046] S3. Input the data to be identified or predicted into a multi-model-layer neural network to obtain the final identification or prediction result.
[0047] In S1, the multi-model-layer neural network includes a neural network module and a multi-model-layer module, such as... Figure 1 As shown.
[0048] The neural network module can be any known neural network, such as the GRU neural network, which includes an input layer, a hidden layer, and an output layer.
[0049] The input layer is used to process input data;
[0050] The hidden layer is connected to the input layer and performs calculations on the data output by the input layer.
[0051] The output layer, used to output identification or prediction results, can have multiple neurons to output a variety of results.
[0052] The multi-model layer module is set between the hidden layer and the output layer of the neural network module. It can have multiple models, each containing multiple units. Each unit takes the output of the hidden layer as input. The outputs of multiple units are merged and regressed, and then passed to the output layer as the output of the model. The output layer outputs the final result.
[0053] By setting up multiple models, the improved multi-model mechanism can handle multi-output problems, greatly expanding the applicability of the method.
[0054] Furthermore, the number of models in the multi-model layer module is the same as the number of neurons in the output layer, so that each model corresponds to one neuron, and the output of the model is only transmitted to its corresponding neuron.
[0055] Preferably, the plurality of units is set to 2 to 10, and the specific number of units can be determined by those skilled in the art through multiple experiments.
[0056] Furthermore, in the multi-model layer module, different state values λ are set for different models and different units. i,j In this context, the subscript i represents the model number, and j represents the unit number.
[0057] Preferably, the state value λ i,j The settings are pre-defined and remain unchanged during training.
[0058] Preferably, the set of different unit state values in the i-th group of models is denoted as the state vector Λ. i ,
[0059] More preferably, the state vector of the i-th model is set as follows:
[0060]
[0061] Among them, D i p represents the maximum output range required for the i-th model, that is, the maximum output range of the neurons in the output layer corresponding to the i-th model. i This indicates the total number of units, and the superscript T indicates transpose. This method ensures that the output of the neural network remains within a reasonable range, thereby accelerating network training and improving network robustness.
[0062] According to the present invention, in each model group, different units have different weights. Preferably, the weight vector of the model is set as follows:
[0063]
[0064] Among them G i Let G represent the weight vector of the i-th model group. i,jh represents the weight of the j-th unit in the i-th model group. last w represents the output of the last hidden layer. out This represents the transfer matrix from the last hidden layer to the multi-model layer module, b. out This represents the bias of the multi-model layer module, and () is the activation function.
[0065] Preferably, the activation function is the softmax function.
[0066] In a preferred embodiment, the regression value O of the i-th model group i express:
[0067]
[0068] Among them, Λ i Let G represent the state vector of the i-th model group. i This represents the weight vector of the i-th model group.
[0069] In this invention, by setting up a multi-model layer module, the idea of transfer learning is introduced into the multi-model mechanism, which reduces the final loss function and improves the recognition or prediction accuracy of the neural network.
[0070] In S2, a multi-model-layer neural network is trained using training samples.
[0071] Unlike traditional neural network training, in this invention, at the beginning of training, since the sum of the weight coefficients of different units in the model is 1, that is, the overall weight coefficient of the model is 1, it is not necessary to train the output of the neural network output layer to the required range first.
[0072] Furthermore, during training, the weight coefficients G of different units in the multi-model layer module are... i,j The training process adjusts itself to ensure that the output of the multi-model layer neural network is within the required range.
[0073] Compared to traditional neural networks, in the training process of this invention, the network's output value is within the required range from the beginning. In contrast, traditional neural networks need to train the output range to a reasonable range first (for example, if the initial output range of the neural network is [0, 100], and the sample label range in the sample library is [0, 0.01], then the traditional neural network needs to adjust the output range to [0, 0.01] before differential training can be performed). At the same time, this neural network allows different units to change, as long as the regression value after training is the required value, which makes the loss function decrease faster during training and the final loss function value obtained is smaller.
[0074] In S3, similar to traditional neural networks, the data to be identified or predicted is input into a multi-model-layer neural network to obtain the final identification or prediction result. The specific process will not be described in detail in this invention.
[0075] On the other hand, the present invention also provides an improved multi-model mechanism for neural network output post-processing identification or prediction, the device comprising a neural network module and a multi-model layer module, wherein the neural network module comprises an input layer, a hidden layer and an output layer.
[0076] The input layer is used to process input data;
[0077] The hidden layer is connected to the input layer and performs calculations on the data output by the input layer.
[0078] The output layer is used to output the identification or prediction results.
[0079] The multi-model layer module is located between the hidden layer and the output layer of the neural network module. The multi-model layer module has a model, which contains multiple units. Each unit takes the output of the hidden layer as input, and the outputs of the multiple units are merged and regressed as the output of the model and passed to the output layer.
[0080] Preferably, the output layer has multiple neurons to output multiple results, so as to output results of different categories;
[0081] Preferably, there are multiple models. More preferably, the number of models is the same as the number of neurons in the output layer, and the models correspond one-to-one with the neurons in the output layer, so that the output of a set of models is only transmitted to the corresponding output layer neurons.
[0082] Preferably, in the multi-model layer module, different models and different units have different state values λ. i,j The state vector of the i-th model is:
[0083]
[0084] Among them, D i p represents the maximum value of the range that the i-th model needs to output. i This indicates the total number of units, and the superscript T indicates transpose.
[0085] Preferably, the model's weight vector is:
[0086]
[0087] Among them G i Let G represent the weight vector of the i-th model group. i,j h represents the weight of the j-th unit in the i-th model group. lastw represents the output of the last hidden layer. out This represents the transfer matrix from the last hidden layer to the multi-model layer module, b. out This represents the bias of the multi-model layer module, and () is the activation function.
[0088] Preferably, the activation function is the softmax function.
[0089] In a preferred embodiment, the output value O of the i-th model group i express:
[0090]
[0091] Among them, Λ i Let G represent the state vector of the i-th model group. i This represents the weight vector of the i-th model group.
[0092] Example
[0093] Example 1
[0094] This paper addresses the problem of rapid identification of enemy interceptor missile guidance laws mentioned in the literature [Wang Yinhan, Fan Shipeng, Wu Guang, Wang Jiang, He Shaoming. A fast identification method for enemy interceptor missile guidance laws based on GRU [J / OL]. Acta Aeronautica Sinica: 1-11 [2021-06-09]. http: / / kns.cnki.net / kcms / detail / 11.1929.v.20210203.1402.013.html.]. By setting up a multi-model layer neural network, the trained multi-model layer neural network is used to identify the input data, thus obtaining the final identification result of the enemy interceptor missile guidance law.
[0095] The sample source and extraction method are consistent with those in the original literature. It should be noted that the original literature is a classification problem. In this embodiment, when extracting data, the sample output needs to be set as guidance law parameters, while the original literature is a non-classification label. Therefore, the sample output range is set to 2.5-5.5, the output range is set to 0.10-0.40, and the rest of the settings are exactly the same as those in the original literature.
[0096] In this embodiment, the multi-model-layer neural network includes a neural network module and a multi-model-layer module, such as... Figure 2 As shown, the neural network module is a GRU neural network, and its input layer, hidden layer and output layer structure is the same as the neural network structure in the original literature, that is, the number of hidden layers is 3, the number of GRU neurons in each hidden layer is 96, and the number of neurons in the output layer is two.
[0097] The multi-model layer module is set between the hidden layer and the output layer of the neural network module. The multi-model layer module has two sets of models, each set of models contains two units. Each unit takes the output of the hidden layer as input, and the output of the two units is merged and regressed as the output of the model and passed to the output layer.
[0098] The model's output value O i for:
[0099]
[0100] The model's state vector is:
[0101]
[0102] The model's weight vector is:
[0103]
[0104] statistics
[0105] Comparative Example 1
[0106] The method described in the original literature was used for identification to obtain the final identification results of the enemy interceptor missile's guidance law.
[0107] Experimental Example
[0108] The loss function is used to characterize the difference between the forward computation result of the neural network in each iteration and the true value. The smaller the loss function, the closer the neural network output is to the true value.
[0109] When the number of training iterations is 1 to 1000, the loss functions of the trained neural networks in Example 1 and Comparative Example 1 are as follows: Figure 3 As shown, the neural network trained by the method in Comparative Example 1 has an initial loss function of 0.3366, while the neural network trained by the method in Example 1 has an initial loss function of 0.0834, meaning that the initial training loss function value is greatly reduced.
[0110] Furthermore, as can be seen from the figure, the loss function in Example 1 decreases faster during the training process, meaning that the method in Example 1 requires fewer training iterations to achieve the same recognition accuracy as the method in Comparative Example 1.
[0111] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," "front," and "rear," etc., indicate the orientation or positional relationship based on the orientation or positional relationship in the working state of this invention, and are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. Furthermore, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0112] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0113] The present invention has been described above with reference to preferred embodiments; however, these embodiments are merely exemplary and illustrative. Various substitutions and modifications can be made to the present invention based on these embodiments, all of which fall within the scope of protection of the present invention.
Claims
1. An improved multi-model mechanism for neural network output post-processing, characterized in that, By setting up a multi-model-layer neural network, the trained multi-model-layer neural network is used to identify or predict input data to obtain the final identification or prediction result. The input data includes the acceleration of our aircraft, the velocity tilt angle of our aircraft, the velocity deflection angle of our aircraft, the relative positions of the enemy and our aircraft, and the relative velocities of the enemy and our aircraft. The final identification result is the guidance law of the enemy interceptor missile. The multi-model-layer neural network identifies the input data to obtain the guidance law of the enemy interceptor missile, and based on the guidance law of the enemy interceptor missile, the guidance law of our aircraft is corrected. The multi-model layer neural network includes a neural network module and a multi-model layer module. The neural network module includes an input layer, a hidden layer, and an output layer. The multi-model layer module is set between the hidden layer and the output layer of the neural network module. The multi-model layer module has a model, which contains multiple units. Each unit takes the output of the hidden layer as its input. The outputs of the multiple units are merged and regressed to become the output of the model and then passed to the output layer. The output layer contains multiple neurons to output results of different categories. The number of models is the same as the number of neurons in the output layer, so that each model corresponds to one neuron, and the output of the model is only transmitted to the neuron corresponding to it. The output of the model is set as follows: , in, Indicates the first The regression values of the group model Indicates the first The state vector of the model group, wherein the state vector is a set of state values of different units in the model. Indicates the first The weight vector of the model group, wherein the weight vector is the weight of different units in the model; The model's state vector is set as follows: , in, Indicates the first The maximum value of the output range of the neurons in the output layer corresponding to the group model. This indicates the total number of units; the superscript T indicates transpose. No. The weight vector of the group model is set as follows: , in Indicates the first The weight vector of the group model, Indicates the first In the group model, the first The weight of each unit, This represents the output of the last hidden layer. This represents the transfer matrix from the last hidden layer to the multi-model layer module. This indicates the bias of the multi-model layer module. This is the activation function.
2. The improved multi-model mechanism neural network output post-processing method according to claim 1, characterized in that, The number of units is set to 2 to 10.
3. The improved multi-model mechanism neural network output post-processing method according to claim 1, characterized in that, Before training a multi-model layer neural network, the state values of different units in the model are pre-set, and the unit state values remain unchanged during the training process.
4. The improved multi-model mechanism neural network output post-processing method according to claim 1, characterized in that, During the training phase, it is not necessary to first train the output of the neural network output layer to the required range.
Citation Information
Patent Citations
Track fusion method based on neural network
CN111582485A
Insect population density prediction method and system based on neural network multi-model combination
CN112862168A