Feed-forward network dynamic learning adjustment training method based on bee colony perception mechanism

By introducing a feedforward network dynamic learning adjustment method based on the swarm awareness mechanism in the Transformer network, the problem of parameters easily falling into local optimal solutions is solved, dynamic optimization of the model and adaptive adjustment of the learning rate are realized, and the training efficiency and generalization ability of the model are improved.

CN119940446APending Publication Date: 2025-05-06METASEQUOIA INTELLIGENCE (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510070655.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

During training, the parameters of the existing Transformer network are easily trapped in the local optimal solution, resulting in insufficient model optimization and dynamic adjustment capabilities, easy to overfit, and feature redundancy affects generalization performance.

Method used

The feedforward network dynamic learning adjustment training method based on the swarm awareness mechanism is adopted, and the feedforward network parameters are optimized through the swarm algorithm, the swarm awareness mechanism is activated, and the local and global information of the bee individual is used to cooperate, the weight parameters are updated, and the learning rate and regularization terms are dynamically adjusted to avoid overfitting.

Benefits of technology

By simulating the collective behavior of bee population, dynamic optimization of feedforward network parameters and adaptive adjustment of learning rate are achieved, the training efficiency and generalization ability of the model are improved, and overfitting and local optimal solutions are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940446A_ABST
    Figure CN119940446A_ABST
Patent Text Reader

Abstract

The invention discloses a feed-forward network dynamic learning adjustment training method based on a bee colony perception mechanism. The method comprises the following steps: initializing parameters of a feed-forward network structure and a bee colony algorithm; performing optimization search on the feedforward network weight through each bee individual in the bee colony algorithm; enabling the bee individuals to cooperate with global information according to local search, and updating the positions of the bee individuals; dynamically adjusting the learning rate of the feed-forward network according to the learning progress of the current feed-forward network; after training is finished, calculating a loss value between an output result of the feedforward network and a real label; and repeating the steps until a preset iteration stopping condition is met. The bee colony perceiving layer is used for optimizing and enhancing global information, the model is allowed to carry out dynamic learning adjustment in the training process, distribution of feedforward network weights is continuously optimized by simulating bee colony behaviors, and the situation that parameters fall into local optimum is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a feedforward network dynamic learning adjustment training method and a pre-trained language model based on a swarm awareness mechanism. Background Art

[0002] At present, there are still many problems in intelligent dialogue systems and related technical fields, which make it difficult to meet the increasingly complex user needs and application scenarios. With the rapid development of artificial intelligence technology, intelligent dialogue systems are widely used in customer service, education, medical care, finance and other fields. However, the existing dialogue systems still have obvious bottlenecks in natural language understanding and generation, task processing efficiency, data support capabilities and other aspects. One of the reasons is that the feedforward network of the large language model built based on the Transformer network is a neural network structure composed of neurons, which plays a role in enhancing the nonlinear processing ability of the model, including weights and bias parameters. When traditional neural networks use the gradient descent method to update parameters, there are gradient vanishing and gradient descent phenomena, as well as the phenomenon that the parameters fall into the local optimal solution after the training is completed, resulting in insufficient model optimization and dynamic adjustment capabilities, easy to fall into the local optimal, and easy to overfit during training, and feature redundancy affects the generalization performance.

[0003] In view of this, it is necessary to provide a feedforward network dynamic learning adjustment training method and a pre-trained language model based on the swarm awareness mechanism to overcome the above defects. Summary of the invention

[0004] The purpose of the present invention is to provide a feedforward network dynamic learning adjustment training method and a pre-trained language model based on a swarm awareness mechanism, aiming to solve the problem that parameters of the existing Transformer network are prone to fall into local optimal solutions during training.

[0005] In order to achieve the above object, the present invention provides a feedforward network dynamic learning adjustment training method based on swarm awareness mechanism, comprising:

[0006] Step S11: Initialize the parameters of the feedforward network structure and the bee colony algorithm;

[0007] Step S12: In each training iteration, the parameters of the feedforward network are optimized using the bee swarm algorithm. The swarm awareness mechanism in the bee swarm algorithm will be activated, and the feedforward network will optimize the feedforward network weights through each individual bee in the bee swarm algorithm according to the output characteristics of the current layer;

[0008] Step S13: enabling individual bees to cooperate with global information based on local search to update their positions;

[0009] Step S14: In each iteration, the learning rate of the feedforward network is dynamically adjusted according to the learning progress of the current feedforward network. If the current optimization progress is lower than the preset condition, the learning rate will be increased; conversely, if the feedforward network is close to the optimal solution, the learning rate will be reduced to avoid overfitting;

[0010] Step S15: After each training, the loss value between the output of the feedforward network and the true label is calculated. The loss function of the feedforward network uses the mean square error to measure the difference between the model output and the true label, and the weight of the regularization term is dynamically updated with the self-supervisory feedback during the training process to ensure that the influence of irrelevant or redundant features is gradually weakened;

[0011] Step S16: Repeat the above steps S12-S15 until the preset stop iteration condition is met, which means that the feedforward network training is completed.

[0012] In a preferred embodiment, in step S11, the data input to the feedforward network is X p , the output of the feedforward network is Y p , the weight parameter of the feedforward network is W p , the bias term of the feedforward network is b p , according to the standard structure of the feedforward network, the way to calculate the activation value of each layer is expressed as:

[0013]

[0014] In the formula, is the weighted input of layer I of the feedforward network; is the weight of layer I of the feedforward network; is the output of the I-1th layer of the feedforward network, which is also the input of the Ith layer of the feedforward network; is the bias term of the Ith layer of the feedforward network.

[0015] In a preferred embodiment, an activation function is used to perform nonlinear mapping on the weighted input of each layer of the feedforward network to obtain the output of the layer, which is expressed as:

[0016]

[0017] In the formula, is the output activation value of the Ith layer of the feedforward network; Sig( ) is the Sigmoid activation function.

[0018] In a preferred embodiment, in step S12, the weight parameters and bias parameters of the feedforward network corresponding to the state of the individual bee are set, and the fitness function is defined as the difference between the output of the feedforward network and the actual label. The specific calculation method is expressed as:

[0019]

[0020] In the formula, F p ( ) is the fitness function; is the weight configuration of the i-th bee individual; is the bias configuration of the i-th bee individual; is the model output of the jth sample, which is calculated by the output of the feedforward network through the preset Softmax function; is the true label of the jth sample; Ns is the number of samples input in the current batch; α p It is the coefficient of the adjustment term to control the influence of different samples on the fitness function.

[0021] In a preferred embodiment, in step S13, the updated position of the individual bees is represented as:

[0022] In the formula, is the weight configuration of the i-th bee individual; α pa is the first hyperparameter controlling the step size; β pa is the second hyperparameter controlling the step size; γ pa is the third hyperparameter that controls the step size; is the weight configuration of the g-th bee individual; is the weight configuration of the rth bee individual; is the current loss of the i-th bee individual; is the global optimal loss.

[0023] In a preferred embodiment, in step S14, the learning rate is updated as follows:

[0024]

[0025] In the formula, is the learning rate of the feedforward network at the tth iteration; is the learning rate of the feedforward network at the t+1th iteration; γ pec Hyperparameters adjusted to control learning rate; is the change in the loss function of the feedforward network; is the loss value of the tth iteration; δ pec is the tuning hyperparameter for the volatility factor; is the variance of the loss value.

[0026] In a preferred embodiment, in step S15, the calculation method of dynamic update is expressed as:

[0027]

[0028] Where, Lp is the loss function of the feedforward network; Ms is the total number of features; is the regularization coefficient of the jth feature; is the value of the jth feature.

[0029] In a preferred embodiment, in step S15, the weight of the regularization term is automatically adjusted according to the learning process of each feature, and the dynamic update rule enables the regularization term to reflect the learning difficulty and importance of each feature; wherein, the calculation method of the regularization coefficient is expressed as:

[0030]

[0031] In the formula, is the regularization coefficient of the j+1th feature; α ptr is the regularization coefficient learning rate; is the gradient of the loss function of the feedforward network with respect to the jth feature, indicating the contribution of this feature to the total loss.

[0032] In a preferred embodiment, α p Set to 0.3, γ pec Set to 0.2, δ pec Set to 0.3, α ptr Set to 0.2.

[0033] The present invention also provides a pre-trained language model, which is used as the master of a large language model to process the natural language understanding and generation of user input, wherein the pre-trained language model is a model based on a Transformer network, including an encoder and a decoder; wherein the encoder is used to map input features into a series of context-aware vector representations, and the decoder is used to gradually decode the vector representations into target features;

[0034] The encoder includes N identical encoding layers, each encoding layer includes two sublayers, a multi-head attention mechanism sublayer and a feedforward network sublayer, and the encoder adds a residual connection and a normalization operation after each sublayer;

[0035] The decoder includes N identical decoding layers, each decoding layer includes two multi-head attention mechanism sublayers and one feedforward network sublayer, and the decoder adds residual connections and normalization operations after each sublayer;

[0036] The feedforward networks of the feedforward network sublayer are all trained using the feedforward network dynamic learning adjustment training method based on the swarm awareness mechanism as described in any of the above implementations.

[0037] The feedforward network dynamic learning adjustment training method and pre-trained language model based on swarm awareness mechanism provided by the present invention have the following beneficial effects:

[0038] By simulating the collective behavior of a colony of bees in the process of searching for nectar, the search for the optimal solution is achieved by collecting and integrating local and global information. In the parameter update process of the feedforward network, a colony awareness layer is used in the feedforward network to optimize and enhance global information, allowing the model to dynamically learn and adjust during the training process, and by imitating the group behavior of bees, the distribution of the weights of the feedforward network is continuously optimized, and different data features can be continuously adjusted and adapted during the training process. In a preferred embodiment, the regularization weights are updated by combining dynamic feedback of feature learning, the model's recognition of the importance of different features is enhanced, and the adjustment coefficient and balance parameter of the entropy term are dynamically adjusted to ensure the diversity and authenticity of the generated samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0040] Figure 1 A flow chart of a feedforward network dynamic learning adjustment training method based on a swarm awareness mechanism provided by the present invention;

[0041] Figure 2 This is an architectural diagram of the pre-trained language model provided by the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and beneficial technical effect of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. It should be understood that the specific implementation methods described in this specification are only for explaining the present invention, not for limiting the present invention.

[0043] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0044] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0045] In an embodiment of the present invention, a feedforward network dynamic learning adjustment training method based on a swarm awareness mechanism is provided, which is used to train the feedforward network of a Transformer network, and the search for the optimal solution is realized by simulating the collective behavior of a bee colony in the process of searching for nectar, and by collecting and integrating local and global information. In the parameter update process of the feedforward network, a swarm awareness layer is used in the feedforward network for optimizing and enhancing global information, allowing the model to dynamically learn and adjust during the training process, and by imitating the group behavior of bees, the distribution of the weights of the feedforward network is continuously optimized, and different data features can be continuously adjusted and adapted during the training process. Therefore, the learning ability of a large language model for complex data features can be improved, and the dynamic learning rate adjustment strategy can ensure the efficiency and stability during the training process.

[0046] like Figure 1 As shown, the feedforward network dynamic learning adjustment training method based on the swarm awareness mechanism includes steps S11-S16.

[0047] Step S11: Initialize the parameters of the feedforward network structure and the bee colony algorithm. Suppose the data input to the feedforward network is X p (i.e. data features), the output of the feedforward network is Y p , the weight parameter of the feedforward network is W p , the bias term of the feedforward network is b p , according to the standard structure of the feedforward network, the way to calculate the activation value of each layer is expressed as:

[0048]

[0049] In the formula, is the weighted input of layer I of the feedforward network; is the weight of layer I of the feedforward network; is the output of the I-1th layer of the feedforward network, which is also the input of the Ith layer of the feedforward network; is the bias term of the Ith layer of the feedforward network.

[0050] Furthermore, the activation function is used to perform nonlinear mapping on the weighted input of each layer of the feedforward network to obtain the output of the layer, which is expressed as:

[0051]

[0052] In the formula, is the output activation value of the Ith layer of the feedforward network; Sig( ) is the Sigmoid activation function.

[0053] Step S12: In each training iteration, the parameters of the feedforward network are optimized using the swarm algorithm. The swarm awareness mechanism in the swarm algorithm will be activated. Specifically, the feedforward network will optimize the feedforward network weights through each individual (i.e., bee) in the swarm algorithm according to the output characteristics of the current layer. The weight parameters and bias parameters of the feedforward network corresponding to the state of the individual bee are set, and the fitness function is defined as the difference between the output of the feedforward network and the actual label. By adopting the adjustment term, the fitness function not only considers the prediction error, but also can dynamically adjust the sensitivity to different samples. The specific calculation method is expressed as:

[0054]

[0055] In the formula, F p ( ) is the fitness function; is the weight configuration of the i-th bee individual; is the bias configuration of the i-th bee individual; is the model output of the jth sample, which is calculated by the output of the feedforward network through the preset Softmax function; is the true label of the jth sample; Ns is the number of samples input in the current batch; α p is the coefficient of the adjustment term to control the influence of different samples on the fitness function. Preferably, α p α p Set to 0.3.

[0056] Step S13: The smaller the calculated fitness value is, the better the performance of the feedforward network is, and the closer the weight configuration of the individual bees is to the optimal solution. Therefore, the individual bees cooperate with the global information based on the local search to update their positions, which is expressed as:

[0057]

[0058] In the formula, is the weight configuration of the i-th bee individual; α pa is the first hyperparameter controlling the step size; β pa is the second hyperparameter controlling the step size; γ pa is the third hyperparameter that controls the step size; is the weight configuration of the g-th bee individual; is the weight configuration of the rth bee individual; is the current loss of the i-th bee individual; is the global optimal loss.

[0059] Step S14: In each iteration, the learning rate of the feedforward network is dynamically adjusted according to the learning progress of the current feedforward network. If the current optimization progress is relatively slow, the learning rate will be appropriately increased; on the contrary, if the feedforward network is close to the optimal solution, the learning rate will be reduced to avoid overfitting. The learning rate update method is expressed as:

[0060]

[0061] In the formula, is the learning rate of the feedforward network at the tth iteration; is the learning rate of the feedforward network at the t+1th iteration; γ pec Hyperparameters adjusted to control learning rate; is the change in the loss function of the feedforward network; is the loss value of the tth iteration; δ pec is the tuning hyperparameter for the volatility factor; is the variance of the loss value. Preferably, γ pec Set to 0.2, δ pec Set to 0.3.

[0062] Step S15: After each training, the loss value between the output of the feedforward network and the true label is calculated. The loss function of the feedforward network uses the mean square error to measure the difference between the model output and the true label, and the weight of the regularization term is dynamically updated with the self-supervisory feedback during the training process to ensure that the influence of irrelevant or redundant features is gradually weakened. The calculation method is expressed as:

[0063]

[0064] Where, L p is the loss function of the feedforward network; Ms is the total number of features; is the regularization coefficient of the jth feature; is the value of the jth feature.

[0065] Furthermore, in step S35, the weight of the regularization term is automatically adjusted according to the learning process of each feature, and the dynamic update rule enables the regularization term to reflect the learning difficulty and importance of each feature. The calculation method of the regularization coefficient is expressed as:

[0066]

[0067] In the formula, is the regularization coefficient of the j+1th feature; α ptr is the regularization coefficient learning rate; is the gradient of the loss function of the feedforward network with respect to the jth feature, which indicates the contribution of this feature to the total loss. Preferably, α ptrSet to 0.2.

[0068] Step S16: Repeat the above steps S12-S15 until the preset stop iteration condition is met, indicating that the feedforward network training is completed. In one embodiment, the preset stop iteration condition is reaching a preset maximum number of iterations, preferably, the preset maximum number of iterations is set to 1000 times.

[0069] The present invention also provides a pre-trained language model, which is used as the master of a large language model to process the natural language understanding and generation of user input. Figure 2 As shown, the pre-trained language model is a model based on a Transformer network, including an encoder and a decoder; wherein the encoder is used to map input features into a series of context-aware vector representations, and the decoder is used to gradually decode the vector representations into target features.

[0070] The encoder consists of N identical encoding layers, each of which includes two sublayers: a multi-head attention mechanism sublayer and a feedforward network sublayer, and the encoder adds residual connections and normalization operations after each sublayer.

[0071] The decoder consists of N identical decoding layers, each of which consists of two multi-head attention mechanism sublayers and one feedforward network sublayer, and the decoder adds residual connections and normalization operations after each sublayer.

[0072] It should be noted that the multi-head attention mechanism sublayer is implemented by a multi-head self-attention mechanism, which is the core module of the encoder and is responsible for learning the importance of each position from the input features and combining the information of all positions to obtain a global context representation, thereby improving the performance and generalization ability of the model. Specifically, the multi-head self-attention mechanism is implemented in that the input features pass through three different linear transformation layers to generate "query", "key" and "value" vectors respectively, and the linear transformation is implemented by multiplying the weight matrix. Further, the generated query, key and value vectors are divided into multiple smaller parts, namely "heads", so that the model processes multiple different representation subspaces in parallel, thereby capturing the diversity information in the input data. Further, for each head, the dot product of the query and key vectors is calculated respectively to obtain an attention score, which reflects the relative importance of different positions in the sequence of input data. In order to make the numerical stability of the attention score, the attention score is divided by a scaling factor, which can be the square root of the dimension of the key vector.

[0073] Furthermore, in both the encoder and decoder, the output of each sub-layer is the result of layer normalization and residual connection, which is expressed as:

[0074] ycgy =LN c (x trans ++Sublayer(x trans )),

[0075] In the formula, LN c ( ) is the layer normalization function, Sublayer( ) represents the output of the residual connected sublayer, x trans Represents the input of the sublayer; y cgy is the feature output after layer normalization.

[0076] The feedforward network sublayer consists of two fully connected layers and an activation function, which is used to perform nonlinear transformation on the vector at each position. In this sublayer, the input vector first passes through the first fully connected layer for linear transformation, then passes through a ReLU activation function for nonlinear transformation, and finally passes through the second fully connected layer for linear transformation to improve the expressiveness and generalization ability of the model, so that the model can better handle complex input data.

[0077] It should be noted that the feedforward networks of the feedforward network sublayer are all trained using the feedforward network dynamic learning adjustment training method based on the swarm awareness mechanism as described in any of the above implementations.

[0078] In summary, the feedforward network dynamic learning adjustment training method and pre-trained language model based on swarm awareness mechanism provided by the present invention have the following beneficial effects:

[0079] By simulating the collective behavior of a colony of bees in the process of searching for nectar, the search for the optimal solution is achieved by collecting and integrating local and global information. In the parameter update process of the feedforward network, a colony awareness layer is used in the feedforward network to optimize and enhance global information, allowing the model to dynamically learn and adjust during the training process, and by imitating the group behavior of bees, the distribution of the weights of the feedforward network is continuously optimized, and different data features can be continuously adjusted and adapted during the training process. In a preferred embodiment, the regularization weights are updated by combining dynamic feedback of feature learning, the model's recognition of the importance of different features is enhanced, and the adjustment coefficient and balance parameter of the entropy term are dynamically adjusted to ensure the diversity and authenticity of the generated samples.

[0080] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0081] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0082] Those of ordinary skill in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0083] In the embodiments provided by the present invention, it should be understood that the disclosed systems or devices / terminal equipment and methods can be implemented in other ways. For example, the system or device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the system or unit can be electrical, mechanical or other forms.

[0084] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0085] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0086] The present invention is not limited to what is described in the specification and implementation modes, and therefore additional advantages and modifications can be easily realized by those skilled in the art. Therefore, without departing from the spirit and scope of the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details, representative devices, and illustrative examples shown and described herein.

Claims

1. A feedforward network dynamic learning adjustment training method based on swarm awareness mechanism, characterized in that: include: Step S11: Initialize the parameters of the feedforward network structure and the bee colony algorithm; Step S12: In each training iteration, the parameters of the feedforward network are optimized using the bee swarm algorithm. The swarm awareness mechanism in the bee swarm algorithm will be activated, and the feedforward network will optimize the feedforward network weights through each individual bee in the bee swarm algorithm according to the output characteristics of the current layer; Step S13: enabling individual bees to cooperate with global information based on local search to update their positions; Step S14: In each iteration, the learning rate of the feedforward network is dynamically adjusted according to the learning progress of the current feedforward network. If the current optimization progress is lower than the preset condition, the learning rate will be increased; On the contrary, if the feedforward network is close to the optimal solution, the learning rate will be reduced to avoid overfitting; Step S15: After each training, the loss value between the output of the feedforward network and the true label is calculated. The loss function of the feedforward network uses the mean square error to measure the difference between the model output and the true label, and the weight of the regularization term is dynamically updated with the self-supervisory feedback during the training process to ensure that the influence of irrelevant or redundant features is gradually weakened; Step S16: Repeat the above steps S12-S15 until the preset stop iteration condition is met, which means that the feedforward network training is completed.

2. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism according to claim 1, characterized in that: In step S11, the data input to the feedforward network is assumed to be X p , the output of the feedforward network is Y p , the weight parameter of the feedforward network is W p , the bias term of the feedforward network is b p , according to the standard structure of the feedforward network, the way to calculate the activation value of each layer is expressed as: In the formula, is the weighted input of layer I of the feedforward network; is the weight of layer I of the feedforward network; is the output of the I-1th layer of the feedforward network, which is also the input of the Ith layer of the feedforward network; is the bias term of the Ith layer of the feedforward network.

3. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism as claimed in claim 2, characterized in that: The activation function is used to perform nonlinear mapping on the weighted input of each layer of the feedforward network to obtain the output of the layer, which is expressed as: In the formula, is the output activation value of the Ith layer of the feedforward network; Sig() is the Sigmoid activation function.

4. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism as claimed in claim 2, characterized in that: In step S12, the weight parameters and bias parameters of the feedforward network corresponding to the state of the individual bee are set, and the fitness function is defined as the difference between the output of the feedforward network and the actual label. The specific calculation method is expressed as: In the formula, F p () is the fitness function; is the weight configuration of the i-th bee individual; is the bias configuration of the i-th bee individual; is the model output of the jth sample, which is calculated by the output of the feedforward network through the preset Softmax function; is the true label of the jth sample; Ns is the number of samples input in the current batch; α p It is the coefficient of the adjustment term to control the influence of different samples on the fitness function.

5. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism according to claim 4, characterized in that: In step S13, the updated position of the individual bee is expressed as: In the formula, is the weight configuration of the i-th bee individual; α pa is the first hyperparameter controlling the step size; β pa is the second hyperparameter controlling the step size; γ pa is the third hyperparameter that controls the step size; is the weight configuration of the g-th bee individual; is the weight configuration of the rth bee individual; is the current loss of the i-th bee individual; is the global optimal loss.

6. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism according to claim 5, characterized in that: In step S14, the learning rate is updated as follows: In the formula, is the learning rate of the feedforward network at the tth iteration; is the learning rate of the feedforward network at the t+1th iteration; γ pec Hyperparameters adjusted to control learning rate; is the change in the loss function of the feedforward network; is the loss value of the tth iteration; δ pec is the tuning hyperparameter for the volatility factor; is the variance of the loss value.

7. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism according to claim 6, characterized in that: In step S15, the calculation method of dynamic update is expressed as: Where, L p is the loss function of the feedforward network; Ms is the total number of features; is the regularization coefficient of the jth feature; is the value of the jth feature.

8. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism according to claim 7, characterized in that: In step S15, the weight of the regularization term is automatically adjusted according to the learning process of each feature, and the dynamic update rule enables the regularization term to reflect the learning difficulty and importance of each feature; wherein, the calculation method of the regularization coefficient is expressed as: In the formula, is the regularization coefficient of the j+1th feature; α ptr is the regularization coefficient learning rate; is the gradient of the loss function of the feedforward network with respect to the jth feature, indicating the contribution of this feature to the total loss.

9. The feedforward network dynamic learning adjustment training method based on swarm awareness mechanism as claimed in claim 8, characterized in that: α p Set to 0.3, γ pec Set to 0.2, δ pec Set to 0.3, α ptr Set to 0.

2.

10. A pre-trained language model used as the master of a large language model to process natural language understanding and generation of user input, characterized in that: The pre-trained language model is a model based on a Transformer network, including an encoder and a decoder; wherein the encoder is used to map input features into a series of context-aware vector representations, and the decoder is used to gradually decode the vector representations into target features; The encoder includes N identical encoding layers, each encoding layer includes two sublayers, a multi-head attention mechanism sublayer and a feedforward network sublayer, and the encoder adds a residual connection and a normalization operation after each sublayer; The decoder includes N identical decoding layers, each decoding layer includes two multi-head attention mechanism sublayers and one feedforward network sublayer, and the decoder adds residual connections and normalization operations after each sublayer; The feedforward networks of the feedforward network sublayer are all trained using the feedforward network dynamic learning adjustment training method based on the swarm awareness mechanism as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Adaptive model optimization method and system based on closed-loop dynamic feedback

    CN121787492A