Decision model construction method and device based on dynamic weight neural network, equipment and storage medium

By introducing dynamic weighting and analytical solution update mechanisms into neural networks, the problem of lack of interpretability in the decision-making process of traditional neural network models is solved, and higher decision transparency and adaptability are achieved.

CN119940429APending Publication Date: 2025-05-06SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411782481.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The lack of interpretability in the decision-making process of traditional neural network models limits their in-depth application in the fields of finance and medical care, especially in tasks that require accurate prediction and explanation of model behavior.

Method used

The decision model construction method based on dynamic weighted neural network is adopted, and the neural network structure is designed by collecting and preprocessing data, and the weights of the output layer are updated using analytical solutions during training to improve the interpretability and adaptability of the model.

Benefits of technology

While maintaining prediction accuracy, this method improves the interpretability of the model, enhances the transparency and reliability of decisions, and is suitable for dealing with scenarios where non-independent homogeneous distribution tasks and new categories appear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940429A_ABST
    Figure CN119940429A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of neural network model construction, and particularly relates to a decision model construction method and device based on a dynamic weight neural network, equipment and a storage medium, and the method comprises the steps: collecting original data, and converting the original data into a format suitable for neural network processing; dividing the processed data into a training set and a verification set; designing a neural network structure; setting hyper-parameters and an optimizer; constructing a neural network working in a feedforward mode, ensuring that the neural network has a continuous learning framework of an analytical solution, training a model by using a training set, and adjusting model parameters by using verification set data; in the training process or after training is completed, the analytical solution of the neural network is utilized to update the weight of the output layer. And the output weight is updated by using the analytic solution, and the method does not need a gradient descent algorithm, so that the learning efficiency and speed are improved. The use of analytic solutions reduces the demand for computing resources, so that the model can quickly adapt to new tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the field of artificial intelligence, decision support systems (DSS) and neural networks (NN) are two crucial technologies. Decision support systems are mainly used to help people make better decisions, while neural networks are a computing model that mimics the structure of human brain neural networks and has learning and adaptive capabilities. With the widespread application of neural networks in various fields, more and more decision support systems have begun to adopt neural network technology to improve decision quality and efficiency.

[0003] However, traditional neural network models often lack interpretability in the decision-making process, which limits their in-depth application in finance, medicine and other fields. These fields require not only accurate predictions, but also the ability to explain model behavior to ensure the transparency and reliability of decision-making. Therefore, how to solve the catastrophic forgetting problem in continuous learning while maintaining the accuracy of neural network predictions, and improve the adaptability and decision-making quality of the model in non-IID tasks is the problem to be solved in this application. Summary of the invention

[0004] In view of the catastrophic forgetting problem existing in the traditional decision model in continuous learning, the present invention provides a decision model construction method, device, equipment and storage medium based on dynamic weight neural network.

[0005] In a first aspect, the technical solution of the present invention provides a decision model construction method based on a dynamic weight neural network, comprising: Collect raw data and convert it into a format suitable for neural network processing; divide the processed data into training set and validation set; Design neural network architecture; Set hyperparameters and optimizer; Build a continuous learning framework that allows neural networks to work in a feed-forward manner and ensure that they have analytical solutions; Use the training set to train the model and use the validation set data to adjust the model parameters; During or after training, the analytical solution of the neural network is used to update the weights of the output layer.

[0006] As a further limitation of the technical solution of the present invention, the steps of designing the neural network structure include: Select the neural network type; Design the network structure, including the number of layers, the number of neurons in each layer, and the activation function; Design a memory function weight matrix to realize the memory function of the input data.

[0007] As a further limitation of the technical solution of the present invention, the step of designing the memory function weight matrix includes: The importance values ​​of the parameters are calculated, and a matrix with the same dimension as the model parameters is constructed as the memory function weight matrix; each element in the weight matrix represents the importance value of the corresponding parameter.

[0008] As a further limitation of the technical solution of the present invention, the step of calculating the importance value of the parameter and constructing a matrix of the same dimension as the model parameter as the memory function weight matrix includes: Create an empty dictionary fisher_matrix to store Fisher information of model parameters; Iterate over all parameters of the model and for each parameter create a tensor in the fisher_matrix dictionary with the same shape as the parameter but with all elements zero; Input data and target labels are fed into the model, the output is calculated by forward propagation through the model, and the loss is calculated using the cross entropy loss function; Calculate gradients and accumulate Fisher information; Traverse the fisher_matrix dictionary and normalize the Fisher information of each parameter; Convert the Fisher information matrix to a memory weight matrix using the square root of the Fisher information as the weight.

[0009] As a further limitation of the technical solution of the present invention, the step of setting the hyperparameters and the optimizer includes: Set hyperparameters; Select a loss function, which is used to evaluate the difference between the model's predictions and the true values, and an optimizer, which is used to update the model's weights.

[0010] As a further limitation of the technical solution of the present invention, the steps of constructing a neural network working in a feedforward manner and ensuring that it has a continuous learning framework with an analytical solution include: Build a feature extractor to extract useful features from the input data; Building memory modules that store and recall knowledge from old tasks; Construct a learning mechanism responsible for learning on new tasks while retaining knowledge from old tasks.

[0011] As a further limitation of the technical solution of the present invention, during or after the training process, the step of updating the weights of the output layer using the analytical solution of the neural network includes: The new weights are calculated using the analytical solution formula for regularized linear regression; Add the calculated new weight to the old weight to get the combined weight; Analytical solution for weights ; Where X is the design matrix or feature matrix, where each row represents a sample and each column represents a feature; y is the target vector, where each element corresponds to the label or output value of each row in X.

[0012] In a second aspect, the technical solution of the present invention also provides a decision model construction device based on a dynamic weight neural network, including a data preprocessing module, a neural network design module, a parameter setting module, a continuous learning framework construction module and an analytical solution weight update module; The data preprocessing module is used to collect raw data and convert the raw data into a format suitable for neural network processing; the processed data is divided into a training set and a validation set; Neural network design module, used to design neural network structure; Parameter setting module, used to set hyperparameters and optimizers; A continuous learning framework building module, which is used to build a neural network that works in a feedforward manner and ensures that it has a continuous learning framework with an analytical solution; The analytical solution weight update module is used to train the model using the training set and adjust the model parameters using the validation set data. During or after the training process, the analytical solution of the neural network is used to update the weights of the output layer.

[0013] As a further limitation of the technical solution of the present invention, the neural network design module includes a network structure design unit and a memory function weight matrix design unit; The network structure design unit is used to select the type of neural network; design the network structure, including the number of layers, the number of neurons in each layer, and the activation function; The memory function weight matrix design unit is used to design the memory function weight matrix to realize the memory function of the input data.

[0014] As a further limitation of the technical solution of the present invention, the memory function weight matrix design unit is used to calculate the importance value of the parameter and construct a matrix of the same dimension as the model parameter as the memory function weight matrix; each element in the weight matrix represents the importance value of the corresponding parameter.

[0015] As a further limitation of the technical solution of the present invention, the memory function weight matrix design unit is specifically used to create an empty dictionary fisher_matrix to store the Fisher information of the model parameters; traverse all the parameters of the model and create a tensor with the same shape as the parameter but all elements are zero in the fisher_matrix dictionary for each parameter; input the input data and the target label into the model, calculate the output through the forward propagation of the model, and calculate the loss using the cross entropy loss function; calculate the gradient and accumulate the Fisher information; traverse the fisher_matrix dictionary and normalize the Fisher information of each parameter; use the square root of the Fisher information as the weight to convert the Fisher information matrix into a memory weight matrix.

[0016] As a further limitation of the technical solution of the present invention, the parameter setting module is specifically used to set hyperparameters; select a loss function and an optimizer, the loss function is used to evaluate the difference between the prediction of the model and the true value, and the optimizer is used to update the weight of the model.

[0017] As a further limitation of the technical solution of the present invention, the continuous learning framework construction module is specifically used to construct a feature extractor for extracting useful features from input data; to construct a memory module for storing and recalling knowledge of old tasks; and to construct a learning mechanism responsible for learning on new tasks while retaining knowledge of old tasks.

[0018] As a further limitation of the technical solution of the present invention, the analytical solution weight update module is specifically used to calculate the new weight using the analytical solution formula of regularized linear regression; add the calculated new weight to the old weight to obtain the combined weight; Analytical solution for weights ; Where X is the design matrix or feature matrix, where each row represents a sample and each column represents a feature; y is the target vector, where each element corresponds to the label or output value of each row in X.

[0019] In a third aspect, the technical solution of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the decision model construction method based on the dynamic weight neural network as described in the first aspect.

[0020] In a fourth aspect, the technical solution of the present invention also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable the computer to execute the decision model construction method based on the dynamic weight neural network as described in the first aspect.

[0021] It can be seen from the above technical solutions that the present invention has the following advantages: by collecting raw data and converting it into a format suitable for neural network processing, a reliable data basis is provided for model training. At the same time, dividing the processed data into a training set and a validation set helps to evaluate the performance of the model during the training process and avoid overfitting. Designing a reasonable neural network structure, including input layer, hidden layer and output layer, as well as the number of nodes and activation function of each layer, helps to improve the prediction ability of the model. By setting reasonable hyperparameters (such as learning rate, batch size, etc.) and optimizers (such as SGD, Adam, etc.), the training process of the model can be accelerated, and the convergence speed and prediction accuracy of the model can be improved.

[0022] Construct a neural network that works in a feedforward manner and ensures that it has a continuous learning framework with analytical solutions. This allows the model to use analytical solutions to update the weights of the output layer during or after training, thereby improving the interpretability of the model. Through dynamic weight adjustment and the application of analytical solutions, the model improves interpretability while maintaining prediction accuracy. This allows the model to give clearer and more reasonable explanations in the decision-making process, enhancing the transparency and reliability of decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0025] Figure 2 is a schematic block diagram of an apparatus according to an embodiment of the present invention.

[0026] Figure 3 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0028] like Figure 1 As shown, an embodiment of the present invention provides a decision model construction method based on a dynamic weight neural network, comprising: Step 1: Collect raw data and convert it into a format suitable for neural network processing; divide the processed data into training set and validation set; This step is responsible for cleaning, standardizing and feature extraction of input data to meet the needs of subsequent model training. It can process data from various sources and types, including structured data and unstructured data, to provide clean and consistent input for the model. This step is key to ensuring that the model can receive high-quality data, which directly affects the effect of model training and the final decision quality; Step 2: Design the neural network structure; Step 3: Set hyperparameters and optimizer; The dynamic weight neural network randomly initializes the input weights and biases, and only needs to solve the output weights analytically, avoiding the gradient vanishing problem during backpropagation. It allows the model to quickly adapt to new tasks while retaining the memory of old tasks. This design reduces the demand for computing resources, speeds up model training, and improves the model's adaptability to new tasks. The update of the output weights is guided by constructing a weight importance matrix. This design enables the model to adaptively adjust network parameters when learning new tasks, reducing the forgetting of old task knowledge. This memory function is critical to maintaining the stability and effectiveness of the model during continuous learning. Step 4: Build a neural network that works in a feedforward manner and ensure that it has a continuous learning framework with an analytical solution. Train the model using the training set and adjust the model parameters using the validation set data. The framework supports incremental learning of the model when facing new tasks, without the need to retrain the entire model. The framework effectively accommodates new categories that appear in new tasks through analytical solutions, achieving the integration of learning and memory. The design of this framework allows the model to maintain continuity and consistency of decision-making in a constantly changing environment, which is particularly important for business scenarios that need to process dynamic data and quickly adapt to new situations.

[0029] Step 5: During or after training, use the analytical solution of the neural network to update the weights of the output layer.

[0030] The importance of model parameters is measured according to the connection weights between nodes of previous tasks, so that the weights can be selectively updated when learning new tasks to reduce the impact on the performance of old tasks. Dynamic weight adjustment mechanism and randomized network structure. This structure can not only achieve rapid adaptation to new tasks, but also achieve learning and memory fusion of new tasks without sacrificing the performance of old tasks. The application of this technology enables the model to perform well in continuous learning scenarios, especially when dealing with non-independent and identically distributed tasks and the emergence of new categories.

[0031] The present invention aims to solve the problem of catastrophic forgetting in continuous learning and improve the adaptability and decision-making quality of the model in non-independent and identically distributed tasks. The method is particularly suitable for decision-making scenarios that need to deal with constantly changing tasks and new categories, such as financial risk assessment, medical diagnosis, resource optimization and personalized recommendation systems. In the financial field, the present invention can help financial institutions adjust investment strategies and risk management models in real time, dynamically adjust weights to adapt to market changes, and improve the accuracy of risk prediction and the efficiency of investment decisions. In the medical field, the method can assist doctors in making more accurate diagnosis and treatment decisions in the ever-changing medical data and treatment plans. In terms of resource optimization, the present invention can dynamically adjust resource allocation strategies according to real-time data to improve resource utilization efficiency. In a personalized recommendation system, the present invention can dynamically adjust the recommendation algorithm according to changes in user behavior to provide recommendation results that better meet user needs.

[0032] In some embodiments, the step of designing a neural network structure includes: Select the neural network type; Design the network structure, including the number of layers, the number of neurons in each layer, and the activation function; Design a memory function weight matrix to realize the memory function of the input data.

[0033] The steps to design the memory function weight matrix include: The importance values ​​of the parameters are calculated, and a matrix with the same dimension as the model parameters is constructed as the memory function weight matrix; each element in the weight matrix represents the importance value of the corresponding parameter.

[0034] In some embodiments, the step of calculating the importance values ​​of the parameters and constructing a matrix of the same dimension as the model parameters as the memory function weight matrix includes: Create an empty dictionary fisher_matrix to store Fisher information of model parameters; Iterate over all parameters of the model and for each parameter create a tensor in the fisher_matrix dictionary with the same shape as the parameter but with all elements zero; Input data and target labels are fed into the model, the output is calculated by forward propagation through the model, and the loss is calculated using the cross entropy loss function; Calculate gradients and accumulate Fisher information; Traverse the fisher_matrix dictionary and normalize the Fisher information of each parameter; Convert the Fisher information matrix to a memory weight matrix using the square root of the Fisher information as the weight.

[0035] In some embodiments, the steps of setting hyperparameters and optimizers include: Set hyperparameters; Select a loss function, which is used to evaluate the difference between the model's predictions and the true values, and an optimizer, which is used to update the model's weights.

[0036] In some embodiments, the first step is to build a general continuous learning framework that works in a feedforward manner and has an analytical solution that is effectively compatible with new categories that appear in new tasks. The design of this framework takes into account the non-independent and identically distributed characteristics of data in the real world, as well as the uncertainty and dynamics of task sequences. Through this framework, we can effectively absorb new knowledge while retaining the memory of old tasks. The construction ideas are as follows: 1) Build a feature extractor: Extract useful features from the input data, which will be used for subsequent tasks.

[0037] Build a feature extractor to extract useful features from the input data; 2) Build memory modules: used to store and recall knowledge from old tasks to reduce forgetting. This can be achieved in many ways, such as using parameterized memory units of neural networks or external storage systems.

[0038] The specific code is as follows: class MemoryModule(nn.Module): def __init__(self): super(MemoryModule, self).__init__() self.memory_bank = [] self.memory_strength = 0.5# Hyperparameter controlling memory strength def store Memories(self, knowledge): #Store knowledge or features self.memory_bank.append(knowledge) def retrieve_memories(self, new_knowledge): # Retrieve relevant memory based on new knowledge relevant_memories = [mem for mem in self.memory_bank if self.is_relevant(mem, new_knowledge)] return relevant_memories def is_relevant(self, mem, new_knowledge): # Define correlation judgment logic, for example, by cosine similarity similarity = nn.functional.cosine_similarity(mem, new_knowledge) return similarity>self.memory_strength 3) Building a learning mechanism: Responsible for learning on new tasks while retaining knowledge of old tasks. This can be achieved through regularization techniques, experience replay, or dynamically adjusting the learning rate.

[0039] Second, building memory modules that store and recall knowledge from old tasks; Based on the synaptic plasticity principle of the biological brain, a weight matrix with memory function is designed. This matrix can adaptively adjust network parameters to avoid forgetting when learning new tasks, thereby protecting the knowledge of old tasks. This design allows the model to dynamically adjust weights during continuous learning, reducing the impact on the performance of old tasks while effectively integrating the knowledge of new tasks.

[0040] Calculate the importance of parameters: Construct a matrix of the same dimension as the model parameters as the initial weight matrix, in which each element represents the importance of the corresponding parameter, and can measure the curvature of the loss function relative to the model parameters to calculate the importance of the parameters.

[0041] The memory function weight matrix constructed using the calculated parameter importance can be directly used as the Fisher information matrix or further processed based on the Fisher information matrix. In the learning of new tasks, the memory function weight matrix is ​​used to adjust the amplitude of parameter updates and protect the parameters that are important to the old tasks.

[0042] The code is as follows: def build_memory_weight_matrix(fisher_matrix): memory_weight_matrix = {} # Convert Fisher information matrix to memory weight matrix for name, fisher in fisher_matrix.items(): # For example, use the square root of the Fisher information matrix as the weight memory_weight_matrix[name] = fisher.sqrt() return memory_weight_matrix def update_parameters_with_memory(model, memory_weight_matrix,learning_rate): with torch.no_grad(): for name, param in model.named_parameters(): if name in memory_weight_matrix: # Adjust the learning rate according to the memory weight matrix adjusted_lr = learning_rate / memory_weight_matrix[name].median() param -= adjusted_lr * param.grad In some embodiments, during or after training, the step of updating the weights of the output layer using the analytical solution of the neural network includes: By updating the output weights through analytical solutions, this method does not require a gradient descent algorithm, improves the efficiency and speed of learning, and reduces the demand for computing resources. This innovative application not only improves the parameter efficiency of the model, but also speeds up the convergence of the model, allowing the model to quickly adapt to new tasks. For linear regression problems, the analytical solution can be obtained through the least squares method, where W is the model parameter, X is the feature matrix, and y is the target value vector:

[0043] In continuous learning, analytical solutions can be used to update model parameters, especially when dealing with new tasks, to avoid catastrophic forgetting of knowledge of old tasks. Suppose we have a simple linear model and we want to update the weights in the new task while retaining the knowledge of the old task. In some cases, the model can be optimized directly by analytical methods, and the optimal regularization parameters can be calculated by analytical methods.

[0044] Where X is the design matrix or feature matrix, where each row represents a sample and each column represents a feature; y is the target vector, where each element corresponds to the label or output value of each row in X.

[0045] The code is as follows: def update_weights_analytically(old_weights, new_data, new_labels,lambda_reg): # Compute the analytical solution for the new task new_weights = np.linalg.inv(new_data.T.dot(new_data) + lambda_reg *np.eye(new_data.shape[1])).dot(new_data.T).dot(new_labels) # Merge old weights with new weights combined_weights = old_weights + new_weights return combined_weights # Analytical method to calculate regularization parameters def optimal_regularization(X, y, lambda Candidates): optimal_lambda = None best_performance = float('inf') for lambda_reg in lambdaCandidates: weights = closed_form_solution(X, y, lambda_reg) performance = np.mean((y - X.dot(weights)) ** 2)# mean square error if performance <best_performance: best_performance = performance optimal_lambda = lambda_reg return optimal_lambda.

[0046] By referring to the principle of prominent plasticity in neuroscience, the network parameters are dynamically adjusted to reduce the forgetting of old task knowledge. The weight importance matrix (weight matrix) is used to quantify the importance of each parameter, and the amplitude of parameter update is adjusted accordingly. By randomly initializing weights and biases, the complex back-propagation process is avoided, simplifying model training. The model can quickly adapt to new tasks while maintaining stability for old tasks. A general continuous learning framework is constructed, which can effectively accommodate new categories that appear in new tasks. The framework provides an analytical solution that enables the model to quickly adjust when facing new tasks without retraining. The analytical solution is used to update the output weights. This method does not require a gradient descent algorithm and improves the efficiency and speed of learning. The use of analytical solutions reduces the demand for computing resources, allowing the model to quickly adapt to new tasks. The robustness of the task sequence is achieved through the combined action of regularization technology and weight adjustment mechanism, so that the model can maintain stable performance when facing tasks of different orders.

[0047] The architecture diagram is as follows Figure 3 As shown, the present invention brings significant beneficial effects in the field of decision support through its innovative integration of dynamic weight neural network and synaptic plasticity ideas. These effects not only improve the efficiency and accuracy of decision-making, but also optimize resource allocation and enhance market adaptability. Through the mechanism of dynamic weight adjustment and re-plasticity inspiration, the catastrophic forgetting problem in continuous learning is effectively avoided, ensuring that the model can maintain memory of old tasks when facing new tasks. This memory fusion capability enables the model to perform well when dealing with non-independent and identically distributed tasks and the emergence of new categories. The model constructed by this method has the characteristics of efficient parameters and fast convergence, which reduces the demand for computing resources and reduces the time cost of model training. This enables enterprises to respond to market changes in a short period of time and quickly adjust decision-making strategies. In addition, the model does not need the ability to store old task data additionally, which reduces storage requirements and reduces the privacy risks caused by storing large amounts of data.

[0048] In practical applications, the invention enables the model to maintain stable performance in different task sequences through its powerful generalization ability, which is particularly important for decision support systems in open environments. These characteristics make it have broad application prospects in the field of auxiliary decision making, especially in those environments that need to quickly adapt to new situations and process dynamic data. Through the application of the present invention, enterprises can achieve smarter and more stable intelligent decision support, improve decision quality, reduce operational risks, and enhance market competitiveness.

[0049] like Figure 2 As shown, an embodiment of the present invention also provides a decision model construction device based on a dynamic weight neural network, including a data preprocessing module, a neural network design module, a parameter setting module, a continuous learning framework construction module and an analytical solution weight update module; The data preprocessing module is used to collect raw data and convert the raw data into a format suitable for neural network processing; the processed data is divided into a training set and a validation set; Neural network design module, used to design neural network structure; Parameter setting module, used to set hyperparameters and optimizers; A continuous learning framework building module, which is used to build a neural network that works in a feedforward manner and ensures that it has a continuous learning framework with an analytical solution; The analytical solution weight update module is used to train the model using the training set and adjust the model parameters using the validation set data. During or after the training process, the analytical solution of the neural network is used to update the weights of the output layer.

[0050] In some embodiments, the neural network design module includes a network structure design unit and a memory function weight matrix design unit; The network structure design unit is used to select the type of neural network; design the network structure, including the number of layers, the number of neurons in each layer, and the activation function; The memory function weight matrix design unit is used to design the memory function weight matrix to realize the memory function of the input data.

[0051] In some embodiments, the memory function weight matrix design unit is used to calculate the importance value of the parameter and construct a matrix with the same dimension as the model parameter as the memory function weight matrix; each element in the weight matrix represents the importance value of the corresponding parameter.

[0052] In some embodiments, the memory function weight matrix design unit is specifically used to create an empty dictionary fisher_matrix to store the Fisher information of the model parameters; traverse all the parameters of the model and create a tensor in the fisher_matrix dictionary for each parameter with the same shape as the parameter but all elements are zero; input the input data and the target label into the model, calculate the output through the forward propagation of the model, and calculate the loss using the cross entropy loss function; calculate the gradient and accumulate the Fisher information; traverse the fisher_matrix dictionary and normalize the Fisher information of each parameter; use the square root of the Fisher information as the weight to convert the Fisher information matrix into a memory weight matrix.

[0053] In some embodiments, the parameter setting module is specifically used to set hyperparameters; select a loss function and an optimizer, wherein the loss function is used to evaluate the difference between the prediction of the model and the true value, and the optimizer is used to update the weight of the model.

[0054] In some embodiments, the continuous learning framework builds modules, specifically for building a feature extractor for extracting useful features from input data; building a memory module for storing and recalling knowledge of old tasks; and building a learning mechanism responsible for learning on new tasks while retaining knowledge of old tasks.

[0055] In some embodiments, the analytical solution weight update module is specifically used to calculate the new weight using the analytical solution formula of regularized linear regression; add the calculated new weight to the old weight to obtain a combined weight; Analytical solution for weights ; Where X is the design matrix or feature matrix, where each row represents a sample and each column represents a feature; y is the target vector, where each element corresponds to the label or output value of each row in X.

[0056] An embodiment of the present invention also provides an electronic device, which includes: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. The communication bus can be used for information transmission between the electronic device and the sensor. The processor can call the logic instructions in the memory to execute the following method: Step 1: Collect raw data and convert the raw data into a format suitable for neural network processing; divide the processed data into a training set and a validation set; Step 2: Design a neural network structure; Step 3: Set hyperparameters and optimizers; Step 4: Build a neural network that works in a feedforward manner and ensure that it has a continuous learning framework with an analytical solution, use the training set to train the model, and use the validation set data to adjust the model parameters; Step 5: During or after the training, use the analytical solution of the neural network to update the weights of the output layer.

[0057] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0058] An embodiment of the present invention provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable a computer to execute the method provided by the above method embodiment, for example, including: Step 1: Collect raw data and convert the raw data into a format suitable for neural network processing; divide the processed data into a training set and a validation set; Step 2: Design a neural network structure; Step 3: Set hyperparameters and optimizer; Step 4: Build a neural network that works in a feedforward manner and ensure that it has a continuous learning framework with an analytical solution, use the training set to train the model, and use the validation set data to adjust the model parameters; Step 5: During or after the training, use the analytical solution of the neural network to update the weights of the output layer.

[0059] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0060] An embodiment of a decision model construction device based on a dynamic weight neural network provided in an embodiment of the present invention belongs to the same inventive concept as the decision model construction method based on a dynamic weight neural network in the above-mentioned embodiments. For details not described in detail in the embodiment of the decision model construction device based on a dynamic weight neural network, reference can be made to the embodiment of the decision model construction method based on a dynamic weight neural network mentioned above.

[0061] The decision model construction device based on the dynamic weight neural network is a unit and algorithm step of each example described in combination with the embodiments disclosed herein, which can be implemented by electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0062] Those skilled in the art will appreciate that various aspects of the decision model construction method based on a dynamic weight neural network can be implemented as a system, method or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which can be collectively referred to as "circuit", "module" or "system".

[0063] Although the present invention has been described in detail with reference to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person of ordinary skill in the art may easily think of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention.

Claims

1. A decision model construction method based on dynamic weight neural network, characterized in that: include: Collecting raw data and converting it into a format suitable for neural network processing; Divide the processed data into training set and validation set; Design neural network architecture; Set hyperparameters and optimizer; Build a neural network that works in a feedforward manner and ensure that it has a continuous learning framework with an analytical solution, train the model using the training set, and adjust the model parameters using the validation set data; During or after training, the analytical solution of the neural network is used to update the weights of the output layer.

2. The decision model construction method based on dynamic weight neural network according to claim 1 is characterized in that: The steps to design a neural network structure include: Select the neural network type; Design the network structure, including the number of layers, the number of neurons in each layer, and the activation function; Design a memory function weight matrix to realize the memory function of the input data.

3. The decision model construction method based on dynamic weight neural network according to claim 2 is characterized in that: The steps to design the memory function weight matrix include: The importance values ​​of the parameters are calculated, and a matrix with the same dimension as the model parameters is constructed as the memory function weight matrix; each element in the weight matrix represents the importance value of the corresponding parameter.

4. The decision model construction method based on dynamic weight neural network according to claim 3 is characterized in that: The steps of calculating the importance values ​​of the parameters and constructing a matrix of the same dimension as the model parameters as the memory function weight matrix include: Create an empty dictionary fisher_matrix to store Fisher information of model parameters; Iterate over all parameters of the model and for each parameter create a tensor in the fisher_matrix dictionary with the same shape as the parameter but with all elements zero; Input data and target labels are fed into the model, the output is calculated by forward propagation through the model, and the loss is calculated using the cross entropy loss function; Calculate gradients and accumulate Fisher information; Traverse the fisher_matrix dictionary and normalize the Fisher information of each parameter; Convert the Fisher information matrix to a memory weight matrix using the square root of the Fisher information as the weight.

5. The decision model construction method based on dynamic weight neural network according to claim 4 is characterized in that: The steps to set hyperparameters and optimizer include: Set hyperparameters; Select a loss function, which is used to evaluate the difference between the model's predictions and the true values, and an optimizer, which is used to update the model's weights.

6. The decision model construction method based on dynamic weight neural network according to claim 5 is characterized in that: The steps to build a continuous learning framework that works in a feed-forward manner and ensures that it has an analytical solution include: Build a feature extractor to extract useful features from the input data; Building memory modules that store and recall knowledge from old tasks; Construct a learning mechanism responsible for learning on new tasks while retaining knowledge from old tasks.

7. The decision model construction method based on dynamic weight neural network according to claim 6 is characterized in that: During or after training, the steps of using the analytical solution of the neural network to update the weights of the output layer include: The new weights are calculated using the analytical solution formula for regularized linear regression; Add the calculated new weight to the old weight to get the combined weight; Analytical solution for weights ; Where X is the design matrix or feature matrix, where each row represents a sample and each column represents a feature; y is the target vector, where each element corresponds to the label or output value of each row in X.

8. A decision model construction device based on a dynamic weight neural network, characterized in that: It includes data preprocessing module, neural network design module, parameter setting module, continuous learning framework construction module and analytical solution weight update module; The data preprocessing module is used to collect raw data and convert the raw data into a format suitable for neural network processing; the processed data is divided into a training set and a validation set; Neural network design module, used to design neural network structure; Parameter setting module, used to set hyperparameters and optimizers; A continuous learning framework building module, which is used to build a neural network that works in a feedforward manner and ensures that it has a continuous learning framework with an analytical solution; The analytical solution weight update module is used to train the model using the training set and adjust the model parameters using the validation set data. During or after the training process, the analytical solution of the neural network is used to update the weights of the output layer.

9. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the decision model construction method based on the dynamic weight neural network as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the decision model construction method based on the dynamic weight neural network as described in any one of claims 1 to 7.