Decision method and prediction model based on brain-like intelligent decision

Through multimodal feature extraction, Transformer spatiotemporal modeling, EW-TOPSIS dynamic evaluation, dynamic memory and reinforcement learning, the adaptability and efficiency problems of traditional intelligent decision-making models in complex environments are solved, and high robustness and efficient decision-making capabilities are achieved.

CN120803265APending Publication Date: 2025-10-17BEIJING NOTHING CHECK TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510906505.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional intelligent decision-making models have poor adaptability, weak generalization ability, low computational efficiency, unreasonable weight distribution, insufficient long-term dependency modeling capabilities, insufficient utilization of historical information, many redundant operations, poor interpretability, strong sample dependence, and high energy consumption in complex dynamic environments, making it difficult to meet high real-time requirements.

Method used

Employing techniques such as multimodal feature extraction, Transformer spatiotemporal modeling, EW-TOPSIS dynamic evaluation, dynamic memory, and reinforcement learning, highly robust decision-making is achieved through multimodal feature extraction and fusion, spatiotemporal modeling, dynamic evaluation, memory storage, and reinforcement learning.

Benefits of technology

It improves the model's multi-objective decision-making ability in complex scenarios, enhances the modeling ability of spatiotemporal correlation characteristics, supports parallel computing, reduces decision latency, optimizes decision paths, and improves the model's adaptability and decision efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803265A_ABST
    Figure CN120803265A_ABST
Patent Text Reader

Abstract

The invention relates to a decision-making method and a prediction model based on brain-like intelligent decision-making, and belongs to the technical field of decision-making generation, and the method comprises the steps that a multi-modal feature extraction and fusion module carries out the feature extraction after receiving the input information of a to-be-optimized decision-making, and generates a fusion feature vector; a Transform space-time modeling module performs space-time modeling on the fusion feature vector to obtain a space-time feature matrix; an EW-TOPSIS dynamic evaluation module processes the spatial-temporal characteristic matrix to generate a decision candidate sequence; the dynamic memory module screens the decision candidate sequence to obtain a decision candidate set; the reinforcement learning module is used for updating the decision gradient of the prediction model based on the brain-like intelligent decision; and the behavior collection module is used for constraining the decision candidate set to obtain an optimal decision. According to the invention, through close cooperation among a plurality of modules, full-link closed-loop optimization from multi-modal to-be-optimized decision input to high-robustness decision is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of decision making, and particularly relates to a decision making method and a prediction model based on brain-like intelligent decision making. BACKGROUND

[0002] The intelligent enterprise decision system based on a large model in the traditional intelligent decision model can integrate market data, financial statements, customer feedback and other multi-source data, and provide real-time decision suggestions through a deep learning model. This system is not only suitable for large enterprises, but also provides an efficient decision support solution for small and medium-sized enterprises. The traditional intelligent decision model training method optimizes the task network and scheduling network through reinforcement learning, improving the accuracy of model output task parameters and training efficiency. This method is suitable for complex task decision making in the field of artificial intelligence. The traditional decision model training in the field of autonomous driving obtains environment models, high-precision maps and vehicle information through a simulation platform, and combines reinforcement learning algorithms to train decision models and evaluation models, improving the training efficiency and decision accuracy of autonomous driving technology. The traditional industrial production decision system includes an IBP consistency planning module, an intelligent master plan decision module and an order life cycle management module. This technology realizes the intelligentization and efficiency of production planning through deep learning and optimization algorithms, significantly improving the decision efficiency of industrial production. The traditional decision tree model processing method focuses on decision optimization in the field of currency financial services. This technology improves the efficiency and accuracy of the model by improving the processing method of the decision tree model. Traditional AI model processing involves model quantization methods and model performance improvement techniques. The former reduces the memory occupied by the model through block quantization matrix, and the latter improves the performance of the AI model through information interaction between nodes.

[0003] However, the traditional intelligent decision making has the following problems or shortcomings:

[0004] 1. Poor adaptability: Traditional decision models rely on heuristic rules and experience accumulation, and their linear decision paths show insufficient adaptability in complex dynamic environments, making it difficult to respond to rapidly changing environments and diverse decision-making needs.

[0005] 2. Weak generalization ability: Traditional models often struggle to effectively handle uncertainty and ambiguity when dealing with multi-dimensional, multi-objective decision-making problems, resulting in poor generalization ability and a large AUC decay rate when applied across scenarios.

[0006] 3. Low computational efficiency: Traditional models typically use serial computing methods, which are computationally inefficient and difficult to meet high real-time requirements, resulting in high decision-making delays and an inability to respond within a short period of time.

[0007] 4. Unreasonable weight distribution: Traditional models often rely on subjective experience or fixed rules in weight distribution, lack dynamic adjustment mechanism, leading to unreasonable weight distribution in multi-dimensional decision indicator evaluation, affecting decision accuracy.

[0008] 5. Insufficient long-range dependency modeling capability: Traditional models have difficulty in effectively capturing spatio-temporal correlation characteristics when processing input data with long-range dependencies, leading to information loss or misjudgment in decision-making process.

[0009] 6. Insufficient use of historical information: Traditional models lack effective dynamic memory mechanism and cannot fully utilize historical decision information, making it difficult to quickly adapt and adjust decision strategy when facing environmental changes.

[0010] 7. Many redundant operations: Traditional models lack effective constraint and optimization mechanism in decision path, prone to produce invalid or redundant operations, affecting decision efficiency and accuracy.

[0011] 8. Poor interpretability: Traditional models often lack interpretability, decision-making process is not transparent, it is difficult to understand and explain the decision result, limiting its application in some high demand fields (such as finance, medical treatment, etc.).

[0012] 9. Strong sample dependence: Traditional models usually need a large amount of sample data for training, and it is difficult to effectively learn and make decisions in small sample cases, limiting its application in data scarce scenarios.

[0013] 10. High energy consumption: Traditional models have high energy consumption in the calculation process, which is difficult to meet the demand of high energy efficiency, especially in the scene that needs to run for a long time, the energy consumption problem is more prominent. SUMMARY

[0014] In view of the above deficiencies of the prior art, the purpose of the invention is to provide a decision-making method and a prediction model based on brain-like intelligent decision-making, an electronic device and a storage medium, which realizes high robustness decision-making through multi-modal feature extraction, spatio-temporal modeling, dynamic evaluation, memory storage and reinforcement learning technologies.

[0015] The first aspect of the present invention proposes a decision-making method applied to a prediction model based on brain-like intelligent decision-making, which comprises:

[0016] The multi-modal feature extraction and fusion module in the prediction model based on brain-like intelligent decision-making receives the input information of the decision to be optimized, extracts the features and generates the fusion feature vector;

[0017] The Transformer spatio-temporal modeling module in the prediction model based on brain-like intelligent decision-making performs spatio-temporal modeling on the fusion feature vector to obtain the spatio-temporal feature matrix;

[0018] The EW-TOPSIS dynamic evaluation module in the brain-like intelligent decision-based prediction model processes the spatio-temporal feature matrix to generate a decision candidate ranking;

[0019] The dynamic memory module in the brain-like intelligent decision-based prediction model filters the decision candidate ranking to obtain a decision candidate set;

[0020] The reinforcement learning module in the brain-like intelligent decision-based prediction model is used to update the decision gradient of the brain-like intelligent decision-based prediction model;

[0021] The behavior convergence module in the brain-like intelligent decision-based prediction model is used to constrain the decision candidate set to obtain an optimal decision.

[0022] Further, in the above decision method, after the multi-modal feature extraction and fusion module in the brain-like intelligent decision-based prediction model receives input information of a decision to be optimized, it extracts features and generates a fusion feature vector, including:

[0023] Image, time series, and text feature vectors are extracted through pre-trained CNN, BiLSTM, and BERT respectively to obtain a feature vector set;

[0024] A gating fusion mechanism is used to dynamically assign weights to the feature vector set to generate a fusion feature vector.

[0025] Further, in the above decision method, the Transformer spatio-temporal modeling module in the brain-like intelligent decision-based prediction model performs spatio-temporal modeling on the fusion feature vector to obtain a spatio-temporal feature matrix, including:

[0026] Long-range dependencies in the fusion feature vector are captured through a self-attention mechanism, and position encoding and timestamp encoding are introduced to obtain a spatio-temporal feature matrix.

[0027] Further, in the above decision method, the EW-TOPSIS dynamic evaluation module in the brain-like intelligent decision-based prediction model processes the spatio-temporal feature matrix to generate a decision candidate ranking, including:

[0028] Index weights are calculated based on the entropy weight method;

[0029] The Euclidean distance between the input information of the decision to be optimized and the positive and negative ideal solutions is calculated through the TOPSIS algorithm;

[0030] The degree of fitting is calculated according to the Euclidean distance;

[0031] The decision candidate ranking is generated according to the degree of fitting.

[0032] Further, in the decision-making method, the dynamic memory module in the brain-like intelligent decision-making prediction model is used for screening the decision candidate ranking to obtain a decision candidate set, and the method comprises the following steps:

[0033] The short-term memory is used for eliminating the decision candidate ranking with a frequency lower than a preset threshold based on an LRU strategy.

[0034] The long-term memory is used for selecting the decision candidate ranking to be persisted according to a decision effect.

[0035] In the step of selecting the decision candidate ranking to be persisted, a vector index is constructed by using a Faiss library to search similar decisions.

[0036] Further, in the decision-making method, the reinforcement learning module in the brain-like intelligent decision-making prediction model is used for updating the decision gradient of the brain-like intelligent decision-making prediction model, and the method comprises the following steps:

[0037] The action distribution of the decision candidate set is output by the Actor network, and the state value of the decision candidate set is evaluated by the Critic network.

[0038] The advantage function is calculated according to the action distribution and the state value.

[0039] The decision gradient of the brain-like intelligent decision-making prediction model is updated according to the advantage function.

[0040] Further, in the decision-making method, the behavior convergence module in the brain-like intelligent decision-making prediction model is used for constraining the decision candidate set to obtain an optimal decision, and the method comprises the following steps:

[0041] The decision candidate set is constrained by using a constraint condition to obtain a constrained decision candidate set.

[0042] The optimal action in the constrained decision candidate set is selected as the optimal decision by using a greedy strategy.

[0043] The second aspect of the application further provides a brain-like intelligent decision-making prediction model, which comprises a multi-modal feature extraction and fusion module, a Transformer space-time modeling module, an EW-TOPSIS dynamic evaluation module, a dynamic memory module, a reinforcement learning module and a behavior convergence module.

[0044] The multi-modal feature extraction and fusion module is used for extracting features from input information of a decision to be optimized and generating a fusion feature vector.

[0045] The Transformer space-time modeling module is used for performing space-time modeling on the fusion feature vector to obtain a space-time feature matrix.

[0046] The EW-TOPSIS dynamic evaluation module is configured to process the spatio-temporal feature matrix to generate a decision candidate ranking;

[0047] The dynamic memory module is configured to filter the decision candidate ranking to obtain a decision candidate set;

[0048] The reinforcement learning module is configured to update a decision gradient of a prediction model based on brain-inspired intelligent decision-making;

[0049] The behavior convergence module is configured to constrain the decision candidate set to obtain an optimal decision.

[0050] The third aspect of the present application also provides an electronic device, comprising a processor and a memory;

[0051] The processor is configured to execute any of the decision-making methods described above by invoking programs or instructions stored in the memory.

[0052] The fourth aspect of the present application also provides a computer-readable storage medium storing programs or instructions, which cause a computer to execute any of the decision-making methods described above.

[0053] The present application has the following beneficial effects: the EW-TOPSIS dynamic evaluation module in the present application optimizes weight distribution by entropy weight method and combines TOPSIS algorithm to realize dynamic evaluation of multi-dimensional decision indicators, effectively handles uncertainty and fuzziness, and is suitable for multi-objective decision-making problems in complex scenarios; the Transformer architecture in the Transformer spatio-temporal modeling module uses self-attention mechanism to capture long-range dependencies in input information of the decision to be optimized, enhances the modeling capability of spatio-temporal correlation characteristics, supports parallel computing to significantly reduce decision delay, and meets high real-time requirements; the dynamic memory module stores and updates historical decision information to realize rapid adaptation to environmental changes and supports multi-module parallel operation to improve the system's ability to handle complex tasks; the behavior convergence module avoids invalid or redundant operations by constraining and optimizing the decision path, and the reinforcement learning module continuously optimizes the decision strategy through interaction with the environment, thereby improving the adaptive ability of the model. BRIEF DESCRIPTION OF DRAWINGS

[0054] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:

[0055] Figure 1 A decision-making method is provided for the embodiments of the present application.

[0056] Figure 2 A method diagram for generating a fusion feature vector is provided for an embodiment of the present application;

[0057] Figure 3 A method diagram for generating a decision candidate ranking is provided for an embodiment of the present application;

[0058] Figure 4 A method diagram for obtaining a decision candidate set is provided for an embodiment of the present application;

[0059] Figure 5 A method diagram for updating a decision gradient of a brain-inspired intelligent decision-based prediction model is provided for an embodiment of the present application;

[0060] Figure 6 A method diagram for obtaining an optimal decision is provided for an embodiment of the present application;

[0061] Figure 7 A brain-inspired intelligent decision-based prediction model schematic diagram is provided for an embodiment of the present application;

[0062] Figure 8 A schematic block diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0063] In order to enable persons skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that these descriptions are only exemplary, and are not used to limit the scope of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of protection of the present application.

[0064] In addition, in the following description, the description of well-known structures and techniques is omitted to avoid unnecessary confusion of the concepts disclosed in the present application.

[0065] In the description of the present application, the terms "first", "second", "third" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance. The terms "mounting", "connecting", "connecting" should be broadly understood, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the communication between two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0066] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, like reference numerals refer to like elements throughout the description. The following exemplary embodiments described herein represent illustrations among the various implementations consistent with the present application. That is, the following exemplary embodiments do not represent all implementations consistent with the present application, but rather they are merely examples of methods and systems consistent with some aspects of the present application as detailed in the appended claims.

[0067] The application provides a decision-making method and a prediction model based on brain-like intelligent decision-making, multi-modal feature extraction, space-time modeling, dynamic evaluation, memory storage and reinforcement learning, and the like, to achieve high-robustness decision-making.

[0068] Before introducing the embodiments of the application, first introduce the professional terms involved in the application.

[0069] EW-TOPSIS (Entropy Weight-Technique for Order Preference by Similarity to Ideal Solution) is a multi-criteria decision-making method that combines entropy weight method and TOPSIS method. It is mainly used to solve multi-attribute decision-making problems, and ranks and selects schemes by considering the weights of each attribute and the advantages and disadvantages of the schemes.

[0070] Transformer (a deep learning architecture based on self-attention mechanism) can dynamically assign different weights to each element in the input sequence, capturing the dependency between elements.

[0071] Multi-Head Self-Attention (Multi-Head Self-Attention) refers to parallel computing through multiple attention heads (Attention Heads) to capture the relationship between different positions in the sequence. Each attention head calculates the correlation between elements in the input sequence and generates a weighted representation.

[0072] Method embodiments

[0073] Figure 1 A decision-making method is provided for the embodiments of the application.

[0074] In a first aspect of the application, a decision-making method is provided for use in a prediction model based on brain-like intelligent decision-making, which combines Figure 1 The method includes six steps S1 to S6:

[0075] S1: After the multi-modal feature extraction and fusion module in the prediction model based on brain-like intelligent decision-making receives the input information to be optimized, the module performs feature extraction and generates a fusion feature vector.

[0076] Specifically, in the embodiment of the present application, after the multi-modal feature extraction and fusion module in the brain-like intelligent decision-based prediction model receives the input information of the decision to be optimized, it extracts and fuses the features of the multi-modal data to generate a fusion feature vector. Here, the features are extracted from image, time series signal, text and other multi-modal data, and a unified feature vector is generated through a gating fusion mechanism. The method of generating a fusion feature vector is described in detail below.

[0077] S2: The Transformer spatiotemporal modeling module in the brain-like intelligent decision-based prediction model performs spatiotemporal modeling on the fusion feature vector to obtain a spatiotemporal feature matrix.

[0078] Specifically, in the embodiment of the present application, the method of the Transformer spatiotemporal modeling module performing spatiotemporal modeling on the fusion feature vector to obtain a spatiotemporal feature matrix is described in detail below. Here, the Transformer spatiotemporal modeling module performs spatiotemporal modeling on the fusion feature to capture long-range dependencies.

[0079] S3: The EW-TOPSIS dynamic evaluation module in the brain-like intelligent decision-based prediction model processes the spatiotemporal feature matrix to generate a decision candidate ranking.

[0080] Specifically, in the embodiment of the present application, the method of the EW-TOPSIS dynamic evaluation module processing the spatiotemporal feature matrix to generate a decision candidate ranking is described in detail below. Here, the decision candidate ranking is generated based on the entropy weight method and the TOPSIS algorithm.

[0081] S4: The dynamic memory module in the brain-like intelligent decision-based prediction model filters the decision candidate ranking to obtain a decision candidate set.

[0082] Specifically, in the embodiment of the present application, the method of the dynamic memory module filtering the decision candidate ranking to obtain a decision candidate set is described in detail below. Here, short-term and long-term memories are stored to support fast case retrieval, and a decision candidate set is obtained.

[0083] S5: The reinforcement learning module in the brain-like intelligent decision-based prediction model is used to update the decision gradient of the brain-like intelligent decision-based prediction model.

[0084] Specifically, in the embodiment of the present application, the method of the reinforcement learning module updating the decision gradient of the brain-like intelligent decision-based prediction model through the Actor-Critic framework to optimize the decision model is described in detail below.

[0085] S6: The behavior convergence module in the brain-like intelligent decision-based prediction model is used to constrain the decision candidate set to obtain an optimal decision.

[0086] Specifically, in an embodiment of the present invention, the behavior convergence module in the prediction model based on brain-like intelligent decision-making is used to constrain the decision candidate set to obtain the optimal decision, ensuring the feasibility and optimality of the decision. The method of constraining the decision candidate set to obtain the optimal decision is introduced in detail below.

[0087] Figure 2 A diagram of a method for generating a fused feature vector provided by an embodiment of the present invention.

[0088] Furthermore, in the above-mentioned decision-making method, the multimodal feature extraction and fusion module in the prediction model based on brain-like intelligent decision-making performs feature extraction and generates a fusion feature vector after receiving the input information of the decision to be optimized. Figure 2 , including two steps S21 to S22:

[0089] S21: Extract image, time series and text feature vectors respectively through pre-trained CNN, BiLSTM and BERT to obtain feature vector sets.

[0090] Specifically, in the embodiment of the present invention, the image feature vector f is extracted by a pre-trained CNN (such as ResNet-50) img , extract the time series feature vector f through pre-trained BiLSTM seq , extract text feature vector f through pre-trained BERT text , image feature vector f img , time series feature vector f seq and text feature vector f text Composed feature vector set f i .

[0091] S22: Use the gated fusion mechanism to dynamically assign weights to the feature vector set to generate a fused feature vector.

[0092] Specifically, in the embodiment of the present invention, the gated fusion mechanism calculates the weight coefficient α i The formula is as follows:

[0093] α i =Softmax(W g ·[h prev ;f i ])

[0094] Among them, the weight coefficient α i Combine the historical state h through the Softmax function prev With learnable parameters W g Calculated, h prev Indicates the historical state, W g represents the learnable parameter, f idenotes the i-th modality feature vector in the feature vector set, for example, i=1 denotes the image feature vector f img , i=2 denotes the time series feature vector f seq , and i=3 denotes the text feature vector f text .

[0095] The formula for generating the fusion feature is as follows:

[0096]

[0097] Here, the image data needs to be normalized to the range [0, 1], the time series signal needs to be aligned with the timestamp, the text data needs to be segmented and stop words need to be removed, and in the gating fusion mechanism, W g needs to be optimized through back propagation.

[0098] Further, in the above decision method, the Transformer spatiotemporal modeling module in the brain-inspired intelligent decision prediction model performs spatiotemporal modeling on the fusion feature vector to obtain a spatiotemporal feature matrix, including:

[0099] The long-range dependency in the fusion feature vector is captured through a self-attention mechanism, and position encoding and timestamp encoding are introduced to obtain the spatiotemporal feature matrix.

[0100] Specifically, in the embodiment of the present application, the fused feature F fused is input into the Transformer spatiotemporal modeling module, the long-range dependency is captured through a self-attention mechanism , and position encoding PE(t) and timestamp encoding are introduced to enhance the spatiotemporal correlation, and a spatiotemporal feature matrix H out ∈R T×D is output.

[0101] Here, the fusion feature: F fused ∈R T×D ;

[0102] R represents the set of real numbers, T represents the number of time steps (sequence length), and D represents the feature dimension. The shape of the fusion feature is TxD, that is, the number of time steps is T and the feature dimension of each time step is D.

[0103] The formula for calculating the attention weight through the self-attention mechanism is as follows:

[0104]

[0105] Q (Query), K (Key), and V (Value) represent matrices obtained by linear transformation of the input feature; d k represents the dimension of the Key, which is used to scale the dot product to prevent gradient explosion or disappearance.

[0106] Position encoding PE(t) and timestamp encoding are introduced to enhance spatiotemporal correlation. Position encoding uses a sine function, and timestamp encoding must be aligned with the data sampling frequency.

[0107] Spatiotemporal feature matrix: H out ∈R T×D .

[0108] R: real number set, T: time step number, D: feature dimension, output spatiotemporal feature matrix H out ∈R T×D The number of time steps and feature dimensions of the input are preserved, but the spatiotemporal correlation is enhanced through self-attention.

[0109] Figure 3 A diagram of a method for generating a decision candidate ranking provided by an embodiment of the present invention.

[0110] Furthermore, in the above-mentioned decision-making method, the EW-TOPSIS dynamic evaluation module in the prediction model based on brain-like intelligent decision-making processes the spatiotemporal feature matrix to generate a ranking of decision candidates, and combines Figure 3 , including four steps from S31 to S34:

[0111] S31: Calculate indicator weights based on entropy weight method;

[0112] Specifically, in the embodiment of the present invention, the entropy value e is calculated by the following formula: j :

[0113]

[0114] i represents the sample index, j represents the feature index, n represents the total number of samples, and the standardized feature value Entropy value e j is the information entropy of feature j, reflecting the amount of information or uncertainty of the feature, and xij represents the jth eigenvalue of the i-th sample;

[0115] The weight is calculated using the following formula:

[0116]

[0117] S32: Calculate the Euclidean distance between the input information of the decision to be optimized and the positive and negative ideal solutions through the TOPSIS algorithm.

[0118] Specifically, in the embodiment of the present invention, the Euclidean distance between the input information of the decision to be optimized and the positive and negative ideal solutions is calculated by the TOPSIS algorithm. and Positive ideal solution distance and negative ideal solution distance is the weighted Euclidean distance from sample i to the positive / negative ideal solution, with weight w jThe entropy weight method is used to calculate the entropy value to determine the weight, and the weight affects the distance calculation.

[0119] S33: Calculate the closeness degree according to the Euclidean distance.

[0120] Specifically, in the embodiment of the present application, the closeness degree C i is calculated according to the following formula:

[0121]

[0122] S34: Generate decision candidate ranking according to closeness degree.

[0123] Specifically, in the embodiment of the present application, the decision candidate ranking is generated according to the closeness degree.

[0124] Figure 4 A method for obtaining a decision candidate set provided by the embodiment of the present application is shown in the figure.

[0125] Further, in the above-mentioned decision method, the dynamic memory module in the brain-like intelligent decision prediction model is used to screen the decision candidate ranking to obtain a decision candidate set, in combination with Figure 4 , including two steps S41 to S42:

[0126] S41: Short-term memory eliminates decision candidate ranking with a frequency lower than a preset threshold based on LRU strategy;

[0127] S42: Long-term memory selects persistent decision candidate ranking according to decision effect;

[0128] In the selection of persistent decision candidate ranking, the Faiss library is also used to construct vector index to search similar decisions.

[0129] Specifically, in the embodiment of the present application, the short-term memory eliminates low-frequency data based on LRU strategy, and the long-term memory selects persistent data according to decision effect △AUC, and the Faiss library is used to accelerate similar case search in the selection of persistent data.

[0130] Figure 5 A method for updating the decision gradient of the brain-like intelligent decision prediction model provided by the embodiment of the present application is shown in the figure.

[0131] Further, in the above-mentioned decision method, the reinforcement learning module in the brain-like intelligent decision prediction model is used to update the decision gradient of the brain-like intelligent decision prediction model, in combination with Figure 5 , including three steps S51 to S53:

[0132] S51: Output the action distribution of the decision candidate set through the Actor network, and evaluate the state value of the decision candidate set through the Critic network.

[0133] Specifically, in the embodiments of the present application, the Actor network outputs an action distribution π(a|s) according to a state s, and the Critic network evaluates a state value V(s), where a represents an action (decision), s represents a state (environment information), and r represents a reward (feedback value of the action).

[0134] S52: Calculate an advantage function according to the action distribution and the state value.

[0135] Specifically, in the embodiments of the present application, the formula for calculating the advantage function is as follows:

[0136] A(s, a) = r + γV(s') - V(s)

[0137] s' represents a next state transferred after the action a is performed, and in the advantage function A(s, a) = r + γV(s') - V(s), s' is used to evaluate the future state value, the learning rate γ can be set to 0.95, and the policy update frequency is matched with the environment dynamics.

[0138] S53: Update the decision gradient of the prediction model based on the brain-like intelligent decision according to the advantage function.

[0139] Specifically, in the embodiments of the present application, the decision gradient of the prediction model based on the brain-like intelligent decision is updated according to the advantage function The update frequency is matched with the environment dynamics, θ represents a trainable parameter of the Actor network, and is used to define the decision strategy of the prediction model based on the brain-like intelligent decision.

[0140] Figure 6 A method for obtaining an optimal decision provided by the embodiments of the present application is shown in the following figure.

[0141] Further, in the above decision method, the behavior convergence module in the prediction model based on the brain-like intelligent decision is used to constrain the decision candidate set to obtain an optimal decision, and in combination with Figure 6 , the method includes two steps S61 and S62:

[0142] S61: Constrain the decision candidate set by a constraint condition to obtain a constrained decision candidate set;

[0143] Specifically, in the embodiments of the present application, the feasible decision set is represented as A valid , and the constraint condition is dynamically adjusted according to the specific application scenario.

[0144] S62: Select an optimal action in the constrained decision candidate set as an optimal decision by a greedy strategy.

[0145] Specifically, in the embodiment of the present invention, the formula for selecting the optimal action in the decision candidate set after the constraints by the greedy strategy as the optimal decision is expressed as follows:

[0146]

[0147] Among them, in the greedy strategy, Q(s,a) is evaluated in real time by the Critic network, a * It represents the solution to the policy optimization problem, and Q(s,a) represents the action-value function (Action-Value Function) or Q function, which is used to quantify the long-term expected benefits of performing action a in a specific state s. It is the key basis for decision-making in the prediction model of brain-like intelligent decision-making.

[0148] Model embodiment

[0149] Figure 7 A schematic diagram of a prediction model based on brain-inspired intelligent decision-making provided in an embodiment of the present invention.

[0150] The second aspect of the present invention also proposes a prediction model based on brain-like intelligent decision-making, combined with Figure 7 , including: multimodal feature extraction and fusion module 71, Transformer spatiotemporal modeling module 72, EW-TOPSIS dynamic evaluation module 73, dynamic memory module 74, reinforcement learning module 75 and behavior convergence module 76,

[0151] The multimodal feature extraction and fusion module 71 is used to extract features from the input information to be optimized and generate a fusion feature vector.

[0152] Specifically, in an embodiment of the present invention, after receiving the input information of the decision to be optimized, the multimodal feature extraction and fusion module 71 performs feature extraction and fusion of multimodal data to generate a fused feature vector. Here, features are extracted from multimodal data such as images, time series signals, and text, and a unified feature vector is generated through a gated fusion mechanism.

[0153] The Transformer spatiotemporal modeling module 72 is used to perform spatiotemporal modeling on the fused feature vector to obtain a spatiotemporal feature matrix.

[0154] Specifically, in the embodiment of the present invention, the Transformer spatiotemporal modeling module 72 performs spatiotemporal modeling on the fused feature vector to obtain a spatiotemporal feature matrix, which performs spatiotemporal modeling on the fused features to capture long-range dependencies.

[0155] The EW-TOPSIS dynamic evaluation module 73 is used to process the spatiotemporal feature matrix to generate decision candidate rankings.

[0156] Specifically, in the embodiment of the present application, the EW-TOPSIS dynamic evaluation module 73 generates the decision candidate ranking based on the entropy weight method and the TOPSIS algorithm.

[0157] The dynamic memory module 74 is used to screen the decision candidate ranking to obtain a decision candidate set.

[0158] Specifically, in the embodiment of the present application, the dynamic memory module 74 screens the decision candidate ranking to obtain a decision candidate set, which stores short-term and long-term memories and supports fast case retrieval to obtain the decision candidate set.

[0159] The reinforcement learning module 75 is used to update the decision gradient of the prediction model based on the brain-like intelligent decision.

[0160] Specifically, in the embodiment of the present application, the reinforcement learning module 75 optimizes the decision model through the Actor-Critic framework and updates the decision gradient of the prediction model based on the brain-like intelligent decision.

[0161] The behavior convergence module 76 is used to constrain the decision candidate set to obtain an optimal decision.

[0162] Specifically, in the embodiment of the present application, the behavior convergence module is used to constrain the decision candidate set to obtain an optimal decision, thereby ensuring the feasibility and optimality of the decision.

[0163] The third aspect of the present application further provides an electronic device, comprising a processor and a memory.

[0164] The processor is used to execute the decision method according to any one of the above aspects by calling the program or instructions stored in the memory.

[0165] The fourth aspect of the present application further provides a computer readable storage medium, which stores programs or instructions, and the programs or instructions make the computer execute the decision method according to any one of the above aspects.

[0166] Figure 8 is a schematic block diagram of an electronic device provided by the embodiment of the present application.

[0167] As shown in Figure 8 the electronic device includes at least one processor 801, at least one memory 802 and at least one communication interface 803. The various components in the electronic device are coupled together through a bus system 804. The communication interface 803 is used for information transmission between the external device. It can be understood that the bus system 804 is used to realize the connection communication between the components. In addition to the data bus, the bus system 804 also includes a power supply bus, a control bus and a state signal bus. However, for the purpose of clear illustration, only the data bus is shown in the figure.Figure 8 In the embodiment, the various buses are marked as bus system 804.

[0168] It can be understood that the memory 802 in the embodiment can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.

[0169] In some embodiments, the memory 802 stores elements, executable units or data structures, or a subset of them, or an extended set of them: an operating system and an application program.

[0170] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program includes various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The program for implementing any of the decision-making methods provided by the embodiments of the application can be included in the application program.

[0171] In the embodiments of the application, the processor 801 processes the steps of each embodiment of the decision-making method provided by the embodiments of the application by calling the programs or instructions stored in the memory 802, specifically, the programs or instructions stored in the application program.

[0172] The multi-modal feature extraction and fusion module in the brain-like intelligent decision-making based prediction model performs feature extraction and generates a fusion feature vector after receiving input information of a decision to be optimized;

[0173] The Transformer spatio-temporal modeling module in the brain-like intelligent decision-making based prediction model performs spatio-temporal modeling on the fusion feature vector to obtain a spatio-temporal feature matrix;

[0174] The EW-TOPSIS dynamic evaluation module in the brain-like intelligent decision-making based prediction model processes the spatio-temporal feature matrix to generate a decision candidate ranking;

[0175] The dynamic memory module in the brain-like intelligent decision-making based prediction model filters the decision candidate ranking to obtain a decision candidate set;

[0176] The reinforcement learning module in the brain-like intelligent decision-making based prediction model is used to update the decision gradient of the brain-like intelligent decision-making based prediction model;

[0177] The behavior convergence module in the brain-like intelligent decision-making based prediction model is used to constrain the decision candidate set to obtain an optimal decision.

[0178] Any of the decision-making methods provided in the embodiments of the present application can be applied in the processor 801 or implemented by the processor 801. The processor 801 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 801. The processor 801 described above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general processor can be a microprocessor or the processor can also be any conventional processor.

[0179] Any of the steps of the decision-making methods provided in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or executed by a combination of hardware and software units in the decoding processor. The software units can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register or other mature storage media in the field. The storage media is located in the storage 802, and the processor 801 reads information in the storage 802 and completes the steps of the method in combination with the hardware thereof.

[0180] Those skilled in the art can understand that although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments means to be within the scope of the present application and form different embodiments.

[0181] Those skilled in the art can understand that the description of each embodiment is focused on, and the parts not described in detail in a certain embodiment can refer to the related description of other embodiments.

[0182] Although the embodiments of the present application are described in conjunction with the accompanying drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes are intended to fall within the scope of the appended claims. Above, only specific embodiments of the present application are described, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0183] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements shall be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A decision-making method, characterized in that: Applied to a prediction model based on brain-inspired intelligent decision-making, the method includes: The multimodal feature extraction and fusion module in the prediction model based on brain-like intelligent decision-making receives the input information of the decision to be optimized, performs feature extraction and generates a fusion feature vector; Based on the Transformer spatiotemporal modeling module in the prediction model of brain-like intelligent decision-making, the fusion feature vector is subjected to spatiotemporal modeling to obtain a spatiotemporal feature matrix; The EW-TOPSIS dynamic evaluation module in the prediction model based on brain-like intelligent decision-making processes the spatiotemporal feature matrix to generate a ranking of decision candidates; The dynamic memory module in the prediction model based on brain-like intelligent decision-making screens the decision candidates to obtain a decision candidate set; The reinforcement learning module in the prediction model based on brain-inspired intelligent decision-making is used to update the decision gradient of the prediction model based on brain-inspired intelligent decision-making; The behavior convergence module in the prediction model based on brain-like intelligent decision-making is used to constrain the decision candidate set to obtain the optimal decision.

2. A decision-making method according to claim 1, characterized in that: The multimodal feature extraction and fusion module in the prediction model based on brain-inspired intelligent decision-making receives the input information of the decision to be optimized, extracts features and generates a fusion feature vector, including: The pre-trained CNN, BiLSTM, and BERT are used to extract image, time series, and text feature vectors to obtain feature vector sets. A gated fusion mechanism is used to dynamically assign weights to the feature vector set to generate a fused feature vector.

3. A decision-making method according to claim 1, characterized in that: The Transformer spatiotemporal modeling module in the prediction model based on brain-like intelligent decision-making performs spatiotemporal modeling on the fused feature vector to obtain a spatiotemporal feature matrix, including: The long-range dependencies in the fused feature vector are captured through the self-attention mechanism, and position encoding and timestamp encoding are introduced to obtain the spatiotemporal feature matrix.

4. A decision-making method according to claim 1, characterized in that: The EW-TOPSIS dynamic evaluation module in the prediction model based on brain-like intelligent decision-making processes the spatiotemporal feature matrix to generate a ranking of decision candidates, including: Calculate indicator weights based on the entropy weight method; The TOPSIS algorithm is used to calculate the Euclidean distance between the input information of the decision to be optimized and the positive and negative ideal solutions; Calculate the progress based on Euclidean distance; Generate decision candidate ranking based on posting progress.

5. A decision-making method according to claim 1, characterized in that: The dynamic memory module in the prediction model based on brain-like intelligent decision-making screens the decision candidate ranking to obtain a decision candidate set, including: Short-term memory sorts decision candidates whose frequency is lower than a preset threshold based on the LRU strategy; Long-term memory selects a persistent ranking of decision candidates based on the decision effect; Among them, when selecting the persistent decision candidate ranking, a vector index is constructed through the Faiss library to retrieve similar decisions.

6. A decision-making method according to claim 1, characterized in that: The reinforcement learning module in the prediction model based on brain-inspired intelligent decision-making is used to update the decision gradient of the prediction model based on brain-inspired intelligent decision-making, including: The Actor network outputs the action distribution of the decision candidate set, and the Critic network evaluates the state value of the decision candidate set; Compute the advantage function based on the action distribution and state value; Update the decision gradient of the prediction model based on brain-like intelligent decision-making according to the advantage function.

7. A decision-making method according to claim 1, characterized in that: The behavior convergence module in the prediction model based on brain-inspired intelligent decision-making is used to constrain the decision candidate set to obtain the optimal decision, including: Constraining the decision candidate set through constraint conditions to obtain the constrained decision candidate set; The optimal action in the constrained decision candidate set is selected through a greedy strategy as the optimal decision.

8. A prediction model based on brain-like intelligent decision-making, characterized by: include: Multimodal feature extraction and fusion module, Transformer spatiotemporal modeling module, EW-TOPSIS dynamic evaluation module, dynamic memory module, reinforcement learning module and behavior convergence module, The multimodal feature extraction and fusion module is used to extract features from the input information to be optimized and generate a fusion feature vector; The Transformer spatiotemporal modeling module is used to perform spatiotemporal modeling on the fused feature vector to obtain a spatiotemporal feature matrix; The EW-TOPSIS dynamic evaluation module is used to process the spatiotemporal feature matrix to generate a decision candidate ranking; The dynamic memory module is used to filter the decision candidate rankings to obtain a decision candidate set; The reinforcement learning module is used to update the decision gradient of the prediction model based on brain-like intelligent decision-making; The behavior convergence module is used to constrain the decision candidate set to obtain the optimal decision.

9. An electronic device, characterized in that: include: processor and memory; The processor is configured to execute a decision-making method according to any one of claims 1 to 7 by calling a program or instruction stored in the memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute a decision-making method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Robot decision control method based on gradient rarefaction and robot

    CN121157058A