Self-evolution method and system of intelligent decision-making large model of fusion OODA ring

By using the monitoring and updating mechanism of the outer autonomous evolutionary loop, the problem of model solidification in intelligent decision-making systems is solved, enabling continuous learning and autonomous adaptation, and improving the system's performance in complex environments.

CN122133741APending Publication Date: 2026-06-02BEIHANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2026-02-10
Publication Date
2026-06-02

Smart Images

  • Figure CN122133741A_ABST
    Figure CN122133741A_ABST
Patent Text Reader

Abstract

This invention provides an autonomous evolution method and system for a large-scale intelligent decision-making model integrating an OODA loop, relating to the field of artificial intelligence technology. The system includes an inner decision loop and an outer autonomous evolution loop; the inner decision loop includes an observation module, a judgment module, a decision module, and an action module; the outer autonomous evolution loop includes a monitoring evaluator and an autonomous evolver. Due to a lack of relevant information in the knowledge graph, the inner decision loop experiences a decrease in overall performance score and an increase in uncertainty. The monitoring evaluator in the outer autonomous evolution loop promptly detects this performance decline and high uncertainty state and triggers the autonomous evolver. The autonomous evolver then uses the collected new data to update the parameters of the recognition model of the observation module, the knowledge graph of the judgment module, and the strategy of the decision module. This overcomes the limitation of traditional systems experiencing significant performance degradation when facing new situations, and continuously optimizes its decision-making capabilities based on new knowledge and challenges encountered in actual operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to an autonomous evolution method and system for intelligent decision-making large-scale models that integrate the OODA loop. Background Technology

[0002] The OODA loop refers to the cycle of "Observation-Orientation-Decision-Action." Victory in dynamic adversarial situations is achieved by completing the OODA loop before the opponent and disrupting, interrupting, or invading their OODA loop. Since its inception, the OODA loop theory has undergone various improvements by researchers to adapt it to more complex environments, driven by technological advancements and diverse scenario requirements.

[0003] In recent years, combining OODA loop theory with artificial intelligence technology to build intelligent decision-making systems has become an important research direction in fields such as command and control and unmanned systems. Existing technologies mainly focus on how to combine artificial intelligence technologies, especially large language models, with the OODA loop. However, in the final system architecture, the large model is fixed for specific tasks, and the core parameters are fixed after system deployment. The system can only make decisions based on its historical best state during training, and cannot continuously and automatically learn and update the model from each decision-making practice and subsequent environmental feedback. This leads to a significant performance degradation when facing new situations. Summary of the Invention

[0004] The problem that this invention aims to solve is that existing technologies suffer from model rigidity and lack the ability to continuously evolve.

[0005] To address the aforementioned issues, in a first aspect, the present invention provides an intelligent decision-making large-scale model autonomous evolution system that integrates an OODA loop, comprising an inner decision loop and an outer autonomous evolution loop; The inner decision loop is used to receive multi-source heterogeneous data and output situation tensors, action vectors, and control command vectors to control the actuators; the inner decision loop includes an observation module, a judgment module, a decision module, and an action module. The outer self-evolving ring includes: The monitoring and evaluation system is used to obtain the moving average, average response time, and task completion rate based on the action vector and the preset optimal action vector under the same situation, and then obtain the comprehensive performance score; it is also used to obtain the uncertainty based on the current situation tensor and the historical situation tensor; and based on the comprehensive performance score and uncertainty, it determines whether to perform an autonomous evolution update on the inner decision loop. An autoevolver is used to update the parameters of the observation module, judgment module, and decision module, and update the knowledge graph in the judgment module, if an autoevolving update is performed on the inner decision loop.

[0006] Secondly, this invention also provides an autonomous evolution method for intelligent decision-making large-scale models that integrate the OODA loop, comprising: The inner decision loop receives heterogeneous data from multiple sources and outputs situation tensors, action vectors, and control command vectors to control the actuators; the inner decision loop includes an observation module, a judgment module, a decision module, and an action module. The outer autonomous evolution loop obtains the moving average, average response time, and task completion rate based on the action vector and the preset optimal action vector under the same situation, and then obtains the comprehensive performance score; it is also used to obtain the uncertainty based on the current situation tensor and the historical situation tensor; based on the comprehensive performance score and uncertainty, it determines whether to perform autonomous evolution update on the inner decision loop; if it is determined to perform autonomous evolution update on the inner decision loop, the parameters of the observation module, judgment module, and decision module are updated, and the knowledge graph in the judgment module is updated.

[0007] This invention provides an autonomous evolution method and system for intelligent decision-making large-scale models that integrates the OODA loop. Compared with existing technologies, it has the following advantages: By introducing an outer autonomous evolutionary loop, continuous and automated learning and model updates of the inner decision-making loop are achieved. The inner decision-making loop may fail to accurately identify or assess threats due to a lack of relevant information in the knowledge graph, leading to a decline in overall performance score and increased uncertainty. In this situation, the monitoring and evaluator in the outer autonomous evolutionary loop promptly captures this performance degradation and high uncertainty state, triggering the autonomous evolver. The autonomous evolver then uses the collected new data to update the parameters of the observation module's identification model, the judgment module's knowledge graph, and the decision-making module's strategy. This dynamic update mechanism enables the system to continuously learn from each decision-making practice and subsequent environmental feedback, thus overcoming the limitation of traditional systems experiencing significant performance degradation when facing new situations. With the intervention of the outer autonomous evolutionary loop, the system no longer relies solely on the historical optimal state before deployment for decision-making, but can continuously optimize its decision-making capabilities based on new knowledge and challenges encountered during actual operation. This not only improves the system's adaptability and robustness in unknown environments but also extends the system's lifespan and scope of application. This system possesses autonomous adaptive and continuously evolving intelligent decision-making capabilities, significantly improving the performance of intelligent systems in complex and dynamic environments. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a schematic diagram of the structure of an intelligent decision-making large model autonomous evolution system that integrates an OODA loop, as provided in an embodiment of the present invention. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application are described clearly and completely. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0012] like Figure 1 As shown in the embodiment of this application, an intelligent decision-making large model autonomous evolution system integrating an OODA loop is provided, including an inner decision loop and an outer autonomous evolution loop; The inner decision loop is used to receive multi-source heterogeneous data and output situation tensors, action vectors and control command vectors to control the actuators; the inner decision loop includes an observation module, a judgment module, a decision module and an action module.

[0013] The outer self-evolving ring includes: The monitoring and evaluation system is used to obtain the moving average, average response time, and task completion rate based on the action vector and the preset optimal action vector under the same situation, and then obtain the comprehensive performance score; it is also used to obtain the uncertainty based on the current situation tensor and the historical situation tensor; and based on the comprehensive performance score and uncertainty, it determines whether to perform an autonomous evolution update on the inner decision loop. An autoevolver is used to update the parameters of the observation module, judgment module, and decision module, and update the knowledge graph in the judgment module, if an autoevolving update is performed on the inner decision loop.

[0014] In this optional embodiment, by introducing an outer autonomous evolutionary loop, continuous and automated learning and model updates of the inner decision-making loop are achieved. For example, when a UAV encounters a new type of stealth aircraft for the first time, the inner decision-making loop may be unable to accurately identify or assess the threat due to a lack of relevant information in the knowledge graph, leading to a decrease in overall performance score and an increase in uncertainty. At this time, the monitoring and evaluation device in the outer autonomous evolutionary loop will promptly capture this performance degradation and high uncertainty state and trigger the autonomous evolutionary device. The autonomous evolutionary device then uses new data about the new aircraft collected by the UAV during actual flight to update the parameters of the recognition model of the observation module, the knowledge graph of the judgment module, and the strategy of the decision-making module. For example, the visual features and radar signal features of the new aircraft are added to the knowledge graph, and the recognition model is updated to improve its ability to recognize the new target. This dynamic update mechanism enables the system to continuously learn from each decision practice and subsequent environmental feedback, thereby overcoming the limitation of traditional systems that experience significant performance degradation when facing new situations. By incorporating an outer autonomous evolutionary loop, the system no longer relies solely on its pre-deployment historical optimal state for decision-making. Instead, it continuously optimizes its decision-making capabilities based on new knowledge and challenges encountered during actual operation. This not only enhances the system's adaptability and robustness in unknown environments but also extends its lifespan and scope of application. This system possesses autonomous adaptive and continuously evolving intelligent decision-making capabilities, significantly improving the performance of intelligent systems in complex and dynamic environments.

[0015] The structure of each part of the system is described in detail below.

[0016] like Figure 1 As shown, the inner decision loop is responsible for quickly executing decision-making tasks in a dynamic environment; the outer autonomous evolution loop is responsible for monitoring, evaluating and optimizing the inner decision loop, forming a closed loop of "decision-evaluation-evolution", which enables the system to have the ability to self-reflect and self-optimize.

[0017] The inner decision-making loop comprises an observation module, a judgment module, a decision-making module, and an action module. The inner decision-making loop is the main execution component of the system, and its core lies in deeply integrating the large model into the four stages of OODA (Object-Oriented Analysis and Decision Making).

[0018] 1. The observation module is used to fuse multi-source heterogeneous data to generate a situation tensor.

[0019] The core task of the Observation module is to achieve multimodal data fusion and spatiotemporal feature extraction. This module constructs a multimodal spatiotemporal fusion network composed of a convolutional neural network (CNN) and a long short-term memory network (LSTM), which fuses different types of data from multiple heterogeneous sensors to generate a unified situation tensor with spatiotemporal dimensions.

[0020] For image data, CNNs are used for feature extraction. The convolution operation slides a set of learnable convolutional kernels across the input feature map, extracting image features through local weighted summation. Let the activation value of the output convolutional feature map of the l-th layer at the two-dimensional coordinates (x, y) be... The calculation is as follows: (1) The spatial size of the convolution kernel is k is a natural number; The value represents the activation value of the convolutional feature map of layer l-1 at coordinates (x+i, y+j), where i and j are the index variables of the convolutional kernel in the horizontal and vertical directions, and their traversal range is determined by the size of the convolutional kernel. This represents the weight parameters of the l-th convolutional kernel at its relative coordinates (i,j). The bias parameters are the kernel bias parameters of the l-th layer. Both the weight parameters and the bias parameters are learnable parameters. To modify the ReLU activation function, nonlinearity is introduced to enable the network to learn more complex feature patterns.

[0021] For continuous frames of sensor data or time-series state data, LSTM is used to capture their temporal dependencies. Let the time step be t, and the input be... The hidden state of the previous time step, i.e., the extracted temporal features, is The internal state update of the LSTM cell is as follows: (2) in, , , These represent the activation values ​​of the input gate, forget gate, and output gate at time t, respectively, which determine the information flow between the current input, historical memory, and final output. The value of each gate is between 0 and 1. The Sigmoid function compresses the input to the (0,1) interval and is used for gating. The candidate cell state at time step t represents new information calculated from the current input and the hidden state at the previous time step; tanh is the hyperbolic tangent activation function that compresses the input to the (-1,1) interval to generate candidate states and output. This represents the current cell state and is the long-term memory of the LSTM. The current hidden state represents the short-term memory of the LSTM, and the symbol is... This represents element-wise multiplication; , , and Here are the weight matrices corresponding to each gating and candidate state. , , and These are the corresponding biases, and these are all learnable parameters of the LSTM network. The spatial feature vector output by the CNN is concatenated with the temporal feature vector output by the LSTM, and then a fully connected layer is used for nonlinear transformation and dimensionality reduction, ultimately outputting a unified situation tensor. : (3) Among them is The weight matrix of the fully connected layer is responsible for learning how to effectively fuse features from different modalities. It is a bias vector. It is the ReLU activation function. A fixed-dimensional situation tensor that contains spatiotemporal information about the current environment.

[0022] 2. The Orientation module is used to obtain the cognitive context vector based on the input text, the situation tensor, and the knowledge graph. The Orientation module is responsible for high-level semantic understanding and intent recognition of the situation tensor, combining the general reasoning capabilities of large-scale language models with the structured knowledge of the domain knowledge graph to achieve deep cognition of the dynamic environment.

[0023] The step of obtaining the cognitive context vector based on the input text, the situation tensor, and the knowledge graph includes: The situation tensor is input into the fully connected layer to obtain the situation embedding vector.

[0024] Specifically, the situation tensor output by the observation module This is transformed into a vector representation semantically aligned with the word embedding space of a large language model, achieved through a trainable fully connected layer: (4) in, and Let be the weight matrix and bias vector of the fully connected projection layer, and be the learnable parameters. The dimension of word embedding in the large model; The generated situation embedding vector contains the semantic encoding of situation information.

[0025] Based on the input text and the situational embedding vector, the primary cognitive embedding vector is obtained.

[0026] Specifically, the input text (such as task instructions, system status, and other natural language text) is fed into the word embedding layer built into the Large Language Model (LLM). This layer converts each lexical token into its corresponding word embedding vector, forming a text embedding sequence. Where m is the number of tags in the text, each .

[0027] Embed the situation vector As a special context representing the current environment, and text embedding sequences The sequences are concatenated to form a final cue word embedding sequence of length (m+1): (5) will sequence The input is fed into a Transformer-based LLM model, which utilizes its internal self-attention mechanism to encode and reason about the entire input sequence. This allows the model to understand the semantic relationship between the current situation and task instructions, and outputs a natural language description or structured summary of the current situation, encoded as a primary cognitive embedding vector. It contains multiple tokens.

[0028] One or more entity references are extracted from the primary cognitive embedding vector. Each entity reference contains at least the entity text and its corresponding context information.

[0029] Specifically, a pre-trained Named Entity Recognition (NER) model is used. This model is based on the classic BERT-BiLSTM-CRF architecture, but its input layer is specially designed to adapt to the output of large language models. This architecture has demonstrated high performance on several publicly available named entity recognition datasets (such as CoNLL-2003 and OntoNotes), capable of recognizing entity mentions belonging to predefined categories in text. The specific steps are as follows: First, Each token in the algorithm is projected onto the input dimension of BERT through a fully connected layer and then fed into the pre-trained BERT model to obtain the contextual representation of each token. Next, the representation of each token output by BERT is fed into a bidirectional LSTM (BiLSTM) to capture contextual information and obtain the BiLSTM hidden state of each token. Finally, a Conditional Random Field (CRF) layer is used to perform sequence labeling on the BiLSTM output, identifying entity mentions in the text according to the BIO labeling system. The predefined entity categories in this system include: target entities such as "debris A", "spacecraft", and "vehicle B"; action entities such as "avoidance", "braking", and "acceleration"; attribute entities such as "high speed", "danger", and "low fuel"; and spatiotemporal entities such as "orbital plane", "100 meters behind", and "2 o'clock direction". The CRF layer outputs the optimal label sequence, and the system extracts the text, type, and location information of entity mentions based on the label sequence to form an entity mention set. .

[0030] Input the text string mentioned by the entity into the Sentence-BERT model to obtain the text semantic information.

[0031] Input the name and alias of each entity in the knowledge graph into the Sentence-BERT model to obtain name semantic information.

[0032] Analyze the cosine similarity between text semantic information and name semantic information, and select graph entities with a cosine similarity greater than a preset similarity as candidate graph entities to obtain a set of candidate graph entities.

[0033] Specifically, the system introduces a domain-specific knowledge graph. This study validates, supplements, and deepens our understanding of LLM, where E is the entity set, R is the relation set, and F is the fact triple set, including: head entity. , indicating the subject in a relation; tail entity , indicating the object or end point in a relation; relation This indicates the semantic relationship connecting the head entity and the tail entity.

[0034] Using a pre-trained Sentence-BERT model, which is based on the Transformer architecture, can map variable-length text into dense 768-dimensional vectors, capturing the deep semantic information of the text: (6) For each entity e in the knowledge graph, pre-compute the semantic vector of its name and main aliases: (7) Calculate the mention vector With each entity vector Cosine similarity: (8) Preserving semantic similarity exceeding a threshold The entity is output as a candidate entity set: (9) The link confidence score is obtained based on the context sentence of the mention obtained from the entity mention information and the attribute field text of the candidate graph entity in the knowledge graph corresponding to each candidate graph entity.

[0035] Specifically, a ranking model based on BERT cross-encoder is used to determine the unique correct link target from the candidate entity set. For each candidate entity... Construct the input sequence in the format: [CLS]Context Sentence [SEP]Entity Description Text [SEP], where the context sentence is... The entity description text is the text of the `description` attribute field of entity `e` in the knowledge graph. The sequence is input into a BERT-base model to obtain the hidden state vector labeled `[CLS]` in the last layer. As a joint representation of the entire input pair. Using a single-layer fully connected neural network, and outputting link confidence scores using the Sigmoid function: (10) in, The weight parameters represent the weights of the fully connected layer. These are the bias parameters for the fully connected layer. This is the Sigmoid activation function.

[0036] If the highest link confidence score is greater than the preset confidence score, the corresponding graph entity is recorded as the final graph entity, the entity mention is linked to the final graph entity, and marked as linked; otherwise, the entity mention is marked as unlinked, and the number of times the entity mention is unlinked is recorded.

[0037] Specifically, the best candidate is selected from the set of candidate entities: (11) like Then it will be mentioned Link to entity It is marked as "linked" if it is not linked otherwise. A threshold is set in this system. .

[0038] Finally, for each mention marked "not linked" The system generates a new entity candidate record: (12) Where count is the initial frequency of the mention, Store in a new entity candidate pool When the same mentioned text reappears in a subsequent decision and is marked as "not linked", it will be considered as... Increment the count value by 1.

[0039] Using the successfully linked set of graph entities as the query seed, a two-hop traversal query algorithm based on breadth-first search is adopted to extract triples related to entity mentions from the knowledge graph, and the two-hop search results are merged to obtain a global knowledge subgraph.

[0040] Specifically, the set of successfully linked entities As a query seed, a two-hop traversal query algorithm based on breadth-first search (BFS) is used to extract relevant triples from the knowledge graph. The specific implementation steps are as follows: For each entity In the first-hop query, retrieve all direct related queries from the knowledge graph. Connected triples: (13) The second-hop query is based on the result of the first-hop query. All of the above except For entities outside the scope, retrieve the triples that are directly connected to them: (14) Merging a single entity subgraph Then, all subgraphs are merged, and duplicate triples are removed to construct a global knowledge subgraph: (15) Where k is the number of entities that were successfully linked.

[0041] The cognitive context vector is obtained by fusing the primary cognitive embedding vector and the global knowledge subgraph.

[0042] Specifically, a graph attention network (GAT) is used to encode knowledge subgraphs, embedding the primary cognition generated by the large model into vectors. The graph embedding vectors of the knowledge subgraph are fused with features from a multilayer perceptron (MLP) to generate a unified cognitive context vector. The specific implementation steps are as follows: A two-layer GAT encoding method is used for the knowledge subgraph. First, node initialization is performed. For each entity node v in the knowledge subgraph, pre-trained GloVe word vectors (300 dimensions) are used as its initial features. If the entity contains multiple words, the average of the word vectors is taken. (16) For the l-th layer (l=0,1), the update formula for node v is: (17) in, This represents the set of neighboring nodes of node v in the knowledge subgraph. It is a learnable weight matrix, where , , , It is the attention weight, calculated as follows: (18) in It is a learnable parameter vector of the attention mechanism, symbol This represents vector concatenation, and LeakyReLU is an activation function with a negative slope of 0.2. The first layer uses the ReLU function, and the second layer uses identity mapping. After two layers of GAT, max pooling is performed on the feature vectors of all nodes in the graph to obtain the graph embedding vector. : (19) A two-layer fully connected neural network is used for feature fusion, embedding the primary cognitive data output by the large model into an embedded vector. With graph embedding vectors Perform the splicing, then input it into the fusion network: (20) in, , These are the parameters of the first layer. , These are the parameters of the second layer, with ReLU and tanh as activation functions, ultimately outputting a cognitive context vector. .

[0043] 3. The decision-making module, based on the cognitive context vector, derives the action vector and state value scalar. The decision-making module is responsible for rapidly generating the optimal action strategy based on the understanding of the situation. It consists of a policy network and a value network, and is based on the cognitive context vector from the decision-making module. Output action vector and state value scalar .

[0044] The step of obtaining the action vector and state value scalar based on the cognitive context vector includes: The cognitive context vector is input into the multilayer perceptron within the policy network. The output layer of the policy network outputs parameters used to define the action probability distribution, i.e., action probability distribution parameters.

[0045] Based on the action probability distribution parameters and the logarithmic standard deviation parameter vector, the following is adopted: ε - Greedy search strategy, sampling to obtain action vectors.

[0046] Specifically, the policy network employs a multilayer perceptron (MLP) with two hidden layers, receiving cognitive context vectors from the decision module. The first hidden layer is a fully connected layer, using the ReLU activation function: (twenty one) in, , These are the weight matrix and the bias vector, respectively.

[0047] The second hidden layer is also a fully connected layer, with an output dimension of 64. (twenty two) in, , These are the weight matrix and the bias vector, respectively.

[0048] The output layer has no activation function; it takes the mean of the actions, with the dimension being the dimension of the continuous action space. This represents the degree of freedom of movement: (twenty three) in, , These are the weight matrix and the bias vector, respectively.

[0049] For continuous action spaces, an independent Gaussian distribution is used for each dimension. (Except for the mean...) In addition, a trainable log-standard deviation parameter vector is defined. This vector is independent of the network's forward propagation. (Using...) ε - A greedy exploration strategy that samples action vectors from the distribution output by the policy network with probability ε at each decision. : (twenty four) The mean of the distribution is output directly with probability 1-ε. As an action vector: (25) 'a' represents the action. This strategy allows the module to inject a small amount of randomness to explore the environment and collect diverse data while ensuring the stability of the subject's behavior.

[0050] The cognitive context vector is input into the multilayer perceptron within the value network, and the output layer of the value network outputs a state value scalar that evaluates the value of the current state.

[0051] Specifically, the value network also employs a two-layer MLP, receiving cognitive context vectors from the judgment module. The first and second layers are fully connected layers with dimensions of 128 and 64 respectively, using the ReLU activation function: (26) (27) in, , and , These are the weight matrix and bias vector of the fully connected layer, respectively.

[0052] The value network is similar to the policy network, but it is completely independent, and the output of the output layer is a one-dimensional scalar. Let the output state value scalar be: (28) in , These are the weight matrix and the bias vector, respectively.

[0053] 4. Action module, used to convert action vectors into control command vectors for the actuator.

[0054] Specifically, the Action module is responsible for processing the abstract, high-order action vectors generated by the decision module. Transformed into low-level control instruction vectors that the specific actuator can understand. And ensure that these signals conform to physical constraints. Since the specific implementation of the action module is highly dependent on the specific physical platform, it is usually determined based on platform parameters during system integration. The specific implementation steps are as follows: First, based on the platform type, higher-order instructions are mapped using predefined mapping functions. Transformed into intermediate motion commands For example, in a spacecraft orbital maneuver model, the decision output... For speed increment The desired terminal velocity vector for: (29) in The current speed is derived from the state. The desired acceleration command This is an intermediate motion command. The model can be simplified to tracking speed.

[0055] Then, the motion commands are assigned using the pre-computed control allocation matrix. Assign to each actuator, taking constraints into account: (30) in The efficiency control matrix is ​​an inherent physical parameter matrix of the platform. Apply actuator upper and lower limit saturation constraints to each component: (31) in The clipping function restricts x to... Within the range.

[0056] Furthermore, when the number of actuators exceeds the control degrees of freedom, optimization objectives such as minimizing fuel consumption or energy consumption are introduced to solve the optimization problem. This results in an optimization problem needing to be solved for each actuator's actual underlying control command, ultimately generating a control command vector. It is sent to the physical actuator through the corresponding hardware interface.

[0057] 5. External Environment Module: This module simulates the interaction between the system and the external environment. It updates the environmental state of the external environment module based on the control command vector and receives immediate rewards.

[0058] Specifically, the external environment module simulates the interaction between the system and the external environment. This module has a built-in dynamics model and a preset reward function, which is implemented upon receiving control command vectors. Then, the environment updates its internal state and calls its reward function to calculate the immediate reward. For example, in the spacecraft avoidance mission of implementation scenario 1 below, the preset reward function is designed as follows: (32) in, Represents a spacecraft collision indication function. For the velocity increment modulus, For the fragment avoidance success indicator function, A preset constant is provided for the indicator function for all distances greater than the safety threshold.

[0059] In addition, the environment will also output corresponding change data, which is... The system's observation module receives the corresponding raw observation data and runs its multimodal fusion network to generate a situation tensor used uniformly within the system. This constitutes one complete OODA loop.

[0060] The outputs of each module are combined to generate a decision data stream. , for use by the outer evolutionary ring.

[0061] The outer autonomous evolutionary loop comprises a monitoring evaluator, autonomous evolvers, and an evolution coordinator. The autonomous evolvers further include model evolvers, policy evolvers, and knowledge evolvers. This outer autonomous evolutionary loop is responsible for monitoring and evaluating the long-term performance of the inner decision-making loop and driving its systematic evolution.

[0062] 1. Monitoring and Evaluation Tool Based on comprehensive performance rating The decision-making quality of the inner loop is quantified. Based on the action vector and the preset optimal action vector under the same situation, the moving average, average response time, and task completion rate are obtained, thus yielding a comprehensive performance score. Specifically, this includes the following:

[0063] The moving average is the average of the decision accuracy rates over a preset number of times prior to the current moment.

[0064] Specifically, define decision accuracy. It represents the degree of consistency between the system's decision and the optimal solution to the task or the expert's decision. At each decision point, the action vector output by the system is... With a decision-making database of historical experts in the same situation tensor The optimal action generated below A comparison is made. Cosine similarity is used to measure the consistency of action direction, and combined with amplitude error, the decision accuracy at time t is... (33) in, Represents the action vector. This indicates the preset optimal action. Indicates the magnitude error weight. This represents the decision accuracy at time t. For the most recent Secondary decision The moving average.

[0065] The average response time is the average time delay from the observation module outputting the situation tensor to the action module issuing the corresponding control command vector, calculated from a preset number of times prior to the current moment.

[0066] Specifically, define the average system response time. This represents the output of new trends from the observation module. The action module issues the corresponding control command. Time delay: (34) For each data stream within the system Add a timestamp. Take the nearest Sub-decision cycle The average value.

[0067] The task completion rate is the success rate of achieving the preset task objective within the second preset number of task cycles prior to the current moment.

[0068] Specifically, define the task completion rate. It refers to the success rate of a system in achieving a preset goal within a complete task cycle. This is determined by analyzing the data flow within a complete task cycle. Analyze the final state of each cycle. Success or failure is determined based on different task objectives. For the most recent Success rate of each task cycle.

[0069] The overall performance score is the weighted sum of the moving average, the reciprocal of the mean response time, and the task completion rate.

[0070] Specifically, the overall performance score is as follows: (35) in, , and For the weighting coefficients, satisfying , The reference response time is used for normalization.

[0071] The degree of "surprise" of the current situation is assessed using the kernel density estimation (KDE) method. A sliding window is maintained to store the most recent... A set of historical state tensors For the current situation tensor (Taking time t as the current time), calculate its kernel density estimate in the historical distribution: (36) in, The state tensor at time t The estimated probability density of the current situation is the probability distribution described by the set of historical situation tensors. The larger the value, the more common the current situation is; the smaller the value, the rarer the current situation is. Let t represent the state tensor at time t. L represents the i-th historical situation tensor in the historical situation tensor set, L represents the total number of tensors in the historical situation tensor set, which is preset to 1000; h represents the bandwidth, which controls the width of the kernel function. For the kernel function, this system uses the Gaussian kernel function, that is: (37) in, The value output by the kernel function represents a historical point. For the current point The contribution weight of density estimation is higher the closer the distance.

[0072] The uncertainty is in, This indicates uncertainty. The larger the value, the more the current situation deviates from historical experience, and the higher the environmental uncertainty.

[0073] Then, based on the comprehensive performance score and uncertainty, it is determined whether to perform an autonomous evolution update on the inner decision-making loop.

[0074] Specifically, when the overall performance score is less than the score threshold or the uncertainty is greater than the uncertainty threshold, it is determined that the inner decision loop should be automatically updated and a trigger command signal should be generated; otherwise, it is determined that the inner decision loop should not be automatically updated.

[0075] Continuous monitoring and A command signal to trigger the autonomous autogenerator is generated when any of the following conditions are met: (38) in, As the scoring threshold, As an uncertainty threshold, based on historical data The percentiles of the distribution are dynamically set.

[0076] 2. Self-evolving incubator If it is determined that an autonomous evolution update is needed for the inner decision loop, then the parameters of the observation module, judgment module, and decision module are updated, and the knowledge graph in the judgment module is updated, including: 1) If the inner decision loop is to be updated autonomously, the model evolver minimizes the joint loss function of the observation module and the decision module, and uses an online stochastic gradient descent algorithm to update the task-specific neural network parameters of the observation module and the decision module in the inner decision loop.

[0077] Specifically, the model evolver uses an online stochastic gradient descent algorithm to update the task-specific neural network parameters of the observation and judgment modules in the inner decision loop, aiming to continuously optimize the system's low-level perception and mid-level cognitive abilities.

[0078] The set of parameters updated in the model evolver Specifically as follows: CNN convolutional kernel weights and bias The gating weight matrices of LSTM , , and and corresponding bias , , and Weight matrix of fully connected layers used for feature concatenation and dimensionality reduction With bias The weights of the fully connected projection layer that project the situation tensor onto the semantic space of the large model With bias In the BERT-BiLSTM-CRF model, the parameters of the BiLSTM and CRF layers, and the weights of the single-layer fully connected neural network used to calculate link confidence in the entity linking module. With bias The learnable weight matrix of two layers in GAT and attention mechanism parameter vector Weight matrix of MLP used for feature fusion With bias .

[0079] The update process follows the principle of online gradient descent, utilizing batch data sampled in real time from recent decision data streams to minimize a multi-task loss function: (39) in, Let L be the learning rate, and L be the total loss function, defined as follows: (40) and This is a hyperparameter that balances the loss weights of the observation module and the judgment module. (Observation module loss) The situation tensor generated by the observation module It can more accurately predict the short-term future state of the environment, thereby improving the quality of its perception and representation, which is determined by the future: (41) Where B represents the training batch size, which comes from the data stream. and The tensor representing the current and next time state of the k-th sample in the batch. It is a future state predictor consisting of a CNN-LSTM and a two-layer fully connected network, whose parameters are included in middle.

[0080] Determine the cognitive loss of the module The model is encouraged to develop a consistent perception of the same or similar situations, while widening the gap between its perception of different situations. It uses the Info NCE loss calculation formula from the contrastive learning loss function to calculate: (42) in Let k be the cognitive context vector of the k-th sample. A positive sample represents a state. Below is the ideal cognitive vector corresponding to the actions generated by the expert strategy. The cognitive vector representing other samples in the batch, Represents cosine similarity. The temperature coefficient hyperparameter is used, and the average of the entire batch is taken in the end.

[0081] 2) The policy evolver performs collaborative optimization of the policy network and value network of the decision module based on the total loss of the decision module; where the total loss is the weighted sum of the policy network loss and the value network loss.

[0082] Specifically, the policy evolver uses the policy network of the inner decision-making module. and value network Collaborative optimization is performed to enhance the system's decision-making capabilities. This system employs the Proximal Policy Optimization (PPO-Clip) algorithm as its core update mechanism to ensure the stability of policy updates. Simultaneously, it utilizes Generalized Advantage Estimation (GAE) to calculate the advantage function, achieving efficient value network updates.

[0083] The strategy evolver operates in real time from the decision data stream of the inner decision loop. Empirical data is collected during this process. First, to calculate the policy gradient and update the value network, the advantage function at each time step needs to be estimated. This system uses GAE, and the calculation formula is as follows: (43) in, Represents the dominance function. This represents the discount factor, which measures the current value of future rewards, and is typically set to 0.99. This represents the GAE smoothing parameter, used to balance bias and variance, and is typically set to 0.95. Let t represent the state value scalar at time t. This represents the state value scalar at time t+1. Indicates an immediate reward. This represents the single-step residual at time t. express The single-step residual at each time step. The advantage function needs to be recursively calculated from the future to the present, therefore, at the time of evolution triggering, it is necessary to collect decision data streams for a future time period T.

[0084] Next, the PPO-Clip algorithm is used to update the policy network in the decision-making module by maximizing the objective function, thereby optimizing the objective function (i.e., the policy network loss) as follows: (44) in, This represents the probability of taking action 'a' in the new policy network given a cognitive context vector. This represents the probability of taking action 'a' in the old policy network, given a cognitive context vector. This represents the probability ratio between the old and new policy networks; To truncate the hyperparameters and limit the range of probability ratio variation, the system sets it to 0.2. By introducing a pruning function, the policy update will not cause performance crashes due to excessively large single update steps.

[0085] The value network is updated by minimizing the mean squared error between the predicted and target values. The target value is expressed as GAE + value baseline. (45) Value network loss is (46) in, This represents the state value scalar obtained under the old value network, used for stable training. This represents the scalar value of the state obtained under the new value network. This represents the target value.

[0086] The total loss function of the policy evolver is the weighted sum of the above two terms: (47) Where c is the value loss coefficient, used to balance the weights of policy updates and value updates, typically set to 0.5. Parameter updates use the stochastic gradient descent method consistent with the model evolver: (48) in and These are the learning rates for the policy network and the value network, respectively, and can be set to the same value.

[0087] 3) The knowledge evolver automatically updates the knowledge graph in the judgment module based on the system's decision-making experience.

[0088] Specifically, the knowledge evolver is responsible for automatically updating the knowledge graph in the judgment module based on the system's decision-making experience, thereby optimizing the system's cognitive capabilities. Information about entities marked as unlinked and their unlinked counts are stored in the entity candidate pool. If the same entity mention reappears in a subsequent decision and is marked as unlinked, the unlinked count is accumulated. For example, in addition to the evolution trigger conditions based on the monitoring evaluator, the knowledge evolver checks the knowledge graph in the judgment module every 100 decision cycles. For those count values ​​exceeding the threshold Upon receiving the record, the knowledge evolver will initiate a new entity creation process, creating new entity nodes and establishing related relationships in the knowledge graph (KG). If no record exceeding the threshold is found, the record with the largest count value in the current pool can be added.

[0089] The knowledge evolver automatically updates the knowledge graph in the judgment module based on the system's decision-making experience, including: When a trigger command signal is received or the number of decision loops reaches a preset number, information mentioned by candidate entities in the entity candidate pool whose number of unlinked times is greater than a preset unlinked time threshold is extracted.

[0090] Based on the information mentioned by the candidate entity, the semantic vector of the candidate entity is obtained using the pre-trained Sentence-BERT model, and compared with the semantic vector of the existing entity in the knowledge graph to calculate the cosine similarity.

[0091] If the cosine similarity is greater than the similarity threshold, the candidate entity is determined to be an alias of an existing entity, and the semantic information and link records of the existing entity are updated.

[0092] If the cosine similarity is less than or equal to the similarity threshold, the candidate entity is identified as a new entity, a new entity is added to the knowledge graph, and the set of existing entities that co-occur with the new entity is extracted. The set of existing entities that co-occur with the new entity includes multiple co-occurring entities.

[0093] Specifically, when the system encounters the same unlinked entity mention in multiple decisions—that is, when it cannot correspond to an entity in the existing knowledge graph—the system will attempt to create a new entity. First, it extracts… The count value exceeds the threshold. The mention history is as follows: (49) Each unlinked mention of text is encoded into a 768-dimensional semantic vector using a pre-trained Sentence-BERT model: (50) in This represents the combined text of each mention and its context information, for each mention. Multiple combined texts generate multiple semantic vectors. The average value is used to obtain the unified semantic representation of the mention: (51) Next, to avoid creating duplicate entities, the system will Compare with the semantic vectors of all existing entities in the knowledge graph (KG) to calculate With each existing entity semantic vector Cosine similarity: (52) If any existing entity exists Make , If the value is 0.85, it is considered that the reference is likely an alias or variant of an existing entity. Instead of creating a new entity, the link record of the existing entity is updated.

[0094] If the cosine similarity of entities does not exceed the threshold Then the unified semantic representation mentioned Define as a new entity node It is added to the knowledge graph.

[0095] For each co-entity, a BERT-based pre-trained relation extraction model is used to predict the relation type and relation confidence score between the new entity and the co-entity. The frequency of the relation in the historical context is also counted. The relation confidence score and the frequency of occurrence are used to calculate the overall confidence score. When the overall confidence score exceeds the confidence threshold, a new relation between the new entity and the co-entity is established in the knowledge graph.

[0096] Specifically, creating new entities Next, the system needs to establish relationships between itself and existing entities in the knowledge graph. From All contexts In, extraction and A set of existing entities that frequently co-occur For each co-real entity Collect context statements that contain both of them.

[0097] For each entity pair The model uses a BERT-based pre-trained relation extraction model to predict relation types. This model is pre-trained on a general relation extraction dataset and fine-tuned on domain-specific data. The model outputs relations. Relationship type and its confidence score The predefined set of domain relationships in this system includes: set relationships such as belonging to, component of, and originating from; spatial relationships such as being located at, near, and far from; causal relationships such as causing, inducing, and influencing; and temporal relationships such as occurring before or simultaneously with.

[0098] For each predicted relation The frequency of system statistical relations appearing in the context: (53) in For relationship The number of times it appears in the co-occurrence context. For entities and The system requires the total number of times they co-occur. The relationship is considered valid only if more than 50% of the co-occurring context supports it; otherwise, it is not accepted. The overall confidence level is calculated by combining the model prediction confidence and the statistical validation frequency. (54) in When the overall confidence level exceeds a threshold, the system adds new relationships to the knowledge graph. and triplet .

[0099] 3. Evolutionary Coordinator The evolution coordinator is used to coordinate the execution order of each module in the autonomous evolutioner and execute the co-evolution protocol: first, the knowledge evolutioner is started to update the knowledge graph; then, the model evolutioner is started to evolve and update the model parameters using the updated knowledge graph; then, the policy evolutioner is started to evolve and update the parameters of the policy network and value network using the updated knowledge graph and updated model parameters; and after each co-evolution is completed, in an offline verification environment, a set of standard test cases are used to evaluate the comprehensive performance score of the evolved system, and based on the comprehensive performance score of the evolved system, the severity of the conflict is determined and a graded rollback is performed.

[0100] To resolve potential conflicts arising from the parallel or sequential execution of multiple evolvers, this application designs a top-level evolution coordinator. It receives evolution trigger instructions and diagnostic information from the monitoring and evaluator, and formulates an ordered, cooperative evolution execution plan. The cooperative evolution protocol is designed as follows: The coordinator first initiates the knowledge evolver. After knowledge evolution is complete, the system adds the newly added entity set to the knowledge graph. New set of relations and the newly added fact triplet The set. Then the model evolver starts, and its loss function includes the cognitive context vectors from the historical data stream. The updated knowledge graph will be used for regeneration. After the model evolver updates its parameters, the policy evolver starts. The cognitive context vectors used in policy evolution require a forward propagation recalculation using the updated model parameters. The PPO algorithm utilizes Update the policy network and value network.

[0101] At the same time, the coordinator maintains a system performance baseline. This includes baseline scores, a knowledge graph, and model parameters. After each co-evolution, the system is used in an offline verification environment to quickly evaluate the overall performance score of the new version using a set of standard test cases. ,like ,in If the threshold is reached, a conflict is identified. Based on the severity of the conflict, the coordinator performs a tiered rollback: if the score drops by 5%-10%, only the policy network parameters are rolled back; if the score drops by 10%-20%, both the policy and model evolver parameters are rolled back; if the score drops by more than 20%, a complete rollback to baseline B is performed. Furthermore, all conflict events and contextual information are recorded in the evolution log for analysis and optimization.

[0102] By introducing an evolution coordinator and phased coordination protocols, this system ensures the orderliness and consistency of the evolution process, avoiding internal system conflicts and performance degradation caused by asynchronous module updates. Offline verification and rollback mechanisms form a safety net, making the complex autonomous evolution process more robust and reliable.

[0103] This application provides an embodiment of an autonomous evolution method for a large-scale intelligent decision-making model that integrates an OODA loop, comprising: The inner decision loop receives heterogeneous data from multiple sources and outputs situation tensors, action vectors, and control command vectors to control the actuators. The inner decision loop includes an observation module, a judgment module, a decision module, and an action module.

[0104] The outer autonomous evolution loop obtains the moving average, average response time, and task completion rate based on the action vector and the preset optimal action vector under the same situation, and then obtains the comprehensive performance score; it is also used to obtain the uncertainty based on the current situation tensor and the historical situation tensor; based on the comprehensive performance score and uncertainty, it determines whether to perform autonomous evolution update on the inner decision loop; if it is determined to perform autonomous evolution update on the inner decision loop, the parameters of the observation module, judgment module, and decision module are updated, and the knowledge graph in the judgment module is updated.

[0105] In this optional embodiment, a dual-loop collaborative mechanism addresses the limitations of traditional static decision-making systems. The inner decision loop focuses on real-time response, ensuring a closed loop from situational awareness to action commands is completed within milliseconds. The outer autonomous evolution loop, through periodic or event-driven evaluation, identifies performance degradation or environmental novelty, thereby driving iterative optimization of model parameters and knowledge structures. For example, when the system encounters a new situation not fully covered in the historical situation tensor set, the uncertainty index increases, triggering the expansion of the knowledge graph and model fine-tuning, integrating new entities, relationships, and related features into the decision-making system. This design not only avoids the adaptability defects caused by the solidification of core parameters but also ensures that updates are performed only when necessary through a joint determination mechanism of performance scoring and uncertainty, balancing system stability and evolutionary efficiency. By constructing the organic linkage between the inner decision loop and the outer autonomous evolution loop, a paradigm shift from "static deployment" to "dynamic evolution" in intelligent decision-making systems is fundamentally achieved. While maintaining the rapid iteration advantage of the OODA loop, the system can continuously improve decision-making accuracy based on actual operating data, effectively overcoming the performance degradation problem caused by the inability to learn autonomously in new situations in the background technology, and significantly enhancing the long-term robustness and task adaptability of the intelligent decision-making system in complex dynamic environments.

[0106] The system in this application, through dynamic interaction between two stages, enables the system to review, learn, and innovate from long-term decision-making practices, ultimately overcoming the fundamental defects of existing technologies such as model rigidity and reliance on manual intervention.

[0107] Scenario 1: Spacecraft Space Debris Avoidance A certain satellite in orbit needs to perform long-term missions in an environment with multiple pieces of space debris. The satellite needs to make autonomous decisions to avoid maneuvers, ensure mission safety and save fuel to the maximum extent. The requirements for the real-time nature of decision-making, safety and fuel efficiency are extremely high.

[0108] The process of applying this application is as follows: The satellite's inner OODA decision loop continues to operate, and the observation module obtains satellite orbital parameters and debris orbital parameters through sensors to generate a situation tensor. The judgment module calls a large model to analyze orbital data and combines it with a knowledge graph to determine the fragmentation threat level context vector. The decision module outputs decision instructions through the policy network. The action module converts the instructions into on / off commands and thrust direction for the satellite thrusters, executing satellite maneuvers. The external environment module simulates the real-world environmental changes after the satellite maneuvers and uses a reward function. Provide feedback on the successful completion of this evasion operation.

[0109] After the maneuver is executed, the outer autonomous evolutionary loop continuously monitors the system. When the number of fragments increases sharply or an unexpected proximity event occurs, the outer evolutionary loop's monitoring and evaluator determines that the decision-making effectiveness score has decreased, triggering autonomous evolution. At this point, the autonomous evolutionary engine starts, and the model evolutionary engine uses this decision data stream to fine-tune the parameters of the large model network in the observation and judgment modules through online learning, enabling it to better predict fragment movement trends. The knowledge evolutionary engine transforms successful avoidance experience into new rules and injects them into the knowledge graph of the judgment module. The policy evolutionary engine generates better strategies by optimizing the weights of the decision network. Through autonomous evolution, when fragments encounter more complex fragment scenarios, the OODA decision loop can complete the observation, judgment, decision, and action cycle more quickly, improving the system's accuracy and security.

[0110] This application's solution, through its outer evolutionary loop, learns from each interaction with fragments, continuously updating its internal OODA decision-making model. This results in continuously improved performance when facing new fragments or complex combinations of scenarios, solving the problem of model solidification. The system, through a knowledge evolver, can automatically abstract new avoidance rules from successful or failed avoidance cases and autonomously update them in the knowledge graph, significantly reducing reliance on manual data analysis and knowledge updates by ground station experts. This addresses the issue of high operational costs. Furthermore, the outer evolutionary layer of this application can comprehensively monitor the overall performance of the inner OODA loop and coordinate the three evolvers for targeted optimization, solving the problem of limited model collaboration capabilities.

[0111] Implementation Scenario 2: Intelligent connected vehicles passing through intersections with mixed traffic When an autonomous vehicle navigates a complex urban intersection without traffic lights, it must respond in real time to unpredictable targets such as pedestrians suddenly crossing the road and vehicles violating traffic rules, posing challenges to the accuracy and adaptability of its decision-making.

[0112] The process of applying this application is as follows: The car's inner decision loop acquires the states of all traffic participants at the intersection in real time and generates a situation tensor. If a pedestrian is observed crossing the road ahead, or a vehicle is merging from the left, the system will then output data using a large model and knowledge graph. Information such as "The pedestrian's intention is uncertain; according to traffic rules, you should stop and yield to vehicles on the left" is then used by the policy network to generate emergency braking driving decision information. The action module, acting on instructions, brought the vehicle to an abrupt stop. While this prevented an accident, it caused passenger discomfort and nearly resulted in a rear-end collision. Following this decision-making cycle, the outer evolutionary cycle assessed a high safety score but extremely low comfort and efficiency scores, resulting in a low overall performance score. If the target is not met, the evolver is triggered. The model evolver fine-tunes the CNN in the observation module, enabling it to recognize pedestrians' subtle movements as they prepare to cross the road earlier and more accurately. The knowledge evolver summarizes new knowledge: "At certain intersections, pedestrians often weave through the gaps between vehicles, requiring advance prediction," and updates the knowledge graph. The policy evolver optimizes the policy network, learning smoother braking strategies such as early and gentle deceleration. After this evolution, the system's inner-layer decision OODA loop will better handle scenarios like mixed-traffic intersections.

[0113] This proposed solution can continuously learn from real-world traffic scenarios, autonomously transforming experience in handling complex road conditions into internal knowledge. Through evolution, it continuously improves driving strategies, greatly enhancing adaptability and reducing maintenance costs. It evaluates the overall performance of each passage from a global perspective and unifies the collaborative evolution of command modules, achieving system-level performance optimization and effectively addressing the shortcomings of existing technical solutions.

[0114] In summary, compared with existing technologies, it has the following beneficial effects: 1. The system achieves lifelong learning and autonomous evolution. It can learn from continuous decision-making practices, automatically adjust and optimize itself, thereby effectively responding to dynamically changing environments and overcoming the drawbacks of existing system models being rigid and performance degrading.

[0115] 2. Significantly reduced system operation and maintenance costs and reliance on human experts. By automating processes such as knowledge updates and model optimization, the system gains self-maintenance and growth capabilities, overcoming the predicament of traditional systems requiring frequent manual upgrades.

[0116] 3. It enhances the system's overall coordination and intelligence. Due to the existence of an overarching, self-evolving evolutionary loop, the modules within the system are no longer isolated functional units, but rather organic components that can be collaboratively optimized, resulting in more efficient and robust overall decision-making capabilities.

[0117] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0118] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop, characterized in that, This includes an inner decision-making loop and an outer autonomous evolution loop; The inner decision loop is used to receive multi-source heterogeneous data and output situation tensors, action vectors, and control command vectors to control the actuators; the inner decision loop includes an observation module, a judgment module, a decision module, and an action module. The outer self-evolving ring includes: The monitoring and evaluation system is used to obtain the moving average, average response time, and task completion rate based on the action vector and the preset optimal action vector under the same situation, and then obtain the comprehensive performance score; it is also used to obtain the uncertainty based on the current situation tensor and the historical situation tensor; and based on the comprehensive performance score and uncertainty, it determines whether to perform an autonomous evolution update on the inner decision loop. An autoevolver is used to update the parameters of the observation module, judgment module, and decision module, and update the knowledge graph in the judgment module, if an autoevolving update is performed on the inner decision loop.

2. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 1, characterized in that, The observation module is used to fuse multi-source heterogeneous data to generate a situation tensor; The judgment module is used to obtain the cognitive context vector based on the input text, the situation tensor, and the knowledge graph; The decision-making module is used to obtain action vectors and state value scalars based on cognitive context vectors; The action module is used to convert action vectors into control command vectors for the actuators; The external environment module is used to simulate the interaction between the system and the external environment. It updates the environmental state of the external environment module according to the control command vector and receives immediate rewards.

3. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 2, characterized in that, The process of obtaining the cognitive context vector based on the input text, the situation tensor, and the knowledge graph includes: The situation tensor is input into the fully connected layer to obtain the situation embedding vector; Based on the input text and the situational embedding vector, the primary cognitive embedding vector is obtained; One or more entity mentions are extracted from the primary cognitive embedding vectors. Each entity mention contains at least the entity text and its corresponding context information. The text string of the entity mention is then input into the Sentence-BERT model to obtain the text semantic information. Input the name and alias of each entity in the knowledge graph into the Sentence-BERT model to obtain name semantic information; Analyze the cosine similarity between text semantic information and name semantic information, and select graph entities with a cosine similarity greater than a preset similarity as candidate graph entities to obtain a set of candidate graph entities; The link confidence score is obtained based on the context sentence of each candidate graph entity obtained from the entity mention information and the attribute field text of the candidate graph entity in the knowledge graph. If the highest link confidence score is greater than the preset confidence score, the corresponding graph entity is recorded as the final graph entity, the entity mention is linked to the final graph entity, and marked as linked; otherwise, the entity mention is marked as unlinked, and the number of times the entity mention is unlinked is recorded. Using the successfully linked set of graph entities as the query seed, a two-hop traversal query algorithm based on breadth-first search is adopted to extract triples related to entity mentions from the knowledge graph, and the two-hop search results are merged to obtain a global knowledge subgraph. The cognitive context vector is obtained by fusing the primary cognitive embedding vector and the global knowledge subgraph.

4. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 2, characterized in that, The process of obtaining the action vector and state value scalar based on the cognitive context vector includes: The cognitive context vector is input into the multilayer perceptron within the policy network, and the output layer of the policy network outputs the action probability distribution parameters. Based on the action probability distribution parameters and the log-standard deviation parameter vector, an ε-greedy exploration strategy is used to sample and obtain action vectors; The cognitive context vector is input into the multilayer perceptron within the value network, and the output layer of the value network outputs a state value scalar that evaluates the value of the current state.

5. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 1, characterized in that, The moving average is the average of the decision accuracy rates of a preset number of times prior to the current moment; The decision accuracy rate is in, Represents the action vector. This indicates the preset optimal action. Indicates the magnitude error weight. Represents the decision accuracy at time t; The average response time is the average time delay from the observation module outputting the situation tensor to the action module issuing the corresponding control command vector, calculated from a preset number of times prior to the current moment. The task completion rate is the success rate of achieving the preset task objective within the second preset number of task cycles prior to the current moment. The overall performance score is the weighted sum of the moving average, the reciprocal of the mean response time, and the task completion rate.

6. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 1, characterized in that, The uncertainty is , in, Indicates uncertainty. The state tensor at time t The estimated probability density appearing in the probability distribution described by the set of historical situation tensors The state tensor at time t represents the state tensor. Let L represent the i-th historical situation tensor in the historical situation tensor set, L represent the total number of tensors in the historical situation tensor set, and h represent the bandwidth. This is the kernel function.

7. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 1, characterized in that, The autonomous evolver includes a model evolver, a policy evolver, and a knowledge evolver; If it is determined that the inner decision loop should be automatically evolved and updated, then the parameters of the observation module, judgment module, and decision module should be updated, and the knowledge graph in the judgment module should be updated, including: If it is determined that the inner decision loop should be automatically evolved and updated, the model evolver will minimize the joint loss function of the observation module and the decision module, and use the online stochastic gradient descent algorithm to update the task-specific neural network parameters of the observation module and the decision module in the inner decision loop. The policy evolver performs collaborative optimization of the policy network and value network of the decision module based on the total loss of the decision module; where the total loss is the weighted sum of the policy network loss and the value network loss. The policy network loss is in, Represents the dominance function. Indicates the discount factor. This represents the GAE smoothing parameter. Let t represent the state value scalar at time t. This represents the state value scalar at time t+1. Indicates an immediate reward. This represents the single-step residual at time t. express The single-step residual at each moment; This represents the probability of taking action 'a' in the new policy network given a cognitive context vector. This represents the probability of taking action 'a' in the old policy network, given a cognitive context vector. This represents the probability ratio between the old and new policy networks; To truncate hyperparameters; Value network loss is in, This represents the state value scalar obtained under the old value network. This represents the scalar value of the state obtained under the new value network. Indicates the target value; The knowledge evolver automatically updates the knowledge graph in the judgment module based on the system's decision-making experience.

8. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 3, characterized in that, The information and number of times the entity mentioned is marked as unlinked are stored in the entity candidate pool. If the same entity mention appears again in a subsequent decision and is marked as unlinked, the number of times it is unlinked is accumulated. The knowledge evolver automatically updates the knowledge graph in the judgment module based on the system's decision-making experience, including: When a trigger command signal is received or the number of decision loops reaches a preset number, information mentioned by candidate entities in the entity candidate pool whose number of unlinked times is greater than a preset unlinked time threshold is extracted. Based on the information mentioned by the candidate entity, the semantic vector of the candidate entity is obtained using the pre-trained Sentence-BERT model, and compared with the semantic vector of the existing entity in the knowledge graph to calculate the cosine similarity. If the cosine similarity is greater than the similarity threshold, the candidate entity is determined to be an alias of an existing entity, and the semantic information and link records of the existing entity are updated. If the cosine similarity is less than or equal to the similarity threshold, the candidate entity is determined to be a new entity, the new entity is added to the knowledge graph, and the set of existing entities that co-occur with the new entity is extracted. The set of existing entities that co-occur with the new entity includes multiple co-occurring entities. For each co-entity, a BERT-based pre-trained relation extraction model is used to predict the relation type and relation confidence score between the new entity and the co-entity, and the frequency of the relation in the historical context is counted. The overall confidence score is calculated using the relation confidence score and the frequency of occurrence. When the overall confidence score exceeds the relation confidence threshold, a new relation between the new entity and the co-entity is established in the knowledge graph.

9. The intelligent decision-making large-scale model autonomous evolution system integrating the OODA loop as described in claim 7, characterized in that, It also includes an evolution coordinator, which coordinates the execution order of the modules in the autonomous evolutionary unit and executes a co-evolution protocol: first, the knowledge evolutionary unit is started to update the knowledge graph; then, the model evolutionary unit is started to evolve and update the model parameters using the updated knowledge graph; then, the policy evolutionary unit is started to evolve and update the parameters of the policy network and the value network using the updated knowledge graph and the updated model parameters; and after each co-evolution is completed, in an offline verification environment, a set of standard test cases are used to evaluate the overall performance score of the evolved system, and based on the overall performance score of the evolved system, the severity of the conflict is determined and a graded rollback is performed.

10. An autonomous evolution method for intelligent decision-making large-scale models integrating the OODA loop, characterized in that, include: The inner decision loop receives heterogeneous data from multiple sources and outputs situation tensors, action vectors, and control command vectors to control the actuators; the inner decision loop includes an observation module, a judgment module, a decision module, and an action module. The outer autonomous evolution loop obtains the moving average, average response time, and task completion rate based on the action vector and the preset optimal action vector under the same situation, and then obtains the comprehensive performance score; it is also used to obtain the uncertainty based on the current situation tensor and the historical situation tensor; based on the comprehensive performance score and uncertainty, it determines whether to perform autonomous evolution update on the inner decision loop; if it is determined to perform autonomous evolution update on the inner decision loop, the parameters of the observation module, judgment module, and decision module are updated, and the knowledge graph in the judgment module is updated.