Dynamic arrangement and continuous learning knowledge inquiry intelligent agent generation method and system

CN122472217BActive Publication Date: 2026-09-11STATE GRID JIANGSU ELECTRIC POWER CO LTD MARKETING SERVICE CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610944134.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-11
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

[0006]本发明旨在解决现有知识问询智能体在处理复杂电力业务时依赖静态预设规则导致任务编排灵活性不足,以及在增量学习新知识时因缺乏底层参数保护机制而导致原有核心领域逻辑发生灾难性遗忘的技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472217B_ABST
    Figure CN122472217B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for generating knowledge query agents through dynamic orchestration and continuous learning. The method utilizes a finite state machine to constrain a large language model, decomposing natural language queries into sequences of atomic operations. Next, it constructs a voltage level adaptation factor and selection neural network, instantiates a toolset, and dynamically generates a parallel directed acyclic graph for task execution using a reinforcement learning strategy network that integrates knowledge from the power sector. Then, it concurrently schedules task nodes based on topological dependencies and verifies multi-source intermediate results through a time-stamp alignment mechanism, fusing them to generate the final solution. Finally, it employs a parameter consolidation strategy based on domain feature weighting, locking core parameters through word frequency and topological betweenness features to achieve continuous incremental learning of the model. This solves the problems of existing technologies relying on static rules for task orchestration and catastrophic forgetting caused by incremental learning, improving the agent's logical reasoning ability and long-term evolutionary stability in complex power business scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method and system for generating knowledge query agents through dynamic orchestration and continuous learning. Background Technology

[0002] With the rise of large language models, knowledge query agents have shown great application potential in scenarios such as enterprise knowledge management, intelligent customer service, and scientific research assistance. Especially in vertical fields such as power and industry, agents not only need to answer general questions, but also need to handle complex business tasks involving multi-step reasoning, conditional branch judgment, and multi-source data retrieval.

[0003] In existing technologies, some methods for constructing knowledge-questioning agents involve building attention fusion networks to calculate the semantic relevance between task instructions and entities in the knowledge graph, supplemented by reinforcement learning modules for dynamic fine-tuning of execution actions. However, these methods essentially still generate a linear or semi-linear sequence of execution actions. When facing complex power business scenarios, this linear structure lacks flexibility, and simple action fine-tuning is insufficient to reconstruct the global path, thus limiting the ability to handle complex logical tasks. On the other hand, some methods use a multi-agent coordination model, distributing questions to the most suitable sub-module. The logic is relatively static, and although there is a coordination module, it is mainly based on preset rules or distribution logic. This distribution architecture lacks the ability to dynamically reconstruct the task topology and cannot, like reinforcement learning-driven dynamic orchestration, adjust the directed acyclic graph structure of subsequent tasks in real time based on feedback at each step during task execution, making it difficult to cope with dynamically changing business needs.

[0004] Furthermore, in terms of the long-term evolution and maintenance of intelligent agents, existing strategies often focus on the static expansion of external knowledge bases (such as knowledge graphs) or the absorption of new knowledge through simple full-parameter fine-tuning. However, due to the lack of protection mechanisms for the underlying parameters of the model, catastrophic forgetting can easily be induced when the intelligent agent performs incremental learning to adapt to new scenarios and rules. That is, in the process of acquiring new knowledge, the model will experience a significant degradation of its original core domain knowledge and logical processing capabilities, seriously hindering the long-term stability and improvement of the intelligent agent's intelligence level.

[0005] In summary, how to construct a knowledge-questioning agent that possesses both the ability to dynamically orchestrate complex tasks and the capacity to simultaneously absorb new knowledge and consolidate old knowledge during continuous learning is a technical problem that urgently needs to be solved in the current technological field. Summary of the Invention

[0006] This invention aims to address the technical problems of existing knowledge-based query agents, which suffer from insufficient flexibility in task orchestration due to reliance on static preset rules when handling complex power business, and catastrophic forgetting of core domain logic due to the lack of underlying parameter protection mechanisms during incremental learning of new knowledge. The invention includes: first, using a finite state machine to constrain a large language model, decomposing natural language queries into sequences of atomic operations; second, constructing a voltage level adaptation factor and selection neural network, instantiating a toolset, and dynamically generating a parallel directed acyclic graph for task execution using a reinforcement learning strategy network that integrates power domain knowledge; third, concurrently scheduling task nodes based on topological dependencies, verifying multi-source intermediate results through a time-stamp alignment mechanism, and fusing them to generate the final solution; and finally, employing a domain feature-weighted parameter consolidation strategy, locking core parameters through word frequency and topological betweenness features to achieve continuous incremental learning of the model. This solves the problems of existing technologies relying on static rules for task orchestration and catastrophic forgetting caused by incremental learning, improving the logical reasoning ability and long-term evolutionary stability of the agent in complex power business.

[0007] The present invention adopts the following technical solution.

[0008] In a first aspect, the present invention provides a method for generating a knowledge query agent with dynamic orchestration and continuous learning, including: The system receives natural language queries from users, performs intent recognition and task decomposition using a large language model, and introduces a finite state machine to mask and constrain the original logic vectors obtained from the large language model during the task decomposition process to obtain a sequence of atomic operations. Based on the atomic operation sequence, the voltage level adaptation factor and context state vector are calculated to select the tool instantiation, resulting in an instantiated atomic operation set; according to the instantiated atomic operation set, a reinforcement learning policy network that integrates power domain knowledge is used to predict topology connection actions, resulting in a parallel directed acyclic graph. Based on topological dependencies, the task nodes in the parallel directed acyclic graph are concurrently scheduled and executed to capture intermediate results from multiple sources; the intermediate results from multiple sources are then time-stamped and logically fused to generate the final business solution.

[0009] Preferably, in S1, the introduction of a finite state machine to mask the original logical vector predicted by the large language model at each generation step specifically includes: A predefined set of states and valid transition paths between states; the set of states includes at least an intent-locked state, an atomic operation selection state, a parameter key name state, a parameter key value state, and an end state; the valid transition paths include: a unidirectional transition from the intent-locked state to the atomic operation selection state; an attribute-filling transition from the atomic operation selection state to the parameter key name state, and then to the parameter key value state; and a cyclic decision transition from the parameter key value state back to the parameter key name state, back to the atomic operation selection state, or to the end state based on parameter integrity. During the decoding process, the current state of the finite state machine is determined based on the generated lexical sequence, and the set of legal lexical terms for the next time step is determined based on the current state. Specifically, if the current state is the intent-locked state, the set of legal lexical terms is limited to intent identifier lexical terms in the predefined intent category set; if the current state is the atomic operation selection state, the corresponding atomic operation name lexical terms are determined as the set of legal lexical terms based on the previously locked intent categories; if the current state is the parameter key name state or the parameter key value state, the corresponding parameter key name or parameter key value lexical terms are determined as the set of legal lexical terms based on the previously selected atomic operations.

[0010] Preferably, in S1, the step of introducing a finite state machine to mask the original logical vector predicted by the large language model at each generation step further includes: constructing a mask vector with the same dimension as the full vocabulary of the large language model; setting the values ​​of the index positions in the mask vector corresponding to the index positions outside the legal word set to negative infinity, and setting the values ​​of the index positions corresponding to the index positions of the legal word set to zero, based on the determined set of legal words; superimposing the mask vector onto the original logical vector output by the large language model at the current time to obtain a corrected probability distribution, and sampling based on the corrected probability distribution to generate the current word.

[0011] Preferably, in S2, the voltage level adaptation factor and context state vector are calculated for tool instantiation selection, specifically including: Calculate the conditional suitability score between the current atomic operation and the candidate tool, the conditional suitability score being determined by the product of the semantic matching degree of the selection neural network output and the voltage level adaptation factor; The selection neural network concatenates the embedding vector of the current atomic operation, the embedding vector of the candidate tool, and the context state vector that records the historical operation path to output the semantic matching degree. The voltage level adaptation factor is generated based on the matching relationship between the rated voltage level and device category of the device involved in the current atomic operation and the applicable scope declared by the candidate tool. When the device attributes of the current atomic operation exceed the applicable scope of the candidate tool, the weight of the voltage level adaptation factor is reset to zero or a preset minimum value to reduce the probability of the candidate tool being selected through a multiplication gating mechanism.

[0012] Preferably, in S2, the step of predicting topology connection actions using a reinforcement learning policy network that incorporates knowledge from the power domain, based on the instantiated set of atomic operations, specifically includes: The process of predicting topological connection actions is modeled as a Markov decision process, and the state vector at the current moment is defined, including: the currently constructed local directed acyclic graph structure, the remaining sequence of atomic operations to be placed consisting of the set of instantiated atomic operations, and the current context state vector. A graph neural network is used to extract the topological features of the local directed acyclic graph structure, and a Transformer encoder is used to extract the semantic features of the remaining atomic operation sequences to be placed. The two are then fused and input into the reinforcement learning policy network that incorporates knowledge from the power domain, and the output is the prediction result of the predecessor node connection of the current atomic operation in the local directed acyclic graph. During the training phase of the reinforcement learning policy network that integrates knowledge from the power sector, a composite reward function based on power rules is introduced to update the network. The composite reward function includes: a constraint matching reward based on the matching of tool and equipment voltage levels, a gain reward based on the current context state vector, and a cost penalty based on the number of execution steps.

[0013] Preferably, in S3, task nodes are concurrently scheduled and results are fused based on topological dependencies, including executing the following rules: The topology concurrent scheduling rule is used to monitor the in-degree state of each task node in the parallel directed acyclic graph in real time, add task nodes with an in-degree of zero to the ready queue, and use an asynchronous concurrent thread pool to trigger the task nodes in the ready queue in parallel; after the task node finishes execution, update the in-degree state of the downstream child node, and store the output data of each task node as the multi-source intermediate result in the context state vector. The time stamp alignment verification rule is used to obtain the power data section timestamp corresponding to the multi-source intermediate results; calculate the time deviation between the multi-source intermediate results participating in the same business logic calculation; if the time deviation exceeds a preset threshold, the current data is determined to be invalid, and a data re-sampling and rescheduling instruction is triggered for the task node that generated the multi-source intermediate results. The logical fusion generation rule is used to extract the multi-source intermediate results from the context state vector as an evidence chain after all task nodes have been executed, dynamically assemble the prompt word template by combining it with the preset power industry constraint criteria, and input the prompt word template into the large language model to generate the final business answer in natural language form.

[0014] Preferably, the parameter consolidation strategy based on domain feature weighting specifically includes: Construct a loss function that includes an elastic weight consolidation penalty term, and introduce a domain correction factor for the parameters of the large language model into the elastic weight consolidation penalty term; The inverse document frequency features of words in the power document set are statistically analyzed. The betweenness centrality features of equipment nodes in the power grid topology in the power equipment knowledge graph are calculated. The product of the inverse document frequency features and the normalized betweenness centrality features is determined as the business importance score of the words. Locate the row vector index corresponding to the core electricity vocabulary in the input embedding layer of the large language model, assign the business importance score to the domain correction factor of the parameter corresponding to the row vector, and keep the domain correction factors of other parameters at their default values.

[0015] Secondly, this invention provides a knowledge query intelligent agent generation system with dynamic orchestration and continuous learning, comprising: The semantic parsing and task decomposition module is used to receive user natural language queries, perform intent recognition and task decomposition using a large language model, and introduce a finite state machine to mask the original logic vector obtained by the large language model during the task decomposition process to obtain an atomic operation sequence. The dynamic topology orchestration module is used to calculate the voltage level adaptation factor and context state vector based on the atomic operation sequence to select tools for instantiation, thereby obtaining an instantiated atomic operation set; based on the instantiated atomic operation set, a reinforcement learning policy network that integrates power domain knowledge is used to predict topology connection actions, thereby obtaining a parallel directed acyclic graph. The task execution and result integration module is used to concurrently schedule and execute task nodes in the parallel directed acyclic graph based on topological dependencies, capture multi-source intermediate results, perform time-stamp alignment verification and logical fusion on the multi-source intermediate results, and generate the final business solution.

[0016] Thirdly, the present invention provides a terminal, including a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method.

[0017] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method.

[0018] The beneficial effects of this invention are as follows: (1) This invention utilizes a reinforcement learning strategy network that integrates knowledge in the power field to transform tasks into parallel directed acyclic graphs, automatically mines implicit dependencies between subtasks, and performs parallel scheduling of dependent nodes through an asynchronous concurrent thread pool, thereby reducing the total response latency when processing complex power business and improving the system's execution efficiency and throughput.

[0019] (2) This invention establishes a rigorous constraint mechanism in the power field. Through voltage level adaptation factor and multiplication gating mechanism, tools with mismatched voltage levels or equipment types are forcibly filtered out, thus avoiding logical fallacies and potential safety risks from the source. At the same time, time stamp alignment verification rules are introduced to automatically verify the data section timestamp when integrating multi-source intermediate results. This effectively solves the calculation error caused by the asynchronous data of SCADA system in the power scenario and ensures the accuracy of the final business solution.

[0020] (3) This invention proposes a parameter consolidation strategy based on domain feature weighting. By combining text statistical features and power grid topology features, it accurately identifies core words that play a pivotal role in power business and locks their corresponding underlying model parameters through embedded layer fixed-point mapping technology. This enables the agent to retain its understanding of the original core power logic to the maximum extent while absorbing new regulations and equipment knowledge, overcoming the catastrophic forgetting problem that is easily caused by traditional fine-tuning methods and ensuring the performance stability of the agent in the long-term evolution process.

[0021] (4) In the task decomposition stage, the present invention introduces a finite state machine to mask the large language model, which forces the model to output atomic operation sequences that conform to strict syntax rules, eliminating the common illusion and format error problems of large models, so that the generated execution plan can be seamlessly connected to industrial-grade business systems, and improving the usability of intelligent agents in actual engineering implementation. Attached Figure Description

[0022] Figure 1 This is an overall flowchart of the knowledge query agent generation method with dynamic arrangement and continuous learning provided by the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this application are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.

[0024] Example 1: like Figure 1 As shown, this invention provides a method for generating knowledge query agents through dynamic orchestration and continuous learning, including: S1: Receive user natural language queries, use a large language model for intent recognition and task decomposition, and introduce a finite state machine to mask the original logic vector obtained from the large language model during the task decomposition process to obtain an atomic operation sequence.

[0025] In this step, a Large Language Model (LLM) is used as the core engine for intent recognition and task decomposition. Using the pre-defined set of user intent categories in the prompt, and leveraging the semantic understanding capabilities of the LLM, intent recognition is performed on the natural language queries input by the user.

[0026] Specifically, the model first performs deep semantic analysis on the user's natural language query to accurately capture its core needs. When the user's question cannot be completed through a single step, the query is decomposed into a series of logically continuous and executable atomic operations of n. Each subtask It corresponds to the smallest unit of operation that can be completed by a single function call or interface call.

[0027] To ensure the executability of the atomic operation sequence and the stability of the generation process, this embodiment employs a task generation strategy that combines a predefined atomic operation library with finite state machine (FSM) constraint decoding. Specifically, this includes: (1) Predefine the set of intent categories Examples include numerical computation and information retrieval; predefined sets of atomic operations. Examples include querying database interfaces, performing mathematical operations, and accessing knowledge bases. Different intent categories are forcibly mapped to corresponding subsets of atomic operation candidates, ensuring that, given an intent, only atomic operations matching that intent are allowed to be generated. Based on this, a state node is constructed containing the following: initial state This is used to represent the initial state of sequence generation; the intention is to lock the state. This is used to generate and lock the intent category corresponding to the user query, and its value is limited to the set of intent categories. In the middle; atomic operation selection state This is used to generate specific atomic operation identifiers, whose values ​​must belong to the set of atomic operations corresponding to the current intent category. ; Parameter name status With parameter value state These are used to generate the parameter key names and parameter key values ​​for atomic operations, respectively; End state , is used to indicate that the generation of the atomic operation sequence is complete.

[0028] To support scenarios with multiple parameters in a single action within complex instructions, the legal transition paths include: a unidirectional transition from the intent-locked state to the atomic operation selection state; an attribute-filling transition from the atomic operation selection state to the parameter key name state, and then to the parameter key value state; and a cyclic decision transition from the parameter key value state back to the parameter key name state, back to the atomic operation selection state, or to the end state based on parameter integrity. The legal transition paths of the state machine are defined as follows:

[0029] Among them, when completing a parameter value After filling, FSM checks the completeness of the parameters according to the predefined atomic operation parameter template: if there are still unassigned parameters, it returns to the parameter name status. If all parameters for the current atomic operation have been generated and there are subsequent atomic operations, then return to the atomic operation state. If the atomic operation sequence has been generated, then transition to the end state. .

[0030] (2) During the decoding process of LLM generating tokens one by one, the current state of FSM is maintained in real time. If, in a certain generation step, the probability distribution of all candidate tokens output by LLM tends to zero after masking constraints (i.e., the model cannot generate valid outputs within the valid set), an exception is thrown and the current generation process is terminated, and the FSM state is reinitialized to ensure the robustness of the generation. In each generation step, the current state of FSM is determined based on the history of the generated token sequences, and then the set of valid tokens (L) for the next moment is locked according to the preset hierarchical constraint relationship: If currently in intent-locked state :L contains only intent identifier terms from a predefined set of intent categories.

[0031] If currently in an atomic operation state Based on the intent category locked in the previous steps (e.g., "numerical calculation"), retrieve the allowed operation subset under that category, and use the term corresponding to the operation name in that subset as L.

[0032] If the current state is parameter key name / key value ( / ): Based on the currently selected atomic operation definition, the token corresponding to the valid parameter name or parameter value range of the operation is taken as L.

[0033] (3) In order to force the model to sample only in the above-mentioned set of legal lexical units L, a mask vector with the same dimension as the full vocabulary of the large language model is constructed; according to the determined set of legal lexical units, the values ​​of the index positions in the mask vector corresponding to the index positions outside the set of legal lexical units are set to negative infinity, and the values ​​of the index positions corresponding to the set of legal lexical units are set to zero; the mask vector is superimposed on the original logical vector output by the large language model at the current time to obtain the corrected logical vector; the corrected logical vector is normalized to obtain the corrected probability distribution, and sampling is performed based on the corrected probability distribution to generate the current lexical unit.

[0034] Specifically, this embodiment introduces a masking mechanism at the model's output layer. Assume the original logical vector output by the LLM at the current time step is... Where V is the size of the full vocabulary, and the system constructs a mask vector M of the same dimension. The numerical rules for the mask vector are as follows:

[0035] Subsequently, the mask vector is superimposed on the original logic vector to obtain the corrected logic vector. Finally, regarding The probability distribution is obtained by performing Softmax normalization. Because the position of the illegal word is incremented by negative infinity, its final probability value... This will approach 0, thus effectively suppressing the generation of hallucinations or non-compliant instructions.

[0036] For example, for a user query: "Please analyze the current load status of main transformer No. 1 and main transformer No. 2 at station A, and determine which has a higher load rate?", this step, after FSM verification, is decomposed into the following sequence of atomic operations. : Atomic operations This is a database query interface; its parameters (corresponding to...) and The combination includes the target object as "Station A No. 1 Main Transformer", and the target attributes as current active power and rated capacity.

[0037] Atomic operations This is a database query interface; its parameters include the target object being "Station A, Main Transformer No. 2", and the target attributes being the current active power and rated capacity.

[0038] Atomic operations For load rate calculation, the parameters include the input source as and The returned result.

[0039] Atomic operations To perform numerical comparisons, the parameters include the load rate calculation results, and numerical logic judgments are made based on preset power equipment health assessment criteria, such as determining a heavy load state if the load rate exceeds 80%.

[0040] This step transforms the complex query into a structured initial state sequence that passes the FSM logic check, providing reliable input for subsequent reinforcement learning decisions. Compared to conventional mechanisms that rely on post-event regularization matching based on prompt word engineering or on the model itself following instructions, this embodiment introduces FSM hard constraints at the stage of calculating the original logic vector at the system's underlying level. This mechanism blocks the generation path of illegal tokens from the system's bottom layer, reducing the computational overhead of complex exception handling and repeated retries for invalid formats in the backend. It ensures that the final output atomic operation sequence conforms to the interface protocol specifications of the downstream scheduler, improving the system's syntax parsing efficiency and the determinism of backend data flow.

[0041] S2, based on the atomic operation sequence, calculate the voltage level adaptation factor and context state vector to select tool instantiation, and obtain an instantiated atomic operation set; according to the instantiated atomic operation set, use a reinforcement learning policy network that integrates power domain knowledge to predict topology connection actions, and obtain a parallel directed acyclic graph.

[0042] S2-1, Selection of Tools and Knowledge Sources Calculate the conditional suitability score between the current atomic operation and the candidate tool, the conditional suitability score being determined by the product of the semantic matching degree of the selection neural network output and the voltage level adaptation factor; The selection neural network concatenates the embedding vector of the current atomic operation, the embedding vector of the candidate tool, and the context state vector that records the historical operation path to output the semantic matching degree. The voltage level adaptation factor is generated based on the matching relationship between the rated voltage level and device category of the device involved in the current atomic operation and the applicable scope declared by the candidate tool. When the device attributes of the current atomic operation exceed the applicable scope of the candidate tool, the weight of the voltage level adaptation factor is reset to zero or a preset minimum value to reduce the probability of the candidate tool being selected through a multiplication gating mechanism.

[0043] Specifically, for each of the decomposed subtasks Evaluate the set of candidate tools and select the most suitable one. The tool executes or determines dependencies between tasks. Specifically, this refers to the software interfaces or functional modules that the intelligent agent can call, including but not limited to: APIs of power business systems (such as SCADA data query interfaces), knowledge base retrieval engines, professional mathematical calculation modules, or logical judgment functions. Successful binding constitutes an instantiated atomic operation. The set of tasks containing explicit tool bindings obtained after processing all subtasks in this step is the set of instantiated atomic operations.

[0044] To address the lack of power equipment adaptability in general models, such as confusion regarding voltage levels and misuse of equipment types, this embodiment calculates a condition suitability score. The score depends not only on the semantic matching between the task and the tool itself, but also strictly on the current context state, defined as:

[0045] Among them, Emb( () represents the embedding vector of the task or tool, extracted by a pre-trained Transformer encoder. Atomic operations... Text descriptions (such as "query the load of the main transformer on site A") and tools Functional declarations (such as "database query interface, supporting reading active power") are mapped to high-dimensional vector spaces respectively; The context state vector is used to record key parameters that have been acquired (such as the voltage value that has been found) and historical operation paths, ensuring that the agent has the ability to make deduplicated decisions. To select the neural network, this embodiment employs a multilayer perceptron (MLP) or an attention mechanism. It uses the concatenation of the task vector, tool vector, and context vector as input features, with the symbol... This indicates splicing. When the network detects... The required information already exists In such cases, the network will output a very low matching weight, thus rejecting the selection of the same query tool. In this situation, the redundant task will not trigger the instantiation of the external tool, and subsequent tasks relying on this information will directly retrieve it from the context state. It reads existing data to achieve efficient deduplication decisions. For example, when... If the rated capacity of the No. 1 main transformer in the current device has already been recorded, and the next atomic operation is still to query the rated capacity, then select the neural network. It will be identified through attention mechanisms The high degree of overlap with the task's semantics reduces the tool's matching weight to an extremely low level. In this case, the system directly... The existing values ​​are injected into subsequent tasks, skipping actual interface calls, which effectively reduces the concurrent access pressure on the power grid monitoring system.

[0046] The voltage level adaptation factor reflects the actual applicability of the tool in power operation scenarios. The matching factor is weighted based on the matching relationship between the rated voltage level (e.g., 500kV, 220kV, 110kV, etc.) and equipment category (main transformer, outgoing line, switch, SVG, line protection, etc.) of the equipment involved in the task and the applicable level declared by the tool.

[0047] To prevent the generation of unexecutable task chains, this embodiment employs a multiplicative gating mechanism to impose hard constraints on tool selection: when the voltage level or equipment category involved in the task exceeds the tool's processing range, for example, a tool that only supports 110kV equipment is attempted to handle a 500kV main transformer task, [the following will occur]. The weight of the factor is reset to zero or a preset minimum value. Since the S value is a product relationship, this operation will force the final score to be lowered to zero or ignored, thereby eliminating the probability of selecting the mismatched tool in the candidate set; conversely, when the tool's capabilities perfectly correspond to the task's equipment attributes, the factor takes a higher value (e.g., 1.0) to increase its selection weight.

[0048] Based on the above calculations, for each task Select the tool with the highest score This generates a set of instantiated atomic operations, providing input for the subsequent topology orchestration in S2-2.

[0049] Step S2-2, Strategy Optimization The process of predicting topological connection actions is modeled as a Markov decision process, and the state vector at the current moment is defined, including: the currently constructed local directed acyclic graph structure, the remaining sequence of atomic operations to be placed consisting of the set of instantiated atomic operations, and the current context state vector. A graph neural network is used to extract the topological features of the local directed acyclic graph structure, and a Transformer encoder is used to extract the semantic features of the remaining atomic operation sequences to be placed. The two are then fused and input into the reinforcement learning policy network that incorporates knowledge from the power domain, and the output is the prediction result of the predecessor node connection of the current atomic operation in the local directed acyclic graph. During the training phase of the reinforcement learning policy network that integrates knowledge from the power domain, a composite reward function based on power rules is introduced to update the network. The composite reward function includes: a constraint matching reward based on the matching of tool and equipment voltage levels, a gain reward based on the current context state vector, and a cost penalty based on the number of execution steps.

[0050] Specifically, based on the set of atomic operation nodes instantiated in step S2-1, the goal of this step is to transform these discrete nodes into an execution plan G=(V,E) in the form of a directed acyclic graph (DAG), where node v∈V represents the specific operation after instantiation (i.e., bound to the optimal tool). Task Let e∈E represent the execution order of operations and the direction of data flow. This embodiment models the graph construction process as a Markov Decision Process (MDP), where the agent learns the optimal policy through interaction with the environment. The specific implementation process is as follows: (1) Definition of state space: In order for the agent to perceive the current orchestration progress and dependencies, the state at time t is defined. A composite vector containing three parts of information:

[0051] This represents the directed acyclic graph structure as of the current moment, including the installed nodes and their topological connections. The policy network needs to determine, based on this state, which predecessor nodes a new node should be attached to. This represents the list of atomic operations that have not yet been placed in the sequence output by S1 (e.g., , The agent needs to extract the semantic features of the task at the head of the queue as the basis for the current decision. This represents the context state vector, which records the data dependencies between tasks (e.g., ...). Clearly require The output of this state is used to constrain the agent to schedule execution nodes only after the data preconditions are met, preventing the generation of logically flawed graphs.

[0052] (2) Action space definition: reinforcement learning strategy network that integrates knowledge from the power field Based on the dependencies between tasks, determine the location of the current task node to be assigned in the graph. The connection method in the code. Specifically, it involves predicting the set of parent nodes of the current node, thus forming the following topology: Serial connection: Select the previously completed node as the parent node and construct the dependency edge; Parallel branching: Select a node that has no data dependency on the current node as the parent node (or connect it directly to the Start node) so that the current node can be executed in parallel with the existing path; Merge waiting: Select multiple predecessor nodes as a common parent node to aggregate data from multiple sources, such as comparing load rates. Reinforcement learning policy network. It can automatically identify implicit dependencies between tasks and arrange nodes with no dependencies to be executed in parallel to the greatest extent possible, thereby improving system response efficiency.

[0053] (3) The optimization objective is to find the optimal strategy. This maximizes the cumulative expected reward obtained after the generated task graph is executed.

[0054] in, It is an execution node v The resulting composite reward is used to comprehensively score the selection results from step S2-1 and the ranking results from this step. Composite Reward Defined as:

[0055] in, This represents a constraint matching reward. Although S2-1 has performed initial screening, during the training phase, the environment needs to provide feedback based on the actual voltage / equipment rules. If the tool and equipment are perfectly matched, +1 point is awarded; if a mismatch occurs, such as using a 500kV tool to check a 220kV station, a logic blocking penalty (-10 points) is triggered. This negative feedback will propagate back to the network to correct the selection strategy of S2-1. This represents a gain reward: +1 point for a node acquiring new information, and 0 points for a duplicate query. This signal is used to train the network to identify and skip redundant steps. This represents the cost per step. For each node executed, a fixed 0.1 point is deducted to help the agent find the shortest path.

[0056] In this embodiment, the reinforcement learning policy network A hybrid architecture combining GNN and Transformer is employed. GNNs, such as GraphSage or GAT, act as topological encoders. Their node feature vectors include the functional labels of the assigned tasks, the completeness of input parameters, and the estimated execution time. Their adjacency matrices represent the existing data flow. Through three layers of aggregation operations, the current local graph is extracted. The current features are used to output the embedding vector of the graph structure. The Transformer is used to extract the remaining task sequence. The semantic features are used to output the semantic embedding vector of the sequence. The concatenated feature vector is used as the input to the decision head, which outputs the probability distribution in the action space.

[0057] To ensure training stability and avoid performance crashes caused by excessive policy updates, this embodiment preferably employs the Proximal Policy Optimization (PPO) algorithm on the policy network. Parameter updates are performed. PPO introduces a truncation mechanism to limit the ratio of new to old policies, ensuring that the magnitude of each parameter update is within a controllable range. Through this mechanism, the reward function R(v) effectively serves as a verification of the agent's decision-making accuracy, ensuring continuous optimization of the selection accuracy of S2-1 and the orchestration efficiency of S2-2.

[0058] S3: Based on topological dependencies, concurrently schedule and execute task nodes in the parallel directed acyclic graph, capture multi-source intermediate results; perform time-stamp alignment verification and logical fusion on the multi-source intermediate results to generate the final business solution.

[0059] In step S2, reinforcement learning has generated an optimized directed acyclic graph for task execution. Based on this, the task execution and result integration module organically integrates the distributed intermediate results into a final solution to the user's problem.

[0060] The core component is the scheduler, whose function is to precisely schedule and execute each task node according to the topological order and dependencies of the task graph. This includes executing the following rules: 1) Topology concurrent scheduling rules are used to monitor the in-degree state of each task node in the parallel directed acyclic graph in real time, add task nodes with an in-degree of zero to the ready queue, and use an asynchronous concurrent thread pool to trigger the task nodes in the ready queue in parallel; after the task node finishes execution, update the in-degree state of the downstream child node, and store the output data of each task node as the multi-source intermediate result in the context state vector.

[0061] Specifically, the scheduler no longer uses a simple linear traversal, but instead monitors the in-degree state of each node in the DAG in real time. The scheduler first scans and simultaneously starts all initial nodes with an in-degree of zero (i.e., tasks with no prerequisites). (Query A station's power) and (Querying Bilibili's power) These two tasks are independent of each other, and the scheduler will use an asynchronous concurrent thread pool to trigger them, thereby reducing the overall response latency. At this point, the specific business data returned after each task node completes execution (such as the specific power value of Bilibili and the specific power value of Bilibili) constitutes the multi-source intermediate result. An event-driven mechanism is adopted, and whenever a preceding node completes execution, the scheduler automatically updates the in-degree of its downstream nodes. Once all the preceding constraints of a node are satisfied, the node immediately enters the ready queue and is scheduled in real time. Each task node corresponds to a well-encapsulated atomic operation (such as an API call, database query, or internal function calculation). The scheduler supports non-blocking I / O, and can continuously schedule other ready nodes while waiting for I / O returns, avoiding system blocking caused by a single long-running task.

[0062] 2) Time stamp alignment verification rules are used to obtain the power data section timestamps corresponding to the multi-source intermediate results; calculate the time deviation between the multi-source intermediate results participating in the same business logic calculation; if the time deviation exceeds a preset threshold, the current data is determined to be invalid, and a data re-sampling and rescheduling instruction is triggered for the task node that generated the multi-source intermediate results.

[0063] Specifically, the output of each task node is captured and stored in the context state vector defined in S2. In this context, the container not only serves as a result storage area but also facilitates data transfer between tasks. For example, and The acquired real-time load data will be transmitted to This is used to further calculate the load factor. Specifically, to address the stringent time synchronization requirements of power services, the context state vector introduces a time-stamp alignment mechanism: when integrating measurement data from different SCADA systems, the timestamps of the intermediate results from each multi-source source are automatically verified. If the time deviation of the data involved in the calculation exceeds a preset threshold (e.g., 5 minutes), the scheduler will determine that the current data is invalid and trigger a data resampling instruction, forcibly rescheduling the relevant query tasks to ensure that the final calculation results (such as the load factor) are based on the grid state at the same time segment, avoiding misjudgments caused by data asynchrony.

[0064] 3) Logical fusion generation rules are used to extract the multi-source intermediate results from the context state vector as evidence chains after all task nodes have been executed, dynamically assemble prompt word templates in combination with preset power industry constraint criteria, and input the prompt word templates into the large language model to generate the final business answer in natural language form.

[0065] Specifically, after all the preceding tasks are completed, the scheduler activates the comprehensive analysis node located at the end of the task graph. This node's role is to summarize, integrate, and logically reason about the intermediate results in the context. It dynamically combines the user's original intent, the acquired chain of evidence, and power industry constraints (such as heavy overload judgment criteria) to reassemble a Prompt template. Finally, this Prompt is input into the large model, generating a complete and structured response.

[0066] Compared to the serial execution and simple data concatenation mechanisms used in conventional large-scale question-answering systems, this embodiment constructs an asynchronous concurrent thread pool based on in-degree state. This allows full utilization of the concurrent processing capabilities of multi-core servers, reducing the total blocking time for background calls to multi-source heterogeneous APIs. Simultaneously, the system-level timescale alignment verification mechanism introduced into the context state vector can automatically identify and intercept discrete data caused by asynchronous transmission from the underlying SCADA system. This mechanism ensures the consistency of data slices entering the final logical fusion module across physical time segments, preventing system inference logic errors caused by asynchronous underlying data links.

[0067] S4: The parameters of the large language model are incrementally updated using a parameter consolidation strategy based on domain feature weighting.

[0068] A parameter consolidation strategy based on domain feature weighting is adopted, specifically including: constructing a loss function containing an elastic weight consolidation penalty term, and introducing a domain correction factor for the parameters of the large language model into the elastic weight consolidation penalty term; statistically analyzing the word frequency inverse document frequency features of words in the power document set, calculating the betweenness centrality features of equipment nodes in the power grid topology in the power equipment knowledge graph, and determining the business importance score of the word by multiplying the word frequency inverse document frequency features with the normalized betweenness centrality features; locating the row vector index corresponding to the core power words in the input embedding layer of the large language model, assigning the business importance score to the domain correction factor of the parameter corresponding to the row vector, and maintaining the domain correction factors of other parameters at their default values.

[0069] Step S4-1, Knowledge Increment Absorption: For newly added power grid equipment or topology changes (such as adding a new 220kV line), an incremental learning algorithm based on graph neural networks is used to update only the embedded local subgraphs affected by the changes, avoiding global retraining. For newly released unstructured knowledge (such as policy documents), a hybrid strategy based on knowledge replay and elastic weight consolidation is used to fine-tune the LLM parameters.

[0070] Step S4-2, Elastic Weight Consolidation: To prevent forgetting, an elastic weight consolidation regularization method is introduced. When learning new knowledge, a penalty term is added to the loss function to limit the update of important parameters. To prevent the model from losing its understanding of core electricity vocabulary during the learning process, a weighted average of core electricity vocabulary is added to the elastic weight consolidation framework. Matrix, modified loss function:

[0071] in, This is the loss function based on new knowledge. The second half is the elastic weight consolidation penalty term, where... This is the regularization coefficient, which controls the strength of protection for existing knowledge. The i-th diagonal element of the Fisher information matrix is ​​used for quantization parameters. The importance of old tasks. For the current parameter, Set parameter values ​​after training for the old task is completed; This is the domain correction factor proposed in this embodiment, used to introduce power business logic constraints.

[0072] Traditional Simply reflecting statistical patterns at the data level is insufficient to understand the logic of power business operations. Therefore, this embodiment introduces a domain correction factor. In order to determine To calculate the value, we first need to calculate the business importance score of the k-th word in the vocabulary. The calculation formula is as follows:

[0073] in, To utilize word frequency features, statistics are performed on power document sets (such as dispatching procedures and maintenance manuals) to calculate the TF-IDF value of word k. The higher the value, the more critical the word is in the power text. This represents the betweenness centrality of node k calculated based on a knowledge graph of power equipment. Specifically, a knowledge graph of power equipment is constructed with power equipment as nodes and electrical connections as edges, and the betweenness centrality of node k is calculated using a graph algorithm. This index reflects the hub status of the equipment in the power grid topology; for example, the hub status of a main transformer is much higher than that of an end meter. For the enhancement coefficient (e.g., 10.0), Norm is the normalization function, such as Min-Max normalization.

[0074] Since the semantics of deep parameters in a large model are distributed and difficult to directly correspond to specific words, this embodiment adopts a point-based augmentation strategy at the embedding layer. Assign to The embedding layer of a large model is a matrix, where each row vector uniquely corresponds to a word in the vocabulary. The mapping rule is as follows: if the parameters Let the embedding vector row corresponding to the core term k in the power industry be... If the parameter For parameters belonging to the general vocabulary or deep network, retain the default values. Through the above mapping, it is possible to accurately lock and protect the underlying parameters that are strongly related to the core concepts of power, preventing them from being significantly modified during fine-tuning.

[0075] Compared to conventional full-parameter fine-tuning strategies, this embodiment integrates the physical characteristics of the power network structure with textual statistical features to generate a domain correction factor. This scheme enables the system to proactively perceive the underlying physical memory blocks corresponding to core topology nodes such as main transformers in the neural network. By specifically increasing the update cost of these parameters in the loss function, the parameter weight distribution of core domain logic is locked while continuously absorbing frequently changing unstructured business data. This reduces the probability of catastrophic forgetting in the backend model and extends the period of stable operation without intervention.

[0076] Step S4-3, Knowledge Replay Mechanism: The knowledge replay mechanism allows the model to "review" old knowledge while learning new knowledge, directly consolidating the model's memory at the data level. The system maintains a fixed-size buffer M that stores representative samples of old knowledge. During training, when a batch of new data is sampled from a new knowledge source... At the same time, a small batch of old data is sampled from buffer M using a mixed sampling method. Of these, 70% were common user questions, such as the five-prevention interlocking logic and main transformer overload handling, and 30% were random questions. The old and new samples were combined to form the final training batch.

[0077] By optimizing on both new and old data simultaneously, the model's memory of historical knowledge is explicitly reinforced.

[0078] Example 2: A knowledge query agent generation system with dynamic orchestration and continuous learning includes: The semantic parsing and task decomposition module is used to receive user natural language queries, perform intent recognition and task decomposition using a large language model, and introduce a finite state machine to mask the original logic vector obtained by the large language model during the task decomposition process to obtain an atomic operation sequence. The dynamic topology orchestration module is used to calculate the voltage level adaptation factor and context state vector based on the atomic operation sequence to select tools for instantiation, thereby obtaining an instantiated atomic operation set; based on the instantiated atomic operation set, a reinforcement learning policy network that integrates power domain knowledge is used to predict topology connection actions, thereby obtaining a parallel directed acyclic graph. The task execution and result integration module is used to concurrently schedule and execute task nodes in the parallel directed acyclic graph based on topological dependencies, capture multi-source intermediate results, perform time-stamp alignment verification and logical fusion on the multi-source intermediate results, and generate the final business solution. The domain-based continuous learning module is used to incrementally update the parameters of the large language model using a parameter consolidation strategy based on domain features.

[0079] Example 3: A terminal includes a processor and a storage medium; the storage medium is used to store instructions. The processor is configured to operate according to the instructions to execute the steps of the method.

[0080] Example 4: A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method.

[0081] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0082] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0083] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0084] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A method for generating knowledge-questioning intelligent agents through dynamic orchestration and continuous learning, characterized in that, include: The system receives natural language queries from users, performs intent recognition and task decomposition using a large language model, and introduces a finite state machine to mask and constrain the original logic vectors obtained from the large language model during the task decomposition process to obtain a sequence of atomic operations. The system includes a predefined set of states and valid transition paths between states. The set of states includes at least an intent-locked state, an atomic operation selection state, a parameter key name state, a parameter key value state, and an end state. The valid transition paths include: a one-way transition from the intent-locked state to the atomic operation selection state; an attribute-filling transition from the atomic operation selection state to the parameter key name state and then to the parameter key value state; and a cyclic decision transition from the parameter key value state back to the parameter key name state, back to the atomic operation selection state, or to the end state based on parameter integrity. During the decoding process, the current state of the finite state machine is determined based on the generated word sequence, and the set of legal words for the next time step is determined based on the current state. Specifically, if the current state is the intent-locked state, the set of legal words is limited to intent identifier words in the predefined intent category set; if the current state is the atomic operation selection state, the corresponding atomic operation name words are determined as the set of legal words based on the previously locked intent categories; if the current state is the parameter key name state or the parameter key value state, the corresponding parameter key name or parameter key value words are determined as the set of legal words based on the previously selected atomic operations. Construct a mask vector with the same dimensions as the full vocabulary of the large language model; based on the determined set of legal lexical units, set the values ​​at index positions outside the set of legal lexical units in the mask vector to negative infinity, and set the values ​​at index positions within the set of legal lexical units to zero; superimpose the mask vector onto the original logical vector output by the large language model at the current time step to obtain a corrected logical vector; normalize the corrected logical vector to obtain a corrected probability distribution, and sample based on the corrected probability distribution to generate the current lexical unit; Based on the atomic operation sequence, the voltage level adaptation factor and context state vector are calculated to select the tool instantiation, resulting in an instantiated atomic operation set; according to the instantiated atomic operation set, a reinforcement learning policy network that integrates power domain knowledge is used to predict topology connection actions, resulting in a parallel directed acyclic graph. Based on topological dependencies, the task nodes in the parallel directed acyclic graph are concurrently scheduled and executed to capture intermediate results from multiple sources; the intermediate results from multiple sources are then time-stamped and logically fused to generate the final business solution.

2. The method for generating a knowledge query agent based on dynamic orchestration and continuous learning according to claim 1, characterized in that, Calculate the voltage level adaptation factor and context state vector for tool instantiation selection, specifically including: Calculate the conditional suitability score between the current atomic operation and the candidate tool, the conditional suitability score being determined by the product of the semantic matching degree of the selection neural network output and the voltage level adaptation factor; The selection neural network concatenates the embedding vector of the current atomic operation, the embedding vector of the candidate tool, and the context state vector that records the historical operation path to output the semantic matching degree. The voltage level adaptation factor is generated based on the matching relationship between the rated voltage level and device category of the device involved in the current atomic operation and the applicable scope declared by the candidate tool. When the device attributes of the current atomic operation exceed the applicable scope of the candidate tool, the weight of the voltage level adaptation factor is reset to zero or a preset minimum value to reduce the probability of the candidate tool being selected through a multiplication gating mechanism.

3. The method for generating a knowledge query agent based on dynamic orchestration and continuous learning according to claim 2, characterized in that, The step of predicting topology connection actions based on the instantiated set of atomic operations using a reinforcement learning policy network that incorporates knowledge from the power domain specifically includes: The process of predicting topological connection actions is modeled as a Markov decision process, and the state vector at the current moment is defined, including: the currently constructed local directed acyclic graph structure, the remaining sequence of atomic operations to be placed consisting of the set of instantiated atomic operations, and the current context state vector. A graph neural network is used to extract the topological features of the local directed acyclic graph structure, and a Transformer encoder is used to extract the semantic features of the remaining atomic operation sequences to be placed. The two are then fused and input into the reinforcement learning policy network that incorporates knowledge from the power domain, and the output is the prediction result of the predecessor node connection of the current atomic operation in the local directed acyclic graph. During the training phase of the reinforcement learning policy network that integrates knowledge from the power sector, a composite reward function based on power rules is introduced to update the network. The composite reward function includes: a constraint matching reward based on the matching of tool and equipment voltage levels, a gain reward based on the current context state vector, and a cost penalty based on the number of execution steps.

4. The method for generating a knowledge query agent based on dynamic orchestration and continuous learning according to claim 3, characterized in that, Concurrent scheduling and result fusion of task nodes are performed based on topological dependencies, including executing the following rules: The topology concurrent scheduling rule is used to monitor the in-degree status of each task node in the parallel directed acyclic graph in real time, add task nodes with an in-degree of zero to the ready queue, and use an asynchronous concurrent thread pool to trigger the task nodes in the ready queue in parallel. After the task node completes its execution, the in-degree state of the downstream child nodes is updated, and the output data of each task node is stored in the context state vector as the multi-source intermediate result. Time stamp alignment verification rules are used to obtain the power data section timestamps corresponding to the multi-source intermediate results; Calculate the time deviation between the multi-source intermediate results participating in the same business logic calculation. If the time deviation exceeds a preset threshold, determine that the current data is invalid and trigger a data re-sampling and rescheduling instruction for the task node that generated the multi-source intermediate results. The logical fusion generation rule is used to extract the multi-source intermediate results from the context state vector as an evidence chain after all task nodes have been executed, dynamically assemble the prompt word template by combining it with the preset power industry constraint criteria, and input the prompt word template into the large language model to generate the final business answer in natural language form.

5. The method for generating a knowledge query agent based on dynamic orchestration and continuous learning according to claim 4, characterized in that, The method also includes a parameter consolidation strategy based on domain feature weighting, specifically including: Construct a loss function that includes an elastic weight consolidation penalty term, and introduce a domain correction factor for the parameters of the large language model into the elastic weight consolidation penalty term; The inverse document frequency features of words in the power document set are statistically analyzed. The betweenness centrality features of equipment nodes in the power grid topology in the power equipment knowledge graph are calculated. The product of the inverse document frequency features and the normalized betweenness centrality features is determined as the business importance score of the words. Locate the row vector index corresponding to the core electricity vocabulary in the input embedding layer of the large language model, assign the business importance score to the domain correction factor of the parameter corresponding to the row vector, and keep the domain correction factors of other parameters at their default values.

6. A knowledge query intelligent agent generation system with dynamic arrangement and continuous learning, characterized in that, include: The semantic parsing and task decomposition module is used to receive user natural language queries, perform intent recognition and task decomposition using a large language model, and introduce a finite state machine to mask the original logic vector obtained by the large language model during the task decomposition process to obtain an atomic operation sequence. The system includes a predefined set of states and valid transition paths between states. The set of states includes at least an intent-locked state, an atomic operation selection state, a parameter key name state, a parameter key value state, and an end state. The valid transition paths include: a one-way transition from the intent-locked state to the atomic operation selection state; an attribute-filling transition from the atomic operation selection state to the parameter key name state and then to the parameter key value state; and a cyclic decision transition from the parameter key value state back to the parameter key name state, back to the atomic operation selection state, or to the end state based on parameter integrity. During the decoding process, the current state of the finite state machine is determined based on the generated word sequence, and the set of legal words for the next time step is determined based on the current state. Specifically, if the current state is the intent-locked state, the set of legal words is limited to intent identifier words in the predefined intent category set; if the current state is the atomic operation selection state, the corresponding atomic operation name words are determined as the set of legal words based on the previously locked intent categories; if the current state is the parameter key name state or the parameter key value state, the corresponding parameter key name or parameter key value words are determined as the set of legal words based on the previously selected atomic operations. Construct a mask vector with the same dimensions as the full vocabulary of the large language model; based on the determined set of legal lexical units, set the values ​​at index positions outside the set of legal lexical units in the mask vector to negative infinity, and set the values ​​at index positions within the set of legal lexical units to zero; superimpose the mask vector onto the original logical vector output by the large language model at the current time step to obtain a corrected logical vector; normalize the corrected logical vector to obtain a corrected probability distribution, and sample based on the corrected probability distribution to generate the current lexical unit; The dynamic topology orchestration module is used to calculate the voltage level adaptation factor and context state vector based on the atomic operation sequence to select tools for instantiation, thereby obtaining an instantiated atomic operation set; based on the instantiated atomic operation set, a reinforcement learning policy network that integrates power domain knowledge is used to predict topology connection actions, thereby obtaining a parallel directed acyclic graph. The task execution and result integration module is used to concurrently schedule and execute task nodes in the parallel directed acyclic graph based on topological dependencies, capture multi-source intermediate results, perform time-stamp alignment verification and logical fusion on the multi-source intermediate results, and generate the final business solution.

7. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Non-signalized intersection traffic decision-making method based on large model enabling reinforcement learning decision-making framework

    CN121849148A

  • Automaton-Based Controller and Method with Generative Language Models for Task Execution

    US20250172913A1