Multi-agent collaborative task dispatching method and system based on cycle perception reinforcement learning and user feedback
Patent Information
- Application Number
- CN202611200632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-10
- Publication Date
- 2026-09-11
AI Technical Summary
离线反馈方案仅利用用户评价数据长周期迭代模型,无法在当次会话内实时干预派发策略
Smart Images

Figure CN122733482A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of multi-agent cooperative scheduling and natural language processing technology, specifically to a multi-agent cooperative task dispatching method and system based on recurrent perceptual reinforcement learning and user feedback. Background Technology
[0002] Multi-agent collaborative architectures are widely used in scenarios such as intelligent customer service, distributed business processing, and multimodal task scheduling. They typically employ a central control agent combined with multiple execution sub-agents, where the central control agent uniformly receives user requests and makes task dispatch decisions. Existing task dispatch solutions mainly include fixed rule matching, pre-trained classification model prediction, offline feedback iteration, and low-confidence multi-round confirmation. Fixed rule solutions dispatch tasks by pre-setting the mapping relationship between keywords and sub-agents, but cannot handle semantically complex or ambiguous user requests. Classification model solutions predict dispatch targets by fine-tuning pre-trained language models, adapting to a certain degree of semantic change, but suffer from a self-locking dead loop problem after erroneous dispatch. Offline feedback solutions only utilize user evaluation data for long-term model iteration, unable to intervene in the dispatch strategy in real time within the current session. Multi-round confirmation solutions only trigger a clarification process when the model confidence is low, failing to address the problem of continuous erroneous dispatch in high-confidence error scenarios. Because user requests are semantically diverse and ambiguous, when a user repeatedly initiates similar requests within the same session after an erroneous dispatch, existing solutions cannot identify the cyclic error state, and feedback signals cannot be applied to the decision-making process of the current session in real time. This can easily lead to dispatch deadlock under high-confidence errors, resulting in insufficient robustness of multi-agent system interactions.
[0003] Therefore, it is necessary to provide a multi-agent collaborative task dispatching method and system based on recurrent perceptual reinforcement learning and user feedback to solve the above-mentioned technical problems. Summary of the Invention
[0004] The purpose of this application is to provide a multi-agent collaborative task dispatching method, system, computer terminal, and computer-readable storage medium based on loop-aware reinforcement learning and user feedback, which can improve task dispatching accuracy, enhance session-level interaction robustness, and effectively avoid error loops.
[0005] Firstly, this application provides a multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback, including: Receive the request text input by the user and obtain the current session identifier; Perform normalization processing on the request text to generate a request semantic hash value; The session memory is queried using the session identifier and semantic hash value as indexes to obtain the historical negative feedback count of the request in the current session and the identifier of the most recently dispatched sub-Agent; Based on the historical negative feedback count, a circuit breaker judgment is performed. When the circuit breaker condition is met, a preset prompt message is returned, and the current dispatch process is terminated. For requests that do not trigger the circuit breaker, perform intent identification and output the dispatch probability distribution and maximum confidence level for each sub-Agent; Based on the maximum confidence level, a threshold judgment is performed. When the confidence level is lower than the preset threshold, the clarification mode is entered. After receiving the user's explicit selection, the application is dispatched to the corresponding sub-Agent. For requests that meet the confidence requirements, a state vector integrating multi-dimensional features is constructed. Based on the historical negative feedback count, the corresponding dispatch strategy is matched, the target sub-Agent is determined, and the corresponding task is forwarded for execution. Collect user feedback signals on task execution results and map the feedback signals into quantified reward values; Update the negative feedback count and distribution trajectory record in the session memory bank based on the quantified reward value; The state transition quadruple is stored in the experience pool, and the network parameters of the reinforcement learning decision model are updated based on the experience replay mechanism.
[0006] Further, the generation of the request semantic hash value includes: The request text is processed sequentially by removing stop words, lowercase conversion, and stemming. The standardized text is segmented to obtain a word sequence, and the TF-IDF weight of each word is calculated. The term frequency is the ratio of the number of times the word appears in the current request to the total number of words in the current request, and the inverse document frequency is the logarithm of the ratio of the total number of historical requests in the system to the number of historical requests containing the word. The MurmurHash3 hash function is used to map each word to an L-bit binary vector, and the binary bits are converted into {-1,+1} encoding to obtain the word hash vector; Based on the TF-IDF weights, a bit-by-bit weighted summation is performed on the hash vectors of all words to obtain an L-dimensional real-number summation vector; Perform a sign determination on each bit of the accumulated vector; set the bit to 1 if the bit value is not less than 0, and set it to 0 if the bit value is less than 0, to obtain an L-bit SimHash value. The semantic similarity of requests is determined based on the Hamming distance between two SimHash values. If the Hamming distance does not exceed a preset distance threshold, the two requests are considered to be semantically similar. The Hamming distance is the sum of the number of bits that are different in corresponding positions in two binary strings of equal length.
[0007] Furthermore, the circuit breaker determination and the confidence threshold determination constitute a two-stage circuit breaker mechanism, including: Secondary negative feedback circuit breaker: When the historical negative feedback count is greater than or equal to the preset circuit breaker threshold, the circuit breaker is triggered and a preset prompt message is returned, terminating the current automatic dispatch process; Level 1 Confidence Circuit Breaker: When the maximum confidence level of the intent recognition output is lower than the preset confidence threshold, the system enters clarification mode and returns a clarification question to the user. After receiving the user's explicit selection, the system is directly dispatched to the corresponding sub-Agent. The execution timing of the secondary negative feedback circuit breaker is earlier than that of the primary confidence circuit breaker. When both triggering conditions are met simultaneously, the secondary negative feedback circuit breaker is executed first.
[0008] Furthermore, the construction of the state vector fusing multi-dimensional features includes: The state vector is obtained by concatenating the BERT embedding vector of the request text, the maximum confidence of intent recognition, the historical negative feedback count, the continuous negative feedback flag, and the one-hot encoding of the previous round of dispatched sub-Agents. The total dimension is 768+3+K, where K is the total number of sub-Agents in the system.
[0009] Furthermore, the matching of the corresponding dispatch strategy based on the historical negative feedback count includes: When the historical negative feedback count is 0 and the previous round of feedback is non-negative feedback, the ε-greedy strategy is adopted, selecting the sub-Agent with the largest Q value with a probability of 1-ε, and randomly selecting the target sub-Agent from all sub-Agents with a probability of ε. When the historical negative feedback count is greater than or equal to 1, switch to the forced exploration strategy; When positive feedback is received from the user, the historical negative feedback count is cleared and the strategy is switched back to ε-greedy. When negative feedback is received from the user, the historical negative feedback count is incremented by 1. When the historical negative feedback count reaches the circuit breaker threshold, the secondary negative feedback circuit breaker is triggered.
[0010] Furthermore, the forced exploration strategy employs a semantically dissimilar exploration mode, including: During the system initialization phase, the Sentence-BERT embedding vector of the function description text of each sub-Agent is pre-calculated and stored in the sub-Agent registry; Calculate the cosine distance between the functional embedding vector of the most recently failed sub-Agent and the functional embedding vectors of the remaining candidate sub-Agents; The target sub-Agent is selected with the largest cosine distance using a first preset probability, and the target sub-Agent is randomly selected from all candidate sub-Agents with equal probability using a second preset probability.
[0011] Furthermore, mapping the feedback signal to a quantified reward value includes: When user feedback is satisfactory, it is mapped to a reward value of +1; when feedback is partially satisfactory, it is mapped to a reward value of 0; and when feedback is unsatisfactory, it is mapped to a reward value of -1. When the current feedback is unsatisfactory and the previous feedback was also unsatisfactory, a continuous negative feedback penalty of -0.5 is added to the base reward value.
[0012] Furthermore, the reinforcement learning decision model employs a dual-tower deep Q-network structure, comprising: Text feature pyramid: The input 768-dimensional request text BERT embedding vector is processed sequentially through 256-dimensional and 128-dimensional fully connected layers, using the ReLU activation function, with the first fully connected layer set to a Dropout rate of 0.1; Scalar Feature Tower: Input 3+K dimensional scalar concatenation features, processed through a 64-dimensional fully connected layer, and activated by the ReLU function; Fusion layer: The text feature tower output and the scalar feature tower output are concatenated into a 192-dimensional fusion vector, which is then processed by a 128-dimensional fully connected layer to output a K-dimensional Q-value vector, with each dimension corresponding to the state and action value of a sub-Agent; Maintain a target Q-network with the exact same structure as the deep Q-network, and copy the parameters of the deep Q-network to the target Q-network every preset number of training steps; update the parameters of the deep Q-network based on the mean squared error loss function.
[0013] Furthermore, the experience pool adopts a fixed-capacity circular buffer structure, including: Add a corresponding environment version number to each experience, and filter out historical experiences that are inconsistent with the current sub-Agent configuration version during training sampling; A time-sensitivity weight is added to each experience, and the time-sensitivity weight decays exponentially based on the number of training steps after the experience is stored. During sampling, weighted sampling is performed according to the time-sensitivity weight. The reward value is clipped to the range of [-2, +2], and the Huber loss is used instead of the mean squared error loss to reduce the impact of abnormal reward values on gradient updates. The experience pool is divided into a normal experience area, a loop experience area, and a circuit breaker experience area, which respectively store the interaction experience of normal state, error loop state, and circuit breaker trigger. During training sampling, samples are drawn from the three partitions in a ratio of 5:3:2.
[0014] Furthermore, the updated reinforcement learning decision model supports both offline and online update modes: In offline update mode, batch training is performed in a background thread after the session ends, and the network parameters obtained from the training take effect when the next new session is initialized. In the online update mode, incremental training is performed after each round of user feedback collection. The deep Q network maintains a double buffer for decision parameters and training parameters. The training parameters are synchronized to the decision parameters only when the difference between the L2 norms of the two sets of parameters exceeds a preset threshold. The sampling range of online training excludes the interaction experience that has not been completed in the current round. The learning rate of online update is set to 1 / 10 of the learning rate of offline update.
[0015] Secondly, based on the same inventive concept, this application also provides a multi-agent collaborative task dispatching system based on recurrent perceptual reinforcement learning and user feedback, including: The central control agent is used to receive user requests and perform intent recognition, dispatch decision, loop detection, and policy update. The central control agent includes an intent recognition module, a confidence evaluator, a loop detector, a circuit breaker controller, a policy selector, and an execution trajectory recorder. Multiple sub-Agents are used to receive tasks dispatched by the central control Agent, execute the corresponding business logic, and return the execution results. The user feedback collection module is used to collect explicit or implicit user feedback on the execution results and map the feedback into corresponding reward signals. The reinforcement learning module is used to maintain the deep Q-network and the target Q-network, and to perform network training and policy updates based on state transition data; The session memory is used to store request semantic hashes, dispatch trajectory records, negative feedback counts, and historical dispatch sub-Agent identifiers at the session level.
[0016] Thirdly, based on the same inventive concept, this application also provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned multi-agent cooperative task dispatching method based on recurrent perceptual reinforcement learning and user feedback, or runs the above-mentioned multi-agent cooperative task dispatching system based on recurrent perceptual reinforcement learning and user feedback.
[0017] Fourthly, based on the same inventive concept, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described multi-agent cooperative task dispatching method based on cyclic perceptual reinforcement learning and user feedback, or runs the above-described multi-agent cooperative task dispatching system based on cyclic perceptual reinforcement learning and user feedback.
[0018] Compared with the prior art, this application has the following advantages: This application provides a multi-agent collaborative task dispatching method, system, computer terminal, and computer-readable storage medium based on loop-aware reinforcement learning and user feedback. By introducing a session-level loop-aware mechanism and real-time feedback loop, combined with two-level circuit breaker judgment and dynamic adjustment of reinforcement learning strategy, it effectively identifies and interrupts erroneous dispatching loop paths, avoids the continuous accumulation of errors within a single session, improves task dispatching accuracy, enhances session-level interaction robustness, and effectively avoids error loops. Attached Figure Description
[0019] Figure 1 A flowchart illustrating the multi-agent collaborative task dispatching method based on cyclic perceptual reinforcement learning and user feedback provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of a multi-agent collaborative task dispatching system based on cyclic perceptual reinforcement learning and user feedback, provided in an embodiment of this application. Detailed Implementation
[0020] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0022] Existing multi-agent task dispatching schemes have limitations in handling semantically complex or ambiguous requests. They cannot effectively handle the self-locking deadlock problem after erroneous dispatching, and feedback signals cannot intervene in the dispatching strategy in real time within the current session. They also cannot solve the problem of continuous erroneous dispatching in high-confidence error scenarios. These problems cause existing schemes to fail to identify cyclic error states, easily leading to dispatch deadlocks under high-confidence errors, resulting in insufficient robustness of multi-agent system interactions.
[0023] To address this, this application proposes a multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback. This method receives user requests and obtains session identifiers, then standardizes the request text to generate semantic hash values. By querying the session memory, it obtains the historical negative feedback count of the request and the identifier of the most recently dispatched sub-Agent, and performs circuit breaker checks based on this count to avoid duplicate erroneous dispatches. For requests that do not trigger circuit breaker checks, it performs intent recognition and outputs the dispatch probability distribution and maximum confidence level, using a confidence threshold to determine whether to enter clarification mode. For requests that meet the confidence requirements, it constructs a state vector incorporating multi-dimensional features, matches the corresponding dispatch strategy based on the historical negative feedback count, determines the target sub-Agent, and forwards the task for execution. Simultaneously, it collects user feedback signals and maps them to quantified reward values, which are used to update the negative feedback count and dispatch trajectory records in the session memory. The state transition quadruple is stored in the experience pool, and the network parameters of the reinforcement learning decision model are updated based on the experience replay mechanism.
[0024] For ease of understanding, the following explains some key terms in this embodiment: A session identifier is used to uniquely identify a continuous interaction between a user and the system. This identifier allows the system to track and manage all user requests and responses within a specific time period, thereby maintaining session continuity.
[0025] A request semantic hash value is a fixed-length number or binary string generated after standardizing the user-input request text. This value can concisely represent the semantic content of the request. By comparing the semantic hash values of different requests, their semantic similarity can be quickly determined.
[0026] A session memory is a storage unit used to store historical data associated with a specific session. It may contain information such as the semantic hash value of the request, past dispatch trajectory records, historical negative feedback counts for a specific request, and the identifier of the target sub-Agent in the most recent dispatch, providing contextual basis for subsequent decisions.
[0027] Historical negative feedback count records the total number of times a user has provided unsatisfactory feedback for a specific request or similar requests in the current session. This count is an important indicator for the system to detect cyclic error states and is used to trigger circuit breakers or adjust dispatch strategies.
[0028] Circuit breaking is an error prevention mechanism. When the system detects a specific condition (such as the historical negative feedback count reaching a preset threshold), it will actively interrupt the normal task dispatch process and instead execute preset error handling or prompts to avoid repeated error dispatch and improve user experience.
[0029] The reinforcement learning decision-making model is the core intelligent agent in the embodiments of this application. It learns the optimal task dispatch strategy through interaction with the environment. This model is typically composed of a deep neural network and can output action values (Q-values) for different sub-agents based on the current state, thereby guiding task dispatch.
[0030] An experience pool is a buffer used to store state transition quadruples (current state, action performed, reward obtained, and next state). By randomly sampling historical experiences from the experience pool for training, temporal correlations between data can be broken, improving the training efficiency and stability of reinforcement learning decision models.
[0031] The experience replay mechanism refers to the process of randomly extracting historical experience data from an experience pool to train a reinforcement learning decision model. This mechanism helps the model learn from diverse historical interactions, avoids getting trapped in local optima, and improves data utilization.
[0032] like Figure 1 As shown in the figure, this application provides a multi-agent collaborative task dispatching method based on cyclic perceptual reinforcement learning and user feedback.
[0033] S1. Receive the request text input by the user and obtain the current session identifier.
[0034] The session identifier can be a unique string or number sequence automatically generated by the system, used to track the user's request context throughout the interaction.
[0035] S2. Perform normalization processing on the request text to generate a request semantic hash value; Specifically, the request text can be normalized in a basic way, such as removing punctuation marks and unifying the case of characters. Then, a hash value can be calculated using a basic hash function (such as a portion of CRC32 or MD5). This hash value is used to represent the semantic features of the request.
[0036] S3. Query the session memory using the session identifier and semantic hash value as indexes to obtain the historical negative feedback count of the request in the current session and the identifier of the most recently dispatched sub-Agent.
[0037] The session memory stores historical interaction data. This query can retrieve the historical negative feedback count for the request within the current session, as well as the identifier of the most recently dispatched sub-Agent. For example, the session memory can be configured as a key-value store, where the key is a combination of the session identifier and a semantic hash value, and the value contains the negative feedback count and the sub-Agent identifier.
[0038] S4. Perform circuit breaker judgment based on historical negative feedback count. If the circuit breaker condition is met, return the preset prompt message and terminate the current dispatch process.
[0039] Based on the acquired historical negative feedback count, a circuit breaker is triggered. When the historical negative feedback count reaches a preset circuit breaker condition, such as exceeding a certain fixed value, a preset prompt message will be returned to the user, and the current task dispatch process will be terminated. This prevents the system from repeatedly dispatching invalid tasks when a problem is known to exist.
[0040] S5. Perform intent recognition on requests that do not trigger circuit breakers, and output the dispatch probability distribution and maximum confidence level for each sub-Agent.
[0041] For requests that do not trigger a circuit breaker, intent recognition is performed to determine the potential intent of the user's request. The intent recognition module can be configured as a component based on rule matching or a base classification model (such as a Naive Bayes classifier), which outputs the dispatch probability distribution for each sub-Agent and the highest confidence value among them.
[0042] S6. Perform threshold judgment based on the maximum confidence level. When the confidence level is lower than the preset threshold, enter the clarification mode and receive the user's explicit selection before dispatching to the corresponding sub-Agent.
[0043] Based on the maximum confidence level of the intent recognition output, a threshold judgment is performed. If the maximum confidence level is lower than the preset threshold, it indicates uncertainty about the recognized intent, and the system will enter clarification mode. In clarification mode, a clarification question will be posed to the user, and after receiving a clear selection from the user, the task will be dispatched to the corresponding sub-Agent specified by the user.
[0044] S7. For requests that meet the confidence requirements, construct a state vector that integrates multi-dimensional features, match the corresponding dispatch strategy based on historical negative feedback counts, determine the target sub-Agent, and forward it to execute the corresponding task.
[0045] For requests that meet the confidence requirements, a state vector integrating multi-dimensional features is constructed. This state vector can be composed of bag-of-words model features of the request text, a simple encoding of the intent recognition result, and historical negative feedback counts. Subsequently, based on these historical negative feedback counts, a corresponding dispatch strategy is matched. For example, when the negative feedback count is zero, a default dispatch strategy can be used; when the negative feedback count is non-zero, an exploratory strategy is switched. Thus, through this strategy, a target sub-Agent is determined, and the task is forwarded to that sub-Agent for execution.
[0046] S8. Collect user feedback signals on task execution results and map the feedback signals into quantified reward values.
[0047] After the task is completed, user feedback on the result is collected. This feedback can be expressed as an explicit satisfaction rating (e.g., "satisfied" or "dissatisfied") or implicit behavior (e.g., whether the user asks further questions). This feedback is then mapped to a quantified reward value; for example, satisfaction is mapped to a positive value, and dissatisfaction is mapped to a negative value.
[0048] S9. Update the negative feedback count and distribution trajectory record in the session memory bank based on the quantified reward value.
[0049] Based on this quantified reward value, the negative feedback count and dispatch trajectory records in the session memory are updated. For example, if the user feedback is unsatisfactory, the corresponding historical negative feedback count is increased; at the same time, the status, action, and result of this dispatch are recorded for subsequent analysis and learning.
[0050] S10. Store the state transition quadruple into the experience pool and update the network parameters of the reinforcement learning decision model based on the experience replay mechanism.
[0051] The state transition quadruple (including the current state, the action performed, the reward obtained, and the next state) formed during this interaction is stored in the experience pool. The experience pool can be configured as a basic list or queue structure. Based on the experience replay mechanism, a batch of historical experience data is randomly drawn from the experience pool to update the network parameters of the reinforcement learning decision model. For example, the gradient descent algorithm can be used to adjust the model's weights based on this experience data to optimize its dispatch strategy.
[0052] This embodiment effectively solves the problems of erroneous assignment self-locking loops, inability to intervene in feedback signals in real time, and continuous assignment of high-confidence errors in traditional multi-agent task dispatch schemes by introducing a loop perception mechanism and user feedback. By perceiving historical negative feedback in real time and combining it with reinforcement learning for policy adjustment, this embodiment can avoid repeated erroneous assignments, improve the robustness of the system in handling semantically complex and ambiguous requests, and thus improve the interactive experience of multi-agent systems.
[0053] In some possible implementations, generating a request semantic hash value includes: The request text is processed sequentially by removing stop words, lowercase conversion, and stemming. The standardized text is segmented to obtain a word sequence, and the TF-IDF weight of each word is calculated. The term frequency is the ratio of the number of times the word appears in the current request to the total number of words in the current request, and the inverse document frequency is the logarithm of the ratio of the total number of historical requests in the system to the number of historical requests containing the word. Mapping each word to an L-bit binary vector using the MurmurHash3 hash function, and converting the binary bits to {-1,+1} encoding to obtain a word hash vector; Performing bit-wise weighted accumulation on hash vectors of all words based on TF-IDF weights to obtain an L-dimensional real-valued accumulation vector; Performing sign judgment on each bit of the accumulation vector, setting the bit to 1 when the bit value is not less than 0, and setting the bit to 0 when the bit value is less than 0, to obtain an L-bit SimHash value; Judging the semantic similarity of requests based on the Hamming distance between two SimHash values, determining that the semantics of two requests are similar when the Hamming distance does not exceed a preset distance threshold, wherein the Hamming distance is the total number of bits with different values at corresponding positions of two binary strings of equal length.
[0054] Specifically, when performing standardization processing on request text, a stop word removal operation is performed first to remove high-frequency words that contribute little to semantics such as "de" (of), "shi" (is), "zai" (in), so as to reduce text noise and make subsequent analysis more focused on core semantics. Then, lowercase processing is performed, converting all uppercase letters to lowercase, so as to eliminate word inconsistency caused by case differences and ensure that "Apple" and "apple" are regarded as the same word. Subsequently, stemming is performed to restore words to their stem or root form, for example, unifying "running" and "runs" into "run", further normalizing word form changes and improving the accuracy of semantic representation.
[0055] After standardization processing, text segmentation is performed on the text to obtain an independent word sequence. Subsequently, the TF-IDF weight is calculated for each word. Among them, term frequency (TF) measures the importance of a word in the current request, calculated as the ratio of the number of occurrences of the word in the current request to the total number of words in the current request. Inverse document frequency (IDF) measures the rarity and discriminability of a word in the entire system's historical request corpus, calculated as the logarithm of the ratio of the total number of historical requests of the system to the number of historical requests containing the word. Through TF-IDF weights, the local importance of words in the current request and their discriminative ability in the global corpus can be comprehensively evaluated, providing more representative features for subsequent semantic hash calculation.
[0056] Next, each word is mapped to a fixed-length L-bit binary vector using the MurmurHash3 hash function. MurmurHash3 is an efficient non-cryptographic hash function that can ensure the rapidity and stability of the mapping from words to binary vectors. Subsequently, each bit of these L-bit binary vectors is converted to {-1,+1} encoding, that is, binary bit 0 is mapped to -1, and binary bit 1 is mapped to +1, so as to obtain word hash vectors. This conversion introduces discrete binary values into the real number domain, which facilitates subsequent weighted accumulation operations and is consistent with the principle of the SimHash algorithm.
[0057] Based on the TF-IDF weights calculated above, a bit-by-bit weighted summation operation is performed on the hash vectors of all words to obtain an L-dimensional real-valued summation vector. Each bit of this summation vector combines the weighted contribution of all words to the corresponding hash bit, and its magnitude and sign reflect the bias of that bit in the entire request text.
[0058] Finally, a sign check is performed on each bit of the accumulated vector. If the bit value is not less than 0, it is set to 1; otherwise, it is set to 0. This operation converts the L-dimensional real-number accumulated vector into an L-bit binary SimHash value. This SimHash value serves as a semantic fingerprint of the request text, compactly representing the semantic features of the text.
[0059] In determining the semantic similarity of requests, this embodiment uses the Hamming distance between two SimHash values. The Hamming distance is defined as the sum of the number of corresponding bits that differ between two binary strings of equal length. A characteristic of the SimHash algorithm is that semantically similar texts typically have a small Hamming distance between their SimHash values. Therefore, by calculating the Hamming distance between the SimHash values of two requests and comparing it to a preset distance threshold, it is possible to efficiently and accurately determine whether the two requests are semantically similar. When the Hamming distance does not exceed the preset distance threshold, the two requests are considered semantically similar.
[0060] In some possible implementations, circuit breaker determination and confidence threshold determination constitute a two-stage circuit breaker mechanism. This mechanism is designed as a safety safeguard to prevent the system from repeatedly attempting to process potentially failed or ambiguous requests. It introduces different levels of checks, each designed to handle specific types of request issues, thereby enhancing the robustness and efficiency of the task dispatch process, preventing resource waste and improving user experience by providing timely feedback or clarification.
[0061] Level 2 Negative Feedback Circuit Breaker: When the historical negative feedback count is greater than or equal to the preset circuit breaker threshold, the circuit breaker is triggered and a preset prompt message is returned, terminating the current automatic dispatch process.
[0062] The historical negative feedback count reflects past failures or dissatisfactions of a specific request in the current session. When this count reaches or exceeds a preset threshold, the system immediately returns a preset message to the user, such as "Sorry, this request has failed multiple times. Please try a different approach or contact customer service," and terminates the current dispatch process. This mechanism aims to prevent the system from endlessly trying to fulfill a request that continuously leads to negative user feedback, thereby saving computing resources and avoiding further user frustration, effectively acting as a "hard stop" for persistent problem requests.
[0063] Level 1 Confidence Circuit Breaker: When the maximum confidence of the intent recognition output is lower than the preset confidence threshold, the system enters clarification mode and returns a clarification question to the user. After receiving the user's explicit selection, the system is directly dispatched to the corresponding sub-Agent.
[0064] When the maximum confidence score output by the intent recognition module for the current request is lower than a preset confidence threshold, it indicates that the system is uncertain about the user's true intent. In this case, the system will not automatically dispatch the request but will instead enter clarification mode, returning a clarifying question to the user, such as "Do you want to process service A or service B?", to solicit a clear choice. Once the user provides a clear choice, the system will directly dispatch the request to the corresponding sub-Agent without further intent recognition or complex policy matching. This mechanism aims to gracefully handle ambiguous requests, improve accuracy by seeking user confirmation, ensure the correct sub-Agent is invoked, and thus enhance user satisfaction.
[0065] The execution timing of the second-level negative feedback circuit breaker precedes that of the first-level confidence circuit breaker. When both triggering conditions are met simultaneously, the second-level negative feedback circuit breaker is executed first.
[0066] The system first checks if the historical negative feedback count has reached the circuit breaker condition. Only if the secondary negative feedback circuit breaker is not triggered will the system further check if the maximum confidence level of intent recognition is lower than a preset threshold. If both conditions are met simultaneously—that is, the historical negative feedback count reaches the circuit breaker threshold and the intent recognition confidence level is low—the system will prioritize executing the secondary negative feedback circuit breaker. This priority setting ensures that the system can immediately stop processing requests that have failed multiple times, avoiding unnecessary intent recognition and clarification processes.
[0067] In some possible implementations, a state vector that fuses multi-dimensional features is constructed, including: The state vector is obtained by concatenating the BERT embedding vector of the request text, the maximum confidence of intent recognition, the historical negative feedback count, the continuous negative feedback flag, and the one-hot encoding of the previous round of dispatched sub-Agents. The total dimension is 768+3+K, where K is the total number of sub-Agents in the system.
[0068] Specifically, the state vector is the foundation for a reinforcement learning agent to perceive its environment and make decisions. It encodes various information related to task assignment at the current moment into a fixed-dimensional numerical vector, which serves as input to the reinforcement learning decision model to evaluate the value of different actions (i.e., assignment to different sub-agents). A comprehensive and information-rich state vector is crucial for the agent to learn efficient assignment strategies. The BERT embedding vector of the request text is used to capture the deep semantic information of the user's request. Through a pre-trained BERT model, the original text request input by the user can be transformed into a fixed-dimensional dense vector representation, which effectively encodes the contextual relationships between words and the overall semantic meaning. This embedding vector provides the reinforcement learning model with a detailed understanding of the user's true intent, going beyond simple keyword matching, thus supporting more accurate assignment decisions. The maximum confidence score of intent recognition reflects the system's certainty in judging the user's request intent. This value is typically output by the intent recognition module, representing its confidence level in the most probable intent identified. Incorporating this confidence level into the state vector allows the reinforcement learning model to perceive the ambiguity or definiteness of the current intent recognition. Therefore, when the confidence level is low, it tends to adopt a more cautious strategy, such as entering a clarifying mode or engaging in exploratory dispatch. The historical negative feedback count records the number of dispatch failures in the current session or for similar requests. This counter quantifies the user's dissatisfaction with the system's past dispatch decisions. Including this as part of the state vector, the reinforcement learning model can learn to avoid repeating dispatch paths that lead to user dissatisfaction and adjust its strategy as negative feedback accumulates, such as triggering a circuit breaker or switching to exploratory dispatch. The continuous negative feedback flag is a binary feature indicating whether the most recent user feedback was negative. This flag captures the immediate trend of user experience, i.e., whether the user has continuously encountered unsatisfactory service. By perceiving this flag, the reinforcement learning model can respond more sensitively to continuous negative interactions, such as triggering a stronger penalty mechanism or immediately switching to a different dispatch strategy to prevent further deterioration of the user experience. The one-hot encoding of the previous dispatch sub-Agent is used to represent the specific sub-Agent selected in the previous task dispatch. Through one-hot encoding, each sub-Agent is represented as a K-dimensional vector, with only one position set to 1 and the rest to 0, where K is the total number of sub-Agents in the system. This feature allows the reinforcement learning model to perceive the impact of the previous action, thereby learning which combinations or sequences of sub-Agents are effective and which are ineffective in a specific context. This helps avoid repeatedly assigning sub-Agents to inappropriate ones or to effectively switch over after a specific sub-Agent fails. The state vector is formed by concatenating the above features.Specifically, the BERT embedding vector of the request text (768-dimensional), the maximum confidence of intent recognition (1-dimensional), the historical negative feedback count (1-dimensional), the continuous negative feedback flag (1-dimensional), and the one-hot encoding of the previous dispatched sub-Agent (K-dimensional) are sequentially connected to form a comprehensive state vector with a total dimension of 768+3+K.
[0069] In some possible implementations, a corresponding dispatch strategy is matched based on historical negative feedback counts, including: When the historical negative feedback count is 0 and the previous round of feedback is non-negative feedback, the ε-greedy strategy is adopted, selecting the sub-Agent with the largest Q value with a probability of 1-ε, and randomly selecting the target sub-Agent from all sub-Agents with a probability of ε. When the historical negative feedback count is greater than or equal to 1, switch to the forced exploration strategy; When positive feedback is received from the user, the historical negative feedback count is cleared and the strategy is switched back to ε-greedy. When negative feedback is received from the user, the historical negative feedback count is incremented by 1. When the historical negative feedback count reaches the circuit breaker threshold, the secondary negative feedback circuit breaker is triggered.
[0070] Specifically, when the historical negative feedback count is 0 and the previous round of feedback was non-negative, the system employs an ε-greedy strategy for task assignment. The ε-greedy strategy is a commonly used exploration-exploitation balancing strategy in reinforcement learning. In this case, the system selects the sub-Agent with the highest current Q-value (i.e., expected reward) with a high probability (1-ε), representing the "exploitation" of the known optimal choice; simultaneously, it randomly selects a target sub-Agent from all available sub-Agents with a lower probability (ε) for "exploration." This mechanism ensures that the system utilizes known experience while also trying new sub-Agents to discover potential better solutions, avoiding getting trapped in local optima. For example, the system can maintain a Q-value table or output the Q-values of each sub-Agent through a deep Q-network. During decision-making, a random number between 0 and 1 is generated. If this random number is greater than ε, the sub-Agent with the highest Q-value is selected; otherwise, a sub-Agent is randomly and uniformly selected from the set of all sub-Agents.
[0071] When the historical negative feedback count is greater than or equal to 1, the system switches to a forced exploration strategy. A historical negative feedback count greater than or equal to 1 indicates that the user was dissatisfied with the previous task execution result at least once. This usually means that the system may have dispatched an unsuitable sub-Agent or that the current strategy has a problem. To avoid repeating mistakes and find the correct sub-Agent as quickly as possible, the system switches from the ε-greedy strategy to a more proactive exploration mode. This mode aims to guide the system out of the current erroneous path and actively try other sub-Agents that are different from the previously failed ones, in order to discover a more effective solution.
[0072] When the system receives positive feedback from a user, it clears the historical negative feedback count to zero and switches back to the ε-greedy policy. Positive user feedback is a crucial signal for system learning and improvement, indicating that the system's previous dispatch decisions were successful. Clearing the historical negative feedback count and reverting to the ε-greedy policy allows the system to continue striking a balance between utilizing known optimal solutions and moderate exploration, consolidating successful experiences. For example, after the user feedback acquisition module receives a positive feedback signal, it triggers an event that updates the historical negative feedback count for the corresponding session in the session memory to 0 and notifies the policy selector to reset the current dispatch policy to ε-greedy mode.
[0073] When the system receives negative feedback from a user, it increments the historical negative feedback count by 1. Each time negative feedback is received, the historical negative feedback count increases, reflecting the user's level of dissatisfaction with the system's performance in the current session. When the historical negative feedback count reaches a preset circuit breaker threshold, the system triggers a secondary negative feedback circuit breaker. This indicates that the system may have experienced multiple consecutive dispatch failures or is consistently unable to meet the user's needs in the current session. Triggering the circuit breaker mechanism is to stop losses promptly, prevent the system from continuing invalid attempts, return a preset prompt message to the user, and terminate the current automatic dispatch process to prevent further damage to the user experience. For example, after updating the historical negative feedback count, the circuit breaker controller checks whether the count has reached the preset circuit breaker threshold. If it has, the circuit breaker controller immediately interrupts the current dispatch process and returns a preset error or prompt message to the user.
[0074] In some possible implementations, the forced exploration strategy employs a semantically dissimilar exploration mode, including: During the system initialization phase, the Sentence-BERT embedding vector of the function description text of each sub-Agent is pre-calculated and stored in the sub-Agent registry; Calculate the cosine distance between the functional embedding vector of the most recently failed sub-agent and the functional embedding vectors of the remaining candidate sub-agents. The calculation formula is: ; in, The functional embedding vector for the failed sub-agent. Functional embedding vectors for candidate sub-Agents; The target sub-Agent is selected with the largest cosine distance using a first preset probability, and the target sub-Agent is randomly selected from all candidate sub-Agents with equal probability using a second preset probability.
[0075] Specifically, during the system initialization phase, the Sentence-BERT embedding vectors for the functional description text of each sub-Agent are pre-computed, and these embedding vectors are stored in the sub-Agent registry. Sentence-BERT is a model that can map sentences to a high-dimensional vector space, and its generated embedding vectors can effectively capture the semantic information of the text. By pre-compiling and storing these functional description embedding vectors, the system can quickly obtain the functional semantic representation of each sub-Agent at runtime, laying the foundation for subsequent semantic comparison.
[0076] When forced exploration is required, the system calculates the cosine distance between the feature embedding vector of the most recently failed sub-Agent and the feature embedding vectors of the remaining candidate sub-Agents. Cosine distance is a metric that measures the cosine of the angle between two non-zero vectors and is commonly used to assess the semantic similarity of text or vectors. It is calculated by dividing the cosine similarity by the product of the vector dot product and the vector magnitude, and then subtracting this value from 1 to obtain the cosine distance. A larger distance indicates greater semantic dissimilarity.
[0077] When determining the target sub-Agent, the system selects the sub-Agent with the largest cosine distance with a first preset probability. This means the system prioritizes exploring the sub-Agent with the greatest semantic difference in functional description from the failed sub-Agents, aiming to avoid the original erroneous dispatch path. Simultaneously, to ensure the breadth of exploration and avoid getting trapped in local optima, the system also randomly selects the target sub-Agent from all candidate sub-Agents with a second preset probability. This strategy, combining semantic dissimilarity-based priority selection with random selection, ensures the effectiveness and comprehensiveness of the exploration.
[0078] In some possible implementations, the feedback signal is mapped to a quantified reward value, including: When user feedback is satisfactory, it is mapped to a reward value of +1; when feedback is partially satisfactory, it is mapped to a reward value of 0; and when feedback is unsatisfactory, it is mapped to a reward value of -1. When the current feedback is unsatisfactory and the previous feedback was also unsatisfactory, a continuous negative feedback penalty of -0.5 is added to the base reward value.
[0079] When a user is satisfied with the task result performed by the sub-Agent, the system converts this positive feedback into a positive reward value of +1. This reward value is used by the reinforcement learning model to encourage it to choose actions that lead to user satisfaction when encountering similar states in the future. This positive incentive mechanism is the foundation for guiding agents to learn optimal policies in reinforcement learning. When a user is partially satisfied or has no clear preference for the task result performed by the sub-Agent, the system maps it to a neutral reward value of 0. This reward value indicates that the current dispatched action neither brings significant positive effects nor causes obvious negative consequences. In reinforcement learning, neutral rewards help the model to evaluate actions with unclear effects impartially during the exploration process, avoiding over-punishment or over-reward, thus maintaining the balance of exploration. When a user is dissatisfied with the task result performed by the sub-Agent, the system converts this negative feedback into a negative reward value of -1. This reward value is used to penalize the reinforcement learning model, prompting it to avoid choosing dispatched actions that lead to user dissatisfaction when encountering similar states in the future. This negative penalty mechanism is key to guiding agents to avoid suboptimal or incorrect policies in reinforcement learning. To more strongly reflect consecutive user dissatisfaction, when the system detects that both the current and previous feedback is unsatisfactory, an additional penalty of -0.5 for consecutive negative feedback is added on top of the base reward of -1. This means that in cases of consecutive dissatisfaction, the total reward will become -1.5. This mechanism aims to send a stronger signal to the reinforcement learning model that consecutive failures are unacceptable and that the model needs to quickly adjust its strategy to avoid falling into a cycle of repeated error distribution.
[0080] In some possible implementations, the reinforcement learning decision model employs a dual-tower deep Q-network, including: Text feature pyramid: The input 768-dimensional request text BERT embedding vector is processed sequentially through 256-dimensional and 128-dimensional fully connected layers, using the ReLU activation function, and the first fully connected layer is set with a Dropout rate of 0.1.
[0081] Scalar Feature Tower: Input 3+K dimensional scalar concatenation features, processed through a 64-dimensional fully connected layer, and activated by the ReLU function; Fusion layer: The text feature tower output and the scalar feature tower output are concatenated into a 192-dimensional fusion vector, which is then processed by a 128-dimensional fully connected layer to output a K-dimensional Q-value vector, with each dimension corresponding to the state and action value of a sub-Agent; Maintain a target Q-network with the exact same structure as the deep Q-network, and copy the parameters of the deep Q-network to the target Q-network every preset number of training steps; update the parameters of the deep Q-network based on the mean squared error loss function. for: ; in To quantify reward value, As a discount factor, For the goal The network's output value, Select an action for the current state of value.
[0082] For example, the dual-tower deep Q-network includes a text feature tower, a scalar feature tower, and a fusion layer, and maintains a target Q-network with the exact same structure as the deep Q-network, updating the deep Q-network parameters based on the mean squared error loss function.
[0083] Specifically, the text feature tower in this dual-tower deep Q-network is designed to process high-dimensional textual semantic information. The input to the text feature tower is a 768-dimensional BERT embedding vector of the request text, obtained by encoding the user request text using a pre-trained BERT model, which captures the deep semantic features of the text. To effectively extract these features and prevent overfitting, the text feature tower is processed sequentially through 256-dimensional and 128-dimensional fully connected layers, with a Dropout rate of 0.1 set in the first fully connected layer to enhance the model's generalization ability. After each fully connected layer, a ReLU activation function is used to introduce non-linearity, enabling the network to learn more complex feature representations.
[0084] Meanwhile, the scalar feature tower is used to process structured scalar features. The input to the scalar feature tower is a 3+K scalar concatenated feature set, where K is the total number of sub-Agents in the system. These scalar features can include maximum confidence in intent recognition, historical negative feedback counts, continuous negative feedback flags, and one-hot encodings of the sub-Agents dispatched in the previous round, providing key information about the current session state and historical interactions. The scalar feature tower is processed through a 64-dimensional fully connected layer and uses the ReLU activation function to perform effective feature transformation and dimensionality compression on these scalar features.
[0085] Subsequently, the fusion layer concatenates the outputs of the text feature pyramid and the scalar feature pyramid to form a 192-dimensional fusion vector. The text feature pyramid outputs 128-dimensional features, and the scalar feature pyramid outputs 64-dimensional features; their concatenation results in a 192-dimensional comprehensive feature representation. This fusion vector is then processed by a 128-dimensional fully connected layer, ultimately outputting a K-dimensional Q-value vector, where each dimension corresponds to the state-action value of a sub-agent. Each element of this K-dimensional vector represents the expected cumulative reward for selecting the corresponding sub-agent as the action in the current state.
[0086] To improve training stability and convergence, this embodiment also maintains a target Q-network with the exact same structure as the deep Q-network. The target Q-network provides a relatively stable Q-value estimation target, avoiding oscillations caused by continuous changes in the Q-value target during training. Specifically, every preset number of training steps, the parameters of the deep Q-network are copied to the target Q-network, allowing it to slowly track the learning progress of the main network. Regarding parameter updates, the deep Q-network parameters are updated based on the mean squared error loss function, where the loss function L is defined as... Where r is the quantized reward value, γ is the discount factor, Q_target is the output value of the target Q network, and Q(s,a) is the Q value of choosing action a in the current state. This loss function is used to measure the difference between the predicted Q value and the target Q value, and minimizes this difference through a gradient descent optimization algorithm, thereby enabling the network to learn a more accurate Q value estimate.
[0087] In some possible implementations, the experience pool employs a fixed-capacity circular buffer structure, including: Add a corresponding environment version number to each experience, and filter out historical experiences that are inconsistent with the current sub-Agent configuration version during training sampling; Add a time-sensitivity weight to each experience. The time-sensitivity weight decays exponentially based on the number of training steps after the experience is stored. Weighted sampling is performed according to the time-sensitivity weight during sampling. The reward value is clipped to the range of [-2, +2], and the Huber loss is used instead of the mean squared error loss to reduce the impact of abnormal reward values on gradient updates. The experience pool is divided into a normal experience area, a loop experience area, and a circuit breaker experience area, which respectively store the interaction experience of normal state, error loop state, and circuit breaker trigger. During training sampling, samples are drawn from the three partitions in a ratio of 5:3:2.
[0088] For example, the circular buffer can store a fixed number of experience samples. When the buffer is full, new experience will overwrite the oldest experience, thereby ensuring that the size of the experience pool is controlled, preventing memory overflow, and ensuring that the training data has a certain timeliness.
[0089] To address the issue of dynamic changes in system configuration, this embodiment adds a corresponding environment version number to each experience, filtering out historical experiences that are inconsistent with the current sub-Agent configuration version during training sampling. This means that when the sub-Agent configuration (e.g., functionality, quantity) in the system is updated, a global environment version number is incremented. When a new interaction experience is stored in the experience pool, the current environment version number is recorded. In subsequent training sampling, only experiences with the same system environment version number are selected for training, thereby preventing the model from learning from outdated or irrelevant experiences and improving the model's adaptability and stability.
[0090] Meanwhile, to make the model focus more on recent experiences, accelerate the learning process, and adapt to environmental changes, this embodiment adds a time-sensitive weight to each experience. This time-sensitive weight decays exponentially based on the number of training steps since the experience was stored, and weighted sampling is performed according to the time-sensitive weight during sampling. Specifically, each experience is assigned an initial weight, such as 1, upon storage. After each training iteration, the weights of all experiences decay exponentially according to a preset decay factor (e.g., 0.99). During sampling, the system performs non-uniform sampling based on these decayed weights; experiences with higher weights have a greater probability of being sampled, thus enabling the model to adapt to the latest user behavior and system state more quickly.
[0091] To enhance the robustness of model training and reduce the drastic impact of outlier reward values on gradient updates, this embodiment prunes reward values to the [-2, +2] range and uses Huber loss instead of mean squared error loss. The reward pruning operation maps user feedback to quantized reward values and then limits them to the preset range of [-2, +2]. For example, if the original reward value exceeds this range, it is truncated to -2 or +2. The Huber loss function uses squared error when the error is small and linear error when the error is large, making it less sensitive to outliers and effectively preventing gradient explosion or training instability caused by outlier reward values.
[0092] Furthermore, to address the uneven distribution of different types of experience and ensure the model can fully learn decision-making strategies under various critical states, especially handling errors and anomalies, this embodiment divides the experience pool into a normal experience zone, a cyclic experience zone, and a circuit breaker experience zone. These zones store interaction experiences in normal states, error cyclic states, and before the circuit breaker is triggered, respectively. During training sampling, samples are drawn from each of the three zones in a 5:3:2 ratio. Specifically, the normal experience zone stores experiences of successful dispatch and obtaining non-negative feedback; the cyclic experience zone stores experiences of users repeatedly requesting or the system entering an error cyclic state due to dispatch failure; and the circuit breaker experience zone stores experiences of system-user interaction before the circuit breaker mechanism is triggered. During training sampling, the system no longer samples uniformly from the entire experience pool, but instead draws samples from each of the three zones according to a preset ratio (e.g., 5:3:2), and then merges these samples for training.
[0093] In some possible implementations, updating the reinforcement learning decision model supports both offline and online update modes: In offline update mode, batch training is performed in a background thread after the session ends, and the network parameters obtained from the training take effect when the next new session is initialized. In the online update mode, incremental training is performed after each round of user feedback collection. The deep Q network maintains a double buffer for decision parameters and training parameters. The training parameters are synchronized to the decision parameters only when the difference between the L2 norms of the two sets of parameters exceeds a preset threshold. The sampling range of online training excludes the interaction experience that has not been completed in the current round. The learning rate of online update is set to 1 / 10 of the learning rate of offline update.
[0094] Specifically, the offline update mode aims to utilize historical session data for stable and comprehensive model learning. When a user session ends, the system starts a separate thread in the background to perform batch training. This training process aggregates experience data accumulated from that session and even multiple sessions, updating the parameters of the deep Q-network through one or more iterations. Because it is executed asynchronously in the background and uses batch processing, computational resources can be effectively utilized, avoiding delays in real-time task dispatch. After training, the newly obtained network parameters do not take effect immediately but are loaded and activated at the start of the next new user session, thus ensuring a smooth transition and stability of model updates.
[0095] Meanwhile, the online update mode aims to enable the reinforcement learning decision model to quickly respond to real-time user feedback in the current session. After each round of user interaction and the collection of user feedback signals, the system immediately performs incremental training. This training is done in small batches at high frequency to fine-tune model parameters to adapt to the current user's preferences or session context. To balance real-time response and model stability, the deep Q-network maintains two sets of parameters: one set of "decision parameters" used for actual decision-making, and the other set of "training parameters" used for online incremental training. The core of this double-buffering mechanism is to decouple real-time decision-making from the training process. The decision parameters are used for current task assignment, while the training parameters are incrementally updated in the background. To avoid frequent and potentially unstable online training results directly affecting real-time decision-making, the system introduces an L2 norm difference threshold. Training parameters are only synchronized to decision parameters when the L2 norm difference between the training parameters and decision parameters (i.e., the Euclidean distance between the two parameter vectors) exceeds a preset threshold. This ensures that only when the training parameters undergo sufficiently significant changes that may lead to improvement are they adopted as new decision parameters, thereby filtering out minor or unstable updates and improving the robustness of the decision.
[0096] Furthermore, in online update mode, to ensure the validity of training data and avoid the model learning incomplete state transition information, the system explicitly excludes interactive experiences from the current round that have not yet completed user feedback collection or task execution during incremental training. This means that experience is only used for online training after a complete state-action-reward-next state sequence is confirmed. This effectively prevents the model from learning based on uncertain or incomplete feedback, thereby improving the accuracy and reliability of online updates. Meanwhile, the learning rate is a key hyperparameter controlling the step size of model parameter updates. The learning rate for online updates is set to one-tenth of the offline update learning rate, a conservative strategy. A smaller learning rate allows for more fine-grained and slower adjustments to model parameters during online incremental training, reducing the risk of overfitting when faced with single or limited real-time feedback. This design helps maintain the stability of the overall model strategy, avoiding drastic fluctuations in model behavior due to local or transient feedback, ensuring both online adaptability and the model's generalization ability.
[0097] It should be noted that the above description describes some embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than that shown in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0098] like Figure 2 As shown, based on the same inventive concept, and corresponding to the methods of any of the above embodiments, this application also provides a multi-agent collaborative task dispatching system based on recurrent perceptual reinforcement learning and user feedback, including: The central control agent is used to receive user requests and perform intent recognition, dispatch decision-making, loop detection, and policy updates. The central control agent includes an intent recognition module, a confidence evaluator, a loop detector, a circuit breaker controller, a policy selector, and an execution trajectory recorder.
[0099] For example, the intent recognition module uses the BERT classification model, takes the user's request text as input, and outputs the probability distribution and highest confidence level of each sub-Agent.
[0100] Confidence estimator: Determines whether the highest confidence level is lower than the preset threshold θ_conf (typical value 0.6). If it is lower than the threshold, clarification mode is triggered; if it is higher than the threshold, the dispatch decision module is entered.
[0101] Cycle detector: Queries the session memory to determine whether the semantic hash of the current request has generated consecutive negative feedback within the session, and returns the negative feedback count K.
[0102] Circuit breaker controller: When K ≥ MAX_NEG (typical value 3), the circuit breaker is triggered, and manual takeover or an error message is returned. Automatic dispatching will no longer be performed.
[0103] Policy selector: Determines whether to adopt a normal policy, an exploration policy, or a forced exploration policy based on the current state, and calls the reinforcement learning module to select a specific action.
[0104] Execution Tracker: Writes the dispatched data (request hash, selected sub-Agent, user feedback) to the session memory.
[0105] Multiple sub-Agents are used to receive tasks dispatched by the central control Agent, execute the corresponding business logic, and return the execution results. The user feedback collection module is used to collect explicit or implicit user feedback on the execution results and map the feedback into corresponding reward signals.
[0106] The reinforcement learning module is used to maintain the deep Q-network and the target Q-network, and to perform network training and policy updates based on state transition data.
[0107] The session memory is used to store request semantic hashes, dispatch trajectory records, negative feedback counts, and historical dispatch sub-Agent identifiers at the session level.
[0108] The system described above is used to implement the corresponding methods in the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0109] Based on the same inventive concept, this application also provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned multi-agent cooperative task dispatching method based on cyclic perceptual reinforcement learning and user feedback, or runs the above-mentioned multi-agent cooperative task dispatching system based on cyclic perceptual reinforcement learning and user feedback.
[0110] Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the above-described multi-agent cooperative task dispatching method based on cyclic perceptual reinforcement learning and user feedback, or runs the above-described multi-agent cooperative task dispatching system based on cyclic perceptual reinforcement learning and user feedback.
[0111] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multi-agent collaborative task dispatching method based on cycle perception reinforcement learning and user feedback, characterized in that, include: Receive the request text input by the user and obtain the current session identifier; Perform normalization processing on the request text to generate a request semantic hash value; The session memory is queried using the session identifier and semantic hash value as indexes to obtain the historical negative feedback count of the request in the current session and the identifier of the most recently dispatched sub-Agent; Based on the historical negative feedback count, a circuit breaker judgment is performed. When the circuit breaker condition is met, a preset prompt message is returned, and the current dispatch process is terminated. For requests that do not trigger the circuit breaker, perform intent identification and output the dispatch probability distribution and maximum confidence level for each sub-Agent; Based on the maximum confidence level, a threshold judgment is performed. When the confidence level is lower than the preset threshold, the clarification mode is entered. After receiving the user's explicit selection, the application is dispatched to the corresponding sub-Agent. For requests that meet the confidence requirements, a state vector integrating multi-dimensional features is constructed. Based on the historical negative feedback count, the corresponding dispatch strategy is matched, the target sub-Agent is determined, and the corresponding task is forwarded for execution. Collect user feedback signals on task execution results and map the feedback signals into quantified reward values; Update the negative feedback count and distribution trajectory record in the session memory bank based on the quantified reward value; The state transition quadruple is stored in the experience pool, and the network parameters of the reinforcement learning decision model are updated based on the experience replay mechanism.
2. The multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback as described in claim 1, characterized in that, The generated request semantic hash value includes: The request text is processed sequentially by removing stop words, lowercase conversion, and stemming. The standardized text is segmented to obtain a word sequence, and the TF-IDF weight of each word is calculated. The term frequency is the ratio of the number of times the word appears in the current request to the total number of words in the current request, and the inverse document frequency is the logarithm of the ratio of the total number of historical requests in the system to the number of historical requests containing the word. The MurmurHash3 hash function is used to map each word to an L-bit binary vector, and the binary bits are converted into {-1,+1} encoding to obtain the word hash vector; Based on the TF-IDF weights, a bit-by-bit weighted summation is performed on the hash vectors of all words to obtain an L-dimensional real-number summation vector; Perform a sign determination on each bit of the accumulated vector; set the bit to 1 if the bit value is not less than 0, and set it to 0 if the bit value is less than 0, to obtain an L-bit SimHash value. The semantic similarity of requests is determined based on the Hamming distance between two SimHash values. If the Hamming distance does not exceed a preset distance threshold, the two requests are considered to be semantically similar. The Hamming distance is the sum of the number of bits that are different in corresponding positions in two binary strings of equal length.
3. The multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback as described in claim 1, characterized in that, The circuit breaker mechanism consisting of the circuit breaker determination and the confidence threshold determination includes: Secondary negative feedback circuit breaker: When the historical negative feedback count is greater than or equal to the preset circuit breaker threshold, the circuit breaker is triggered and a preset prompt message is returned, terminating the current automatic dispatch process; Level 1 Confidence Circuit Breaker: When the maximum confidence level of the intent recognition output is lower than the preset confidence threshold, the system enters clarification mode and returns a clarification question to the user. After receiving the user's explicit selection, the system is directly dispatched to the corresponding sub-Agent. The execution timing of the secondary negative feedback circuit breaker is earlier than that of the primary confidence circuit breaker. When both triggering conditions are met simultaneously, the secondary negative feedback circuit breaker is executed first.
4. The multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback as described in claim 1, characterized in that, The construction of the state vector that integrates multi-dimensional features includes: The state vector is obtained by concatenating the BERT embedding vector of the request text, the maximum confidence of intent recognition, the historical negative feedback count, the continuous negative feedback flag, and the one-hot encoding of the previous round of dispatched sub-Agents. The total dimension is 768+3+K, where K is the total number of sub-Agents in the system.
5. The multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback according to claim 1, characterized in that, The matching of the corresponding dispatch strategy based on the historical negative feedback count includes: When the historical negative feedback count is 0 and the previous round of feedback is non-negative feedback, the ε-greedy strategy is adopted, selecting the sub-Agent with the largest Q value with a probability of 1-ε, and randomly selecting the target sub-Agent from all sub-Agents with a probability of ε. When the historical negative feedback count is greater than or equal to 1, switch to the forced exploration strategy; When positive feedback is received from the user, the historical negative feedback count is cleared and the strategy is switched back to ε-greedy. When negative feedback is received from the user, the historical negative feedback count is incremented by 1. When the historical negative feedback count reaches the circuit breaker threshold, the secondary negative feedback circuit breaker is triggered.
6. The multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback according to claim 5, characterized in that, The forced exploration strategy employs a semantically dissimilar exploration mode, including: During the system initialization phase, the Sentence-BERT embedding vector of the function description text of each sub-Agent is pre-calculated and stored in the sub-Agent registry; Calculate the cosine distance between the functional embedding vector of the most recently failed sub-Agent and the functional embedding vectors of the remaining candidate sub-Agents; The target sub-Agent is selected with the largest cosine distance using a first preset probability, and the target sub-Agent is randomly selected from all candidate sub-Agents with equal probability using a second preset probability.
7. The multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback according to claim 1, characterized in that, The reinforcement learning decision model employs a dual-tower deep Q-network structure, including: Text feature pyramid: The input 768-dimensional request text BERT embedding vector is processed sequentially through 256-dimensional and 128-dimensional fully connected layers, using the ReLU activation function, with the first fully connected layer set to a Dropout rate of 0.1; Scalar Feature Tower: Input 3+K dimensional scalar concatenation features, processed through a 64-dimensional fully connected layer, and activated by the ReLU function; Fusion layer: The text feature tower output and the scalar feature tower output are concatenated into a 192-dimensional fusion vector, which is then processed by a 128-dimensional fully connected layer to output a K-dimensional Q-value vector, with each dimension corresponding to the state and action value of a sub-Agent; Maintain a target Q-network with the exact same structure as the deep Q-network, and copy the parameters of the deep Q-network to the target Q-network every preset number of training steps; update the parameters of the deep Q-network based on the mean squared error loss function.
8. The multi-agent cooperative task dispatching method based on recurrent perceptual reinforcement learning and user feedback according to claim 7, characterized in that, The experience pool adopts a fixed-capacity circular buffer structure, including: Add a corresponding environment version number to each experience, and filter out historical experiences that are inconsistent with the current sub-Agent configuration version during training sampling; A time-sensitivity weight is added to each experience, and the time-sensitivity weight decays exponentially based on the number of training steps after the experience is stored. During sampling, weighted sampling is performed according to the time-sensitivity weight. The reward value is clipped to the range of [-2, +2], and the Huber loss is used instead of the mean squared error loss to reduce the impact of abnormal reward values on gradient updates. The experience pool is divided into a normal experience area, a loop experience area, and a circuit breaker experience area, which respectively store the interaction experience of normal state, error loop state, and circuit breaker trigger. During training sampling, samples are drawn from the three partitions in a ratio of 5:3:
2.
9. The multi-agent collaborative task dispatching method based on recurrent perceptual reinforcement learning and user feedback according to claim 7, characterized in that, The updated reinforcement learning decision model supports both offline and online update modes. In offline update mode, batch training is performed in a background thread after the session ends, and the network parameters obtained from the training take effect when the next new session is initialized. In the online update mode, incremental training is performed after each round of user feedback collection. The deep Q network maintains a double buffer for decision parameters and training parameters. The training parameters are synchronized to the decision parameters only when the difference between the L2 norms of the two sets of parameters exceeds a preset threshold. The sampling range for online training excludes interactive experiences that have not been completed in the current round, and the learning rate for online updates is set to 1 / 10 of the learning rate for offline updates.
10. A multi-agent cooperative task dispatching system based on recurrent perceptual reinforcement learning and user feedback, used to execute the multi-agent cooperative task dispatching method based on recurrent perceptual reinforcement learning and user feedback as described in any one of claims 1-9, characterized in that, include: The central control agent is used to receive user requests and perform intent recognition, dispatch decision-making, loop detection, and policy updates. The overall control agent includes an intent recognition module, a confidence evaluator, a loop detector, a circuit breaker controller, a policy selector, and an execution trajectory recorder; Multiple sub-Agents are used to receive tasks dispatched by the central control Agent, execute the corresponding business logic, and return the execution results. The user feedback collection module is used to collect explicit or implicit user feedback on the execution results and map the feedback into corresponding reward signals. The reinforcement learning module is used to maintain the deep Q-network and the target Q-network, and to perform network training and policy updates based on state transition data; The session memory is used to store request semantic hashes, dispatch trajectory records, negative feedback counts, and historical dispatch sub-Agent identifiers at the session level.