Multi-agent self-adaptive questioning method and device

By introducing ReAct architecture and reinforcement learning mechanism into the large language model, the agent is trained to ask questions independently, which solves the problem that the large language model falls into thinking dilemma in complex tasks, and improves the success rate of task completion and planning abstraction ability.

CN120371964APending Publication Date: 2025-07-25CHONGQING XIANXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510464371.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing large language model agents cannot ask questions and questions independently when dealing with complex tasks, resulting in a dilemma of thinking and being unable to effectively learn and complete tasks.

Method used

Add ReAct architecture on the basis of the general big model, build a multi-agent reinforcement learning enhancement large language model, train the agent to ask questions independently through the reinforcement learning mechanism, including single-agent and multi-agent training, and design reward function to optimize the timing, frequency and content of the question.

Benefits of technology

The agent can ask questions independently, avoid falling into thinking dilemma, improve the success rate of task completion and the ability to abstract planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371964A_ABST
    Figure CN120371964A_ABST
Patent Text Reader

Abstract

The invention relates to the field of multi-agent reinforcement learning, in particular to a multi-agent self-adaptive questioning method and device.The method comprises the steps that information input from the outside is sent to a built reinforcement learning enhanced large language model, and a proposed question is output; wherein the reinforcement learning enhanced large language model comprises a general large model and a LoRA framework; and the reinforcement learning enhanced large language model is trained through a reinforcement learning mechanism, so that the large model agent can propose a problem which is more practical, and the ability of self-planning abstraction and the success rate of task completion are further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-intelligent reinforcement learning, and in particular to a multi-agent adaptive questioning method and device. Background Art

[0002] In the field of large language models, there are currently three types of improvement directions: The first type of improvement direction focuses on improving the planning ability of large model agents. By designing the combination of the state modules of the internal multi-turn chain of thought (Multi-turn CoT), the reasoning and task completion abilities of large model agents are enhanced. The ReAct architecture (ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al., 2022) endows large language models with the ability to perform intelligent actions (Action); by using external tools, the lack of knowledge in the large model itself is supplemented. The Reflexion architecture (Reflexion: Language Agents with Verbal Reinforcement Learning, Noah Shinn et al., 2023) also introduces a reflection module for agents. Through reflecting on their past behaviors, agents enhance their reasoning abilities. The second type of improvement direction is to enhance the interaction ability between agents and humans by designing the language dialogue (Dialog) mode of the agents themselves, especially in the field of robotics. For example, the Language to Action architecture (Language to Action: Towards Interactive Task Learning with Physical Agents, Joyce Y. Chai et al., 2018) converts the dialogue language between agents and humans into the actual actions of the agents, enabling the agents to understand human language through human-computer interaction and abstract it into corresponding actions.

[0003] The third type of improvement direction is to implement the learning of multi-turn interactions of large model agents through reinforcement learning algorithms. For example, the Archer framework (Archer: Training language model agents via hierarchical multi-turn rl, Yifei Zhou et al., 2024) uses reinforcement learning to control the multi-turn interaction actions of the minimum semantic units (Tokens) of large language models.

[0004] The first type of improvement direction assumes that the agent can only observe the external environment through the internal thought chain mode and cannot interact with the external environment independently, which limits the ability of the large model agent to interact with the outside world. This "head-down hard work" mode also makes it impossible for the large model agent to break out of the logical dilemma of its existing thought chain when dealing with complex tasks, ultimately resulting in a large number of ineffective iterations and task failures. True autonomy is reflected in the large model agent's ability to understand tasks through its own cognition, question tasks, and interact with the outside world through question-and-answer to solve the doubts, and then continuously learn and complete the tasks.

[0005] The second type of improvement direction adopts an external question-and-answer human-computer interaction mode. However, this mode cannot be used as an internal thought chain and is more of an external multi-turn chat mode (Multi-Turn Chat). Therefore, this mode cannot control the agent in terms of actions, and the question-and-answer interaction method shown by the agent in this mode lacks the ability of learning in abstraction.

[0006] The third type of improvement direction, from the perspective of the model architecture of the large model, controls the generation of tokens at the language level of multi-turn interaction (Multi-Turn Interaction) through reinforcement learning. Current research also attempts a single-agent state module (i.e., converting the control at the level of the smallest semantic unit (Token) into the control of a complex internal thought chain). And no work has attempted to apply this control in the field of multi-agent interaction. Summary of the Invention

[0007] Regarding the problems existing in the three existing improvement directions in the field of large language models, from the perspective of the internal thought chain of the large model agent, on the basis of the general large model, the ReAct architecture is added to form a multi-agent reinforcement learning enhanced large language model. A multi-agent adaptive questioning method and device are proposed. When interacting with humans, if the agent encounters problems during the task completion process, it can ask questions autonomously and improve its own thought chain in real time through external feedback, and will not fall into its own thought dilemma, further enhancing the agent's ability of planning and abstraction.

[0008] To achieve the above concept, the technical solution proposed by the present invention is: A multi-agent adaptive questioning method, comprising the following steps: Send the information input from the outside world to the constructed reinforcement learning enhanced large language model, and output the questions proposed; wherein, the reinforcement learning enhanced large language model includes a general large model and a LoRA architecture; the reinforcement learning enhanced large language model is trained through a reinforcement learning mechanism.

[0009] Further, the training of the reinforcement learning enhanced large language model includes single-agent training and multi-agent training.

[0010] Further, the single-agent training includes: the general large model interacts with people to generate an expert dataset, and uses the expert dataset for offline training.

[0011] Further, the multi-agent training includes offline-to-Online training and Online training. The offline-to-Online training is used to enable the Worker agent to self-explore, and the Online training is used to train the Supervisor.

[0012] Further, the reinforcement learning enhanced large language model includes a sentence-level value network and a minimum semantic unit-level network, which are used to convert the sentence-level value network into the minimum semantic unit-level network.

[0013] Further, converting the sentence-level value network into the minimum semantic unit-level network includes converting sentence-level information into minimum semantic unit-level information. The conversion formula for converting sentence-level information into minimum semantic unit-level information is ; where is a high-level sentence, au is the complete sentence, represents the current state, t is the timestamp, represents the Actor network policy. The right side of the formula is the underlying minimum semantic unit in the form of the product of each minimum semantic unit. is each minimum semantic unit, l represents the low level, represents the action to the last minimum semantic unit of the sentence. i is the serial number of the minimum semantic unit, and k is the total number of minimum semantic units.

[0014] Further, the sentence-level value network includes a Critic network, and the minimum semantic unit-level network includes an Actor network.

[0015] Further, converting the sentence-level value network into the minimum semantic unit-level network: bringing the minimum semantic unit into the Actor-Critic architecture to obtain the minimum semantic unit-level Actor network.

[0016] Further, the calculation formula of the minimum semantic unit-level Actor network is: ; Further, the reward function used by the reinforcement learning mechanism to train the reinforcement learning enhanced large language model is: ; Among them, rV(au(t), S(t)) is the question timing factor, rC(au(t), S(t)) is the question content factor, rT(au(t), S(t)) is the question frequency factor, and rF(au(t), S(t)) is the task completion factor. For this complete sentence, It represents the current state.

[0017] Furthermore, the calculation formula for the question timing factor is: ; Among them, V(S(t)) is the value, representing the current degree of dilemma. For this complete sentence, It represents the current state, exp is the exponential expression, t is the timestamp, representing the first weight coefficient of the agent reward function.

[0018] Furthermore, the calculation formula for the question content factor is: , where d((au(t), H)) is the distance and H is the historical question content.

[0019] Furthermore, the calculation method for the question frequency factor includes: the action of the sentence "Asking" appears twice consecutively, and the question frequency factor , where φ represents the third weight coefficient of the agent reward function.

[0020] Furthermore, the calculation formula for the task completion factor is: ; Among them, For this complete sentence, It represents the current state, representing the fourth weight coefficient of the agent reward function.

[0021] Furthermore, use a central supervisor agent to replace the person in the human-computer interaction for training and actual measurement, forming an architecture with the central supervisor agent replacing the person as the center and other workers distributed.

[0022] A multi-agent adaptive question device includes at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of the above.

[0023] A storage medium having instructions stored thereon that are executable by a processor, and when the instructions are executed by the processor, the processor is caused to execute the method described in any one of the above.

[0024] The beneficial effects brought about by the improvement of the technical solution of the present invention are as follows: By training the reinforcement learning enhanced large language model through a reinforcement learning mechanism, the large model agent can propose a more practical problem, thereby further enhancing its own ability of planning abstraction and the success rate of task completion. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is an overall architecture diagram of the concept of a multi-agent adaptive questioning method; Figure 2 are the steps of single-agent - multi-agent synchronous and simultaneous training; Figure 3 is a snippet of pseudocode during the training process Figure 1 ; Figure 4 is a snippet of pseudocode during the training process Figure 2 ; Figure 5 A schematic diagram of the conversion from the sentence level to the smallest semantic unit level; Figure 6 is an overall architecture diagram of the concept of a multi-agent adaptive questioning method; Figure 7 is an example of the Alfworld dataset; Figure 8 is the result under the ReAct framework; Figure 9 is the result of a specific example using the method of the present invention; Figure 10 A schematic diagram of a multi-agent reinforcement learning device. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present application will be further described in detail below in conjunction with test examples and specific embodiments. However, this should not be construed as limiting the scope of the above subject matter of the present application to the following embodiments. All technologies implemented based on the content of the present application fall within the scope of protection of the present application.

[0027] Unless otherwise specified, in the description of the specific embodiments of the present application, the expression terms indicating the orientation or positional relationship such as "upper", "lower", "left", "right", "center", "inner", "outer", "side", etc. are all based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship when the product / device / device is usually used. These terms of orientation or positional relationship are only for the convenience of describing the solution of the present application or simplifying the description in the specific embodiments, so as to facilitate technicians to quickly understand the solution, rather than indicating or implying that a specific device / component / element must have a specific orientation, or be constructed and operated in a specific positional relationship. Therefore, it should not be construed as a limitation to the present application.

[0028] In the description of the embodiments of the present application, technical terms such as "first" and "second" only distinguish one entity or operation from another entity or operation, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "plurality" is two or more, unless otherwise specifically limited.

[0029] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0030] Embodiment 1 A multi-agent adaptive questioning method, the overall architecture diagram of the concept is as Figure 1 shown, mainly including the following steps: Send the information input from the outside world to the constructed reinforcement learning enhanced large language model, and output the questions proposed; wherein, the reinforcement learning enhanced large language model includes a general large model and a LoRA architecture; train the reinforcement learning enhanced large language model through a reinforcement learning mechanism, and the introduction of the LoRA architecture and the training of the reinforcement learning enhanced large language model optimize the timing, frequency and content relevance of the questions proposed. Among them, the general large model can be any open-source general large model. In the LoRA architecture, B = 0, A = N(0, ), A represents the A matrix (dimensionality reduction matrix); B represents the B matrix (dimensionality increase matrix). The general large model can also be any closed-source large model. Figure 1 Only the open-source general large model is used as an example for display herein, and it is not a limitation to the present invention.

[0031] Preferably, the training of the reinforcement learning enhanced large language model includes single-agent training and multi-agent training. Single-agent training is for the base large model to interact with people, generate an expert dataset, and use the expert data to perform offline training on the agent to enhance the ability to ask questions. Multi-agent training (where a / b are trained simultaneously): a. Offline-to-Online enables the Worker agent to explore itself and further strengthen the parameters in the first step. b. Online training of the Supervisor to train its generalization ability to answer questions. The steps of single-agent - multi-agent synchronous and simultaneous training are as Figure 2 shown, and the screenshot of the pseudocode during the training process is as Figure 3 and Figure 4 shown.

[0032] Preferably, in the reinforcement learning enhanced large language model, a sentence-level value network (Critic network) and a minimum semantic unit-level Actor network are defined. The sentence-level value network and the minimum semantic unit-level Actor network embody the idea of hierarchical reinforcement learning. At a high level (High-level), the sentence-level value network is responsible for generating a complete sentence (i.e., an action Action) based on the information input from the outside world. For example, a complete high-level sentence "Go to the room and pick up the object" can be disassembled into low-level minimum semantic units: "Go", "room", "pick up", "object". Since the latest semantic unit (token i) generated by the large model is the result of stacking the probabilities of all tokens, and the result of stacking the probabilities of all tokens is used to represent the semantic features of the high-level (High-level) sentence, therefore, the high-level sentence is gradually converted into low-level minimum semantic units through the conversion formula from the sentence level to the minimum semantic unit level. The schematic diagram of the conversion from the sentence level to the minimum semantic unit level is as Figure 5 shown, and the conversion formula from the sentence to the minimum semantic unit is: (1); where, is the high-level sentence, is the complete sentence, represents the current state, t is the timestamp, represents the Actor network policy, is the action of each minimum semantic unit, l represents the low level (low-level), Denote the action to the last minimal semantic unit of the sentence. Here, i is the serial number of the minimal semantic unit, and k is the total number of minimal semantic units. The right side of the formula equals the low-level minimal semantic units, and the high-level sentence is represented in the form of the product of each minimal semantic unit. After converting the sentence into minimal semantic units through formula (1), the minimal semantic units are input into the Actor-Critic architecture to obtain the Actor network at the minimal semantic unit level: (2); Among them, π θ represents policy, represents the low-level action, S(t) represents the current state; and δ behind π represents the representation method of the Advantage function in reinforcement learning. The critic network only judges at the sentence level, which is a judgment of the sentence level in the Actor network. Therefore, it remains at the sentence level. The input is the Act at the sentence level of the Actor network. The value network at the sentence level (Critic network) is used to judge , and no low-level conversion is required. Therefore, the value Critic network remains at the sentence level (high-level), that is, a sentence is an Action: (3) The trained Actor network is the LoRA architecture. The Critic network is not used after training. The Critic network is only used to judge the quality of the Actor's actions. After the Actor network becomes LoRA, it is embedded in the general large model and provides services together with the general large model. At this time, this Actor network has the ability of adaptive Asking, and can not only generate sentences from the perspective of minimal semantic units, but also make judgments on sentences.

[0033] Furthermore, the present invention designs a reward function for the agent to enable the agent to learn how to ask questions. The reward function is used to train the Actor-Critic framework. When designing the reward function, it is considered from four aspects: the content of the question, the timing of the question, the frequency of the question, and whether the task is completed. The formula of the reward function is: (4). Among them, rV(au(t),S(t)) is the timing factor of the question, which is used to reward the agent for asking a question in a difficult situation (timing), (5) Among them, V(S(t)) is the value, representing the current degree of dilemma, passing in the current state S(t), exp is the exponential expression, and t is the timestamp. It represents the first weight coefficient of the agent reward function.

[0034] The design of the question-asking timing factor uses the theory of Reward Shaping. By passing in the value of the value to judge whether it is a "dilemma" at present. If it is in a dilemma, use the Asking action, that is, Provide rewards.

[0035] Among them, rC(au(t), S(t)) is the question content factor. , d((au(t), H)) is the distance, and H is the historical question content. This part is mainly used to make the reward agent ask questions that meet the task and do not repeat questions. λ represents the second weight coefficient of the agent reward function. Therefore, if the question content of the agent is similar to the question of the task, that is, the distance between the vector of the question and the vector of the question asked is close enough, we provide a positive reward d(au(t), q). If the question proposed by the agent exists in his historical question space, that is, repeated questions, we provide a negative reward, that is, −d(au(t), H). So that the agent can be improved in terms of question content.

[0036] Among them, rT(au(t), S(t)) is the question frequency factor. It is used to punish the agent for continuous question-asking and optimize the question-asking frequency. φ represents the third weight coefficient of the agent reward function. That is, if the action of Asking appears in two consecutive sentences, we define this as too high a question-asking frequency and punish the agent. .

[0037] Among them, rF(au(t), S(t)) is the task completion factor. ; rF(au(t), S(t)) rewards for completing the task, that is, getting rewards for completing the task, otherwise not getting rewards. Among them, represents the fourth weight coefficient of the agent reward function, F represents whether the task is completed, 1 means the task is completed, and 0 means the task is not completed.

[0038] Through the design of the reward function r(au(t), S(t)), the training of the question-asking timing, frequency, and content relevance enables the agent to understand how to ask questions and how to ask good questions. The weight coefficients of the agent reward function Are determined according to the actual training situation.

[0039] Embodiment 2 Based on the reinforcement learning-enhanced large language model of the present invention, a further improved solution is given on the basis of Embodiment 2. The human-computer interaction mode of the reinforcement learning-enhanced large language model is improved, and the participation of humans plays an absolute guiding and controlling role for the agent. However, it also consumes a huge amount of manpower. In the supervisor-worker mode in the prior art, the human in the human-computer interaction is not used as a supervisor. In this embodiment, a central supervisor agent is used to replace the human in the human-computer interaction for training and actual measurement, forming an architecture centered on the central supervisor agent that replaces humans, with other workers distributed. The worker agents are responsible for actual tasks and ask questions to the supervisor agent when necessary.

[0040] For the supervisor, the supervisor needs to monitor the status of all workers and answer the questions raised by the workers. However, the workers only need to focus on their own work and do not need to care about the work of other workers or the supervisor. Therefore, the present invention assumes that the multi-agent workers are only responsible for their own tasks and do not need to observe any other agents. In the actor-critic network, the states / observations (S(t) / O(t)) of all people need to be passed to the supervisor. The gradient policy of the low-level Actor network of the minimum semantic unit level of the agent can be generally expressed as: (6) In a certain agent of the multi-agent, it can be either a supervisor or any worker. And the state of the actor critic of the worker is the same as the state when it is a single worker. The above formula (6) is the gradient expression of the Actor network. If it is a supervisor, the simplified formula (6) will have the states of all worker agents; if it is a worker agent, the simplified one is the same as the minimum semantic unit level Actor network in the above text.

[0041] In order to enable the supervisor to better answer the questions of the workers, the present invention also designs a reward function for the supervisor to improve their ability to answer the questions of the workers: The reward function of the supervisor can be designed as: (7) In this function: The reward is calculated based on the distance between the supervisor's observation Op(t) and the answer au(t), which is used to reflect the supervisor's understanding of the current situation of the worker. Here, the supervisor's observation is the question of the worker. Therefore, we hope that the answer of the supervisor is closer to the content of the worker's question to avoid the situation where the supervisor answers irrelevantly.

[0042] Reward the supervisor's response au(t) based on its proximity to the target solution G (assuming the supervisor has access to the correct trajectory or answer for the task). Thus, we hope that the supervisor's response is close to the correct path that the worker should take, thereby enhancing the supervisor's guidance for the worker.

[0043] Penalize the score for the exact match between the supervisor's response and the expected solution G, aiming to prevent the supervisor from directly providing the answer to the worker without any guidance.

[0044] Penalize when the supervisor's response length exceeds 100 tokens to prevent the supervisor from over-responding.

[0045] Similarly, Let be the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient of the supervisor agent reward function respectively. The four weight coefficients of the supervisor agent reward function are determined according to the actual training situation.

[0046] Supervised learning is carried out in the way of the central supervisor-distributed worker multi-agent model, and experiments have proved its effectiveness, providing a theoretical basis for the construction of the model cluster of "small models at the edge - large models at the center" in the future.

[0047] A specific overall framework flowchart of the multi-agent adaptive questioning method is as Figure 6 shown.

[0048] Example 3 An example of a specific method of the present invention is an instance for solving hallucinations caused by large models (from the Alfworld public dataset).

[0049] The dataset is as Figure 7 shown. It can be seen from Figure 7 that in this instance, the worker large language agent needs to find soap (an object) in the room, then go to the sink to clean the soap, and finally put the cleaned soap back into the cabinet. Figure 5 shows this task and attaches a feasible path. Besides the countertop, the soap may be among other objects, so the worker agent needs to explore in the room.

[0050] Figure 8 Shows the results under the ReAct framework. After the exploration fails, the agent falls into hallucinations caused by the large model, believing that there is no soap in the room and stopping to continue the task (no longer choosing actions), resulting in the failure of the task.

[0051] Figure 9Shows the results of a specific example under the framework of the present invention. After three consecutive initial exploration failures, the worker agent immediately autonomously asks the outside world based on the current situation and obtains feedback from the human / supervisor agent. After continuing the exploration without finding the object, the agent asks questions autonomously again and obtains feedback from the human / supervisor agent once more. According to the feedback clues, the agent successfully finds the object, returns to the correct path, and completes the task within the limited number of steps. It can be seen that after learning the questioning skills, the agent can accurately grasp the timing of asking questions (asking immediately), the content (i.e., elaborating on the current predicament and not asking repeated questions), the frequency (avoiding excessive questions), and complete the task through new clues.

[0052] Figure 9 Among them, pink represents exploration (incorrect / non-optimal) actions; orange represents the agent's question-asking actions and content; blue represents the feedback from the human / supervisor agent after the agent's question-asking actions; green represents correct actions.

[0053] Example 4 Please refer to Figure 10 , Figure 10 which is a schematic diagram of the multi-agent reinforcement learning device provided by the embodiments of the present application. The multi-agent reinforcement learning device may include: A processing module, configured to send the information input from the outside world to the constructed reinforcement learning enhanced large language model and output the questions raised.

[0054] It should be understood that when each module of the multi-agent reinforcement learning device provided in the above embodiments performs processing, only the division of each functional module in the above description content is used as an example. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0055] Each functional module in the above embodiments may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present application.

[0056] Example 5 Based on the same inventive concept, this embodiment also provides a computer device, which may include a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method described in any one of the above embodiments 1-4.

[0057] Embodiment 6 Based on the same inventive concept, this embodiment also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method described in any one of the above-described Embodiments 1-4 is implemented.

[0058] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A multi-agent adaptive questioning method, characterized in that, It includes the following steps: Send the information input from the outside world to the constructed reinforcement learning enhanced large language model, and output the questions proposed; wherein, the reinforcement learning enhanced large language model includes a general large model and a LoRA architecture; Train the reinforcement learning enhanced large language model through a reinforcement learning mechanism.

2. The multi-agent adaptive questioning method according to claim 1, wherein The training of the reinforcement learning enhanced large language model includes single-agent training and multi-agent training.

3. The multi-agent adaptive questioning method according to claim 2, wherein The single-agent training includes: the general large model interacts with people to generate an expert dataset, and uses the expert dataset for offline training.

4. The multi-agent adaptive questioning method according to claim 2, characterized in that, The multi-agent training includes offline-to-Online training and Online training. The offline-to-Online training is used to enable the Worker agent to explore itself, and the Online training is used to train the Supervisor.

5. The multi-agent adaptive questioning method according to claim 2, wherein The reinforcement learning enhanced large language model includes a sentence-level value network and a minimum semantic unit-level network, which are used to convert the sentence-level value network into the minimum semantic unit-level network.

6. The multi-agent adaptive questioning method according to claim 5, characterized in that Converting the sentence-level value network into the minimum semantic unit-level network includes converting sentence-level information into minimum semantic unit-level information. The conversion formula for converting sentence-level information into minimum semantic unit-level information is: Among them, is a high-level sentence, a u is the complete sentence, represents the current state, t is the timestamp, represents the Actor network policy, etc. The right side of the formula is the lowest-level minimum semantic unit, in the form of the product of each minimum semantic unit, is each minimum semantic unit, l represents the low level, represents the action to the last minimum semantic unit of the sentence, i is the serial number of the minimum semantic unit, and k is the total number of minimum semantic units.

7. The multi-agent adaptive questioning method according to claim 5, wherein, The sentence-level value network includes a Critic network, and the minimum semantic unit-level network includes an Actor network.

8. The multi-agent adaptive questioning method according to claim 6, wherein, Converting the sentence-level value network into the minimum semantic unit-level network: bringing the minimum semantic unit into the Actor-Critic architecture to obtain the minimum semantic unit-level Actor network.

9. The multi-agent adaptive questioning method according to claim 8, wherein, The calculation formula of the minimum semantic unit-level Actor network is: 。 10. A multi-agent adaptive questioning method according to any one of claims 1-9, characterized in that, The reward function adopted by the reinforcement learning mechanism to train the reinforcement learning enhanced large language model is: ; Among them, rV(au(t), S(t)) is the question timing factor, rC(au(t), S(t)) is the question content factor, rT(au(t), S(t)) is the question frequency factor, and rF(au(t), S(t)) is the task completion factor. For this complete sentence, represents the current state.

11. The multi-agent adaptive questioning method according to claim 10, wherein The calculation formula of the question timing factor is: ; Among them, V(S(t)) is the value, representing the current degree of predicament. For this complete sentence represents the current state, exp is the exponential expression, and t is the timestamp. represents the first weight coefficient of the agent reward function.

12. The multi-agent adaptive questioning method according to claim 10, wherein The calculation formula of the question content factor is: , Wherein, d((au(t), H)) is the distance, and H is the historical question content.

13. A multi-agent adaptive questioning method according to claim 10, characterized in that, The calculation method of the question frequency factor includes: the action of Asking appears in two consecutive sentences, Question frequency factor , φ represents the third weight coefficient of the agent reward function.

14. A multi-agent adaptive questioning method according to claim 10, characterized in that, The calculation formula of the task completion factor is: ; Among them, is the complete sentence, represents the current state, represents the fourth weight coefficient of the agent reward function.

15. A multi-agent adaptive questioning method according to any one of claims 1-9, characterized in that, Use a central supervisor agent to replace the person in the human-computer interaction for training and actual measurement, and form an architecture with the central supervisor replacing the person as the center and other workers distributed.

16. A multi-agent adaptive questioning device, characterized in that It includes at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 15.

17. A storage medium, on which instructions executable by a processor are stored, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 15.