A multi-agent interaction method and device based on organization control and constraint adjustment

CN122837620APending Publication Date: 2026-09-29AIMAI TECHNOLOGY (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610663044.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0007]本发明要解决的技术问题是:使用现有技术回复的信息不一致导致无法产生有效或合规的谈判、采购及对话结果,和由于严格过滤导致无法生成复杂约束范围内的有效回复,以及无法灵活动态地交互导致产生不稳定行为的问题

Benefits of technology

本发明将“如何根据用户输入生成候选输出”与“候选输出是否可作为最终的输出”分离,从而能够按照用户需求对候选输出进行指定的约束,避免在无约束下执行高风险操作,回复的信息不一致导致产生无效或违规结果。通过约束强度参数,将现有技术的复杂静态规则转化为约束强度参数的参数化的连续控制面,使得约束连续且可调节,从而能够生成复杂约束范围内的有效回复,且通过调节约束强度参数,能够实现灵活动态交互,避免系统出现不稳定行为。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837620A_ABST
    Figure CN122837620A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer software, and provides a multi-agent interaction method and device based on organization control and constraint adjustment. The present application screens an agent set participating in interaction through a routing mechanism; each sub-agent in the agent set generates a candidate output based on user input; constraint evaluation is performed on the candidate output according to an adjustable constraint strength parameter, and a response result of the candidate output is obtained; comprehensive scoring is performed based on the response result, and an optimal output is determined according to the result of the comprehensive scoring to obtain an execution result of the optimal output. The present application solves the problems that inconsistent information using the prior art reply cannot produce effective or compliant negotiation, procurement and dialogue results, strict filtering cannot generate effective replies within a complex constraint range, and unstable behavior is produced due to unresponsive interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer software technology, and in particular to a multi-agent interaction method and apparatus based on organizational control and constraint regulation. Background Technology

[0002] In task scenarios with real economic impact (such as negotiation, procurement, and conversational commerce), multi-agent systems built on large language models are increasingly widely used. In these application scenarios, intelligent agent clusters collaboratively complete the entire process management, including product retrieval, preference acquisition, price comparison, commitment generation, and after-sales support.

[0003] However, intelligent agents are increasingly becoming proactive: they can anticipate user needs, invoke external tools, and initiate operations autonomously without human prompting. Without structured oversight mechanisms, proactive agents may make binding commitments, trigger irrevocable transactions, or violate platform policies before human review, severely impacting actual business operations. Existing multi-agent systems, when applied to negotiation, procurement, and conversational business scenarios, primarily suffer from the following problems: Inconsistent information from multiple buyers or sellers can prevent the formation of effective or compliant negotiation, procurement, and dialogue outcomes. For example, multiple buyers may submit independent quotes to the seller, leading to price confusion, contradictory information (e.g., one party claims a certain model is in stock while another claims it is not), conflicting commitments (e.g., a seller accepts the final transaction price from two buyers simultaneously), and triggering violations (e.g., a seller's quote is below cost, but the system violates regulations by directly accepting it, resulting in subsequent non-performance).

[0004] Existing technologies often only achieve strict filtering. For constraints in real-world application scenarios (e.g., negotiation or procurement rules manually specified by the user), effective outputs within complex constraints cannot be generated during negotiations, procurement, and dialogues. For example, if the system pre-sets "quotes must strictly not exceed the budget," and the buyer's budget is 5000 yuan while the seller's quotation is 5001 yuan, any quotation exceeding 5000 yuan will be prohibited, preventing the generation of a valid counter-offer, leading to a stalemate, and the system outputs an empty response. Another example is a system setting "supplier quotations cannot be lower than 10% of cost price." If a supplier is willing to sell at cost price (below the 10% threshold) due to clearance sales, the system will directly reject the quotation with no room for negotiation, causing the buyer to miss the optimal deal, resulting in no contract and wasted time and resources. Yet another example is a system stipulating that "any commitment must be manually confirmed before being sent." If the user's budget is on the edge of the limit, and the price commitment is involved, they cannot make any statement and can only reply "I cannot answer," ending the dialogue.

[0005] In negotiation, procurement, and dialogue processes, the lack of flexible dynamic adjustment mechanisms can easily lead to unstable behavior in the system. For example, if the negotiation rounds are fixed at 10, and both parties are close to reaching an agreement in the third round, but the system cannot terminate or adjust the pace in advance, the subsequent seven rounds will repeatedly confirm the same price, ultimately resulting in contradictory conclusions due to timeouts or redundant interactions (e.g., one party suddenly backing out). In procurement scenarios, if market prices suddenly fluctuate during the procurement process, fixed constraints (such as "quotes are valid for 24 hours") cannot adapt to the changes, causing previously accepted quotes to become invalid before they take effect, leaving the buyer and supplier in a state of constant cancellation and resubmission of quotes. Furthermore, if users continuously inquire about price details, the system can only re-query the database each round according to a fixed strategy, leading to a gradual increase in response delays.

[0006] Therefore, overcoming the shortcomings of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention

[0007] The technical problems to be solved by this invention are: inconsistent information in responses using existing technologies leads to the inability to produce effective or compliant negotiation, procurement, and dialogue results; strict filtering prevents the generation of effective responses within complex constraints; and the inability to interact flexibly and dynamically results in unstable behavior.

[0008] Firstly, a multi-agent interaction method based on organizational control and constraint regulation is provided, including: A set of intelligent agents participating in the interaction is selected through a routing mechanism; each sub-agent in the set of intelligent agents generates candidate outputs based on user input; The candidate output is constrained and evaluated based on the adjustable constraint strength parameter, and the response result of the candidate output is obtained. A comprehensive score is performed based on the response results, and the optimal output is determined according to the comprehensive score results to obtain the execution result of the optimal output.

[0009] Further, the step of constraining the candidate output according to the adjustable constraint strength parameter and obtaining the response result of the candidate output includes: The constraint strength parameters are mapped to a correction threshold, a rejection threshold, and a tolerance threshold. The risk score of the candidate output is determined based on the tolerance threshold to constrain the candidate output; The risk score, the correction threshold, and the rejection threshold are compared to determine the range of the risk score, and the response result is obtained.

[0010] Furthermore, mapping the constraint strength parameter to a correction threshold, a rejection threshold, and a tolerance threshold includes: Obtain the preset constraint strength parameters; The product of the preset correction lower limit and the constraint strength parameter is determined as the first intermediate value, and the difference between the preset correction upper limit and the first intermediate value is limited to a preset range to obtain the correction threshold. The product of the preset rejection lower limit and the constraint strength parameter is determined as the second intermediate value, and the difference between the preset rejection upper limit and the second intermediate value is limited to a preset range to obtain the rejection threshold. The product of the preset tolerance lower limit and the constraint strength parameter is determined as the third intermediate quantity, and the difference between the preset tolerance upper limit and the third intermediate quantity is limited to a preset range to obtain the tolerance threshold.

[0011] Further, determining the risk score of the candidate output based on the tolerance threshold includes: Calculate the deviation of the candidate output from the preset constraints; If the deviation value is less than or equal to the tolerance threshold, the risk score is obtained based on the basic risk value and the small risk increment; wherein, the basic risk value is adjusted by the constraint strength parameter; Otherwise, the risk score is obtained based on the basic risk value and the significant risk increment.

[0012] Further, the comparison of the risk score, the correction threshold, and the rejection threshold to perform an interval judgment on the risk score and obtain the response result includes: When the risk score is less than the correction threshold, a release operation is performed on the candidate output to determine the response result; When the risk score is greater than or equal to the correction threshold and less than the rejection threshold, a correction operation is performed on the candidate output to determine the response result. When the risk score is greater than or equal to the rejection threshold, a rejection operation is performed on the candidate output to determine the response result.

[0013] Furthermore, the step of performing a comprehensive score based on the response result, and determining the optimal output based on the comprehensive score result to obtain the execution result of the optimal output includes: Obtain the individual scores of the response result on each preset indicator, and sum all the individual scores by weight to obtain the action score of the response result; The optimal output is determined from all the candidate outputs corresponding to the user input, based on the action score. Execute the optimal output to obtain the execution result.

[0014] Furthermore, the set of agents selected through the routing mechanism to participate in the interaction includes: When the user input is received, based on the task characteristics of the user input, the K agents with the highest current scores are selected to respond to the user input; Among the selected agents, user agents are initialized for preferences and authorization, platform coordinators are initialized for routing and governance, seller agents are initialized for pricing and policy constraints, and expert agents are initialized for interpretation and comparison, to obtain the agent set.

[0015] Furthermore, after obtaining the response result of the candidate output, the method further includes: Record the constraint evaluation results, response results, and audit logs for each candidate output to adjust the constraint strength parameters.

[0016] Secondly, a multi-agent interaction device based on organizational control and constraint regulation is provided, the multi-agent interaction device based on organizational control and constraint regulation includes: a processor and a memory for storing processor-executable instructions; The processor is configured to execute the multi-agent interaction method based on organizational control and constraint regulation.

[0017] Thirdly, a non-volatile computer storage medium is provided, the computer storage medium storing computer-executable instructions, which are executed by one or more processors to perform the multi-agent interaction method based on organizational control and constraint regulation described in the first aspect.

[0018] Fourthly, a computer program product containing instructions is provided that, when executed on a computer or processor, causes the computer or processor to perform the multi-agent interaction method based on organizational control and constraint regulation as described in the first aspect.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention separates "how to generate candidate outputs based on user input" from "whether the candidate outputs can be used as the final output," thereby enabling the specification of constraints on candidate outputs according to user needs. This avoids high-risk operations performed without constraints and prevents invalid or illegal results due to inconsistent response information. By using constraint strength parameters, the complex static rules of existing technologies are transformed into a parameterized continuous control surface of constraint strength parameters. This makes the constraints continuous and adjustable, enabling the generation of effective responses within complex constraint ranges. Furthermore, by adjusting the constraint strength parameters, flexible dynamic interaction can be achieved, preventing unstable system behavior. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating a multi-agent interaction method based on organizational control and constraint regulation provided in an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating a specific example of a prior art system architecture provided by an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating a specific example of a system architecture provided in an embodiment of the present invention; Figure 4 This is a flowchart illustrating step 20 provided in an embodiment of the present invention; Figure 5 This is a flowchart illustrating step 201 provided in an embodiment of the present invention; Figure 6 This is a flowchart illustrating step 202 provided in an embodiment of the present invention; Figure 7 This is a flowchart illustrating step 203 provided in an embodiment of the present invention; Figure 8 This is a flowchart illustrating step 30 provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of a multi-agent interaction device based on organizational control and constraint regulation provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0023] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted as openly inclusive, meaning "including, but not limited to." In the description of the specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," "specific example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples; that is, although they may be incorporated into embodiments or examples using the above terms for reasons such as order and position, it does not limit them to be incorporated in combination by a single embodiment or example.

[0024] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more. Furthermore, for example, the description may use the prefix "A" or "B" to describe the same type of nouns as two independent entities. In this case, the corresponding features defined with "A" and "B" are used only to distinguish between similar entities and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features.

[0025] In describing some embodiments, the terms "coupled," "coupled," and "connected," and their derivative expressions, may be used. For example, the term "connected" may be used in describing some embodiments to indicate that two or more components have direct physical or electrical contact with each other. Similarly, the term "coupled" may be used in describing some embodiments to indicate that two or more components have direct physical or electrical contact. However, the terms "connected" or "coupled" may also refer to two or more components that do not have direct contact with each other but still cooperate or interact with each other, such as "optical coupling," "wireless connection," etc. The embodiments disclosed herein are not necessarily limited to the scope of this invention.

[0026] Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0027] Existing research focuses on the collaboration dimension in multi-agent systems, i.e., how agents can work together more efficiently. These solutions do not address the orthogonal governance dimension: what happens after collaboration generates candidate operations, including permission checks, constraint enforcement, auditing, and problem reporting mechanisms.

[0028] In recent years, technologies related to multi-agent systems have shown a trend of shifting from passively responsive to proactive large language model agents, enabling agents to anticipate user needs and take proactive actions. This proactivity significantly exacerbates governance challenges: when agents autonomously invoke tools, make commitments, or initiate transactions, structured oversight mechanisms must be established to prevent unauthorized or harmful behavior.

[0029] Existing research on multi-agent large language model systems has made significant progress in collaboration. Breakthroughs have been continuously achieved in coordination topology, communication protocols, role assignment strategies, and automatic routing techniques. However, these studies have explored how agents can achieve more efficient coordination, but have not addressed how to handle candidate actions generated during coordination, nor have they addressed the legality of these actions, whether they need modification or prevention, or the method of decision recording. The governance problem of achieving such control is fundamentally different from the design of multi-agent collaborative systems and cannot be solved simply by improving the agents themselves. Even using agents with more fluent negotiation skills in existing technologies, there is still a lack of mechanisms to judge the rationality of the solutions proposed by agents.

[0030] Existing multi-agent systems, when applied to negotiation, procurement, and conversational business scenarios, mainly suffer from the following problems: Most systems employ broadcasting or simple scheduling, resulting in a lack of unified constraints among agents. For example, in negotiation scenarios, multiple buyer agents simultaneously send price requests to the same seller agent. Without a central coordination mechanism, the seller agent can only respond individually, leading to price confusion and conflicting commitments (e.g., accepting the final price from two buyers simultaneously). In procurement scenarios, the purchasing agent broadcasts its needs to all supplier agents, each submitting an independent quote. Without a unified compliance check layer, a supplier's quote, even if below cost (i.e., potentially illegal), might be directly accepted by the system, leading to subsequent non-fulfillment. Furthermore, in conversational commerce scenarios, a user inputs "recommend a mobile phone under 5000 yuan," and multiple sales agents respond simultaneously, some recommending brand A, others brand B. These responses may contradict each other (e.g., one agent says brand B is in stock, another says it's out of stock), providing inconsistent information and causing significant confusion for the user.

[0031] Existing constraint mechanisms are too rigid, often only able to employ strict filtering strategies for agent actions. Actions that violate constraints are directly rejected, leading to system stagnation in complex constraint scenarios where the system cannot generate effective output. For example, in a negotiation scenario, if the system's preset constraint is "quotes must strictly not exceed the budget," the buyer agent's budget is 5000 yuan, and the seller's quotation is 5001 yuan, any quotation exceeding 5000 yuan is prohibited, preventing the agent from generating any effective counter-offers, resulting in a deadlock and an empty system output. Similarly, in a procurement scenario, if the system sets "supplier quotations cannot be lower than 10% of cost price," and a supplier is willing to sell at cost price (i.e., below the 10% threshold) to clear inventory, the system directly rejects the quotation without any room for negotiation. The buyer misses the optimal deal, ultimately resulting in no effective contract generated, and wasting time and effort. For example, in a conversational business scenario, the system stipulates that "any promise must be confirmed by a human before it can be sent." The agent recognizes that the user wants to buy a certain product, but the budget is just at the limit. Because it involves price commitments, it cannot make any statement other than "not recommending to buy". The agent can only reply "I cannot answer", the conversation ends, and it cannot provide any effective information feedback to the user.

[0032] In multi-agent competition or collaboration, the lack of dynamic adjustment mechanisms can easily lead to unstable behavior. For example, in a negotiation scenario, if the negotiation rounds are fixed at 10, and both parties are close to reaching an agreement in the third round, but the system cannot terminate or adjust the pace in advance, the subsequent seven rounds will repeatedly confirm the same price, ultimately resulting in contradictory conclusions due to timeouts or redundant interactions (e.g., one party suddenly backs out). In a procurement scenario, if market prices suddenly fluctuate during the procurement process, fixed constraints (such as "quotes are valid for 24 hours") cannot adapt to the changes, causing previously accepted quotes to become invalid before they take effect, leaving the procurement agent and supplier agent in a state of constant cancellation and resubmission of quotes. Furthermore, in a conversational business scenario, if a user continuously inquires about price details, the agent will re-query the database each round according to a fixed strategy, leading to a gradual increase in response delays. Simultaneously, due to the lack of an escalation mechanism, when a user complains, the system cannot transfer the complaint to a human agent and can only repeatedly apologize, causing the user's emotions to escalate and their behavior to spiral out of control.

[0033] Furthermore, there is no room for exploration, meaning that the agent can only produce safe but suboptimal actions, resulting in low welfare and efficiency. Moreover, it cannot distinguish the severity of violations, treating minor, temporary violations (e.g., approaching the budget) and serious violations (e.g., exceeding the budget by 30%) equally.

[0034] The existing AgenticPay benchmark evaluated the performance of large language models in economically constrained interaction scenarios involving private value, multi-round negotiations, and outcome indicators such as feasibility, efficiency, and welfare. Research shows that even the most powerful models exhibit significant shortcomings in negotiation quality and execution reliability. In the following embodiments, AgenticPay is also used as the primary benchmark platform for evaluating the multi-agent interaction method based on organizational control and constraint regulation, with its negotiation environment considered as an extreme test of governance mechanisms under constrained conditions. The AgentPay benchmark specifically verified this when commitments exceeded the scope of role permissions or transactions required manual confirmation: even skilled negotiators exhibited low feasibility rates and poor welfare outcomes in constrained business environments. After relevant experimental analysis, the applicant believes that the bottleneck is not insufficient reasoning ability, but rather the lack of a structured mediating layer connecting agent intentions and system execution.

[0035] In multi-agent systems, especially in interactive scenarios with economic consequences such as negotiation, procurement, and e-commerce shopping guides, the lack of a dynamically adjustable governance middle layer leads to: agents potentially performing actions beyond their authority or violating constraints (e.g., exceeding budget commitments or engaging in illegal transactions); the system's inability to strike a balance between exploration (i.e., flexibly generating actions) and constraint satisfaction (i.e., compliant execution); inefficiency (e.g., numerous negotiation rounds and high latency); and welfare loss (different benefits for buyers and sellers).

[0036] Those skilled in the art often believe that in multi-agent systems, governance decisions and capabilities (i.e., agents) cannot be processed separately. However, after relevant experimental analysis, the applicant argues that while the single-agent baseline model employs the same large-scale language model and context budget as the organizational control layer system, it exhibits significantly lower welfare levels and efficiency. This performance gap does not stem from superior reasoning ability or a larger model, but rather from the structural flaws resulting from the separation of action proposals and execution processes within the governance structure. Governance decisions and capabilities should be studied and addressed as independent topics.

[0037] After conducting relevant experimental analysis, the applicant believes that the correct solution is not to make individual intelligent agents more cautious, but to introduce a dedicated intermediate layer, referred to as the organization control layer in this embodiment of the invention. This layer is responsible for handling governance issues orthogonal to action generation, and executes action strategies independently of the intelligent agents it supervises.

[0038] This invention provides an architecture based on an organizational control layer, which is located between the collaborative agents that generate candidate actions and the runtime environment that executes these candidate actions.

[0039] The method of this invention constructs a system architecture by clearly defining role divisions, permission relationships, and coordination protocols, and implements this by formalizing the governance mechanism as an explicit middle layer with adjustable execution strength. Unlike traditional organizational models that assume fixed role divisions, the organizational control layer separates role allocation from the execution mechanism and treats constraint strength as a continuous control variable.

[0040] This invention classifies operations by risk level, sets access conditions or rewrites rules based on role permissions, reports issues for manual review when necessary, and records the decision-making process for auditing purposes; all of the aforementioned operations do not require modification of the underlying proxy strategy. The execution granularity is controlled by a single scalar τ∈[0, 1], which provides methodological support for studying the trade-off between constraint satisfaction and exploration capabilities in system research. Empirical studies show that this relationship exhibits non-monotonicity: both insufficient and excessive constraints lead to performance degradation, while the organizational control layer defines a robust operating range between the two.

[0041] Specifically, this embodiment proposes a multi-agent interaction method based on organizational control and constraint regulation. In one embodiment, such as... Figure 1 As shown, it includes: Step 10: Select a set of intelligent agents to participate in the interaction through a routing mechanism; each sub-agent in the set of intelligent agents generates candidate outputs based on user input.

[0042] The routing mechanism is determined by those skilled in the art based on the specific use case and is not limited here.

[0043] The user input is a natural language request; a specific example of user input is: "Help me buy a mobile phone with a budget of 5,000 yuan."

[0044] Among them, the candidate output is the candidate action generated by the agent, which can be a candidate response or a policy output.

[0045] Multiple agents in a set of agents generate multiple candidate actions (i.e., candidate outputs) in parallel. In one embodiment, user input, each agent's role context, and current environmental state (e.g., dialogue turn, historical quotes) are input to the multiple agents, and each selected agent independently generates a candidate output (e.g., a language proposal or tool call); the candidate output may include: price proposal, commitment statement, tool call intent, etc.

[0046] The method described in this invention is applicable to scenarios such as e-commerce shopping guides, multi-agent collaborative decision-making, and intelligent service platforms.

[0047] like Figure 2As shown, existing multi-agent LLM systems typically lack a dedicated governance layer; agents directly generate and execute actions, or rely solely on simple coordination. For example... Figure 3 As shown, this embodiment of the invention proposes an organizational control layer as an intermediate layer, located between the collaborative intelligent agents (upper layer) and the execution environment (lower layer). The upper layer is the collaborative intelligent agent layer, where multiple agents fulfill various roles, including user agents, platform coordinators, seller agents, and expert agents. The intermediate layer is the organizational control layer, which decomposes roles: checks whether actions are within the agent's permissions; calculates risk, applies constraint strength parameters and thresholds to decide whether to allow, correct, or reject actions, achieving risk gating; in one embodiment, it can also record all decisions and constraint checks, and handle high-risk or momentary violations, triggering replanning / human intervention; and outputs controlled executable actions (i.e., the optimal output in this embodiment of the invention). The lower layer is the execution environment layer, which in one embodiment may include an order system, payment interface, dialogue engine, and database. This organizational control layer operates independently of the collaboration of the multiple intelligent agents it manages and independently of the underlying action generation strategy, and can be uniformly applied to intelligent agent backends based on prompting, retrieval enhancement, and learning.

[0048] In one embodiment, for an agent whose role is a user agent, the candidate output could be "I offer 4800 yuan to buy this phone"; for an agent whose role is a platform coordinator, the candidate output could be "Transaction locked, awaiting payment confirmation"; for an agent whose role is a seller agent, the candidate output could be "Our store has the first model of the first brand phone, priced at 4599 yuan, which fits your budget"; for an agent whose role is an expert agent, the candidate output could be "Based on your budget of 5000 yuan, we recommend the second model of the first brand phone (4699 yuan) or the third model of the second brand phone (4899 yuan), which offer high cost-performance."

[0049] Step 20: Based on the adjustable constraint strength parameter, perform constraint evaluation on the candidate output and obtain the response result of the candidate output.

[0050] The constraint strength parameters are determined by those skilled in the art based on the specific application scenario.

[0051] This invention models the constraint mechanism as an adjustable control variable, namely, the constraint strength parameter, to achieve dynamic and stable control of a multi-agent system. By parameterizing the governance strategy using the scalar constraint strength parameter, the applicant has experimentally determined that it has a non-monotonic impact on feasibility, efficiency, and welfare. This ensures that the execution of the corresponding constraints has an optimal operating mechanism, rather than a simple "the more the better" relationship.

[0052] The following provides a specific example of constraining candidate outputs based on constraint strength parameters. The specific type of response result is determined by those skilled in the art based on the specific application scenario. The response result is the system's determination of the processing of the candidate output; in one embodiment, it can be a process of allowing, correcting, or rejecting the candidate output.

[0053] Step 30: Perform a comprehensive score based on the response results, determine the optimal output based on the comprehensive score results, and obtain the execution result of the optimal output.

[0054] The optimal output is defined as the best candidate output among all candidate outputs and their responses corresponding to the user input. The optimal output is determined by a comprehensive score across all candidate outputs. The optimal output is then executed to obtain the execution result. The specific methods for comprehensive scoring, determining the optimal output, and execution are determined by those skilled in the art based on the specific use case.

[0055] In one embodiment, after step 30, the processed candidate outputs are input into the adjudication module, and the optimal output is selected as the system execution result based on the scoring function. The execution result is then fed back to the user or the environment. If the task is not completed (e.g., the negotiation has not reached an agreement), the process returns to step 10, where user requests are received and the sub-agent is routed to proceed to the next round of interaction. If an upgrade is triggered or the task is terminated, the final result is recorded and resources are released. In an optional embodiment, the selected action is sent to the execution environment (e.g., an order system), the environment status (budget, round, transaction price) is updated, and it is determined whether the task is completed. If not, the process returns to step 10 to continue interaction and conduct multiple rounds of negotiation.

[0056] This invention separates "how to generate candidate outputs based on user input" from "whether the candidate outputs can be used as the final output," thereby enabling the specification of constraints on candidate outputs according to user needs. This avoids high-risk operations performed without constraints and prevents invalid or illegal results due to inconsistent response information. By using constraint strength parameters, the complex static rules of existing technologies are transformed into a parameterized continuous control surface of constraint strength parameters. This makes the constraints continuous and adjustable, enabling the generation of effective responses within complex constraint ranges. Furthermore, by adjusting the constraint strength parameters, flexible dynamic interaction can be achieved, preventing unstable system behavior.

[0057] This invention provides an effective management method based on an organizational control layer for proactive large language model agents. The corresponding governance mechanism achieves effective management without constraining the proactive capabilities of the underlying agents. By separating governance from capabilities, "how the agent generates actions" and "whether the actions are executable" are decomposed into two orthogonal modules, preventing agents from performing high-risk operations without constraints. This allows agents to focus on "what to do," i.e., providing capabilities. An organizational control layer is introduced as an intermediary between the agent and the environment, responsible for role decomposition and risk gating, focusing on "whether it is allowed," i.e., governing the agent. By introducing the organizational control layer and constraint strength parameters, static rules are transformed into a parameterized continuous control surface of constraint strength parameters, making constraints continuous and adjustable. This avoids the system stagnation problem caused by traditional strong constraints, while improving system stability and execution efficiency. Thus, without modifying the underlying agent policies, dynamic gating, correction, upgrading, and auditing are implemented, achieving the effects of separating governance from capabilities and non-monotonic controllable performance, realizing unified constraints on multiple agents.

[0058] The following is a further description of the multi-agent interaction method based on organizational control and constraint adjustment according to embodiments of the present invention: A specific example of an applicable scenario is evaluating the organizational control layer on the AgentPay benchmark, which provides a multi-round negotiation task that includes private buyer budgets, seller reserve prices, and economic outcome metrics. Each round involves buyer agents and seller agents (or, a seller-side multi-agent system) negotiating product prices. The task terminates when the agents reach an agreement, the maximum number of rounds is completed, or an upgrade exit mechanism is triggered.

[0059] First, initialize the system configuration by inputting: user scenario type (e.g., e-commerce negotiation, customer service) and security constraint set (budget cap, price floor, compliance rules).

[0060] Then, the request intent and key constraints (e.g., budget, product type) are parsed. In one embodiment, in step 10, the selection of the set of agents participating in the interaction through a routing mechanism includes: Upon receiving the user input, based on the task characteristics of the user input, the K agents with the highest current scores are selected to respond to the user input. All agents below this standard are excluded. The specific method for determining the current score (K) is determined by those skilled in the art based on existing technology and the specific application scenario, and is not limited here.

[0061] Based on reputation history or role matching, select the K candidate agents with the highest current ratings from the agent pool. Output the set of agents participating in this round of interaction.

[0062] This invention, based on role decomposition (user agent, platform coordinator, seller agent, expert agent) and reputation / historical performance, filters the set of agents participating in the interaction, ensuring that each agent only generates actions within its authorized scope. In one embodiment, among the filtered agents, the user agent is initialized for preferences and authorization, the platform coordinator is initialized for routing and governance, the seller agent is initialized for pricing and policy constraints, and the expert agent is initialized for interpretation and comparison, thus obtaining the set of agents.

[0063] This invention proposes an organizational control layer as a formalized intermediate layer between the platform-level collaborative intelligent agent and the downstream execution environment. The organizational control layer parameterizes a series of governance strategies using a constraint strength parameter τ∈[0, 1] of scalar control strength. This strength can simultaneously adjust risk classification, action gating, and escalation behavior. By uniformly adjusting risk classification, gating thresholds, and escalation sensitivity through a single scalar parameter τ∈[0,1], continuous control from "loose" to "strict" is achieved.

[0064] Specifically, embodiments of the present invention model the economic multi-agent task as a tuple: ; in, Indicates the status of the environment (including market, orders, inventory, and conversation status). Represents intelligent agents The action space (including language proposals and instrumental intentions). This represents the observation space (including private constraints such as budget, reservation price, or platform rules). It represents the set of hard and soft constraints (including permissions, compliance rules, and risk strategies). Indicates utility specifications (including feasibility, benefits, and cost-adjusted benefits).

[0065] The aforementioned abstract model aligns closely with the negotiation benchmarks used in practical applications of existing technologies, and also covers downstream operational processes such as order placement, payment confirmation, and issue escalation. While theoretically universal, the specific implementation mechanisms vary depending on the application scenario and are not limited here.

[0066] In existing instantiation schemes, state variables and constraints mainly revolve around budget allocation, seller minimum bid threshold, transaction rounds, and pricing strategies. However, this invention defines the organizational control layer as a set of control strategies: ; and the corresponding set of mappings: .

[0067] in, Represents the trajectory of messages and observations. This represents the original sequence of actions of the intelligent agent. It is a sequence of controlled, executable actions, and It includes audit signals, constraint check records, and attribution tracking information.

[0068] The core architectural principle of this invention is that the organizational control layer dominates the mapping process from language behavior to economic behavior without altering the underlying agent's strategy. Each component undertakes different governance functions, as detailed below: Role decomposition (i.e., This is used to assign responsibilities to specialized agents: user agents represent preferences and authorization, the platform coordinator is responsible for routing and governance, seller agents represent pricing and policy constraints, and expert agents are used for interpretation and comparison. Role policies ensure that each agent operates within its designated authority.

[0069] Risk management (i.e., This is used to classify candidate operations into risk levels and implement differentiated controls. High-risk operations (financial commitments, order submissions, policy-sensitive commitments) require explicit verification, while low-risk operations are automatically approved. The gating system generates one of four decisions for each operation: approve, rewrite, modify, or reject.

[0070] Audit logs (i.e., Audit trails are used for: recording quote trajectories, constraint checks, control decisions, and execution results. This facilitates post-event debugging, repeatability verification, and compliance reviews. Importantly, audit trails also enable systems to detect and recover from transient violations—that is, to capture and correct constraint failures before they reach the execution boundary.

[0071] Upgrading and replanning (i.e., This escalation mechanism is used when negotiations reach an impasse, constraints conflict, or risks exceed tolerable limits. It can trigger task refactoring, expert intervention, user confirmation, or a safe exit. This component is critical; removing the escalation mechanism will lead to catastrophic system failure.

[0072] This invention introduces a scalar control parameter: the constraint strength parameter τ∈[0,1], which is used to characterize a series of organizational control strategies. The constraint strength parameter τ here is not a single threshold, but rather applies to the overall governance strength of candidate actions. Specifically, τ collectively controls: , , and

[0073] in, This indicates the set of operations considered high-risk, which increases with the constraint strength parameter. The larger the constraint strength parameter, the more cautious the system, and more types of actions will be classified as high-risk (for example, originally only "payment" was considered high-risk, but after the constraint strength parameter is increased, "confirm order" may also be classified as high-risk).

[0074] It is the threshold for rewriting or confirmation, which decreases as the constraint strength parameter increases; a lower threshold means that even actions with low risk scores may trigger rewriting or manual confirmation, i.e., "easier to intervene".

[0075] It is the threshold for blocking or escalating, which decreases as the constraint strength parameter increases; similarly, it decreases, making it possible for actions with low risk scores to be blocked or escalated, i.e., more likely to be rejected.

[0076] This indicates the tolerance for near-miss constraint violations, which tightens as the constraint strength parameter increases. Tolerance refers to a near-miss but not complete violation of the constraint (e.g., a budget of 1 yuan remaining, with the quoted price exceeding the limit by just 0.5 yuan). Increasing the constraint strength parameter lowers the tolerance, meaning even very small deviations will be considered violations.

[0077] In one embodiment, such as Figure 4 As shown, step 20 includes: Step 201: Map the constraint strength parameters to a correction threshold, a rejection threshold, and a tolerance threshold.

[0078] In this embodiment of the invention, the constraint strength parameter is not directly used to determine candidate actions. Instead, it serves as a meta-parameter, mapped to three specific executable thresholds (i.e., a correction threshold, a rejection threshold, and a tolerance threshold). The risk score of the candidate output is then range-based using the gating logic in step 20. This maintains the simplicity of parameterization while providing sufficient expressive power. By transforming the constraint mechanism from static rules into a dynamically adjustable governance layer with the constraint strength parameter as the single control variable, a balance between efficiency and constraint feasibility can be achieved.

[0079] Step 202: Determine the risk score of the candidate output based on the tolerance threshold to constrain the candidate output.

[0080] A specific example of risk assessment and gating in an intermediate layer is as follows: Given the candidate output set, the current environmental state, and preset hard constraints (e.g., budget, floor price), calculate a risk score for each candidate output using the modified threshold, rejection threshold, and tolerance threshold related thresholds calculated in the steps below (e.g., based on action type, role permissions, historical violation records, etc.).

[0081] Step 203: Compare the risk score, the correction threshold, and the rejection threshold to perform an interval judgment on the risk score and obtain the response result.

[0082] The response result is not the result of the execution of the candidate output, but a processing judgment result after risk assessment.

[0083] Specifically, in one embodiment, such as Figure 5 As shown, step 201 includes: Step 2011: Obtain the preset constraint strength parameters.

[0084] The preset constraint strength parameters are determined by those skilled in the art based on the specific application scenario.

[0085] In this embodiment of the invention, the relationship between τ and system welfare and efficiency is not linear, and it is not a monotonic control curve. There exists an optimal operating range (τ≈0.25–0.5), and both excessive leniency and excessive strictness will reduce performance. In an optional embodiment, if the system performs poorly in actual operation (e.g., welfare decreases or violation rate increases), the performance of the constraint strength parameter within the [0,1] interval can be scanned on the validation set. The constraint strength parameter value that maximizes the average cost-adjusted welfare (CAW) can be selected. The constraint strength parameter can be dynamically adjusted according to scenario or user type. The larger the constraint strength parameter, the stricter the governance: the high-risk action set expands, the intervention threshold decreases, tolerance tightens, and the system's sensitivity to violations increases across the board.

[0086] In one embodiment, the applicant demonstrated through experiments that constraint enforcement is not simply a matter of the more the better. When the constraint strength parameter is 1.0, the system becomes overly conservative: actions that might lead to efficient protocol achievement may be marked, rewritten, or unnecessarily upgraded, thereby increasing latency and reducing welfare. When the constraint strength parameter is 0.0, the governance mechanism is insufficient to allow more exploratory actions, but misses early correction opportunities. The optimal state occurs within the constraint strength parameter τ∈[0.25,0.50] interval, achieving peak welfare while maintaining full feasibility and zero enforcement violations by balancing exploration and constraint satisfaction. The constraint strength parameter is non-monotonic; by treating the constraint strength parameter as an adjustable parameter and using validation data, the operating point for maximizing the objective can be determined. The solution of this embodiment provides a way to handle this trade-off. In summary, the relationship between enforcement strength and system performance is non-monotonic: neither insufficient nor excessive constraints result in the optimal solution, while the method of this embodiment can achieve the best balance in the form of an intermediate operating interval.

[0087] Step 2012: Determine the product of the preset correction lower limit and the constraint strength parameter as the first intermediate value, and limit the difference between the preset correction upper limit and the first intermediate value within a preset range to obtain the correction threshold.

[0088] The preset upper limit and the preset lower limit are determined by those skilled in the art based on the specific application scenario. In one embodiment, the preset upper limit can be 0.92 and the preset lower limit can be 0.27.

[0089] Step 2013: Determine the product of the preset rejection lower limit and the constraint strength parameter as the second intermediate value, and limit the difference between the preset rejection upper limit and the second intermediate value within a preset range to obtain the rejection threshold.

[0090] The preset rejection upper limit and preset rejection lower limit are determined by those skilled in the art based on the specific application scenario. In one embodiment, the preset rejection upper limit can be 1.02 and the preset rejection lower limit can be 0.32.

[0091] Step 2014: Determine the product of the preset tolerance lower limit and the constraint strength parameter as the third intermediate quantity, and limit the difference between the preset tolerance upper limit and the third intermediate quantity within a preset range to obtain the tolerance threshold.

[0092] The preset tolerance upper limit and preset tolerance lower limit are determined by those skilled in the art based on the specific use case. In one embodiment, the preset tolerance upper limit can be 0.22 and the preset tolerance lower limit can be 0.14.

[0093] In one specific implementation, the mapping relationship between a constraint strength parameter and a specific control parameter is as follows: ; The clip(·) function restricts the value to the range [0, 1]. A constraint strength parameter approaching 0 corresponds to a relaxed control state, while a constraint strength parameter approaching 1 corresponds to strict control execution. As the constraint strength parameter increases, all control parameters tighten uniformly.

[0094] Specifically, in one embodiment, such as Figure 6 As shown, step 202 includes: Step 2021: Calculate the deviation of the candidate output from the preset constraints.

[0095] Among them, the preset constraints, namely the preset hard constraints (such as budget, base price), and the specific methods for calculating the deviation value by referring to the corresponding constraints are determined by those skilled in the art according to the specific use case, and are not limited here.

[0096] Step 2022: If the deviation value is less than or equal to the tolerance threshold, the risk score is obtained based on the basic risk value and the small risk increment; wherein the basic risk value is adjusted by the constraint strength parameter.

[0097] The basic risk value and the minor risk increment are determined by those skilled in the art based on the specific use case, and are not limited here.

[0098] In this embodiment of the invention, the total risk of an action (characterized by a risk score) is equal to its basic risk (i.e., deviation value) plus the additional risk arising from violation or proximity to the constraint (i.e., the small risk increment in step 2022 or the large risk increment in step 2023 below).

[0099] In one embodiment, the basic risk can be determined by the action type, for example, a basic risk of 0.05+0.25τ, where τ is a constraint strength parameter.

[0100] Step 2023: Otherwise, obtain the risk score based on the basic risk value and the significant risk increment.

[0101] The definition of a significant risk increment is determined by those skilled in the art based on the specific use case and is not limited here. A significant risk increment is much greater than a minor risk increment.

[0102] Specifically, in one embodiment, such as Figure 7 As shown, step 203 includes: Step 2031: When the risk score is less than the correction threshold, perform a release operation on the candidate output to determine the response result.

[0103] Step 2032: When the risk score is greater than or equal to the correction threshold and less than the rejection threshold, perform a correction operation on the candidate output to determine the response result.

[0104] Step 2033: When the risk score is greater than or equal to the rejection threshold, perform a rejection operation on the candidate output to determine the response result.

[0105] In one specific implementation, a risk score rt∈[0,1] is estimated for each proposed action (i.e., candidate output). ; in, This indicates that the candidate output should be allowed. This indicates that a correction operation will be performed on the candidate output. This indicates that a rejection operation will be performed on the candidate output.

[0106] In one specific instance, if the risk score is less than the correction threshold: proceed to the adjudication module. If the risk score is greater than or equal to the correction threshold but less than the rejection threshold: correct (e.g., limit the price to within the budget, rewrite sensitive statements), and then proceed to the adjudication module. If the risk score is greater than or equal to the rejection threshold: escalate (transfer to manual intervention, request expert assistance, or trigger a secure exit process) or reject. Finally, a set of controlled, executable actions is output (i.e., each candidate output and its response).

[0107] Existing technologies lack a recovery mechanism; once a violation occurs, the system fails immediately and lacks the ability to learn from momentary deviations or retry. In one embodiment, this invention, when the risk score is greater than or equal to the rejection threshold, does not directly execute the action or simply reject it. Instead, it escalates the process: transferring the violation to manual review, triggering expert AI intervention, re-decomposing the task (e.g., allowing a budget checker to adjust it first), and / or safely exiting and recording the reason. This avoids situations where repeated apologies are necessary, leading to escalating user emotions, loss of control, and user complaints. Following this approach, momentary violations are intercepted and recovered from, preventing them from becoming execution violations; the system will not crash due to a single candidate action violation, significantly improving stability.

[0108] The applicant determined through experiments that a significant feature of the method used in the embodiments of this invention is its ability to distinguish between transient violations and execution violations. In the experiments, when a budget overrun check was triggered at least once per round, the organizational control layer exhibited a 100% transient violation rate, but a 0% execution violation rate (no violations entered the environment). This distinction is significant in practical operation; a system capable of capturing and correcting all violations is safer than a system that never encounters violations, because the former embodies a proactive governance mechanism, while the latter may only reflect the distribution of occasional benign scenarios.

[0109] After step 203, the environment is then determined according to... Evolution occurs; due to Depending on the constraint strength parameter, constraint enforcement directly shapes the system dynamics, rather than acting as a passive safety filter.

[0110] In one embodiment, such as Figure 8 As shown, step 30 includes: Step 301: Obtain the individual scores of the response result on each preset indicator, and sum all the individual scores by weight to obtain the action score of the response result.

[0111] In one embodiment, the adjudication module scores and selects the optimal action, inputs all released or corrected candidate outputs, and uses a weighted scoring function to score each candidate action to obtain an action score.

[0112] In one embodiment, the action score for calculating the response result according to preset indicators includes: determining the action score as the sum of the products of each preset indicator and its corresponding weight.

[0113] For example, when the preset indicators include matching degree indicators, cost indicators, compliance indicators, and agent reputation values, the action score for calculating the response result according to the preset indicators includes: ; Among them, the matching degree weight Cost metrics are weighted in measuring the fit between the agent's capabilities and task requirements. Used to measure the economic cost of calling the smart agent (e.g., token consumption, fees, etc.), compliance weight. Reputation score weight is used to measure whether an intelligent agent meets compliance requirements such as security, privacy, and regulations. Used to measure the historical performance, user evaluation, success rate, etc. of intelligent agents.

[0114] Step 302: Determine the optimal output from all the candidate outputs corresponding to the user input, which is the candidate output with the action score.

[0115] The action with the highest score is selected as the final output of the system.

[0116] Step 303: Execute the optimal output to obtain the execution result.

[0117] Actions (i.e., candidate outputs) are sent to downstream execution environments (such as payment interfaces, order systems, and dialogue engines).

[0118] In one embodiment, after step 303, the environment state is updated (e.g., budget is deducted, transaction price is recorded, round count is increased).

[0119] In one embodiment, after step 30, the method further includes: recording the constraint evaluation result, response result, and audit log of each candidate output to adjust the constraint strength parameter.

[0120] Record the following for each action: raw output, gating decision, corrections, risk score, and timestamp. Used for subsequent debugging, compliance review, or training improvement.

[0121] This embodiment of the invention implements a constraint exploration loop through an organizational control layer, which processes each action proposal: (1) Propose a suggestion: The agent generates raw language operations or tool intents. (2) Check: The control layer evaluates each operation based on role permissions, budget constraints, and risk strategies. (3) Gating: Operations are approved, rewritten, upgraded, or rejected based on parameterized thresholds of constraint strength parameters. (4) Execute: Only managed operations are allowed to enter the environment. (5) Record: Result data is recorded for auditing, attribution analysis, and future control. This loop separates constraint execution from action generation, enabling the same organizational control layer to manage different types of agent backends (e.g., prompt-based, retrieval-enhanced, or learning-based) without modification.

[0122] This invention also provides a multi-agent interaction platform system, including an agent management module, a dynamic routing module, a constraint control module, an adjudication and execution module, and an interaction log and feedback module, for executing the methods of this invention.

[0123] This invention avoids system crashes, improves interaction efficiency (reducing the number of rounds), enhances system stability, and achieves a balance between exploration and constraints, making it applicable to various business scenarios (e.g., e-commerce, customer service, decision-making). The applicant evaluated the method of this invention's embodiments in the AgentPay benchmark test. Compared to the unconstrained benchmark model, the structured governance mechanism of this invention's embodiments can continuously improve feasibility and welfare levels. The organizational control layer algorithm achieves a 6.0-fold improvement in welfare indicators after cost adjustment compared to the single-agent benchmark model, while reducing negotiation rounds by 44% and latency by 47%. Regardless of whether the underlying action generation mechanism employs immediate response, retrieval enhancement, or learning-based methods, the method of this invention's embodiments is universally applicable.

[0124] The foregoing embodiments provided a multi-agent interaction method based on organizational control and constraint regulation. In this embodiment, another multi-agent interaction device based on organizational control and constraint regulation will be proposed. The multi-agent interaction device based on organizational control and constraint regulation includes: a processor and a memory for storing processor-executable instructions; wherein, the processor is configured to execute the multi-agent interaction method based on organizational control and constraint regulation described in the foregoing embodiments.

[0125] like Figure 9 As shown, the multi-agent interaction device based on organizational control and constraint regulation includes a processor 21 and a memory 22, wherein the processor 21 and the memory 22 can be connected by a bus or other means.

[0126] Processor 21 can be a CPU. Processor 21 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0127] The memory 22, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the multi-agent interaction method based on organizational control and constraint regulation in the aforementioned embodiments. The processor executes various functional applications and training processes by running the non-transitory software programs, instructions, and modules stored in the memory.

[0128] The memory 22 may include a program storage area and a training storage area. The program storage area may store the operating system and applications required for at least one function; the training storage area may store training data created by the processor. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 22 may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. The one or more modules stored in the memory 22, when executed by the processor 21, perform the multi-agent interaction method based on organizational control and constraint regulation as shown in the embodiments of the present invention. Specific details of the multi-agent interaction method based on organizational control and constraint regulation can be understood by referring to the corresponding descriptions and effects in the embodiments of the present invention, and will not be repeated here.

[0129] This embodiment also provides a computer storage medium storing a computer program that can be executed by a processor to perform the multi-agent interaction method based on organizational control and constraint regulation described in the foregoing embodiments.

[0130] The computer storage medium stores computer-executable instructions, which can execute the multi-agent interaction method based on organizational control and constraint regulation in any of the above method embodiments. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.

[0131] The specific steps of the multi-agent interaction method based on organizational control and constraint regulation are described in the foregoing embodiments and will not be repeated in this embodiment.

[0132] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-agent interaction method based on organizational control and constraint regulation, characterized in that, include: The set of intelligent agents participating in the interaction is selected through a routing mechanism; Each sub-agent in the set of intelligent agents generates candidate outputs based on user input; The candidate output is constrained and evaluated based on the adjustable constraint strength parameter, and the response result of the candidate output is obtained. A comprehensive score is performed based on the response results, and the optimal output is determined according to the comprehensive score results to obtain the execution result of the optimal output.

2. The multi-agent interaction method based on organizational control and constraint regulation according to claim 1, characterized in that, The method includes: The constraint strength parameters are mapped to a correction threshold, a rejection threshold, and a tolerance threshold. The risk score of the candidate output is determined based on the tolerance threshold to constrain the candidate output; The risk score, the correction threshold, and the rejection threshold are compared to determine the range of the risk score, and the response result is obtained.

3. The multi-agent interaction method based on organizational control and constraint regulation according to claim 2, characterized in that, The method includes: Obtain the preset constraint strength parameters; The product of the preset correction lower limit and the constraint strength parameter is determined as the first intermediate value, and the difference between the preset correction upper limit and the first intermediate value is limited to a preset range to obtain the correction threshold. The product of the preset rejection lower limit and the constraint strength parameter is determined as the second intermediate value, and the difference between the preset rejection upper limit and the second intermediate value is limited to a preset range to obtain the rejection threshold. The product of the preset tolerance lower limit and the constraint strength parameter is determined as the third intermediate quantity, and the difference between the preset tolerance upper limit and the third intermediate quantity is limited to a preset range to obtain the tolerance threshold.

4. The multi-agent interaction method based on organizational control and constraint regulation according to claim 2, characterized in that, The method includes: Calculate the deviation of the candidate output from the preset constraints; If the deviation value is less than or equal to the tolerance threshold, the risk score is obtained based on the basic risk value and the small risk increment; wherein, the basic risk value is adjusted by the constraint strength parameter; Otherwise, the risk score is obtained based on the basic risk value and the significant risk increment.

5. The multi-agent interaction method based on organizational control and constraint regulation according to claim 2, characterized in that, The method includes: When the risk score is less than the correction threshold, a release operation is performed on the candidate output to determine the response result; When the risk score is greater than or equal to the correction threshold and less than the rejection threshold, a correction operation is performed on the candidate output to determine the response result. When the risk score is greater than or equal to the rejection threshold, a rejection operation is performed on the candidate output to determine the response result.

6. The multi-agent interaction method based on organizational control and constraint regulation according to claim 1, characterized in that, The method includes: Obtain the individual scores of the response result on each preset indicator, and sum all the individual scores by weight to obtain the action score of the response result; The optimal output is determined from all the candidate outputs corresponding to the user input, based on the action score. Execute the optimal output to obtain the execution result.

7. The multi-agent interaction method based on organizational control and constraint regulation according to any one of claims 1-6, characterized in that, The method includes: When the user input is received, based on the task characteristics of the user input, the K agents with the highest current scores are selected to respond to the user input; Among the selected agents, user agents are initialized for preferences and authorization, platform coordinators are initialized for routing and governance, seller agents are initialized for pricing and policy constraints, and expert agents are initialized for interpretation and comparison, to obtain the agent set.

8. The multi-agent interaction method based on organizational control and constraint regulation according to any one of claims 1-6, characterized in that, The method further includes: Record the constraint evaluation results, response results, and audit logs for each candidate output to adjust the constraint strength parameters.

9. A multi-agent interaction device based on organizational control and constraint regulation, characterized in that, The multi-agent interaction device based on organizational control and constraint regulation includes: a processor and a memory for storing processor-executable instructions; The processor is configured to execute the multi-agent interaction method based on organizational control and constraint regulation as described in any one of claims 1 to 8.

10. A non-volatile computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, which are executed by one or more processors to perform the multi-agent interaction method based on organizational control and constraint regulation as described in any one of claims 1 to 8.