Adaptive multi-modal interaction and multi-agent collaboration investment decision method and system

By using adaptive multimodal interaction and multi-agent collaborative learning, the interaction mode and task allocation are dynamically adjusted, solving the problems of single agent interaction and static decision-making in existing intelligent agents, and realizing efficient and personalized user interaction and investment decision-making.

CN121052683BActive Publication Date: 2026-03-27SHANGHAI GREAT WISDOM SHENJIU INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing intelligent agents have limited and unadaptable interaction modes, insufficient task allocation and agent collaboration capabilities, and a lack of dynamism and personalization in decision generation, resulting in cumbersome user operations, delayed task execution, and low resource utilization.

Method used

An adaptive multimodal interaction method is adopted, which dynamically switches between voice, text and image input modes, and combines user preferences, environmental status and sentiment analysis to identify user intentions and perform task decomposition and priority calculation. Personalized investment suggestions are generated by using multi-agent collaborative learning and reinforcement learning algorithms.

Benefits of technology

It achieves optimal interaction without manual user intervention, improves the accuracy and smoothness of interaction, avoids resource waste and task blockage, generates personalized investment suggestions, and improves system resource utilization and real-time adaptability of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052683B_ABST
    Figure CN121052683B_ABST
Patent Text Reader

Abstract

The application provides an adaptive multi-modal interaction and multi-agent cooperation investment decision method and system, comprising the following steps: S1: obtaining a multi-modal user request, preprocessing the obtained multi-modal user request to obtain a preprocessed user request; S2: identifying a user query intention based on the preprocessed user request; S3: task decomposition based on the identified user query intention, and dynamic allocation of an optimal agent based on the decomposed subtasks; S4: the optimal agent generates an optimal investment suggestion based on the allocated task content, according to a user portrait, historical behavior and real-time market data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and intelligent agent technology, in particular, to an investment decision-making method and system based on adaptive multi-modal interaction and multi-agent collaboration. BACKGROUND

[0002] Intelligent agent technology has been widely applied in various automated tasks, data analysis and human-computer interaction. The limitations of existing intelligent agents mainly lie in the following aspects:

[0003] Single interaction mode and lack of adaptability: existing intelligent agents mostly support only text or voice single input mode, and cannot dynamically adjust the interaction mode according to the user's environment, user state and task demand, resulting in tedious user operation and low demand transmission accuracy;

[0004] Insufficient task allocation and intelligent agent collaboration ability: in the face of multi-dimensional tasks in the financial field, existing solutions mostly rely on static rules to allocate tasks, without considering the real-time load, capacity adaptation and system global resource state of intelligent agents, which may lead to problems such as "high priority tasks allocated to high load intelligent agents" and "single intelligent agent unable to undertake complex tasks", resulting in task execution delay and low system resource utilization;

[0005] Lack of dynamic and personalized decision-making: existing intelligent agent decision-making mostly relies on static models trained by historical data or pre-set strategies, which is difficult to adapt to the high volatility of the financial market in real time, and does not deeply integrate user portrait and real-time behavior, resulting in serious homogenization of generated investment recommendations, which cannot meet the individual needs of different users.

[0006] Patent document CN109389259A (application number: 201710653642.5) discloses a knowledge property right interaction system and method, wherein the system includes a sending module for receiving knowledge property right errands sent by a sending person; an errand module for displaying key information of the knowledge property right errands received by the sending module to assist the knowledge property right service personnel to select the desired knowledge property right errands and call the order grabbing module; and an order grabbing module for displaying detailed information of the knowledge property right errands selected by the knowledge property right service personnel to assist the knowledge property right service personnel to complete the order grabbing. This patent focuses on information retrieval and lacks multi-modal input and dynamic decision support, and does not involve data security design. The present application improves the interaction experience and security of intelligent agents by introducing adaptive multi-modal interaction, real-time dynamic decision-making and differential privacy technology.

[0007] Patent document CN101923669A (application number: 200910161341.6) discloses an intelligent adaptive design, which combines hardware and software to produce various physical and non-physical SweetSpotsTM (sweet spots) to adapt to the needs of users. Sweet spots and resources can be linked to thousands of electronic categories, including: artificial intelligence, communication, educational purposes, environmental protection, guides, interaction, interface, Internet, positioning, predictive, identifier, remote control, security tracking, recreation, motion sensor, infrared and vibration sensor, sound recognizer and synthesizer, toys for all age groups, visual identifier, wireless messaging, etc. The patent focuses on interface dynamic adjustment, and the present application not only focuses on dynamic adjustment of the interactive interface, but also includes deep mining of user portraits and decision optimization of real-time market data, which is more complex and deep in technical level.

[0008] Patent document US20200265301A1 (application number: US16277735) discloses an e-commerce recommendation system based on a collaborative filtering algorithm for user recommendation. The present application combines reinforcement learning and multi-agent collaborative learning, and is not limited to e-commerce recommendation, but also supports market analysis, investment suggestions and other more extensive applications. SUMMARY

[0009] In view of the defects in the prior art, the purpose of the present application is to provide an adaptive multi-modal interaction and multi-agent collaborative investment decision-making method and system.

[0010] According to the adaptive multi-modal interaction and multi-agent collaborative investment decision-making method provided by the present application, the following steps are included:

[0011] Step S1: Obtain a multi-modal user request, and pre-process the obtained multi-modal user request to obtain a pre-processed user request;

[0012] Step S2: Identify the user query intent based on the pre-processed user request;

[0013] Step S3: Task decomposition based on the identified user query intent, and dynamic allocation of optimal agents based on the decomposed sub-tasks;

[0014] Step S4: The optimal agent generates the optimal investment suggestion based on the allocated task content, according to the user portrait, historical behavior and real-time market data.

[0015] Preferably, the step S1 includes:

[0016] Step S1.1: Dynamically selecting an input mode according to user preferences, environmental changes, task requirements and user emotional states; wherein the input mode includes: voice input mode, text input mode and image input mode;

[0017] Step S1.2: When the input mode is a voice input mode or an image input mode, the acquired voice data or image data is converted into text data.

[0018] Preferably, the step S1.1 comprises:

[0019] acquiring user preferences based on historical interaction habits, determining an input mode based on the user preferences, and setting the current input mode as a default interaction mode;

[0020] acquiring environmental parameters through device sensors, when the noise decibel is greater than a threshold value and the current default interaction mode is a voice input mode, automatically switching the interaction mode to a text input mode or an image input mode;

[0021] when the noise decibel is less than or equal to the threshold value, maintaining the current default interaction mode;

[0022] obtaining a target task based on the current input user request, and analyzing the current target task, when the current target task requires structured information meeting preset requirements, setting the current interaction mode to a text input mode or an image input mode;

[0023] performing sentiment analysis based on the current input user request or the collected user facial expressions, judging the emotional state of the current user, when the user emotion is in a depressed state, setting the current interaction mode to a voice input mode.

[0024] Preferably, the step S2 comprises: performing deep semantic analysis on the acquired text data through a BERT model or a GPT model, and comprehensively judging in combination with historical user behaviors to identify user query intent.

[0025] Preferably, the step S3 comprises:

[0026] Step S3.1: performing task decomposition based on the identified user query intent to obtain decomposed subtasks;

[0027] Step S3.2: calculating subtask priorities based on user portraits, historical behavior data, and real-time market data, respectively;

[0028]

[0029] wherein, is the final priority score of the i-th subtask; is the matching degree of the subtask and the user portrait; is the influence coefficient of real-time market data on the subtask; is the historical interaction frequency of the subtask corresponding to the user; is the weight coefficient of each index;

[0030] Step S3.3: generating a task feature vector containing task category, priority, resource requirement, real-time requirement based on the decomposed subtasks;

[0031] Step S3.4: matching the task feature vector with the capability vector of the agent, obtaining the agent whose matching degree is greater than the preset value, and constructing an agent candidate set based on the agent whose matching degree is greater than the preset value;

[0032] Step S3.5: the agent candidate set selects the optimal agent to execute the current task through a multi-agent collaborative learning strategy MARL.

[0033] Preferably, the step S3.5 includes:

[0034] Step S3.5.1: for each candidate agent, constructing a subtask-agent fusion state vector ;

[0035]

[0036] wherein, is the subtask number; is the candidate agent number; is the subtask category; is the subtask priority; is the subtask resource requirement; is the subtask real-time requirement; is the agent's favorite task category; is the agent's resource processing upper limit; is the agent's same type task completion rate; is the agent's current load rate, is the agent's historical response delay;

[0037] Step S3.5.2: preprocessing the subtask-agent fusion state vector including normalization processing, removing outliers, to obtain a matrix wherein, is the number of candidate agents;

[0038] Step S3.5.3: the actor network of each agent in the candidate set obtains the willingness value of undertaking the current subtask based on the corresponding fusion state vector ;

[0039] Step S3.5.4: the critic network evaluates the global benefit of each agent-subtask matching combination based on the matrix , combined with the willingness value of all candidate agents ;​​

[0040] Step S3.5.5: based on the will value and global benefits Calculate the comprehensive score, sort in descending order of the comprehensive score, and select the agent with the highest score as the optimal execution agent;

[0041] .

[0042] Preferably, the step S4 comprises:

[0043] Step S4.1: constructing an agent based on a Deep Q-Learning reinforcement learning algorithm;

[0044] Step S4.2: the agent dynamically obtains an investment strategy through Bayesian optimization dynamic tuning based on real-time market data, user portrait, historical behavior and user demand.

[0045] According to the adaptive multi-modal interaction and multi-agent collaborative investment decision system provided by the application, comprising:

[0046] Module M1: obtaining a multi-modal user request, preprocessing the obtained multi-modal user request to obtain a preprocessed user request;

[0047] Module M2: identifying a user query intent based on the preprocessed user request;

[0048] Module M3: task decomposition based on the identified user query intent, and dynamic allocation of optimal agents based on the decomposed subtasks;

[0049] Module M4: the optimal agent generates the optimal investment suggestion according to the user portrait, historical behavior and real-time market data based on the allocated task content.

[0050] Preferably, the module M1 comprises:

[0051] Module M1.1: dynamically selecting an input mode according to user preferences, environmental changes, task requirements and user emotional states; wherein the input mode includes: voice input mode, text input mode and image input mode;

[0052] Module M1.2: when the input mode is voice input mode or image input mode, the obtained voice data or image data is converted into text data;

[0053] The module M1.1 comprises:

[0054] Based on the historical interaction habits, the user preferences are obtained, the input mode is determined based on the user preferences, and the current input mode is set as the default interaction mode;

[0055] Obtain environmental parameters through device sensors, when the noise decibel is greater than a threshold value, and the current default interaction mode is a voice input mode, then automatically switch the interaction mode to a text input mode or an image input mode;

[0056] When the noise decibel is less than or equal to the threshold value, the current default interaction mode is maintained;

[0057] Obtain a target task based on the current input user request, and analyze the current target task, when the current target task needs to meet a structured information meeting a preset requirement, set the current interaction mode to a text input mode or an image input mode;

[0058] Perform emotion analysis based on the current input user request or collected user facial expressions, judge the emotional state of the current user, when the user emotion is in a depressed state, set the current interaction mode to a voice input mode.

[0059] Preferably, the module M3 comprises:

[0060] Module M3.1: obtain decomposed subtasks based on the identified user query intent for task decomposition;

[0061] Module M3.2: calculate subtask priorities based on user portraits, historical behavior data, and real-time market data, respectively;

[0062]

[0063] wherein, is the final priority score of the ith subtask; is the matching degree of the subtask and the user portrait; is the influence coefficient of real-time market data on the subtask; is the historical interaction frequency of the subtask corresponding to the user; is the weight coefficient of each index;

[0064] Module M3.3: generate a task feature vector containing task categories, priorities, resource requirements, and real-time requirements based on the decomposed subtasks;

[0065] Module M3.4: match the task feature vector with the ability vector of the agent, obtain an agent with a matching degree greater than a preset value, and construct an agent candidate set based on the agent with a matching degree greater than the preset value;

[0066] Module M3.5: the agent candidate set selects the optimal agent to execute the current task through a multi-agent collaborative learning strategy MARL;

[0067] The module M3.5 comprises:

[0068] Module M3.5.1: for each candidate agent, construct a subtask-agent fusion state vector ;

[0069]

[0070] wherein, is the subtask number; is the candidate agent number; is the subtask category; is the subtask priority; is the subtask resource requirement; is the subtask real-time requirement; is the agent's preferred task category; is the agent's resource processing upper limit; is the agent's same-type task completion rate; is the agent's current load rate, is the agent's historical response delay;

[0071] Module M3.5.2: the subtask-agent fusion state vector is preprocessed including normalization processing and removing outliers, to obtain matrix wherein, is the number of candidate agents;

[0072] Module M3.5.3: the actor network of each agent in the candidate set is based on its corresponding fusion state vector obtains the willingness value of undertaking the current subtask ;

[0073] Module M3.5.4: the critic network is based on matrix , combined with the willingness value of all candidate agents , to evaluate the global benefit of each agent-subtask matching combination ;

[0074] Module M3.5.5: based on the willingness value and the global benefit , calculate the comprehensive score, sort in descending order according to the comprehensive score, and select the agent with the highest score as the optimal execution agent;

[0075] .

[0076] Compared with the prior art, the present application has the following beneficial effects:

[0077] 1、The application dynamically combines user preferences, environmental status, task requirements and emotion analysis to select input modes, solves the problem of "single interaction mode and passive input reception" in the prior art, realizes an optimal interaction mode without manual intervention of the user, and improves the accuracy of requirement transmission and the smoothness of operation;

[0078] 2、The application solves the problem of "static rule allocation of tasks and ignoring the state of agents and the global load of the system" in the prior art through task decomposition, priority calculation, score matching and MARL strategy optimization;

[0079] 3、The application avoids "resource waste" and "task blocking" by matching scores to filter adaptive agents and evaluating global benefits through a network of critics;

[0080] 4、The application solves the problem of "homogenization of investment suggestions and ignoring personalized user needs" in the prior art by constructing agents based on Deep Q-Learning, combining real-time market data, user portraits, historical behavior and user needs, and dynamically optimizing with Bayesian optimization. BRIEF DESCRIPTION OF DRAWINGS

[0081] Other features, objects and advantages of the application will become more apparent through a detailed description of the non-limiting embodiments with reference to the following drawings:

[0082] Figure 1 Flow chart of the adaptive multi-modal interaction and multi-agent collaborative investment decision-making method. DETAILED DESCRIPTION

[0083] The application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the application, but do not limit the application in any form. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the application. These are within the scope of the application.

[0084] Example 1

[0085] According to the adaptive multi-modal interaction and multi-agent collaborative investment decision-making method and system provided by the application, through multi-modal input processing, user portrait construction, real-time intent recognition, task priority calculation, multi-agent task allocation, adaptive reinforcement learning and Bayesian optimization combined decision-making generation mechanism, efficient, personalized, privacy-safe human-computer interaction and decision support in a complex dynamic environment are realized.

[0086] The adaptive multi-modal interaction and multi-agent collaborative investment decision-making method comprises:

[0087] Step S1: obtaining a multi-modal user request, preprocessing the obtained multi-modal user request to obtain a preprocessed user request;

[0088] Specifically, the step S1 comprises:

[0089] Step S1.1: dynamically selecting an input mode according to user preferences, environmental changes, task requirements and user emotional states; wherein the input mode comprises: a voice input mode, a text input mode and an image input mode;

[0090] More specifically, the step S1.1 comprises:

[0091] Based on historical interaction habits, the user preferences are obtained, the input mode is determined based on the user preferences, and the current input mode is set as the default interaction mode;

[0092] The environmental parameters are obtained through the device sensor, when the noise decibel is greater than the threshold value, and the current default interaction mode is the voice input mode, the interaction mode is automatically switched to the text input mode or the image input mode;

[0093] When the noise decibel is less than or equal to the threshold value, the current default interaction mode is maintained;

[0094] Based on the current input user request, the target task is obtained, and the current target task is analyzed, when the current target task needs to meet the structured information of the preset requirement, the current interaction mode is set to the text input mode or the image input mode;

[0095] Based on the current input user request or the collected user facial expression, the emotional analysis is performed, the emotional state of the current user is judged, when the user emotion is in a depressed state, the current interaction mode is set to the voice input mode.

[0096] Step S1.2: when the input mode is the voice input mode or the image input mode, the obtained voice data or image data is converted into text data.

[0097] Compared with the single interface scheme passively accepting any input type, in complex or unfavorable conditions, the application can dynamically select the most suitable interaction mode according to the current interaction state and behavior of the user, improve the interaction efficiency and accuracy through adaptive multi-modal interaction, reduce the user burden, and enhance the user experience;

[0098] Step S2: identifying a user query intent based on the preprocessed user request;

[0099] Specifically, the step S2 comprises: performing deep semantic analysis on the obtained text data through a BERT model or a GPT model, and then comprehensively judging the user historical behavior to identify the user query intent.

[0100] Step S3: task decomposition based on the identified user query intent, and dynamic allocation of optimal agents based on the decomposed subtasks;

[0101] Specifically, the step S3 comprises:

[0102] Step S3.1: task decomposition based on the identified user query intent to obtain decomposed subtasks;

[0103] Step S3.2: calculation of subtask priority based on user portrait, historical behavior data, and real-time market data respectively;

[0104]

[0105] wherein, is the final priority score of the ith subtask; is the matching degree of the subtask and the user portrait; is the influence coefficient of real-time market data on the subtask; is the historical interaction frequency of the subtask corresponding to the user; is the weight coefficient of each index;

[0106] More specifically, the user portrait is constructed by combining multiple data sources to comprehensively portrait the user; specifically including: behavior data, sentiment analysis, market relationship data; wherein the behavior data includes purchase history, investment preference; the sentiment analysis includes: identifying the emotional state of the user based on historical input voice and text, and forming an emotional trend; the market relationship data analyzes the correlation between the user's historical investment portfolio and the market fluctuation in the same period, calculates indicators such as fluctuation sensitivity and correlation coefficient, to reflect the user's risk characteristics. The constructed user portrait is dynamically updated, for example: when the user shows interest in a certain field, such as stock A, within a certain period of time, the agent will automatically enhance the weight of that field, and adjust the user portrait in real time, in order to provide more accurate recommendations and suggestions for the user.

[0107] Step S3.3: generation of a task feature vector containing task category, priority, resource demand, and real-time requirement based on the decomposed subtasks;

[0108] Step S3.4: matching of the task feature vector with the capability vector of the agent, obtaining an agent with a matching degree greater than a preset value, and constructing an agent candidate set based on the agent with a matching degree greater than the preset value;

[0109] Step S3.5: selection of the optimal agent for the current task by the multi-agent collaborative learning strategy MARL of the agent candidate set.

[0110] More specifically, the step S3.5 comprises:

[0111] Step S3.5.1: For each candidate agent, construct a subtask-agent fusion state vector ;

[0112]

[0113] wherein, is the subtask number; is the candidate agent number; is the subtask category; is the subtask priority; is the subtask resource requirement; is the subtask real-time requirement; is the agent's proficient task category; is the agent's resource processing upper limit; is the agent's same-type task completion rate; is the agent's current load rate, is the agent's historical response delay;

[0114] Step S3.5.2: The subtask-agent fusion state vector is preprocessed including normalization processing, removing outliers, to obtain a matrix wherein, is the number of candidate agents;

[0115] Step S3.5.3: The actor network of each agent in the candidate set is based on its own corresponding fusion state vector obtains the willingness value of undertaking the current subtask ;

[0116] Step S3.5.4: The critic network is based on the matrix , combined with the willingness value of all candidate agents , to evaluate the global benefit of each agent-subtask matching combination ;

[0117] Step S3.5.5: Based on the willingness value and the global benefit , calculate the comprehensive score, sort in descending order according to the comprehensive score, and select the agent with the highest score as the optimal execution agent;

[0118] .

[0119] Step S4: The optimal agent generates the optimal investment suggestion based on the assigned task content, according to the user portrait, historical behavior and real-time market data;

[0120] Specifically, the step S4 includes:

[0121] Step S4.1: Constructing an agent based on Deep Q-Learning reinforcement learning algorithm;

[0122] Step S4.2: The agent dynamically obtains investment strategies through Bayesian optimization dynamic tuning based on real-time market data, user portraits, historical behavior, and user demand.

[0123] In this embodiment, the agent combines Deep Q-Learning reinforcement learning algorithm to learn how to adjust decisions based on real-time market data, user behavior, historical decision results, and other factors. This algorithm enables the agent to make optimal decisions in a constantly changing environment. The reinforcement learning model can continuously adjust its strategy through continuous training and feedback to achieve the best returns.

[0124] Reward mechanism: The model optimizes the decision path through rewards, such as profit growth, and penalties, such as loss mechanisms. Whenever the agent makes a wrong decision, the system will be punished according to the loss, prompting the agent to make better decisions in the future.

[0125] Combined with the Bayesian optimization algorithm, the agent can handle hyperparameter tuning under the reinforcement learning framework to improve the performance of the model. Bayesian optimization can help the agent automatically optimize the decision-making process by continuously exploring and adjusting model parameters in the face of uncertainty and high-dimensional decision space.

[0126] Based on real-time data feedback and changes in user demand, such as investment style and risk tolerance, the agent adjusts the decision-making model through Bayesian optimization to improve the accuracy and execution effect of the task solution.

[0127] Combined with adaptive reinforcement learning and Bayesian optimization, the agent can dynamically adjust investment strategies according to market fluctuations and changes in user demand to ensure that it adapts to the latest market conditions. This method enables decisions to be based not only on historical data but also on real-time market dynamics, improving the agent's ability to respond to changes.

[0128] The method also includes a data security mechanism, which includes:

[0129] A distributed database architecture is adopted to achieve efficient data storage and retrieval.

[0130] Data security is ensured through multi-layer encryption, differential privacy technology, and federated learning.

[0131] The differential privacy technology prevents user data leakage, so even if the agent's model parameters are obtained by attackers, they cannot infer any user's private data.

[0132] The multi-layer encrypted user data is encrypted by using AES-256 in the transmission and storage process, so that the security of the data is ensured.

[0133] The federated learning is used for training the model in multiple distributed nodes, direct sharing of user data is avoided, and the privacy protection capability is improved.

[0134] The application also provides an adaptive multi-modal interaction and multi-agent collaborative investment decision system, which can be realized by executing the process steps of the adaptive multi-modal interaction and multi-agent collaborative investment decision method, that is, the adaptive multi-modal interaction and multi-agent collaborative investment decision method can be understood as the preferred implementation of the adaptive multi-modal interaction and multi-agent collaborative investment decision system by those skilled in the art.

[0135] Those skilled in the art know that, in addition to implementing the system, device and each module thereof provided by the application in the form of pure computer readable program code, the same program can also be realized in the form of logic gate, switch, application specific integrated circuit, programmable logic controller and embedded microcontroller by logically programming the method steps. Therefore, the system, device and each module thereof provided by the application can be considered as a hardware component, and the modules included therein for realizing various programs can also be considered as structures in the hardware component; the modules for realizing various functions can also be considered as both software programs for realizing the method and structures in the hardware component.

[0136] The specific embodiments of the application are described above. It should be understood that the application is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essential content of the application. In the case of no conflict, the embodiments of the application and the features in the embodiments can be combined with each other at will.

Claims

1. An adaptive multi-modal interaction and multi-agent collaboration investment decision-making method, characterized in that, The application relates to a method for generating an optimal investment suggestion based on a multimodal user request. The method comprises the following steps: Step S1: obtaining a multimodal user request, preprocessing the obtained multimodal user request, and obtaining a preprocessed user request; Step S2: identifying a user query intention based on the preprocessed user request; Step S3: task decomposition based on the identified user query intention, and dynamic allocation of an optimal intelligent agent based on the decomposed subtasks; Step S4: the optimal intelligent agent generates an optimal investment suggestion based on the allocated task content, a user portrait, historical behavior and real-time market data. The step S3 comprises: Step S3.1: task decomposition based on the identified user query intention to obtain decomposed subtasks; wherein, is the final priority score of the i-th subtask; is the matching degree of the subtask and the user portrait; is the influence coefficient of real-time market data on the subtask; is the historical interaction frequency of the user corresponding to the subtask; is the weight coefficient of each index; Step S3.2: calculation of subtask priorities based on a user portrait, historical behavior data and real-time market data; Step S3.3: generation of a task feature vector comprising a task category, a priority, a resource requirement and a real-time requirement based on the decomposed subtasks; Step S3.4: matching of the task feature vector with an ability vector of an intelligent agent, obtaining an intelligent agent with a matching degree greater than a preset value, and constructing an intelligent agent candidate set based on the intelligent agent with the matching degree greater than the preset value; 2.The adaptive multi-modal interaction and multi-agent collaboration based investment decision-making method of claim 1, wherein, Step S3.5: selection of an optimal intelligent agent for executing a current task by the intelligent agent candidate set through a multi-agent collaborative learning strategy MARL. The step S1 comprises: Step S1.1: dynamic selection of an input mode according to a user preference, an environmental change, a task requirement and a user emotional state; wherein the input mode comprises a voice input mode, a text input mode and an image input mode; Step S1.2: when the input mode is the voice input mode or the image input mode, the obtained voice data or image data is converted into text data; The step S1.1 comprises: acquisition of a user preference based on historical interaction habits, determination of an input mode based on the user preference, and setting of a current input mode as a default interaction mode; acquisition of environmental parameters through a device sensor, automatic switching of an interaction mode to a text input mode or an image input mode when a noise decibel is greater than a threshold value and a current default interaction mode is the voice input mode; maintaining the current default interaction mode when the noise decibel is less than or equal to the threshold value; acquisition of a target task based on a currently input user request, analysis of the current target task, and setting of a current interaction mode to a text input mode or an image input mode when the current target task needs to satisfy structured information of a preset requirement; 3.The adaptive multi-modal interaction and multi-agent collaboration based investment decision-making method of claim 1, wherein, emotional analysis based on a currently input user request or a collected user facial expression, judgment of a current user emotional state, and setting of a current interaction mode to a voice input mode when the user emotional state is in a low state. 4.The adaptive multi-modal interaction and multi-agent collaboration based investment decision-making method of claim 1, wherein, The step S2 comprises: deep semantic analysis of obtained text data through a BERT model or a GPT model, and comprehensive judgment in combination with historical behavior of a user to identify a user query intention. Step S3.5.1: For each candidate agent, construct a subtask-agent fused state vector ; wherein, is a subtask sequence number; is a candidate agent sequence number; is a subtask category; is a subtask priority; is a subtask resource requirement; is a subtask real-time requirement; is an agent's proficient task category; is an agent's resource processing upper limit; is an agent's same-type task completion rate; is an agent's current load rate, is an agent's historical response delay; Step S3.5.2: Subtask-agent fusion state vector Pretreatment including normalization processing, removing outliers, to obtain matrix wherein, is the number of candidate agents; Step S3.5.3: the actor network of each agent in the candidate set is updated based on the corresponding fusion state vector of itself obtain the willingness value of the current subtask ; Step S3.5.4: The critic network evaluates each agent-subtask match combination based on the matrix , combining the will value of all candidate agents , evaluating the global benefit of each agent-subtask match combination ; Step S3.5.5: calculating the comprehensive score based on the will value and global benefits calculating a comprehensive score, ranking the comprehensive scores in descending order, and selecting an agent with the highest score as an optimal execution agent; 。 5.The adaptive multi-modal interaction and multi-agent collaboration based investment decision-making method of claim 1, wherein, The step S3.5 comprises: The step S4 comprises: Step S4.1: construction of an intelligent agent based on a Deep Q-Learning reinforcement learning algorithm; Step S4.2: The intelligent agent dynamically obtains the investment strategy through Bayesian optimization dynamic tuning based on real-time market data, user portrait, historical behavior and user demand.

6. An adaptive multi-modal interaction and multi-agent collaboration investment decision system, characterized in that, Comprise: Module M1: obtaining a multimodal user request, preprocessing the obtained multimodal user request to obtain a preprocessed user request; Module M2: identifying a user query intention based on the preprocessed user request; Module M3: task decomposition based on the identified user query intention, and dynamic allocation of an optimal intelligent agent based on the decomposed subtasks; Module M4: the optimal intelligent agent generates the optimal investment suggestion based on the allocated task content, user portrait, historical behavior and real-time market data; The module M3 comprises: Module M3.1: task decomposition based on the identified user query intention to obtain decomposed subtasks; Module M3.2: calculating the subtask priority based on the user portrait, historical behavior data and real-time market data; wherein, is the final priority score of the i-th subtask; is the matching degree of the subtask and the user portrait; is the influence coefficient of real-time market data on the subtask; is the historical interaction frequency of the user corresponding to the subtask; is the weight coefficient of each index; Module M3.3: generating a task feature vector containing task category, priority, resource demand and real-time requirement based on the decomposed subtasks; Module M3.4: matching the task feature vector with the capability vector of the intelligent agent to obtain intelligent agents with a matching degree greater than a preset value, and constructing an intelligent agent candidate set based on the intelligent agents with a matching degree greater than the preset value; Module M3.5: the intelligent agent candidate set selects the optimal intelligent agent to execute the current task through the multi-agent collaborative learning strategy MARL.

7. The adaptive multi-modal interaction and multi-agent collaboration based investment decision system, as claimed in claim 6, wherein, The module M1 comprises: Module M1.1: dynamically selecting an input mode according to user preferences, environmental changes, task requirements and user emotional states; wherein the input mode includes voice input mode, text input mode and image input mode; Module M1.2: when the input mode is voice input mode or image input mode, the obtained voice data or image data is converted into text data; The module M1.1 comprises: Obtaining user preferences based on historical interaction habits, determining the input mode based on the user preferences, and setting the current input mode as the default interaction mode; Obtaining environmental parameters through device sensors, when the noise decibel is greater than the threshold value, and the current default interaction mode is voice input mode, the interaction mode is automatically switched to text input mode or image input mode; When the noise decibel is less than or equal to the threshold value, the current default interaction mode is maintained; Obtaining a target task based on the current input user request, and analyzing the current target task, when the current target task needs to meet the structured information of the preset requirement, the current interaction mode is set to text input mode or image input mode; Performing emotional analysis based on the current input user request or the collected user facial expressions to determine the emotional state of the current user, when the user emotion is in a depressed state, the current interaction mode is set to voice input mode.

8. The adaptive multi-modal interaction and multi-agent collaboration based investment decision system, as claimed in claim 6, wherein, The module M3.5 comprises: Module M3.5.1 : For each candidate agent, construct a subtask-agent fusion state vector ; wherein, is a subtask sequence number; is a candidate agent sequence number; is a subtask category; is a subtask priority; is a subtask resource requirement; is a subtask real-time requirement; is an agent's proficient task category; is an agent's resource processing upper limit; is an agent's same-type task completion rate; is an agent's current load rate, is an agent's historical response delay; Module M3.5.2: Subtask-agent fusion state vector Pretreatment including normalization processing, removing outliers, obtaining matrix wherein, is the number of candidate agents; Module M3.5.3: the actor network of each agent in the candidate set is based on its own corresponding fusion state vector obtain a willingness value of undertaking the current subtask ; Module M3.5.4: Critic network based on matrix , combining the will values of all candidate agents , evaluating the global benefit of each agent-subtask matching combination ; Module M3.5.5: based on the willingness value and global benefits Calculate the comprehensive score, sort the comprehensive score in descending order, and select the agent with the highest score as the optimal execution agent; 。

Citation Information

Patent Citations

  • Intelligent adaptive design

    CN101923669A

  • An intellectual property interactive system and method

    CN109389259A

  • Incremental training of machine learning tools

    US20200265301A1

  • Intelligent financial consultation system and method based on large language model

    CN119444260A

  • Multi-agent interactive efficient data analysis system

    CN119988421A