Multi-agent-driven multi-mode cognitive method and device, electronic equipment and medium

Through the multi-modal cognitive method driven by multi-agents, users' intentions are analyzed, appropriate big models are selected, and decision-making is solved, and the data dependence, high training cost and lack of independent decision-making capabilities of traditional visual cognitive algorithms are solved, achieving more efficient and stable visual cognition and decision-making.

CN119961683AInactive Publication Date: 2025-05-09北京衔远有限公司

Patent Information

Application Number
CN202510442991.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing visual cognitive algorithms have problems such as strong data dependence, high training cost, poor generalization ability, low model stability and accuracy, and lack of independent decision-making ability.

Method used

Multi-modal cognitive method driven by multi-agents is adopted. By obtaining the instructions input by the user, using the user to analyze the agent's analysis intention, selecting to call a general or dedicated large model to perform cognitive tasks. The result analysis agent analyzes the output results, and the decision-making agent performs independent decision-making processing.

Benefits of technology

It reduces data dependence, reduces training costs, improves the generalization ability, stability and accuracy of the model, and has the ability to make independent decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961683A_ABST
    Figure CN119961683A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent-driven multi-mode cognitive method and device, electronic equipment and a medium. The method comprises the following steps: analyzing an instruction input by a user by utilizing a user analysis agent, extracting a user intention, and determining a cognitive task to be executed according to the user intention; based on an analysis result of the user analysis agent, selecting and calling at least one large model to execute the cognitive task; analyzing the output result of the large model by using a result analysis agent, if the output result does not meet the intention of the user, prompting the user to supplement information or re-describe the problem, and re-sending the updated user input to the user analysis agent, otherwise, generating an analysis result; and based on an analysis result of the result analysis agent, performing decision processing on the user problem by using a decision agent, and outputting a processing result. The data dependence can be reduced, the training cost can be reduced, the generalization ability, the stability and the accuracy of the model can be improved, and the autonomous decision-making ability can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual cognition technology, and in particular to a multi-agent driven multimodal cognition method, device, electronic device and medium. Background Art

[0002] With the rapid development of computer vision and deep learning technology, traditional visual cognition algorithms (such as target detection, image classification and image recognition) have been widely used in security monitoring, smart construction sites, industrial automation, medical imaging and autonomous driving. However, the functions of the above traditional visual cognition algorithms are often relatively single, which makes it difficult to meet the needs of complex scenes for diversified visual analysis, and there are still obvious deficiencies in the following aspects: The model relies heavily on large-scale labeled data during the training phase and is very sensitive to changes in data distribution, resulting in insufficient coverage of long-tail data and reduced model accuracy.

[0003] Models are usually closely bound to specific usage scenarios. Once the application environment changes, the generalization performance of the model will drop sharply, requiring a lot of retraining and model optimization.

[0004] Traditional deep learning algorithms mostly rely on pixel-level feature extraction and lack a deep understanding of objective physical laws and context, resulting in weak scene understanding and reasoning capabilities.

[0005] Most existing systems remain at the "cognitive" level and are unable to form a closed loop of automatic decision-making. Human intervention is still required to complete strategy selection and execution. Summary of the invention

[0006] In view of this, the embodiments of the present application provide a multi-agent driven multimodal cognitive method, device, electronic device and medium to solve the problems of strong data dependence, high training cost, poor generalization ability, low model stability and accuracy, and lack of autonomous decision-making ability in the prior art.

[0007] In a first aspect of an embodiment of the present application, a multi-agent driven multimodal cognitive method is provided, comprising: obtaining instructions input by a user; parsing the instructions input by the user using a user analysis agent, extracting the user's intention, and determining the cognitive task to be performed based on the user's intention; based on the parsing results of the user analysis agent, selecting to call at least one large model to perform the cognitive task, wherein the large model includes a general large model and a special large model; analyzing the output results of the large model using a result analysis agent, and if the output results do not meet the user's intention, prompting the user to supplement information or re-describe the problem, and sending the updated user input to the user analysis agent again, otherwise generating an analysis result; based on the analysis results of the result analysis agent, using a decision agent to make a decision on the user's problem, and outputting the processing result.

[0008] According to a second aspect of an embodiment of the present application, a multi-agent driven multimodal cognitive device is provided, comprising: an acquisition module for acquiring instructions input by a user; a parsing module for parsing the instructions input by the user using a user analysis agent, extracting the user's intention, and determining the cognitive task to be performed based on the user's intention; a calling module for selecting to call at least one large model to perform cognitive tasks based on the parsing result of the user analysis agent, wherein the large model includes a general large model and a special large model; an analysis module for analyzing the output result of the large model using a result analysis agent, and if the output result does not meet the user's intention, prompting the user to supplement information or re-describe the problem, and sending the updated user input to the user analysis agent again, otherwise generating an analysis result; a processing module for making a decision on the user's problem using a decision agent based on the analysis result of the result analysis agent, and outputting the processing result.

[0009] According to a third aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.

[0010] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0011] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: By obtaining the instructions input by the user; using the user analysis agent to parse the instructions input by the user, extract the user's intention, and determine the cognitive task to be performed according to the user's intention; based on the analysis results of the user analysis agent, select and call at least one large model to perform cognitive tasks, wherein the large model includes a general large model and a special large model; using the result analysis agent to analyze the output results of the large model, if the output results do not meet the user's intention, prompt the user to supplement the information or re-describe the problem, and send the updated user input to the user analysis agent again, otherwise generate the analysis results; based on the analysis results of the result analysis agent, use the decision agent to make decisions on the user's problem and output the processing results. This application can reduce data dependence, reduce training costs, improve the generalization ability, stability and accuracy of the model, and have the ability to make decisions independently. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0013] Figure 1 It is a schematic diagram of the overall implementation process of the multi-agent driven multimodal cognitive method provided in the embodiment of the present application in an actual scenario; Figure 2 It is a flowchart of a multi-agent driven multimodal cognitive method provided in an embodiment of the present application; Figure 3 is a schematic diagram of the structure of a multi-agent driven multimodal cognitive device provided in an embodiment of the present application; Figure 4 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0014] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0015] Visual cognition algorithms have extremely wide application demands in many fields such as industry, security, and autonomous driving. However, most of the current traditional visual cognition algorithms (including detection, classification, and recognition) are based on deep learning frameworks and have the following major problems or shortcomings: 1. Strong data dependence and sensitive data distribution This type of algorithm requires a massive amount of labeled training data to achieve more reliable performance.

[0016] During training, if the data distribution that the algorithm relies on deviates from the actual usage scenario, or if there is insufficient coverage and annotation of long-tail data, the performance of the model will decline significantly.

[0017] Therefore, in order to obtain better algorithm results, it is necessary to continuously collect and label data, which is costly and inefficient.

[0018] 2. Poor adaptability to dynamic environment Traditional large visual models are often "bound" to specific scene distributions or specific tasks, and it is difficult to maintain stable and accurate performance in different scenes or dynamically changing environments.

[0019] Once the usage environment changes (such as changes in lighting, weather, target appearance, etc.), the generalization ability of traditional algorithms often drops sharply, and the model needs to be trained or fine-tuned again to meet the needs of new scenarios.

[0020] 3. Limited semantic understanding ability Traditional deep learning-based visual algorithms can usually only process a single visual signal and perform recognition or detection through "pixel-level" feature extraction. This method cannot truly understand the objective physical laws or contextual information behind the scene.

[0021] The lack of the ability to think about the environment and context makes the model unable to cope with deeper and more abstract semantic understanding.

[0022] This limitation not only reduces the intelligence of the model, but also raises the threshold for model use and deployment, hindering the popularization and inclusiveness of AI technology.

[0023] 4. Lack of decision-making ability Traditional visual cognition algorithms can only provide detection and recognition results for scenes or targets, but are unable to make further autonomous decisions or take actions.

[0024] In actual systems, if certain operations or controls need to be performed automatically (such as autonomous navigation, automated industrial production process scheduling, etc.), human intervention or additional decision-making modules are often still required to complete them.

[0025] In summary, existing technologies focus more on the detection and recognition of specific targets in images or videos, and have not yet been deeply integrated at the semantic level and decision-making level. In practical applications, the above problems lead to the need for models to constantly rely on data, cannot flexibly adapt to various complex environments, lack high-level intelligent analysis and decision-making capabilities, resulting in high application barriers, high maintenance costs, and limited scope of application.

[0026] In view of the problems existing in the prior art, in order to solve the above-mentioned problems existing in traditional visual cognition algorithms, the present application provides a multi-agent driven general-specialized fusion visual cognition system (Agentic Multi-modal Cognition AMC). The multi-agent driven visual cognition method of the present application integrates visual cognition, multimodal cognition, autonomous decision-making, and continuous learning capabilities to build an autonomous system of "intention analysis-cognition-result analysis-decision-making"; the multimodal large model with visual and text alignment improves the ability of visual cognitive task understanding and interaction; the use of automated cleaning processes for massive data saves a lot of labeling costs, and after learning from massive unlabeled data, visual object cognition is achieved through zero-sample learning; the autonomous learning ability of the agent is used to dynamically expand the cognitive ability of the system, and the cognitive and decision-making capabilities are quickly generalized to vertical fields.

[0027] Before describing the technical solution of the present application in detail, the overall implementation process of the multi-agent driven multimodal cognitive method of the present application in actual scenarios is first explained with reference to the accompanying drawings. Figure 1 1 is a schematic diagram of the overall implementation process of the multi-agent driven multi-modal cognitive method provided in the embodiment of the present application in an actual scenario. Figure 1 As shown in the figure, this method introduces multiple agents into the system and combines the general large model with the special large model to achieve autonomous closed-loop control from user command input to decision feedback. Specifically, the process can be divided into the following steps or links: 1) User input In actual usage scenarios, users often input instructions to the system through natural language or other forms. The instructions may include the user's task requirements, target descriptions, scenario constraints, etc. Users can continuously interact with the system and complete the requirement description or provide additional background information through multiple rounds of input or inquiry.

[0028] 2) User Analysis Agent The user analysis agent receives and parses the instructions input by the user, identifies the core elements and extracts the user's intention. Based on the pre-built knowledge base or user behavior data, the user analysis agent can further match or expand the user's intention, providing a basis for subsequent cognitive task selection and model calling. If the system detects that the user input information is insufficient or ambiguous, the user analysis agent will prompt the user to provide additional information to ensure that sufficient context or data is available in subsequent cognitive tasks.

[0029] 3) General large model and special large model General large model: This model is trained based on automated data cleaning and unlabeled data learning, and has general multimodal visual analysis capabilities, including image classification, object detection, image segmentation, scene understanding, etc. It is suitable for processing most common scenes or general types of visual tasks, and can complete efficient visual recognition and analysis with less manual annotation.

[0030] Dedicated large model: When the user analysis agent detects the needs of a specific vertical field (such as medical imaging, smart construction sites, smart cities, etc.), the corresponding dedicated large model will be called. This type of model has been deeply trained for specific fields and has strong pertinence and professionalism. For example, it can be used for lesion detection in the medical field and for construction safety analysis in smart construction sites. Through the "dynamic mounting" method, the corresponding dedicated large model can be flexibly enabled in actual operation to meet differentiated field needs.

[0031] 4) Cognitive task execution Based on the analysis results output by the user analysis agent, the system will choose to call the general large model or the dedicated large model of the corresponding vertical field to perform visual cognition and analysis on the input data to obtain preliminary model output results. For scenarios with higher scene complexity or more precise requirements, the general large model and the dedicated large model can also be used in combination to obtain richer and deeper analysis.

[0032] 5) Result Analysis Agent The model output results will be handed over to the result analysis agent for comprehensive processing and judgment. If the output results cannot meet the user's expectations or more information is needed for confirmation, the result analysis agent will interact with the user, prompting the user to add additional input or further explain the problem, thereby realizing the iterative execution of cognitive tasks. If the model results can meet the user's needs, it will proceed to the next step for decision making.

[0033] 6) Decision-making Agent The qualified output results verified by the result analysis agent will be passed to the decision-making agent, which will evaluate the various decision options available based on the user's initial intention and environmental context. The decision-making process no longer relies on human intervention, but is completed automatically by the system, and finally forms an executable decision result (such as issuing control instructions to the device, outputting suggestions or conclusions to the user, etc.).

[0034] 7) Closed-loop feedback After the decision result is output, if the user confirms the result or puts forward new requirements, the system will return to the above process again, and continue to improve or adjust the decision through the collaborative work of the user analysis agent, the big model, the result analysis agent and the decision agent. As user behavior and needs continue to evolve, each agent will gradually learn user preferences and usage habits, and dynamically update the model's calling strategy and parameter settings, realizing the continuous evolution of the system and improving its adaptive capabilities.

[0035] Through the integration of the above multi-agent architecture and multimodal large models, the technical solution of this application can not only perform effective visual analysis for general scenarios, but also switch or mount different special large models according to actual needs, and flexibly respond to various vertical field challenges, thereby significantly improving visual cognition and decision-making efficiency while reducing manual annotation costs. This multi-agent driven multimodal cognitive method has good scalability and universality in practical applications, and provides effective technical support for realizing intelligent analysis and autonomous decision-making in multimodal and highly complex scenarios.

[0036] The contents of the technical solution of the present application are described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] Figure 2 Schematic diagram of the process of multi-agent driven multi-modal cognitive method provided in the embodiment of the present application. Figure 2 As shown, the multi-agent driven multimodal cognitive method may specifically include: S201, obtaining a command input by a user; S202, using a user analysis agent to parse the instructions input by the user, extract the user's intention, and determine the cognitive task to be performed according to the user's intention; S203, based on the analysis result of the user analysis agent, select and call at least one large model to perform the cognitive task, wherein the large model includes a general large model and a special large model; S204, using the result analysis agent to analyze the output results of the large model, if the output results do not meet the user's intention, prompt the user to supplement information or re-describe the problem, and send the updated user input to the user analysis agent again, otherwise generate an analysis result; S205, based on the analysis result of the result analysis agent, use the decision agent to make a decision on the user question and output the processing result.

[0038] In some embodiments, a user analysis agent is used to parse the user input command, extract the user intent, and determine the cognitive task to be performed according to the user intent, including: Perform semantic and context analysis on the instructions input by the user to identify the target information and related constraints in the instructions; Extract key elements based on semantic and context analysis results and generate or update corresponding user intent; Based on the pre-built knowledge base or user historical behavior data, user intentions are matched and expanded to form a candidate cognitive task set; Combine user preferences, system resource availability, and user intent priorities to filter and sort candidate cognitive task sets; Determine the cognitive tasks to be performed and generate corresponding task routing information. The task routing information is used to call the large model to complete the corresponding visual analysis.

[0039] Specifically, this embodiment uses the user analysis agent to parse the instructions input by the user to extract the user's intention and determine the cognitive task to be performed based on the user's intention. The user analysis agent first performs semantic and context analysis on the instructions input by the user to identify the target information and related constraints in the instructions. For example, if the user inputs "detect the type of vehicle in the image", the user analysis agent can parse out that the core goal of the instruction is vehicle identification, and identify its data type (image) and task type (classification), thereby extracting the key elements of the instruction.

[0040] After extracting the key elements, the user analysis agent combines the pre-built knowledge base or user historical behavior data to match and expand the user's intentions to form a candidate cognitive task set. For example, the system's built-in knowledge base can store the mapping relationship between different task types and large models, such as: If common vehicle types are identified, a general large model can be called to classify the vehicle; If the task involves identifying a specific brand or new energy vehicle, a dedicated large model may be required for more refined classification in a specific field; If the user has queried for commercial vehicles (such as buses and trucks) many times in the past, the system can prioritize recommending models of related categories to optimize task execution efficiency.

[0041] The user analysis agent then screens and sorts the candidate cognitive task sets based on user preferences, system resource availability, and the priority of user intent. For example: If the user wants the results to be returned quickly, the system can choose a model with lower computational cost; If users are more concerned about recognition accuracy, they can give priority to using models with high accuracy but high computational complexity; When computing resources are tight, the user analysis agent can dynamically adjust the task scheduling strategy to give priority to tasks that occupy less resources.

[0042] Finally, the user analysis agent determines the optimal cognitive task and generates task routing information, which includes the calling path of the large model, input format, storage method of output results, etc. The task routing information is used to call the appropriate large model to complete the corresponding visual analysis. For example, for the vehicle recognition task, the system can generate the following task routing information: Task type: Vehicle identification Data input: User-provided images Calling model: General large model Result output: Vehicle type classification results In addition, the user analysis agent can also guide the user to enter more information to optimize the performance of cognitive tasks. For example, when the user only enters "analyze this image", the user analysis agent can proactively ask the user for more specific requirements, such as: “Do you want to detect vehicles, pedestrians, or other objects?” “Do you need to identify the vehicle brand or model?” Through this interactive optimization mechanism, the user analysis agent can ensure the completeness of the task input and improve the applicability and execution efficiency of the model.

[0043] At the same time, the user analysis agent can learn based on the user's behavior data and continuously optimize the interaction method. For example, the system can record the user's task history and automatically recommend similar tasks that have been performed in the past to reduce the user's input burden. In addition, if the user repeatedly adjusts the task during multiple rounds of interaction (such as first selecting vehicle detection and then changing it to new energy vehicle identification), the system can automatically adjust the task recommendation logic and directly provide task options that are more in line with the user's preferences in subsequent interactions.

[0044] In summary, this embodiment ensures that the system can accurately understand user needs and intelligently select the most appropriate cognitive tasks through semantic analysis, intent parsing, task matching, screening optimization, task routing generation and user interaction optimization of the user analysis agent, thereby improving the execution efficiency of visual cognitive tasks and user experience.

[0045] In some embodiments, based on the parsing results of the user analysis agent, selecting to call at least one large model to perform a cognitive task includes: Obtain the cognitive tasks to be performed and the corresponding user intentions output by the user analysis agent; Determine whether cognitive tasks belong to a general type or a specific vertical domain based on user intent and pre-established model selection strategies; When the cognitive task is of a general type, the general large model is called to process the cognitive task; When the cognitive task belongs to a specific vertical field, dynamically load and call a dedicated large model to process the cognitive task; The output results generated by the general large model or the special large model are passed to the result analysis agent.

[0046] Specifically, in this embodiment, based on the analysis results of the user analysis agent, at least one large model is selected to perform the cognitive task. The user analysis agent first analyzes the user input instruction, extracts the cognitive task to be performed and the user's intention, and then determines whether the cognitive task should be processed by a general large model or a dedicated large model according to the preset model selection strategy.

[0047] When the cognitive task is of a general type, the system calls the general large model to perform visual analysis. For example, if the user inputs the command "identify the object in this picture", the user analysis agent determines that the task belongs to a regular target recognition task after analysis, and the system calls the general large model for processing. The general large model is pre-trained based on the automated data cleaning process and unlabeled data learning technology, and has the capabilities of visual positioning, target classification, image content understanding, etc., so it can directly process the input image and output the recognition result.

[0048] When the cognitive task involves a specific vertical field, the system will dynamically load and call a dedicated large model. For example, if the user enters the command "analyze whether this medical image has abnormalities", the user analysis agent will parse and determine that the task belongs to the field of medical image analysis. At this time, the system will dynamically load a dedicated large model for medical imaging based on the task type and model selection strategy. This dedicated large model is trained based on medical imaging data and can perform tasks such as CT / MRI image lesion detection and X-ray abnormality screening, thereby providing more accurate analysis results.

[0049] In addition, when dynamically loading a dedicated large model, the system also considers the compatibility of computing resources with the task, for example: If the user requests to analyze a large number of medical images, the system can choose high-performance computing resources in the cloud to load a dedicated model; If the user requests on-site construction monitoring, the system may choose to load a lightweight smart construction site monitoring model on the local device to ensure real-time performance.

[0050] After completing the cognitive task, the output results of the general or special large model will be passed to the result analysis agent for further analysis and optimization. For example, if the medical image analysis model detects a possible lesion area, the result analysis agent can evaluate the confidence of the result and decide whether to provide additional information, such as a detailed segmentation map or a risk assessment report, based on the task requirements or user needs.

[0051] In summary, this embodiment achieves efficient processing, intelligent matching and dynamic adaptation of visual cognitive tasks through the process of user analysis agent parsing task types, automatically selecting applicable large models, dynamically loading special models, performing visual analysis and outputting results, thereby improving the applicability and scalability of the system.

[0052] In some embodiments, when the cognitive task is of a general type, calling the general large model to process the cognitive task includes: The general large model is trained based on the automated data cleaning process and unlabeled data learning technology, so that the trained general large model has multi-modal visual positioning, visual classification and image content understanding capabilities; receiving input data for a general type of cognitive task provided by a user analysis agent; A general large model is used to perform visual feature extraction and multimodal semantic analysis on the input data of general types of cognitive tasks, generating general analysis results for cognitive tasks.

[0053] Specifically, in this embodiment, when the cognitive task belongs to a general type, the general large model is called to process the cognitive task. The general large model is a multimodal visual large model trained based on the automated data cleaning process and unlabeled data learning technology. It has the capabilities of visual positioning, visual classification, image content understanding, etc., and can be used for various general visual analysis tasks.

[0054] During the training process, the system first pre-processes large-scale image data through an automated data cleaning process to remove low-quality data, duplicate data, and abnormal data, and extracts effective features from unlabeled data based on self-supervised learning methods. Specifically, the system uses contrastive learning, autoencoders, generative adversarial networks (GANs), and other technologies to enable the model to learn common visual features without the need for a large amount of labeled data.

[0055] When the system receives the input data of the general type of cognitive task provided by the user analysis agent, the general large model begins to extract visual features and perform multimodal semantic analysis on the input data. For example: Visual localization task: If the user inputs "identify the object in this picture", the system will call the object detection module of the general large model, extract features of the image, and return the detection box and category label.

[0056] Visual classification task: If the user inputs "identify the type of plant in this picture", the system calls the image classification module of the general large model, classifies based on the learned visual features, and outputs the plant category.

[0057] Image content understanding task: If the user inputs "analyze the content of this poster", the system uses the multimodal alignment technology of the general large model to combine visual and text information to generate semantic analysis results.

[0058] In the process of visual feature extraction and multimodal semantic analysis, the general large model can further extract low-level features (such as edges, colors), intermediate features (such as shapes, textures) and high-level features (such as object categories, scene semantics) step by step through a hierarchical feature extraction mechanism, and combine the attention mechanism to improve the analysis accuracy. For example, in an autonomous driving environment, the system can combine visual positioning and image content understanding functions to identify vehicles, pedestrians, and traffic signs on the road, and perform risk assessment based on environmental information.

[0059] Finally, the general large model generates general analysis results for cognitive tasks and passes them to the result analysis agent for further processing or decision-making. For example, in a security monitoring scenario, if the model detects abnormal behavior, the result analysis agent can further evaluate the risk level and decide whether to trigger an alarm mechanism.

[0060] In summary, this embodiment trains a general large model through automated data cleaning and unlabeled data learning technology, combined with multimodal visual analysis capabilities, to achieve efficient processing of general visual cognition tasks and ensure that the system can maintain high generalization capabilities in environments with limited or changing data.

[0061] In some embodiments, when the cognitive task belongs to a specific vertical field, a dedicated large model is dynamically loaded and called to process the cognitive task, including: When the cognitive task belongs to a specific vertical field, a dedicated big model is retrieved, wherein the dedicated big model is a model generated by training based on professional data of a specific vertical field to perform visual analysis on professional tasks; According to the scene information or user intention provided by the user analysis agent, the dedicated large model is matched with parameters and task initialization is performed; Use a dedicated large model to perform field-specific visual analysis and feature extraction on input data, and output analysis results for specific vertical fields; During the execution process, the user analysis agent adjusts the calling strategy of the dedicated large model online according to user behavior data or scene changes, so as to dynamically mount and call the dedicated large model in a timely manner.

[0062] Specifically, in this embodiment, when the cognitive task belongs to a specific vertical field, the system will dynamically load and call a dedicated large model to process the cognitive task. The dedicated large model is a visual analysis model trained based on professional data in a specific vertical field, and provides more accurate visual cognitive capabilities for professional tasks such as medical image analysis, smart construction site monitoring, industrial testing, and security monitoring.

[0063] In the specific implementation process, the user analysis agent first analyzes the user input command, and combines the scene information or user intention to determine whether a dedicated large model needs to be called. For example: If the user inputs the command "analyze whether there are lung nodules in this medical image", the system will determine that the task belongs to the field of medical image analysis and needs to call a large model dedicated to medical images.

[0064] If the user inputs the command "Monitor safety hazards at the construction site", the system will determine that the task belongs to the field of smart construction site monitoring and needs to call a large model dedicated to construction safety.

[0065] After determining the cognitive task, the system will dynamically load the dedicated large model and perform parameter matching and task initialization. Task initialization includes: Load domain-specific trained weights, e.g. a medical imaging model may contain network parameters specific to CT, MRI, or X-ray.

[0066] Adjust the input format according to the scene information. For example, industrial inspection tasks may require preprocessing images to enhance defect areas.

[0067] Adjust the allocation of computing resources. For example, large-scale medical image analysis tasks can choose cloud computing, while real-time monitoring tasks can use edge computing.

[0068] The system then uses a dedicated large model to perform domain-specific visual analysis and feature extraction on the input data, and outputs analysis results for specific vertical fields. For example: In medical image analysis, the model can detect lesion areas in images and provide detailed reports on lesion size, location, and category.

[0069] In smart construction site monitoring, the model can analyze the construction progress and identify whether there are safety hazards, such as not wearing a safety helmet or illegal operations.

[0070] During the model execution process, the user analysis agent adjusts the calling strategy of the dedicated large model online according to user behavior data or scene changes, so as to achieve dynamic mounting and timely calling. For example: If users frequently adjust the detection target in medical image analysis tasks (such as switching from lung nodule detection to brain tumor detection), the system can automatically optimize the model selection strategy and give priority to recommending relevant models in future requests.

[0071] If the environment at the construction site changes (such as different lighting conditions during the day and at night), the system can automatically adjust the model parameters or preprocessing methods to ensure the stability of the analysis results.

[0072] In addition, the dedicated large model maintains real-time interaction with the user analysis agent, and the user analysis agent can learn the usage scenarios and capabilities of the dedicated model online to achieve more accurate task matching and call optimization. For example: When using a dedicated model for the first time, the system can provide recommended tasks to help users get started quickly.

[0073] If the user repeatedly switches between multiple tasks, the system can adjust the task priority, reduce unnecessary loading time, and improve task response speed.

[0074] In summary, this embodiment achieves efficient execution and dynamic adaptation of visual cognition tasks in professional fields by analyzing task types through user analysis agents, dynamically loading dedicated large models, optimizing parameter matching, performing domain-specific analysis, and adjusting model calling strategies based on user behaviors, thereby improving the intelligence and applicability of the system.

[0075] In some embodiments, when the output results of the large model do not meet the user's intent, the result analysis agent generates guidance information to prompt the user to supplement the data or redescribe the requirements, and sends the updated user input to the user analysis agent again to iteratively optimize the analysis of the user's intent.

[0076] Specifically, in this embodiment, when the output result of the large model does not meet the user's intention, the result analysis agent analyzes the model output and generates guidance information to prompt the user to supplement data or redescribe the requirements, thereby optimizing the task execution effect.

[0077] First, the result analysis agent receives the output of the general or special large model and evaluates the accuracy, completeness, and relevance of the results. For example: In the vehicle recognition task, if the model returns "unknown model", it may be because the image is blurred or there is no matching model in the database. At this time, the result analysis agent can determine that the output does not meet the user's intention.

[0078] In medical image analysis tasks, if the model's confidence in the lesion area is low (such as less than 70%), the system may determine that the result is not reliable enough and requires the user to provide clearer images or additional patient information.

[0079] Furthermore, when it is detected that the result does not satisfy the user's intention, the result analysis agent generates guidance information and provides suggestions for supplementary input to the user. For example: In the target detection task, the system may prompt: "The detected target is not clear, please upload a higher-resolution picture." In the semantic analysis task, the system may prompt: "Your question is rather vague. Please specify the object or scenario you want to focus on." In medical image analysis, the system may prompt: "The current test result has low confidence. Do you want to provide medical history information to improve the accuracy of diagnosis?" Guidance information can be presented in the form of text prompts, optional items, guided interactions, etc. For example: Provide candidate questions: "Would you like to identify the make or model of vehicle? Please select from the options below." Smart completion of user commands: "Your input is 'Analyze this picture', do you want the system to automatically detect the object category?" Furthermore, after the user supplements the data or adjusts the requirements according to the guidance information, the system will send the updated user input to the user analysis agent again to re-analyze the user's intention and optimize the task execution process.

[0080] In addition, the result analysis agent can also learn based on the user's behavior habits to improve the interactive experience. For example: If a user supplements the same type of information after a specific task multiple times (such as always providing medical history after a medical image analysis task), the system can proactively prompt the user to enter relevant data in advance in future tasks to reduce the number of interaction steps.

[0081] If the user frequently modifies the input content (such as adjusting the detection object multiple times), the system can adjust the initial task recommendation logic to prioritize the most commonly used task types.

[0082] If the system detects that the user fails to correctly understand the guidance information, the prompt method can be gradually adjusted, such as using more interactive guidance, such as intelligent question-and-answer style supplementary information interaction, rather than simple text prompts.

[0083] In summary, this embodiment evaluates the model output through the result analysis agent, intelligently generates user guidance information, optimizes the task parsing process, and combines user behavior learning to improve the interactive experience, thereby achieving dynamic optimization and adaptive improvement of visual cognitive tasks, and improving the intelligence, user-friendliness and task processing accuracy of the system.

[0084] In some embodiments, based on the analysis results of the result analysis agent, the decision agent is used to make a decision on the user question and output the processing result, including: If the analysis results do not meet the user's intention, the possible decision options are evaluated according to the user's needs, system resources and cognitive task results, and corresponding control instructions or prompt information are generated and returned to the user analysis agent or result analysis agent for further optimization; When the analysis results meet the user's intentions, a comprehensive decision is made on the feasible operations related to the user's problem based on the output data of the knowledge base and the large model to generate the final decision result.

[0085] Specifically, in this embodiment, based on the analysis results of the result analysis agent, the decision agent is used to make a decision on the user's question and output the final processing result. The core role of the decision agent is to comprehensively evaluate the analysis results and decide whether to further optimize the cognitive task or directly generate the final decision result.

[0086] First, the decision agent receives the output of the result analysis agent and evaluates whether it meets the user's intention. If the analysis result does not meet the user's needs, the decision agent will conduct a comprehensive evaluation of possible decision options based on user needs, system resources, and the results of cognitive tasks, and generate control instructions or prompt information, which will be returned to the user analysis agent or result analysis agent for further optimization. For example: In the vehicle recognition task, if the recognition result returned by the system has a low confidence level (for example, there may be multiple candidates for the recognized vehicle model), the decision-making agent can provide the user with multiple possible vehicle category options, or prompt the user to upload clearer pictures to improve recognition accuracy.

[0087] In medical image analysis tasks, if the confidence of lesion detection is insufficient, the decision agent can suggest that the user provide additional medical examination data (such as medical history, blood test results) and reroute the task to the result analysis agent to combine the new information for more in-depth analysis.

[0088] In smart security monitoring tasks, if the system detects suspicious behavior but is unsure of the risk level, the decision-making agent can decide whether to trigger an advanced analytical model or notify a human auditor.

[0089] When the analysis results satisfy the user's intention, the decision agent will make a comprehensive decision on the user's question based on the output data of the knowledge base and the large model, and generate the final decision result. For example: In an autonomous driving scenario, if the large model correctly identifies the obstacle ahead and calculates the safe stopping distance, the decision-making agent can directly generate a braking command and send it to the vehicle control system.

[0090] In the smart construction site monitoring task, if the large model successfully identifies abnormal construction progress, the decision-making agent can automatically send adjustment plans to managers or suggest reallocation of resources to optimize the construction progress.

[0091] In smart home control tasks, if the system recognizes that there is no one in the room and there has been no activity for a long time, the decision-making agent can generate energy-saving decision instructions such as turning off lights and reducing air conditioning power, and perform corresponding operations.

[0092] In addition, the decision-making agent can also adaptively optimize the decision-making strategy based on historical user behavior data and environmental changes. For example: In repetitive tasks (such as daily scheduled inspections of security camera images), the system can automatically adjust the detection threshold based on the user's past decision-making habits, reduce unnecessary alarm triggers, and improve system response efficiency.

[0093] During multiple rounds of interaction, if the user repeatedly changes the task requirements (such as modifying medical analysis parameters), the system can predict the user's preferences and actively recommend the optimal parameter configuration in future tasks.

[0094] In resource-constrained scenarios (such as edge computing devices), decision-making agents can combine the system computing power to decide whether to use cloud processing or local execution to optimize task execution efficiency.

[0095] In summary, this embodiment evaluates the analysis results through the decision-making agent, dynamically optimizes the task execution process, makes comprehensive decisions based on the knowledge base, and adapts to user behavior patterns, thereby realizing an intelligent decision-making closed loop, task execution optimization, and autonomous operation, thereby improving the system's automation level and task processing accuracy.

[0096] In some embodiments, after outputting the processing result, the method further includes: After receiving user feedback on the processing results, the user analysis agent and the result analysis agent are used to conduct continuous learning based on the user behavior and interaction history data, and dynamically adjust the parameters or calling strategies of the general large model and / or the special large model based on the learning results.

[0097] Specifically, in this embodiment, after outputting the processing results, the system supports continuous learning based on user feedback and dynamically optimizes the parameters and calling strategies of the large model to improve the long-term adaptability of the model and user experience.

[0098] First, the system receives user feedback on the processing results. The feedback can be in the form of direct user input, operation behavior, or system logs. For example: In the target recognition task, if the user modifies or confirms the system's recognition results (such as manually selecting the correct vehicle type category), the system can record the operation and use it to optimize the recognition strategy for subsequent tasks.

[0099] In medical image analysis tasks, if the doctor modifies or confirms the preliminary diagnosis results given by the system, the system can use this information to adjust the parameters of the dedicated large model to improve the recognition accuracy of similar cases in the future.

[0100] In intelligent monitoring tasks, if the user marks the system's abnormal alarms (such as false alarms or real alarms), the system can optimize the sensitivity of the anomaly detection model and reduce the false alarm rate.

[0101] Subsequently, the user analysis agent and the result analysis agent conduct continuous learning based on user feedback data, combined with user behavior patterns and historical interaction data, and optimize cognitive task execution strategies. For example: If users frequently adjust the parameters of the same type of tasks (such as adjusting the detection threshold of medical image analysis), the system can learn the user's preferences and use the optimal parameter configuration by default in future tasks, reducing the user's need for manual adjustments.

[0102] If certain tasks are often ignored or canceled by users, the system can analyze the possible reasons and optimize the task recommendation strategy to avoid providing irrelevant analysis results.

[0103] If the user's decision-making pattern changes over time (such as a construction site safety monitoring system has different alarm levels during the day and at night), the system can automatically adjust the model calling method based on time or environmental factors.

[0104] After learning from user feedback, the system will further dynamically adjust the parameters or call strategies of the general large model and / or the special large model to improve the adaptability of the model. For example: In the optimization of the general large model, the system can use feedback data to fine-tune the model weights, or use incremental training methods to continuously adapt the model to the new data distribution. For example, if the vehicle recognition system performs differently in different countries, the system can gradually learn localized vehicle features to improve recognition accuracy.

[0105] In the optimization of dedicated large models, the system can adjust the task adaptation rules. For example, according to the medical image analysis needs of specific medical institutions, the system can automatically adjust the image processing process to adapt to the diagnostic habits and data formats of different hospitals.

[0106] In the optimization of model calling strategy, the system can adjust the dynamic loading logic based on user feedback data, for example: If the usage frequency of a special model increases, local caching can be prioritized to reduce loading delays.

[0107] If the recognition error rate of a general model increases, a higher-level dedicated model or cloud computing resources can be automatically triggered to improve analysis accuracy.

[0108] In addition, the system can be based on federated learning or distributed model optimization technology to decentralize the feedback information of different users without leaking user privacy data, thereby improving the overall adaptability and generalization ability of the model. For example, medical image analysis models deployed in different hospitals can be personalized based on doctors' annotation feedback, while using federated learning to improve the general performance of the model without directly sharing sensitive medical data.

[0109] In summary, this embodiment achieves adaptive optimization, personalized adjustment and long-term evolution of the visual cognitive system through user feedback collection, continuous learning of intelligent agents, task strategy optimization and dynamic adjustment of large models, thereby improving the intelligence of the system and user experience, while ensuring that the model can continuously improve its task processing capabilities and accuracy over time.

[0110] According to the technical solutions of the above embodiments of the present application, the present application has at least the following advantages: 1. Autonomous decision-making and closed-loop system: An autonomous closed-loop system of "task analysis-routing-cognition-result analysis-decision-making" has been built, realizing full process automation from environmental cognition to decision-making execution. This design enables the system to respond in real time in a dynamic environment.

[0111] 2. Multimodal cognition and interaction: The system integrates multiple modal data such as vision and text, improving the ability to understand and interact with complex scenes. Through a multimodal large model that aligns vision and text, the system can more accurately parse and respond to complex instructions or situations, enhancing the naturalness and effectiveness of human-computer interaction. Interaction with users based on natural language greatly reduces the threshold for using the system and the difficulty of system integration.

[0112] 3. Efficient data utilization and cost savings: The use of automated data cleaning processes and unlabeled data learning technology has greatly reduced the reliance on manually labeled data, saving a lot of manpower and time costs. Through zero-sample learning capabilities, the system can quickly adapt to new tasks or scenarios without the need for additional data collection and labeling, greatly improving the flexibility and adaptability of the system.

[0113] 4. Dynamic expansion and continuous learning: The system has the ability to learn autonomously and can dynamically expand its cognitive and decision-making capabilities to adapt to changing environments and needs. This continuous learning mechanism enables the system to be continuously optimized and improved over time, especially in rapidly changing fields.

[0114] 5. Cross-domain application and adaptability: General and specialized integration, dynamic mounting of specialized models, can be easily expanded to different vertical fields. Through dynamic adjustment and optimization, the system can be efficiently applied in multiple fields such as medical care, smart construction sites, smart cities, smart agriculture, etc., promoting the popularization and popularization of visual cognition technology.

[0115] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.

[0116] Figure 3 Schematic diagram of the structure of a multi-agent driven multi-modal cognitive device provided in an embodiment of the present application. Figure 3As shown, the multi-agent driven multimodal cognitive device includes: The acquisition module 301 is used to acquire the instruction input by the user; The parsing module 302 is used to parse the instructions input by the user using the user analysis agent, extract the user's intention, and determine the cognitive task to be performed according to the user's intention; A calling module 303 is used to select and call at least one large model to perform cognitive tasks based on the parsing result of the user analysis agent, wherein the large model includes a general large model and a special large model; The analysis module 304 is used to analyze the output results of the large model using the result analysis agent. If the output results do not meet the user's intention, the user is prompted to provide additional information or restate the problem, and the updated user input is sent to the user analysis agent again. Otherwise, an analysis result is generated. The processing module 305 is used to make decisions on user questions based on the analysis results of the result analysis agent using the decision agent and output the processing results.

[0117] In some embodiments, Figure 3 The parsing module 302 performs semantic and contextual analysis on the instructions input by the user to identify the target information and related constraints in the instructions; extracts key elements based on the results of semantic and contextual analysis, and generates or updates the corresponding user intentions; matches and expands the user intentions based on a pre-built knowledge base or user historical behavior data to form a candidate cognitive task set; screens and sorts the candidate cognitive task set based on user preferences, system resource availability, and the priority of user intentions; determines the cognitive tasks to be executed and generates corresponding task routing information, which is used to call the large model to complete the corresponding visual analysis.

[0118] In some embodiments, Figure 3 The calling module 303 obtains the cognitive task to be executed and the corresponding user intention output by the user analysis agent; determines whether the cognitive task belongs to a general type or a specific vertical field according to the user intention and the pre-established model selection strategy; when the cognitive task is of a general type, calls the general big model to process the cognitive task; when the cognitive task belongs to a specific vertical field, dynamically loads and calls the special big model to process the cognitive task; and passes the output result generated by the general big model or the special big model to the result analysis agent.

[0119] In some embodiments, Figure 3The calling module 303 trains the general big model based on the automated data cleaning process and unlabeled data learning technology, so that the trained general big model has multimodal visual positioning, visual classification and image content understanding capabilities; receives input data for general type cognitive tasks provided by the user analysis agent; uses the general big model to extract visual features and perform multimodal semantic analysis on the input data of the general type cognitive tasks, and generates general analysis results for cognitive tasks.

[0120] In some embodiments, Figure 3 The calling module 303 calls the dedicated big model when the cognitive task belongs to a specific vertical field, wherein the dedicated big model is a model for visual analysis of professional tasks generated by training based on professional data of the specific vertical field; according to the scene information or user intention provided by the user analysis agent, the dedicated big model is parameter matched and task initialized; the dedicated big model is used to perform domain-specific visual analysis and feature extraction on the input data, and the analysis results for the specific vertical field are output; during the execution process, the user analysis agent adjusts the calling strategy of the dedicated big model online according to the user behavior data or scene changes, so as to dynamically mount and call the dedicated big model in a timely manner.

[0121] In some embodiments, Figure 3 When the analysis result does not meet the user's intention, the processing module 305 evaluates possible decision options based on user needs, system resources and cognitive task results, generates corresponding control instructions or prompt information, and returns them to the user analysis agent or result analysis agent for further optimization; when the analysis result meets the user's intention, based on the output data of the knowledge base and the large model, a comprehensive decision is made on the feasible operations related to the user's problem to generate a final decision result.

[0122] In some embodiments, Figure 3 The adjustment module 306 is used to utilize the user analysis agent and the result analysis agent to continuously learn according to the user behavior and interaction history data after outputting the processing result and receiving the user's feedback on the processing result, and dynamically adjust the parameters or calling strategies of the general large model and / or the special large model based on the learning result.

[0123] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0124] Figure 4 Schematic diagram of the structure of the electronic device 4 provided in the embodiment of the present application. Figure 4As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0125] Exemplarily, the computer program 403 may be divided into one or more modules / units, which are stored in the memory 402 and executed by the processor 401 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 403 in the electronic device 4.

[0126] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 4 and does not constitute a limitation of the electronic device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0127] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0128] The memory 402 may be an internal storage unit of the electronic device 4, for example, a hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Further, the memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 may also be used to temporarily store data that has been output or is to be output.

[0129] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0130] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0131] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0132] In the embodiments provided in the present application, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0133] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0134] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0135] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0136] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A multi-agent driven multimodal cognitive method, characterized in that: include: Get the command input by the user; Utilizing a user analysis agent to parse the instructions input by the user, extract the user's intention, and determine the cognitive task to be performed according to the user's intention; Based on the analysis result of the user analysis agent, select and call at least one large model to perform the cognitive task, wherein the large model includes a general large model and a special large model; Analyze the output results of the large model using a result analysis agent. If the output results do not meet the user's intention, prompt the user to provide additional information or restate the problem, and send the updated user input to the user analysis agent again. Otherwise, generate an analysis result. Based on the analysis results of the result analysis agent, the decision-making agent is used to make decisions on the user's questions and output the processing results.

2. The method according to claim 1, characterized in that The utilizing of the user analysis agent to parse the instruction input by the user, extract the user intention, and determine the cognitive task to be performed according to the user intention, includes: Performing semantic and context analysis on the instruction input by the user to identify target information and related constraints in the instruction; Extract key elements based on semantic and context analysis results and generate or update corresponding user intent; Based on a pre-built knowledge base or user historical behavior data, the user intention is matched and expanded to form a candidate cognitive task set; Filtering and sorting the candidate cognitive task set based on user preferences, system resource availability, and the priority of the user's intention; The cognitive tasks to be performed are determined, and corresponding task routing information is generated, where the task routing information is used to call the large model to complete the corresponding visual analysis.

3. The method according to claim 1, characterized in that The selecting and calling at least one large model to perform the cognitive task based on the parsing result of the user analysis agent includes: Obtaining the cognitive tasks to be performed and the corresponding user intentions output by the user analysis agent; Determining whether the cognitive task belongs to a general type or a specific vertical field based on the user intent and a pre-established model selection strategy; When the cognitive task is of a general type, calling the general large model to process the cognitive task; When the cognitive task belongs to a specific vertical field, dynamically loading and calling the dedicated large model to process the cognitive task; The output result generated by the general large model or the dedicated large model is delivered to the result analysis agent.

4. The method according to claim 3, characterized in that When the cognitive task is of a general type, calling the general large model to process the cognitive task includes: The general large model is trained based on an automated data cleaning process and unlabeled data learning technology, so that the trained general large model has multimodal visual positioning, visual classification and image content understanding capabilities; Receiving input data for a general type of cognitive task provided by the user analysis agent; The general large model is used to perform visual feature extraction and multimodal semantic analysis on the input data of the general type of cognitive task, so as to generate a general analysis result for the cognitive task.

5. The method according to claim 3, characterized in that: When the cognitive task belongs to a specific vertical field, dynamically loading and calling the dedicated large model to process the cognitive task includes: When the cognitive task belongs to a specific vertical field, the dedicated big model is retrieved, wherein the dedicated big model is a model generated by training based on professional data of the specific vertical field and performing visual analysis on professional tasks; According to the scene information or user intention provided by the user analysis agent, parameter matching and task initialization are performed on the dedicated large model; Using the dedicated large model to perform field-specific visual analysis and feature extraction on input data, and output analysis results for a specific vertical field; During the execution process, the user analysis agent adjusts the calling strategy of the dedicated large model online according to user behavior data or scene changes, so as to dynamically mount and call the dedicated large model in a timely manner.

6. The method according to claim 1, characterized in that Based on the analysis result of the result analysis agent, the decision agent is used to make a decision on the user question and output the processing result, including: In the case where the analysis result does not meet the user's intention, possible decision options are evaluated according to user needs, system resources and cognitive task results, and corresponding control instructions or prompt information are generated and returned to the user analysis agent or result analysis agent for further optimization; When the analysis result meets the user's intention, a comprehensive decision is made on feasible operations related to the user's question based on the knowledge base and the output data of the large model to generate a final decision result.

7. The method according to claim 1, characterized in that After outputting the processing result, the method further includes: After receiving user feedback on the processing results, the user analysis agent and the result analysis agent are used to perform continuous learning based on user behavior and interaction history data, and dynamically adjust the parameters or calling strategies of the general large model and / or the special large model based on the learning results.

8. A multi-agent driven multimodal cognitive device, characterized in that: include: An acquisition module is used to obtain instructions input by the user; A parsing module, used to parse the instructions input by the user using a user analysis agent, extract the user's intention, and determine the cognitive task to be performed according to the user's intention; A calling module, configured to select and call at least one large model to perform the cognitive task based on the parsing result of the user analysis agent, wherein the large model includes a general large model and a special large model; An analysis module is used to analyze the output result of the large model using a result analysis agent, and if the output result does not meet the user's intention, prompt the user to provide additional information or restate the problem, and send the updated user input to the user analysis agent again, otherwise generate an analysis result; The processing module is used to make decisions on user questions based on the analysis results of the result analysis agent using the decision agent and output the processing results.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Fault diagnosis device and method based on multi-agent system and wavelet analysis

    CN102508076A

  • Data prediction system and method based on double-layer model structure

    CN111752556A

  • Intelligent brain system for hydropower station monitoring data analysis and construction method thereof

    CN115456216A

  • Intelligent data analysis system and method based on intelligent agent

    CN116975042A

  • Professional document generation method and device, computer equipment and storage medium

    CN117454863A

Cited By

  • Incremental learning-based self-evolution method and system for multi-mode general-purpose model

    CN120611769A

  • Interaction method and device based on agent cooperation, agent and storage medium

    CN120654731A

  • Interaction method and device based on agent cooperation, agent, and storage medium

    CN120654731B

  • Industrial hidden danger troubleshooting decision-making method based on multi-agent cooperation

    CN120725625A

  • Home monitoring method and system based on multi-agent large model

    CN120751089A