Multi-modal information processing method and device, equipment and storage medium
Through the multimodal large model, the intent identification and planning of multimodal information is solved, and the intent analysis and operator call difficulties of the intelligent system under multimodal input is achieved, and more efficient and accurate intelligent services are achieved, suitable for fields such as autonomous driving, medical diagnosis and intelligent teaching aids.
Patent Information
- Application Number
- CN202510273062.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-04
AI Technical Summary
When facing multimodal input, existing agent systems are difficult to accurately identify user intentions and effectively call corresponding operator capabilities, which limits the depth and breadth of intelligent services, especially in general visual problems.
The multimodal large model is used for intention recognition and planning. By classifying the multimodal information, disassembly the problem and analyzing the thinking process, and calling appropriate operators to improve the accuracy and efficiency of the output results.
The intention analysis and planning capabilities of the intelligent system in multimodal information processing have been improved, the identification of user needs and operator calls have been enhanced, and the quality and efficiency of intelligent services have been improved, especially in complex visual interaction scenarios such as autonomous driving, medical diagnosis and intelligent teaching aids.
Smart Images

Figure CN120258132A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technologies, and particularly to technical fields such as computer vision, deep learning, and large models. Background Art
[0002] With the rapid development of artificial intelligence technologies, technologies such as large language models (LLMs), multimodal large language models (MLLMs), and agents are becoming increasingly important. Large language models can be pre-trained based on deep learning, such as the Transformer architecture, through large-scale corpus data, so as to achieve in-depth understanding and generation of human language. Multimodal large models can process and understand various types of information, can fuse various modal data such as text, images, audio, and video, and perform comprehensive understanding and reasoning, ultimately achieving more powerful capabilities. An agent is a system that can autonomously perceive the environment, make decisions, and execute actions. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, and storage medium for processing multimodal information.
[0004] According to one aspect of this disclosure, a method for processing multimodal information is provided, including:
[0005] Performing intent recognition on multimodal information to obtain the intent category corresponding to the multimodal information;
[0006] Obtaining thinking process information according to the multimodal information and its corresponding intent category;
[0007] Obtaining the operator information to be called according to the thinking process information; wherein, the operator information includes information about the application programs used in the process of processing the multimodal information;
[0008] Obtaining the output result of the multimodal information according to the intent category, the thinking process information, and the operator information.
[0009] According to another aspect of this disclosure, a method for training a multimodal large model is provided, including:
[0010] Inputting the multimodal information in the training dataset into the multimodal large model to be trained to obtain an output result;
[0011] According to the output result and the output label corresponding to the multimodal information in the training dataset, perform supervised fine-tuning on the multimodal large model to be trained to obtain a multimodal large model with a planning function; wherein, the multimodal large model with a planning function is used to execute the planning function of the agent system.
[0012] According to another aspect of the present disclosure, there is provided a processing device for multimodal information, including:
[0013] An intention recognition module, configured to perform intention recognition on the multimodal information to obtain the intention category corresponding to the multimodal information;
[0014] A thinking module, configured to obtain thinking process information according to the multimodal information and its corresponding intention category;
[0015] An invocation module, configured to obtain the operator information to be invoked according to the thinking process information; wherein, the operator information includes information about the application program used in the process of processing the multimodal information;
[0016] An output module, configured to obtain the output result of the multimodal information according to the intention category, the thinking process information, and the operator information.
[0017] According to another aspect of the present disclosure, there is provided a training device for a multimodal large model, including:
[0018] An output input module, configured to input the multimodal information in the training dataset into the multimodal large model to be trained to obtain an output result;
[0019] A supervised fine-tuning module, configured to perform supervised fine-tuning on the multimodal large model to be trained according to the output result and the output label corresponding to the multimodal information in the training dataset to obtain a multimodal large model with a planning function; wherein, the multimodal large model with a planning function is used to execute the planning function of the agent system.
[0020] According to another aspect of the present disclosure, there is provided an agent system, including:
[0021] A planning module, configured to perform intention recognition on the multimodal information to obtain the intention category corresponding to the multimodal information; obtain thinking process information according to the multimodal information and its corresponding intention category; obtain the operator information to be invoked according to the thinking process information;
[0022] An execution module, configured to obtain the output result of the multimodal information according to the intention category, the thinking process information, and the operator information.
[0023] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0024] at least one processor; and
[0025] a memory communicatively connected to the at least one processor; wherein,
[0026] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0027] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0028] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements any method in the embodiments of the present disclosure when executed by a processor.
[0029] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0031] Figure 1 is a schematic flowchart of a method for processing multi-modal information according to an embodiment of the present disclosure;
[0032] Figure 2 is a schematic flowchart of a method for processing multi-modal information according to another embodiment of the present disclosure;
[0033] Figure 3 is a system framework diagram of a multi-modal large model;
[0034] Figure 4 is a schematic flowchart of a method for training a multi-modal large model according to an embodiment of the present disclosure;
[0035] Figure 5 is a schematic structural diagram of a device for processing multi-modal information according to an embodiment of the present disclosure;
[0036] Figure 6 is a schematic structural diagram of a device for training a multi-modal large model according to an embodiment of the present disclosure;
[0037] Figure 7 is a schematic structural diagram of an agent system according to an embodiment of the present disclosure;
[0038] Figure 8 It is a framework diagram of an Agent intelligent agent;
[0039] Figure 9 It is a framework diagram of an intention analysis and task planning system that supports image and text input;
[0040] Figure 10 It is a flowchart of multi-modal large model supervised fine-tuning;
[0041] Figure 11 It is a block diagram of an electronic device for implementing the embodiments of the present disclosure. Detailed implementation manners
[0042] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0043] Figure 1 It is a schematic flowchart of a method 100 for processing multi-modal information according to an embodiment of the present disclosure. The method includes:
[0044] S110. Perform intention recognition on the multi-modal information to obtain the intention category corresponding to the multi-modal information;
[0045] S120. Obtain thinking process information according to the multi-modal information and its corresponding intention category;
[0046] S130. Obtain operator information to be called according to the thinking process information; wherein, the operator information includes information on application programs that may be used in the process of processing the multi-modal information;
[0047] S140. Obtain the output result of the multi-modal information according to the intention category, the thinking process information, and the operator information.
[0048] In the embodiments of the present disclosure, multimodal technology may include ways of expressing, communicating, and understanding using information in various different forms or perceptual channels, such as various sensory input and output methods including vision, audition, text, touch, etc. By integrating data from different modalities, such as images, texts, audio, or videos, the understanding ability and reasoning ability of the model can be enhanced. Among them, images and videos belong to visual information. Multiple intention categories can be preset, and through the large model for intention recognition of the input multimodal information, the intention category corresponding to the multimodal information can be output (which can also be referred to as type, category, name, etc.). The intention categories may include, but are not limited to, one or more of the following: question answering based on pictures, image reasoning, text information processing, picture translation, picture creation, solving liberal arts problems, solving science problems, and code generation. For example, if the content input to the multimodal large model includes an image and text, and the text content includes a question about the image, the intention recognized by the multimodal large model based on the image and the text can be question answering based on pictures. Another example, if the content input to the multimodal large model includes an image and speech, and the speech content includes translation of the language in the image, the intention output by the multimodal large model based on the image and the text can be picture translation. Another example, if the content input to the multimodal large model includes an image with a science problem, the intention recognized by the multimodal large model based on the image can be solving science problems. There are many examples of different intention categories, which can be determined according to the actual application scenario.
[0049] In the embodiments of the present disclosure, as Figure 2 shown, the agent system may include a planning module based on a multimodal large model (or referred to as MLLM Planner, multimodal large model planning module, multimodal large model with planning function, etc.) and an actuator module, etc. The training method of the planning module based on the multimodal large model can refer to the specific description of the training method of the following multimodal large model. After obtaining the intention category corresponding to the multimodal information using the planning module based on the multimodal large model, the chain of thought (CoT) reasoning can be performed by combining the intention category and the input multimodal information to obtain the thought process information. The thought process may include the process of disassembling and analyzing the multimodal information. Through the thought process, problems with relatively complex logic in the multimodal information can be disassembled, and through a series of logically related thoughts, a complete thought process is formed. The thought logic, operators to be called, and other information generated during the thought process can be referred to as thought process information. Some application programs may be required during the process of processing the multimodal information, and these application programs can be understood as operators. From the thought process information, the category and other operator information of the operators to be called can be extracted. There can be various categories of operators, such as: animal recognition, plant recognition, vehicle recognition, landmark recognition, encyclopedia search, logo recognition, text recognition, dish recognition, celebrity recognition, etc.
[0050] In an embodiment of the present disclosure, using a planning module based on a multimodal large model, after the multimodal information undergoes the above steps such as intention recognition, thinking process, and operator analysis, the intention category, thinking process information, and operator information corresponding to the multimodal information can be transmitted to the actuator module. The actuator module can call an appropriate multimodal large model according to the intention category and thinking process information, set reasonable prompt words, and can also call relevant operators to process the multimodal information to obtain a final output result.
[0051] In the embodiment of the present disclosure, by performing intention recognition on multimodal information, planning can be continued according to the recognized intention classification to obtain thinking process information and operator information, improving the accuracy of the output result, reducing the processing time, and increasing the processing speed.
[0052] Figure 3 FIG. 300 is a schematic flowchart of a method 300 for processing multimodal information according to another embodiment of the present disclosure. The method 300 can be used to implement step S110 in the method 100 for processing multimodal information. In one embodiment, the method 300 includes: performing intention recognition on multimodal information to obtain the intention category corresponding to the multimodal information, and further including:
[0053] S310: Performing intention recognition on the image and / or text in the multimodal information to obtain the intention category corresponding to the problem in the image and / or text.
[0054] In the embodiment of the present disclosure, taking the image and / or text included in the multimodal information as an example, multimodal information such as the image and / or text can be input into a planning module based on a multimodal large model to perform intention classification on the image and / or text. The specific content of the problem can be recognized from the text, and intention classification can be performed according to the specific content of the problem. The text can be recognized from the image, and then the specific content of the problem can be recognized, and intention classification can be performed according to the specific content of the problem. By performing intention recognition on multimodal information such as the image and / or text, planning can be continued according to the recognized intention classification to obtain thinking process information and operator information, improving the accuracy of the output result for multimodal information such as the image and / or text, and reducing the processing time for multimodal information such as the image and / or text.
[0055] In one embodiment, step S120 obtains thinking process information according to the multimodal information and its corresponding intention category, including:
[0056] S320. According to the multimodal information and its corresponding intention category, disassemble and analyze the multimodal information, and describe the category of application programming interface (API) operators that need to be called to obtain the thinking process information.
[0057] In the embodiments of the present disclosure, the input multimodal information may include one or more problems to be solved. Using the planning module based on the multimodal large model, these problems can be disassembled and analyzed through the chain of thought, and then the information about the possible applications that can be used to process one piece of information, multiple pieces of information, one problem, or multiple problems, etc. in the multimodal information, such as the category of API operators, can be obtained. For example, if the text of the multimodal information includes problems Q1 and Q2, problems Q1 and Q2 can be disassembled, and it can be analyzed which API operators may be needed to solve problems Q1 and Q2. For example, operator O1 and O2 are used for problem Q1, and operator O3 is used for problem Q2. For these problems, through a series of logically related thoughts, a complete thinking process can be formed. For example, the logical relationship between problems Q1 and Q2 is to solve problem Q1 first, and then solve problem Q2 based on the solution corresponding to problem Q1. The thinking process information may include: first, solve problem Q1 through operators O1 and O2, and then use the solution corresponding to problem Q1 as the input content of problem Q2, and solve problem Q2 through operator O3. Obtaining the thinking process information through the ability of the chain of thought can reduce the difficulty of problem-solving, improve the accuracy of the output result for multimodal information, and reduce the processing time of multimodal information.
[0058] In one implementation manner, step S130 obtains the operator information that needs to be called according to the thinking process information, including:
[0059] S330. According to the thinking process information, obtain the API analysis result; the API analysis result includes the category of the API operator that needs to be called and the input parameters of the API operator.
[0060] In the embodiments of the present disclosure, a planning module based on a multimodal large model can perform API analysis on the thinking process information to obtain the category (or name) of one or more API operators to be called. Further, the input parameters (or response parameters) of the API operators can also be obtained. The input parameters of the API operators can include the necessary input parameters for requesting the operator service. For example, when the operator to be called includes a search operator, input parameters such as the corresponding search question need to be input. When the operator to be called includes a detection operator, input parameters such as the name of the object to be detected need to be input. According to the categories and input parameters of these API operators, the actuator module can call these API operators and input the input parameters of these API operators into the corresponding operators to execute the calculation process. Through the categories and input parameters of the API operators, appropriate API operators can be quickly and reasonably called to perform calculations on multimodal information, thereby improving the processing speed and the accuracy of the processing results.
[0061] In one implementation, step S140 obtains the output result of the multimodal information according to the intention category, the thinking process information, and the operator information, including:
[0062] S340. Perform operator call, model call, and prompt setting according to the intention category, the thinking process information, and the operator information to obtain the output result of the multimodal information.
[0063] In the embodiments of the present disclosure, the actuator module of the agent system can distribute the calling strategies of downstream multimodal large models according to the intention category, including prompt design, operator combination methods for calling, etc. Select and call an appropriate model according to the intention category corresponding to the multimodal information, such as a multimodal large model with a certain function, and call an appropriate operator according to the operator information to obtain a calculation result. Then, set appropriate prompts according to one or more of the intention category, thinking process information, and calculation results of the operator. For example, obtain a matching prompt template according to the intention category, and add the calculation results of one or more operators to the prompt template to obtain a prompt that can be input into the model, so as to input the calculation results of the operators as part of the knowledge into the model. Inputting the prompt into the called model can obtain the output result of the multimodal information. For example, if the intention category is science problem-solving, a multimodal large model with the ability to think about long texts can be called, and a prompt template that is more proficient in science problem-solving can be set. When calling operators, operators such as Optical Character Recognition (OCR) and graph search for problem-solving tend to be called.
[0064] By performing operator call, model call, and prompt setting according to the intention category, thinking process information, and operator information, a more accurate output result can be obtained.
[0065] In one embodiment, as Figure 2 shown, in addition to including a planning module (MLLM Planning or Planner) based on a multimodal large model, the agent system may further include one or more of the following: an agent, a toolset, and a memory network. The toolset may include a collection of external tools that can be invoked, and can also be referred to as an API or an operator set, etc.
[0066] Figure 4 is a schematic flowchart of a training method 400 for a multimodal large model according to an embodiment of the present disclosure. The method includes:
[0067] S410: Input the multimodal information in the training dataset into the multimodal large model to be trained, and obtain an output result;
[0068] S420: According to the output result and the output label corresponding to the multimodal information in the training dataset, perform supervised fine-tuning on the multimodal large model to be trained, and obtain a multimodal large model with a planning function.
[0069] In one embodiment, the multimodal large model with a planning function is used to execute the planning function of the agent system.
[0070] In the embodiments of the present disclosure, the training dataset may include several training labels, including input labels and output labels of the training labels. For example, the input labels may include multimodal information such as images and / or texts. The output labels may include the content expected to be output by the modal large model. After inputting the input labels of each training label in the training dataset into the multimodal large model to be trained, the loss function can be calculated according to the output result of the model and the output label corresponding to each input label. Based on the loss function, the parameters of the model are supervised and fine-tuned. This supervised fine-tuning process can be cyclic, and the training datasets used each time can be different or the same. After the loss function meets the convergence condition, a trained multimodal large model can be obtained. The trained multimodal large model is a multimodal large model with a planning function. This multimodal large model can be used as the planning module of the agent system to execute the planning function of the agent system. The specific functions of this planning module can be referred to the relevant descriptions in any of the above-described multimodal information processing methods in the embodiments. By performing supervised fine-tuning on the multimodal large model to obtain a multimodal large model with a planning function, and thus obtaining the planning module of the agent system, it is beneficial to the planning rationality of the multimodal information processing process and improves the accuracy of the processing results.
[0071] In one embodiment, the output label corresponding to the multimodal information includes one or more of the following: intention category; thinking process information; operator information.
[0072] In the embodiments of the present disclosure, in the training dataset, the input labels of the training labels may include multimodal information. The output labels of the training labels may include the intention categories, thinking process information, and operator information expected to be output by the multimodal large model. After inputting the input labels of each training label in the training dataset into the multimodal large model to be trained, the loss function can be calculated based on the intention categories, thinking process information, and operator information output by the model, as well as the intention categories, thinking process information, and operator information corresponding to each input label. Based on the loss function, the parameters of the model are supervised and fine-tuned. After the loss function satisfies the convergence condition, the trained multimodal large model can be obtained. By using the multimodal information, intention categories, thinking process information, and operator information in the training labels to supervise and fine-tune the multimodal large model, a multimodal large model with a planning function is obtained, thereby obtaining the planning module of the intelligent agent system, which is beneficial to more reasonable thinking and accurately finding the appropriate operator after intention recognition of multimodal information, and improving the accuracy of the processing results.
[0073] Figure 5 FIG. 5 is a schematic structural diagram of a processing device 500 for multimodal information according to an embodiment of the present disclosure. The device 500 may include:
[0074] An intention recognition module 510, configured to perform intention recognition on multimodal information to obtain the intention category corresponding to the multimodal information;
[0075] A thinking module 520, configured to obtain thinking process information according to the multimodal information and its corresponding intention category; wherein, the operator information includes information about application programs that may be used in the process of processing the multimodal information;
[0076] An invocation module 530, configured to obtain the operator information to be invoked according to the thinking process information;
[0077] An output module 540, configured to obtain the output result of the multimodal information according to the intention category, the thinking process information, and the operator information.
[0078] In one implementation, the intention recognition module 510 may be configured to perform intention recognition on the images and / or texts in the multimodal information to obtain the intention category corresponding to the problems in the images and / or texts.
[0079] In one implementation, the thinking module 520 may be configured to disassemble and analyze the multimodal information according to the multimodal information and its corresponding intention category, and describe the category of API operators to be invoked to obtain the thinking process information.
[0080] In one implementation, the call module 530 can be used to obtain an API analysis result based on the thinking process information; the API analysis result includes the category of the API operator to be called and the input parameters of the API operator.
[0081] In one implementation, the output module 540 can be used to perform operator calls, model calls, and prompt word settings based on the intent category, the thinking process information, and the operator information, and obtain the output result of the multimodal information.
[0082] In one implementation, the multimodal large model further includes one or more of the following: an agent, a tool set, and a memory network.
[0083] Figure 6 FIG. 10 is a schematic structural diagram of a training device 600 for a multimodal large model according to an embodiment of the present disclosure. The device 600 may include:
[0084] An output input module 610, configured to input multimodal information in a training data set into a multimodal large model to be trained, and obtain an output result;
[0085] A supervised fine-tuning module 620, configured to perform supervised fine-tuning on the multimodal large model to be trained according to the output result and the output label corresponding to the multimodal information in the training data set, and obtain a multimodal large model with a planning function.
[0086] In one implementation, the multimodal large model with a planning function is used to execute the planning function of the agent system.
[0087] In one implementation, the output label corresponding to the multimodal information includes one or more of the following: an intent category; thinking process information; operator information.
[0088] Figure 7 FIG. 26 is a schematic structural diagram of an agent system 700 according to an embodiment of the present disclosure. The agent system 700 may include:
[0089] A planning module 710, configured to perform intent recognition on multimodal information to obtain the intent category corresponding to the multimodal information; obtain thinking process information according to the multimodal information and its corresponding intent category; obtain operator information to be called according to the thinking process information;
[0090] An execution module 720, configured to obtain the output result of the multimodal information according to the intent category, the thinking process information, and the operator information.
[0091] Further, as Figure 2 shown, the agent system may further include a tool set, a memory network, etc.
[0092] For the specific functions and examples of each module and sub-module of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0093] Agents possess basic characteristics such as autonomy, interactivity, reactivity, and adaptability, and can independently complete tasks in complex and changing environments. The emergence of agents marks the advancement of artificial intelligence from simple rule matching and computational simulation to a higher level of autonomous intelligence. However, when faced with multi-modal inputs, relatively ambiguous or contextual user inputs, agents often have difficulty accurately identifying user intentions, thereby effectively invoking corresponding operator capabilities, which limits the depth and breadth of intelligent services.
[0094] A framework diagram of an agent based on a large language model is as Figure 8 shown. Among them, the planning module (planning) is responsible for thinking about the user's problem intention, decomposing the problem, and making corresponding planning and scheduling. The agent framework can include planning capabilities. The planning module of the agent framework can be implemented based on a large language model and has capabilities such as problem decomposition and operator call planning. However, since the intention analysis and planning capabilities of agents are basically realized with the help of large language models, they have advantages such as powerful text understanding capabilities. However, some large language models do not have multi-modal capabilities and cannot directly understand image information; moreover, the deployment and invocation costs of large language models are relatively high, making it difficult to actually implement the agent system. The large language models in related technologies have insufficient intention analysis and a narrow range of task planning operators, etc.
[0095] The embodiments of the present disclosure can improve the capabilities of agents in intention analysis and planning, especially for agent solutions in general vision problem processing. Specifically, the present disclosure proposes an intention analysis and planning system based on a multi-modal large model, which can integrate various modal information such as text and images, deeply understand the true intention behind the user input, accurately judge the set of operator capabilities required, and give the solution steps to solve the user's problem, etc. The embodiments of the present disclosure can improve the capabilities of agents in intention recognition and operator call planning for user needs in the field of general vision problems, laying a solid foundation for providing more personalized and efficient intelligent services.
[0096] In the system of the embodiments of the present disclosure, the agent can undertake the tasks of intention analysis and task planning. It can automatically analyze the problem attributes according to various forms of user input, such as copywriting creation, code writing, etc., and intelligently select the most suitable operator for subsequent processing. The application scope of this solution is extensive. In scenarios involving complex visual interactions such as autonomous driving, medical diagnosis, and intelligent teaching aids, it can greatly improve the service quality and user experience of the agent, and promote the in-depth application and development of artificial intelligence technology in a wider range of fields.
[0097] Based on multimodal large models, Agent technology, etc., the present disclosure realizes an intention analysis and task planning system that supports image and text input.
[0098] An exemplary framework diagram is as Figure 9 shown. For the input image (Image) and text (including questions (Question)) of the user, the planning module (MLLM Planner) of the system will first analyze the intention category of the input question. The intention category mainly refers to the classification of the user's question for the divide-and-conquer processing of the system. Examples of intention categories can include: question answering based on pictures, picture reasoning, text information processing, picture translation, picture creation, solving liberal arts problems, solving science problems, and code generation, etc.
[0099] Secondly, the system will disassemble and analyze the problem based on the user's question, and describe the API operator categories (or types, names, etc.) that need to be called. This part of the content is a manifestation of the Chain of Thought (CoT), which helps to reduce the difficulty of the large model in solving problems. For example, the question includes a query request "Query: Which anime does the character in the picture come from?" The thought process information can include "Thought: The user's question is to ask about the anime character in the picture. Therefore, it is necessary to first use the image recognition ability, combined with knowledge retrieval, to identify the anime character in the picture and further determine the anime work to which the character belongs. Call the celebrity recognition and encyclopedia operators to achieve this. The execution steps are: v1. Use the celebrity recognition ability to identify the anime character in the picture. v2. Based on the identified anime character, determine the anime work to which the character belongs through encyclopedia knowledge retrieval."
[0100] Then, based on the thinking process (or deliberation process), the system gives the API analysis results, mainly including the names of the operators to be called and input parameters, etc. For example, examples of supported API operators can include: animal recognition, plant recognition, vehicle recognition, landmark recognition, encyclopedia search, logo recognition, text recognition, dish recognition, celebrity recognition, etc. These API operators can assist the system in image perception and information search, etc. Based on the above-mentioned intention analysis and task planning results, the executor module (Executor) of the system can perform operator calls and information integration, and finally output the answer. For example, according to the name and input parameters of the operator, call the operator and input the input parameters into the called operator to obtain the calculation result. Find a suitable multimodal large model according to the intention category, design prompt words according to the input multimodal information, intention category and calculation results of the operator, and input the prompt words into the multimodal large model to process and obtain the output result of the multimodal information. The executor module of the present disclosure can adopt a multimodal large model, or can also adopt an ordinary large model or other models.
[0101] As can be seen from the above processing process, the planning module is a multimodal large model with intention analysis and task planning capabilities. This multimodal large model can be trained based on the Supervised Fine-Tuning (SFT) method. Compared with the large model solution using instruction tuning, it has the advantages of a smaller number of large model parameters and higher accuracy.
[0102] The steps of the SFT method usually include: data engineering, model training, and model evaluation. Among them, data engineering includes the production processes of the training set and the test set, mainly including processes such as data cleaning, large model joint annotation, training set production, data quality evaluation, and data optimization. For the specific process, refer to Figure 10 .
[0103] Based on the above data engineering, a certain number of pseudo-label datasets, such as 100k, can be produced. An example of the format of a training label is as follows:
[0104] Input: base64 format image + user query
[0105] Output:
[0106]
[0107] In the model training part, a certain multimodal large model, such as LLaVA-1.5, can be used as the baseline model for supervised fine-tuning training. When training, all the parameters of the model are turned on, and the training cycle can include multiple, such as 1-3 epochs.
[0108] In the model evaluation part, the accuracy of intent classification and API operator call classification can be measured by precision, recall, etc., and the quality of the thinking process can be evaluated using a certain large language model.
[0109] The system of the embodiments of the present disclosure can be used as an intent analysis and task planning system, which has important planning and divide-and-conquer thinking capabilities for an agent that solves general vision problems. Compared with conventional agents that rely on large models and prompt solutions, this system has the characteristics of a small number of model parameters and higher accuracy. As part of the agent, the embodiments of the present disclosure can reduce the problems of redundant system reasoning and long time consumption. At the same time, because of capabilities such as Chain of Thought (COT), the difficulty of problem-solving is also reduced to a certain extent. In addition, the solution of the embodiments of the present disclosure is general in the field of vision and can be applied to multi-modal large models that support SFT.
[0110] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0111] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0112] Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0113] As Figure 11 shown, the device 1100 includes a computing unit 1101, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1102 or the computer program loaded from the storage unit 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.
[0114] Multiple components in device 1100 are connected to I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0115] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the method for processing multimodal information and / or the method for training a multimodal large model. For example, in some embodiments, the method for processing multimodal information and / or the method for training a multimodal large model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the method for processing multimodal information and / or the method for training a multimodal large model described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the method for processing multimodal information and / or the method for training a multimodal large model in any other suitable manner (e.g., by means of firmware).
[0116] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0117] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.
[0118] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0120] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0121] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.
[0122] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0123] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for processing multimodal information, comprising: Performing intent recognition on the multimodal information to obtain the intent category corresponding to the multimodal information; Obtaining thinking process information according to the multimodal information and its corresponding intent category; Obtaining the operator information to be called according to the thinking process information; wherein, the operator information includes information on the application programs used in the process of processing the multimodal information; Obtaining the output result of the multimodal information according to the intent category, the thinking process information, and the operator information.
2. The method according to claim 1, wherein, Performing intent recognition on the multimodal information to obtain the intent category corresponding to the multimodal information, including: Performing intent recognition on the image and / or text in the multimodal information to obtain the intent category corresponding to the problem in the image and / or text.
3. The method according to claim 1 or 2, wherein Obtaining thinking process information according to the multimodal information and its corresponding intent category, including: Decomposing and analyzing the problem of the multimodal information according to the multimodal information and its corresponding intent category, and describing the category of the application programming interface (API) operator to be called to obtain the thinking process information.
4. The method according to claim 3, wherein Obtaining the operator information to be called according to the thinking process information, including: Obtaining the API analysis result according to the thinking process information; the API analysis result includes the category of the API operator to be called and the input parameters of the API operator.
5. The method according to any one of claims 1 to 4, wherein Obtaining the output result of the multimodal information according to the intent category, the thinking process information, and the operator information, including: Performing operator call, model call, and prompt word setting according to the intent category, the thinking process information, and the operator information to obtain the output result of the multimodal information.
6. A training method for a multimodal large model, comprising: Inputting the multimodal information in the training dataset into the multimodal large model to be trained to obtain an output result; Performing supervised fine-tuning on the multimodal large model to be trained according to the output result and the output label corresponding to the multimodal information in the training dataset to obtain a multimodal large model with a planning function; wherein, the multimodal large model with a planning function is used to execute the planning function of the intelligent agent system.
7. The method according to claim 6, wherein, The output label corresponding to the multimodal information includes one or more of the following: intent category; thinking process information; operator information.
8. A processing device for multimodal information, comprising: An intent recognition module for performing intent recognition on multimodal information to obtain the intent category corresponding to the multimodal information; A thinking module for obtaining thinking process information according to the multimodal information and its corresponding intent category; A calling module for obtaining the operator information to be called according to the thinking process information; wherein, the operator information includes information on the application programs used in the process of processing the multimodal information; An output module for obtaining the output result of the multimodal information according to the intent category, the thinking process information, and the operator information.
9. A training device for a multimodal large model, comprising: An output-input module for inputting multimodal information in a training dataset into a multimodal large model to be trained and obtaining an output result; A supervised fine-tuning module for performing supervised fine-tuning on the multimodal large model to be trained according to the output result and the output label corresponding to the multimodal information in the training dataset, and obtaining a multimodal large model with a planning function; wherein, the multimodal large model with a planning function is used to execute the planning function of an agent system.
10. An agent system, comprising: A planning module for performing intent recognition on multimodal information to obtain an intent category corresponding to the multimodal information; Obtaining thinking process information according to the multimodal information and its corresponding intent category; Obtaining operator information to be called according to the thinking process information; An execution module for obtaining an output result of the multimodal information according to the intent category, the thinking process information, and the operator information.
11. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
13. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.
14. A large model, obtained by training according to the method of claims 6-7, and when executed by a processor, implements the method according to any one of claims 1-5.