Multi-modal task execution method and device and storage medium

By introducing large language models and task feedback mechanisms into multimodal task execution technology, dynamically adjusting the execution strategy, the problems of modal imbalance and long inference time in the existing technology are solved, and efficient execution and personalized services of multimodal tasks are achieved.

CN120045291APending Publication Date: 2025-05-27BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311597053.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing multimodal task allocation and execution technology has problems such as modal imbalance, long inference time, and difficulty in achieving modal alignment and modal fusion.

Method used

Through the large language model (LLM) and task feedback mechanism, the execution strategy of multimodal tasks is dynamically adjusted, the execution process is optimized, and the execution efficiency is improved. The specific method includes determining the execution result of the current single-modal task, continuing to execute the next task in response to the correct result, re-adjusting the execution strategy in response to the wrong result, and ensuring the accuracy of the execution result through comparison of the initial and secondary detection results.

Benefits of technology

It realizes efficient execution of multimodal tasks, optimizes execution processes, improves execution efficiency, and can provide personalized suggestions and services according to user needs to meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045291A_ABST
    Figure CN120045291A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal task execution method and device and a storage medium. The multi-modal task execution method comprises the following steps: determining a single-modal task currently executed in a multi-modal task; and obtaining an execution result of the currently executed single-mode task. And in response to the fact that the execution result is a correct result, continuing to execute the next single-mode task to be executed based on the execution strategy. And in response to the fact that the execution result is an error result, re-determining an execution strategy of the multi-modal task, and executing the multi-modal task according to the re-determined execution strategy. By means of the method and device, the execution process of the multi-modal task can be optimized, and the accuracy of obtaining the execution result of the multi-modal task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of multimodal algorithms, algorithmic reasoning and task allocation, and related artificial intelligence technologies, and in particular to a multimodal task execution method, device, and storage medium. Background Art

[0002] Multimodal task allocation execution technology refers to the technology of assigning tasks of multiple different modalities (such as vision, speech, text, etc.) to different executors (such as different algorithms executing different modalities) for processing. At present, multimodal task allocation execution technology has been widely used in many fields, such as smart home, smart transportation, health care, etc. However, related multimodal task allocation execution technology has problems such as modality imbalance, long reasoning time, and difficulty in achieving modality alignment and modality fusion. Summary of the invention

[0003] In order to overcome the problems existing in the related art, the present disclosure provides a multimodal task execution method, device and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, a modal task execution method is provided, including:

[0005] Determine a currently executed unimodal task in a multimodal task; obtain an execution result of the currently executed unimodal task; in response to the execution result being a correct result, continue to execute the next unimodal task to be executed based on an execution strategy; in response to the execution result being an incorrect result, redetermine the execution strategy of the multimodal task, and execute the multimodal task according to the redetermined execution strategy.

[0006] In one implementation, obtaining the execution result of the currently executed single-modal task includes:

[0007] Obtain output content of the currently executing unimodal task, cache the output content, and based on the output content, initially obtain the execution result of the currently executing unimodal task to obtain an initial detection result; based on the cached output content, obtain the execution result of the unimodal task for a second time to obtain a secondary detection result; use the detection result that is consistent with the initial detection result and the secondary detection result as the execution result of the currently executing unimodal task.

[0008] In one embodiment, the method further comprises:

[0009] In response to the execution result being an erroneous result, the erroneous result is temporarily stored as an intermediate result.

[0010] In one implementation, re-determining the execution strategy of the multimodal task includes:

[0011] Re-executing the unimodal task whose execution result is an error result; executing the multimodal task according to the re-determined execution strategy, including: in response to the number of executions of the unimodal task whose execution result is an error result reaching a threshold, and the execution result is still an error result, caching the error result, and continuing to execute the next unimodal task to be executed until all the multimodal tasks are executed; generating an output result of the multimodal task.

[0012] In one implementation, generating the execution result of the multimodal task includes:

[0013] The execution results of each single-modal task in the multi-modal task are sorted in a manner that correct results are given priority over incorrect results or intermediate results, and the output result of the multi-modal task is obtained according to the output content of the single-modal task corresponding to the sorted execution results.

[0014] In one embodiment, the method further comprises:

[0015] Based on the process of acquiring the output content, execution results, initial detection results and secondary detection results of multiple single-modal tasks and detecting the execution results, execution process information is generated; based on the execution process information and the multi-modal execution results, a log is generated.

[0016] In one implementation, the output content, execution results, primary detection results, and secondary detection results are generated based on a fixed format.

[0017] According to a second aspect of an embodiment of the present disclosure, there is provided a multimodal task execution device, comprising:

[0018] A determination unit is used to determine the currently executed unimodal task in the multimodal task. A detection unit is used to obtain the execution result of the currently executed unimodal task. A processing unit is used to respond to the execution result being a correct result and continue to execute the next unimodal task to be executed based on the execution strategy. The processing unit is also used to: respond to the execution result being an incorrect result, re-determine the execution strategy of the multimodal task, and execute the multimodal task according to the re-determined execution strategy.

[0019] In one implementation, the detection unit obtains the execution result of the currently executed single-modal task in the following manner:

[0020] Obtain output content of the currently executing unimodal task, cache the output content, and based on the output content, initially obtain the execution result of the currently executing unimodal task to obtain an initial detection result; based on the cached output content, obtain the execution result of the unimodal task for a second time to obtain a secondary detection result; use the detection result that is consistent with the initial detection result and the secondary detection result as the execution result of the currently executing unimodal task.

[0021] In one implementation, the processing unit is further configured to:

[0022] In response to the execution result being an erroneous result, the erroneous result is temporarily stored as an intermediate result.

[0023] In one implementation, the processing unit redetermines the execution strategy of the multimodal task in the following manner:

[0024] Re-executing the unimodal task whose execution result is an error result; executing the multimodal task according to the re-determined execution strategy, including: in response to the number of executions of the unimodal task whose execution result is an error result reaching a threshold, and the execution result is still an error result, caching the error result, and continuing to execute the next unimodal task to be executed until all the multimodal tasks are executed; generating an output result of the multimodal task.

[0025] In one implementation, the processing unit generates the execution result of the multimodal task in the following manner:

[0026] The execution results of each single-modal task in the multimodal task are sorted in a manner that correct results are given priority over incorrect results or intermediate results, and the execution results of the multimodal task are obtained according to the output content of the single-modal task corresponding to the sorted execution results.

[0027] In one embodiment, the detection unit is further used for:

[0028] Based on the execution results, the initial detection results and the secondary detection results of the multiple single-modal tasks, the execution process information is generated; based on the execution process information and the multi-modal execution results, a log is generated.

[0029] In one implementation, the execution result, the primary detection result, and the secondary detection result are generated based on a fixed format.

[0030] According to a third aspect of an embodiment of the present disclosure, there is provided a multimodal task execution device, comprising:

[0031] A processor; a memory for storing processor executable instructions; wherein the processor is configured to: execute the multimodal task execution method described in the first aspect or any one of the embodiments of the first aspect.

[0032] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided, in which instructions are stored. When the instructions in the storage medium are executed by a processor of a terminal, the terminal can perform the method described in the first aspect or any one of the implementations of the first aspect.

[0033] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: achieving efficient execution of multimodal tasks, optimizing the execution process of multimodal tasks, improving the execution efficiency of multimodal tasks, and providing personalized suggestions and services according to user needs, thereby better meeting user needs.

[0034] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0036] Figure 1 The figure is a flowchart of a multimodal task execution method according to an exemplary embodiment.

[0037] Figure 2 The figure is a flowchart of a method for obtaining a single-modal task execution result according to an exemplary embodiment.

[0038] Figure 3 The figure is a flowchart of a method for setting an intermediate result according to an exemplary embodiment.

[0039] Figure 4 The figure is a flowchart of a method for determining an execution strategy according to an exemplary embodiment.

[0040] Figure 5 The figure is a flowchart of a method for determining output content of a multimodal task according to an exemplary embodiment.

[0041] Figure 6 The figure is a flowchart of a method for generating a multimodal task log according to an exemplary embodiment.

[0042] Figure 7 is a schematic diagram of multimodal task execution according to an exemplary embodiment.

[0043] Figure 8It is a block diagram of a multimodal task execution device according to an exemplary embodiment.

[0044] Fig. 9 It is a block diagram of a device for executing a multimodal task according to an exemplary embodiment. DETAILED DESCRIPTION

[0045] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure.

[0046] In the accompanying drawings, the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limitations on the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present disclosure. The embodiments of the present disclosure are described in detail below in conjunction with the accompanying drawings.

[0047] The multimodal task execution method provided by the embodiment of the present disclosure can be applied to application scenarios that require efficient allocation and execution of multimodal tasks, such as autonomous driving: In an autonomous vehicle, multiple sensors (such as cameras, radars, lidars, etc.) can capture different information at the same time. By using multimodal task allocation reasoning rules, this information can be integrated to more accurately understand the surrounding environment and better control the vehicle. Smart home: In a smart home, multiple sensors (such as temperature sensors, humidity sensors, light sensors, etc.) can capture different information at the same time. By using multimodal task allocation reasoning rules, this information can be integrated to better control home appliances, such as automatically adjusting temperature, lights, etc. Speech recognition: In speech recognition, multiple modalities (such as sound, mouth shape, language intonation, etc.) can capture different information at the same time. By using multimodal task allocation reasoning rules, this information can be integrated to more accurately recognize speech. The rules of AI multimodal task allocation reasoning also have many application scenarios on mobile phones and tablets, some of which include: Smart assistants: Smart assistants can use multiple modalities (such as voice, text, images, etc.) to understand user requests and provide corresponding responses. By using multimodal task assignment reasoning rules, smart assistants can understand users' requests more accurately and provide better responses. Virtual Reality: Virtual reality applications can use multiple modalities (such as vision, hearing, touch, etc.) to provide a more realistic experience. By using multimodal task assignment reasoning rules, virtual reality applications can better integrate these modalities to provide a better experience. Smart Album: Smart Album applications can use multiple modalities (such as images, audio, location information, etc.) to organize and manage users' photos. By using multimodal task assignment reasoning rules, smart album applications can better understand users' needs and provide better organization and management functions. Health Monitoring: Health monitoring applications can use multiple modalities (such as heart rate, number of steps, sleep quality, etc.) to monitor users' health. By using multimodal task assignment reasoning rules, health monitoring applications can more accurately monitor users' health and provide better suggestions and guidance. By allocating processing based on multimodal tasks and feedback mechanisms, multimodal tasks can be performed efficiently and accurately.

[0048] In the related technology, multimodal task allocation execution technology refers to the technology of assigning tasks of multiple different modes (such as vision, voice, text, etc.) to different executors (such as different algorithms executing different modes) for processing. At present, multimodal task allocation execution technology has been widely used in many fields, such as smart home, smart transportation, health care, etc. At present, there are mainly the following rule technologies for AI multimodal task allocation reasoning: Rule technology based on expert system: This technology is to define rules through expert system, and then assign tasks and executors according to the rules. The advantage of this technology is that it can quickly realize task allocation, but it requires experts to define rules and is difficult to adapt to complex scenarios. Rule technology based on knowledge graph: This technology is to define rules through knowledge graph, and then assign tasks and executors according to the rules. The advantage of this technology is that it can adapt to complex scenarios, but it requires a large amount of knowledge graph data and computing resources. Rule technology based on logical reasoning: This technology is to define rules through logical reasoning, and then assign tasks and executors according to the rules. The advantage of this technology is that it can adapt to complex scenarios, but it requires more computing resources and longer reasoning time. Therefore, the rule technology of AI multimodal task allocation reasoning has different applications in different fields, and it is necessary to select appropriate technology according to specific scenarios.

[0049] However, related technologies suffer from problems such as modal imbalance, modal time-space misalignment, modal fusion errors, and poor real-time performance. For example, there may be imbalances in the amount of training data, task processing capabilities, controllability, etc. of different modalities, resulting in poor performance of the model in certain modalities; data of different modalities may be misaligned in time and space, making it difficult for the model to integrate these data; data of different modalities may have different characteristics and representations, and therefore are difficult to integrate; multimodal task allocation reasoning requires processing a large amount of data, which may result in a long reasoning time, thus affecting real-time performance.

[0050] In view of this, a multimodal task execution method is provided in an embodiment of the present disclosure. In the multimodal task execution method provided by the present disclosure, based on a large language model (LLM) and a task feedback mechanism, multimodal tasks are allocated and processed based on feedback information, the execution process of multimodal tasks is optimized, and the execution efficiency of multimodal tasks is improved.

[0051] Figure 1 is a flowchart of a multimodal task execution method according to an exemplary embodiment. Figure 1 As shown, the following steps are included.

[0052] In step S11 , the currently executed unimodal task in the multimodal task is determined.

[0053] In the disclosed embodiments, a multimodal task refers to a task involving data in multiple modalities (such as vision, voice, text, etc.). It should be understood that the content of the multimodal task itself can be a single modality. For example, a user issues a multimodal task instruction of "generate a video of a child singing", but the instruction itself only has one modality of natural language. Based on the instruction, multimodal tasks such as drawing multiple pictures of children singing, combining multiple pictures into a video, and generating audio of children singing are obtained.

[0054] In the disclosed embodiment, initialization input parameters may also be set for each single-modal task, which is beneficial for the smooth execution of multi-modal tasks.

[0055] In step S12, the execution result of the currently executed single-modal task is obtained.

[0056] In the disclosed embodiment, in response to the completion of the execution of a single-modal task, the execution result of the single-modal task is obtained, and the execution result is used to indicate whether the corresponding single-modal task is executed correctly. By obtaining the execution result of each single-modal task during the execution of the multimodal task, it is determined whether the execution result of each single-modal task is correct, and it is convenient to perform corresponding operations based on the execution result.

[0057] In step S13a, in response to the execution result being a correct result, the next unimodal task to be executed continues to be executed based on the execution strategy.

[0058] In the disclosed embodiment, if the execution result is judged to be a correct execution result, it means that the currently executed unimodal task has output the execution result normally, and the next unimodal task is continued to be executed according to the original execution strategy.

[0059] In the disclosed embodiment, the execution strategy is used to plan the execution process of a single modal task in a multimodal task.

[0060] In the disclosed embodiment, LLM can be used to plan the execution strategy of the unimodal tasks included in the multimodal tasks, set the execution order of the unimodal tasks in the multimodal tasks, and realize the control of global reasoning planning. For example, the planning process is started by obtaining input from multiple types of information: different types of information will form instruction information by combining them, and the fixed format is set as: [[text: picture][historical information]] composed of instruction prompts and other visual and audio input summaries created by auxiliary inspections. Then, LLM will generate appropriate output prompts for the next step of execution. The prompts can be composed of two parts: Task execution steps: describe what should be done next in text language. Although this text language does not directly affect the call of the module or external API, it helps the LLM planning process and has a prompting effect on the feedback survey; Task execution operation: Generate a fixed format structure string prompt (based on a predefined template instruction template), the main purpose of which is to specify which model or external tool to call and what parameters to input. For example, a unimodal task can be: ["What does this picture describe", picture [1]].

[0061] In step S13b, in response to the execution result being an erroneous result, the execution strategy of the multimodal task is re-determined, and the multimodal task is executed according to the re-determined execution strategy.

[0062] In the disclosed embodiment, if the execution result is judged to be an erroneous result, the execution strategy of the multimodal task is re-determined, and the execution process is optimized in real time by dynamically adjusting the execution strategy of the multimodal task, thereby reducing the occurrence of erroneous execution result output. It should be understood that the detection of the execution result of a single-modal task does not interrupt the reasoning execution process of the entire multimodal task, but adjusts the execution strategy in real time based on the execution result of each single-modal task.

[0063] In the disclosed embodiment, the output content of the currently executed single-modal task may be tested twice, and the execution result may be obtained by comparing the results of the two tests.

[0064] Figure 2 is a flowchart of a method for obtaining a single-mode task execution result according to an exemplary embodiment. Figure 2 As shown, the following steps are included.

[0065] In step S21, the output content of the currently executed unimodal task is obtained, the output content is cached, and based on the output content, the output content of the currently executed unimodal task is obtained for the first time to obtain an initial detection result.

[0066] In the disclosed embodiment, the output content of the currently executed unimodal task is checked to obtain the corresponding initial detection result, which is used to indicate whether the corresponding output content is normal content or an error code. If the output content is an execution problem or a running error that occurs when the current unimodal task is executed, the corresponding initial detection result is an error result. For example, when an error occurs during the processing of the i-th unimodal task in the multimodal task, the returned content is a string of error information, that is, the output content of the i-th unimodal task is a string of error information, then the corresponding i-th initial detection result is wrong.

[0067] In the disclosed embodiment, the output content of the current unimodal task can be detected based on the output content of multiple completed unimodal tasks, and the initial detection result can be obtained. For example, after obtaining the output content of the 10th unimodal task, if it is detected that the output content is too different from the output content of the first 9 unimodal tasks, the execution result of the unimodal task is set to an error result, and the initial detection result is obtained.

[0068] In the disclosed embodiment, after setting the execution strategy and initializing the input parameters, it is determined that the relevant model or external interface is called, and after each call to the external tool, the output content of the corresponding modality will be generated, and the non-text modal results will be returned in the form of natural language. It should be understood that the generated output result can be used as an intermediate result of the modal output, such as an intermediate video editing result / image generation result, etc. The intermediate result of the modal output can be temporarily stored, and other single-modal tasks can generate the output content of the current single-modal task modality based on the intermediate result of the modal output and the current single-modal task, so that the output content has good context relevance.

[0069] In step S22, the execution result of the single-modal task is acquired for a second time based on the cached output content to obtain a secondary detection result.

[0070] In the disclosed embodiment, by performing a secondary detection on the cached output content, the output result content is detected based on the first detection result. Compared with obtaining the execution result of the single-modal task for the first time, the secondary detection result obtained by the secondary detection is more reliable.

[0071] In step S23, the detection result that is consistent with the initial detection result and the secondary detection result is used as the execution result of the currently executed single-modal task.

[0072] In the disclosed embodiment, by detecting whether the initial detection result and the secondary detection result are consistent, if they are consistent, the correct detection result can be determined based on the initial detection result and the secondary detection result, and set as the execution result of the currently executed single-modal task.

[0073] In the embodiments of the present disclosure, if it is detected that the initial detection result and the secondary detection result are inconsistent, a corresponding action can be set based on the needs. In an exemplary embodiment, it can be set that when the initial detection result and the secondary detection result are inconsistent, the secondary detection result is determined as the detection result. In another exemplary embodiment, it can be set that when the initial detection result and the secondary detection result are inconsistent, the current single-modal task is re-executed and the initial detection and the secondary detection are performed again.

[0074] In the embodiment of the present disclosure, the execution result determined based on the initial detection result and the secondary detection result includes two situations: a correct result and an incorrect result.

[0075] Figure 3 is a flowchart of a method for setting an intermediate result according to an exemplary embodiment. Figure 3 As shown, the following steps are included.

[0076] because Figure 3 The steps S31, S32 and S33 in Figure 2 Steps S21, S22 and S23 are the same and will not be described in detail here. Please refer to the relevant description of the above embodiment.

[0077] In step S34, in response to the execution result being an erroneous result, the erroneous result is temporarily stored as an intermediate result.

[0078] In the disclosed embodiment, the execution result includes two kinds of contents: a correct result and an error result. When the execution result is an error result, it indicates that the output content is an abnormal output such as error information of a single-modal task, and cannot be output as the output content of a single-modal task. Therefore, it needs to be stored as an intermediate result to avoid the error information being output as the output content, which affects the output result of the multi-modal task.

[0079] In the disclosed embodiment, if the execution result is an error result, it means that the current unimodal task is executed incorrectly and generates incorrect output content. Therefore, the execution strategy needs to be adjusted to ensure that each unimodal task can obtain correct output content.

[0080] Figure 4 is a flowchart of a method for determining an execution strategy according to an exemplary embodiment. Figure 4 As shown, the following steps are included.

[0081] In step S41, the single-mode task whose execution result is an error result is re-executed.

[0082] In the disclosed embodiment, if it is detected that the execution result of a unimodal task being executed in a multimodal task is an erroneous result, the unimodal task is re-run to ensure that each unimodal task in the multimodal task can output correct output content as much as possible.

[0083] In step S42, in response to the number of times the unimodal task whose execution result is an error result is re-executed reaching a threshold, and the execution result is still an error result, the error result is cached, and the next unimodal task to be executed continues to be executed until all multimodal tasks are executed.

[0084] In the disclosed embodiment, when it is detected that the number of times a unimodal task with an erroneous execution result is re-executed reaches a threshold and still the correct execution result cannot be output, the erroneous result is cached and the current unimodal task is no longer executed. Other unimodal tasks to be executed continue to be executed based on the execution strategy, thereby avoiding the phenomenon that multimodal tasks are blocked due to the inability to output the output content of a unimodal task.

[0085] In step S43, the output result of the multimodal task is generated.

[0086] In the disclosed embodiment, the output result of the multimodal task may also be multimodal. The context perception capability can be improved based on the multimodal task, and the generation of multimodal output results can increase the content richness of the restored information.

[0087] Figure 5 is a flowchart of a method for determining output content of a multimodal task according to an exemplary embodiment. Figure 5 As shown, the following steps are included.

[0088] because Figure 5 The steps S51 and S52 in Figure 4 Steps S41 and S42 are the same and will not be described in detail here. Please refer to the relevant description of the above embodiment.

[0089] In step S53, the execution results of each unimodal task in the multimodal task are sorted in a manner that correct results take precedence over erroneous results or intermediate results, and the output result of the multimodal task is obtained according to the output content of the unimodal task corresponding to the sorted execution results.

[0090] In the disclosed embodiment, for the execution results and output contents of all unimodal tasks, the execution results of the unimodal tasks are comprehensively sorted according to the execution strategy of LLM, and finally a reply is given according to the execution strategy to generate an answer description of the multimodal task and output it, so as to optimize the effect of the output results of the multimodal task.

[0091] In the disclosed embodiment, the task execution process of the entire multimodal task can be recorded, and learning can be performed based on the current execution process to optimize the execution strategy.

[0092] Figure 6 is a flowchart of a method for generating a multimodal task log according to an exemplary embodiment. Figure 6 As shown, the following steps are included.

[0093] In step S61, execution process information is generated based on obtaining output contents, execution results, primary detection results, and secondary detection results of multiple single-modal tasks.

[0094] In the disclosed embodiments, by recording information generated during the execution of multiple single-modal information included in a multi-modal task, execution process information is generated, which can facilitate operations such as performance analysis, debugging, troubleshooting, maintenance, and updating.

[0095] In step S62, a log is generated based on the execution process information and the multimodal execution result.

[0096] In the disclosed embodiments, the execution process of the multimodal task is recorded, an execution log is generated, and learning is performed based on the execution log to optimize the processing flow of the multimodal task, thereby better laying the foundation for the subsequent reasoning process. The execution log may include the execution results of the modal task, the corresponding detection results and secondary detection results, etc.

[0097] In the disclosed embodiment, the output content, execution result, primary detection result and secondary detection result can be generated based on a fixed format. It should be understood that all data streams in the multimodal task execution process can be formatted.

[0098] In an exemplary embodiment, the single-mode task can be preprocessed, and the execution strategy and initialization input parameters can be preprocessed to achieve standardization of the input content, facilitate processing of the input content, and improve execution efficiency. For example, when establishing multi-task execution, various modules or APIs are standardized into a unified interface by using a code style template: [[module name][instruction prompt]["related data flow"][output cache]], where the module name is the name of the module, the instruction prompt is the instruction or operation step required to call the module, the related data flow refers to the data flow transmitted between modules, and the output cache refers to the place where the module output results are temporarily stored. In addition, other required parameters or configuration values ​​can be included in the template according to requirements.

[0099] In another exemplary embodiment, the output results can be post-processed and output in a unified format to improve the subsequent processing efficiency. For example, the processing results of each modal task are simplified into a code format: [[task name][status][result][label]], wherein non-text result content is labeled when output, which improves data processing efficiency, facilitates visual analysis and insight, error detection and correction, and facilitates understanding of data.

[0100] In the disclosed embodiments, a multimodal initialization method is established, such as supporting input of different modalities. However, in order to provide a unified input format for LLM and more effectively understand and utilize data derived from complex modalities (such as images, videos, audio, and text), basic text descriptions can be generated by adopting various basic open source models.

[0101] Based on the same concept, an embodiment of the present disclosure also provides a multimodal task execution device.

[0102] It is understandable that in order to realize the above functions, the multimodal task execution device provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function. In combination with the units and algorithm steps of each example disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.

[0103] Figure 7 is a schematic diagram of multimodal task execution according to an exemplary embodiment.

[0104] In this exemplary embodiment, A-planning group, B-execution group, C-feedback inspection group, D-summary output group, and E-feedback learning group are only used to represent different execution processes of multimodal tasks, not to represent specific physical structures. Among them, A-planning group is mainly used to plan the execution process and methods of different modes, and use natural language to plan what the B-execution group should do next according to the current reasoning progress. B-execution group mainly includes the execution process of different modes, and can temporarily store the output content of single-modal tasks. C-feedback inspection group performs cache management on the current task execution process, and at the same time coordinates between different execution groups to establish connections between multiple modes. D-summary output group summarizes the results of different modes and outputs the final result. E-feedback learning group enables the model to autonomously explore and discover the optimal solution after each task is executed and the output is generated, and establishes memory ability to lay the foundation for subsequent task execution.

[0105] It should be understood that in this exemplary embodiment, the execution result of the single-modal task can also be called the reasoning progress, and the execution result can also be called the reasoning script.

[0106] The A-planning group thinks about what needs to be done next based on the current reasoning progress and calls external models. The innovation of this technical solution is that it allows the use of reasoning scripts to call external models. The B-execution group encapsulates the external toolkit into a special input and output format, allowing the use of structured commands to call related models. The output results of other models can be directly cached in the C-feedback check group and passed to subsequent models. The intermediate results obtained in multimodal tasks are usually segmented representations of image or text information at a certain point in time. In complex situations, it is difficult for the feedback check group to manage which information should be retained and which information needs to be sent to the next module, so the C-feedback check group manages the input and intermediate result caches during the reasoning process of other models, checks whether the prediction process is reasonable or judges the correctness of the prediction results based on annotations, and provides all currently available metadata, intermediate results and reasoning results for the subsequent D-summary output group, and generates the final output results. The combination of ABC enables the model to efficiently implement complex reasoning. In order to provide a process basis for the next task allocation, this technical solution sets up an E-feedback learning group to learn the calling process of this reasoning process, that is, to optimize the cache and establish a multi-task allocation path.

[0107] In this technical solution, the preset multimodal processing capabilities include video character detection, image detection, automatic subtitles, narration, OCR module, speech recognition, speech-to-text, speech separation, image segmentation, image generation and other modes, and allow the call of external models. By combining these functions, multimodal tasks can accomplish a wide range of functions.

[0108] In this technical solution, the A-planning group can use open source LLM as a control system to control global reasoning planning. It starts the planning process by obtaining input from 4 (or more) types of information: different types of information will be combined to form instruction information, and the fixed format is set as: [[text: picture][historical information]] composed of instruction prompts and other visual and audio input summaries created by auxiliary inspections. Then, LLM will generate appropriate output prompts for the next step of execution. The prompt consists of two parts: Main idea: Use text language to describe what should be done next. Although this "text" does not directly affect the call of the module or external API, it helps the LLM planning process and has a prompt effect on the feedback survey; Execution operation: Generate a fixed format structure string prompt (based on a predefined template instruction template), the main purpose of which is to specify which model or external tool to call and what parameters to input. For example, ["What does this picture describe", picture [1]].

[0109] This exemplary embodiment supports input of different modalities by establishing a multimodal initialization method as mentioned above, but in order to provide a unified input format for LLM and more effectively understand and utilize data derived from complex modalities (such as images, videos, audio, and text), basic text descriptions are generated by adopting various basic open source models.

[0110] Through the call plan set by A and the initialization input parameters, the B-execution group will call the relevant model or external interface. After each call to the external tool, the output result of the corresponding modality will be generated. At the same time, the non-text modality result will be returned in the form of natural language through the same tool as A. If the model generates intermediate results, such as intermediate video editing results / image generation results, etc., it will be passed to the C-feedback inspection group to store it and generate new prompt information for it. At the same time, the execution results will also be passed to the D-summary output group for temporary storage, waiting for the completion of other tasks to output the comprehensive execution results. Generally speaking, the important reason for improving execution efficiency lies in standard data preprocessing. Here, when establishing multi-task execution, various modules or APIs are standardized into a unified interface by using code style templates: [[module name][instruction prompt]["related data flow"][output cache][…]]. Each module input information is designed to accept multiple content prompts as input, and of course other input content needs to be customized according to different tasks. Finally, the task is executed according to the content provided by the input interface and the final result is obtained.

[0111] The purpose of the C-Feedback Check Group is to manage the execution results and intermediate results of different tasks, and to coordinate the timing issues of different tasks during execution. And decide which result should be directed to which module. During the execution of a task, if the execution result is wrong or the code contains errors during execution, the entire reasoning process will not be interrupted. The error information will be output as an intermediate result to the summary output group and the planning group, so that they can re-optimize the planning process in real time. If there is still an error in the subsequent execution, the result sent to the summary output group will be used as the execution output result of the modality, and the execution problem and error will be fed back to the user. The feedback detection group needs to be post-processed before being sent to other groups, and it will be packaged into a unified format (in order to improve the processing efficiency of subsequent groups). First, the "results (correct results, or errors)" of all current modalities are simplified into code format: [[task name][status][result][…][label]], and secondly, the information other than the text results is labeled, and the label information is bound to the text information to facilitate the subsequent alignment of "task: result".

[0112] The D - summary output group processes the results of the previous execution group and the temporary results of the feedback detection group. The summary output is based on three points: 1. According to the "suggestions" of the plan group, the results of different modalities are comprehensively sorted; 2. According to the results of the feedback detection group, the correct results and complete results are given priority and placed in the front, while the problematic results or intermediate results to be output are placed behind; 3. According to the instruction prompt order in the plan group (e.g., what does the picture describe, generate a piece of music based on the picture description), a reasonable answer description is generated, mainly to establish normal interaction content. The integrated results are then output. Additionally, when the temporary results are not satisfactory, the system will repeatedly attempt to provide an answer until the correct answer is given (when the real - world situation is available) or the predefined maximum number of attempts is reached.

[0113] In processing multi - modal tasks, it is still easy to encounter various problems and errors. Therefore, establishing a self - learning mechanism can better lay the foundation for subsequent reasoning processes (such as which tasks are simple and can be quickly reasoned, which tasks are error - prone, which tasks fail to be called, etc.). By collecting the execution logs (log) of the current task, including successful prediction results, failed prediction results, etc. By storing the "experience" in the plan group, this historical information can be referred to during subsequent execution processes. Of course, during the model reasoning process, after N explorations, if the model gives the correct answer on the first attempt, it indicates that the current problem can be effectively solved well, and such experience needs to be accumulated; but if the model gets the correct answer after n attempts (1 < n ≤ N), it means that there is still room for improvement in the model's planning ability, and such experience logs need to be accumulated. Therefore, when the model encounters a similar task assignment in the future, the experience can be used as a context example prompt for the plan group to plan tasks.

[0114] Through the mutually restrictive operations of ABCDE, the reasoning effect of multi - modal task allocation can be effectively improved, and the efficiency of multi - modal work can also be greatly enhanced.

[0115] In the embodiments of the present disclosure, by detecting the output content of multi - modal tasks and obtaining the execution results, and adaptively adjusting the execution strategy based on the execution result content, it is possible to prevent the execution results of multi - modal tasks from being affected by the incorrect output content of a single - modal task. And by executing multiple single - modal tasks based on the execution strategy, the efficient execution of multi - modal tasks is achieved, improving the execution efficiency under multiple tasks, which is beneficial to improving the user experience. And a comprehensive intelligent assistant can be established through the multi - modal task execution method of the embodiments of the present disclosure to provide a more natural and intuitive human - machine interaction.

[0116] Figure 8 It is a block diagram of a multi - modal task execution device 100 shown according to an exemplary embodiment. Refer to Figure 8The device includes a determination unit 101, a detection unit 102 and a processing unit 103.

[0117] A determination unit 101, used to determine a currently executed single-modal task in a multi-modal task;

[0118] The detection unit 102 is used to obtain the execution result of the currently executed single-modal task;

[0119] The processing unit 103 is used for continuing to execute the next single-modal task to be executed based on the execution strategy in response to the execution result being a correct result;

[0120] In one embodiment, the processing unit 103 is further configured to: in response to the execution result being an erroneous result, redetermine the execution strategy of the multimodal task, and execute the multimodal task according to the redetermined execution strategy.

[0121] In one embodiment, the detection unit 102 obtains the execution result of the currently executing unimodal task in the following manner: obtains the output content of the currently executing unimodal task, caches the output content, and based on the output content, obtains the output content of the currently executing unimodal task for the first time to obtain an initial detection result; based on the cached output content, obtains the output content of the unimodal task for a second time to obtain a secondary detection result; and uses the detection result that is consistent with the initial detection result and the secondary detection result as the execution result of the currently executing unimodal task.

[0122] In one embodiment, the processing unit 103 is further configured to: in response to the execution result being an erroneous result, temporarily store the erroneous result as an intermediate result.

[0123] In one embodiment, the processing unit 103 redetermines the execution strategy of the multimodal task in the following manner: re-executing a unimodal task whose execution result is an error result; executing the multimodal task according to the re-determined execution strategy, including: in response to the number of executions of the unimodal task whose execution result is an error result reaching a threshold, and the execution result is still an error result, caching the error result, and continuing to execute the next unimodal task to be executed until all unimodal tasks are executed; generating the output result of the multimodal task.

[0124] In one embodiment, the processing unit 103 generates the execution result of the multimodal task in the following manner:

[0125] The execution results of each single-modal task in the multi-modal task are sorted in a manner that correct results take precedence over incorrect results or intermediate results, and the execution result of the multi-modal task is obtained according to the output content of the single-modal task corresponding to the sorted execution result.

[0126] In one embodiment, the detection unit 102 is further configured to:

[0127] Based on obtaining the output content, execution results, initial detection results and secondary detection results of multiple single-modal tasks, the execution process information is generated; based on the execution process information and the multi-modal execution results, a log is generated.

[0128] In one embodiment, the output content, execution results, primary detection results, and secondary detection results are generated based on a fixed format.

[0129] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0130] Fig. 9 2 is a block diagram of an apparatus 200 for multimodal task execution according to an exemplary embodiment. For example, the apparatus 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0131] Reference Fig. 9 , the device 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .

[0132] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.

[0133] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0134] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.

[0135] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0136] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC), and when the device 200 is in an operation mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 204 or sent via the communication component 216. In some embodiments, the audio component 210 also includes a speaker for outputting audio signals.

[0137] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0138] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200, the sensor assembly 214 can also detect the position change of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200 and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 214 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor or a temperature sensor.

[0139] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0140] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0141] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, and the instructions can be executed by the processor 220 of the device 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0142] It is to be understood that in the present disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include plural forms, unless the context clearly indicates other meanings.

[0143] It is further understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a specific order or degree of importance. In fact, the expressions "first", "second", etc. can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.

[0144] It can be further understood that, unless otherwise specified, “connection” includes a direct connection without other components between the two, and also includes an indirect connection with other components between the two.

[0145] It is further understood that, although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring the execution of all the operations shown to obtain the desired results. In certain environments, multitasking and parallel processing may be advantageous.

[0146] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modifications, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure.

[0147] It should be understood that the present disclosure is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the scope of the appended claims.

Claims

1. A multimodal task execution method, It is characterized in that include: Determine the unimodal task currently being performed in the multimodal task; Obtaining the execution result of the currently executed single-mode task; In response to the execution result being a correct result, continuing to execute the next single-modal task to be executed based on the execution strategy; In response to the execution result being an erroneous result, the execution strategy of the multimodal task is re-determined, and the multimodal task is executed according to the re-determined execution strategy.

2. The multimodal task execution method according to claim 1, It is characterized in that Obtaining the execution result of the currently executed single-mode task, including: Obtaining output content of the currently executed single-modal task, caching the output content, and obtaining an execution result of the currently executed single-modal task for the first time based on the output content to obtain an initial detection result; Re-acquiring the execution result of the single-modal task based on the cached output content to obtain a secondary detection result; The detection result that is consistent between the initial detection result and the secondary detection result is used as the execution result of the currently executed single-modal task.

3. The multimodal task execution method according to claim 1 or 2, It is characterized in that The method further comprises: In response to the execution result being an erroneous result, the erroneous result is temporarily stored as an intermediate result.

4. The multimodal task execution method according to claim 1, It is characterized in that The re-determining the execution strategy of the multimodal task includes: Re-executing the single-mode task whose execution result is an error result; The executing the multimodal task according to the re-determined execution strategy includes: In response to the number of times the single-modal task whose execution result is an error result is re-executed reaching a threshold, and the execution result is still an error result, caching the error result, and continuing to execute the next single-modal task to be executed until all the multi-modal tasks are executed; Generate output results for multimodal tasks.

5. The multimodal task execution method according to claim 3 or 4, It is characterized in that The output result of generating the multimodal task includes: The execution results of each single-modal task in the multimodal task are sorted in a manner that correct results are given priority over incorrect results or intermediate results, and the output result of the multimodal task is obtained according to the output content of the single-modal task corresponding to the sorted execution results.

6. The multimodal task execution method according to claim 1, It is characterized in that The method further comprises: Generate execution process information based on the output contents, execution results, initial detection results and secondary detection results of the plurality of single-modal tasks; A log is generated based on the execution process information and the multimodal execution result.

7. The multimodal task execution method according to claim 2, It is characterized in that The output content, execution results, primary detection results and secondary detection results are generated based on a fixed format.

8. A multi-modal task execution device, It is characterized in that include: A determination unit, used to determine the currently executed single-modal task in the multi-modal task; A detection unit, used to obtain the execution result of the currently executed single-modal task; A processing unit, configured to, in response to the execution result being a correct result, continue to execute the next single-modal task to be executed based on the execution strategy; The processing unit is also used for: In response to the execution result being an erroneous result, the execution strategy of the multimodal task is re-determined, and the multimodal task is executed according to the re-determined execution strategy.

9. The multimodal task execution device according to claim 8, It is characterized in that The detection unit obtains the execution result of the currently executed single-modal task in the following manner: Acquire the output content of the currently executed single-modal task, cache the output content, and acquire the output content of the currently executed single-modal task for the first time based on the output content to obtain an initial detection result; Re-acquiring the output content of the single-modal task based on the cached output content to obtain a secondary detection result; The detection result that is consistent between the initial detection result and the secondary detection result is used as the execution result of the currently executed single-modal task.

10. The multimodal task execution device according to claim 8 or 9, It is characterized in that The processing unit is also used for: In response to the execution result being an erroneous result, the erroneous result is temporarily stored as an intermediate result.

11. The multimodal task execution device according to claim 8, It is characterized in that The processing unit re-determines the execution strategy of the multimodal task in the following manner: Re-executing the single-mode task whose execution result is an error result; The executing the multimodal task according to the re-determined execution strategy includes: In response to the number of times the single-modal task whose execution result is an error result is re-executed reaching a threshold, and the execution result is still an error result, the error result is cached, and the next single-modal task to be executed is continued to be executed until all the single-modal tasks are executed; Generate output results for multimodal tasks.

12. The multimodal task execution device according to claim 10 or 11, It is characterized in that The processing unit generates the execution result of the multimodal task in the following manner: The execution results of each single-modal task in the multimodal task are sorted in a manner that correct results are given priority over incorrect results or intermediate results, and the output result of the multimodal task is obtained according to the output content of the single-modal task corresponding to the sorted execution results.

13. The multimodal task execution device according to claim 8, It is characterized in that The detection unit is also used for: Generate execution process information based on the output contents, execution results, initial detection results and secondary detection results of the plurality of single-modal tasks; A log is generated based on the execution process information and the multimodal execution result.

14. The multimodal task execution device according to claim 9, It is characterized in that The output content, execution results, primary detection results and secondary detection results are generated based on a fixed format.

15. A multi-modal task execution device, It is characterized in that include: processor; a memory for storing processor-executable instructions; The processor is configured to: execute the method described in any one of claims 1 to 7.

16. A storage medium, It is characterized in that The storage medium stores instructions, and when the instructions in the storage medium are executed by a processor of the terminal, the terminal is enabled to execute the method described in any one of claims 1 to 7.