Multi-modal large model reasoning system, reasoning method and application

By introducing distributed deployment and load balancing modules into the multimodal large-modal large-modal inference system, the problems of large-scale parameters and high computing requirements in the existing technology are solved, multimodal load balancing and efficient inference are realized, and hardware resource utilization and multimodal data processing effects are optimized.

CN119990301APending Publication Date: 2025-05-13SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410617684.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When existing multimodal large models process multimodal data, the large amount of parameters and high computing requirements lead to insufficient utilization of hardware resources, data bandwidth becomes a bottleneck, and a single model has poor effect on processing multimodal data.

Method used

A multimodal large model inference system is proposed to realize multimodal load balancing and efficient inference through distributed deployment and load balancing modules. The system includes input module, input preprocessing module, load balancing module, parallel computing module, evaluation feedback module and output module, and supports a variety of models and input and output forms, and uses evaluation feedback module to perform iterative optimization of results.

Benefits of technology

It realizes load balancing and efficient inference of multiple models on large systems, optimizes hardware resource utilization, reduces data bandwidth requirements, and improves the effect of multimodal data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0004845779770000011
    Figure HDA0004845779770000011
Patent Text Reader

Abstract

The invention discloses a multi-modal large model reasoning system which comprises an input module, an input preprocessing module, a load balancing module, a parallel computing module, an evaluation feedback module and an output module. The input module is used for receiving external input information; the input preprocessing module is used for preprocessing input information in the input module and performing iteration according to feedback given by the evaluation feedback module; the load balancing module is used for monitoring and scheduling system hardware and distributing system resources; the parallel computing module is used for reasoning the task to generate a reasoning result; the evaluation feedback module is used for receiving the reasoning result generated by the parallel computing module and performing verification feedback and result evaluation on the reasoning result; and the output module is used for outputting one or a combination of more of characters, pictures, voices, videos and codes. The invention further discloses a multi-modal large model reasoning method and application, and the multi-modal large model reasoning method has a wide application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large models and distributed deployment, and relates to a multimodal large model reasoning system, a reasoning method and an application. Background Art

[0002] The input and output forms supported by the current large model reasoning framework are becoming more and more diverse. Many companies will launch multimodal large models to connect with various user inputs. However, the number of parameters of multimodal large models is generally astonishingly large, and the effect of a single model processing multimodal data simultaneously may not be good. In addition, the demand for hardware computing power is also astonishing. Even if a high-performance parallel computing system is used, the demand for bandwidth between systems is also particularly large, and data bandwidth often becomes a huge bottleneck. Summary of the invention

[0003] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a multimodal large model reasoning system, reasoning method and application. The system or method of the present invention can be distributedly deployed for multimodal scenarios, process multimodal input and output, understand user intentions and select appropriate large models for services, achieve load balancing and efficient reasoning of multiple models on the reasoning system, and introduce an evaluation feedback module to iterate the results repeatedly.

[0004] The present invention centrally deploys single models in multiple different fields and supports multiple models and input and output forms to meet the diverse needs of users. For example, the Stable-diffusion model is used to specifically process the needs of text images, the LLaMA or ChatGLM dialogue models are used to process the needs of daily information acquisition, the Zeroscope large model is used to process the needs of text videos, and the Deepseek Coder large model is used to process programmers' needs for AI code generation.

[0005] The present invention proposes a multi-modal large model reasoning system, which includes: an input module, an input preprocessing module, a load balancing module, a parallel computing module, an evaluation feedback module, and an output module;

[0006] The input module, the input preprocessing module, the load balancing module, the parallel computing module, the evaluation feedback module, and the output module are connected in sequence;

[0007] The input module is the input end of the entire system, processing the input information of the human-computer interaction interface, and is used to receive one or more combinations of input text, pictures, voice, documents, etc.;

[0008] Text input includes single text or prompt keywords; image input is inferred together with the input text; voice input converts the voice into text before inference; document input is inferred together with the input text;

[0009] The input preprocessing module is connected with the input module and the evaluation feedback module, receives inputs from the input module and the evaluation feedback module, preprocesses the input information in the input module, and iterates the reasoning input information according to the feedback given in the subsequent evaluation feedback module;

[0010] The preprocessing refers to task classification of the input text intent by using a large language model fine-tuned through localized supervised learning using a labeled dataset.

[0011] Specifically, the input preprocessing module uses a fine-tuned pre-trained large language model with a small number of parameters to perform task recognition on the input information through localized supervised learning;

[0012] The load balancing module monitors hardware resources in real time and makes scheduling decisions based on resource status to schedule system hardware and reasonably allocate system resources;

[0013] The load balancing module can monitor the operating status and idle status of hardware resources in real time, such as but not limited to the usage rate, memory usage, task queue length and other information of AI acceleration units such as GPU, TPU, NPU, etc., and ensure that the load balancing module can make scheduling decisions based on the latest resource status by real-time monitoring of relevant information;

[0014] Based on the AI ​​reasoning task requirements received from the input preprocessing module and the resource status monitored in real time, the load balancing module selects the most appropriate resource to allocate to the current task; in the resource allocation process, multiple factors may be considered, such as the urgency of the task, the computing power of the resources, and the specific requirements of the task for resources (for example, some tasks may be more suitable to run on a GPU, while other tasks may prefer a TPU or NPU).

[0015] If the current hardware resources are idle, select the hardware that can meet the running of the latest reasoning task model; if the current hardware resources are not idle, use resource recycling strategies including LRU (Least Recently Used) to eliminate the least recently used large language model and replace it with the latest reasoning task model.

[0016] The parallel computing module is responsible for the reasoning of AI tasks, and includes one or more hardware device units, wherein the hardware device units include one or more multi-card multi-servers and / or one or more multi-card single-servers, wherein the multi-card multi-servers and / or the multi-card multi-servers include one or more boards with AI computing units;

[0017] Based on the parallel hardware device units in the parallel computing module, AI accelerators such as multiple GPUs (graphics processing units), TPUs (tensor processing units), and NPUs (neural network processing units) can work in parallel to process complex AI reasoning tasks;

[0018] The server accesses the big model database through the network or direct connection, and the big model database contains the AI ​​big model required for specific business and the corresponding deployment code; the AI ​​big model and its code, and the deployment environment have been adjusted to optimal performance.

[0019] The evaluation feedback module receives the preliminary reasoning results generated by the parallel computing module, performs iterative reasoning tasks, and evaluates single or multiple model results based on the set Prompt keyword; that is, it iterates repeatedly according to the Prompt keyword and single or multiple model results. If the result is less than the preset threshold, it is sent back to the input preprocessing module for re-reasoning, and finally the merged reasoning result is sent to the output module.

[0020] The output module is used to output the reasoning results evaluated by the evaluation feedback module, and the output includes one or a combination of text, pictures, voice, video, code, etc.

[0021] The present invention also provides a multi-modal large model reasoning method, the method comprising the following steps:

[0022] Step 1: User input processing: The user submits multimodal input through the input module and transmits it to the input preprocessing module;

[0023] Step 2: The input preprocessing module preprocesses the received multimodal input, understands and classifies the input intent, and selects a suitable large model for preliminary reasoning processing;

[0024] Step 3: The load balancing module allocates resources based on the current status of hardware resources and the requirements of the inference task;

[0025] Step 4: The parallel computing module executes the reasoning task, and one or more acceleration units and / or the large model database work in parallel; during the reasoning process, the parallel computing module accesses and executes the reasoning task stored in the large model database according to demand;

[0026] Step 5: The evaluation feedback module uses the Agent to evaluate and provide feedback on the preliminary reasoning results; based on the evaluation results, the evaluation feedback module determines whether it is necessary to pass the reasoning evaluation results to the input preprocessing module for further iterative optimization;

[0027] Step 6: After one or more iterations, the output module outputs the final inference results to the user in the form of text, pictures, voice, and video.

[0028] The present invention also provides the above-mentioned multimodal large model reasoning system, or the application of the reasoning method in product consultation reply generation, product intelligent recommendation, automatic generation of document summaries, image screening, image intelligent commentary, etc.

[0029] The present invention also provides a hardware system for implementing the above-mentioned reasoning method, and the hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned reasoning method is implemented.

[0030] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned reasoning method is implemented.

[0031] The beneficial effects of the present invention include:

[0032] This invention achieves load balancing and efficient reasoning of multiple models on a large system:

[0033] Load balancing: Large model reasoning based on the Transformer framework consumes a lot of GPU computing resources and video memory resources. How to make full use of hardware resources to reason about large models is a very challenging task. The high-performance parallel computing system that the inference system of the present invention relies on in the background is a high-computing computing center composed of many AI computing modules such as GPU, TPU, NPU, etc. After the load balancing module obtains the inference task from the input preprocessor module, it will obtain the dynamic load of all AI computing modules in real time, including the idle status of video memory, and use the polling allocation algorithm to reasonably allocate hardware resources. When the video memory resources are insufficient, the LRU algorithm is used to replace the longest unused model.

[0034] Efficient reasoning: The models contained in the large model database in the high-performance parallel computing module and the reasoning code matching the model are all optimally debugged and introduce a combination of the latest technologies for large model reasoning. Taking the domestic open source large models "Baichuan" and "Yi" as examples, their officially released reasoning codes are based on the hugging face reasoning framework of python. Their 7B / 6B (model parameter quantity) FP16 (model quantization accuracy) reasoning speed is roughly 10+tokens / s. After local transplantation using the C++ reasoning framework of LLaMA.CPP, its reasoning speed can be increased to 30~40tokens / s. Using the recently released TensorRT-LLM framework of NVIDIA for transplantation reasoning, its speed can be increased by 30%~50% on the basis of the original LLaMA.CPP.

[0035] The present invention also introduces an evaluation feedback module, which can iterate the results repeatedly:

[0036] The reasoning process of all the above-mentioned officially provided open source large models is one-time. For large models with high divergence, such as Stable-diffusion, the quality of the generated pictures has a great deal of randomness, and users often need to interact repeatedly to obtain a satisfactory picture. The reasoning system in the present invention leaves the iterative process to the large model itself. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 It is the overall framework and data flow diagram of the reasoning system of the present invention. DETAILED DESCRIPTION

[0039] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.

[0040] In the present invention, for end users, the multimodal large model reasoning system of the present invention provides an input and output system as an interface for human-computer interaction. After the reasoning system obtains the input, it identifies the user's intention through the input preprocessing module, selects a suitable AI large model for the first round of reasoning, and performs one to multiple iterations according to the feedback results given by the evaluation feedback module. Its specific reasoning task will be carried out in each registered hardware server containing an AI acceleration module. The reasoning system of the present invention has prepared the corresponding AI large model and the reasoning code with great performance optimization on the local disk. Each server will download the model and code when reasoning a new task for the first time. After that, its model and code can be resident in the memory of AI accelerators such as GPU, TPU or NPU to speed up the reasoning process. In the event of insufficient memory, the LRU algorithm is used to execute the model replacement strategy. The modules in each system will be described in detail below.

[0041] Input module: Processes the input of the human-computer interaction interface, which can be a single form of text and prompt keywords, pictures, voice, or the entire document, or a combination of one or more input forms such as text + picture or text + document. The more typical scenarios are as follows: Scenario 1 (text), text input "Please introduce the achievements of Qin Shihuang Yingzheng"; Scenario 2 (text + document), text input "Please summarize this document and tell me who is the father of Qiao Feng in Tian Long Ba Bu", and upload the document "Tian Long Ba Bu"; Scenario 3 (voice + picture), voice input "Please describe the information in this picture in detail", and upload a picture.

[0042] Generally speaking, "single-form text" usually refers to any text information directly input by the user, without a specific structure or predefined usage restrictions. The input text information can be a question, an instruction, a description, or a user's random expression. Its content and format are diverse and entirely depend on the user's intention;

[0043] The "prompt keyword" is a pre-set word or phrase with a specific meaning, which is used to trigger specific operations or responses, guide instructions or queries, help the system quickly identify the user's request type, and start the corresponding processing logic.

[0044] Input preprocessing module: The input preprocessing module is connected to the input module and the evaluation feedback module, and can preprocess the input received by the input module and perform one or more iterations on the feedback provided by the evaluation feedback module. The input preprocessing module receives the input of the input module and the evaluation feedback module, and inputs the text information (including single-form text or prompt keywords) into finetuned large models including BERT and GPT2 for task recognition, recognizes the instruction information in the text, and selects a suitable large model to reason about a specific task. The model in the input preprocessing module does not perform specific tasks, so a pre-trained large language model with a small parameter amount can be used through localized supervised learning to complete multi-classification tasks such as task classification and large model selection. For example, a 0.1B parameter Bert or GPT2 pre-trained large language model can be used to complete fine-tuning and deployment, which is used to implement input preprocessing and classify text intent;

[0045] When the input is in the form of an image, it will be inferred and classified together with the relevant text of the input; when the input is in the form of a document, it will be inferred and classified together with the relevant text of the input; when the input is in the form of voice, the voice will be converted into text and then inferred and classified.

[0046] Based on a preprocessed large model with a certain semantic understanding ability, such as Bert or GPT2 mentioned above, it is fine-tuned using labeled data so that the model learns to correctly assign task category labels according to the input text; for example, there are four models A, B, C, and D, which represent dialogue, text image, text video, and 2D to 3D image, etc., and their label data can be: ① Dialogue task "Please tell me", label "A"; ② Text image task "Please draw a picture", label "B"; ③ Text video task "Please generate a video", label "C"; ④ 2D to 3D image task "Please convert the following 2D image into a 3D image", label "D".

[0047] Load balancing module: The load balancing module is responsible for monitoring and scheduling the hardware and allocating resources reasonably. After obtaining the AI ​​reasoning task from the input preprocessing module, it obtains the operation and idle status of the hardware resources in real time, selects a reasonable resource, and allocates it to the current task. Its hardware resources include but are not limited to GPU, TPU, NPU, etc. In the absence of suitable hardware resources, such as when the memory of all AI acceleration unit modules is already fully loaded, the LRU algorithm is enabled to eliminate the longest unused LLM large model and replace it with the model of the latest reasoning task.

[0048] In a specific implementation, the load balancing module will give priority to obtaining the memory / video memory utilization of each resource. For example, the video memory of a single NVIDIARTX4090 graphics card is 8GB / 24GB (used / total), and the video memory of AMD RX7900 XTX is 16GB / 24GB (used / total). Module B selects a LLaMA27B FP16 dialogue model according to the task type, and its model size is 13GB. The algorithm will choose to use NVDIARTX4090 for the inference task.

[0049] Parallel computing module: The parallel computing module is responsible for the reasoning of AI tasks. It is composed of a multi-card multi-server composed of multiple boards with AI computing units, or a multi-card single-server hardware device unit. All servers can access the large model database through the network or direct connection; the large model database contains the AI ​​large models and corresponding deployment codes required for specific businesses, and each model and its related code are the optimal performance combination with the deployment environment. For example, for the scene of Wensheng map, the output time of the Stable-diffusion large model with TensorRT on NVIDIA RTX4090 is 500ms (steps20), and the reasoning speed of the LLama2 large model with TensorRT-LLM is as high as 205tokens / s.

[0050] In the present invention, the parallelization in parallel computing can be data parallelism (the same model is run in parallel on different data subsets), model parallelism (different parts of the model are run on different hardware), or a combination of the two, depending on the task characteristics and hardware configuration;

[0051] The large model database also includes a model version management function, which allows storage and retrieval of different versions of models and codes, facilitating model upgrades, rollbacks, or A / B testing when necessary. In addition, a distributed storage architecture can be used to meet the storage needs of massive model data, and a cache mechanism can be used to accelerate access to model data and reduce network transmission delays.

[0052] Evaluation feedback module: The evaluation feedback module uses Langchain's Agent, which contains a 6B dialogue model. It iterates repeatedly based on the Prompt keyword and single or multiple model results. If the result is not as expected, it is sent back to the input preprocessing module for multiple reasonings, and finally the combined result is sent to the output module.

[0053] Langchain is used to iterate the model generation results. The overall process is as follows:

[0054] 1. Accept user input.

[0055] 2. Preprocess the large model to decide whether to use a deployed AI large model for reasoning tasks.

[0056] 3. Perform reasoning tasks and record observation results (i.e., the output obtained after reasoning using the AI ​​model).

[0057] 4. The model, inputs, and observations are passed back to the agent, which decides what steps to take next.

[0058] 5. Repeat the above process until the agent decides that it no longer needs to perform the AI ​​reasoning task and then responds directly to the user.

[0059] The large model is pre-trained using labeled data. The large model directly performs a comprehensive score based on the generated results and the source of the problem. The present invention provides a threshold, such as 0.6. If the score exceeds the threshold, the iterative process is exited.

[0060] Output module: It belongs to the output end of the human-computer interaction interface and presents the final result of large model reasoning to the end user. The result can be a single form of text, picture, voice, video, or code, or one or more of text + picture or text + video.

[0061] The present invention also provides a method for generating and outputting a large model reasoning result based on the large model deployment system with multi-modal input and output, the method comprising the following steps:

[0062] Step 1: User input processing: The user submits multimodal input through the input module and transmits it to the input preprocessing module;

[0063] Step 2: The input preprocessing module preprocesses the received multimodal input, understands and classifies the input intent, and selects a suitable AI large model for preliminary reasoning processing;

[0064] Step 3: The load balancing module allocates resources based on the current status of hardware resources and the requirements of the inference task. When resources are insufficient, the LRU algorithm is used to dynamically replace the model to optimize resource usage.

[0065] Step 4: The parallel computing module executes the AI ​​reasoning task, and one or more AI acceleration units and / or the large model database work in parallel; during the reasoning process, the parallel computing module accesses and executes the reasoning task stored in the large model database according to demand;

[0066] Step 5: The evaluation feedback module uses the Agent to evaluate and provide feedback on the preliminary reasoning results; based on the evaluation results, the evaluation feedback module determines whether it is necessary to pass the reasoning evaluation results to the input preprocessing module for further iterative optimization;

[0067] Step 6: After one or more iterations, the output module outputs the final inference result to the user in a form that suits the user's needs (text, picture, voice, video, etc.).

[0068] The present invention inputs the multi-modal inputs into the corresponding processing large models in the multi-modal input and output large model deployment system respectively.

[0069] Example

[0070] 1. Aiming at user needs, reasoning and deployment characteristics of large models, and referring to hardware performance, the embodiment of the present invention selects multiple large models of different types, such as the Stable-diffusion basic model of SD1.5 and SDXL versions. In order to improve the expressiveness of the lines and overall shape of the generated pictures, 15 Controlnet models and LORA models are deployed at the same time; the text dialogue models include 7B, 13B Baichuan2, 6B, 34B Yi models; the code assistance models include 7B, 33Bdeepseek; the voice dialogue models include 7B, 13B TalkLLaMA; the cultural video models include Zeroscope, etc.

[0071] 2. For language models, select the LLaMA.CPP framework with relatively high reasoning performance, or TensorRT-LLM. For text and graph models, use the Stable-diffusion webui framework with acceleration plug-ins such as xFormer and TensorRT to achieve image output within 1 second.

[0072] 3. The hardware platform, also known as the high-performance parallel computing system, is a multi-card and multi-server cluster composed of NVIDIAA100, RTX4090, Orin 64G and Apple M2 ULTRA. Its resources are uniformly scheduled by the load balancing module.

[0073] 4. The iteration of the big model results is completed by LangchainAgent connected to a 6B language model. It comprehensively applies the pre-trained big model to control and judge common sense questions, and observes and judges the results generated by the model.

[0074] 5. The human-computer interaction system, whose software framework is completed by Gradio, which can easily realize the visual deployment of AI algorithms. It conveniently realizes the separation of front-end and back-end through the HTTP protocol, and can provide rich input and output methods in conjunction with HTML, JS, CSS and other scripts to connect with various user interaction methods.

[0075] Take the Vincent map as an example:

[0076] 1. Input "Please use the following text to generate a picture 'A cute little Chinese girl, wearing a red Hanfu, two pigtails on her head, photography style'"

[0077] 2. Preprocessing module identification uses stable diffusion model to reason about this task

[0078] 3. The prompt keyword "a cute little Chinese girl, wearing a red Hanfu, with two pigtails on her head, photography style" is passed to the load balancing module, which selects an idle GPU. If there is no idle GPU, the model that has not been used for the longest time is replaced to perform reasoning on the stable diffusion task.

[0079] 4. Currently, the evaluation and feedback system mainly scores the generated text information, and only performs simple image recognition on the image content. If the image generation is normal, it will be handed over to the output module. If the image generation is abnormal or not generated, it will continue to iterate once.

[0080] In addition to implementing the client and server in a purely computer-readable program code, the client and server can also implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a client and server can be considered as a hardware component, and the means for implementing various functions included therein can also be considered as a structure within the hardware component. Or even, the means for implementing various functions can be considered as both a software module for implementing the method and a structure within the hardware component.

[0081] It can be seen from the above description of the implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation method of the present application or some parts of the implementation method.

[0082] Each implementation in this specification is described in a progressive manner, and the same or similar parts between the various implementations can be referred to each other, and each implementation focuses on the differences from other implementations. In particular, for the implementation of the client and the server, both can refer to the introduction of the implementation of the aforementioned method for comparative explanation.

[0083] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0084] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.

Claims

1. A multimodal large model reasoning system, characterized in that: The reasoning system includes: an input module, an input preprocessing module, a load balancing module, a parallel computing module, an evaluation feedback module, and an output module; The input module is the input end of the entire reasoning system, and is used to receive one or more of the input text, pictures, voice, and documents; The input preprocessing module is used to preprocess the input information in the input module and iterate the reasoning input information according to the feedback given in the evaluation feedback module; The load balancing module is used to monitor the scheduling system hardware and allocate system resources; The parallel computing module is used to reason about the task and generate reasoning results; The evaluation feedback module is used to receive the reasoning results generated by the parallel computing module, and to perform verification feedback and result evaluation on the reasoning results; The output module is used to output one or a combination of text, pictures, voice, video, and code.

2. The inference system according to claim 1, characterized in that The input module is used to process the information input by the human-computer interaction interface; the text input includes a single text or prompt keyword; For input in the form of images, reasoning is performed together with the input text; for input in the form of voice, the voice is converted into text before reasoning is performed; for input in the form of documents, reasoning is performed together with the input text.

3. The reasoning system according to claim 1, characterized in that The input preprocessing module interfaces with the input module and the evaluation feedback module, receives inputs from the input module and the evaluation feedback module, preprocesses the inputs received by the input module, and iterates the feedback provided by the evaluation feedback module once or multiple times; The preprocessing refers to performing task recognition on input information by using a large language model fine-tuned through localized supervised learning using a labeled dataset.

4. The reasoning system according to claim 1, characterized in that The load balancing module monitors hardware resources in real time and makes scheduling decisions based on resource status; If the current hardware resources are idle, select hardware that can meet the needs of running the latest inference task model; If the current hardware resources are not idle, the resource recycling strategy is used to eliminate the least recently used large language model and replace it with the latest reasoning task model.

5. The reasoning system according to claim 1, characterized in that The parallel computing module includes one or more hardware device units, wherein the hardware device units include one or more multi-card multi-servers and / or one or more multi-card single-servers, wherein the multi-card multi-servers and / or the multi-card multi-servers include one or more boards with computing units; Parallel task processing is performed through one or more computing units in the parallel computing module; the server accesses the large model database through the network or direct connection, and the large model database contains the large models required for specific businesses and the corresponding deployment codes.

6. The reasoning system according to claim 1, characterized in that The evaluation feedback module repeatedly iterates according to the Prompt keyword and single or multiple model results. If the result is less than the preset threshold, it is sent back to the input preprocessing module for re-inference, and finally the combined inference result is sent to the output module.

7. A method for generating and outputting multimodal large model reasoning results, characterized in that: The method comprises the following steps: Step 1: User input processing: The user submits multimodal input through the input module and transmits it to the input preprocessing module; Step 2: The input preprocessing module preprocesses the received multimodal input, understands and classifies the input intent, and selects a suitable large model for preliminary reasoning processing; Step 3: The load balancing module allocates resources based on the current status of hardware resources and the requirements of the inference task; Step 4: The parallel computing module executes the reasoning task, and one or more acceleration units and / or the large model database work in parallel; during the reasoning process, the parallel computing module accesses and executes the reasoning task stored in the large model database according to demand; Step 5: The evaluation feedback module uses the Agent to evaluate and provide feedback on the preliminary reasoning results; based on the evaluation results, the evaluation feedback module determines whether it is necessary to pass the reasoning evaluation results to the input preprocessing module for further iterative optimization; Step 6: After one or more iterations, the output module outputs the final inference results to the user in the form of text, pictures, voice, and video.

8. Application of the multimodal large model reasoning system as described in any one of claims 1 to 6, or the reasoning method as described in claim 7 in product consultation reply generation, product intelligent recommendation, automatic document summary generation, image screening, and image intelligent commentary.

9. A hardware system for implementing the method as claimed in claim 7, characterized in that: The hardware system comprises: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to claim 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to claim 7 is implemented.

Citation Information

Cited By

  • Hardware reliability scheduling method, device and equipment for large model reasoning and medium

    CN122044900A