ME-RAG-based data center large model intelligent operation and maintenance method and system
By adopting a multi-expert RAG-based multi-expert RAG method in data center operation and maintenance, a multi-modal database and expert agent system is built, and the existing large language model is difficult to provide practical and effective solutions in data center operation and maintenance, achieving more efficient and intelligent operation and maintenance management, and providing more intuitive and comprehensive answers.
Patent Information
- Application Number
- CN202510480917.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing large language models are difficult to provide practical and effective solutions in data center operation and maintenance. Due to the huge database scale, high computing cost and poor scalability, it is difficult to provide intuitive and clear guidance, and is limited by context length, it is difficult to deal with problems in multiple knowledge fields.
Using the intelligent operation and maintenance method of large-scale data centers based on ME-RAG, by dividing data center data into multiple fields, creating an expert agent in each field, building a multi-modal database as an embedded knowledge base for expert agents, using the manager agent to dynamically select the expert fields involved in combination with expert skill list and user questions, and assign them to relevant expert agents for answering.
The correctness and comprehensiveness of answers to user questions are improved. Through the multimodal database retrieval mechanism, users are provided with more intuitive multimodal information, reducing resource consumption, and avoiding large language models that ignore weak correlation but related fields due to context limitations.
Smart Images

Figure CN120011523A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent operation and maintenance of data centers, and specifically relates to a method and system for intelligent operation and maintenance of a large model of a data center based on Multi-Experts RAG (ME-RAG). Background Art
[0002] In recent years, the demand for computing power has increased dramatically, and data centers (DCs) have become the focus of global attention. The stability and security of data centers are directly related to the operation of many industries, and their importance is self-evident. For this reason, it is extremely important to carry out effective operation and maintenance management of data center facilities. At present, the operation and maintenance of data center facilities mainly rely on professional technicians to provide 7×24 hours of uninterrupted service. This not only places extremely high demands on the professional quality and energy of operation and maintenance personnel, but also requires a lot of costs in talent training and operation and maintenance resource investment.
[0003] In recent years, large conversational language models such as GPT, LLama, Mistral, and ChatGLM have brought hope for solving the above problems. They have strong emergence and generalization capabilities and have great potential in the field of data center operation and maintenance. However, the data center field is highly professional and rapidly developing. Existing large language models can hardly provide effective solutions for data centers intuitively with their own knowledge reserves. Retrieval-Augmented Generation (RAG) can deploy large specialized models in the DC field at low cost. The current RAG method usually treats all text data as a knowledge base, which results in the database being too large, the computational cost of the retrieval module being greatly increased, and the scalability being unsatisfactory. And due to the lack of image information to assist understanding, these methods are difficult to provide intuitive and clear guidance during the maintenance of data center equipment. The data center environment is complex and large in scale. Such environments require experts from multiple fields to work together to solve problems efficiently. However, due to the limitation of the context length of the large language model, when a user asks a question involving multiple knowledge fields, the fragments of fields that are closely related to the question will be given priority, while the fragments of fields with lower relevance may be ignored due to their low ranking. As a result, the answers provided to users are often neither complete nor accurate enough.
[0004] Therefore, there is an urgent need for a more efficient and intelligent operation and maintenance solution to make artificial intelligence technology easier to understand and apply, so as to improve operation and maintenance efficiency, ensure operation and maintenance safety, and ultimately realize automated operation and maintenance of data centers. Summary of the invention
[0005] Purpose of the invention: The purpose of the present invention is to provide a method and system for intelligent operation and maintenance of a large data center model based on ME-RAG, which effectively improves the correctness and comprehensiveness of answers to user questions through a multi-expert operation and maintenance mechanism, and provides users with more intuitive multimodal information through a multimodal database retrieval mechanism.
[0006] Technical solution: To achieve the above-mentioned invention object, the present invention adopts the following technical solution: In a first aspect, the present invention provides a data center large model intelligent operation and maintenance method based on ME-RAG, comprising the following steps: Divide the data center data into multiple fields and create an expert agent for each field; For each field, a multimodal database including text database, image database and table database is constructed through unstructured documents as the embedded knowledge base of expert agents; Based on the ME-RAG framework, the manager agent combines the expert skill list with the user's question to dynamically select the expert field involved and assigns it to the relevant expert agent for answering; Each expert agent who receives the answer instruction searches his or her own knowledge base according to the question and answers the question in text. If there are relevant pictures and tables, they will be returned together. Finally, the reporting specialist will summarize the answers of all the expert agents.
[0007] Preferably, the text database is constructed in a multi-vector manner, including: extracting the text portion of the unstructured document and dividing it into blocks according to subsections, each block is put into a large language model for summarization to obtain a summary block, embedding the summary block, and linking it to the source text block through a multi-vector.
[0008] Preferably, the image database is constructed in a multi-vector manner, including: extracting images from unstructured documents, and putting the names of the images and the context around the images into a large language model, summarizing the functions of the images to obtain summary blocks, embedding the summary blocks, and linking the summary blocks with the image saving path through multi-vector.
[0009] Preferably, the table database is constructed in a multi-vector manner, including: extracting the table part of the unstructured document to form a csv document, then summarizing the table context, table name and table content in a large language model to obtain a summary block, and linking the summary block with the csv file through multi-vector.
[0010] Preferably, the manager agent is composed of a large language model and a single classifier. The expert skill list is input into the large language model by means of prompt words so that the large language model knows the skills possessed by each expert. The large language model combines the expert skill list with the user's question and dynamically selects an expert agent suitable for answering the user's question. The single classifier is trained using business texts in various fields as input. The single classifier and the large language model jointly select an expert agent, and the union of the two is taken as the expert agent selected by the manager agent.
[0011] Preferably, the expert agent is composed of a retrieval module and a large language model corresponding to three modal databases, namely, a text database, an image database, and a table database; the retrieval module obtains a summary block with the highest semantic similarity to the user's question by calculating in the semantic space, and returns the source linked to the summary block to the large language model; the large language model answers the user's question in combination with the information returned by the retrieval module, and presents relevant image and / or table information.
[0012] Preferably, according to the actual data center operation and maintenance process, the data center data is divided into multiple fields according to different devices, and the expert skill list in each field is provided to the large language model in the form of a directory through prompt words.
[0013] In the second aspect, the present invention provides a data center large model intelligent operation and maintenance system based on ME-RAG, which is built based on the ME-RAG framework and includes a manager agent, multiple expert agents, and a reporting specialist; each expert agent corresponds to a field of data center information, and through unstructured documents, a multimodal database including a text database, an image database and a table database is constructed as an embedded knowledge base for the expert agent; the manager agent is used to dynamically select the expert field involved in combination with the expert skill list and user questions, and assign them to relevant expert agents for answering; each expert agent that receives the answer instruction searches its own knowledge base according to the question, answers the question in text, and returns relevant pictures and tables if there are any; the reporting specialist is used to summarize the answers of all expert agents.
[0014] Preferably, the system is built using the AutoGen open source framework and consists of a manager agent, multiple expert agents and a reporting specialist.
[0015] In a third aspect, the present invention provides a computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the steps of the ME-RAG-based data center large model intelligent operation and maintenance method are implemented.
[0016] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention integrates multimodal databases into multi-expert retrieval, so that ordinary non-multimodal large models can process multimodal information, including text, tables and images, and obtain more intuitive and easy-to-understand answers. 2. The present invention makes full use of the thinking ability of the large language model. The large language model assisted by the classifier acts as a manager. It can coordinate domain experts and handle fuzzy queries at the same time, effectively reducing the risk of hallucinations; a series of expert models divide the database into multiple fields, improving the scalability of the database, which avoids the resource consumption caused by traversing large-scale databases, and the situation where the large language model may ignore weakly related but related fields due to contextual limitations, effectively improving the accuracy and comprehensiveness of the answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the ME-RAG framework and expert model in an embodiment of the present invention.
[0018] Figure 2 It is a schematic diagram of the operation mode of the manager agent in the embodiment of the present invention.
[0019] Figure 3 It is a schematic diagram of the thought chain for constructing a question-answering data set in an embodiment of the present invention.
[0020] Figure 4 It is a comparison chart of the answer effects of manager agents with different composition modes in the embodiments of the present invention.
[0021] Figure 5 It is a diagram showing the effect of the picture answer given as an example in the embodiment of the present invention.
[0022] Figure 6 It is a diagram showing the effect of the table answer given as an example in the embodiment of the present invention.
[0023] Figure 7 It is a diagram showing the effect of the comprehensive answer given as an example in the embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the embodiments of the present invention are described in detail below in conjunction with the accompanying drawings: This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given. It should be understood that the specific examples described here are only used to explain the present invention, but the protection scope of the present invention is not limited to the following embodiments.
[0025] An embodiment of the present invention discloses a data center large model intelligent operation and maintenance method based on ME-RAG, which mainly includes: dividing data center data into multiple fields, creating an expert agent for each field, and constructing a multimodal database including a text database, an image database and a table database through unstructured documents as an embedded knowledge base for the expert agent; based on the ME-RAG framework, the manager agent dynamically selects the expert field involved in combination with the expert skill list and the user question, and assigns it to the relevant expert agent for answering; each expert agent that receives the answer instruction searches its own knowledge base according to the question, answers the question with text, and returns relevant pictures and tables if there are any; finally, the reporting specialist summarizes the answers of all expert agents, so as to achieve a more comprehensive, intuitive and intelligent answer to the data center user question.
[0026] The embodiment of the present invention builds a manager agent and a multi-expert agent, clarifies the division of labor, and effectively alleviates the contextual limitations of large language models and the problem of incomplete retrieval; through the ME-RAG framework, the manager agent assigns questions to experts in all relevant fields and considers the problem from multiple angles; the expert agent converts private unstructured data into a multimodal index database to realize data center dialogue question and answer. The embodiment of the present invention adopts a multi-expert agent mechanism to subdivide the responsibilities of experts in various fields, avoiding the problem of a single searcher retrieving multiple similar knowledge blocks in highly relevant fields while ignoring knowledge in weaker relevant fields, resulting in the occurrence of a single retrieval knowledge field, and at the same time provides scalability for the introduction of subsequent functions. Specifically, for the operation and maintenance of a data center, the real data center operation and maintenance process can be simulated, and a ME-RAG intelligent operation and maintenance system can be built according to different equipment (such as air-conditioning equipment and refrigeration equipment).
[0027] This embodiment also discloses a data center large model intelligent operation and maintenance system based on ME-RAG, the purpose of which is to establish a data center privatization large model, improve the operation and maintenance efficiency of the data center, reduce the pressure of manual operation and maintenance, and ensure the comprehensiveness and accuracy of the answers. The system is built based on the ME-RAG framework, including a manager agent, multiple expert agents, and a reporting specialist; each expert agent corresponds to a field of data center information, and through unstructured documents, a multimodal database including a text database, an image database, and a table database is constructed as an embedded knowledge base for the expert agent; the manager agent is used to dynamically select the expert field involved in combination with the expert skill list and the user's question, and assign it to the relevant expert agent for answering; each expert agent who receives the answer instruction searches its own knowledge base according to the question, answers the question with text, and returns relevant pictures and tables if there are any; the reporting specialist is used to summarize the answers of all expert agents.
[0028] Combine the following Figure 1The ME-RAG framework shown in the figure further illustrates the detailed steps of the embodiment of the present invention. Figure 1 As shown in the figure, the ME-RAG framework is used for multi-expert retrieval enhancement generation in data centers. It consists of three types of intelligent agents: manager agent, expert agent, and reporting specialist. In the data center scenario, the manager agent is responsible for assigning user questions to all experts in the relevant field. Each expert represents a professional in a field. When receiving the command from the manager agent, they will use their knowledge in a specific field to answer the user's questions. The reporting specialist will summarize the answers of all the experts who have responded. Compared with traditional RAG, ME-RAG will get more comprehensive and accurate answers to questions involving multiple fields.
[0029] Traditional RAG mainly consists of a retrieval module and a generation module, which are usually used to process text information, as shown below:
[0030] Among them, the retrieval module In the document Search to extract the most relevant document fragments , and then compare these fragments to the problem Combined together as a generation module The input finally generates the answer .
[0031] For the ME-RAG framework proposed in this embodiment, all text files They are classified into multiple text repositories according to their fields, i.e. ,in Represents the number of fields. In the manager proxy Selected with question Related After the expert, report to the specialist Integrate the answers. Therefore, the above formula can be rewritten as follows:
[0032] in, , , Represent the document, retrieval module and generation module of the i-th field respectively. This method of storing the database separately effectively alleviates the scalability limitations of large-scale databases, significantly expanding the horizontal coverage across different fields while ensuring that the vertical depth of retrieval is maintained. This multi-expert mechanism not only better fits the workflow of the real world, but also effectively alleviates the contextual limitations of large language models through collaborative responses, thereby ensuring the comprehensiveness and accuracy of the answers.
[0033] Agent Manager The main structure of Figure 1 As shown in the left part of the figure, in this embodiment, the agent manager will use a large language model and a single classifier to analyze user queries and select experts in related fields to respond. In this process, the large language model acts as a cognitive tool to evaluate the problem and determine the relevant field. However, the large language model only judges based on the names of the experts and lacks understanding of the skills that each expert is good at. This is because detailed information about the data center may not be included in the initial training stage of the large language model. In order to enhance the large language model's understanding of the specific skills of each expert, a directory is introduced in this embodiment, which lists the capabilities of each expert, and the directory is entered into the prompt word in the format of List. This structured presentation method allows the large language model to better understand the relevant expertise available in the system.
[0034] Despite the superior capabilities of the large language model, it occasionally produces difficult-to-explain "hallucinations", that is, generating erroneous or unreasonable content, which may have an adverse impact on the normal operation of the system. To mitigate the possible impact of these "hallucinations", a single classifier framework is proposed in this embodiment, which combines machine learning technology with a large language model to support system operation and build a powerful manager agent. Specifically, the classifier uses business texts from various fields as input to train a support vector machine (SVM) model. The user's query is simply processed by this classification model to determine the most relevant field. The manager agent uses the union of the large language model and the single classifier output as the selection result. By collaborating with the single classifier, the large language model enables the manager agent to further identify and select a suitable list of experts. This integrated approach significantly improves the accuracy and reliability of expert selection, thereby ensuring the smooth and efficient operation of the system.
[0035] One of the main challenges facing RAG is the fuzzy query of users. Fuzzy descriptions often lead to the wrong search direction or inability to obtain relevant information completely. To solve this problem, in this embodiment, the agent manager will select all possible relevant domain experts based on his understanding of the user's query and the professional field of each expert. These experts will give a response together to ensure that a comprehensive answer is provided from multiple perspectives. The specific workflow example is as follows: Figure 2As shown in the figure, when users ask "What features should be considered when inspecting a battery pack?", they may actually also want to know information related to the uninterruptible power supply (UPS) system, although this intention is not explicitly stated in the question. Traditional RAG may retrieve information related to battery equipment based on semantic similarity because this question seems to be more related to battery equipment, which may lead to misleading results. The ME-RAG framework proposed in this embodiment can solve this problem. The agent manager will select a battery expert and an uninterruptible power supply expert to jointly answer the user's query. This multi-expert collaboration approach enables the system to respond to the user's query more comprehensively, ensuring that all relevant aspects of the question are fully answered. By properly soliciting the insights of multiple experts, the quality of the information provided is improved, the perspective of problem solving is broadened, and ultimately users are helped to make more informed decisions.
[0036] Due to the complexity and large number of components of data center equipment, text descriptions alone are often not enough to convey information intuitively in daily maintenance operations. In many cases, appropriate images and tables need to be added to give a clearer and more comprehensive explanation. Therefore, as shown in the right part of Figure 1, experts must have the ability to process multimodal information. In this embodiment, each expert consists of three parts: a database, a retrieval module, and a generation module. We divide knowledge into three modes: tables, images, and text, and use a large language model to build a comprehensive multi-vector. This method can effectively retrieve different types of data and generate responses. By integrating multiple modes, the system can provide users with richer and more valuable answers, improving overall understanding and usability. The system has the ability to access and process information in various forms, which ensures that actual users can fully understand the current problems, thereby improving the efficiency and effectiveness of data center equipment operation and maintenance.
[0037] The specific method of constructing a multimodal external knowledge base is as follows: After the data center related information is subdivided into multiple fields, it is constructed into a text database. , Image Database and table database The database of each field contains multiple document blocks, and each document block is stored in a multi-vector form after some embedding processing. , , ,Right now:
[0038]
[0039]
[0040] Where k represents the index of the vector.
[0041] The specific source text-multi-vector conversion form is as follows:
[0042]
[0043]
[0044] For text content, this embodiment uses a large language model to analyze each text block. Summarize and get the summary block , and then through the multivectorizer Compare it with the corresponding Mapping and matching are performed. The retrieval module retrieves the most similar , and the corresponding Transmitted to the generation module; for the image content, we use the text context of the image And the name of the picture , prompting the large language model to generate summary blocks that describe the intent of the image Then through the multi-vector Will and The retrieval module finds the corresponding image according to the image intent and then feeds it back to the generation module; for tables, a similar method to image processing is used. And the entire table (table content and table name) Input into the large language model, which generates a description summary block of the table and its details , and finally through the multi-vectorizer Considering the differences in databases, the formula of the ME-RAG framework is rewritten as:
[0045] Each vector database acts as a retrieval module, outputting snippets that match the user's question to the generation module, specifically to the large language model, and ultimately the expert gives the answer. The multi-vector search and Jaccard search methods are combined to improve the search quality from a semantic level while reducing the impact of word frequency. The output results generated by the search process will be consistent with the user's question. Together, they are then fed into the generation module Finally, the reporting specialists will summarize the experts’ answers through prompt words to generate the answer .
[0046] The effect of the present invention is explained below in conjunction with a simulation experiment.
[0047] Simulation experiment settings: This experiment was built using the open source framework Autogen and the ChatGLM3-6B model. In addition, the embedding model used is text2vec-base-chinese, which can effectively process Chinese text representation. The data set is based on the key equipment operation and maintenance guide of Part 2 of "How far are you from operation and maintenance rookie to expert: Data center facility operation and maintenance guide". In order to compare the experimental results, a question-answering data set was constructed through the way of thinking chain, such as Figure 3 As shown. By outputting the intermediate steps of reasoning, the reasoning ability of the large language model is enhanced, making the reasoning process more transparent and comprehensive. The specific process is as follows: First, the chapters in the book are divided into several non-overlapping paragraphs. Using the generative power of the large language model, the key points of each text paragraph are summarized and refined.
[0048] Then, for each identified key point, the large language model is prompted to ask a corresponding question and give an answer based on information from the relevant paragraph.
[0049] Finally, a professional review was conducted to screen and improve the generated content, resulting in a high-quality question-answering dataset with 441 records in total, some of which contained unclear questions, to test the ME-RAG framework.
[0050] This systematic approach ensures the comprehensiveness and reliability of the question-answering dataset, thereby increasing its usefulness for data center operations.
[0051] According to the constructed dataset, the number of experts in the ME-RAG framework is set to 10, including uninterruptible power supply (UPS) system, data center infrastructure management (DCIM), data center operations management (DCOM), air-cooled air-conditioning system, integrated wiring system, battery, backup generator, KVM system, chilled water air-conditioning system, and high and low voltage systems.
[0052] In order to compare the effects, the experiment uses ROUGE and cosine similarity as the main evaluation indicators, which focus on the semantic space of text overlap and answer quality, respectively. ROUGE-1 emphasizes the matching of single words, while ROUGE-L considers the preservation of word order in the text. On the other hand, cosine similarity is used to measure the similarity between word vector pairs, which strongly reflects the semantic consistency between the corresponding texts. In addition, in order to evaluate the decision accuracy of the manager agent in the ME-RAG framework, the indicator "Expert Selection Accuracy (Expert_Accuracy)" is introduced. At the same time, when the large language model produces "hallucinations" and no experts are selected, all indicators of the corresponding samples are set to 0 in the experiment. This comprehensive evaluation framework can strictly evaluate the performance of the ME-RAG framework in generating relevant and accurate answers in the context of data center operations.
[0053] Table 1 shows the advantages of the proposed multi-expert mechanism, comparing the results of the ME-RAG framework with those of a single-expert RAG that integrates all knowledge. The results show that the ME-RAG framework outperforms the single-expert agent in all evaluation metrics.
[0054] Table 1 Comparison of single-expert / multi-expert RAG results method Rouge-1 Rouge-L Cosine Single Expert RAG 0.5990 0.5439 0.7926 ME-RAG 0.6534 0.5920 0.8150 At the same time, in order to prove the superiority of the ME-RAG framework in decision-making ability, the experiment used different decision-making methods and conducted a comprehensive comparison on various evaluation indicators, such as Figure 4 As shown. In this experiment, "NC (No content)" means that the agent manager only knows the expert's name, and the expert's specific skill content is not provided in the prompt. "LC (Content generated by the large language model)" means that the expert content generated by the large language model is included in the agent manager as a skill list. "MC (Manually created content)" refers to integrating manually created content (such as the table of contents of a book) into the agent manager as a skill list. "MC+Class" means adding a single classifier to handle the situation when the large language model produces "hallucinations".
[0055] from Figure 4It is obvious from the results that relying solely on the embedding knowledge of the large language model can only roughly classify the user's query, and it is difficult to accurately grasp the specific skills of each expert. This highlights the limitations of relying solely on the understanding ability of the large language model and emphasizes the necessity of introducing expert skill content. However, it is worth noting that expert content generated using the large language model leads to worse results. In the comparative experiments of the "LC" group, we observed that in many cases, the agent manager was unable to effectively classify the user's query. This is because the text directly input into the large language model, and the content generated by it, is too detailed and lengthy. Such a lengthy context hinders the ability of the large language model to effectively understand the user's query, significantly increases the possibility of "hallucination", and ultimately fails to generate accurate answers. In contrast, the decision mechanism of ME-RAG performs well in solving the "hallucination" problem. By using concise, manually curated, and easier to understand directory points, and with the support of a single classifier, it achieves the best decision performance.
[0056] exist Figures 5 to 7 In this paper, we simulated the actual results into a user interface to demonstrate the capabilities of the ME-RAG framework in multimodal retrieval and multi-expert analysis. Due to the limited length of the answer, we only included the key information of the answer in the figure. Figure 5 The image retrieval capability of the ME-RAG framework is demonstrated, which is able to locate relevant images from the dataset based on the user's query and provide an intuitive visual presentation. Figure 6 The ME-RAG framework is shown to be able to retrieve answers from table contents and reply to user queries. In order to give users a more intuitive understanding, the ME-RAG framework directly displays the retrieved tables to users. Figure 7 The comprehensive retrieval capability of the ME-RAG framework is demonstrated. When a user asks a question about evaporators, the evaporator equipment in both air-cooled and water-cooled air conditioning systems is relevant to the question. Therefore, the ME-RAG framework selects experts in these two fields to provide relevant information about evaporators, retrieves relevant images for better visualization, and synthesizes the responses of the two experts into a coherent summary.
[0057] Experimental results show that ME-RAG has good application prospects, and has the advantages of strong scalability and comprehensive answers. It solves the problems of single traditional RAG retrieval, alleviates the contextual limitations of large language models, and introduces multimodal information to make processing data center problems more intuitive, greatly alleviating the operation and maintenance pressure of data center maintenance personnel.
[0058] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the aforementioned ME-RAG-based data center large model intelligent operation and maintenance method are implemented.
[0059] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A data center large model intelligent operation and maintenance method based on ME-RAG, characterized in that: The following steps are involved: Divide the data center data into multiple fields and create an expert agent for each field; For each field, a multimodal database including text database, image database and table database is constructed through unstructured documents as the embedded knowledge base of expert agents; Based on the ME-RAG framework, the manager agent combines the expert skill list with the user question to dynamically select the expert field involved and assigns it to the relevant expert agent for answering; Each expert agent who receives the answer instruction searches his or her own knowledge base according to the question and answers the question in text. If there are relevant pictures and tables, they will be returned together. Finally, the reporting specialist will summarize the answers of all the expert agents.
2. According to claim 1, a data center large model intelligent operation and maintenance method based on ME-RAG is characterized by: The text database is constructed in a multi-vector manner, including: extracting the text part of the unstructured document and dividing it into blocks according to sections, each block is put into a large language model for summarization to obtain a summary block, embedding the summary block, and linking it to the source text block through a multi-vector.
3. According to claim 1, a data center large model intelligent operation and maintenance method based on ME-RAG is characterized by: The image database is constructed in a multi-vector manner, including: extracting images from unstructured documents, and putting the names of the images and the context around the images into a large language model, summarizing the functions of the images to obtain summary blocks, embedding the summary blocks, and linking the summary blocks with the image saving paths through multi-vector.
4. According to claim 1, a data center large model intelligent operation and maintenance method based on ME-RAG is characterized by: The table database is constructed in a multi-vector manner, including: extracting the table of the unstructured document to form a csv document, then summarizing the table context, table name and table content in a large language model to obtain a summary block, and linking the summary block with the csv file through multi-vector.
5. According to claim 1, a data center large model intelligent operation and maintenance method based on ME-RAG is characterized by: The manager agent is composed of a large language model and a single classifier. The expert skill list is input into the large language model by means of prompt words so that the large language model knows the skills possessed by each expert. The large language model combines the expert skill list with the user's question and dynamically selects an expert agent suitable for answering the user's question. The single classifier is trained using business texts in various fields as input. The single classifier and the large language model jointly select an expert agent, and the union of the two is taken as the expert agent selected by the manager agent.
6. According to claim 1, a data center large model intelligent operation and maintenance method based on ME-RAG is characterized by: The expert agent is composed of a retrieval module and a large language model corresponding to three modal databases, namely a text database, an image database, and a table database; the retrieval module obtains a summary block with the highest semantic similarity to the user's question by calculating in the semantic space, and returns the source linked to the summary block to the large language model; the large language model answers the user's question in combination with the information returned by the retrieval module, and presents relevant image and / or table information.
7. The method for intelligent operation and maintenance of a large data center model based on ME-RAG according to claim 1, characterized in that: According to the actual data center operation and maintenance process, the data center data is divided into multiple fields according to different devices. The expert skill list in each field is provided to the large language model in the form of a catalog through prompt words.
8. A data center large model intelligent operation and maintenance system based on ME-RAG, characterized by: Based on the ME-RAG framework, it includes manager agents, multiple expert agents, and reporting specialists. Each expert agent corresponds to a field of data center information. Through unstructured documents, a multimodal database including text database, image database and table database is constructed as the embedded knowledge base of the expert agent. The manager agent is used to dynamically select the expert field involved in combination with the expert skill list and the user question, and assign it to the relevant expert agent for answering; each expert agent that receives the answer instruction searches its own knowledge base according to the question, answers the question in text, and returns relevant pictures and tables if there are any; the reporting specialist is used to summarize the answers of all expert agents.
9. The ME-RAG-based data center large model intelligent operation and maintenance system according to claim 8, characterized in that: The system is built using the AutoGen open source framework and consists of a manager agent, multiple expert agents and a reporting specialist.
10. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is loaded into the processor, the steps of the ME-RAG-based data center large model intelligent operation and maintenance method are implemented according to any one of claims 1 to 7.
Citation Information
Patent Citations
Tacit knowledge acquisition method based on HWME (Hall for Workshop of Metasynthetic Engineering)
CN103425774A
Visual question and answer implementation method and method based on visual question and answer test model
CN115221369A
Unsupervised cross-modal hash retrieval method based on dynamic multi-expert knowledge distillation
CN116894120A
Construction method of expert system based on RAG technology
CN119227792A
Retrieval optimization method based on hierarchical expert routing model and CoT reasoning
CN119336900A
Cited By
Knowledge-intensive visual question and answer automatic data generation method and device
CN120930746A