Information processing method
The proposed method addresses the challenge of input data exceeding maximum processing length in large language models by using intent recognition and a knowledge graph-based index database to ensure accurate and efficient log data processing, enhancing system reliability and response precision.
Patent Information
- Application Number
- CN202510308113.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, in log analysis, the retrieved document data exceeds the maximum processing length of the large model if it is too long, resulting in system operation abnormalities and the accuracy of the result decreases.
Through the intent identification model, the information retrieval type of query information is identified, and the pre-constructed index database and dynamic search strategy are used to generate context reference information that meets the input requirements of large language models, and reason through the first large language model to ensure the appropriate length of the input data.
It avoids system exceptions caused by excessive input data, enhances system reliability and accuracy, can flexibly respond to diversified query needs, and improves the logical reasoning ability of large models for unknown fields.
Smart Images

Figure CN120316073A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of artificial intelligence technology, and particularly relates to an information processing method. Background Art
[0002] With the rapid development of large artificial intelligence models, various efficient system solutions for processing massive log data have emerged in the field of log analysis. However, due to the domain specificity of log data, traditional methods often face challenges in processing such data. For this reason, the Retrieval-Augmented Generation (RAG) technology has emerged, effectively solving the "hallucination" problem that may occur when large models process unknown data.
[0003] In the log parsing process based on large models, RAG first constructs an index database with text data. When a user initiates a query, the index database is retrieved to obtain relevant documents, and then the original query is merged with the retrieved document data and passed as input to the large model, and finally the corresponding answer is generated. However, this method has at least the problem that when the retrieved document data is too long, the merged input data may exceed the maximum processing length limit of the large model, thus affecting the normal operation of the system and the accuracy of the results. Summary of the Invention
[0004] Embodiments of this application are expected to provide an information processing method.
[0005] The technical solution of this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide an information processing method, and the method includes:
[0007] Input the query information for the target object into an intent recognition model for intent recognition to obtain the information retrieval type corresponding to the query information; wherein, the query information is used to query all relevant log information or partial relevant log information of the target object;
[0008] Utilize the pre-constructed index database, and through the retrieval strategy corresponding to the information retrieval type, retrieve the query information to obtain a retrieval result that meets the information length condition corresponding to the input requirements of the first large language model; wherein, different information retrieval types correspond to different retrieval strategies;
[0009] Perform fusion processing on the query information and the retrieval result to generate context reference information that meets the input requirements of the first large language model;
[0010] Use the first large language model to reason about the context reference information to generate the reply content corresponding to the query information.
[0011] In a second aspect, an information processing apparatus provided by an embodiment of the present application includes:
[0012] An intent recognition module, configured to input query information for a target object into an intent recognition model for intent recognition, and obtain an information retrieval type corresponding to the query information; wherein, the query information is used to query all relevant log information or partial relevant log information of the target object;
[0013] A retrieval module, configured to use a pre-constructed index database to retrieve the query information through a retrieval strategy corresponding to the information retrieval type, and obtain a retrieval result that meets the information length condition required for input to a first large language model; wherein, different information retrieval types correspond to different retrieval strategies;
[0014] A fusion processing module, configured to perform fusion processing on the query information and the retrieval result to generate context reference information that meets the input requirements of the first large language model;
[0015] An inference module, configured to use the first large language model to perform inference on the context reference information to generate a reply content corresponding to the query information.
[0016] In a third aspect, an electronic device provided by an embodiment of the present application includes: a processor and a memory;
[0017] The memory stores a computer program that can run on the processor;
[0018] The processor executes the computer program stored in the memory to implement the steps of the above information processing method.
[0019] In a fourth aspect, a storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by at least one processor, the steps of the above information processing method are implemented.
[0020] In a fifth aspect, a computer program product provided by an embodiment of the present application includes a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above information processing method are implemented.
[0021] An embodiment of the present application provides an information processing method. The query information for the target object is input into an intent recognition model for intent recognition to obtain the information retrieval type corresponding to the query information. Among them, the query information is used to query all relevant log information or partial relevant log information of the target object. Using the pre-constructed index database, through the retrieval strategy corresponding to the information retrieval type, the query information is retrieved to obtain a retrieval result that meets the information length condition required for the input of the first large language model. Among them, different information retrieval types correspond to different retrieval strategies. The query information and the retrieval result are fused to generate context reference information that meets the input requirements of the first large language model. The first large language model is used to reason about the context reference information to generate a reply content corresponding to the query information. In this way, first, through intent recognition and dynamic retrieval strategies, it is ensured that the length of the retrieval result always meets the input requirements of the large model, avoiding system operation anomalies caused by too long input data and enhancing the reliability of the system. Then, different information retrieval types correspond to different retrieval strategies, enabling the system to flexibly respond to diverse query requirements, and different retrieval strategies can accurately screen relevant content according to the query intent, reducing interference from irrelevant information, helping the large model better understand the query intent and generate accurate answers, and balancing retrieval accuracy and complexity. Finally, based on the data index database of the knowledge graph, the end-to-end process from user input to the output of the large model can have logical reasoning ability, solving the hallucination problem faced by the large model in unknown fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic flowchart of an optional information processing method provided by an embodiment of the present application;
[0023] Figure 2 It is a structural block diagram of an optional information processing method provided by an embodiment of the present application;
[0024] Figure 3 It is a partial structural schematic diagram of the intent recognition model provided by an embodiment of the present application;
[0025] Figure 4 It is a schematic flowchart of an optional information processing method provided by an embodiment of the present application;
[0026] Figure 5 It is a schematic flowchart of an optional information processing method provided by an embodiment of the present application;
[0027] Figure 6 It is a schematic flowchart of an optional information processing method provided by an embodiment of the present application;
[0028] Figure 7 It is a flowchart of the retrieval strategy corresponding to the global retrieval type provided by an embodiment of the present application;
[0029] Figure 8 A flowchart of an alternative information processing method provided for an embodiment of the present application;
[0030] Figure 9 A flowchart of a retrieval strategy corresponding to a local retrieval type provided for an embodiment of the present application;
[0031] Figure 10 A flowchart of an alternative information processing method provided for an embodiment of the present application;
[0032] Figure 11 A flowchart of an alternative information processing method provided for an embodiment of the present application;
[0033] Figure 12 A flowchart of an alternative information processing method provided for an embodiment of the present application;
[0034] Figure 13 A structural block diagram of an index database construction provided for an embodiment of the present application;
[0035] Figure 14 A structural diagram of an alternative information processing apparatus provided for an embodiment of the present application;
[0036] Figure 15 A structural diagram of an electronic device provided for an embodiment of the present application. Detailed implementation manners
[0037] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0038] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0039] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase occurs in various places in the specification and is not necessarily referring to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0040] An embodiment of the present application provides a method for information processing, which is applied to an electronic device. Referring to Figure 1 as shown, the method includes the following steps:
[0041] Step 101: Input the query information for the target object into the intent recognition model for intent recognition to obtain the information retrieval type corresponding to the query information.
[0042] Among them, the query information is used to query all relevant log information or partial relevant log information of the target object.
[0043] In the embodiments of the present application, the target object can be understood as a specific entity or object to be queried in the log system. Exemplarily, in the scenario of server log analysis, the target object can be specific, such as the target object can be a server, and the target object can also be abstract, such as a test task or an event. Of course, the target object can also be a specific range, and the range of the target object can be global (such as all relevant logs) or local (such as the logs of a certain day).
[0044] In the embodiments of the present application, the query information can be a request input by the user, used to retrieve log data related to the target object from the log system. The query information usually includes filtering conditions such as the identifier of the target object, time range, event type, etc. The query information determines the granularity of log retrieval, such as all logs or partial logs.
[0045] In one case, the query information can be used to query all relevant log information of the target object. All relevant log information refers to all log information that exactly matches the target object without adding any filtering conditions. In this way, when a comprehensive analysis of the target object is required, querying all relevant log information can provide a complete context for summarizing the analysis. Exemplarily, the query information for a server can be "Please comprehensively analyze all the test results of the r1b43 server and give a summary".
[0046] In another case, the query information can also be partial relevant log information for querying the target object. The partial relevant log information refers to filtering out partial log information from all the relevant logs of the target object according to the filtering conditions in the query information. In this way, when only a specific part of the target object needs to be analyzed, querying the partial relevant log information can effectively improve the efficiency. Exemplarily, the query information for the server can be "Please analyze whether there are any errors in the Power-Cycle test performed by the r1b43 server on July 18, 2022. If you want to get a reward, please give a highly feasible solution."
[0047] In the embodiments of the present application, the information retrieval type is the type of global retrieval or local retrieval for the query information, and the information retrieval type includes the global retrieval type or the local retrieval type.
[0048] In the embodiments of the present application, the intent recognition model is used to parse the query information for the target object input by the user, accurately capture the user's intent through semantic understanding, and map it to the corresponding information retrieval type. In practical applications, this intent recognition model is a dedicated model obtained by performing supervised fine-tuning on a pre-trained large language model (LLM) based on training samples of large-scale log data and optimizing it in combination with a specific structure. Through this process, the model can better adapt to the specific requirements of the log analysis field, so as to achieve accurate recognition and efficient response to the user's query intent. It should be noted that the large language model is an artificial intelligence model based on deep learning, which has powerful language understanding and generation capabilities after being trained with a large amount of text data.
[0049] In an implementable scenario, referring to Figure 2 as shown, after obtaining the query information 201 for the target object, the query information 201 is input into the trained intent recognition model 202. The intent recognition model 202 accurately recognizes the user's intent by performing semantic parsing and intent classification on the query information, and outputs the information retrieval type corresponding to the query information. In this way, the accuracy and efficiency of log retrieval are ensured, providing a clear direction for subsequent log analysis and result generation.
[0050] Step 102: Use the pre-constructed index database to retrieve the query information through the retrieval strategy corresponding to the information retrieval type, and obtain a retrieval result that meets the information length condition required for the input of the first large language model.
[0051] Among them, different information retrieval types correspond to different retrieval strategies.
[0052] In the embodiments of the present application, the index database is a knowledge graph database constructed by preprocessing and structuring the original log data, which is used to efficiently store and retrieve log-related information to support fast querying.
[0053] In the embodiments of the present application, the retrieval strategy is a strategy determined according to the information retrieval type, which is used to extract relevant retrieval results of log data that meet the conditions from the index database. It should be noted that different information retrieval types correspond to different retrieval strategies, that is, the global retrieval type and the local retrieval type respectively correspond to different retrieval strategies.
[0054] In the embodiments of the present application, the first large language model is the core model that infers based on the query information input by the user and generates the final answer. The input requirements of the first large language model include but are not limited to the maximum input information processing length defined by the first large language model. To ensure that the model can efficiently process the input data, the information length of the retrieval result needs to meet the input requirements of the first large language model.
[0055] In the embodiments of the present application, the information length condition corresponding to the input requirements of the first large language model refers to constraining the information length of the retrieval result according to the input limitations of the first large language model during the retrieval process. If the maximum input length of the first large language model is N (N is a positive integer), in one case, if the retrieval result is directly input into the first language model, the total length of the retrieval result shall not exceed the maximum input length N; in another case, if at least part of the retrieval result, query information, and historical record are input into the first large language model, the total length of the retrieval result does not exceed the maximum input length N. In this way, by limiting the information length of the retrieval result through the input requirements of the first large language model, it can effectively avoid the decline in model performance or processing failure caused by too long input data, thus ensuring that the first large language model can efficiently and accurately complete the inference task.
[0056] In one realizable manner, continue to refer to Figure 2 , the retrieval process can be divided into two cases:
[0057] The first case is to use the pre-constructed index database, combined with the retrieval strategy corresponding to the information retrieval type (such as Figure 2 's global retrieval strategy 204 or local retrieval strategy 205), to retrieve the query information input by the user for the target object. Through the exact matching and screening of the retrieval strategy, the retrieval result that meets the information length condition corresponding to the input requirements of the first large language model is finally obtained.
[0058] In the second case, in the presence of session context, first obtain the historical conversation record 203 related to the query information in the current session; then, use the pre-constructed index database and combine the retrieval strategy corresponding to the information retrieval type (such as Figure 2 the global retrieval strategy 204 or the local retrieval strategy 205 in ), and jointly retrieve the query information and the historical conversation record at the same time, and finally obtain a retrieval result that meets the information length condition required for the input of the first large language model. In this way, not only the current query information is considered, but also the context information of the historical conversation is fully utilized, thereby improving the accuracy and relevance of the retrieval result.
[0059] Step 103: Perform a fusion process on the query information and the retrieval result to generate context reference information that meets the input requirements of the first large language model.
[0060] In the embodiment of the present application, the context reference information is information that meets the input requirements of the first large language model, that is, the information length of the context reference information is less than the maximum input length of the first large language model.
[0061] In the embodiment of the present application, the context reference information may be information obtained by integrating, fusing, and optimizing the query information and the retrieval result. The context reference information may also be information obtained by integrating, fusing, and optimizing the query information, the historical conversation record corresponding to the query information, and the retrieval result, and the information length of the context reference information meets the input requirements of the first large language model; of course, in order to guide the first large language model to perform accurate reasoning, the context reference information may also be information obtained by integrating, fusing, and optimizing based on the query information, the historical conversation record corresponding to the query information, the retrieval result, and the prompt of the first large language model's prompt engineering. It can be understood that the context reference information can provide a complete context for the first large language model, helping the model better understand the user's intention and generate accurate answers.
[0062] In an implementable scenario, continue to refer to Figure 2 , after using the pre-constructed index database and retrieving the query information through the retrieval strategy corresponding to the information retrieval type to obtain a retrieval result that meets the information length condition required for the input of the first large language model, perform a fusion process 206 on the query information and the retrieval result, or perform a fusion process 206 on the query information, the retrieval result, and the historical conversation record to generate context reference information 207 that meets the input requirements of the first large language model. In this way, a clear context is provided for the large language model, avoiding reasoning errors caused by information loss, and improving the accuracy and relevance of the answer generated by the model.
[0063] Step 104: Use the first large language model to reason about the context reference information to generate a reply content corresponding to the query information.
[0064] In the embodiments of the present application, continue to refer to Figure 2 As shown, after fusing the query information and the retrieval results to generate context reference information that meets the input requirements of the first large language model, the context reference information 207 generated after the fusion process is input into the first large language model 208. Through the semantic understanding and reasoning capabilities of the first large language model 208, high-quality response content that matches the query information of the target object input by the user is generated, thereby meeting the user's needs and enhancing the interaction experience.
[0065] The embodiments of the present application provide an information processing method. The query information for the target object is input into an intent recognition model for intent recognition to obtain the information retrieval type corresponding to the query information. Among them, the query information is used to query all relevant log information or partial relevant log information of the target object. Using a pre-constructed index database, the query information is retrieved through the retrieval strategy corresponding to the information retrieval type to obtain a retrieval result that meets the information length condition required for input to the first large language model. Among them, different information retrieval types correspond to different retrieval strategies. The query information and the retrieval result are fused to generate context reference information that meets the input requirements of the first large language model. The first large language model is used to reason about the context reference information to generate the response content corresponding to the query information. In this way, first, through intent recognition and dynamic retrieval strategies, it is ensured that the length of the retrieval result always meets the input requirements of the large model, avoiding system operation anomalies caused by overly long input data and enhancing the reliability of the system. Second, different information retrieval types correspond to different retrieval strategies, enabling the system to flexibly respond to diverse query needs, and different retrieval strategies can accurately screen relevant content according to the query intent, reducing interference from irrelevant information, helping the large model better understand the query intent and generate accurate answers, and balancing retrieval accuracy and complexity. Finally, based on the data index database of the knowledge graph, the end-to-end process from user input to the output of the large model can have logical reasoning capabilities, solving the hallucination problem of the large model in the face of unknown fields.
[0066] In some embodiments, before performing step 101 of inputting the query information for the target object into the intent recognition model for intent recognition to obtain the information retrieval type corresponding to the query information, the training process of the intent recognition model is described.
[0067] Step 111, perform supervised fine-tuning on the second large language model using a sample data set.
[0068] In the embodiments of the present application, the sample data set may be log data generated for testing a target object such as a server. The second large language model is an artificial intelligence model based on deep learning. After being trained with a large amount of text data, it has powerful language understanding and generation capabilities.
[0069] In the embodiments of the present application, the sample data set is used to perform supervised fine-tuning on the second large language model to form a second large language model specific to the test field of the target object such as a server, that is, the second large language model after supervised fine-tuning, which improves the generation accuracy and reduces the hallucination rate, enabling the model to more accurately understand the key data and context information of the log information.
[0070] Step 112: Add a binary classification model including a fully connected layer and an activation layer after the output layer of the second large language model after supervised fine-tuning to obtain an improved second large language model.
[0071] Among them, the binary classification model is used to classify the retrieval types of the retrieval information output by the second large language model after supervised fine-tuning.
[0072] In the embodiments of the present application, the binary classification model is used to classify the retrieval types of the retrieval information output by the second large language model. The binary classification model includes a fully connected layer and an activation layer; among them, the fully connected layer is responsible for linearly combining and feature-transforming the data input to the fully connected layer. The activation layer is used to introduce non-linearity, enhance the model's expression ability, and control the output range, thereby generating the probability distribution of each category.
[0073] In practical applications, since it is a binary classification task, the activation layer can use the softmax function. For a given information retrieval type x = [x1, x2 ··· x k , for a binary classification problem, k = 2, and the corresponding softmax output σ(x) i , can be represented by formula (1).
[0074]
[0075] In an implementable scenario, the basic structure of the second large language model includes, but is not limited to, an input layer, an embedding layer, an encoder (based on the Transformer architecture), and an output layer. Refer to Figure 3As shown, first, the second largest language model is supervised and fine-tuned using the sample data set to obtain the second largest language model 302 after supervised fine-tuning. Further, in order to further improve the binary classification performance of the model, after the output layer of the second largest language model 302 after supervised fine-tuning, a binary classification model 303 including a fully connected layer and an activation layer is added to improve the second largest language model after supervised fine-tuning. By introducing the binary classification model and using the fully connected layer and the activation layer to work together, it is possible to further learn from the high-level features extracted by the second largest language model after supervised fine-tuning, thus efficiently completing the binary classification task.
[0076] Step 113: Use each training sample in the training sample set and the sample information retrieval type corresponding to the training sample to perform reinforcement learning fine-tuning on the fully connected layer in the improved second largest language model to obtain an intent recognition model.
[0077] In the embodiment of the present application, the training sample set is a sample set specifically used to train the information retrieval type of the query information for the target object. The training sample set includes multiple training samples, and the sample information retrieval type corresponding to each training sample includes a sample global retrieval type and a sample local retrieval type. Here, the training sample can be problem text information, and the sample information retrieval type is the information retrieval type corresponding to the problem text. Through this structured training data, the model can learn the mapping relationship from the problem to the information retrieval type, thereby improving the accuracy and efficiency of the information retrieval type for the query information.
[0078] In the embodiment of the present application, in order to further improve the performance of the improved second largest language model, a phased fine-tuning strategy is adopted. Exemplarily, in the process of performing reinforcement learning fine-tuning on the improved second largest language model using each training sample in the training sample set and its corresponding sample information retrieval type, first, the parameters of the large language model part in the improved second largest language model are frozen to keep them unchanged during the fine-tuning process. Subsequently, the network parameters of the fully connected layer in the improved second largest language model are fine-tuned by reinforcement learning. In this way, the fully connected layer can be optimized for specific tasks (such as intent recognition) without affecting the original feature extraction ability of the large language model. Finally, after reinforcement learning fine-tuning, the improved second largest language model can more accurately complete the intent recognition task, thus obtaining a trained intent recognition model. This phased fine-tuning strategy not only improves the training efficiency of the model but also ensures the performance improvement of the model in the binary classification task.
[0079] In some embodiments, the retrieval strategy corresponding to the information retrieval type in step 102 is combined Figure 4 for illustration.
[0080] Step 401: Perform vectorization processing on the query information corresponding to the information retrieval type to obtain a query information vector.
[0081] In the embodiments of the present application, performing vectorization processing on the query information corresponding to the information retrieval type can be understood as follows: If the information retrieval type is a global retrieval type, perform vectorization processing on the entire information of the query information to obtain a query information vector; if the information retrieval type is a local retrieval type, extract the entity information at each data granularity in the query information and perform vectorization processing to obtain a query entity information vector, where the query information vector includes the query entity information vector.
[0082] It can be understood that for the vectorization processing of the query information corresponding to the information retrieval type, an embedding model can be used to convert some or all of the text information in the query information into a vector form to obtain a query information vector. This query information vector can capture the semantic information of the text and be represented in the vector space. In this way, through vectorization processing, the query information is converted into a query information vector comparable to the reference vector in the index database, and the correlation between the two is measured in a quantitative manner to achieve precise matching and improve the fit between the retrieval result and the query.
[0083] Step 402: Perform correlation matching between the query information vector and the reference vector of the reference object in the index database to obtain multiple matching results.
[0084] Among them, the matching result includes the correlation score between the query information vector and the reference vector.
[0085] It can be understood that since the index database is a knowledge graph database constructed by preprocessing and structuring the original log data, the reference object can be a subgraph in the index database or an entity in the index database. The reference vector can be a vector representing the summary information of the subgraph or an entity information vector representing the entity information. The present application does not make specific limitations on this.
[0086] In the embodiments of the present application, the matching result includes the correlation score between the query information vector and the reference vector, and the correlation score can be obtained based on the correlation or similarity between the query information vector and the reference vector. It should be noted that various methods can be used to calculate the correlation or similarity, such as cosine similarity, Euclidean distance, or dot product, etc. The embodiments of the present application do not make specific limitations on this.
[0087] Step 403: Screen out the reference objects corresponding to the correlation scores that meet the score conditions from the reference objects.
[0088] Among them, the score condition includes that the correlation score is less than or equal to the target correlation threshold.
[0089] In the embodiments of the present application, the scoring condition is used to select reference objects with high relevance. The scoring condition includes that the relevance score is less than or equal to the target relevance threshold. Among them, the target relevance threshold can be preset or a threshold dynamically determined based on the information length of the retrieval result and the input requirements of the first large language model. The present application does not make specific limitations on this.
[0090] In the embodiments of the present application, first, the relevance or similarity between the query information vector and the reference vectors of each reference object in the index database is calculated to obtain the relevance score between the query information vector and each reference vector. Then, according to the target relevance threshold, the reference objects with a relevance score less than or equal to the threshold are screened out from the reference objects. By setting a reasonable target relevance threshold, the system can effectively filter out reference objects with low relevance, retain reference objects highly relevant to the query information, avoid interference from redundant information, focus on the core relevant content, and thus ensure the quality and accuracy of the retrieval result.
[0091] Step 404: Based on the screened reference objects, obtain a retrieval result that meets the information length condition required for the input of the first large language model.
[0092] In the embodiments of the present application, after screening out the reference objects corresponding to the relevance scores that meet the scoring conditions from the reference objects, based on the screened reference objects, the query information is retrieved to obtain a retrieval result that meets the information length condition required for the input of the first large language model. In this way, based on the screened reference objects, focus on the core relevant content, prevent the input data from being too long and causing abnormal model processing, ensure that the large language model can effectively utilize the retrieval information, improve the system stability and the accuracy of the output result, balance the retrieval precision and complexity, and improve the retrieval efficiency.
[0093] In some embodiments, the determination process of the target relevance threshold in step 403 is combined with Figure 5 for illustration.
[0094] Step 501: Obtain the first retrieval result.
[0095] Among them, the first retrieval result is obtained by performing a single retrieval on the query information vector based on the reference relevance threshold using the retrieval strategy corresponding to the information retrieval type with the index database.
[0096] In the embodiments of the present application, the reference relevance threshold can be a preset initial relevance threshold.
[0097] In an embodiment of the present application, after obtaining the information retrieval type corresponding to the query information, first, perform vectorization processing on the query information corresponding to the information retrieval type to obtain a query information vector; second, perform a correlation match between the query information vector and the reference vectors of the reference objects in the index database to obtain the correlation scores between the query information vector and each reference vector; then, screen out the reference objects corresponding to the correlation scores less than or equal to the reference correlation threshold from the reference objects; finally, based on the screened reference objects, obtain a first retrieval result obtained by performing a first retrieval on the query information vector.
[0098] Step 502: If the information length of the first retrieval result is greater than or equal to the length threshold corresponding to the input requirement of the first large language model, obtain a threshold adjustment coefficient and an information length attenuation coefficient.
[0099] In an embodiment of the present application, the length threshold corresponding to the input requirement of the first large language model can be understood as that for different retrieval strategies, the length thresholds of the retrieval results corresponding to the input requirements of the first large language model may be the same or different.
[0100] In an embodiment of the present application, the threshold adjustment coefficient is the maximum amplitude for controlling the change of the correlation threshold when the information length of the retrieval result exceeds the length threshold.
[0101] In an embodiment of the present application, the information length attenuation coefficient is the rate for controlling the growth rate and slowing down as the gap between the information length of the retrieval result and the length threshold increases.
[0102] It should be noted that the threshold adjustment coefficient and the information length attenuation coefficient corresponding to different retrieval strategies are different.
[0103] Step 503: Based on the information length of the first retrieval result, the length threshold, the threshold adjustment coefficient, and the information length attenuation coefficient, adjust the reference correlation threshold one or more times.
[0104] It can be understood that adjusting the reference correlation threshold once based on the information length of the first retrieval result, the length threshold, the threshold adjustment coefficient, and the information length attenuation coefficient can be achieved through the following process:
[0105] Obtain the first difference between the information length of the first retrieval result and the length threshold; obtain the maximum value between the first difference and the first value such as 0; determine the attenuation rate based on the maximum value and the information length attenuation coefficient; calculate the second difference between the second value such as 1 and the attenuation rate; calculate the product of the second difference and the threshold adjustment coefficient to obtain the threshold change amplitude; based on the threshold change amplitude, perform a single adjustment on the reference relevance threshold to obtain the adjusted reference relevance threshold. If the reference relevance threshold is adjusted multiple times, the obtained adjusted reference relevance threshold can be used as the current reference relevance threshold, and the above steps can be repeated to obtain the next adjusted reference relevance threshold. For this, the embodiments of this application will not be described in detail.
[0106] In an implementable scenario, the adjustment of the reference relevance threshold can be obtained through the following formula (2):
[0107]
[0108] where T adjusted is the adjusted reference relevance threshold, T0 is the reference relevance threshold, L is the information length of the retrieval result, and L taget is the length threshold corresponding to the input requirements of the first large language model, α is the threshold adjustment coefficient, and γ is the information length attenuation coefficient.
[0109] As can be seen from the above, by monitoring the length of the first retrieval result in real time, once it exceeds the length threshold required for the input of the large language model, the reference relevance threshold is dynamically adjusted using the threshold adjustment coefficient and the information length attenuation coefficient, so as to balance the retrieval accuracy and time complexity, and make the subsequent generated intermediate retrieval results adapt to the model input length, avoiding the model being unable to process normally due to excessive data and ensuring the smooth operation of the system.
[0110] Step 504, after each adjustment of the reference relevance threshold, use the index database to perform a secondary retrieval on the query information vector through the retrieval strategy corresponding to the information retrieval type to obtain an intermediate retrieval result.
[0111] Step 505, if the information length of the intermediate retrieval result is less than the length threshold, determine the adjusted reference relevance threshold as the target relevance threshold.
[0112] In the embodiments of the present application, after adjusting the reference relevance threshold one or more times based on the information length of the first retrieval result, the length threshold, the threshold adjustment coefficient, and the information length attenuation coefficient, each time the reference relevance threshold is adjusted, it is necessary to use the index database to retrieve the query information vector again through the retrieval strategy corresponding to the information retrieval type to obtain an intermediate retrieval result; compare the information length of the intermediate retrieval result with the length threshold; if the information length of the intermediate retrieval result is less than the length threshold, at this time, there is no need to adjust the adjusted reference relevance threshold again, and the currently adjusted reference relevance threshold can be determined as the target relevance threshold, so that the relevance threshold remains unchanged and is used to obtain the final retrieval result. In this way, when facing first retrieval results of different lengths, the system can automatically trigger the threshold adjustment process without manual intervention. After one or more adjustments, it can quickly find the intermediate retrieval result that meets the length requirements and determine the target relevance threshold. In this way, the retrieval efficiency is greatly improved, the system instability factors caused by improper data processing are reduced, and the robustness of the system is enhanced.
[0113] In some embodiments, different information retrieval types correspond to different retrieval strategies. If the query information is used to query all relevant log information of the target object, the information retrieval type is the global retrieval type, and the retrieval strategy corresponding to the global retrieval type is combined Figure 6 for description.
[0114] Step 601: Vectorize the query information to obtain a query information vector.
[0115] In the embodiments of the present application, since the query information is used to query all relevant log information of the target object, and the information retrieval type corresponding to the query information is the global retrieval type, that is, a comprehensive query of the target object, it is necessary to vectorize the entire information of the query information to obtain a query information vector.
[0116] Step 602: Perform correlation matching between the query information vector and the summary information vectors of each sub-graph in the index database to obtain multiple first matching results.
[0117] Among them, the first matching result includes the first correlation score between the query information vector and the summary information vector of the sub-graph, and the relevant intermediate response information obtained based on the summary information vector of the sub-graph and the query information vector.
[0118] In the embodiments of the present application, the relevant intermediate response information is the response information obtained by performing information fusion processing on the summary information vector and the query vector.
[0119] In a feasible implementation manner, since the index database is a knowledge graph database constructed by preprocessing and structuring the original log data, refer to Figure 7As shown, in one case, the query information vector 701 can directly perform a relevance match with the summary information vectors 703 of each of the N subgraphs in the index database; in another case, in the presence of session context, a historical dialogue record vector 702 of the historical dialogue records related to the query information in the current session is obtained, and the historical dialogue record vector 702 and the query information vector 701 are subjected to a fusion extraction process to obtain a processed information vector; the processed information vector can perform a relevance match with the summary information vectors 703 of each of the N subgraphs in the index database; thus, a first relevance score between the query information vector and the summary information vector of the subgraph, and a first matching result 704 of the relevant intermediate response information obtained based on the summary information vector of the subgraph and the query information vector are obtained.
[0120] Step 603: From each of the subgraphs in the index database, filter out the subgraphs corresponding to the relevance scores that meet the first score condition.
[0121] Among them, the first score condition includes that the first relevance score is less than or equal to the first relevance threshold, and the target relevance threshold includes the first relevance threshold.
[0122] In the embodiments of the present application, the first score condition is used to select subgraphs with high relevance. The first score condition includes that the first relevance score is less than or equal to the first relevance threshold; among them, the first relevance threshold can be preset or a threshold dynamically determined based on the information length of the retrieval result and the input requirements of the first large language model. For this, the present application does not make specific limitations.
[0123] In the embodiments of the present application, first, the query information vector is subjected to a relevance match with the summary information vectors of each of the subgraphs in the index database to obtain a first relevance score between the query information vector and the summary information vector of the subgraph, and a first matching result of the relevant intermediate response information obtained based on the summary information vector of the subgraph and the query information vector. Then, all the relevant intermediate response information is sorted according to the first relevance score, such as in descending or ascending order, and according to the first relevance threshold, the subgraphs with the first relevance score less than or equal to the threshold are filtered out from each of the subgraphs in the index database. By setting a reasonable first relevance threshold, the system can effectively filter out subgraphs with low relevance and retain subgraphs highly relevant to the query information, thereby ensuring the quality and accuracy of the retrieval result.
[0124] Step 604: Perform a fusion process on the relevant intermediate response information corresponding to the filtered subgraphs to obtain a retrieval result that meets the first information length condition.
[0125] Among them, the first information length condition includes that the information length of the retrieval result is less than or equal to the first length threshold, and the first length threshold is determined based on the maximum input length of the first large language model.
[0126] In the embodiments of the present application, the first information length condition is used to limit the information length of the retrieval result output by the retrieval strategy corresponding to the global retrieval type. The first information length condition includes: the information length of the retrieval result is less than or equal to the first length threshold, and the first length threshold is determined based on the maximum input length of the first large language model.
[0127] In the embodiments of the present application, continue to refer to Figure 7 , after screening out the subgraphs corresponding to the relevance scores that meet the first score condition from each subgraph in the index database, the relevant intermediate response information 704 corresponding to the screened subgraphs is fused to obtain a retrieval result 705 that meets the first information length condition. In this way, by fusing the summary information of the screened subgraphs, it is ensured that the information length after fusion is less than the first length threshold determined based on the maximum input length of the large language model, effectively avoiding the problem that the large language model cannot process due to too long input data, ensuring the stable operation of the system, and at the same time integrating key information to provide high-quality input for the large language model, so that the large language model generates more accurate and effective reply content, improving the user query experience and the overall performance of the system.
[0128] In some embodiments, different information retrieval types correspond to different retrieval strategies. If the query information is used to query some relevant log information of the target object, the information retrieval type is the local retrieval type, and the retrieval strategy corresponding to the local retrieval type will be described in conjunction with Figure 8 for illustration.
[0129] Step 801: Extract the query entity information of the query information and perform vectorization processing on the query entity information to obtain a query entity information vector.
[0130] Among them, the query information vector includes the query entity information vector.
[0131] In the embodiments of the present application, the query entity information includes, but is not limited to, one or more data granularities such as entities, relationships, and covariates.
[0132] In the embodiments of the present application, since the query information is used to query some relevant log information of the target object, the information retrieval type corresponding to the query information is the local retrieval type, that is, a specific part of the target object is queried. Referring to Figure 9 as shown, extract the query entity information 903 at different data granularities in the query information 901, and perform vectorization processing on the extracted query entity information 903 to obtain a query entity information vector.
[0133] Step 802: Perform a relevance match between the query entity information vector and each entity information vector in the index database to obtain multiple second matching results.
[0134] Among them, the second matching result includes the second relevance score between the query entity information vector and the entity information vector.
[0135] In an implementable scenario, since the index database is a knowledge graph database constructed by preprocessing and structuring the original log data, continue to refer to Figure 9 , in one case, the query entity information vector can directly perform a relevance match with each entity information vector in the index database 904; in another case, in the presence of session context, obtain the historical conversation record 902 of the historical conversation records related to the query information in the current session, respectively extract the historical entity information in the historical conversation record and the query entity information in the query information, fuse and vectorize the historical entity information and the query entity information to obtain the processed information vector; the processed information vector can perform a relevance match with each entity information vector in the index database 904; thus, the second relevance score between the query entity information vector and the entity information vector is obtained.
[0136] Step 803: From the entity information vectors in the index database, filter out the entity information vectors corresponding to the second relevance scores that meet the second score condition.
[0137] Among them, the second score condition includes that the second relevance score is less than or equal to the second relevance threshold, and the relevance threshold includes the first relevance threshold.
[0138] In the embodiments of the present application, the second score condition is used to select entity information vectors with high relevance. The second score condition includes that the second relevance score is less than or equal to the second relevance threshold; among them, the second relevance threshold can be preset or a threshold dynamically determined based on the information length of the retrieval result and the input requirements of the first large language model. The present application does not make specific limitations on this.
[0139] In the embodiments of the present application, after performing a relevance match between the query information vector and the summary information vectors of each subgraph in the index database to obtain multiple first matching results, for all entity information vectors, sort them according to the second relevance score in descending or ascending order, and according to the second relevance threshold, filter out the entity information vectors whose second relevance scores are less than or equal to the threshold from the entity information vectors in the index database. By setting a reasonable second relevance threshold, the system can effectively filter out entity information vectors with low relevance and retain entity information vectors highly relevant to the query entity information, thereby ensuring the quality and accuracy of the retrieval results.
[0140] Step 804: Perform fusion processing on the filtered entity information vectors and the query vector to obtain a retrieval result that meets the second information length condition.
[0141] Among them, the second information length condition includes that the information length of the retrieval result is less than or equal to a second length threshold, and the second length threshold is determined based on the maximum input length of the first large language model.
[0142] In the embodiments of the present application, the second information length condition is used to limit the information length of the retrieval result output by the retrieval strategy corresponding to the local retrieval type. The second information length condition includes: the information length of the retrieval result is less than or equal to a second length threshold, and the second length threshold is determined based on the maximum input length of the first large language model. The second length threshold and the first length threshold may be the same or different.
[0143] In the embodiments of the present application, continue to refer to Figure 9 , after filtering out the entity information vectors corresponding to the second relevance score that meets the second score condition from the entity information vectors in the index database, perform fusion processing on the filtered entity information vectors and the query vector to obtain a retrieval result that meets the second information length condition. In this way, perform fusion processing on the filtered entity information vectors to ensure that the information length after fusion is less than the second length threshold determined based on the maximum input length of the large language model, effectively avoiding the problem that the large language model cannot process due to too long input data, ensuring the stable operation of the system, and at the same time integrating key information to provide high-quality input for the large language model, so that the large language model generates more accurate and effective response content, improving the user query experience and the overall performance of the system.
[0144] In some embodiments, step 103 is described for the process of performing fusion processing on the query information and the retrieval result to generate context reference information that meets the input requirements of the first large language model.
[0145] Step 131: Perform fusion processing on the query information and the retrieval result to generate first context reference information that does not meet the input requirements of the first large language model.
[0146] In the embodiments of the present application, the first context reference information may be information obtained by integrating, fusing, and optimizing the query information and the retrieval results. The first context reference information may also be information obtained by integrating, fusing, and optimizing the query information, the historical conversation record corresponding to the query information, and the retrieval results, and the information length of the context reference information meets the input requirements of the first large language model. Of course, in order to guide the first large language model to perform accurate reasoning, the first context reference information may also be information obtained by integrating, fusing, and optimizing based on the query information, the historical conversation record corresponding to the query information, and the retrieval results, in combination with the prompt engineering prompt of the first large language model. It can be understood that the first context reference information can provide a complete context for the first large language model, helping the model better understand the user's intention and generate accurate answers.
[0147] In the embodiments of the present application, the first context reference information is information that does not meet the input requirements of the first large language model, that is, the information length of the first context reference information is greater than the maximum input length of the first large language model.
[0148] It can be understood that after fusing the query information and the retrieval results, the obtained first context reference information is still too long and exceeds the maximum input length of the first large language model. At this time, the feedback mechanism can be used to perform information fusion compression on the query information and the retrieval results so that the compressed information is less than the maximum input length of the first large language model.
[0149] Step 132: Perform information compression on the query information and the retrieval results without changing the semantics to obtain context reference information that meets the input requirements.
[0150] In the embodiments of the present application, after generating the first context reference information that does not meet the input requirements of the first large language model, the feedback mechanism can be used to generate feedback information for performing information fusion compression on the query information and the retrieval results again. The feedback information may include information compression on the query information and the retrieval results without changing the semantics. The feedback information may also include the information length of the output after fusion compression. In this way, it is ensured that the final input data does not exceed the maximum input length of the model, enhancing the reliability of the system.
[0151] In some embodiments, the construction process of the index database in step 102 is described in combination with Figure 10 for illustration.
[0152] Step 1001: Preprocess the obtained log data, and extract entity information from the preprocessed log data to obtain an entity node set and the entity information of each entity node.
[0153] Among them, the entity information includes entity attribute information, entity relationships related to entity nodes, and the weights of associated entities.
[0154] In the embodiments of this application, preprocessing the log data can be understood as splitting the log data into text chunks to complete data loading. Here, the size of the text chunks can be set based on a fixed length or based on modularity; that is, the size of the text chunks can be set according to specific circumstances. Since each text chunk generally contains hundreds to thousands of tokens, subsequent analysis of the text chunks will not result in low processing efficiency due to overly long text.
[0155] In the embodiments of this application, the entity information includes the relevant information of each entity node. When preprocessing the log data to split it into multiple log data text chunks; further, entity information extraction can be performed on each log data text chunk. When performing entity information extraction on each log data text chunk, large language models can be used for analysis to extract the entities in the log data text chunk that represent nodes, the entity relationships that represent the relationships between nodes, and the covariates that represent entity attribute information. In this way, using large language models to effectively process log data text chunks avoids potential performance bottlenecks or memory limitations that large language models may encounter when processing large amounts of text. Exemplarily, in the process of extracting entities, an event is used as an entity, and an example of the entity of the entity information is shown in Table 1. In Table 1, the event is the entity, and the attribute information of the event, which is the covariate of the entity, includes metadata and a timestamp. Among them, the metadata includes, but is not limited to, event type, description information, result, and server identifier; the timestamp is the time.
[0156]
[0157] Table 1
[0158] Step 1002: Construct a knowledge graph based on the entity node set and the entity information of the entity nodes, and partition the knowledge graph. Classify the entity nodes with the same representative meaning in the knowledge graph into one subgraph to obtain multiple subgraphs.
[0159] It can be understood that a knowledge graph is a semantic network that shows the relationships between entities and the attributes of entities in a graphical way. Among them, each entity corresponds to a node, and the entity relationships are represented by edges. The metadata of the nodes is used to represent entity attributes or entity information, including but not limited to event type, description information, result, server identifier, and time, etc.; the metadata of the edges contains detailed information about the relationships, such as the relationship type and the description of the relationship.
[0160] In the embodiments of the present application, a complete knowledge graph can be constructed based on the entity node set, i.e., nodes, and the entity information of the entity nodes.
[0161] It can be understood that entity nodes with the same representative meaning in the knowledge graph can be understood as that some entity nodes existing in the knowledge graph are similar or relevant in a certain specific dimension. A subgraph is a part of the knowledge graph, which is a graph structure composed of some nodes selected from the entire knowledge graph and the edges between these nodes.
[0162] In the embodiments of the present application, when partitioning the knowledge graph, multiple graph structures formed after classifying entity nodes with the same representative meaning can be used, and each graph structure corresponds to a subgraph; in this way, the entire knowledge graph is split into multiple subgraphs, and the entity nodes within each subgraph have certain commonalities.
[0163] It should be noted that partitioning the knowledge graph and classifying entity nodes with the same representative meaning in the knowledge graph into a subgraph can be obtained by using a community detection algorithm or other clustering and classification methods.
[0164] Step 1003: Input each subgraph into the third large language model to generate summary information of the subgraph.
[0165] In the embodiments of the present application, the third large language model is an artificial intelligence model based on deep learning. After being trained with a large amount of text data, it has powerful language understanding and generation capabilities.
[0166] In the embodiments of the present application, each subgraph is input into the third large language model, and summary information of each subgraph is generated through the third large language model.
[0167] Step 1004: Respectively perform vectorization processing on the summary information of each subgraph and the entity information of each entity node to obtain vector indexes of each subgraph and vector indexes of each entity node.
[0168] Among them, the vector index of the subgraph includes a summary information vector, and the vector index of the entity node includes an entity information vector.
[0169] Step 1005: Construct an index database based on the subgraph, the vector index of the subgraph, the entity node, and the vector index of the entity node.
[0170] In the embodiments of the present application, an embedding model can be used to perform vectorization processing on the summary information of each sub-graph and the entity information of each entity node respectively, so as to obtain the vector index of each sub-graph and the vector index of each entity node; further, an index database is constructed based on the sub-graph, the vector index of the sub-graph, the entity node, and the vector index of the entity node. In this way, in applications related to knowledge graphs, when a specific sub-graph or entity node needs to be queried, the similarity between the vector of the query content and the vector index generated in the index database can be calculated to quickly locate the relevant sub-graph or entity node. Compared with the traditional text-matching-based retrieval method, the retrieval based on vector index can better capture semantic similarity. Even if the query statement is not exactly the same as the stored text, as long as the semantics are similar, relevant results can be retrieved, greatly improving the efficiency and accuracy of retrieval.
[0171] In some embodiments, the process of dividing the knowledge graph in step 1002, classifying entity nodes with the same representative meaning in the knowledge graph into one sub-graph, and obtaining multiple sub-graphs is described in conjunction with Figure 11 the following.
[0172] Step 1101: Create a community for each entity node in the knowledge graph to obtain a community set and obtain the global modularity of the knowledge graph.
[0173] In the embodiments of the present application, the global modularity is used to measure the quality of community division.
[0174] In a realizable scenario, for the knowledge graph A = (V, E), the knowledge graph can be an undirected graph or a weighted graph. Among them, V represents the set of entity nodes, such as V = {v1, v2, v3, v4, v5, v6}, and E represents the relationship between entity nodes, that is, the edge set, such as E = {(v1, v2), (v2, v3), (v3, v4), (v4, v5), (v5, v6), (v1, v6)}. Here, the knowledge graph A can be represented by an adjacency matrix or an adjacency edge list, and the elements in the matrix or list are used to represent the connection weights between entity nodes.
[0175] Here, for each entity node v ∈ V in the knowledge graph A, a community can be created for it, that is, C v = {v}, so as to obtain multiple communities, forming a community set, such as C = {{C1 = {v1}}, {C2 = {v2}}, {C3 = {v3}}, {C4 = {v4}}, {C5 = {v5}}, {C6 = {v6}}}, and initialize the community set. Further, according to the following formula (3), combined with the knowledge graph A, calculate the global modularity of the knowledge graph.
[0176]
[0177] Among them, M represents the global modularity of the knowledge graph; A ij represents the connection weight, i.e., the weight, between entity node i and entity node j. It should be noted that for an undirected graph, A ij = A ji ; D i represents the degree of entity node i, i.e., the sum of the weights of the edges connected to entity node i; δ(c i , c j ) is an indicator function used to determine whether the community c i to which entity node i belongs and the community c j to which entity node j belongs are in the same community. If so, it is 1; otherwise, it is 0.
[0178] Step 1102: Add all entity nodes to an initially empty queue.
[0179] It can be understood that a queue is a special linear data structure. The data in the queue follows the first-in, first-out rule. The data or elements that enter the queue first will be processed or taken out first, and the data or elements that enter later need to wait. It should be noted that when new data or elements are added to the queue, they enter the queue from the tail of the queue, that is, they are placed behind the last data or element in the queue and become the new tail data or element. When data or elements are removed from the queue, they are taken out from the head of the queue, that is, the data or element located at the head of the queue is removed or processed, and then the head pointer of the queue moves one position backward, pointing to the second data or element in the original queue, making it the new head of the queue.
[0180] In the embodiment of the present application, continuing with the above example for illustration, an empty queue Q is established and initialized in advance, and all nodes v ∈ V in the knowledge graph A are added to the empty queue Q, that is, the six entity nodes v1 - v6 are added to an empty queue Q. At this time, there are six entity nodes v1, v2, v3, v4, v5, v6 in the queue Q, preparing for subsequent operations.
[0181] Step 1103: In each iteration, if the queue is not empty, based on the global modularity, obtain the first global modularity gain by moving the first entity node in the queue to any first community.
[0182] Among them, the first community is other communities in the community set except the community where the first entity node is located.
[0183] It can be understood that the first entity node can be the entity node located at the head of the queue.
[0184] In the embodiment of the present application, the first community is other communities in the community set except the community where the first entity node is located.
[0185] In the embodiments of the present application, the first global modularity gain can be understood as the change in modularity before and after the movement of the first entity node.
[0186] In the embodiments of the present application, during each iteration, it is determined whether the queue Q is empty. If the queue Q is not empty, the first entity node located at the head of the queue can be moved to any first community outside the current community where it is located; when the first entity node is moved to a first community each time, the first global modularity of the current knowledge graph is calculated through the above formula (3), so as to obtain multiple first global modularities corresponding to when the first entity node is moved to different first communities; further, a difference operation is performed between the multiple first global modularities and the global modularity respectively, so as to obtain multiple first global modularity gains corresponding to when the first entity node is moved to different first communities.
[0187] In an implementable scenario, continuing with the above example for illustration, it is assumed that during the first iteration, the queue Q is not empty. The first entity node v1 dequeued from the queue Q is moved to other communities such as C2, C3, C4, C5, and C6 except the community C1, so as to obtain five first global modularities; the five first global modularities are respectively subtracted from the global modularity to obtain the first global modularity gain ΔM corresponding to when the first entity node v1 is moved to C2, C3, C4, C5, and C6.
[0188] Step 1104: Determine the first target community corresponding to the first global modularity gain that satisfies the first gain condition, and move the first entity node to the first target community to update the community set.
[0189] In the embodiments of the present application, the first gain condition may be that the first global modularity gain is greater than the second gain threshold. The second gain threshold may be used to filter out the first global modularity gain greater than 0. For example, the second gain threshold may be 0. The second gain threshold may also be used to filter out the maximum first global modularity gain. Of course, the second gain threshold may also be used to filter out the maximum first global modularity gain greater than 0. In this way, it is determined through the second gain value that the quality of the community after the current division is better than the quality of the community in the previous division.
[0190] In the embodiments of the present application, determine the remaining communities in the community set except the community where the first entity node is located, determine the community in the remaining communities where the first global modularity gain is greater than the first gain threshold, and determine this community as the first target community. Further, move the first entity node to the first target community, thereby realizing the update of the community set.
[0191] In an implementable scenario, continuing with the above example for illustration, if it is determined that the first entity node v1 moves to community C6 and the first global modularity gain ΔM > 0, it means that moving the first entity node v1 to community C6 will improve the overall quality of community division. At this time, the first entity node v1 is divided from the original community C1 to community C6 to update the community set. In this way, deciding the movement of entity nodes based on the global modularity gain can gradually gather the entity nodes in the knowledge graph into more appropriate communities, and the obtained community division is more in line with the connection relationship and semantic connection between nodes, improving the rationality of community division.
[0192] Step 1105: Add the associated entity nodes of the first entity node to the queue and update the global modularity.
[0193] In the embodiments of the present application, the associated entity nodes may be the adjacent entity nodes of the first entity node, and the adjacent entity node and the first entity node are not in the same community and the first target community.
[0194] In an implementable scenario, the adjacent entity nodes not in the same community as the first entity node v1 are v2 and v6. Since the adjacent entity node v6 is in the first target community, the adjacent entity node v2 that is not in the first target community is added to the queue Q as an associated entity node to prepare for subsequent analysis of the situation where the associated entity node moves to other communities. At the same time, the global modularity M of the current knowledge graph is updated to reflect the modularity situation under the current new community division. In this way, each time an entity node is moved, its associated entity nodes are added to the queue, ensuring that the nodes related to the entity node can be processed in a timely manner, fully considering the association relationship between nodes, and making the community division result more holistic and coherent.
[0195] Step 1106: Stop the iteration when the iteration condition is met.
[0196] Among them, the divided community corresponds to a subgraph.
[0197] Among them, the iteration condition includes at least one of the following: the queue is empty and no community optimization can be performed, the global modularity gain is less than the first gain threshold, and the preset maximum number of iterations is reached.
[0198] In the embodiments of the present application, after the knowledge graph is divided when the iteration condition is met, each divided community obtained corresponds to a subgraph.
[0199] In the embodiments of the present application, the queue is empty and no community optimization can be performed can be understood as the queue is empty and there is no community that can be optimized.
[0200] In the embodiments of the present application, it can be understood that the global modularity gain being less than the first gain threshold means that the global modularity gain between the global modularity obtained in the current iteration round and the global modularity obtained in the previous iteration round is less than the first gain threshold. The first gain threshold can be determined based on the actual situation.
[0201] In the embodiments of the present application, the preset maximum number of iterations can be the total number of iterations determined based on the actual situation. For example, the maximum number of iterations can be 20, or it can be other values. In this regard, the embodiments of the present application do not make specific limitations.
[0202] In the embodiments of the present application, after each iteration is completed, it is determined whether one or more of the current queue, the divided communities, the global modularity gain, and the number of iterations up to the current iteration meet the iteration conditions. If the iteration conditions are not met, the next entity node is taken out from the queue, and steps 1103 to 1105 are repeated until the iteration conditions are met; if the iteration conditions are met, the iteration is stopped, thereby obtaining the divided communities obtained by partitioning the knowledge graph, and each divided community corresponds to a subgraph. In this way, by setting multiple iteration stop conditions, such as the queue being empty and no community optimization can be performed, the global modularity gain being less than the threshold, reaching the maximum number of iterations, etc., the algorithm can flexibly stop according to the actual situation, avoiding unnecessary calculations, and at the same time ensuring the quality of the community partitioning results.
[0203] In some embodiments, after step 1105 updates the global modularity, as shown in 12, the following steps can also be performed.
[0204] Step 1201: After iterating a first number of times, use the union-find set to detect whether there are non-connected communities in the current community set.
[0205] In the embodiments of the present application, the first number can be greater than 0 and less than the maximum number of iterations. Exemplarily, the first number can be 3, 4, 5, or other values. In this regard, the present application does not make specific limitations.
[0206] In the embodiments of the present application, a non-connected community is an independently existing community, that is, it has no connection relationship or connectivity with other communities.
[0207] In the embodiments of the present application, after the first number of iterations, the union-find set can be used to detect the connectivity of each community in the community set and determine whether there are unconnected communities. In this way, by introducing a detection and processing mechanism for unconnected communities after a specific number of iterations, the algorithm can adapt to the complex situations that may occur in the knowledge graph during the iterative process. Secondly, using the union-find set to detect unconnected communities can timely discover the disconnection problems existing in the community division of the knowledge graph, ensure that there is good connectivity between the entity nodes within each community in the finally obtained community set, the community structure is more reasonable, avoid the emergence of isolated communities, and improve the quality of community division.
[0208] Step 1202: If there are unconnected communities, based on the global modularity corresponding to the current iteration round, obtain the second global modularity gain of assigning the unconnected communities to any second community.
[0209] Wherein, the second community is other connected communities adjacent to the unconnected communities in the community set.
[0210] In the embodiments of the present application, the second global modularity gain can be understood as the change in modularity before and after moving all the entity nodes in the unconnected communities to other communities.
[0211] In the embodiments of the present application, when it is detected by the union-find set that there are unconnected communities in the current community set, the unconnected communities are assigned to any second community outside the current community where they are located, that is, all the entity nodes in the unconnected communities are simultaneously moved to the second community adjacent to the unconnected communities; when each unconnected community is assigned to the second community, calculate the second global modularity of the current knowledge graph, so as to obtain multiple second global modularities corresponding to when the unconnected communities are assigned to different second communities; further, perform a difference operation on the multiple second global modularities and the global modularity corresponding to the current iteration round respectively, to obtain multiple second global modularity gains corresponding to when the unconnected communities are assigned to different second communities. In this way, after determining that there are unconnected communities, calculate the second global modularity gain of assigning them to adjacent connected communities based on the global modularity, and select the community corresponding to the maximum gain for assignment. In this way, while solving the connectivity problem, the global modularity of the knowledge graph is further optimized.
[0212] Step 1203: If multiple second global modularity gains meet the second gain condition, assign the unconnected communities to the second target community corresponding to the maximum second global modularity gain among the multiple second global modularity gains.
[0213] In the embodiments of the present application, the second gain condition includes: the maximum second global modularity gain among the multiple second global modularity gains is greater than the third gain threshold, and the third gain threshold can be 0.
[0214] In the embodiments of the present application, the maximum second global modularity gain among multiple second global modularity gains is obtained. If the maximum second global modularity gain is greater than the third gain threshold, it can be known that the multiple second global modularity gains meet the second gain condition. At this time, the second target community corresponding to the maximum second global modularity gain among the multiple second global modularity gains is determined, and the unconnected community is assigned to the second target community, that is, all entity nodes in the unconnected community are moved to the second target community, thereby achieving the connectivity of the community. In this way, for the second global modularity gains that meet the second gain condition, the community corresponding to the maximum gain is selected to assign the unconnected community, avoiding the randomness and uncertainty in dealing with the assignment of unconnected communities, ensuring the stability and repeatability of the community division process, and improving the quality of community division.
[0215] Step 1204: If multiple second global modularity gains meet the third gain condition, assign the unconnected community to the third target community corresponding to the minimum second global modularity gain among the multiple second global modularity gains.
[0216] In the embodiments of the present application, the third gain condition includes: some of the second global modularity gains among the multiple second global modularity gains are less than the fourth gain threshold, and the fourth gain threshold can be 0.
[0217] In the embodiments of the present application, each second global modularity gain among the multiple second global modularity gains is compared with the fourth gain threshold. If there is a second global modularity gain less than the fourth gain threshold among the multiple second global modularity gains, it can be known that the multiple second global modularity gains meet the third gain condition. At this time, the third target community corresponding to the minimum second global modularity gain among the multiple second global modularity gains is determined, and the unconnected community is assigned to the third target community, that is, all entity nodes in the unconnected community are moved to the third target community. In this way, the connectivity of the unconnected community is achieved.
[0218] Step 1205: For any second entity node in the third target community, obtain the third global modularity gain of moving the second entity node to any third community based on the global modularity corresponding to the current iteration round.
[0219] Wherein, the third community is other connected communities in the community set except the third target community.
[0220] In the embodiments of the present application, the third global modularity gain can be understood as the change in modularity before and after the movement of the second entity node.
[0221] In an embodiment of the present application, for any second entity node in the third target community, move the second entity node to any third community in the community set except the third target community; when moving the second entity node to each third community, calculate the third global modularity of the current knowledge graph, so as to obtain multiple third global modularities corresponding to when the second entity node moves to different third communities; further, perform a difference operation on the multiple third global modularities and the global modularity corresponding to the current iteration round to obtain multiple third global modularity gains corresponding to when the second entity node moves to different third communities, so as to obtain multiple third global modularity gains corresponding to each second entity node among all second entity nodes in the third target community.
[0222] Step 1206: Determine the fourth target community corresponding to the third global modularity gain that meets the fourth gain condition, and move the second entity node to the fourth target community.
[0223] In an embodiment of the present application, the fourth gain condition may be that the third global modularity gain is greater than the fifth gain threshold, and the fifth gain threshold may be 0, so as to determine that the community quality after the current division is better than that of the previous division.
[0224] In an embodiment of the present application, for any second entity node in the third target community, determine the remaining communities in the community set except the third target community, determine the connected community with a third global modularity gain greater than the fifth gain threshold from the remaining communities, and determine the connected community as the fourth target community, and move the second entity node to the fourth target community until there is no second entity node in the third target community; in this way, the connectivity of non-connected communities is achieved, and at the same time, the connected communities are optimized. In this way, the community structure is fine-tuned within a local range, further optimizing the global modularity. By continuously adjusting the attribution of entity nodes, the connection relationship between communities in the knowledge graph becomes more reasonable, enhancing the cohesion within the community and the distinction between communities.
[0225] It should be noted that step 1203 and steps 1204 to 1206 can be executed synchronously, step 1203 can also be executed before steps 1204 to 1206, and step 1203 can also be executed after steps 1204 to 1206. In this regard, the present application does not make specific limitations.
[0226] Here, in an implementable scenario, an information processing method provided in an embodiment of the present application is described.
[0227] With the emergence of large artificial intelligence models, there are already feasible system solutions in the field of log analysis to achieve batch parsing of massive logs. However, most logs belong to specific domains, and the Retrieval-Augmented Generation (RAG) technology needs to be used to solve the hallucination phenomenon of large models when facing unknown data. In this parsing process, based on the large model method, an index database will be constructed with text data. Through the user's Query query, relevant documents will be retrieved from the database. Finally, the original Query query and the retrieved data will be merged and input into the large model as input, and finally the answer will be generated. However, in this process, an inevitable problem will occur, that is, the retrieved data is too long, which will cause the final input to exceed the maximum input processing length of the large model.
[0228] To solve the above problems, the embodiments of the present application provide an information processing method, and the process is as follows:
[0229] Step A1: Construct a knowledge graph database.
[0230] Here, the knowledge graph database corresponds to the above-mentioned index database, and the process of constructing the knowledge graph database is as follows:
[0231] Refer to Figure 13 As shown, first, the log data, that is, the source log file 1301, is split into text blocks to complete data loading 1302. Then, for each text block, with the help of a large language model analysis, entities and relationships in each text block are extracted 1303, such as extracting information such as nodes (entities), edges (relationships), and covariates (metadata, timestamps) in the text block to construct a knowledge graph. Each entity and relationship will also have several attribute information. In the process of extracting entities, an event is used as an entity, and the entity information is shown in Table 1.
[0232] After all entities and relationships are constructed to obtain the knowledge graph, the community detection algorithm is used to divide the entire knowledge graph into multiple subgraphs 1304, and the large language model analysis is used to generate summary information for each subgraph 1305, that is, all information (including entities, subgraphs, metadata, and relationships) is used to construct a vector index through an embedding model, thereby completing the construction of the knowledge graph database 1306. In this way, the subsequent retrieval process is accelerated. It should be noted that the above-mentioned large language model is a large language model specific to the server test field formed by using the log data output from the CSP server test to perform supervised fine-tuning on the established large language model in related technologies.
[0233] Here, the community detection algorithm is improved, and the algorithm implementation process is as follows.
[0234] Input:
[0235] A = (V, E): An undirected graph or weighted graph, where V is the set of nodes (corresponding to the above-mentioned set of entity nodes), and E is the set of edges (corresponding to the above-mentioned node relationships); A: The adjacency matrix or edge list representation of the graph, used to calculate the connection weights between nodes.
[0236] Output:
[0237] C: The result of community (subgraph) partitioning, which assigns the nodes in the graph to different communities.
[0238] Algorithm steps:
[0239] Step 1-1: For each node v ∈ V, create a community, i.e., C v = {v}, thus obtaining the community set C and initializing this community set C;
[0240] Step 1-2: According to the above formula (3), combined with the graph A, calculate the global modularity M of A.
[0241] Step 2-1: Randomly add all nodes v ∈ V and initialize an empty queue Q;
[0242] Step 2-2: If Q is not empty, execute the following steps:
[0243] 1) Dequeue Q and remove node i;
[0244] 2) For the communities in the community set other than the community where node i is located C i calculate the global modularity gain of moving node v i to each community, that is, the change in global modularity before and after moving ΔM. If ΔM > 0, update the community set C, add all adjacent nodes (not in the same community as v i affected by the movement of v) to the queue Q, and update the global modularity M; i
[0245] 3) After each iteration of a certain number of times, check the connectivity of all communities C i ∈ C through the union-find set. If there are non-connected communities, optimize them in the following order:
[0246] Under the maximum connectivity gain, reassign the communities of non-connected components to other adjacent connected communities;
[0247] For some communities with negative connectivity gain in the previous step, use the local search algorithm for the nodes in them, evaluate with the modularity gain, and partition them into other locally optimal communities.
[0248] Step 3-1: Repeat Step 2-2 until one or more of the following conditions are met: the queue Q is empty and no further community optimization can be performed, the change in the global modularity M is less than a preset threshold ΔQ threshold ; the preset maximum number of iterations E is reached max .
[0249] Step 3-2: Output the final community partition result C.
[0250] Step A2: Retrieve the query input by the user based on the knowledge graph database.
[0251] First step, perform intent recognition on the query input by the user.
[0252] Here, the query information for the server test log can be divided into two categories. One is global, for example: Please help me comprehensively analyze all the test results of the r1b43 server and give a summary. The other is local, for example: Please help me analyze whether there are errors in the Power-Cycle test executed by the r1b43 server on July 18, 2022. If you want to get a reward, please give a highly feasible solution. For the above two types of queries, a version of the intent recognition classification model based on the large language model (corresponding to the above-mentioned intent recognition model) is trained. The query information is input into the intent recognition classification model to determine the query category corresponding to the query information (corresponding to the above-mentioned information retrieval type). Here, the model structure is as Figure 3 shown.
[0253] Among them, the large language model is composed of the basic structures of the traditional input layer, embedding layer, and Transformer encoder. By freezing the parameters of the large language model, the parameters of the fully connected layer and / or activation function are updated during the training process. Since it is a binary classification task, the softmax function is used as the activation function. For the given vector x = [x1, x2 ··· x k , for the binary classification problem, k = 2, and the corresponding softmax output σ(x) i is defined as shown in formula (1).
[0254] It should be noted that this intent recognition classification model is to use the log data produced by the CSP server test to perform supervised fine-tuning on the established large language model in the related technology to form a large language model specific to the server test field.
[0255] Second step, after completing the intent recognition task under the above method and obtaining the query category corresponding to the query information, adapt the first-order retriever (corresponding to the retrieval strategy corresponding to the above-mentioned global retrieval type) or the second-order retriever (corresponding to the retrieval strategy corresponding to the above-mentioned local retrieval type).
[0256] Among them, for the first-order retriever process, as Figure 7 shown, the first-order retriever performs a relevance match on the abstract information of each sub-graph in the knowledge graph index database according to the query information input by the user and the historical conversation (corresponding to the above historical conversation record), and outputs a relevant intermediate response with a score; then, based on a given relevance threshold, filters out the sub-graphs with scores less than the relevance threshold, obtains the sub-graphs relevant to the query information input by the user, and then summarizes and reviews the abstract information of the sub-graphs relevant to the query information as the final output of this first-order retriever to obtain the retrieval result.
[0257] Among them, for the second-order retriever process, as Figure 9 shown, the second-order retriever extracts entity information according to the query information input by the user and the historical conversation (corresponding to the above historical conversation record), performs a relevance match between the entity information and the entity information in the knowledge graph index database (including entities, metadata, and relationships at various data granularities), and obtains a relevance score; then, based on a given relevance threshold, filters out the entity information in the knowledge graph index database with a relevance score less than the relevance threshold, and then summarizes and reviews the remaining entity information as the final output of this second-order retriever to obtain the retrieval result.
[0258] It should be noted that the relevant matching mentioned above adopts the scheme of using a large language model according to the prompt "limit the relevance value range to -1 to 1 and give the relevance score of two texts" to complete the matching.
[0259] It should also be noted that there is a dynamic threshold feedback mechanism added to the input and output ends of the retriever. Because both have relevance matching, evaluate the output length of the retriever, and dynamically adjust the relevance threshold based on the maximum input length of the final large language model, so as to limit the output of the retriever and avoid the problem of too long input when finally accessing the log analysis large language model, so as to limit the retrieval accuracy and time complexity. Here, in order to formalize the dynamic threshold feedback mechanism, a non-linear exponential decay adjustment strategy is proposed, which can be represented by the above formula (2). It should be emphasized that when the length L of the retriever output is less than the length threshold L taget corresponding to the input requirement of the final access log analysis large language model, the adjusted relevance threshold T adjusted = T0, that is, the relevance threshold remains unchanged.
[0260] It should be noted that a self-regressive algorithm is adopted in the retrieval stage and a large language model is used to complete text compression to ensure that the final input data does not exceed the maximum input length of the model, so as to complete the final output of the retriever.
[0261] In the third step, combine the query information input by the user, the historical conversation, and the retrieval results. Through prompt engineering, synthesize the context reference information for input into the final large language model according to the prompt template, where the length of the context reference information is less than the maximum input length of the final large language model.
[0262] Here, combine the query information input by the user, the historical conversation, and the retrieval results. Through prompt engineering, synthesize the context reference information according to the prompt template. Continue to refer to Figure 2 , there is also a feedback mechanism here. If the context reference information is still too long and exceeds the maximum input length of the model, use the large language model to extract the summary of the three components of the query information, historical conversation, and retrieval results again for semantic-invariant compression, and in the summary prompt, specify the length of the summary output, and finally obtain the context reference information with a length less than the maximum input length. It should be noted that an autoregressive algorithm is used in this stage and the large language model is used to complete text compression to ensure that the final input data does not exceed the maximum input length of the model.
[0263] As can be seen from the above, through this solution, based on the data index database of the knowledge graph, the end-to-end process from user input to the output of the large model can have logical reasoning ability, solve the hallucination problem of the large model facing unknown fields; balance the retrieval accuracy and complexity, and solve the phenomenon that the input data of the large model exceeds the model.
[0264] The embodiments of the present application provide an information processing device. Refer to Figure 14 As shown, the information processing device 14 includes:
[0265] An intention recognition module 1401, configured to input the query information for the target object into an intention recognition model for intention recognition to obtain the information retrieval type corresponding to the query information; wherein, the query information is used to query all relevant log information or part of the relevant log information of the target object;
[0266] A retrieval module 1402, configured to use the pre-constructed index database to retrieve the query information through the retrieval strategy corresponding to the information retrieval type to obtain a retrieval result that meets the information length condition of the first large language model input requirement; wherein, different information retrieval types correspond to different retrieval strategies;
[0267] A fusion processing module 1403, configured to perform fusion processing on the query information and the retrieval result to generate context reference information that meets the first large language model input requirement;
[0268] An inference module 1404, configured to use the first large language model to perform inference on the context reference information to generate a reply content corresponding to the query information.
[0269] Based on the foregoing embodiments, an embodiment of the present application provides an electronic device. Referring to Figure 15 as shown, the electronic device 15 includes: a processor 1501 and a memory 1502. Among them,
[0270] the memory 1502 stores a computer program that can run on the processor 1501;
[0271] the processor 1501 executes the computer program stored in the memory 1502 to implement the following steps:
[0272] Input the query information for the target object into the intent recognition model for intent recognition to obtain the information retrieval type corresponding to the query information; wherein, the query information is used to query all relevant log information or partial relevant log information of the target object;
[0273] Utilize the pre-constructed index database, through the retrieval strategy corresponding to the information retrieval type, retrieve the query information to obtain a retrieval result that meets the information length condition required for input to the first large language model; wherein, different information retrieval types correspond to different retrieval strategies;
[0274] Perform a fusion process on the query information and the retrieval result to generate context reference information that meets the input requirements of the first large language model;
[0275] Utilize the first large language model to reason about the context reference information to generate a reply content corresponding to the query information.
[0276] The method provided by the embodiment of the present application can be directly embodied as a combination of software modules executed by the processor 1501. The software modules can be located in a storage medium, and the storage medium is located in the memory 1502. The processor 1501 reads the executable instructions included in the software modules in the memory 1502 and combines the necessary hardware to complete the method provided by the embodiment of the present application.
[0277] As an example, the processor 1501 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0278] It should be noted that for the specific implementation process of the steps executed by the processor in this embodiment, reference can be made to the steps in the method provided by the foregoing embodiments, and details are not described herein again.
[0279] An embodiment of the present application provides a storage medium that stores a computer program. When the computer program is executed by at least one processor, the steps in the method provided in the above embodiment are implemented, which will not be elaborated here.
[0280] An embodiment of the present application provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps in the method provided in the above embodiment are implemented, which will not be elaborated here.
[0281] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0282] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0283] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0284] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0285] The above is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. An information processing method, the method comprising: Inputting query information for a target object into an intent recognition model for intent recognition to obtain an information retrieval type corresponding to the query information; wherein, the query information is used to query all relevant log information or partial relevant log information of the target object; Using a pre-constructed index database, retrieving the query information through a retrieval strategy corresponding to the information retrieval type to obtain a retrieval result that meets the information length condition corresponding to the input requirements of the first large language model; wherein, different information retrieval types correspond to different retrieval strategies; Performing fusion processing on the query information and the retrieval result to generate context reference information that meets the input requirements of the first large language model; Using the first large language model to reason about the context reference information to generate a reply content corresponding to the query information.
2. The method according to claim 1, the method further comprising: Supervising and fine-tuning a second large language model using a sample data set; Adding a binary classification model including a fully connected layer and an activation layer after the output layer of the second large language model after supervision and fine-tuning to obtain an improved second large language model; wherein, the binary classification model is used to classify the retrieval type of the retrieval information output by the second large language model after supervision and fine-tuning; Using each training sample in the training sample set and the sample information retrieval type corresponding to the training sample to perform reinforcement learning fine-tuning on the fully connected layer in the improved second large language model to obtain the intent recognition model.
3. The method according to claim 1, the retrieval strategy corresponding to the information retrieval type includes: Performing vectorization processing on the query information corresponding to the information retrieval type to obtain a query information vector; Performing correlation matching between the query information vector and the reference vector of the reference object in the index database to obtain a plurality of matching results; wherein, the matching result includes the correlation score between the query information vector and the reference vector; Screening out the reference object corresponding to the correlation score that meets the score condition from the reference objects; wherein, the score condition includes that the correlation score is less than or equal to a target correlation threshold; Based on the screened reference objects, obtaining a retrieval result that meets the information length condition corresponding to the input requirements of the first large language model.
4. The method according to claim 3, the determination process of the target correlation threshold includes: Obtaining a first retrieval result, wherein the first retrieval result is obtained by using the index database, through the retrieval strategy corresponding to the information retrieval type, and performing a single retrieval on the query information vector based on a reference correlation threshold; If the information length of the first retrieval result is greater than or equal to the length threshold corresponding to the input requirements of the first large language model, obtaining a threshold adjustment coefficient and an information length attenuation coefficient; Based on the information length of the first retrieval result, the length threshold, the threshold adjustment coefficient, and the information length attenuation coefficient, adjusting the reference correlation threshold one or more times; After adjusting each pair of the reference relevance thresholds, using the index database, through the retrieval strategy corresponding to the information retrieval type, the query information vector is retrieved again to obtain an intermediate retrieval result; If the information length of the intermediate retrieval result is less than the length threshold, the adjusted reference relevance threshold is determined as the target relevance threshold.
5. The method according to claim 3, wherein the query information is used to query all relevant log information of the target object, the information retrieval type is a global retrieval type, and the retrieval strategy corresponding to the global retrieval type includes: Performing vectorization processing on the query information to obtain the query information vector; Performing relevance matching between the query information vector and the abstract information vectors of each sub-graph in the index database to obtain a plurality of first matching results; wherein, the first matching result includes a first relevance score between the query information vector and the abstract information vector of the sub-graph, and relevant intermediate response information obtained based on the abstract information vector of the sub-graph and the query information vector; Filtering out the sub-graphs corresponding to the relevance scores that meet the first score condition from each sub-graph in the index database, wherein the first score condition includes that the first relevance score is less than or equal to a first relevance threshold, and the relevance threshold includes the first relevance threshold; Performing fusion processing on the relevant intermediate response information corresponding to the filtered sub-graphs to obtain a retrieval result that meets the first information length condition; wherein, the first information length condition includes that the information length of the retrieval result is less than or equal to a first length threshold, and the first length threshold is determined based on the maximum input length of the first large language model.
6. The method according to claim 3, wherein the query information is used to query partial relevant log information of the target object, the information retrieval type is a local retrieval type, and the retrieval strategy corresponding to the local retrieval type includes: Extracting the query entity information of the query information and performing vectorization processing on the query entity information to obtain a query entity information vector; The query information vector includes the query entity information vector; Performing relevance matching between the query entity information vector and each entity information vector in the index database to obtain a plurality of second matching results; wherein, the second matching result includes a second relevance score between the query entity vector and the entity information vector; Filtering out the entity information vectors corresponding to the second relevance scores that meet the second score condition from the entity information vectors in the index database; wherein, the second score condition includes that the second relevance score is less than or equal to a second relevance threshold, and the target relevance threshold includes the first relevance threshold; Performing fusion processing on the filtered entity information vectors to obtain a retrieval result that meets the second information length condition; wherein, the second information length condition includes that the information length of the retrieval result is less than or equal to a second length threshold, and the second length threshold is determined based on the maximum input length of the first large language model.
7. The method according to any one of claims 1 to 6, wherein the fusion processing of the query information and the retrieval results to generate context reference information that meets the input requirements of the first large language model includes: Performing fusion processing on the query information and the retrieval results to generate first context reference information that does not meet the input requirements of the first large language model; Performing information compression on the query information and the retrieval results without semantic change to obtain context reference information that meets the input requirements.
8. The method according to any one of claims 1 to 6, wherein the construction process of the index database includes: Preprocessing the obtained log data, and extracting entity information from the preprocessed log data to obtain a set of entity nodes and the entity information of each entity node, wherein the entity information includes entity attribute information, entity relationships related to the entity node, and weights of associated entities; Constructing a knowledge graph based on the set of entity nodes and the entity information of the entity nodes, and partitioning the knowledge graph, classifying entity nodes with the same representative meaning in the knowledge graph into one subgraph to obtain multiple subgraphs; Inputting each of the subgraphs into a third large language model to generate summary information of the subgraph; Performing vectorization processing on the summary information of each subgraph and the entity information of each entity node respectively to obtain a vector index of each subgraph and a vector index of each entity node, wherein the vector index of the subgraph includes a summary information vector, and the vector index of the entity node includes an entity information vector; Constructing the index database based on the subgraphs, the vector indexes of the subgraphs, the entity nodes, and the vector indexes of the entity nodes.
9. The method according to claim 8, wherein the partitioning of the knowledge graph, classifying entity nodes with the same representative meaning in the knowledge graph into one subgraph to obtain multiple subgraphs, includes: Creating a community for each entity node in the knowledge graph to obtain a set of communities, and obtaining the global modularity of the knowledge graph; Adding all entity nodes to an initially empty queue; In each iteration process, if the queue is not empty, based on the global modularity, obtaining a first global modularity gain for moving a first entity node in the queue to any first community; the first community is other communities in the set of communities except the community where the first entity node is located; Determining a first target community corresponding to the first global modularity gain that meets the first gain condition, moving the first entity node to the first target community, and updating the set of communities; Adding the associated entity nodes of the first entity node to the queue and updating the global modularity; Stopping the iteration when the iteration condition is met; wherein the partitioned community corresponds to a subgraph, and the iteration condition includes at least one of the following: the queue is empty and no community optimization can be performed, the global modularity gain is less than the first gain threshold, and a preset maximum number of iterations is reached.
10. The method according to claim 9, after updating the global modularity, the method includes: After iterating a first number of times, detecting whether there is a non-connected community in the current community set through the union-find set; If there is the non-connected community, obtaining a second global modularity gain for allocating the non-connected community to any second community based on the global modularity corresponding to the current iteration round; wherein, the second community is other connected communities adjacent to the non-connected community in the community set; If multiple second global modularity gains meet the second gain condition, allocating the non-connected community to the second target community corresponding to the maximum second global modularity gain among the multiple second global modularity gains.
11. The method according to claim 10, the method further includes: If multiple second global modularity gains meet the third gain condition, allocating the non-connected community to the third target community corresponding to the minimum second global modularity gain among the multiple second global modularity gains; For any second entity node in the third target community, obtaining a third global modularity gain for moving the second entity node to any third community based on the global modularity corresponding to the current iteration round; The third community is other connected communities in the community set except the third target community; Determining the fourth target community corresponding to the third global modularity gain that meets the fourth gain condition, and moving the second entity node to the fourth target community.
Citation Information
Cited By
Enhanced retrieval generation optimization method fusing knowledge graph
CN121350280A
An enhanced retrieval generation optimization method based on knowledge graph
CN121350280B
Multi-document question and answer method and device based on artificial intelligence and knowledge graph and medium
CN121388179A