Intelligent question answering method and device based on large language model and computer equipment

By segmenting and enhancing the knowledge documents of specific fields, generating high-dimensional vector representations, and combining large language models to achieve dynamic answer generation, the existing fine-tuning large language model method solves the problem of high development costs and inability to deal with rapidly changing business needs in the field of power IT services, achieving efficient, dynamic and accurate intelligent question-and-answer effects.

CN120086328APending Publication Date: 2025-06-03CYG SUNRI CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510088351.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing fine-tuning large language model methods have problems such as high development costs, long update cycles, inability to deal with rapidly changing business needs in real time, and decision-making errors caused by model illusions in the field of power IT services.

Method used

By dividing a specific domain knowledge document into semantic independent text blocks, semantic enhancement processing is performed, and high-dimensional vector representations are generated through the embedding model and stored in the vector database. The query statements entered by the user are converted into query vectors through the embedding model, and the relevant text block collection is retrieved using vector similarity calculation, and the input of the large language model is generated in combination with the prompt template to achieve dynamic answer generation.

Benefits of technology

It effectively improves the timeliness and accuracy of the intelligent question-and-answer system, reduces the dependence on fine-tuning of large language models, reduces computing resource consumption, and improves the scalability and flexibility of the question-and-answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086328A_ABST
    Figure CN120086328A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of natural language processing, and provides an intelligent question and answer method based on a large language model, which comprises the following steps: segmenting a knowledge document in a specific field into a plurality of text blocks with independent semantics, and performing semantic enhancement processing on the text blocks in combination with knowledge characteristics in the specific field; generating a high-dimensional vector representation for each text block through an embedded model, and storing the high-dimensional vector representation in a vector database; a query statement input by a user is obtained, the query statement is converted into a query vector through an embedded model, and a text block set related to the query vector is retrieved from a vector database through vector similarity calculation; reordering the text blocks in the text block set according to correlation priorities queried by the user, generating input of a large language model in combination with a prompt template and the reordered text blocks, and transmitting the input to the large language model to perform answer generation processing; and returning an answer generated by the large language model to the user. Therefore, the timeliness and accuracy of intelligent questions and answers in the specific field are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of natural language processing, and particularly relates to an intelligent question-answering method, device, and computer device based on a large language model. Background Art

[0002] Existing power IT services are facing increasingly complex operation and maintenance requirements and diverse service requirements from business departments. There are numerous systems and applications involved in power IT operation and maintenance work, and the maintenance workload of business documents and configuration operations is large and the updates are not timely. For new employees, traditional training methods such as video tutorials and course content are updated slowly and at high cost, and cannot effectively solve operation problems in a timely manner. With the popularization of the Large Language Model (LLM), a power IT service question-answering system based on the large language model has emerged. This question-answering system can quickly understand users' questions about power IT services and respond in a timely manner. Moreover, the large language model can also understand the intention according to the context, achieve more natural and accurate interaction, and greatly improve the efficiency and accuracy of the question-answering system.

[0003] However, although existing methods for fine-tuning large language models can provide customized question-answering capabilities for specific scenarios, there are still the following disadvantages: First, the fine-tuning process requires a large amount of labeled data and high computing resources, resulting in high development costs and long update cycles; Second, the fine-tuned model depends on static data and cannot respond to rapidly changing business needs in real time, resulting in a lagging knowledge base and inability to handle new query scenarios; Moreover, due to the existence of model hallucinations, in a complex power IT operation and maintenance environment, the fine-tuned model often cannot accurately respond to the actual situation, resulting in decision-making errors or inability to adapt flexibly. Therefore, traditional methods for fine-tuning large language models are difficult to meet the question-answering needs of specific fields such as efficient, dynamic, and accurate power IT services. Summary of the Invention

[0004] The embodiments of this application provide an intelligent question-answering method, device, and computer device based on a large language model, which improve the timeliness and accuracy of intelligent question-answering in specific fields.

[0005] In a first aspect, the embodiments of this application provide an intelligent question-answering method based on a large language model, including:

[0006] Segment a specific domain knowledge document into multiple semantically independent text blocks, and perform semantic enhancement processing on the multiple text blocks in combination with the knowledge characteristics of the specific domain;

[0007] Generate a high-dimensional vector representation for each text block through an embedding model, and store the generated vectors in a vector database;

[0008] Obtain the query statement input by the user, convert the query statement into a query vector through an embedding model, and retrieve a set of text blocks related to the above query vector from the vector database using vector similarity calculation;

[0009] Re-sort the text blocks in the above set of text blocks according to the relevance priority to the user's query, combine the prompt template with the re-sorted text blocks to generate the input of the large language model, and pass it to the large language model for answer generation processing;

[0010] Return the answer generated by the large language model to the above user.

[0011] In a possible implementation manner of the first aspect, the specific domain knowledge document is segmented into multiple semantically independent text blocks, including:

[0012] Adopt a dynamic multi-granularity segmentation strategy to segment the specific domain knowledge document into multiple semantically independent text blocks; the above dynamic multi-granularity segmentation strategy includes: through semantic analysis technology, segment the document content according to the paragraph level, sentence semantic similarity, and context logical relationship.

[0013] In a possible implementation manner of the first aspect, the semantic enhancement processing includes: by analyzing the semantic relevance between text blocks, dynamically supplementing the missing context information.

[0014] In a possible implementation manner of the first aspect, generating a high-dimensional vector representation for each text block through an embedding model, including:

[0015] Combining at least one general pre-trained language model with the above specific domain fine-tuned language model to generate a high-dimensional vector representation for each text block.

[0016] In a possible implementation manner of the first aspect, the metric method for vector similarity calculation is selected in the following way:

[0017] Select the metric method according to the task characteristics and data distribution.

[0018] In a possible implementation manner of the first aspect, the method further includes: combining the search results of at least one open-source search engine, and improving the retrieval recall rate through cross-validation and fusion ranking.

[0019] In a possible implementation manner of the first aspect, the prompt template is obtained in the following way:

[0020] According to the characteristics of the above query statement, dynamically match the corresponding prompt template through keyword extraction and corpus distribution analysis.

[0021] In the second aspect, an intelligent question-answering device based on a large language model provided by an embodiment of the present application includes:

[0022] A text block generation module for splitting a specific domain knowledge document into multiple semantically independent text blocks and performing semantic enhancement processing on the multiple text blocks in combination with the knowledge characteristics of the specific domain;

[0023] A vector generation module for generating high-dimensional vector representations for each text block through an embedding model and storing the generated vectors in a vector database;

[0024] A text block query module for obtaining a query statement input by a user, converting the query statement into a query vector through an embedding model, and retrieving a set of text blocks related to the query vector from the vector database using vector similarity calculation;

[0025] An answer generation module for reordering the text blocks in the set of text blocks according to the relevance priority to the user query, generating an input for a large language model in combination with a prompt template and the reordered text blocks, and passing it to the large language model for answer generation processing;

[0026] An answer return module for returning the answer generated by the large language model to the above user.

[0027] In a third aspect, an embodiment of the present application provides a computer device, including:

[0028] A memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the method according to any one of the first aspect when executing the computer program.

[0029] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method according to any one of the first aspect.

[0030] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a computer device, it causes the computer device to execute the method according to any one of the first aspect.

[0031] The beneficial effects of the embodiments of the present application compared with the prior art are:

[0032] By dynamically retrieving a specific domain knowledge document, generating an input for a large language model based on the retrieval result, and then obtaining the answer generation result of the large language model, the present application can effectively avoid the disadvantages of traditional large language model fine-tuning relying on static data, high computing resources, and large language model hallucinations. Moreover, based on the present application, the answer can be updated in a timely manner by updating the specific domain knowledge document without fine-tuning the large language model, effectively improving the timeliness and accuracy of the answer, and also greatly enhancing the scalability and flexibility of the question-answering system.

[0033] It can be understood that for the beneficial effects of the above-mentioned second to fifth aspects, reference can be made to the relevant descriptions in the above-mentioned first aspect, and details will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 is a flowchart of an existing method for implementing question and answer by fine-tuning a large language model;

[0036] Figure 2 is a schematic flowchart of an intelligent question and answer method based on a large language model provided by an embodiment of the present application;

[0037] Figure 3 is a schematic flowchart of an intelligent question and answer method based on a large language model provided by another embodiment of the present application;

[0038] Figure 4 is a schematic structural diagram of an intelligent question and answer device based on a large language model provided by an embodiment of the present application;

[0039] Figure 5 is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0041] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0042] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0043] As used in the specification of this application and the appended claims, the term "if" may be construed as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".

[0044] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0045] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0046] To facilitate understanding of the improvements of this application over the prior art, the prior art will first be described by way of example as follows:

[0047] Figure 1 A flowchart showing an existing method for implementing question answering by fine-tuning a large language model is shown. As Figure 1 shown, the process specifically includes: collecting a large amount of text data from a large-scale dataset as the basic data for model pre-training; then, performing large-scale unsupervised training on the large language model through the pre-training stage to enable the model to master a wide range of general language knowledge; subsequently, fine-tuning the pre-trained model in combination with a dataset in a specific domain, and adjusting the model parameters through supervised learning to adapt to the task requirements of a specific scenario, thereby generating a fine-tuned large language model. In actual use, when a user inputs a query, the system will call the fine-tuned large language model, understand and reason about the user's query, generate a targeted answer, and finally return the generated answer to the user.

[0048] It can be seen that the existing methods for fine-tuning large language models have the following deficiencies: on the one hand, they rely on static training data and cannot dynamically update knowledge, making it difficult to meet the rapidly changing domain requirements; on the other hand, fine-tuning requires a large amount of computing resources and labeled data, with high development costs and poor scalability. Therefore, the existing methods for fine-tuning large language models are difficult to meet the requirements of efficient, dynamic, and accurate question answering in specific fields such as power IT services.

[0049] To address the above problems, this application combines an improved retrieval-enhanced generation technique with a large language model for intelligent question answering in specific fields, which not only improves the timeliness and accuracy of information but also greatly enhances the scalability and flexibility of the question answering system. It should be noted that this application is applicable to a variety of specific fields and can be used for a single field or a combination of multiple fields. The following mainly takes power IT services as an example for illustration, but this field should not be regarded as a limitation on the scope of application of this application.

[0050] The technical solutions in the embodiments of this application will be described in detail below.

[0051] Figure 2 The flowchart of the intelligent question answering method based on a large language model provided by this application is shown. In some embodiments, this method can be applied to a computer device. As Figure 2 shown, the process includes:

[0052] Step 201, split a specific domain knowledge document into multiple semantically independent text blocks, and perform semantic enhancement processing on the text blocks in combination with the knowledge characteristics of the specific domain.

[0053] A specific domain knowledge document is a document containing relevant knowledge in a specific domain, which can usually provide in-depth details and professional guidance. It should be understood that there are various ways to obtain this specific domain knowledge document in actual applications. Exemplarily, when this method is applied to a question answering system in power IT services, the document uploaded to the local knowledge base of power IT services can be used as the specific domain knowledge document. In one embodiment, before splitting the document, it can be first cleaned and normalized to improve the standardization and reliability of the data.

[0054] In one embodiment, in the text chunking operation, a dynamic multi-granularity chunking strategy can be adopted to split the document into text blocks. Specifically, the dynamic multi-granularity chunking strategy is as follows: through semantic analysis technology, the document content is split according to the paragraph level, sentence semantic similarity, and context logical relationship. The dynamic multi-granularity chunking strategy can ensure that the text blocks after splitting can express clear context semantics while having a moderate length (convenient for calculation). For example, when processing power IT operation and maintenance documents, the document is split into clear operation step blocks, configuration description blocks, and fault troubleshooting blocks to meet the needs of operation and maintenance personnel for quick positioning and learning.

[0055] In one embodiment, the semantic enhancement processing includes a cross-block semantic reinforcement mechanism. Specifically, the cross-block semantic reinforcement mechanism includes: dynamically supplementing missing context information by analyzing the semantic relevance between text blocks. Through the cross-semantic reinforcement mechanism, the problem of possible semantic disconnection between adjacent text blocks can be effectively solved. For example, when an operation and maintenance personnel queries solutions for a certain type of fault, the system will automatically supplement missing configuration parameters or operation precautions according to adjacent text blocks to improve the applicability and integrity of the document segmentation results.

[0056] In one embodiment, after text chunking, vector generation optimization association can also be performed. Specifically, the chunked text is directly linked to the vector generation step of the embedding model to ensure that the segmentation strategy can improve the representation quality of the vectors, thereby optimizing the relevance and efficiency of subsequent retrieval.

[0057] Through the combination of text segmentation and semantic enhancement in step 201, the segmentation results can not only meet the requirements of subsequent high-quality vector generation, but also improve the flexibility and adaptability of the system in multi-domain applications.

[0058] Step 202, generate high-dimensional vector representations for each text block through an embedding model, and store the generated vectors in a vector database.

[0059] In one embodiment, to solve the problem that the embedding model may not be able to capture deep semantic information, a hybrid embedding mechanism can be introduced. The hybrid embedding mechanism generates richer and more accurate text representations by combining different embedding techniques or models. In an example, the hybrid embedding mechanism specifically includes: combining at least one general pre-trained language model (such as BERT, GPT) with a fine-tuned language model in a specific domain (such as the power IT operation and maintenance domain) to improve the accuracy of semantic vector representations. For example, for some operation instruction documents, the fine-tuned model can more accurately capture the semantic association between "voltage parameter adjustment" and "load monitoring".

[0060] In one embodiment, during the text block generation process of step 201 or the query statement processing of step 203, a semantic expansion model based on the attention mechanism can be used to deeply expand the text to improve the coverage and semantic consistency of the embedding representation. For example, when a user searches for "switch tripping fault", the system can associate with content related to "overload protection setting" or "short circuit fault troubleshooting" and perform semantic expansion.

[0061] Step 203, obtain the query statement input by the user, convert the query statement into a query vector through an embedding model, and retrieve a set of text blocks related to the query vector from the vector database using vector similarity calculation.

[0062] In one embodiment, the search results of at least one open-source search engine can also be combined, and the retrieval recall rate can be improved through cross-validation and fusion ranking to ensure that as many query-related documents as possible are recalled. Specifically, the open-source search engine can be the existing Elasticsearch; cross-validation is a model evaluation technique that divides the dataset into multiple subsets, trains and tests the model multiple times to obtain a more reliable performance evaluation; fusion ranking is a method of combining multiple sorted result sets into a single result set, and by fusing multiple sorted results, the retrieval performance and the robustness of the results can be improved.

[0063] In one embodiment, when calculating the vector similarity, a suitable metric method can be dynamically selected according to the task characteristics and data distribution. Exemplarily, for sparse vectors (such as word-level embeddings), cosine similarity is used to measure the directional similarity of vectors; for dense vectors (such as sentence-level embeddings), Euclidean distance is used to measure the proximity between specific values; for tasks that need to measure the strength of correlation (such as sorting problems), statistical methods such as Pearson correlation coefficient are used. In one example, a dynamic selection mechanism based on a feedback mechanism can also be adopted, which specifically includes: by real-time evaluating the retrieval recall rate and accuracy, dynamically switching the vector similarity metric method to ensure the continuous optimization of the retrieval effect.

[0064] Step 204: Re-rank the text blocks in the text block set according to the relevance priority to the user query, combine the prompt template with the re-ranked text blocks to generate the input of the large language model, and pass it to the large language model for answer generation processing.

[0065] The relevance priority between the text block and the user query can be determined by combining multiple factors, such as keyword matching degree, semantic similarity, etc. In one embodiment, the relevance scores corresponding to different factors can be set to assist in determining the relevance priority.

[0066] In one embodiment, in the text block screening in step 203 and the text block re-ranking in step 204, a clustering pre-screening mechanism can be introduced, which specifically includes: constructing a text block pre-screening module based on clustering and semantic association, performing multiple rounds of screening on the potential retrieval results, and trying to make the text blocks related to the query in a specific domain (such as the power IT scenario) appear with a higher priority.

[0067] In one embodiment, in order to improve the quality of the answers generated by the large language model, a multi-template strategy and / or a fine-tuning strategy can be introduced.

[0068] Among them, the multi-template strategy can achieve dynamic selection of prompt templates. A prompt template is a pre-set template for guiding the large language model to understand the nature and requirements of a task. Usually, a prompt template will clearly specify factors such as the role the model needs to play (such as an expert user), the task (such as answering questions), and the requirements for the output (such as a comprehensive and informative answer), etc. The multi-template strategy specifically includes: according to the characteristics of the query statement, dynamically matching the corresponding prompt template through keyword extraction and corpus distribution analysis. For example, select a fault troubleshooting template for the "power dispatching" scenario and a step description template for the "configuration operation" scenario. Among them, corpus distribution analysis refers to the statistical and analysis of text data in the corpus to understand the distribution of features such as vocabulary, syntax, and semantics in the corpus, which helps to understand the language characteristics of a specific domain.

[0069] The fine-tuning strategy specifically includes: using a large-scale operation and maintenance Q&A corpus in a specific domain (such as the power IT domain) to fine-tune the large language model, expanding the diversity of training data through random sampling and data augmentation (such as synonym replacement, sentence pattern transformation), optimizing the performance of the generation model in a specific domain, and reducing generation bias.

[0070] In the answer generation stage, in one embodiment, the user query context and retrieval results can be combined to optimize the answer generation result in multiple rounds to ensure the accuracy and domain relevance of the generated answer.

[0071] Step 205, return the answer generated by the large language model to the user.

[0072] In one embodiment, the large language model generates an answer text according to the input content and returns the answer text to the user.

[0073] So far, the description of Figure 2 the shown process is completed.

[0074] To facilitate a better understanding of the main process of this application, Figure 3 a schematic flowchart of an intelligent question-answering method based on a large language model provided by another embodiment of this application is shown, which exemplarily shows the main process of this application in a more intuitive way.

[0075] This application can effectively avoid the disadvantages of traditional large language model fine-tuning relying on static data, high computing resources, and large language model hallucinations by dynamically retrieving knowledge documents in a specific domain, generating the input of the large language model according to the retrieval results, and then obtaining the answer generation result of the large language model. And based on this application, the answer can be updated in a timely manner by updating the knowledge documents in a specific domain without fine-tuning the large language model, effectively improving the timeliness and accuracy of the answer, and also greatly enhancing the scalability and flexibility of the question-answering system.

[0076] To verify the effectiveness of the intelligent question - answering method of the large - language model provided in this application, the Llama 3.18b Instruct large - language model was selected for experimental testing. After testing, the comparison between the method provided in the embodiments of this application and the traditional large - language model fine - tuning method is as follows:

[0077] The adopted solution Accuracy rate Average response time Computing resource consumption Large language model fine-tuning 87.47% 1.42s High Embodiments of the present application 96.21% 1.09s Low

[0078] According to the test results, compared with the traditional large - language model fine - tuning method, the method provided in the embodiments of this application has an accuracy improvement of 8.74 percentage points, the average response time is shortened by 23.24%, and the consumption of computing resources is further reduced.

[0079] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0080] Corresponding to the intelligent question - answering method based on the large - language model described in the above embodiments, Figure 4 The structural block diagram of the intelligent question - answering device based on the large - language model provided in the embodiments of this application is shown. For the sake of convenience of description, only the parts related to the embodiments of this application are shown.

[0081] Refer to Figure 4 , the device includes:

[0082] A text - block generation module 401, configured to split a specific - domain knowledge document into multiple semantically independent text blocks, and perform semantic enhancement processing on the multiple text blocks in combination with the knowledge characteristics of the specific domain;

[0083] A vector generation module 402, configured to generate high - dimensional vector representations for each text block through an embedding model, and store the generated vectors in a vector database;

[0084] A text - block query module 403, configured to obtain a query statement input by a user, convert the query statement into a query vector through an embedding model, and retrieve a set of text blocks related to the query vector from the vector database by using vector similarity calculation;

[0085] An answer generation module 404, configured to re - sort the text blocks in the text - block set according to the relevance priority with the user's query, combine a prompt template with the re - sorted text blocks to generate an input for the large - language model, and pass it to the large - language model for answer generation processing;

[0086] An answer return module 405, configured to return the answer generated by the large - language model to the user.

[0087] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned modules, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought about, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0088] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0089] The embodiments of the present application also provide a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, and the processor implements the steps in any of the above method embodiments when executing the computer program.

[0090] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and the computer program, when executed by a processor, can implement the steps in the above method embodiments.

[0091] The embodiments of the present application provide a computer program product, which, when running on a computer device, enables the computer device to implement the steps in the above method embodiments when executed.

[0092] Figure 5 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 5 shown, the computer device of this embodiment includes: at least one processor 50 ( Figure 5 only one is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, and the processor 50 implements the steps in any of the above visual programming method embodiments when executing the computer program 52.

[0093] The computer device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 merely examples of the computer device, which do not constitute a limitation on the computer device, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0094] The so-called processor 50 may be a central processing unit (CPU), and the processor 50 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0095] The memory 51 may be an internal storage unit of the computer device in some embodiments, such as the hard disk or memory of the computer device. The memory 51 may also be an external storage device of the computer device in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory 51 may also include both the internal storage unit and the external storage device of the computer device. The memory 51 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program, etc. The memory 51 may also be used to temporarily store data that has been output or will be output.

[0096] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the device / computer equipment, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0097] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0098] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0099] In the embodiments provided in this application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0100] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0101] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. An intelligent question answering method based on a large language model, characterized in that: include: Segmenting a domain-specific knowledge document into a plurality of semantically independent text blocks, and performing semantic enhancement processing on the plurality of text blocks in combination with the knowledge characteristics of the domain-specific knowledge; Generate a high-dimensional vector representation for each text block through the embedding model, and store the generated vector in the vector database; Obtaining a query statement input by a user, converting the query statement into a query vector through an embedding model, and retrieving a set of text blocks related to the query vector from a vector database using vector similarity calculation; Reordering the text blocks in the text block set according to the relevance priority to the user query, combining the prompt template with the reordered text blocks to generate input for the large language model, and passing it to the large language model for answer generation processing; The answer generated by the large language model is returned to the user.

2. The method according to claim 1, characterized in that The segmentation of the domain-specific knowledge document into multiple semantically independent text blocks includes: A dynamic multi-granularity segmentation strategy is adopted to segment domain-specific knowledge documents into multiple semantically independent text blocks; the dynamic multi-granularity segmentation strategy includes: using semantic analysis technology to segment document content into text blocks according to paragraph level, sentence semantic similarity and contextual logical relationship.

3. The method according to claim 1, characterized in that The semantic enhancement process includes: dynamically supplementing missing context information by analyzing the semantic relevance between text blocks.

4. The method according to claim 1, characterized in that The method of generating a high-dimensional vector representation for each text block through an embedding model includes: At least one general pre-trained language model is combined with the domain-specific fine-tuned language model to generate a high-dimensional vector representation for each text block.

5. The method according to claim 1, characterized in that The measurement method for calculating the vector similarity is selected in the following way: Select the measurement method based on the task characteristics and data distribution.

6. The method according to claim 1, characterized in that The method further comprises: Combine the search results of at least one open source search engine and improve the retrieval recall rate through cross-validation and fusion ranking.

7. The method according to claim 1, characterized in that The prompt template is obtained in the following way: According to the characteristics of the query sentence, the corresponding prompt template is dynamically matched through keyword extraction and corpus distribution analysis.

8. An intelligent question-answering device based on a large language model, characterized in that: include: A text block generation module, used to segment a specific domain knowledge document into a plurality of semantically independent text blocks, and perform semantic enhancement processing on the plurality of text blocks in combination with the knowledge characteristics of the specific domain; A vector generation module is used to generate a high-dimensional vector representation for each text block through an embedding model and store the generated vector in a vector database; A text block query module is used to obtain a query statement input by a user, convert the query statement into a query vector through an embedding model, and retrieve a text block set related to the query vector from a vector database using vector similarity calculation; An answer generation module is used to reorder the text blocks in the text block set according to the relevance priority to the user query, combine the prompt template and the reordered text blocks to generate input for the large language model, and pass it to the large language model for answer generation processing; The answer returning module is used to return the answer generated by the large language model to the user.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that When the computer program product is executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Natural language query method and device

    CN120723888A

  • Chemical enterprise knowledge base construction method and device, electronic equipment and storage medium

    CN120822595A