Method, device and equipment for generating reply information based on large language model
By performing event classification and argument extraction on user question text, combining semantic vectors and event categories, candidate documents in specific domain document question-and-answer systems are recalled, which solves the problem of low accuracy of document recall in the existing technology, and achieves higher quality response information generation.
Patent Information
- Application Number
- CN202311596336.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively obtain multi-dimensional information about user problems in document question-and-answer systems, resulting in low accuracy of document recall.
By performing event classification and argument extraction on the user's problem text, multi-dimensional information of the problem is obtained; combining semantic vectors, event categories and argument information, candidate documents are recalled in document libraries in specific fields, and target documents are determined based on event categories and quality evaluation information.
Improve the accuracy of document recall in the domain-specific document Q&A system to ensure the quality and relevance of reply information.
Smart Images

Figure CN120067243A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to the fields of document retrieval, natural language processing, and large language models. Specifically, it relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating reply information based on a large language model. Background Art
[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] As a key technology in the field of natural language, document question answering has a wide range of application prospects, such as intelligent customer service, search engines, blog systems, and local knowledge bases in specific fields. For example, in the financial field, it is often used for document question answering in vertical fields related to asset investment consulting and research report writing.
[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating reply information based on a large language model.
[0006] According to one aspect of the present disclosure, there is provided a method for generating reply information based on a large language model, including: in response to receiving a question text from a user, obtaining a semantic vector of the question text and event information related to a specific field, where the event information includes an event category queried by the question text and at least one argument information in the question text; obtaining a plurality of candidate documents in a document library of the specific field based on at least two of the semantic vector of the question text, at least one argument information, and the event category; for a candidate document among the plurality of candidate documents, determining quality evaluation information of the candidate document based on the event category; and determining at least one target document among the plurality of candidate documents based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document, so as to obtain reply information for replying to the question text based on the at least one target document.
[0007] According to another aspect of the present disclosure, there is provided a device for generating reply information based on a large language model, including: a first acquisition unit configured to, in response to receiving a question text of a user, acquire a semantic vector of the question text and event information related to a specific domain, where the event information includes an event category queried by the question text and at least one argument information in the question text; a second acquisition unit configured to acquire a plurality of candidate documents in a document library of the specific domain based on at least two of the semantic vector of the question text, at least one argument information, and the event category; a first determination unit configured to determine quality evaluation information of a candidate document for the candidate documents among the plurality of candidate documents based on the event category; and a second determination unit configured to determine at least one target document among the plurality of candidate documents based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document, so as to obtain reply information for replying to the question text based on the at least one target document.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned method for generating reply information based on a large language model.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the above-mentioned method for generating reply information based on a large language model.
[0010] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the above-mentioned method for generating reply information based on a large language model.
[0011] According to one or more embodiments of the present disclosure, it is possible to classify events and extract arguments from the question text of the user to obtain multi-dimensional information of the question text; comprehensively utilize multi-dimensional information such as semantic vectors, event categories, and argument information to recall documents; and then determine quality evaluation information corresponding to the event category according to the event category of the question text, and comprehensively consider the relevance and quality evaluation information to determine target documents from the recalled documents, thereby further improving the document recall accuracy of a document question-answering system in a specific domain.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0013] The accompanying drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram of an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure is shown;
[0015] Figure 2 A flowchart of a method for generating response information based on a large language model according to an embodiment of the present disclosure is shown;
[0016] Figure 3 A flowchart of obtaining a document semantic vector according to an embodiment of the present disclosure is shown;
[0017] Figure 4 A flowchart of obtaining a plurality of candidate documents according to an embodiment of the present disclosure is shown;
[0018] Figure 5 A flowchart of determining candidate document quality evaluation information according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A flowchart of determining a comprehensive score of a candidate document according to an embodiment of the present disclosure is shown;
[0020] Figure 7 A structural block diagram of a device for generating response information based on a large language model according to an embodiment of the present disclosure is shown;
[0021] Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. Detailed Embodiments
[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] In the present disclosure, unless otherwise specified, the use of terms such as "first", "second", etc. to describe various elements does not intend to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0024] In the description of the various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0025] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0026] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to embodiments of the present disclosure. Referring Figure 1 to, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0027] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of the above-described method for generating reply information based on a large language model.
[0028] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0029] In Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from system 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0030] Users may use client devices 101, 102, 103, 104, 105, and / or 106 to input question text. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.
[0031] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computer devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smart phones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. The client device is capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0032] Network 110 can be any type of network well-known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0033] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 can run one or more services or software applications that provide the functions described below.
[0034] The computing units in server 120 can run one or more operating systems including any of the above operating systems as well as any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0035] In some embodiments, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0036] In some embodiments, server 120 can be a server of a distributed system, or a server incorporating blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, which solves the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.
[0037] System 100 may also include one or more databases 130. In some embodiments, these databases can be used to store data and other information. For example, one or more of the databases 130 can be used to store information such as audio files and video files. The databases 130 can reside in various locations. For example, the database used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The databases 130 can be of different types. In some embodiments, the database used by the server 120 can be, for example, a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.
[0038] In some embodiments, one or more of the databases 130 can also be used by an application to store application data. The databases used by the application can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0039] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described in this disclosure.
[0040] According to some embodiments, as Figure 2 shown, a method for generating reply information based on a large language model is provided, including:
[0041] Step S201, in response to receiving the question text of the user, obtain the semantic vector of the question text and event information related to a specific domain, where the event information includes the event category asked by the question text and at least one argument information in the question text;
[0042] Step S202, based on at least two of the semantic vector of the question text, at least one argument information, and the event category, obtain multiple candidate documents in the document library of the specific domain;
[0043] Step S203, for the candidate documents among the multiple candidate documents, determine the quality evaluation information of the candidate documents based on the event category; and
[0044] Step S204, based on the relevance between the candidate documents and the question text and the quality evaluation information of the candidate documents, determine at least one target document among the multiple candidate documents, so as to obtain reply information for replying to the question text based on the at least one target document.
[0045] Thus, first, by performing event classification and argument extraction on the user's question text, multi-dimensional information of the question text is obtained; by integrating multi-dimensional information such as semantic vectors, event categories, and argument information, documents are recalled; and then, according to the event category of the question text, quality evaluation information corresponding to the event category is determined, and by integrating the relevance and quality evaluation information, target documents are determined from the recalled documents, thereby further improving the document recall accuracy of the document question-answering system in a specific domain.
[0046] In some embodiments, after receiving the user's question text, first, event classification in a specific domain can be performed on the question text to obtain the event category asked by the question text. For example, in the financial field, if the question text input by the user is used to ask about the event of shareholder change, then its event category is shareholder change.
[0047] In some embodiments, the event category obtained for a question text can be one or more.
[0048] In some embodiments, the event category obtained for a question text can be event categories at multiple levels.
[0049] In some embodiments, event classification can be implemented based on a pre-trained text classification model. For event classification at multiple levels, it can be implemented based on pre-trained event classification models corresponding to each major category respectively. It can be understood that event classification can also be implemented based on other methods (such as text matching, etc.), which are not limited herein.
[0050] In some embodiments, argument extraction can be performed on the question text simultaneously. For example, in the financial field, if the question text input by the user is used to ask about the event of shareholder change, then key arguments such as "shareholder equity share" and "bond scale" may be extracted from the question text.
[0051] In some embodiments, by performing methods such as part-of-speech tagging, named entity recognition, and syntactic analysis on the question text, combined with a deep learning framework for training, an argument extraction model corresponding to a specific domain can be obtained, and then argument extraction is performed based on this model.
[0052] In some embodiments, methods based on rules, templates, and machine learning can also be used for argument extraction, which are not limited herein.
[0053] In some embodiments, before performing document recall, a document library in a specific domain can be pre-constructed. Among them, first, the documents can be parsed to parse document data in various formats into plain text format for storage.
[0054] Taking the documents in the financial field as an example, they mainly include three categories of data: financial reports, research reports, and news sentiment. For financial reports, listed companies usually publicly release them in the form of PDF, which mainly contains content related to the three major financial statements and mainly exists in the form of tables in the document. This requires that the document parsing module should have the ability to extract tables. For research reports, the research trends of financial institutions and related practitioners on the macro economy or listed companies are usually also released in the form of PDF. The difference from the document form of financial reports is that the document elements of research reports are rich and diverse, and the layout is relatively complex, often containing two columns, three columns, etc. Using common PDF parsing tools will result in text garbled and sentence disorder. For news sentiment, it is often publicly released in the form of web content, so document parsing needs to meet the function of html parsing.
[0055] For the documents in the financial field, they can be realized through a document parsing module with the abilities of web page parsing, PDF parsing, document column splitting, table extraction, etc. For a given financial document, first parse the document to obtain the paragraph text content in the document and save it in pure text format; then extract the basic structured fields of the document, such as document source, document author, document release date, document title, etc., and associate the above basic structured fields with the index information of the document, and store them in the document library together with the document content in pure text format.
[0056] In some embodiments, event classification and argument extraction can be further performed on the document content.
[0057] In some embodiments, a method similar to the event classification and argument extraction of the above problem text can be used to perform event classification and argument extraction on the document content, and associate one or more document event categories and one or more document argument information of the document content with the index information of the document, and store them in the document library together. Taking the documents in the financial field as an example, a total of 10 major categories and 65 minor categories of financial events have been sorted out for the events in the financial field. Each category of financial event contains basic arguments such as event name, event subject, event date, etc. In addition, each category of financial event also has additional key arguments (such as the event "shareholder change" contains "shareholder equity share", "bond default" contains "bond scale", etc.).
[0058] In some embodiments, the semantic vector of the document content can be further obtained.
[0059] In some embodiments, the document content can be directly input into the semantic understanding model to obtain the semantic vector of the text content.
[0060] In some embodiments, such as Figure 3As shown in the figure, the acquisition of the document semantic vector of the preset document includes: for each preset document among multiple preset documents, perform the following operations: Step S301, segment the preset document to obtain at least one document paragraph; Step S302, obtain at least one paragraph semantic vector corresponding to at least one document paragraph; and Step S303, based on at least one paragraph semantic vector, obtain the document semantic vector of the preset document.
[0061] Thus, by segmenting the document and obtaining the semantic vectors of each paragraph respectively, it is possible to obtain a document semantic vector with richer semantic information based on the paragraph semantic vectors, further improving the accuracy of subsequent document recall and ranking.
[0062] In some embodiments, at least one paragraph semantic vector can be concatenated to obtain the document semantic vector of the preset document.
[0063] In some embodiments, at least one paragraph semantic vector can also be superimposed to obtain the document semantic vector of the preset document.
[0064] In some embodiments, the semantic vector of the document can be associated with the index information of the document and stored in the document library together.
[0065] In some embodiments, to obtain the text content and related structured fields from the original document above, ElasticSearch can be used as the data storage. Thus, it is possible to quickly respond to the relevance retrieval of the user's question text.
[0066] In some embodiments, the semantic vector can be stored using milvus and associated with the document content through the document index, thereby further improving the efficiency of semantic vector recall.
[0067] In some embodiments, based on at least two of the semantic vector of the question text, at least one argument information, and the event category, multiple candidate documents can be obtained from the document library in a specific domain. Document recall can be performed respectively through the semantic vector, document recall can be performed by matching the argument information, and document recall can be performed by matching the event category, and the results recalled by each part are intersected. Thus, the recall results can be filtered through multi-dimensional information, improving the accuracy of the recall results.
[0068] In some embodiments, based on at least two of the semantic vector of the question text, at least one argument information, and the event category, multiple candidate documents are retrieved from a document library in a specific domain. First, a preliminary recall of documents can be performed based on the event category and the argument information, and then multiple candidate documents with the highest semantic relevance are retrieved from the preliminary recall results based on the semantic vector. Thus, through the two-stage precise recall method, on the one hand, the scope of semantic retrieval can be greatly reduced, and on the other hand, the most accurate response information can be generated based on the most core documents.
[0069] In some embodiments, as Figure 4 shown, retrieving multiple candidate documents from a document library in a specific domain based on at least two of the semantic vector of the question text, at least one argument information, and the event category includes: Step S401, retrieving multiple first candidate documents with the highest semantic similarity from the document library based on the semantic vector of the question text and the document semantic vectors of the preset documents in the document library; Step S402, retrieving multiple second candidate documents in the document library that match at least one of the document event category and the document argument information based on the event category and at least one argument information; and Step S403, obtaining multiple candidate documents based on the multiple first candidate documents and the multiple second candidate documents.
[0070] Thus, by retrieving based on the semantic vector respectively and retrieving based on the time category and / or argument information, a richer recall result can be obtained, further improving the accuracy of the subsequent target document.
[0071] In some embodiments, the multiple second candidate documents can be retrieved based on the event category or the argument information.
[0072] In some embodiments, the multiple second candidate documents can be retrieved based on the event category and the argument information respectively, and the results retrieved from each part are intersected, so that the recall result can be filtered by multi-dimensional information and the accuracy of the recall result can be improved.
[0073] In some embodiments, the multiple second candidate documents can also be retrieved based on the event category and the argument information at the same time, and the documents that match both the event category and the argument information are retrieved, so as to further improve the accuracy of the recall result.
[0074] In some embodiments, obtaining multiple candidate documents based on the multiple first candidate documents and the multiple second candidate documents can be to take the intersection of the multiple first candidate documents and the multiple second candidate documents, so that the recall result can be filtered by multi-dimensional information and the accuracy of the recall result can be improved.
[0075] In some embodiments, obtaining a plurality of candidate documents based on a plurality of first candidate documents and a plurality of second candidate documents may be taking the union of the plurality of first candidate documents and the plurality of second candidate documents. For the merged plurality of documents, duplicate removal and preliminary sorting work (e.g., sorting based on semantic relevance) may be further performed to obtain a plurality of candidate documents ranked at the top.
[0076] In some embodiments, after obtaining a plurality of candidate documents, for each candidate document among the plurality of candidate documents, quality evaluation information of the candidate document may be determined based on the event category.
[0077] In some embodiments, corresponding quality evaluation strategies may be set for each event category, and based on the quality evaluation strategy corresponding to the problem event category, the quality evaluation information (e.g., quality score) of the document may be determined. For example, if the event category involved in the user's question is operating income, since the user preferably obtains answers from professional and standardized financial reports, followed by public opinions with high timeliness, the quality scores of three types of documents, namely research reports, financial reports, and public opinions, may be set to 70, 90, and 80 respectively. Thus, the quality evaluation information of each candidate document is determined according to the problem event category.
[0078] In some embodiments, as Figure 5 shown, determining the quality evaluation information of a candidate document among a plurality of candidate documents based on the event category includes: Step S501, determining at least one document evaluation dimension corresponding to the event category based on the event category, where the document evaluation dimension includes a plurality of document categories; Step S502, determining the quality score corresponding to each document category in the document evaluation dimension based on the event category; and Step S503, for the candidate document among the plurality of candidate documents, determining the quality evaluation information of the candidate document based on the corresponding document category in each document evaluation dimension of the at least one document evaluation dimension of the candidate document and the quality score of the document category corresponding to the event category.
[0079] Thus, based on the event category of the problem text, the evaluation dimension of the document is determined, and then the quality score of each document category corresponding to each document in each evaluation dimension is determined based on the event category, and the quality evaluation information of the document is determined based on the quality scores of each dimension, so as to be able to perform a quality evaluation of the document that better meets the user's needs according to the event category of the user's problem text and improve the accuracy of subsequent sorting.
[0080] In some embodiments, one event category of the problem text may correspond to one or more quality evaluation dimensions, and for each evaluation dimension, the quality score corresponding to the event category can be determined for the document.
[0081] In some exemplary embodiments, the event category queried by the user's question text may be policy information, and this event category may respectively correspond to two quality evaluation dimensions: authority and timeliness. Among them, for the authority evaluation dimension, documents can be classified based on the document source, for example, classified into research reports, financial reports, and public opinions, and a higher authority quality score can be set for the news and public opinion category; for the timeliness evaluation dimension, documents can be classified based on the release time of the document, and a higher timeliness quality score can be set for the more recent release time.
[0082] In some embodiments, for each candidate document, one or more quality scores can be obtained based on the above method. For multiple quality scores, different weights can be further set for each quality evaluation dimension based on the question event category, and each quality score can be weighted to obtain the comprehensive quality score of the candidate document, which can be used as the quality evaluation information of the candidate document.
[0083] In some embodiments, determining at least one target document among multiple candidate documents based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document may include: determining the comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document; and determining at least one candidate document whose comprehensive score meets the preset conditions among the multiple candidate documents as at least one target document.
[0084] Thus, by combining the comprehensive quality score and the relevance between the document content and the question text, the comprehensive score of each candidate document can be determined, and then the candidate documents can be re - sorted and the target document can be determined, so that a target document with better quality and higher relevance to the question text can be obtained.
[0085] In some embodiments, the semantic relevance between the candidate document and the question text can be weighted based on the comprehensive quality score to obtain the comprehensive score of each candidate document.
[0086] In some embodiments, the quality evaluation information may include the quality scores corresponding to each document evaluation dimension in at least one document evaluation dimension of the corresponding candidate document. For example, for the quality evaluation information of multiple document evaluation dimensions, the quality scores of each dimension can be sorted into a feature vector in a preset order.
[0087] In some embodiments, such as Figure 6As shown in the figure, determining the comprehensive score of a candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document may include: Step S601, inputting at least the question text, the candidate document, and the quality evaluation information of the candidate document into a trained ranking model; Step S602, using the ranking model to determine the relevance between the candidate document and the question text based on at least the question text and the candidate document; and Step S603, using the ranking model to determine the comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document.
[0088] Thus, by inputting the relevant information of the question text and the document, as well as the quality evaluation information of each document relative to the current question text, into the ranking model, it is possible to comprehensively consider the document relevance information and quality information to obtain a more accurate ranking result that better meets the user's needs.
[0089] In some embodiments, the quality evaluation information (such as the above-mentioned comprehensive quality score or quality feature vector) can be input into a pre-trained and finely ranked ranking model together with the question text and the document content of each candidate document. The model comprehensively evaluates the relevance and document quality, and then outputs the comprehensive score of each candidate document.
[0090] In some embodiments, the event category information and argument information of the question text and the candidate document can be further input into the above-mentioned ranking model to further improve the ranking effect of the model.
[0091] In some embodiments, at least one candidate document whose comprehensive score among multiple candidate documents meets a preset condition is determined as at least one target document, and the at least one target document is used as the reply information for answering the question text. The preset condition can be that the comprehensive score is greater than a certain preset threshold, or it can be to determine the target document as the one or more documents with the highest comprehensive score, or it can be to determine the target document as the one or more documents with the highest comprehensive score and whose comprehensive score is greater than a certain preset threshold.
[0092] In some embodiments, the above method for generating reply information based on a large language model may further include: organizing the question text and at least one target document into an instruction text according to a preset instruction template; and inputting the instruction text into a reply information generation model to obtain the reply information output by the reply information generation model for answering the question text.
[0093] Thus, by integrating the question text and the target document into an instruction text and inputting them into the reply information generation model together, the readability of the reply information can be further improved, and the user experience can be optimized.
[0094] In some exemplary embodiments, the preset instruction templates can be, for example, "Please try to avoid generating response content that is irrelevant to the question, self-contradictory, or semantically repetitive", "Please comprehensively refer to the search results related to the question and clearly, smoothly, and in detail answer the question [question text] in combination with the retrieved documents [target document 1], [target document 2], [target document 3], [target document 4]", etc.
[0095] In some embodiments, the response information generation model can be a knowledge-enhanced large language model for conversations (such as ERNIE bot, etc.). The response information generation model is trained using a vast amount of knowledge resources and conversation data. Applying such a model as the response information generation model can not only directly process chitchat-style conversation information but also directly generate response information for logical reasoning-style, common sense-style, and image generation-style conversation information, capable of further improving the generation efficiency while generating higher-quality response information.
[0096] In some embodiments, as Figure 7 shown, a generation device 700 for response information based on a large language model is provided, including:
[0097] A first acquisition unit 710, configured to, in response to receiving the question text of the user, acquire the semantic vector of the question text and event information related to a specific domain, where the event information includes the event category queried by the question text and at least one argument information in the question text;
[0098] A second acquisition unit 720, configured to acquire a plurality of candidate documents in a document library of a specific domain based on at least two of the semantic vector of the question text, at least one argument information, and the event category;
[0099] A first determination unit 730, configured to, for a candidate document among the plurality of candidate documents, determine quality evaluation information of the candidate document based on the event category;
[0100] A second determination unit 740, configured to determine at least one target document among the plurality of candidate documents based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document, so as to obtain response information for answering the question text based on the at least one target document.
[0101] Among them, the operations performed by units 710 - 740 in the above generation device 700 for response information based on a large language model are similar to the operations of steps S201 - S204 in the above generation method for response information based on a large language model, and will not be elaborated here.
[0102] In some embodiments, the first determination unit may include: a first determination subunit configured to determine, based on an event category, at least one document evaluation dimension corresponding to the event category, where the document evaluation dimension includes multiple document categories; a second determination subunit configured to determine, based on the event category, a quality score corresponding to the document category; and a third determination subunit configured to, for a candidate document among multiple candidate documents, determine quality evaluation information of the candidate document based on the corresponding document category in each document evaluation dimension of at least one document evaluation dimension of the candidate document and the quality score of the document category corresponding to the event category.
[0103] In some embodiments, the second determination unit may include: a fourth determination subunit configured to determine a comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document; and a fifth determination subunit configured to determine at least one candidate document whose comprehensive score among multiple candidate documents meets a preset condition as at least one target document.
[0104] In some embodiments, the quality evaluation information may include quality scores corresponding to each document evaluation dimension in at least one document evaluation dimension of the corresponding candidate document, and the fourth determination subunit may be further configured to: input at least the question text, the candidate document, and the quality evaluation information of the candidate document into a trained ranking model; use the ranking model to determine the relevance between the candidate document and the question text based on at least the question text and the candidate document; and use the ranking model to determine the comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document.
[0105] In some embodiments, the document library includes multiple preset documents, and the preset documents among the multiple preset documents may include corresponding document semantic vectors, at least one document event category, and at least one document argument information. The second acquisition unit may include: a first acquisition subunit configured to recall, based on the semantic vector of the question text and the document semantic vectors of the preset documents in the document library, multiple first candidate documents with the highest semantic similarity in the document library; a second acquisition subunit configured to acquire, based on the event category and at least one argument information, multiple second candidate documents in the document library whose at least one of the document event category and the document argument information matches; and a third acquisition subunit configured to acquire multiple candidate documents based on the multiple first candidate documents and the multiple second candidate documents.
[0106] In some embodiments, the acquisition of the document semantic vector of the preset document may include: segmenting the preset document to obtain at least one document paragraph; acquiring at least one paragraph semantic vector corresponding to the at least one document paragraph; and acquiring the document semantic vector of the preset document based on the at least one paragraph semantic vector.
[0107] In some embodiments, the second determination unit may include: an arrangement subunit configured to arrange the question text and at least one target document into an instruction text according to a preset instruction template; and a fourth acquisition subunit configured to input the instruction text into a reply information generation model to obtain reply information for replying to the question text output by the reply information generation model.
[0108] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.
[0109] Referring Figure 8 , a block diagram of an electronic device 800 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0110] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0111] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device capable of presenting information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include but are not limited to a magnetic disk, an optical disc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0112] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the method for generating reply information based on the large language model described above. For example, in some embodiments, the method for generating reply information based on the large language model described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method for generating reply information based on the large language model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the method for generating reply information based on the large language model described above in any other suitable manner (e.g., by means of firmware).
[0113] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0114] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0115] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0117] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0118] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.
[0119] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.
[0120] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. A method for generating reply information based on a large language model, the method comprises: In response to receiving a user's question text, obtaining a semantic vector of the question text and event information related to a specific domain, where the event information includes an event category queried by the question text and at least one argument information in the question text; Based on at least two of the semantic vector of the question text, the at least one argument information, and the event category, obtaining a plurality of candidate documents in a document library of the specific domain; For a candidate document among the plurality of candidate documents, determining quality evaluation information of the candidate document based on the event category; and Based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document, determining at least one target document among the plurality of candidate documents, so as to obtain reply information for replying to the question text based on the at least one target document.
2. The method according to claim 1, wherein, The determining quality evaluation information of the candidate document for the candidate document among the plurality of candidate documents based on the event category includes: Based on the event category, determining at least one document evaluation dimension corresponding to the event category, where the document evaluation dimension includes a plurality of document categories; Based on the event category, determining a quality score corresponding to the document category; and For the candidate document among the plurality of candidate documents, based on the corresponding document category in each document evaluation dimension of at least one document evaluation dimension of the candidate document and the quality score of the document category corresponding to the event category, determining the quality evaluation information of the candidate document.
3. The method according to claim 2, wherein, The determining at least one target document among the plurality of candidate documents based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document includes: Based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document, determining a comprehensive score of the candidate document; and Determining at least one candidate document whose comprehensive score meets a preset condition among the plurality of candidate documents as the at least one target document.
4. The method according to claim 3, wherein, The quality evaluation information includes quality scores corresponding to each document evaluation dimension in at least one document evaluation dimension of the corresponding candidate document, and the determining a comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document includes: At least inputting the question text, the candidate document, and the quality evaluation information of the candidate document into a trained ranking model; Using the ranking model, determining the relevance between the candidate document and the question text based on at least the question text and the candidate document; and Using the ranking model, determining a comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document.
5. The method according to any one of claims 1-4, wherein, The document library includes multiple preset documents. The preset documents among the multiple preset documents include corresponding document semantic vectors, at least one document event category, and at least one document argument information. Obtaining multiple candidate documents in the document library of the specific domain based on at least two of the semantic vector of the question text, the at least one argument information, and the event category includes: Based on the semantic vector of the question text and the document semantic vectors of the preset documents in the document library, recalling multiple first candidate documents with the highest semantic similarity in the document library; Based on the event category and the at least one argument information, obtaining multiple second candidate documents in the document library that match at least one of the document event category and the document argument information; and Based on the multiple first candidate documents and the multiple second candidate documents, obtaining the multiple candidate documents.
6. The method according to claim 5, wherein, Obtaining the document semantic vector of the preset document includes: Segmenting the preset document to obtain at least one document paragraph; Obtaining at least one paragraph semantic vector corresponding to the at least one document paragraph; and Based on the at least one paragraph semantic vector, obtaining the document semantic vector of the preset document.
7. The method according to any one of claims 1-6, wherein, Obtaining reply information for replying to the question text based on the at least one target document includes: Organizing the question text and the at least one target document into an instruction text according to a preset instruction template; and Inputting the instruction text into a reply information generation model to obtain the reply information output by the reply information generation model for replying to the question text.
8. A device for generating reply information based on a large language model, the device comprises: A first acquisition unit, configured to, in response to receiving a question text of a user, acquire the semantic vector of the question text and event information related to a specific domain, where the event information includes the event category asked by the question text and at least one argument information in the question text; A second acquisition unit, configured to obtain multiple candidate documents in the document library of the specific domain based on at least two of the semantic vector of the question text, the at least one argument information, and the event category; A first determination unit, configured to, for a candidate document among the multiple candidate documents, determine quality evaluation information of the candidate document based on the event category; and A second determination unit, configured to determine at least one target document among the multiple candidate documents based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document, so as to obtain reply information for replying to the question text based on the at least one target document.
9. The device according to claim 8, wherein, The first determination unit includes: A first determination subunit, configured to determine at least one document evaluation dimension corresponding to the event category based on the event category, where the document evaluation dimension includes multiple document categories; A second determination subunit, configured to determine a quality score corresponding to the document category based on the event category; and A third determination subunit, configured to, for a candidate document among the multiple candidate documents, determine quality evaluation information of the candidate document based on each document evaluation dimension in at least one document evaluation dimension of the candidate document, the corresponding document category in the document evaluation dimension, and the quality score of the document category corresponding to the event category.
10. The apparatus according to claim 9, wherein, the second determination unit includes: A fourth determination subunit, configured to determine a comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document; and A fifth determination subunit, configured to determine at least one candidate document whose comprehensive score among the multiple candidate documents meets a preset condition as the at least one target document.
11. The apparatus according to claim 10, wherein, the quality evaluation information includes quality scores corresponding to each document evaluation dimension in at least one document evaluation dimension of the corresponding candidate document, and the fourth determination subunit is further configured to: input at least the question text, the candidate document, and the quality evaluation information of the candidate document into a trained ranking model; utilize the ranking model to determine the relevance between the candidate document and the question text based on at least the question text and the candidate document; and utilize the ranking model to determine the comprehensive score of the candidate document based on the relevance between the candidate document and the question text and the quality evaluation information of the candidate document.
12. The apparatus according to any one of claims 8-11, wherein, the document library includes multiple preset documents, and the preset documents in the multiple preset documents include corresponding document semantic vectors, at least one document event category, and at least one document argument information. The second acquisition unit includes: A first acquisition subunit, configured to recall multiple first candidate documents with the highest semantic similarity in the document library based on the semantic vector of the question text and the document semantic vectors of the preset documents in the document library; A second acquisition subunit, configured to acquire multiple second candidate documents in the document library whose at least one of the document event category and the document argument information matches based on the event category and the at least one argument information; and A third acquisition subunit, configured to acquire the multiple candidate documents based on the multiple first candidate documents and the multiple second candidate documents.
13. The apparatus according to claim 12, wherein, the acquisition of the document semantic vector of the preset document includes: segmenting the preset document to obtain at least one document paragraph; acquiring at least one paragraph semantic vector corresponding to the at least one document paragraph; and acquiring the document semantic vector of the preset document based on the at least one paragraph semantic vector.
14. The apparatus according to any one of claims 8-13, wherein, the second determination unit includes: A sorting subunit, configured to sort the problem text and the at least one target document into an instruction text according to a preset instruction template; and A fourth acquisition subunit, configured to input the instruction text into a reply information generation model to obtain reply information for replying to the problem text output by the reply information generation model.
15. An electronic device comprising:[[]] At least one processor; and A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions wherein The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product comprising a computer program wherein The computer program, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Cited By
Multi-modal speech language large model training method and device, equipment and medium
CN120472888A
Data processing method and device
CN121412362A
Information recall method and device, electronic equipment and storage medium
CN121501922A