Session response method and device based on large language model, equipment and medium
By combining a large language model with a three-layer screening logic of pre-computation screening and vector retrieval algorithm, the problems of cross-turn information loss and high computational complexity in historical memory management in multi-turn dialogue systems are solved, and efficient and accurate conversation response is achieved.
Patent Information
- Application Number
- CN202511064602.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-04
AI Technical Summary
In multi-turn dialogue systems, the management of historical dialogue memory suffers from problems such as loss of key information across turns, difficulty in identifying semantic variants, and high computational complexity, which affect the continuity of interaction and the accuracy of response.
By employing a large language model combined with pre-computation screening and sparse and dense vector retrieval algorithms, historical sessions related to the current user's request are screened out through keyword matching, topic matching, and logical relevance analysis, thereby reducing computational load and resource consumption.
It improves the accuracy and consistency of the session system's response, reduces computational complexity and resource consumption, and enhances the user experience.
Smart Images

Figure CN120892529A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to a conversation response method and device based on a large language model, equipment and medium. BACKGROUND
[0002] In a multi-turn dialogue system, effective management of historical dialogue memory is a core link to ensure the continuity of interaction and the accuracy of response. As the number of dialogue turns increases, accumulated historical information, such as user requests and model answers, gradually forms a large context. If the relevant content cannot be accurately filtered, it will not only cause an increase in computational complexity and response delay during model reasoning, but also may cause the answer to deviate from the topic due to irrelevant information interference, which seriously affects the user experience. Early multi-turn dialogue memory filtering methods have obvious limitations: 1. Sliding window mechanism: By retaining the last N turns of dialogue as context to control the context size, but it cannot identify key information across turns, for example, the "meeting location" mentioned in the first turn still needs to be referenced in the fifth turn, which may be truncated by the window, resulting in loss of important information and affecting the continuity of the dialogue. 2. Keyword matching method: Relies on manually set keywords or entities for matching, which can only handle literal consistent expressions and cannot handle semantic variants such as "mobile phone" and "smart terminal", "call a car" and "call a car", which are likely to misjudge relevant information or miss implicit associated content, resulting in low filtering accuracy.
[0003] With the development of deep learning technology, the Transformer architecture and self-attention mechanism provide a new direction for memory filtering, which to some extent solves the problem of semantic variants, but as the number of dialogue turns increases linearly, the size of historical memory continues to expand, resulting in an exponential increase in the complexity of similarity calculation and a significant decrease in model reasoning efficiency. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a conversation response method, device, equipment and medium based on a large language model, which reduces the interference of irrelevant memory on model answers by combining pre-computed preliminary screening with a large language model, improves the accuracy and continuity of the content generated by the conversation system, reduces the computational complexity and resource occupation, and improves the user experience. The specific solutions are as follows:
[0005] In a first aspect, the present application provides a conversation response method based on a large language model, comprising:
[0006] obtaining a current user request and determining the request keywords and request topics corresponding to the current user request;
[0007] determine, based on the request keyword, a first historical session that matches the keyword of the current user request by a first preset matching degree threshold from historical session records, to construct a first historical session candidate set according to the first historical session, determine, based on the request topic, a second historical session that matches the topic of the current user request by a second preset matching degree threshold from the first historical session candidate set, to construct a second historical session candidate set according to the second historical session; wherein the historical session includes a historical user request and a corresponding historical model answer;
[0008] analyze, by using a preset large language model, logical relevance between each of the second historical sessions in the second historical session candidate set and the current user request, to obtain an analysis result, and label the second historical sessions with corresponding relevance labels according to the analysis result; the relevance labels include a strong relevance label representing that a preset strong relevance is met, a partial relevance label representing that a preset partial relevance is met, and an irrelevant label representing no relevance.
[0009] determine, based on the relevance labels, a target historical session associated with the current user request, to respond to the current user request by using the target historical session.
[0010] Optionally, the large language model-based session response method further includes:
[0011] obtain an initial user request sent by a user terminal, and perform a preprocessing operation on the initial user request to obtain a current user request; the preprocessing operation includes irrelevant information removal, sentence clarification processing, sentence correction, and omission information filling.
[0012] Optionally, the large language model-based session response method further includes:
[0013] extract, by using a preset large language model, core session keywords and session topics corresponding to each historical session record, and save the historical session records and the corresponding core session keywords and session topics into a target database; the core session keywords are keywords that meet a preset core keyword standard.
[0014] Optionally, the determining, based on the request keyword, of a first historical session that matches the keyword of the current user request by a first preset matching degree threshold from historical session records includes:
[0015] perform keyword matching degree analysis on the request keyword corresponding to the current user request and the core session keywords corresponding to each historical session record saved in the target database based on a preset sparse vector retrieval algorithm;
[0016] if there is a historical session in which at least one of the core session keywords is contained in the request keywords, the historical session in which at least one of the core session keywords is contained in the request keywords is determined as a first historical session in which the keyword matching degree with the current user request reaches a first preset matching degree threshold.
[0017] Optionally, the determining, from the first historical session candidate set, a second historical session in which the topic matching degree with the current user request reaches a second preset matching degree threshold based on the request topic, comprises:
[0018] performing topic matching degree analysis on the request topic corresponding to the current user request and the session topic corresponding to each first historical session in the first historical session candidate set based on a preset dense vector retrieval algorithm;
[0019] if there is a first historical session in the first historical session candidate set that is consistent with the request topic, the first historical session that is consistent with the request topic is determined as a second historical session in which the topic matching degree with the current user request reaches the second preset matching degree threshold.
[0020] Optionally, the analyzing, by using a preset large language model, the logical correlation between each second historical session in the second historical session candidate set and the current user request to obtain an analysis result, comprises:
[0021] determining a target dialogue scenario corresponding to the current user request, and analyzing, in combination with the target dialogue scenario and by using a preset large language model, the logical correlation between each second historical session in the second historical session candidate set and the current user request;
[0022] if there is a target historical session in the second historical session candidate set that covers all information required by the current user request, the target historical session is determined as a strongly related session corresponding to the current user request;
[0023] if there is a target historical session in the second historical session candidate set that covers part of the information required by the current user request, the target historical session is determined as a partially related session corresponding to the current user request;
[0024] if there is a target historical session in the second historical session candidate set that is irrelevant to the current user request, the target historical session is determined as an irrelevant session corresponding to the current user request.
[0025] Optionally, the marking, according to the analysis result, of a corresponding relevance label for the second historical session, comprises:
[0026] Assign strongly relevant labels to strongly relevant sessions in the second historical session candidate set, assign partially relevant labels to partially relevant sessions in the second historical session candidate set, and assign irrelevant labels to irrelevant sessions in the second historical session candidate set.
[0027] Accordingly, determining the target historical session associated with the current user request based on the relevance label includes:
[0028] Sessions marked with strong and partial relevant tags in the second set of historical session candidates are selected as target historical sessions associated with the current user request.
[0029] Secondly, this application provides a conversational response device based on a large language model, comprising:
[0030] The data acquisition module is used to acquire the current user request and determine the request keywords and request topic corresponding to the current user request;
[0031] The set determination module is used to determine, based on the requested keywords, a first historical session from historical session records whose keyword matching degree with the current user's request reaches a first preset matching degree threshold, so as to construct a first historical session candidate set based on the first historical session; and to determine, based on the requested topic, a second historical session from the first historical session candidate set whose topic matching degree with the current user's request reaches a second preset matching degree threshold, so as to construct a second historical session candidate set based on the second historical session; wherein, historical sessions include historical user requests and corresponding historical model answers;
[0032] The relevance analysis module is used to analyze the logical relevance between each of the second historical sessions in the second historical session candidate set and the current user request using a preset large language model, obtain the analysis results, and mark the second historical sessions with corresponding relevance tags according to the analysis results; the relevance tags include strongly relevant tags that represent a preset strong relevance, partially relevant tags that represent a preset partial relevance, and irrelevant tags that represent no relevance;
[0033] The request-response module is used to determine the target historical session associated with the current user request based on the relevance label, so as to respond to the current user request using the target historical session.
[0034] Thirdly, this application provides an electronic device, comprising:
[0035] Memory, used to store computer programs;
[0036] A processor is used to execute the computer program to implement the aforementioned conversational response method based on a large language model.
[0037] In a fourth aspect, the present application provides a computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the foregoing large language model-based conversation response method.
[0038] In the present application, the current user request is obtained, and the request keyword and the request topic corresponding to the current user request are determined; a first historical conversation in which the keyword matching degree with the current user request reaches a first preset matching degree threshold is determined from historical conversation records based on the request keyword, to construct a first historical conversation candidate set according to the first historical conversation, a second historical conversation in which the topic matching degree with the current user request reaches a second preset matching degree threshold is determined from the first historical conversation candidate set based on the request topic, to construct a second historical conversation candidate set according to the second historical conversation; wherein the historical conversation includes a historical user request and a corresponding historical model answer; the logical correlation between each second historical conversation in the second historical conversation candidate set and the current user request is analyzed by using a preset large language model, to obtain an analysis result, and the second historical conversation is labeled with a corresponding correlation label according to the analysis result; the correlation label includes a strong correlation label representing that a preset strong correlation is met, a partial correlation label representing that a preset partial correlation is met, and an irrelevant label representing no correlation; a target historical conversation associated with the current user request is determined based on the correlation label, to respond to the current user request by using the target historical conversation. As can be seen from the above, the present application breaks through the round limit through the three-layer filtering logic of “keyword matching-topic matching-logical correlation”, no matter which round the historical conversation is in, as long as at least one of the core conversation keywords is contained in the request keyword, belongs to the same topic field as the current request, and has a logical correlation with the current request, the historical conversation will be labeled as “strong correlation” or “partial correlation” and filtered out, avoiding the loss of cross-round key information. On the other hand, the present application first calculates the request keyword and the request topic corresponding to the current user request to preliminarily screen all historical conversations, greatly compresses the size of the candidate set, and only uses the large language model to analyze the logical correlation of the compressed candidate set, avoiding the direct processing of all historical conversations by the large model, significantly reducing the amount of calculation and resource consumption, while ensuring the filtering accuracy, and solving the contradiction between the calculation efficiency and the filtering effect. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0040] Figure 1 A flow chart of a conversation response method based on a large language model disclosed in the present application;
[0041] Figure 2 A structural schematic diagram of a conversation response device based on a large language model disclosed in the present application;
[0042] Figure 3 A structural schematic diagram of an electronic device disclosed in the present application. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0044] Early multi-turn dialogue memory filtering methods have obvious limitations, such as inability to identify cross-turn key information, inability to cope with semantic variants, and low filtering precision. With the development of deep learning technology, the Transformer architecture and self-attention mechanism have solved the semantic variant problem to some extent, but as the number of dialogue turns increases linearly, the size of historical memory continues to expand, resulting in an exponential increase in the complexity of similarity calculation, and a significant decrease in model inference efficiency. Therefore, the present application provides a conversation response method based on a large language model, which reduces the interference of irrelevant memory on model answers by combining pre-computed preliminary screening with a large language model, improves the accuracy and continuity of the content generated by the conversation system, and reduces the computational complexity and resource occupation, thereby improving the user experience of the conversation.
[0045] Referring to Figure 1 The embodiments of the present application disclose a conversation response method based on a large language model, which comprises:
[0046] Step S11, obtaining a current user request and determining a request keyword and a request topic corresponding to the current user request.
[0047] In this embodiment, an initial user request sent by a user terminal can be obtained first, and then a current user request is obtained by performing a preprocessing operation on the initial user request; wherein the preprocessing operation includes but is not limited to irrelevant information removal, sentence clarification processing, sentence correction, and omitted information filling.
[0048] It can be understood that in the multi-round dialogue, the historical memory mainly includes two element information of a user request Q (Query) and a model answer A (Answer). The user request here is not the original request of the user, but an optimized rewritten user request, so as to improve the knowledge retrieval accuracy and the model answer accuracy in the multi-round dialogue scenario.
[0049] In step S12, a first historical session in which the keyword matching degree with the current user request reaches a first preset matching degree threshold is determined from the historical session records based on the request keywords, a first historical session candidate set is constructed according to the first historical session, a second historical session in which the theme matching degree with the current user request reaches a second preset matching degree threshold is determined from the first historical session candidate set based on the request theme, and a second historical session candidate set is constructed according to the second historical session. The historical session includes a historical user request and a corresponding historical model answer.
[0050] In this embodiment, first, the request keywords corresponding to the current user request and the core session keywords corresponding to each historical session record saved in the target database can be analyzed for keyword matching degree based on a preset sparse vector retrieval algorithm. If there is a historical session in which the core session keywords are at least contained by the request keywords, the historical session in which the core session keywords are at least contained by the request keywords is taken as the first historical session in which the keyword matching degree with the current user request reaches the first preset matching degree threshold, and a first historical session candidate set is constructed according to the first historical session. Further, the request theme corresponding to the current user request and the session theme corresponding to each first historical session in the first historical session candidate set can be analyzed for theme matching degree based on a preset dense vector retrieval algorithm. If there is a first historical session in the first historical session candidate set that is consistent with the request theme, the first historical session that is consistent with the request theme is taken as the second historical session in which the theme matching degree with the current user request reaches the second preset matching degree threshold, and a second historical session candidate set is constructed according to the second historical session.
[0051] The core session keywords and the session theme corresponding to each historical session record can be extracted by using a preset large language model, and the historical session record and the corresponding core session keywords and session theme are saved in the target database. The core session keywords meet the preset core keyword standard, and the target database can be a vector database.
[0052] It should be noted that the above process of analyzing the keyword matching degree between the request keywords corresponding to the current user request and the core session keywords corresponding to each historical session record saved in the target database is a keyword matching stage. In the keyword matching stage, the historical round Q n , A ncore keywords such as entities, actions, attributes, etc. in the historical session are compared with the keywords in the current round user question Q. For example, assuming that in the historical session, Q n is “recommend a Sichuan restaurant in XX area”, A n is “recommend XX Sichuan restaurant, located in XX district, XX road, per capita 80 yuan”, the core session keywords extracted therefrom include “XX area”, “Sichuan restaurant”, “XX road”, and “per capita 80 yuan”, the current user request Q is “does the Sichuan restaurant in XX area need to be reserved?”, the request keywords extracted therefrom are “XX area” and “Sichuan restaurant”, and the matching degree of the two is calculated by the sparse vector retrieval algorithm. Since Q contains the core keywords “XX area” and “Sichuan restaurant” in the historical session, the keyword matching degree reaches the first preset matching degree threshold, so the historical session is included in the first historical session candidate set and enters the next stage. At the same time, the historical sessions in the target database that do not match the keywords of the current user request to the first preset matching degree threshold can be marked as “weakly related”.
[0053] The above process of subject matching degree analysis of the request subject corresponding to the current user request and the session subject corresponding to each first historical session in the first historical session candidate set is the subject consistency judgment stage. In the subject consistency judgment stage, it can be judged whether Q n / A n and Q belong to the same field or topic, that is, whether the subjects are consistent. For example, assuming that in the historical session, Q n is “recommend a love movie suitable for couples to watch”, A n is “Love Before Dawn is a classic love film, suitable for couples to watch, telling the romantic encounter of a pair of strangers in Vienna”, the session subject extracted therefrom is “love movie recommendation”, the current user request Q is “what other works does the director of the movie have”, and the session subject extracted therefrom is “other works of the director of Love Before Dawn”, which still belongs to the field of “love movies and related creators”, the subjects are consistent, so the historical session is included in the second historical session candidate set and enters the next logical relevance judgment stage. At the same time, the historical sessions in the first historical session candidate set that do not match the subject of the current user request to the second preset matching degree threshold can be marked as “irrelevant”.
[0054] The above keyword matching stage and subject consistency judgment stage are similar to the external knowledge retrieval capability in the RAG (Retrieval-Augmented Generation) system, which provides the initial screening capability for historical sessions.
[0055] It should be noted that the keyword matching stage and the theme consistency judgment stage can use the model vector method to calculate the similarity, compared with the traditional matching algorithm, the model vector method can provide stronger semantic relevance, and capture the implicit relationship between synonyms and near synonyms. The keyword matching stage uses a sparse vector matching algorithm such as BM25 (Best Matching 25), and the theme consistency judgment stage uses a dense vector matching algorithm such as HNSW (Hierarchical Navigable Small World), and the algorithm similarity return value is used as the judgment basis. The keywords of the historical round Q n / A n The keyword vector and the theme information vector are pre-computed, that is, when the historical session is saved, the system has already converted the keywords and themes in it into vectors and stored them in the database. When the user asks a question, only the keywords and themes of Q need to be converted into vectors and compared directly with the "pre-computed historical vectors" in the database. No temporary processing of historical session text is needed, and only one calculation is needed. Compared with large language models, the use of vector models requires less computing resources and is faster. At the same time, the pre-computation and storage of vectors are completed before the user initiates the current question, so the user will not feel additional delays such as waiting for the system to analyze historical data to ensure smooth conversation.
[0056] In step S13, the logical relevance between each second historical session in the second historical session candidate set and the current user request is analyzed using a pre-set large language model to obtain an analysis result, and a corresponding relevance label is marked for the second historical session according to the analysis result. The relevance label includes a strong relevance label representing a pre-set strong relevance, a partial relevance label representing a pre-set partial relevance, and an irrelevant label representing no relevance.
[0057] In this embodiment, the target dialogue scenario corresponding to the current user request can be determined first, and the logical relevance between each second historical session in the second historical session candidate set and the current user request is analyzed in combination with the target dialogue scenario and using a pre-set large language model. If there is a target historical session in the second historical session that covers all the information required by the current user request, the target historical session is taken as a strong relevant session corresponding to the current user request. If there is a target historical session in the second historical session that covers part of the information required by the current user request, the target historical session is taken as a partially relevant session corresponding to the current user request. If there is a target historical session in the second historical session that is irrelevant to the current user request, the target historical session is taken as an irrelevant session corresponding to the current user request.
[0058] Further, a strong correlation label can be assigned to a strongly related session in the second historical session candidate set, a partially related label can be assigned to a partially related session in the second historical session candidate set, and an irrelevant label can be assigned to an irrelevant session in the second historical session candidate set.
[0059] It can be understood that, in the logical correlation judgment stage, if all information required by A n to cover Q, the A n is marked as “strongly related”, if only part of the information required by A n to cover Q or additional supplementary information is needed, the A n is marked as “partially related”, and if A n is irrelevant to Q, the A n is marked as “irrelevant”.
[0060] In step S14, a target historical session associated with the current user request is determined based on the correlation labels, so as to respond to the current user request by using the target historical session.
[0061] In this embodiment, the sessions in the second historical session candidate set marked with the strong correlation label and the partially related label can be used as the target historical session associated with the current user request, so as to respond to the current user request by using the target historical session.
[0062] In this way, by focusing on the “strongly related” and “partially related” target historical sessions, the system can accurately locate the semantic demand of the current request, ensure that the answer is based on the historical context and is not disturbed by irrelevant information, and improve the coherence and accuracy of multi-round dialogue.
[0063] As can be seen from the above, in this embodiment, first, keyword matching filtering is performed on historical session records according to a current user request and by using a sparse vector retrieval algorithm, to obtain a first historical session candidate set. Then, theme consistency retrieval filtering is performed on the first historical session candidate set according to the current user request and by using a dense vector retrieval algorithm, to further filter to obtain a second historical session candidate set. Finally, logical coherence judgment filtering is performed on the second historical session candidate set by using a large language model fine-tuned for a scene + scene prompt word, according to the second historical session candidate set and the current user request. In this way, by using the pre-computed preliminary screening combined with the large language model, compared with the traditional historical session filtering method, the accuracy of the model historical session memory is improved, the interference of irrelevant memory on the model answer is reduced, and the accuracy and continuity of the content generated by the dialogue system reasoning are improved. At the same time, compared with the filtering method relying on a single large model, the computational amount and resource occupancy rate are reduced, and the user experience of the session is improved.
[0064] Referring to Figure 2As shown, the embodiment of the present application also discloses a large language model-based conversation response device, which comprises:
[0065] The data acquisition module 11 is configured to acquire a current user request and determine a request keyword and a request topic corresponding to the current user request.
[0066] The set determination module 12 is configured to determine, based on the request keyword, a first historical conversation with which the keyword matching degree of the current user request reaches a first preset matching degree threshold from historical conversation records, to construct a first historical conversation candidate set according to the first historical conversation, and determine, based on the request topic, a second historical conversation with which the topic matching degree of the current user request reaches a second preset matching degree threshold from the first historical conversation candidate set, to construct a second historical conversation candidate set according to the second historical conversation; wherein the historical conversation comprises a historical user request and a corresponding historical model answer.
[0067] The correlation analysis module 13 is configured to analyze, by using a preset large language model, the logical correlation between each second historical conversation in the second historical conversation candidate set and the current user request, to obtain an analysis result, and to mark the second historical conversation with a corresponding correlation label according to the analysis result; the correlation label comprises a strong correlation label representing that a preset strong correlation is met, a partial correlation label representing that a preset partial correlation is met, and an irrelevant label representing that no correlation is met.
[0068] The request response module 14 is configured to determine, based on the correlation label, a target historical conversation associated with the current user request, and to respond to the current user request by using the target historical conversation.
[0069] As can be seen from the above, the present application breaks through the round limit through the three-layer screening logic of “keyword matching-topic matching-logical correlation”, and no matter which round the historical conversation is in, as long as at least one of the core conversation keywords is contained in the request keyword, belongs to the same topic field as the current request, and has a logical correlation with the current request, the historical conversation will be marked as “strong correlation” or “partial correlation” and screened out, thereby avoiding the loss of cross-round key information. On the other hand, the present application first calculates the request keyword and the request topic corresponding to the current user request to preliminarily screen the full-amount historical conversation, greatly compresses the size of the candidate set, and only uses the large language model to analyze the logical correlation of the compressed candidate set, thereby avoiding the direct processing of the full-amount historical conversation by the large model, significantly reducing the calculation amount and resource consumption, while ensuring the screening accuracy, and solving the contradiction between the calculation efficiency and the filtering effect.
[0070] In some specific embodiments, the large language model-based conversation response device further comprises:
[0071] The request acquisition unit is configured to acquire an initial user request sent by a user terminal, and perform a preprocessing operation on the initial user request to obtain a current user request. The preprocessing operation includes irrelevant information removal, sentence clarification processing, sentence correction, and omission information filling.
[0072] In some embodiments, the large language model-based conversation response apparatus further includes:
[0073] The data extraction unit is configured to extract core conversation keywords and conversation topics corresponding to each historical conversation record by using a preset large language model, and save the historical conversation record and the corresponding core conversation keywords and conversation topics into a target database. The core conversation keywords are keywords that meet preset core keyword standards.
[0074] In some embodiments, the set determination module 12 includes:
[0075] The first matching degree analysis unit is configured to perform keyword matching degree analysis on the request keywords corresponding to the current user request and the core conversation keywords corresponding to each historical conversation record saved in the target database based on a preset sparse vector retrieval algorithm.
[0076] The first conversation determination unit is configured to, if there is a historical conversation in which at least one of the core conversation keywords is contained by the request keywords in the historical conversation record, determine the historical conversation in which at least one of the core conversation keywords is contained by the request keywords as a first historical conversation with a keyword matching degree reaching a first preset matching degree threshold with the current user request.
[0077] In some embodiments, the set determination module 12 includes:
[0078] The second matching degree analysis unit is configured to perform topic matching degree analysis on the request topic corresponding to the current user request and the conversation topics corresponding to each first historical conversation in the first historical conversation candidate set based on a preset dense vector retrieval algorithm.
[0079] The second conversation determination unit is configured to, if there is a first historical conversation in the first historical conversation candidate set that is consistent with the request topic, determine the first historical conversation that is consistent with the request topic as a second historical conversation with a topic matching degree reaching a second preset matching degree threshold with the current user request.
[0080] In some embodiments, the relevance analysis module 13 includes:
[0081] a logic analysis unit configured to determine a target dialogue scenario corresponding to the current user request, analyze logical relevance between each of the second historical conversations and the current user request in the second historical conversation candidate set in combination with the target dialogue scenario and by using a preset large language model;
[0082] a first processing unit configured to, if there is a target historical conversation in the second historical conversation that covers all information required by the current user request, take the target historical conversation as a strongly relevant conversation corresponding to the current user request;
[0083] a second processing unit configured to, if there is a target historical conversation in the second historical conversation that covers part of the information required by the current user request, take the target historical conversation as a partially relevant conversation corresponding to the current user request;
[0084] a third processing unit configured to, if there is a target historical conversation in the second historical conversation that is irrelevant to the current user request, take the target historical conversation as an irrelevant conversation corresponding to the current user request.
[0085] In some embodiments, the relevance analysis module 13 comprises:
[0086] a label assigning unit configured to assign a strongly relevant label to a strongly relevant conversation in the second historical conversation candidate set, assign a partially relevant label to a partially relevant conversation in the second historical conversation candidate set, and assign an irrelevant label to an irrelevant conversation in the second historical conversation candidate set;
[0087] Correspondingly, the request response module 14 comprises:
[0088] a conversation determining unit configured to take a conversation in the second historical conversation candidate set that is marked with a strongly relevant label and a partially relevant label as a target historical conversation associated with the current user request.
[0089] Further, the embodiments of the present application also disclose an electronic device, Figure 3 is a structural diagram of an electronic device 20 according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation on the use range of the present application.
[0090] Figure 3A structural schematic diagram of an electronic device 20 is provided in the embodiments of the present application. The electronic device 20 can specifically include at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25 and a communication bus 26. The memory 22 is configured to store a computer program, and the processor 21 is configured to load and execute the computer program to implement the related steps in the large language model-based session response method disclosed in any of the preceding embodiments. In addition, the electronic device 20 in the embodiments can be an electronic computer.
[0091] In the embodiments, the power supply 23 is configured to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol followed by the communication interface 24 can be any communication protocol applicable to the technical solution of the present application, which is not limited specifically herein; the input / output interface 25 is configured to obtain external input data or output data to the outside, and the specific interface type can be selected according to the specific application needs, which is not limited specifically herein.
[0092] In addition, the memory 22 as a carrier for resource storage can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage mode can be temporary storage or permanent storage.
[0093] The operating system 221 is configured to manage and control each hardware device on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the large language model-based session response method executed by the electronic device 20 disclosed in any of the preceding embodiments, the computer program 222 can further include a computer program capable of completing other specific work.
[0094] Further, the present application also discloses a computer readable storage medium for storing a computer program; wherein the computer program is executed by a processor to implement the large language model-based session response method disclosed in the preceding embodiments. The specific steps of the method can refer to the corresponding contents disclosed in the preceding embodiments, which will not be repeated here.
[0095] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can refer to the method part.
[0096] Those skilled in the art will further appreciate that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their functionality, which has been described generally and symbolically in flow charts. Having thus described the functionality of the examples, a person of ordinary skill in the art will be able to implement such functions in hardware and / or software, using the means and methods available to those skilled in the art. The examples described herein are not meant to limit the scope of the application, but merely to provide examples of the methods and systems being described.
[0097] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0098] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are more especially used for the purpose of distinction from other elements in the specification. Also, the terms "comprise", "include" or "contain" or any other variant thereof are intended to encompass non-exclusive inclusions, such that processes, methods, articles, or apparatuses that comprise, include, or contain a list of elements are not limited to those elements, but can include other elements not expressly listed or inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element defined by the phrase "comprising a... " does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0099] The above has introduced the technical solutions provided by the present application in detail, and the principles and implementation manners of the present application have been described by applying specific examples; the above example descriptions are only for helping to understand the method and core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will have changes; in conclusion, the content of the present description should not be understood as limiting the present application.
Claims
1. A conversational response method based on a large language model, characterized in that, include: Obtain the current user request and determine the request keyword and request topic corresponding to the current user request; Based on the requested keywords, a first historical session is determined from historical session records whose keyword matching degree with the current user's request reaches a first preset matching degree threshold, so as to construct a first historical session candidate set based on the first historical session. Based on the requested topic, a second historical session is determined from the first historical session candidate set whose topic matching degree with the current user's request reaches a second preset matching degree threshold, so as to construct a second historical session candidate set based on the second historical session. The historical session includes historical user requests and corresponding historical model answers. The logical correlation between each of the second historical sessions in the second historical session candidate set and the current user request is analyzed using a preset large language model to obtain analysis results. Based on the analysis results, the second historical sessions are labeled with corresponding correlation tags. The correlation tags include strongly correlated tags that represent a preset strong correlation, partially correlated tags that represent a preset partial correlation, and irrelevant tags that represent no correlation. Based on the relevance tags, a target historical session associated with the current user request is determined, so as to respond to the current user request using the target historical session.
2. The conversational response method based on a large language model according to claim 1, characterized in that, Also includes: Obtain the initial user request sent by the user terminal, and perform preprocessing operations on the initial user request to obtain the current user request; The preprocessing operations include removing irrelevant information, clarifying statements, correcting statements, and filling in omitted information.
3. The conversational response method based on a large language model according to claim 1, characterized in that, Also includes: The core conversation keywords and conversation topics corresponding to each historical conversation record are extracted using a preset large language model, and the historical conversation records and the corresponding core conversation keywords and conversation topics are saved to the target database; the core conversation keywords are keywords that meet the preset core keyword criteria.
4. The conversational response method based on a large language model according to claim 3, characterized in that, The step of determining the first historical session from historical session records based on the requested keyword, where the keyword matching degree with the current user's request reaches a first preset matching degree threshold, includes: Based on a preset sparse vector retrieval algorithm, keyword matching degree analysis is performed on the request keywords corresponding to the current user request and the core session keywords corresponding to each historical session record stored in the target database. If there is a historical session in the historical session record where at least one core session keyword is contained in the requested keyword, then the historical session in which at least one core session keyword is contained in the requested keyword is regarded as the first historical session whose keyword matching degree with the current user's request reaches a first preset matching degree threshold.
5. The conversational response method based on a large language model according to claim 1, characterized in that, The step of determining a second historical session from the first historical session candidate set based on the request topic, where the topic matching degree with the current user's request reaches a second preset matching degree threshold, includes: Based on a preset dense vector retrieval algorithm, the topic matching degree analysis is performed on the request topic corresponding to the current user request and the session topics corresponding to each first historical session in the first historical session candidate set; If there is a first historical session in the first historical session candidate set that matches the topic of the request, then the first historical session that matches the topic of the request is regarded as the second historical session whose topic matching degree with the current user's request reaches the second preset matching degree threshold.
6. The conversational response method based on a large language model according to claim 1, characterized in that, The step of using a preset large language model to analyze the logical correlation between each of the second historical sessions in the second historical session candidate set and the current user request, and obtaining the analysis results, includes: Determine the target dialogue scenario corresponding to the current user request, and analyze the logical correlation between each of the second historical sessions in the second historical session candidate set and the current user request by combining the target dialogue scenario and using a preset large language model; If there exists a target historical session in the second historical session that covers all the information required for the current user request, then the target historical session is regarded as a strongly related session corresponding to the current user request. If there is a target historical session in the second historical session that covers some of the information required for the current user request, then the target historical session is regarded as a partially related session corresponding to the current user request; If there is a target historical session in the second historical session that is not related to the current user request, then the target historical session is regarded as an unrelated session corresponding to the current user request.
7. The conversational response method based on a large language model according to claim 6, characterized in that, The step of assigning relevant tags to the second historical session based on the analysis results includes: Assign strongly relevant labels to strongly relevant sessions in the second historical session candidate set, assign partially relevant labels to partially relevant sessions in the second historical session candidate set, and assign irrelevant labels to irrelevant sessions in the second historical session candidate set. Accordingly, determining the target historical session associated with the current user request based on the relevance label includes: Sessions marked with strong and partial relevant tags in the second set of historical session candidates are selected as target historical sessions associated with the current user request.
8. A conversational response device based on a large language model, characterized in that, include: The data acquisition module is used to acquire the current user request and determine the request keywords and request topic corresponding to the current user request; The set determination module is used to determine, based on the requested keywords, a first historical session from historical session records whose keyword matching degree with the current user's request reaches a first preset matching degree threshold, so as to construct a first historical session candidate set based on the first historical session; and to determine, based on the requested topic, a second historical session from the first historical session candidate set whose topic matching degree with the current user's request reaches a second preset matching degree threshold, so as to construct a second historical session candidate set based on the second historical session; wherein, historical sessions include historical user requests and corresponding historical model answers; The relevance analysis module is used to analyze the logical relevance between each of the second historical sessions in the second historical session candidate set and the current user request using a preset large language model, obtain the analysis results, and mark the second historical sessions with corresponding relevance tags according to the analysis results; the relevance tags include strongly relevant tags that represent a preset strong relevance, partially relevant tags that represent a preset partial relevance, and irrelevant tags that represent no relevance; The request-response module is used to determine the target historical session associated with the current user request based on the relevance label, so as to respond to the current user request using the target historical session.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the large language model-based conversational response method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the conversational response method based on a large language model as described in any one of claims 1 to 7.