Response reference extraction method and device based on LCS algorithm, equipment and storage medium
By using the LCS algorithm in the Q&A system of the large language model, the reference information in the response generated by the large language model is extracted, and the problem of the model generation answering original documents that cannot be clearly cited is solved, which accurately locates and traces the response content, and improves the credibility of the response.
Patent Information
- Application Number
- CN202510117382.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
AI Technical Summary
The answers generated by the large language model cannot clarify the original document it refers to, and it is difficult to accurately locate the referenced fragments corresponding to the original document in the generated free text.
The response reference extraction method based on the LCS algorithm is adopted to receive the response content generated by the large language model, divide it into a response statement list, traverse all document fragments, and determine the longest common substring between the response statement and the document fragment through the LCS algorithm, thereby determining the document fragment associated with the response statement and outputting the reference list.
It realizes finding the longest common substring between the generated response and the original document fragment, accurately positioning the reference content, providing the original document position of its reference for each sentence generated, and improving the credibility of the response.
Smart Images

Figure CN120030125A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large language model question answering technology, and in particular to a response reference extraction method, device, equipment and storage medium based on an LCS algorithm. Background Art
[0002] With the development of large language models, they have shown powerful capabilities in areas such as dialogue generation and question-answering systems. However, these models are often regarded as "black boxes", and it is difficult for users to know where the generated answers come from. Especially when the model is combined with retrieval enhancement (RAG) technology, how to associate the generated responses with the original documents becomes a major challenge.
[0003] Currently, a common approach is to input the retrieved relevant documents into the model and let the model generate answers containing this information. However, this approach has the following shortcomings: Lack of precise reference location: The text generated by the model may contain the content of the original document, but it is impossible to determine which parts are specifically referenced. Failure to meet traceability requirements: In some application scenarios, users need to know the source of the answer to verify its reliability and authority.
[0004] The above contents are only used to assist in understanding the technical solution of the present invention and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of the present invention is to provide a response reference extraction method, device, equipment and storage medium based on the LCS algorithm, aiming to solve the technical problem that the answers generated by the current large language model cannot clearly identify the original document they refer to, and it is difficult to accurately locate the reference fragment corresponding to the original document in the generated free text.
[0006] To achieve the above object, the present invention provides a response reference extraction method based on the LCS algorithm, and the response reference extraction method based on the LCS algorithm comprises the following steps:
[0007] Receiving response content generated by the large language model, and obtaining a response sentence list based on the response content;
[0008] For each response statement in the response statement list, traverse all retrieved document fragments;
[0009] Determine the longest common substring between each response statement and each document fragment based on the LCS algorithm;
[0010] The document fragment associated with each response statement is determined according to the longest common substring, and a corresponding reference list is output.
[0011] In some embodiments, obtaining a list of response statements based on the response content includes:
[0012] Split the response content into lines to obtain a number of response statements;
[0013] Filter out blank lines from the obtained number of response statements to obtain a number of processed response statements;
[0014] Obtain a response statement list based on the number of processed response statements.
[0015] In some embodiments, the determining the longest common substring between each response statement and each document fragment based on the LCS algorithm includes:
[0016] Use the dynamic programming algorithm to determine the corresponding longest common substring based on the length, start index, and end index of the longest common substring in each response statement and each document fragment. The longest common substring is the longest consecutive matching part between two strings in the response statement and the document fragment, and the two strings correspond to the start index and the end index respectively.
[0017] In some embodiments, the determining the document fragments associated with each response statement according to the longest common substring includes:
[0018] Determine the length of the longest common substring of any response statement;
[0019] Compare the length of the longest common substring with a preset length threshold;
[0020] Based on the comparison result, screen out the document fragments associated with each response statement from all document fragments.
[0021] In some embodiments, the screening out the document fragments associated with each response statement from all document fragments based on the comparison result includes:
[0022] If the comparison result is that the length of the longest common substring is greater than or equal to the preset length threshold, record the corresponding document fragment.
[0023] In some embodiments, the outputting the corresponding reference list includes:
[0024] Obtain the association information between the response sentence and the document fragment. The association information includes the identifier of the fragment, the matching content, the matching length, and the start and end positions of the match in the response sentence;
[0025] Based on the association information, associate the response sentence with the document fragment to obtain a mapping relationship, store the association information in a list, and output it.
[0026] In some embodiments, the method further includes:
[0027] If the comparison result is that the length of the longest common substring is less than the preset length threshold, the corresponding document fragment is ignored.
[0028] In addition, to achieve the above object, the present invention further proposes a response reference extraction device based on the LCS algorithm, the response reference extraction device based on the LCS algorithm comprising:
[0029] A receiving module, used to receive the response content generated by the large language model, and obtain a response sentence list based on the response content;
[0030] A query module, configured to traverse all retrieved document fragments for each response statement in the response statement list;
[0031] A calculation module, used for determining the longest common substring between each response statement and each document fragment based on an LCS algorithm;
[0032] The extraction module is used to determine the document fragments associated with each response statement according to the longest common substring, and output a corresponding reference list.
[0033] In addition, to achieve the above-mentioned purpose, the present invention also proposes a response reference extraction device based on the LCS algorithm, and the response reference extraction device based on the LCS algorithm includes: a memory, a processor, and a response reference extraction program based on the LCS algorithm stored in the memory and executable on the processor, and the response reference extraction program based on the LCS algorithm is configured to implement the steps of the response reference extraction method based on the LCS algorithm as described above.
[0034] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which a response reference extraction program based on the LCS algorithm is stored. When the response reference extraction program based on the LCS algorithm is executed by a processor, the steps of the response reference extraction method based on the LCS algorithm as described above are implemented.
[0035] The present invention receives the response content generated by the large language model, and obtains a response statement list based on the response content; traverses all retrieved document fragments for each response statement in the response statement list; determines the longest common substring between each response statement and each document fragment based on the LCS algorithm; determines the document fragment associated with each response statement based on the longest common substring, and outputs the corresponding reference list. In the above manner, by using the LCS algorithm, the longest common substring can be found between the generated response and the original document fragment, the reference content can be accurately located, and the original document location referenced can be provided for each generated sentence, thereby improving the credibility of the response. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flowchart of the first embodiment of the response reference extraction method based on the LCS algorithm of the present invention;
[0037] Figure 2 It is a schematic diagram of the overall technical solution flow in the response reference extraction method based on the LCS algorithm of the present invention;
[0038] Figure 3 It is a structural block diagram of the first embodiment of the response reference extraction device based on the LCS algorithm of the present invention.
[0039] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0040] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0041] The embodiment of the present invention provides a response reference extraction method based on the LCS algorithm, referring to Figure 1 , Figure 1 The figure is a flow chart of a first embodiment of a response reference extraction method based on an LCS algorithm of the present invention.
[0042] In this embodiment, the response reference extraction method based on the LCS algorithm includes the following steps:
[0043] Step S10: receiving the response content generated by the large language model, and obtaining a response sentence list based on the response content.
[0044] In this embodiment, the executor of this embodiment is a response reference extraction device based on the LCS algorithm, wherein the response reference extraction device based on the LCS algorithm has functions such as data processing, data communication and program running. The response reference extraction device based on the LCS algorithm can be a computer terminal device or other network device, and of course it can also be other devices with similar functions, and this embodiment does not limit this.
[0045] It should be noted that with the development of large-scale language models, they have demonstrated powerful capabilities in areas such as dialogue generation and question-answering systems. However, these models are often regarded as "black boxes", and it is difficult for users to know where the generated answers come from. Especially when the model is combined with retrieval enhancement (RAG) technology, how to associate the generated responses with the original documents becomes a major challenge. At present, the common method is to input the retrieved relevant documents into the model and let the model generate answers containing this information. However, this method has the following shortcomings: Lack of precise reference positioning: The text generated by the model may contain the content of the original document, but it is impossible to determine which parts are specifically quoted. Unable to meet traceability requirements: In some application scenarios, users need to know the source of the answer to verify its reliability and authority.
[0046] In order to solve the above technical problems, the present embodiment receives the response content generated by the large language model, and obtains a response statement list based on the response content; traverses all retrieved document fragments for each response statement in the response statement list; determines the longest common substring between each response statement and each document fragment based on the LCS algorithm; determines the document fragment associated with each response statement based on the longest common substring, and outputs the corresponding reference list. In the above manner, by using the LCS algorithm, the longest common substring can be found between the generated response and the original document fragment, the reference content can be accurately located, and the original document location referenced can be provided for each generated sentence, thereby improving the credibility of the response.
[0047] In this embodiment, this embodiment first combines Figure 2 The overall process of this solution is described in detail. Figure 2 As shown, first obtain the response sentence fragments output by the model, traverse the response sentence and the list of document fragments retrieved one by one, then use the LCD algorithm to calculate the longest common substring, determine whether the length of the longest common substring exceeds the length threshold, record the referenced document fragments according to the judgment result, and finally output the reference list after completing all traversals. Among them, the threshold setting and comparison is to set the matching length threshold: to ensure the reliability of the match, set a minimum matching length threshold (such as 6 characters). Matching result screening: If the calculated longest common substring length is less than the threshold, it is considered that the match has no practical significance and the fragment is ignored. Otherwise, it is considered that the response sentence references the document fragment. Reference relationship recording: Information storage: For matches that meet the threshold conditions, record the association information between the response sentence and the document fragment, including the fragment identifier, the matched content, the matching length, and the starting and ending positions of the match in the response sentence.
[0048] To facilitate understanding, the above process is further illustrated with an example. Assume that the response content generated by the model is: The equipment makes abnormal noise during operation, which may be caused by bearing wear. It is recommended to check the lubrication condition. The retrieved document fragments include: 1. "Possible faults during equipment operation include bearing wear." 2. "Poor lubrication may cause abnormal noise in the equipment." Through the above algorithm, the longest common substring between each response sentence and the retrieval fragment is calculated. For the first response sentence "The equipment makes abnormal noise during operation, which may be caused by bearing wear." The longest common substring with the first retrieval fragment is "Possible faults during equipment operation include bearing wear", the matching length exceeds the threshold, and the reference relationship is recorded. The longest common substring length with the second retrieval fragment does not reach the threshold, so it is ignored. For the second response sentence "It is recommended to check the lubrication condition." The longest common substring with the second retrieval fragment is "lubrication", the matching length does not reach the threshold, so it is ignored. Finally, the system recognizes that the first sentence references the first retrieval fragment and records the relevant information.
[0049] In a specific implementation, in this embodiment, the response content generated by the large language model is first received, and a response statement list can be obtained based on the response content. Specifically, in this embodiment, the response needs to be preprocessed first, and the preprocessing is content segmentation and cleaning. First, the response content is segmented according to lines to obtain a number of response statements, and then the blank lines are filtered for the several response statements obtained by the segmentation to obtain the processed response statements. After filtering the blank lines, a more refined sentence set can be obtained, which improves the accuracy of subsequent association analysis. The processed response statements constitute the response statement list.
[0050] Step S20: traverse all retrieved document fragments for each response statement in the response statement list.
[0051] It should be noted that the matching is double-iterative: for each response sentence, the system traverses all retrieved document fragments. This process aims to compare the response content with possible reference sources and establish potential reference relationships.
[0052] Step S30: Determine the longest common substring between each response statement and each document fragment based on the LCS algorithm.
[0053] In a specific implementation, a dynamic programming algorithm is used to determine the corresponding longest common substring based on the length, start index and end index of the longest common substring in each response statement and each document fragment. The longest common substring is the longest continuous matching part between two strings in the response statement and the document fragment, and the two strings correspond to the start index and the end index respectively. The above dynamic programming method avoids repeated calculations, improves the efficiency of the algorithm, and can find the longest continuous matching fragment, thereby improving the accuracy of reference recognition.
[0054] Step S40: Determine the document fragment associated with each response statement according to the longest common substring, and output a corresponding reference list.
[0055] In a specific implementation, the length of the longest common substring of any response statement is determined; the length of the longest common substring is compared with a preset length threshold; based on the comparison result, the document fragments associated with each response statement are filtered out from all document fragments, wherein the preset length threshold can be set accordingly according to actual needs, and this is not limited in this embodiment.
[0056] Further, if the comparison result is that the length of the longest common substring is greater than or equal to the preset length threshold, the corresponding document fragment is recorded, wherein outputting the corresponding reference list specifically obtains the association information between the response sentence and the document fragment, the association information includes the fragment identifier, the matched content, the match length, and the start and end positions of the match in the response sentence; based on the association information, the response sentence is associated with the document fragment to obtain a mapping relationship, and the association information is stored in a list and output. It should be noted that in this embodiment, the response sentence is associated with the fragment to associate each response sentence with the information of all document fragments it references to form a mapping relationship. The final data organization uses an appropriate data structure (such as a dictionary or list) to store the association information for subsequent query and display.
[0057] Furthermore, if the comparison result is that the length of the longest common substring is less than a preset length threshold, the corresponding document fragment is ignored.
[0058] What is finally generated is a structured result, specifically a list containing the response sentence and the corresponding quoted fragment information is output as the final result, and supports subsequent processing. For example, the result can be used to display to the user, or for further analysis, such as generating citation annotations or highlighting the quoted content.
[0059] In this embodiment, the response content generated by the large language model is received, and a response statement list is obtained based on the response content; all retrieved document fragments are traversed for each response statement in the response statement list; the longest common substring between each response statement and each document fragment is determined based on the LCS algorithm; the document fragment associated with each response statement is determined based on the longest common substring, and the corresponding reference list is output. In the above manner, by using the LCS algorithm, the longest common substring can be found between the generated response and the original document fragment, the reference content can be accurately located, and the original document location referenced can be provided for each generated sentence, thereby improving the credibility of the response.
[0060] In addition, an embodiment of the present invention further proposes a storage medium, on which a response reference extraction program based on the LCS algorithm is stored. When the response reference extraction program based on the LCS algorithm is executed by a processor, the steps of the response reference extraction method based on the LCS algorithm as described above are implemented.
[0061] Reference Figure 3 , Figure 3 It is a structural block diagram of the first embodiment of the response reference extraction device based on the LCS algorithm of the present invention.
[0062] like Figure 3 As shown, the response reference extraction device based on the LCS algorithm proposed in the embodiment of the present invention includes:
[0063] A receiving module 10, configured to receive response content generated by a large language model, and obtain a response statement list based on the response content;
[0064] A query module 20, configured to traverse all retrieved document fragments for each response statement in the response statement list;
[0065] A calculation module 30, configured to determine the longest common substring between each response statement and each document fragment based on an LCS algorithm;
[0066] The extraction module 40 is used to determine the document fragment associated with each response statement according to the longest common substring, and output a corresponding reference list.
[0067] In this embodiment, the response content generated by the large language model is received, and a response statement list is obtained based on the response content; all retrieved document fragments are traversed for each response statement in the response statement list; the longest common substring between each response statement and each document fragment is determined based on the LCS algorithm; the document fragment associated with each response statement is determined based on the longest common substring, and the corresponding reference list is output. In the above manner, by using the LCS algorithm, the longest common substring can be found between the generated response and the original document fragment, the reference content can be accurately located, and the original document location referenced can be provided for each generated sentence, thereby improving the credibility of the response.
[0068] In some embodiments, the receiving module 10 is used to divide the response content into lines to obtain a plurality of response statements;
[0069] Filter the obtained response statements by blank lines to obtain the processed response statements;
[0070] A response statement list is obtained based on the processed response statements.
[0071] In some embodiments, the calculation module 30 is used to use a dynamic programming algorithm to determine the corresponding longest common substring based on the length, starting index and ending index of the longest common substring in each response statement and each document fragment, wherein the longest common substring is the longest continuous matching part between two strings in the response statement and the document fragment, and the two strings correspond to the starting index and the ending index respectively.
[0072] In some embodiments, the extraction module 40 is used to determine the length of the longest common substring of any response statement;
[0073] Comparing the length of the longest common substring with a preset length threshold;
[0074] Based on the comparison result, the document fragments associated with each response statement are filtered out from all the document fragments.
[0075] In some embodiments, the extraction module 40 is configured to record a corresponding document segment if the comparison result is that the length of the longest common substring is greater than or equal to a preset length threshold.
[0076] In some embodiments, the extraction module 40 is used to obtain association information between the response sentence and the document fragment, the association information including the fragment identifier, the matched content, the matched length, and the start and end positions of the match in the response sentence;
[0077] The response sentence is associated with the document fragment based on the association information to obtain a mapping relationship, and the association information is stored in a list and outputted.
[0078] In some embodiments, the extraction module 40 is configured to ignore the corresponding document segment if the comparison result shows that the length of the longest common substring is less than a preset length threshold.
[0079] An embodiment of the present application also provides a response reference extraction device based on the LCS algorithm, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus, and the memory is used to store a response reference extraction program based on the LCS algorithm; the processor is used to implement the above-mentioned response reference extraction method based on the LCS algorithm when executing the program stored in the memory.
[0080] The communication bus mentioned in the above-mentioned LCS algorithm-based response reference extraction device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0081] The communication interface is used for communication between the above-mentioned response reference extraction device based on the LCS algorithm and other devices.
[0082] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0083] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components.
[0084] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.
[0085] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0086] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0087] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0088] It should be understood that the above is only an example and does not constitute any limitation on the technical solution of the present invention. In specific applications, technicians in this field can make settings as needed, and the present invention does not limit this.
[0089] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of the present invention. In practical applications, technicians in this field can select part or all of them according to actual needs to achieve the purpose of the present embodiment, and no limitation is made here.
[0090] In addition, for technical details not described in detail in this embodiment, please refer to the response reference extraction method based on the LCS algorithm provided in any embodiment of the present invention, which will not be repeated here.
[0091] In addition, it should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0092] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory (ROM) / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0094] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
[0095] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above method.
Claims
1. A response reference extraction method based on LCS algorithm, characterized in that: The response reference extraction method based on the LCS algorithm includes: Receiving response content generated by the large language model, and obtaining a response sentence list based on the response content; For each response statement in the response statement list, traverse all retrieved document fragments; Determine the longest common substring between each response statement and each document fragment based on the LCS algorithm; The document fragment associated with each response statement is determined according to the longest common substring, and a corresponding reference list is output.
2. The response reference extraction method based on the LCS algorithm as claimed in claim 1, characterized in that: The obtaining a response statement list based on the response content includes: Splitting the response content into rows to obtain a plurality of response statements; Filter the obtained response statements by blank lines to obtain the processed response statements; A response statement list is obtained based on the processed response statements.
3. The response reference extraction method based on the LCS algorithm as claimed in claim 1, characterized in that: The determining the longest common substring between each response statement and each document fragment based on the LCS algorithm includes: A dynamic programming algorithm is used to determine the corresponding longest common substring based on the length, starting index and ending index of the longest common substring in each response statement and each document fragment. The longest common substring is the longest continuous matching part between two strings in the response statement and the document fragment, and the two strings correspond to the starting index and the ending index respectively.
4. The response reference extraction method based on the LCS algorithm as claimed in claim 1, characterized in that: The step of determining the document fragments associated with each response statement according to the longest common substring includes: Determine the length of the longest common substring of any response statement; Comparing the length of the longest common substring with a preset length threshold; Based on the comparison result, the document fragments associated with each response statement are filtered out from all the document fragments.
5. The response reference extraction method based on the LCS algorithm as claimed in claim 4, characterized in that: The step of selecting document fragments associated with each response statement from all document fragments based on the comparison result includes: If the comparison result is that the length of the longest common substring is greater than or equal to the preset length threshold, the corresponding document fragment is recorded.
6. The response reference extraction method based on the LCS algorithm as claimed in claim 5, characterized in that: The output corresponds to a reference list, including: Acquire association information between the response sentence and the document fragment, the association information including an identifier of the fragment, matched content, match length, and start and end positions of the match in the response sentence; The response sentence is associated with the document fragment based on the association information to obtain a mapping relationship, and the association information is stored in a list and outputted.
7. The response reference extraction method based on the LCS algorithm as claimed in claim 5, characterized in that: The method further comprises: If the comparison result is that the length of the longest common substring is less than the preset length threshold, the corresponding document fragment is ignored.
8. A response reference extraction device based on LCS algorithm, characterized in that: The response reference extraction device based on the LCS algorithm comprises: A receiving module, used to receive the response content generated by the large language model, and obtain a response sentence list based on the response content; A query module, configured to traverse all retrieved document fragments for each response statement in the response statement list; A calculation module, used for determining the longest common substring between each response statement and each document fragment based on an LCS algorithm; The extraction module is used to determine the document fragments associated with each response statement according to the longest common substring, and output a corresponding reference list.
9. A response reference extraction device based on LCS algorithm, characterized in that: The response reference extraction device based on the LCS algorithm includes: a memory, a processor, and a response reference extraction program based on the LCS algorithm stored in the memory and executable on the processor, wherein the response reference extraction program based on the LCS algorithm is configured to implement the steps of the response reference extraction method based on the LCS algorithm as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a response reference extraction program based on the LCS algorithm, and when the response reference extraction program based on the LCS algorithm is executed by the processor, the steps of the response reference extraction method based on the LCS algorithm as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Soft, prioritised early packet discard system
CA2311410A1
Target text determination method, device and equipment
CN111401031A
Large language model question and answer optimization method and device, electronic equipment and storage medium
CN117851575A
Question and answer result tracing method and device, equipment, medium and program product
CN117909451A
Question and answer information processing method and system, electronic equipment and storage medium
CN118535712A