Text fine tuning data set construction method and system based on retrieval enhancement generation and medium

By combining the methods of retrieval enhanced generation and dialogue-based large language models, the problem of building high-quality text fine-tuning data sets is solved, efficient and accurate data generation and screening is achieved, and the cost of manual intervention is reduced.

CN119988975APending Publication Date: 2025-05-13CAS SUZHOU INSTITUTE OF INTELLIGENT COMPUTING TECHNOLOGY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510108961.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently construct high-quality text fine-tuning data sets, especially in processing multi-source data, filtering noise and ensuring data consistency.

Method used

A high-quality text fine-tuning data set is constructed by using a method based on retrieval enhancement, combining dialogue-based large language model and vector model, and by generating data problems, building an index vector library, similarity search, data improvement, filtering and correction, clustering and other steps.

Benefits of technology

It significantly improves the efficiency of data generation and screening, ensures high quality and accuracy of data, reduces the need for manual intervention, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988975A_ABST
    Figure CN119988975A_ABST
Patent Text Reader

Abstract

The invention discloses a text fine-tuning data set construction method and system based on retrieval enhancement generation and a medium. The method comprises the steps that text data are collected and preprocessed to form first data, and the first data form a first data set; each piece of first data in the first data set generates at least one data problem, and the data problems form a second data set; the first data set and the second data set construct an index vector library through a vector model, and first data corresponding to the data problem is searched through similarity retrieval; perfecting the first data corresponding to each data problem, forming a data pair by the perfected first data and the corresponding data problem, and forming a third data set by the data pair; filtering and correcting data pairs in the third data set to form a fourth data set; clustering the fourth data set to form a fifth data set; and verifying the fifth data set to form a text fine tuning data set. And combining retrieval enhancement generation and a dialogue type large language model to construct a text fine tuning data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a method, system and medium for constructing a text fine-tuning dataset based on retrieval-enhanced generation. Background Art

[0002] With the increasing demand for large models in fields such as natural language processing and computer vision, architectures based on models such as Transformer, GPT, and BERT have become mainstream. In recent years, the scale and number of parameters of models have increased significantly, and training ultra-large-scale models has become common. The core of large language model fine-tuning technology is to adjust the general large model in a specific task or field to meet the task requirements. For many organizations and institutions that lack computing power resources for large model training, large language model fine-tuning has become a key step in the implementation of the model. The effect of fine-tuning depends on high-quality massive annotated data sets. However, as the diversity and complexity of tasks increase, the difficulty of building and processing text fine-tuning data sets is also increasing. For example, how to effectively process multi-source data and filter noise has become a key issue.

[0003] In the existing technology, although the use of GPT to generate data can quickly generate a large amount of content and reduce human intervention, there are problems with data quality and accuracy. It is often difficult to maintain consistency in the generated content, and factual errors may occur when facing complex scenarios or professional fields. In addition, the data generated by GPT is easy to be patterned, lacks sufficient diversity, and the generation process is difficult to fully control, especially in semantically complex tasks, the verification and traceability of the content are weak. On the other hand, traditional clustering methods can assist manual annotation to reduce some workload through grouping, but the efficiency improvement is limited. Clustering algorithms often cannot accurately reflect the semantic similarity of data, and their processing capabilities for complex tasks and long-tail data are insufficient, resulting in manual review that is still arduous. In addition, subjective differences are prone to occur in the annotation process, and the clustering results have poor adaptability when dealing with dynamic data, making it difficult to meet the rapid expansion needs of large-scale data sets. Therefore, how to efficiently and high-quality construct a text fine-tuning data set has become a technical problem to be solved. Summary of the invention

[0004] In order to overcome the above-mentioned shortcomings, the purpose of the present invention is to provide a method, system and medium for constructing a text fine-tuning dataset based on retrieval-enhanced generation, which adopts a method combining retrieval-enhanced generation and a conversational large language model to realize the construction of a high-quality, less-artificial text fine-tuning dataset.

[0005] In order to achieve the above objectives, the technical solution adopted by the present invention is: a method for constructing a text fine-tuning dataset based on retrieval enhancement generation, comprising: Collecting text data and preprocessing the text data to form first data, wherein the first data forms a first data set; Each of the first data in the first data set generates at least one data question through a conversational large language model, and the data questions form a second data set; The first data set and the second data set construct an index vector library through a vector model, and search for the first data corresponding to each of the data problems in the second data set through similarity retrieval; The conversational large language model improves the first data corresponding to each of the data questions, and the improved first data and the corresponding data question form a data pair, and the data pair forms a third data set; The conversational large language model filters and modifies the data pairs in the third data set to form a fourth data set; Clustering the fourth data set using a similarity discriminator to form a fifth data set; The fifth data set is verified to form the text fine-tuning data set after verification. The beneficial effects of the present invention are: Instead of relying on large text models to directly generate questions, it makes full use of its powerful semantic understanding capabilities to optimize and constrain the knowledge base content obtained from text retrieval technology. The questions generated by the conversational large language model can accurately reflect the core content of the text, thereby improving the overall accuracy of the data.

[0006] A combination of a conversational large language model and traditional text similarity matching methods is used to assist experts in data screening, reducing the need for manual intervention and significantly reducing labor costs.

[0007] It not only improves the efficiency of data generation and screening, but also ensures the high quality and accuracy of the data.

[0008] Furthermore, the similarity discriminator includes a text embedding model and agglomerative hierarchical clustering, and the algorithm uses the similarity discriminator to cluster the fourth data set to form the fifth data set, which specifically includes: The data pairs in the fourth data set are semantically analyzed by a text embedding model to identify core intents, and the fourth data set is classified according to core intent categories, and the frequency and weight of each core intent are quantified; The text embedding model pairs the data in the fourth dataset and generates a high-dimensional vector representation according to the fourth dataset and the core intent category; The agglomerative hierarchical clustering algorithm clusters the fourth data set according to the frequency and weight of the core intent to form the fifth data set.

[0009] Furthermore, the agglomerative hierarchical clustering algorithm clusters the fourth data set according to the frequency and weight of the core intent, specifically including: The obtained data of the same core intent is used to calculate the distance with the center of the cluster to which it belongs through Euclidean distance; Select the data with the smallest distance as the characteristic data point of the corresponding cluster; The two most similar clusters are merged until the similarity between all clusters is less than the set threshold.

[0010] Specifically, the first data set and the second data set construct an index vector library through a vector model, and searching for text data corresponding to each data problem in the second data set through similarity retrieval includes: Inputting the first data set into the vector model, generating a first vector corresponding to each first data, and storing the first vector as an index vector library; Input each data problem in the second data set into the vector model to generate a second vector corresponding to each data problem; The second vector is sequentially input into the index vector library as request data, and the first vector in the index vector library closest to the request data is obtained through Euclidean distance calculation of the similarity retrieval index, that is, the first data corresponding to the data problem in the second data set.

[0011] Furthermore, filtering and correcting the data pairs in the third data set to form a fourth data set specifically includes: Preliminarily filtering the data pairs in the third data set through a conversational large language model, storing the data pairs with problems as data set A, and storing the data pairs without problems as data set B; Correcting the data set A by using the conversational large language model, and storing the corrected data set A as a fourth data set; Perform semantic judgment on each data pair in data set B through a conversational large language model; If the data pair is reasonable, the data pair is stored in the fourth data set.

[0012] Furthermore, a conversational large language model is used to perform semantic judgment on each of the data pairs in data set B. If the data pair is unreasonable, the large language model is used to correct the data pair, and the corrected data pair is stored in the fourth data set.

[0013] Further, preprocessing the text data to form the first data specifically includes: The text data is cleaned using natural language processing technology to filter out emoticons, emoticons, and chapter headers to form the first data.

[0014] Furthermore, the text data can be stored in a variety of different formats.

[0015] The present invention also discloses a text fine-tuning dataset construction system based on retrieval enhancement generation, which adopts the above-mentioned text fine-tuning dataset construction method.

[0016] The present invention also discloses a computer-readable storage medium, on which a text data construction method program is stored. When the text data construction method program is executed by a processor, the steps of the above-mentioned text fine-tuning data set construction method are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 The process of the method in the embodiment of the present invention is Figure 1 ; Figure 2 The process of the method in the embodiment of the present invention is Figure 2 ; Figure 3 The process of the method in the embodiment of the present invention is Figure 3 ; Figure 4 The process of the method in the embodiment of the present invention is Figure 4 . DETAILED DESCRIPTION

[0018] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0019] In the text fine-tuning dataset construction method based on retrieval enhancement generation of the present invention, a conversational large language model has been deployed before construction. For the convenience of calling, the conversational large language model is deployed in the cloud. Exemplarily, the conversational large language model is based on the llama-7b framework.

[0020] See attached Figure 1 As shown in FIG. 1 , the method for constructing a text fine-tuning dataset based on retrieval enhancement generation includes: S100: Collect text data and pre-process the text data to form first data, which forms a first data set.

[0021] Collect large-scale text data from various sources, such as public databases, literature, and the Internet, and purchase through third-party channels or download directly. Text data can be saved in any format, such as books, Excel spreadsheets, papers, news releases, etc. Exemplarily, text data is saved in JSON or TXT format.

[0022] Preprocessing the text data to form the first data specifically includes: The text data is cleaned using natural language processing (NLP) technology to filter out useless information such as emoticons, emoticons, chapter headers, etc., to form the first data. The first data formed after cleaning is all useful information and is recorded as the first data set.

[0023] S200. Each first data in the first data set generates at least one data question through a conversational large language model, and the data questions form a second data set.

[0024] Based on the content of the first data set, a prompt is designed and a conversational large language model is used to generate a set of related questions. At least one data question is generated for each piece of first data, and the set of data questions is recorded as the second data set.

[0025] Exemplarily, a first data is: "Drinking too much water may dilute electrolytes in the body, leading to hyponatremia and muscle cramps, especially during high-intensity exercise. Sweating profusely during exercise will cause loss of water and electrolytes, but drinking too much water may cause electrolyte imbalance. Athletes should maintain a balance in water and electrolyte intake and drink electrolyte-containing beverages in moderation to support normal body functioning and prevent muscle cramps." The first data and prompt are input into the conversational large language model, and the conversational large language model outputs the corresponding data question: "Why does drinking too much water may lead to electrolyte imbalance?".

[0026] The first data in the first data set is stored in the first data set in a row-by-row or segment-by-segment manner.

[0027] S300, the first data set and the second data set construct an index vector library through a vector model, and search for the first data corresponding to each data problem in the second data set through similarity retrieval.

[0028] See attached Figure 2 As shown, step S300 specifically includes: S31. Input the first data set into the vector model, generate a first vector corresponding to each first data, and store the first vector as an index vector library.

[0029] Exemplarily, the vector model is a BGE (BAAI General Embedding) model, which first encodes each first data in the first data set to generate a high-dimensional vector. The first vector is indexed using the similarity retrieval FAISS to construct an efficient index vector library for subsequent similarity retrieval.

[0030] S32: Input each data problem in the second data set into the vector model to generate a second vector corresponding to each data problem.

[0031] Each data question in the second data set is encoded through the BGE model to generate a high-dimensional vector.

[0032] S33. The second vector is sequentially input into the index vector library as the request data, and the first vector in the index vector library closest to the request data is obtained through similarity retrieval, that is, the first data corresponding to the data problem in the second data set.

[0033] The second vector is input into the index vector library, and the closest first data is found through the Euclidean distance calculation of the FAISS index.

[0034] S400. The conversational large language model improves the first data corresponding to each data question, and the improved first data and the corresponding data question form a data pair, and the data pair forms a third data set.

[0035] The conversational large language model improves the content of the first data corresponding to the data question found to ensure that its answer is complete and reasonable. Finally, a question-answer data pair is formed to construct a preliminary question-answer pair data set, which is recorded as the third data set.

[0036] For example, the question: "Why can excessive drinking of water lead to electrolyte imbalance?", the answer: "Excessive drinking of water may dilute the electrolytes in the body, leading to hyponatremia and muscle cramps, especially during high-intensity exercise. Excessive sweating during exercise will cause loss of water and electrolytes, but drinking too much water may cause electrolyte imbalance. Athletes should maintain a balance in water and electrolyte intake and drink electrolyte-containing beverages in moderation to support normal body function and prevent muscle cramps" is a data pair.

[0037] S500: The conversational large language model filters and modifies the data pairs in the third data set to form a fourth data set.

[0038] There may be anomalies in the data pairs in the third data set, so a conversational large language model is needed to filter and correct the data in the third data set.

[0039] See attached Figure 3 As shown, step S500 specifically includes: S51. Preliminarily filter the data pairs in the third data set through the conversational large language model to determine whether there are problems with the data pairs.

[0040] The conversational large language model classifies the data pairs in the third data set, and returns the label 0 to the problematic data pairs, and otherwise returns the label 1. The labels are used to determine whether the data pairs have problems.

[0041] For example, for the data pair "Question: Apples grow in the ground, Answer: Please enter the text you want to optimize", the conversational language model returns 0, and this data pair is stored as dataset A. For the data pair "Question: Apples grow in the ground, Answer: No, they grow on apple trees", the conversational language model returns 1, and this data pair is stored as dataset B. This step is the error output caused by the conversational language model's misunderstanding of the prompt instruction.

[0042] S52: Store the data pairs with problems as data set A, and store the data pairs without problems as data set B.

[0043] S53: Correct the data set A by using the conversational large language model, and store the corrected data set A as a fourth data set.

[0044] The data pairs in dataset A are reconstructed or re-answered through the conversational large language model, so that the revised dataset A meets the question-answering requirements, and can therefore be stored as the fourth dataset for subsequent calls.

[0045] For example, the data pair "Question: Apples grow in the ground, Answer: Please enter the text you want to optimize" is corrected through the conversational large language model to the data pair "Question: Apples grow in the ground, Answer: Apples do not grow in the ground, but on trees. The apple tree is a deciduous fruit tree, and the apple fruits grow and ripen on the branches. Apple trees need sunlight, water and suitable soil conditions to grow healthily, and the fruits can be picked when they are ripe. Different plants, such as potatoes, grow directly underground. Therefore, apples are fruits that grow on trees, not in the ground."

[0046] S54. Determine whether the semantics of each data pair in data set B is reasonable through a conversational large language model.

[0047] S55: If the data pair is reasonable, store the data pair in the fourth data set.

[0048] S56: If the data pair is unreasonable, the data pair is corrected using the large language model, and the corrected data pair is stored in the fourth data set.

[0049] This step is to prevent mismatch between questions and answers during the retrieval process.

[0050] Steps S54-S56 are for processing data set B, and step S53 is for processing data set A. Steps S54-S56 and step S53 can be performed simultaneously. Through step S500, data set A and data set B are integrated as an optimized data pair to form a fourth data set.

[0051] S600: Clustering the fourth data set using a similarity discriminator to form a fifth data set.

[0052] The similarity discriminator includes a text embedding model and an agglomerative hierarchical clustering (AHC) algorithm. Exemplarily, the text embedding model is gte-Qwen1.5-7B-instruct, which is specially optimized for complex instruction-driven tasks and has excellent semantic reasoning and instruction execution capabilities. AHC is a bottom-up method that can form a hierarchical structure by iteratively merging the most similar data points, and adjust the weights in combination with the intent frequency information to reflect the semantic similarity and intent importance, ensuring that the AHC clustering process is based on semantic similarity. The similarity discriminator uses gte-Qwen1.5-7B-instruct for semantic analysis and label application, and clusters in combination with intent frequency and weight ratio. The AHC algorithm is used in combination with Euclidean distance calculation to cluster data points, and then the fourth data set is optimized and clustered.

[0053] See attached Figure 4 As shown, step S600 specifically includes: S61. The data pairs in the fourth data set are semantically analyzed through a text embedding model to identify core intents, and the fourth data set is classified according to core intent categories, and the frequency and weight of each core intent are quantified.

[0054] S62. The text embedding model pairs the data in the fourth dataset according to the fourth dataset and the core intent category and generates a high-dimensional vector representation.

[0055] High-dimensional vectors can preserve the original semantic information and reflect the semantic similarity between texts.

[0056] S63. Clustering the fourth data set according to the frequency and weight of the core intent using an agglomerative hierarchical clustering algorithm to form a fifth data set.

[0057] Agglomerative hierarchical clustering (AHC) is used to classify the data, and the weights are adjusted based on the intention distribution information to ensure that the clustering is based on semantic similarity. The most representative data point is selected in each cluster and its distance to the category center is calculated to ensure the diversity and representativeness of the data set.

[0058] Step S63 specifically includes: The obtained data in the same core intent are compared with the center of the cluster to which they belong using the Euclidean distance calculation.

[0059] The data with the smallest distance is selected as the characteristic data point of the corresponding cluster. The selection of this characteristic data point determines the diversity and representativeness of the data to prevent the deletion of data content with the same text description but different intentions.

[0060] The two most similar clusters are merged until the similarity between all clusters is less than the set threshold. The threshold can be set, and the threshold is 0.75 by way of example.

[0061] Ultimately, the diversity and representativeness of the fifth dataset are ensured, thereby optimizing the quality and expressiveness of the fifth dataset.

[0062] S700 , verifying the fifth data set, and forming a text fine-tuning data set after verification.

[0063] Data verification is an indispensable step in data engineering and is carried out by experts.

[0064] Steps S100-S700 generate text fine-tuning datasets based on retrieval enhancement generation and conversational large language models, reducing the labor costs required to build a prediction library from books, articles, and news data in the early stage; processing multi-source, multi-dimensional text datasets with unified specifications; ensuring the diversity of the dataset, the accuracy of the dataset, and the quality of the dataset. The text fine-tuning dataset constructed through the above process can be used for fine-tuning training of large models to improve the performance of the model in specific tasks (such as question answering, dialogue generation, etc.). Through this series of processing, it is ensured that the text fine-tuning dataset achieves the best balance in quality and diversity, making the training of subsequent models more effective.

[0065] Unlike traditional methods, this embodiment does not rely on large text models to directly generate questions, but makes full use of its powerful semantic understanding capabilities to optimize and constrain the knowledge base content obtained from text retrieval technology. This process ensures that the generated questions accurately reflect the core content of the text, thereby improving the overall accuracy of the data. During the review of the text fine-tuning dataset, a combination of a conversational large language model and a traditional text similarity matching method was used to assist experts in data screening. This innovation reduces the need for manual intervention and significantly reduces labor costs. In addition, by optimizing the data generation process, this embodiment also effectively saves the cost of generating data using the GPT API. Through the above improvements, this embodiment not only improves the efficiency of data generation and screening, but also ensures the high quality and accuracy of the data.

[0066] The present invention also discloses a text fine-tuning dataset construction system based on retrieval enhanced generation, including a conversational large language model, a vector model and a similarity discriminator deployed in the cloud. The text fine-tuning dataset construction system adopts the above-mentioned text fine-tuning dataset construction method to construct a fifth dataset, verifies the fifth dataset, and forms a text fine-tuning dataset after verification.

[0067] The present invention also discloses a computer-readable storage medium, on which a text fine-tuning data set construction method program is stored. When the text fine-tuning data set construction method program is executed by a processor, the steps of the above-mentioned text fine-tuning data set construction method are implemented. Based on such understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of each of the above-mentioned method embodiments when executed by a processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0068] The above implementation modes are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with this technology to understand the content of the present invention and implement it, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for constructing a text fine-tuning dataset based on retrieval-enhanced generation, characterized by: include: Collecting text data and preprocessing the text data to form first data, wherein the first data forms a first data set; Each of the first data in the first data set generates at least one data question through a conversational large language model, and the data questions form a second data set; The first data set and the second data set construct an index vector library through a vector model, and search for the first data corresponding to each of the data problems in the second data set through similarity retrieval; The conversational large language model improves the first data corresponding to each of the data questions, and the improved first data and the corresponding data question form a data pair, and the data pair forms a third data set; The conversational large language model filters and modifies the data pairs in the third data set to form a fourth data set; Clustering the fourth data set using a similarity discriminator to form a fifth data set; The fifth data set is verified to form the text fine-tuning data set after verification.

2. The method for constructing a text fine-tuning dataset according to claim 1, characterized in that: The similarity discriminator includes a text embedding model and agglomerative hierarchical clustering. The algorithm uses the similarity discriminator to cluster the fourth data set to form the fifth data set, which specifically includes: The data pairs in the fourth data set are semantically analyzed by a text embedding model to identify core intents, and the fourth data set is classified according to core intent categories, and the frequency and weight of each core intent are quantified; The text embedding model pairs the data in the fourth dataset and generates a high-dimensional vector representation according to the fourth dataset and the core intent category; The agglomerative hierarchical clustering algorithm clusters the fourth data set according to the frequency and weight of the core intent to form the fifth data set.

3. The method for constructing a text fine-tuning dataset according to claim 2, characterized in that: The agglomerative hierarchical clustering algorithm clusters the fourth data set according to the frequency and weight of the core intent, specifically including: The obtained data of the same core intent is used to calculate the distance with the center of the cluster to which it belongs through Euclidean distance; Select the data with the smallest distance as the characteristic data point of the corresponding cluster; The two most similar clusters are merged until the similarity between all clusters is less than the set threshold.

4. The method for constructing a text fine-tuning dataset according to claim 1, characterized in that: The first data set and the second data set construct an index vector library through a vector model, and searching for text data corresponding to each data problem in the second data set through similarity retrieval specifically includes: Inputting the first data set into the vector model, generating a first vector corresponding to each first data, and storing the first vector as an index vector library; Input each data problem in the second data set into the vector model to generate a second vector corresponding to each data problem; The second vector is sequentially input into the index vector library as request data, and the first vector in the index vector library closest to the request data is obtained through Euclidean distance calculation of the similarity retrieval index, that is, the first data corresponding to the data problem in the second data set.

5. The method for constructing a text fine-tuning dataset according to claim 1, characterized in that: Filtering and correcting the data pairs in the third data set to form a fourth data set specifically includes: Preliminarily filtering the data pairs in the third data set through a conversational large language model, storing the data pairs with problems as data set A, and storing the data pairs without problems as data set B; Correcting the data set A by using the conversational large language model, and storing the corrected data set A as a fourth data set; Perform semantic judgment on each data pair in data set B through a conversational large language model; If the data pair is reasonable, the data pair is stored in the fourth data set.

6. The method for constructing a text fine-tuning dataset according to claim 5, characterized in that: A semantic judgment is performed on each data pair in data set B through a conversational large language model. If the data pair is unreasonable, the data pair is corrected through the large language model, and the corrected data pair is stored in the fourth data set.

7. The method for constructing a text fine-tuning dataset according to claim 1, characterized in that: Preprocessing the text data to form first data specifically includes: The text data is cleaned using natural language processing technology to filter out emoticons, emoticons, and chapter headers to form the first data.

8. The method for constructing a text fine-tuning dataset according to claim 2, characterized in that: The text data can be stored in a variety of different formats.

9. A system for constructing a text fine-tuning dataset based on retrieval-enhanced generation, characterized by: The text fine-tuning dataset construction method according to any one of claims 1 to 8 is adopted.

10. A computer-readable storage medium, characterized in that: The readable storage medium stores a text data construction method program, and when the text data construction method program is executed by a processor, the steps of the text fine-tuning dataset construction method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Automatic question answering system and method based on large language model

    CN120541183A