Purchase data retrieval model fine tuning method and system
By automatically generating search problems and high-quality positive and negative sample pairs, and using large language models to fine-tune the procurement data retrieval model, the problem of inaccurate high-cost annotation and negative sample construction in the government procurement field is solved, and efficient model training and optimization are achieved.
Patent Information
- Application Number
- CN202411886145.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-16
AI Technical Summary
In areas with extremely high professionalism such as government procurement, the existing procurement data retrieval model fine-tuning methods face the problems of high positive sample annotation cost, inaccurate negative sample construction strategies and lack of effective utilization of large amounts of unlabeled data, resulting in limited training efficiency and retrieval performance.
By obtaining a large number of procurement documents and segmenting them, using pre-trained large language models for field recognition, and building a procurement field data set. Then, search questions related to text fragments are automatically generated, problem text sets are constructed, positive sample candidate sets and negative sample candidate sets are further constructed through large language models, and high-quality positive sample sets and negative sample sets are constructed through refining process, and finally fine-tuning the procurement data retrieval model.
This method greatly reduces labor costs and thresholds, improves the quality and difficulty of negative samples, enhances the understanding and discrimination of semantics in the procurement field, and enables the procurement data retrieval model to quickly adapt to regulatory updates and data changes, and improves training efficiency and retrieval performance.
Smart Images

Figure CN120012912A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of procurement data retrieval, and in particular to a procurement data retrieval model fine-tuning method and system. Background Art
[0002] In the field of natural language processing and information retrieval, with the widespread application of pre-trained language models (such as BERT and RoBERTa), text retrieval technology has made significant progress. Traditional keyword matching methods (such as BM25) have gradually given way to dense vector retrieval methods based on pre-trained language models; pre-trained language models vectorize queries and documents and use cosine similarity and other metrics for semantic retrieval, laying the foundation for efficient and accurate information extraction in large-scale knowledge bases. In recent years, a series of dense retrieval models (retrievers) based on BERT-like architectures have become mainstream, such as BGE, DPR, ANCE, and ColBERT. These retrieval models mostly use a dual encoder structure to generate vector representations for queries and documents, and then complete retrieval and ranking through vector similarity; these dense retrieval models are pre-trained on massive general corpora and fine-tuned for specific tasks, and have strong language understanding and representation capabilities, and can generate high-quality text embeddings for text matching and retrieval tasks.
[0003] The current training process of the retrieval model usually includes the following main stages: pre-training, fine-tuning, and evaluation. The training process of BGE, which is currently open source and widely used in the industry, is used as an example to illustrate:
[0004] 1. The process of the pre-training phase is as follows Figure 4 As shown in the figure, the pre-training of the retrieval model usually uses large-scale unlabeled corpus to build an encoder-decoder architecture through a method similar to RetroMAE; the retrieval model masks the input text, the encoder generates text embedding representation, and the lightweight decoder reconstructs the input text, with the goal of minimizing the reconstruction error. This step aims to enhance the semantic fidelity and robustness of the embedding representation.
[0005] 2. Fine-tuning stage: First, general fine-tuning is performed to improve the universality and discrimination of embedding. The process is as follows: Figure 5 As shown in Figure 2. The fine-tuning process uses an unlabeled dataset in a general domain, which contains a large number of semantically related text pairs, such as questions and answers, titles and texts. Contrastive learning is widely used in this stage, and its goal is to optimize the representation ability of the retrieval model by maximizing the similarity of positive sample pairs and minimizing the similarity with negative sample pairs. In order to generate high-quality negative samples, in-batch sampling is usually used, and large batch training (such as a batch size of 19,200) is used to further improve the performance of the retrieval model.
[0006] 3. Then, task-specific fine-tuning is performed on the labeled dataset of the specific task, that is, task-specific fine-tuning is performed to further optimize the performance of the retrieval model in specific tasks such as retrieval, ranking, and semantic similarity. The process is as follows: Figure 6 As shown in the figure. Instructions are introduced during training to enhance the task adaptability of the retrieval model by generating input text in a specific format for each task. Taking the retrieval task as an example, during the fine-tuning process, BGE adds the prefix "Generate representation for this sentence to retrieve related articles:" to the question to help the retrieval model better understand the current task and improve the retrieval effect of the retrieval model. At the same time, the retrieval model also mines hard negative samples (samples with high similarity scores but incorrect) from task-related corpora to further improve its ability to understand and distinguish task-specific semantics.
[0007] 4. In the evaluation phase, the trained retrieval model is comprehensively evaluated using standard benchmarks, using indicators such as NDCG and RECALL to comprehensively measure the performance of the retrieval model in order to generate highly versatile and high-quality text embeddings, while also having the ability to be further fine-tuned to adapt to specific application scenarios.
[0008] However, in practical applications, although general pre-trained retrieval models perform well in general scenarios, they cannot achieve ideal results in specialized, highly complex, and context-unique vertical fields (such as government procurement, law, finance, and medicine). Therefore, there is a need to fine-tune the retrieval model in vertical fields so that the retrieval model can better understand the domain-specific semantic features, professional terms, and expressions. The core idea of fine-tuning the retrieval model in a vertical field is to collect and annotate professional corpora related to the field, and then retrain the retrieval model using the same model as the specific task fine-tuning stage, including:
[0009] 1. Collection and cleaning of domain data: Collect a large amount of domain data (text data) from the target domain, such as legal and regulatory documents, government procurement regulations and cases, medical papers or clinical records, etc., and clean and preprocess the domain data to ensure the corpus quality and domain specificity of the domain data.
[0010] 2. Construction and expansion of annotated data sets: With the participation of domain experts, the collected domain data is manually annotated to construct annotated data sets; the annotations may include question-answer pairs (Q&A pairs), relevance judgments of text pairs, and relevance level markings for retrieval tasks. Since there are often no ready-made large-scale annotated data sets in vertical fields, it requires a large investment in annotation. This step is considered a key link in improving the performance of retrieval models in existing practices.
[0011] 3. Fine-tuning for specific tasks in vertical fields: The labeled datasets are used for the fine-tuning process for specific tasks. Similar to the labeled fine-tuning in general scenarios, in vertical fields, the retrieval model uses supervised learning, comparative training with labeled datasets, mining of difficult negative samples, or the introduction of imperative prompts, etc., to focus on improving the retrieval model's understanding and discrimination capabilities for domain-specific problems.
[0012] 4. Model evaluation based on domain validation set: After training, the retrieval model needs to be evaluated on a small-scale validation set or benchmark dataset related to the field to determine the recall rate, NDCG value and other indicators of the fine-tuned retrieval model on domain-specific tasks.
[0013] Compared with general scenarios, the outstanding feature of vertical domain fine-tuning is the high reliance on annotated datasets and domain knowledge. The final performance of the retrieval model depends largely on the quality, quantity and domain expertise of the annotated dataset. In addition, specific tasks in different fields (such as "searching for matching specific policy provisions in government procurement data" or "searching for research related to specific symptoms in medical literature") require careful design and annotation strategies to allow the retrieval model to capture the unique semantic relationships in the field.
[0014] Although it has become a consensus to use the method of specific task fine-tuning to retrain the retrieval model in vertical fields, it still faces many challenges and shortcomings in highly professional fields such as government procurement. The following is an analysis from the aspects of positive and negative sample construction, corpus utilization and data dynamics:
[0015] 1. Difficulty in constructing positive samples: In the field of government procurement, the knowledge base contains a large number of legal and regulatory provisions, policy documents, operating guidelines and case studies. These texts are usually complex in structure and require deep professional background knowledge to accurately label them. Due to the scarcity and high cost of domain expert resources, in order to obtain high-quality positive samples, a large number of experts are often required to interpret and judge. This process not only raises the threshold for data construction, but also causes a huge waste of resources, limiting the possibility of rapid iteration and updating of retrieval models.
[0016] 2. Negative sample construction is not accurate enough: In the government procurement scenario, different regulations or policies may be similar in semantics, but their applicable conditions and actual effects are significantly different. For the retrieval model, it is particularly important to identify these "superficially related but essentially irrelevant" document fragments; however, existing negative sample construction methods (such as random sampling or simple reliance on similarity sorting) are difficult to capture such subtle differences, resulting in uneven quality of negative samples; if the negative sample is too simple, it cannot effectively train the retrieval model's ability to distinguish difficult-to-distinguish samples; if it is too difficult, it may lead to unstable retrieval model training. The lack of a sophisticated negative sample selection strategy weakens the retrieval model's ability to deeply understand the semantics of vertical fields.
[0017] 3. Low efficiency in data acquisition and utilization: Although there is a large amount of unlabeled text data in the field of government procurement, how to efficiently utilize this unlabeled text data remains a difficult problem. Existing methods mainly rely on manually labeled positive and negative samples for supervised fine-tuning, but lack automated and systematic means of utilizing massive amounts of unlabeled text data. As a result, the retrieval model cannot fully absorb the background knowledge of the field and it is difficult to achieve ideal performance in the absence of large-scale labeled data. In the long run, this model that relies on manual annotation is not only costly, but also weakens the scalability and rapid adaptability of the retrieval model.
[0018] To sum up, the core problems faced by the current fine-tuning methods of procurement data retrieval models in the field of government procurement are that the positive sample labeling is costly and relies on experts, the negative sample construction strategy is inaccurate, and a large amount of unlabeled data is not effectively utilized, which in turn limits the training efficiency and retrieval performance of the procurement data retrieval model.
[0019] Therefore, how to provide a procurement data retrieval model fine-tuning method and system to improve the training efficiency and retrieval performance of the procurement data retrieval model has become a technical problem that needs to be solved urgently. Summary of the invention
[0020] The technical problem to be solved by the present invention is to provide a method and system for fine-tuning a procurement data retrieval model, so as to improve the training efficiency and retrieval performance of the procurement data retrieval model.
[0021] In a first aspect, the present invention provides a method for fine-tuning a procurement data retrieval model, comprising the following steps:
[0022] Step S1, obtaining a large number of procurement documents, segmenting each procurement document based on content category to obtain a number of text segments, and performing domain recognition on each text segment through a pre-trained large language model to construct a procurement domain dataset;
[0023] Step S2: Generate a search question for each text segment in the procurement domain data set through the large language model, and construct a question text set based on each search question and the corresponding text segment;
[0024] Step S3: constructing a positive sample candidate set and a negative sample candidate set based on the question text set and the procurement domain data set through the large language model;
[0025] Step S4: Refining the positive sample candidate set and the negative sample candidate set respectively by using the large language model to construct a positive sample set and a negative sample set;
[0026] Step S5: Create a purchase data retrieval model, and fine-tune the purchase data retrieval model using the positive sample set and the negative sample set.
[0027] Furthermore, the step S1 is specifically as follows:
[0028] Obtain a large number of procurement documents, and construct a procurement document set D = {d1, d2, ..., d n}, d n Indicates the nth purchase document;
[0029] Each of the procurement documents is segmented based on the content category to obtain a segment subset S including several text segments. i ={s i1 ,s i2 ,…,s im},S i represents the subset of fragments of the i-th purchase document, s im represents the mth text segment of the i-th procurement document;
[0030] Construct a fragment set based on each of the fragment subsets
[0031] The pre-trained large language model performs domain recognition on each text segment in the segment set through the input domain prompt words to obtain the text segments in the procurement domain and non-procurement domain. After normalization and cleaning of each text segment in the procurement domain, a procurement domain dataset is constructed:
[0032] rel im =M(s im ,P domain );
[0033] S proc ={s im ∈S|M(s im ,P domain )=1};
[0034] Among them, rel im represents the domain of the mth text segment of the ith procurement document. A value of 1 indicates the procurement domain, and a value other than 1 indicates the non-procurement domain. M() represents the large language model. P domain Indicates domain prompt words; S proc Represents the procurement domain dataset.
[0035] Furthermore, the step S2 is specifically as follows:
[0036] The large language model generates a search question for each text segment in the procurement domain dataset through the input question prompt words, and constructs a question text set based on each search question and the corresponding text segment:
[0037] q im =M(s im ,P question );
[0038] QD={(q im ,s im )|s im ∈S proc};
[0039] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; M() represents the large language model; s im represents the mth text segment of the i-th procurement document; P question represents the question prompt word; QD represents the question text set; S proc Represents the procurement domain dataset.
[0040] Furthermore, the step S3 is specifically as follows:
[0041] Construct a positive sample candidate set and a negative sample candidate set with an initial state of an empty set:
[0042]
[0043] Among them, C pos represents the positive sample candidate set; C neg represents the negative sample candidate set;
[0044] The retrieval questions and text fragments in the question text set are included in the positive sample candidate set:
[0045]
[0046] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; s im represents the mth text segment of the i-th procurement document; QD represents the question text set; ∪ represents the merge operation;
[0047] Based on cosine similarity, select N candidate text segments in the procurement domain dataset that are most similar to each search question in the question text set to construct a candidate text segment set:
[0048]
[0049] R im = {r im1,r im2 ,…,r imN};
[0050] Among them, R im Indicates that im The most similar candidate text segment set; S proc represents the procurement domain dataset; r imN Indicates that im The Nth most similar candidate text segment;
[0051] The large language model uses the input positive and negative sample judgment prompt words and adopts a one-time input strategy to perform positive and negative sample judgment on each candidate text segment in the candidate text segment set to obtain positive samples and negative samples:
[0052]
[0053] in, represents a positive sample; represents a negative sample; M() represents a large language model; P judge Indicates the positive and negative sample judgment prompt words;
[0054] Each of the positive samples Include the positive sample candidate set and add each negative sample Include negative sample candidates.
[0055] Furthermore, the step S4 is specifically as follows:
[0056] The large language model refines the positive sample candidate set and the negative sample candidate set respectively through the input positive sample judgment prompt word and the negative sample judgment prompt word to construct a positive sample set and a negative sample set:
[0057] C' pos =C pos \{(q im ,s p )}ifM(q im ,s p ,P ifpos )≠Yes;
[0058] C' neg =C neg \{(q im ,s n )}ifM(q im ,s n ,P ifneg )≠irrelevant;
[0059] Among them, C' pos represents the positive sample set; C' neg represents the negative sample set; Cpos represents the positive sample candidate set; C neg represents the negative sample candidate set; \ represents removing samples from the set; q im represents the retrieval problem of the mth text segment of the i-th procurement document; s p represents the positive sample in the positive sample candidate set; s n represents the negative samples in the negative sample candidate set; M() represents the large language model; P ifpos Indicates the positive sample judgment prompt word; P ifneg Indicates the negative sample judgment prompt word;
[0060] The step S5 is specifically as follows:
[0061] A procurement data retrieval model is created, and a loss function of the procurement data retrieval model is set. For each retrieval question in the question text set, corresponding positive samples and negative samples are searched from the positive sample set and the negative sample set respectively. The procurement data retrieval model is fine-tuned through each of the retrieval questions, positive samples and negative samples. During the fine-tuning process, the loss function is used to maximize the similarity between the retrieval question and the positive sample, and minimize the similarity between the retrieval question and the negative sample.
[0062] In a second aspect, the present invention provides a procurement data retrieval model fine-tuning system, comprising the following modules:
[0063] A procurement domain data set construction module is used to obtain a large number of procurement documents, segment each procurement document based on content category to obtain a number of text segments, and perform domain recognition on each text segment through a pre-trained large language model to construct a procurement domain data set;
[0064] A question text set construction module is used to generate a search question for each text segment in the procurement field data set through the large language model, and to construct a question text set based on each search question and the corresponding text segment;
[0065] A positive and negative sample candidate set construction module, used to construct a positive sample candidate set and a negative sample candidate set based on the question text set and the procurement domain data set through the large language model;
[0066] A positive and negative sample set construction module, used to refine the positive sample candidate set and the negative sample candidate set respectively through the large language model to construct a positive sample set and a negative sample set;
[0067] The procurement data retrieval model fine-tuning module is used to create a procurement data retrieval model and fine-tune the procurement data retrieval model through the positive sample set and the negative sample set.
[0068] Furthermore, the procurement domain dataset construction module is specifically used for:
[0069] Obtain a large number of procurement documents, and construct a procurement document set D = {d1, d2, ..., d n}, d n Indicates the nth purchase document;
[0070] Each of the procurement documents is segmented based on the content category to obtain a segment subset S including several text segments. i ={s i1 ,s i2 ,…,s im},S i represents the subset of fragments of the i-th purchase document, s im represents the mth text segment of the i-th procurement document;
[0071] Construct a fragment set based on each of the fragment subsets
[0072] The pre-trained large language model performs domain recognition on each text segment in the segment set through the input domain prompt words to obtain the text segments in the procurement domain and non-procurement domain. After normalization and cleaning of each text segment in the procurement domain, a procurement domain dataset is constructed:
[0073] rel im =M(s im ,P domain );
[0074] S proc ={s im ∈S|M(s im ,P domain )=1};
[0075] Among them, rel im represents the domain of the mth text segment of the ith procurement document. A value of 1 indicates the procurement domain, and a value other than 1 indicates the non-procurement domain. M() represents the large language model. P domain Indicates domain prompt words; S proc Represents the procurement domain dataset.
[0076] Furthermore, the question text set construction module is specifically used for:
[0077] The large language model generates a search question for each text segment in the procurement domain dataset through the input question prompt words, and constructs a question text set based on each search question and the corresponding text segment:
[0078] q im =M(sim ,P question );
[0079] QD={(q im ,s im )|s im ∈S proc};
[0080] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; M() represents the large language model; s im represents the mth text segment of the i-th procurement document; P question represents the question prompt word; QD represents the question text set; S proc Represents the procurement domain dataset.
[0081] Furthermore, the positive and negative sample candidate set construction module is specifically used for:
[0082] Construct a positive sample candidate set and a negative sample candidate set with an initial state of an empty set:
[0083]
[0084] Among them, C pos represents the positive sample candidate set; C neg represents the negative sample candidate set;
[0085] The retrieval questions and text fragments in the question text set are included in the positive sample candidate set:
[0086]
[0087] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; s im represents the mth text segment of the i-th procurement document; QD represents the question text set; ∪ represents the merge operation;
[0088] Based on cosine similarity, select N candidate text segments in the procurement domain dataset that are most similar to each search question in the question text set to construct a candidate text segment set:
[0089]
[0090] R im = {r im1 ,r im2 ,…,r imN};
[0091] Among them, R im Indicates that im The most similar candidate text segment set; Sproc represents the procurement domain dataset; r imN Indicates that im The Nth most similar candidate text segment;
[0092] The large language model uses the input positive and negative sample judgment prompt words and adopts a one-time input strategy to perform positive and negative sample judgment on each candidate text segment in the candidate text segment set to obtain positive samples and negative samples:
[0093]
[0094] in, represents a positive sample; represents a negative sample; M() represents a large language model; P judge Indicates the positive and negative sample judgment prompt words;
[0095] Each of the positive samples Include the positive sample candidate set and add each negative sample Include negative sample candidates.
[0096] Furthermore, the positive and negative sample set construction module is specifically used for:
[0097] The large language model refines the positive sample candidate set and the negative sample candidate set respectively through the input positive sample judgment prompt word and the negative sample judgment prompt word to construct a positive sample set and a negative sample set:
[0098] C' pos =C pos \{(q im ,s p )}ifM(q im ,s p ,P ifpos )≠Yes;
[0099] C' neg =C neg \{(q im ,s n )}ifM(q im ,s n ,P ifneg )≠irrelevant;
[0100] Among them, C' pos represents the positive sample set; C' neg represents the negative sample set; C pos represents the positive sample candidate set; C neg represents the negative sample candidate set; \ represents removing samples from the set; q im represents the retrieval problem of the mth text segment of the i-th procurement document; s prepresents the positive sample in the positive sample candidate set; s n represents the negative samples in the negative sample candidate set; M() represents the large language model; P ifpos Indicates the positive sample judgment prompt word; P ifneg Indicates the negative sample judgment prompt word;
[0101] The procurement data retrieval model fine-tuning module is specifically used for:
[0102] A procurement data retrieval model is created, and a loss function of the procurement data retrieval model is set. For each retrieval question in the question text set, corresponding positive samples and negative samples are searched from the positive sample set and the negative sample set respectively. The procurement data retrieval model is fine-tuned through each of the retrieval questions, positive samples and negative samples. During the fine-tuning process, the loss function is used to maximize the similarity between the retrieval question and the positive sample, and minimize the similarity between the retrieval question and the negative sample.
[0103] The advantages of the present invention are:
[0104] A large number of procurement documents are obtained and segmented to obtain several text fragments, and each text fragment is identified by a pre-trained large language model to construct a procurement domain data set; then, a retrieval question is generated for each text fragment in the procurement domain data set through the large language model, and a question text set is constructed based on each retrieval question and the corresponding text fragment; then, a positive sample candidate set and a negative sample candidate set are constructed based on the question text set and the procurement domain data set through the large language model; then, the positive sample candidate set and the negative sample candidate set are refined by the large language model to construct a positive sample set and a negative sample set; finally, a procurement data retrieval model is created, and the procurement data retrieval model is fine-tuned by the positive sample set and the negative sample set; that is, the large language model is used to automatically generate retrieval questions that are highly related to the text fragment, and then the cosine similarity and the large language model are used to determine whether the text fragment can answer the retrieval question, without the participation of experts, it can be obtained from massive unlabeled text. Automatically obtain high-quality positive sample candidate sets (positive sample pairs) from text fragments, greatly reducing labor costs and thresholds; through a one-time input strategy, input retrieval questions and candidate text fragments to the large language model, so that it can distinguish positive and negative sample candidates in the input comparison context. On this basis, the positive sample candidate sets and negative sample candidate sets are again strictly refined and reviewed one by one, so as to ensure that the quality and difficulty of negative samples are more reasonable, and effectively enhance the understanding and discrimination ability of the semantics in the procurement field; starting directly from unlabeled text fragments, the large language model automatically generates retrieval questions and related positive and negative sample pairs (positive samples-negative samples), and converts unlabeled text fragments into positive sample sets and negative sample sets that can be used for training in a zero-labeled manner. This not only reduces the dependence on manual annotation, but also enables the procurement data retrieval model to quickly adapt to regulatory updates and data changes, ensuring scalability and rapid iteration capabilities, and ultimately greatly improving the training efficiency and retrieval performance of the procurement data retrieval model. BRIEF DESCRIPTION OF THE DRAWINGS
[0105] The present invention will be further described below in conjunction with embodiments with reference to the accompanying drawings.
[0106] Figure 1 It is a flow chart of a method for fine-tuning a procurement data retrieval model of the present invention.
[0107] Figure 2 It is a structural schematic diagram of a procurement data retrieval model fine-tuning system of the present invention.
[0108] Figure 3 It is a schematic diagram comparing the present invention with the existing method.
[0109] Figure 4 It is a flowchart of the pre-training stage of the existing method.
[0110] Figure 5It is a flow chart of the fine-tuning phase of the existing method.
[0111] Figure 6 It is a flowchart of task-specific fine-tuning of existing methods. DETAILED DESCRIPTION
[0112] The technical solution in the embodiment of the present application has the following general idea: a large language model is used to automatically generate retrieval questions that are highly relevant to the text fragment, and then cosine similarity and the large language model are used to determine whether the text fragment can answer the retrieval question. Without the participation of experts, a high-quality positive sample candidate set can be automatically obtained from a large number of unlabeled text fragments, greatly reducing labor costs and thresholds; a one-time input strategy is used to input the retrieval question and the candidate text fragment to the large language model, so that it can distinguish the positive and negative sample candidates under the input comparison context. On this basis, the positive sample candidate set and the negative sample candidate set are again strictly refined and reviewed one by one, so as to ensure that the quality and difficulty of the negative samples are more reasonable, and effectively enhance the understanding and discrimination ability of the semantics in the procurement field; starting directly from the unlabeled text fragment, the retrieval question and the related positive and negative sample pairs are automatically generated by the large language model, and the unlabeled text fragment is converted into a positive sample set and a negative sample set that can be used for training in a zero-labeled manner, which not only reduces the dependence on manual annotation, but also enables the procurement data retrieval model to quickly adapt to regulatory updates and data changes, ensure scalability and rapid iteration capabilities, and thus improve the training efficiency and retrieval performance of the procurement data retrieval model.
[0113] Please refer to Figures 1 to 6 As shown, a preferred embodiment of a method for fine-tuning a procurement data retrieval model of the present invention comprises the following steps:
[0114] Step S1, obtaining a large number of procurement documents, segmenting each procurement document based on content category to obtain a number of text segments, and performing domain recognition on each text segment through a pre-trained large language model to construct a procurement domain dataset;
[0115] Step S2: Generate a search question for each text segment in the procurement domain data set through the large language model, and construct a question text set based on each search question and the corresponding text segment;
[0116] Step S3: constructing a positive sample candidate set and a negative sample candidate set based on the question text set and the procurement domain data set through the large language model;
[0117] Step S4: Refining the positive sample candidate set and the negative sample candidate set respectively by using the large language model to construct a positive sample set and a negative sample set;
[0118] Step S5: Create a purchase data retrieval model, and fine-tune the purchase data retrieval model using the positive sample set and the negative sample set.
[0119] The core idea of the present invention is to automatically generate retrieval questions and corresponding high-quality positive and negative sample pairs from original unlabeled corpus (purchased documents / text fragments) with the help of a large language model (LLM) without the need for manual annotation intervention.
[0120] The step S1 is specifically as follows:
[0121] Obtain a large number of procurement documents, and construct a procurement document set D = {d1, d2, ..., d n}, d n Indicates the nth purchase document;
[0122] Each of the procurement documents is segmented based on the content category to improve the subsequent processing efficiency, and a segment subset S including several text segments is obtained. i ={s i1 ,s i2 ,…,s im},S i represents the subset of fragments of the i-th purchase document, s im represents the mth text segment of the i-th procurement document;
[0123] Construct a fragment set based on each of the fragment subsets
[0124] The pre-trained large language model performs domain recognition on each text segment in the segment set through the input domain prompt words to obtain the text segments in the procurement domain and non-procurement domain. After normalization and cleaning of each text segment in the procurement domain, a procurement domain dataset is constructed:
[0125] rel im =M(s im ,P domain );
[0126] S proc ={s im ∈S|M(s im ,P domain )=1};
[0127] Among them, rel im represents the domain of the mth text segment of the ith procurement document. A value of 1 indicates the procurement domain, and a value other than 1 indicates the non-procurement domain. M() represents the large language model. P domain Indicates domain prompt words; S proc Represents the procurement domain dataset.
[0128] An example of the field prompt word is: Please judge whether the following text fragment belongs to the procurement field based on its content. If it does, output 1; if it does not, output 0.
[0129] The goal of this step is to automatically and effectively screen out procurement documents that are highly relevant to the procurement field from the huge and complex initial text set (procurement documents), and to clean and fine-grained slice them. Traditional methods often require manual intervention to determine field relevance, which is relatively inefficient and easily misses "procurement-irrelevant" segments in "procurement-related" procurement documents. The present invention utilizes the powerful semantic understanding capabilities of a large language model to determine the relevance of text segments to the procurement field through field prompt words, which can significantly reduce the overhead of processing irrelevant data in subsequent links, while more meticulously ensuring fine-tuning of professional field knowledge in the training set.
[0130] Corpus screening and preprocessing eliminate noise corpus irrelevant to procurement, ensuring the domain-specific nature of subsequent retrieval question generation and sample construction; each text fragment in the procurement domain data set will serve as the basic unit for subsequent automatic retrieval question generation and sample construction.
[0131] The step S2 is specifically as follows:
[0132] The large language model generates a search question for each text segment in the procurement domain dataset through the input question prompt words, and constructs a question text set based on each search question and the corresponding text segment:
[0133] q im =M(s im ,P question );
[0134] QD={(q im ,s im )|s im ∈S proc};
[0135] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; M() represents the large language model; s im represents the mth text segment of the i-th procurement document; P question represents the question prompt word; QD represents the question text set; S proc Represents the procurement domain dataset.
[0136] An example of the question prompt word is: Please generate a search question that a user may ask based on the following purchased text fragment, and the search question needs to be answered by the text fragment.
[0137] When constructing the training set of the procurement data retrieval model, "query" is needed to simulate the user's retrieval needs; this step aims to automatically generate corresponding retrieval questions for each text fragment without annotation. The present invention uses a large language model to automatically mine the retrieval questions that potential users are concerned about from the semantics of the text fragment, forming a highly relevant query corpus, so that the fine-tuning of the procurement data retrieval model is closer to the real retrieval scenario.
[0138] The semantic features of the retrieval question should be consistent with the key information, terms and context contained in the text fragment, so as to truly simulate the retrieval question that the user may ask when querying; unlike the traditional method that requires manual annotation or expert design of questions, the present invention implements this process under zero annotation conditions through a large language model, which significantly reduces labor costs and time investment.
[0139] Through the question text set, positive and negative sample mining can be performed on the retrieval questions later, and at the same time, it can be clearly known which text segment each automatically generated retrieval question comes from, ensuring that when the positive and negative samples are automatically constructed later, the text segment corresponding to the retrieval question can be easily located.
[0140] The step S3 is specifically as follows:
[0141] Construct a positive sample candidate set and a negative sample candidate set with an initial state of an empty set:
[0142]
[0143] Among them, C pos represents the positive sample candidate set; C neg represents the negative sample candidate set;
[0144] The retrieval questions and text fragments in the question text set are included in the positive sample candidate set:
[0145]
[0146] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; s im represents the mth text segment of the i-th procurement document; QD represents the question text set; ∪ represents the merge operation;
[0147] Since the retrieval problem in the question text set is generated based on text fragments of the procurement domain dataset, (q im ,s im ) is naturally a strong positive sample initial candidate pair and can be included in the positive sample candidate set without additional judgment;
[0148] Based on cosine similarity, select N candidate text segments in the procurement domain dataset that are most similar to each search question in the question text set to construct a candidate text segment set:
[0149]
[0150] R im = {r im1 ,r im2 ,…,r imN};
[0151] Among them, R im Indicates that im The most similar candidate text segment set; S proc represents the procurement domain dataset; r imN Indicates that im The Nth most similar candidate text segment;
[0152] The large language model uses the input positive and negative sample judgment prompt words and adopts a one-time input strategy to perform positive and negative sample judgment on each candidate text segment in the candidate text segment set to obtain positive samples and negative samples:
[0153]
[0154] in, Represents a positive sample, which can answer the retrieval question; represents a negative sample, which cannot answer the retrieval question; M() represents a large language model; P judge Indicates the positive and negative sample judgment prompt words; the positive and negative sample judgment prompt words are exemplified as: Please help me determine which documents can answer or partially answer the search question based on the search question and document list I gave you, and which documents are irrelevant to the search question, and return them in JSON format;
[0155] The one-time input strategy can guide the third language model to make comparisons in the input samples, thereby reducing the large model hallucination;
[0156] Each of the positive samples Include the positive sample candidate set and add each negative sample Include negative sample candidates.
[0157] That is, by using cosine similarity and the third language model, positive and negative sample candidate sets are mined for each search question without manual annotation. Traditionally, this is usually done by experts when designing search questions, or expert annotation is required to distinguish which text fragments can answer the given search questions and which text fragments are irrelevant to the search questions. The present invention uses the third language model to determine the relevance, automatically matches the search questions and unlabeled text fragments into positive and negative sample candidate pairs (positive sample candidate set, negative sample candidate set), and significantly reduces labor costs and human intervention.
[0158] The step S4 is specifically as follows:
[0159] The large language model refines the positive sample candidate set and the negative sample candidate set respectively through the input positive sample judgment prompt word and the negative sample judgment prompt word to construct a positive sample set and a negative sample set:
[0160] C' pos =C pos \{(q im ,s p )}ifM(q im ,s p ,P ifpos )≠Yes;
[0161] C' neg =C neg \{(q im ,s n )}ifM(q im ,s n ,P ifneg )≠irrelevant;
[0162] Among them, C' pos represents the positive sample set; C' neg represents the negative sample set; C pos represents the positive sample candidate set; C neg represents the negative sample candidate set; \ represents removing samples from the set; q im represents the retrieval problem of the mth text segment of the i-th procurement document; s p represents the positive sample in the positive sample candidate set; s n represents the negative samples in the negative sample candidate set; M() represents the large language model; P ifpos Indicates the positive sample judgment prompt word; P ifneg Indicates the negative sample judgment prompt word;
[0163] That is, there is a certain risk of misjudgment in the positive sample candidate set and the negative sample candidate set, and the quality of positive and negative samples may be uneven. Therefore, on the basis of assuming the preliminary category (positive / negative) of each sample, each sample pair (retrieval question-text fragment pair) is judged separately, and misjudged positive and negative sample pairs are eliminated to ensure the high quality and consistency of the final positive and negative sample sets.
[0164] An example of the positive sample judgment prompt word is: You are an expert in the procurement field. You are given a procurement question and an answer. If the answer can answer or partially answer the question, please return 'yes'; if the answer cannot answer the question or the question is not a procurement question, please return 'no'.
[0165] An example of the negative sample judgment prompt word is: You are an expert in the field of procurement. You are given a procurement question and an answer. If the answer is irrelevant to this question, please return 'irrelevant'. If the answer is relevant to this question, please return 'relevant'. If the question is not a procurement question, please return 'no'.
[0166] The step S5 is specifically as follows:
[0167] A procurement data retrieval model is created, and a loss function of the procurement data retrieval model is set. For each retrieval question in the question text set, the corresponding positive sample and negative sample are searched from the positive sample set and the negative sample set, respectively. The procurement data retrieval model is fine-tuned through each retrieval question, positive sample and negative sample. During the fine-tuning process, the loss function is used to maximize the similarity between the retrieval question and the positive sample, and minimize the similarity between the retrieval question and the negative sample. With multiple rounds of training iterations, the procurement data retrieval model gradually learns to distinguish subtle differences between difficult-to-distinguish domain documents, thereby improving retrieval accuracy and robustness.
[0168] A preferred embodiment of a procurement data retrieval model fine-tuning system of the present invention includes the following modules:
[0169] A procurement domain data set construction module is used to obtain a large number of procurement documents, segment each procurement document based on content category to obtain a number of text segments, and perform domain recognition on each text segment by pre-training the large language model to construct a procurement domain data set;
[0170] A question text set construction module is used to generate a search question for each text segment in the procurement field data set through the large language model, and to construct a question text set based on each search question and the corresponding text segment;
[0171] A positive and negative sample candidate set construction module, used to construct a positive sample candidate set and a negative sample candidate set based on the question text set and the procurement domain data set through the large language model;
[0172] A positive and negative sample set construction module, used to refine the positive sample candidate set and the negative sample candidate set respectively through the large language model to construct a positive sample set and a negative sample set;
[0173] The procurement data retrieval model fine-tuning module is used to create a procurement data retrieval model and fine-tune the procurement data retrieval model through the positive sample set and the negative sample set.
[0174] The core idea of the present invention is to automatically generate retrieval questions and corresponding high-quality positive and negative sample pairs from original unlabeled corpus (purchased documents / text fragments) with the help of a large language model (LLM) without the need for manual annotation intervention.
[0175] The procurement domain dataset construction module is specifically used for:
[0176] Obtain a large number of procurement documents, and construct a procurement document set D = {d1, d2, ..., d n}, d n Indicates the nth purchase document;
[0177] Each of the procurement documents is segmented based on the content category to improve the subsequent processing efficiency, and a segment subset S including several text segments is obtained. i ={s i1 ,s i2 ,…,s im},S i represents the subset of fragments of the i-th purchase document, s im represents the mth text segment of the i-th procurement document;
[0178] Construct a fragment set based on each of the fragment subsets
[0179] The pre-trained large language model performs domain recognition on each text segment in the segment set through the input domain prompt words to obtain the text segments in the procurement domain and non-procurement domain. After normalization and cleaning of each text segment in the procurement domain, a procurement domain dataset is constructed:
[0180] rel im =M(s im ,P domain );
[0181] S proc ={s im ∈S|M(s im ,P domain )=1};
[0182] Among them, rel imrepresents the domain of the mth text segment of the ith procurement document. A value of 1 indicates the procurement domain, and a value other than 1 indicates the non-procurement domain. M() represents the large language model. P domain Indicates domain prompt words; S proc Represents the procurement domain dataset.
[0183] An example of the field prompt word is: Please judge whether the following text fragment belongs to the procurement field based on its content. If it does, output 1; if it does not, output 0.
[0184] The goal of this step is to automatically and effectively screen out procurement documents that are highly relevant to the procurement field from the huge and complex initial text set (procurement documents), and to clean and fine-grained slice them. Traditional methods often require manual intervention to determine field relevance, which is relatively inefficient and easily misses "procurement-irrelevant" segments in "procurement-related" procurement documents. The present invention utilizes the powerful semantic understanding capabilities of a large language model to determine the relevance of text segments to the procurement field through field prompt words, which can significantly reduce the overhead of processing irrelevant data in subsequent links, while more meticulously ensuring fine-tuning of professional field knowledge in the training set.
[0185] Corpus screening and preprocessing eliminate noise corpus irrelevant to procurement, ensuring the domain-specific nature of subsequent retrieval question generation and sample construction; each text fragment in the procurement domain data set will serve as the basic unit for subsequent automatic retrieval question generation and sample construction.
[0186] The question text set building module is specifically used for:
[0187] The large language model generates a search question for each text segment in the procurement domain dataset through the input question prompt words, and constructs a question text set based on each search question and the corresponding text segment:
[0188] q im =M(s im ,P question );
[0189] QD={(q im ,s im )|s im ∈S proc};
[0190] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; M() represents the large language model; s im represents the mth text segment of the i-th procurement document; P question represents the question prompt word; QD represents the question text set; S proc Represents the procurement domain dataset.
[0191] An example of the question prompt word is: Please generate a search question that a user may ask based on the following purchased text fragment, and the search question needs to be answered by the text fragment.
[0192] When constructing the training set of the procurement data retrieval model, "query" is needed to simulate the user's retrieval needs; this step aims to automatically generate corresponding retrieval questions for each text fragment without annotation. The present invention uses a large language model to automatically mine the retrieval questions that potential users are concerned about from the semantics of the text fragment, forming a highly relevant query corpus, so that the fine-tuning of the procurement data retrieval model is closer to the real retrieval scenario.
[0193] The semantic features of the retrieval question should be consistent with the key information, terms and context contained in the text fragment, so as to truly simulate the retrieval question that the user may ask when querying; unlike the traditional method that requires manual annotation or expert design of questions, the present invention implements this process under zero annotation conditions through a large language model, which significantly reduces labor costs and time investment.
[0194] Through the question text set, positive and negative sample mining can be performed on the retrieval questions later, and at the same time, it can be clearly known which text segment each automatically generated retrieval question comes from, ensuring that when the positive and negative samples are automatically constructed later, the text segment corresponding to the retrieval question can be easily located.
[0195] The positive and negative sample candidate set construction module is specifically used for:
[0196] Construct a positive sample candidate set and a negative sample candidate set with an initial state of an empty set:
[0197]
[0198] Among them, C pos represents the positive sample candidate set; C neg represents the negative sample candidate set;
[0199] The retrieval questions and text fragments in the question text set are included in the positive sample candidate set:
[0200]
[0201] Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; s im represents the mth text segment of the i-th procurement document; QD represents the question text set; ∪ represents the merge operation;
[0202] Since the retrieval problem in the question text set is generated based on text fragments of the procurement domain dataset, (q im ,s im) is naturally a strong positive sample initial candidate pair and can be included in the positive sample candidate set without additional judgment;
[0203] Based on cosine similarity, select N candidate text segments in the procurement domain dataset that are most similar to each search question in the question text set to construct a candidate text segment set:
[0204]
[0205] R im = {r im1 ,r im2 ,…,r imN};
[0206] Among them, R im Indicates that im The most similar candidate text segment set; S proc represents the procurement domain dataset; r imN Indicates that im The Nth most similar candidate text segment;
[0207] The large language model uses the input positive and negative sample judgment prompt words and adopts a one-time input strategy to perform positive and negative sample judgment on each candidate text segment in the candidate text segment set to obtain positive samples and negative samples:
[0208]
[0209] in, Represents a positive sample, which can answer the retrieval question; represents a negative sample, which cannot answer the retrieval question; M() represents a large language model; P judge Indicates the positive and negative sample judgment prompt words; the positive and negative sample judgment prompt words are exemplified as: Please help me determine which documents can answer or partially answer the search question based on the search question and document list I gave you, and which documents are irrelevant to the search question, and return them in JSON format;
[0210] The one-time input strategy can guide the third language model to make comparisons in the input samples, thereby reducing the large model hallucination;
[0211] Each of the positive samples Include the positive sample candidate set and add each negative sample Include negative sample candidates.
[0212] That is, by using cosine similarity and the third language model, positive and negative sample candidate sets are mined for each search question without manual annotation. Traditionally, this is usually done by experts when designing search questions, or expert annotation is required to distinguish which text fragments can answer the given search questions and which text fragments are irrelevant to the search questions. The present invention uses the third language model to determine the relevance, automatically matches the search questions and unlabeled text fragments into positive and negative sample candidate pairs (positive sample candidate set, negative sample candidate set), and significantly reduces labor costs and human intervention.
[0213] The positive and negative sample set construction module is specifically used for:
[0214] The large language model refines the positive sample candidate set and the negative sample candidate set respectively through the input positive sample judgment prompt word and the negative sample judgment prompt word to construct a positive sample set and a negative sample set:
[0215] C' pos =C pos \{(q im ,s p )}ifM(q im ,s p ,P ifpos )≠Yes;
[0216] C' neg =C neg \{(q im ,s n )}ifM(q im ,s n ,P ifneg )≠irrelevant;
[0217] Among them, C' pos represents the positive sample set; C' neg represents the negative sample set; C pos represents the positive sample candidate set; C neg represents the negative sample candidate set; \ represents removing samples from the set; q im represents the retrieval problem of the mth text segment of the i-th procurement document; s p represents the positive sample in the positive sample candidate set; s n represents the negative samples in the negative sample candidate set; M() represents the large language model; P ifpos Indicates the positive sample judgment prompt word; P ifneg Indicates the negative sample judgment prompt word;
[0218] That is, there is a certain risk of misjudgment in the positive sample candidate set and the negative sample candidate set, and the quality of positive and negative samples may be uneven. Therefore, on the basis of assuming the preliminary category (positive / negative) of each sample, each sample pair (retrieval question-text fragment pair) is judged separately, and misjudged positive and negative sample pairs are eliminated to ensure the high quality and consistency of the final positive and negative sample sets.
[0219] An example of the positive sample judgment prompt word is: You are an expert in the procurement field. You are given a procurement question and an answer. If the answer can answer or partially answer the question, please return 'yes'; if the answer cannot answer the question or the question is not a procurement question, please return 'no'.
[0220] An example of the negative sample judgment prompt word is: You are an expert in the field of procurement. You are given a procurement question and an answer. If the answer is irrelevant to this question, please return 'irrelevant'. If the answer is relevant to this question, please return 'relevant'. If the question is not a procurement question, please return 'no'.
[0221] The procurement data retrieval model fine-tuning module is specifically used for:
[0222] A procurement data retrieval model is created, and a loss function of the procurement data retrieval model is set. For each retrieval question in the question text set, the corresponding positive sample and negative sample are searched from the positive sample set and the negative sample set, respectively. The procurement data retrieval model is fine-tuned through each retrieval question, positive sample and negative sample. During the fine-tuning process, the loss function is used to maximize the similarity between the retrieval question and the positive sample, and minimize the similarity between the retrieval question and the negative sample. With multiple rounds of training iterations, the procurement data retrieval model gradually learns to distinguish subtle differences between difficult-to-distinguish domain documents, thereby improving retrieval accuracy and robustness.
[0223] In summary, the advantages of the present invention are:
[0224] A large number of procurement documents are obtained and segmented to obtain several text fragments, and each text fragment is identified by a pre-trained large language model to construct a procurement domain data set; then, a retrieval question is generated for each text fragment in the procurement domain data set through the large language model, and a question text set is constructed based on each retrieval question and the corresponding text fragment; then, a positive sample candidate set and a negative sample candidate set are constructed based on the question text set and the procurement domain data set through the large language model; then, the positive sample candidate set and the negative sample candidate set are refined by the large language model to construct a positive sample set and a negative sample set; finally, a procurement data retrieval model is created, and the procurement data retrieval model is fine-tuned by the positive sample set and the negative sample set; that is, the large language model is used to automatically generate retrieval questions that are highly related to the text fragment, and then the cosine similarity and the large language model are used to determine whether the text fragment can answer the retrieval question, without the participation of experts, it can be obtained from massive unlabeled text. Automatically obtain high-quality positive sample candidate sets (positive sample pairs) from text fragments, greatly reducing labor costs and thresholds; through a one-time input strategy, input retrieval questions and candidate text fragments to the large language model, so that it can distinguish positive and negative sample candidates in the input comparison context. On this basis, the positive sample candidate sets and negative sample candidate sets are again strictly refined and reviewed one by one, so as to ensure that the quality and difficulty of negative samples are more reasonable, and effectively enhance the understanding and discrimination ability of the semantics in the procurement field; starting directly from unlabeled text fragments, the large language model automatically generates retrieval questions and related positive and negative sample pairs (positive samples-negative samples), and converts unlabeled text fragments into positive sample sets and negative sample sets that can be used for training in a zero-labeled manner. This not only reduces the dependence on manual annotation, but also enables the procurement data retrieval model to quickly adapt to regulatory updates and data changes, ensuring scalability and rapid iteration capabilities, and ultimately greatly improving the training efficiency and retrieval performance of the procurement data retrieval model.
[0225] Although the specific implementation modes of the present invention are described above, those skilled in the art should understand that the specific implementation modes described are only illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included in the scope of protection of the claims of the present invention.
Claims
1. A method for fine-tuning a procurement data retrieval model, characterized by: The steps include: Step S1, obtaining a large number of procurement documents, segmenting each procurement document based on content category to obtain a number of text segments, and performing domain recognition on each text segment through a pre-trained large language model to construct a procurement domain dataset; Step S2: Generate a search question for each text segment in the procurement domain data set through the large language model, and construct a question text set based on each search question and the corresponding text segment; Step S3: constructing a positive sample candidate set and a negative sample candidate set based on the question text set and the procurement domain data set through the large language model; Step S4: Refining the positive sample candidate set and the negative sample candidate set respectively by using the large language model to construct a positive sample set and a negative sample set; Step S5: Create a purchase data retrieval model, and fine-tune the purchase data retrieval model using the positive sample set and the negative sample set.
2. A method for fine-tuning a procurement data retrieval model as claimed in claim 1, characterized in that: The step S1 is specifically as follows: Obtain a large number of procurement documents, and construct a procurement document set D = {d1, d2, ..., d n }, d n Indicates the nth purchase document; Each of the procurement documents is segmented based on the content category to obtain a segment subset S including several text segments. i ={s i1 ,s i2 ,…,s im },S i represents the subset of fragments of the i-th purchase document, s im represents the mth text segment of the i-th procurement document; Construct a fragment set based on each of the fragment subsets The pre-trained large language model performs domain recognition on each text segment in the segment set through the input domain prompt words to obtain the text segments in the procurement domain and non-procurement domain. After normalization and cleaning of each text segment in the procurement domain, a procurement domain dataset is constructed: rel im =M(s im ,P domain ); S proc ={s im ∈S∣M(s im ,P domain )=1}; Among them, rel im represents the domain of the mth text segment of the ith procurement document. A value of 1 indicates the procurement domain, and a value other than 1 indicates the non-procurement domain. M() represents the large language model. P domain Indicates domain prompt words; S proc Represents the procurement domain dataset.
3. The method for fine-tuning a procurement data retrieval model according to claim 1, characterized in that: The step S2 is specifically as follows: The large language model generates a search question for each text segment in the procurement domain dataset through the input question prompt words, and constructs a question text set based on each search question and the corresponding text segment: q im =M(s im ,P question ); QD={(q im ,s im )|s im ∈S proc }; Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; M() represents the large language model; s im represents the mth text segment of the i-th procurement document; P question Indicates question prompt words; QD represents the question text set; S proc Represents the procurement domain dataset.
4. The method for fine-tuning a procurement data retrieval model according to claim 1, characterized in that: The step S3 is specifically as follows: Construct a positive sample candidate set and a negative sample candidate set with an initial state of an empty set: Among them, C pos represents the positive sample candidate set; C neg represents the negative sample candidate set; The retrieval questions and text fragments in the question text set are included in the positive sample candidate set: Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; s im represents the mth text segment of the i-th procurement document; QD represents the question text set; ∪ represents the merge operation; Based on cosine similarity, select N candidate text segments in the procurement domain dataset that are most similar to each search question in the question text set to construct a candidate text segment set: R im ={r im1 ,r im2 ,...,r imN}; Among them, R im Indicates that im The most similar candidate text segment set; S proc represents the procurement domain dataset; r imN Indicates that im The Nth most similar candidate text segment; The large language model uses the input positive and negative sample judgment prompt words and adopts a one-time input strategy to perform positive and negative sample judgment on each candidate text segment in the candidate text segment set to obtain positive samples and negative samples: in, represents a positive sample; represents a negative sample; M() represents a large language model; P judge Indicates the positive and negative sample judgment prompt words; Each of the positive samples Include the positive sample candidate set and add each negative sample Include negative sample candidates.
5. The method for fine-tuning a procurement data retrieval model according to claim 1, characterized in that: The step S4 is specifically as follows: The large language model refines the positive sample candidate set and the negative sample candidate set respectively through the input positive sample judgment prompt word and the negative sample judgment prompt word to construct a positive sample set and a negative sample set: C' pos = C pos {(q im , s p )} if M(q im , s p , P ifpos ) ≠ yes; C' neg =C neg \{(q im ,s n )}ifM(q im ,s n ,P ifneg )≠irrelevant; Among them, C' pos represents the positive sample set; C' neg represents the negative sample set; C pos represents the positive sample candidate set; C neg represents the negative sample candidate set; \ represents removing samples from the set; q im represents the retrieval problem of the mth text segment of the i-th procurement document; s p represents the positive sample in the positive sample candidate set; s n represents the negative samples in the negative sample candidate set; M() represents the large language model; P ifpos Indicates the positive sample judgment prompt word; P ifneg Indicates the negative sample judgment prompt word; The step S5 is specifically as follows: A procurement data retrieval model is created, and a loss function of the procurement data retrieval model is set. For each retrieval question in the question text set, corresponding positive samples and negative samples are searched from the positive sample set and the negative sample set respectively. The procurement data retrieval model is fine-tuned through each of the retrieval questions, positive samples and negative samples. During the fine-tuning process, the loss function is used to maximize the similarity between the retrieval question and the positive sample, and minimize the similarity between the retrieval question and the negative sample.
6. A procurement data retrieval model fine-tuning system, characterized by: Includes the following modules: A procurement domain data set construction module is used to obtain a large number of procurement documents, segment each procurement document based on content category to obtain a number of text segments, and perform domain recognition on each text segment through a pre-trained large language model to construct a procurement domain data set; A question text set construction module is used to generate a search question for each text segment in the procurement field data set through the large language model, and to construct a question text set based on each search question and the corresponding text segment; A positive and negative sample candidate set construction module, used to construct a positive sample candidate set and a negative sample candidate set based on the question text set and the procurement domain data set through the large language model; A positive and negative sample set construction module, used to refine the positive sample candidate set and the negative sample candidate set respectively through the large language model to construct a positive sample set and a negative sample set; The procurement data retrieval model fine-tuning module is used to create a procurement data retrieval model and fine-tune the procurement data retrieval model through the positive sample set and the negative sample set.
7. A method for fine-tuning a procurement data retrieval model as claimed in claim 1, characterized in that: The procurement domain dataset construction module is specifically used for: Obtain a large number of procurement documents, and construct a procurement document set D = {d1, d2, ..., d n }, d n Indicates the nth purchase document; Each of the procurement documents is segmented based on the content category to obtain a segment subset S including several text segments. i ={s i1 ,s i2 ,…,s im },S i represents the subset of fragments of the i-th purchase document, s im represents the mth text segment of the i-th procurement document; Construct a fragment set based on each of the fragment subsets The pre-trained large language model performs domain recognition on each text segment in the segment set through the input domain prompt words to obtain the text segments in the procurement domain and non-procurement domain. After normalization and cleaning of each text segment in the procurement domain, a procurement domain dataset is constructed: rel im =M(s im ,P domain ); S proc ={s im ∈S∣M(s im ,P domain )=1}; Among them, rel im represents the domain of the mth text segment of the ith procurement document. A value of 1 indicates the procurement domain, and a value other than 1 indicates the non-procurement domain. M() represents the large language model. P domain Indicates domain prompt words; S proc Represents the procurement domain dataset.
8. A procurement data retrieval model fine-tuning system as claimed in claim 6, characterized in that: The question text set building module is specifically used for: The large language model generates a search question for each text segment in the procurement domain dataset through the input question prompt words, and constructs a question text set based on each search question and the corresponding text segment: q im =M(s im ,P question ); QD={(q im ,s im )∣s im ∈S proc }; Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; M() represents the large language model; s im represents the mth text segment of the i-th procurement document; P question represents the question prompt word; QD represents the question text set; S proc Represents the procurement domain dataset.
9. A procurement data retrieval model fine-tuning system as claimed in claim 6, characterized in that: The positive and negative sample candidate set construction module is specifically used for: Construct a positive sample candidate set and a negative sample candidate set with an initial state of an empty set: Among them, C pos represents the positive sample candidate set; C neg represents the negative sample candidate set; The retrieval questions and text fragments in the question text set are included in the positive sample candidate set: Among them, q im represents the retrieval problem of the mth text segment of the i-th procurement document; s im represents the mth text segment of the i-th procurement document; QD represents the question text set; ∪ represents the merge operation; Based on cosine similarity, select N candidate text segments in the procurement domain dataset that are most similar to each search question in the question text set to construct a candidate text segment set: R im ={r im1 ,r im2 ,…,r imN}; Among them, R im Indicates that im The most similar candidate text segment set; S proc represents the procurement domain dataset; r imN Indicates that im The Nth most similar candidate text segment; The large language model uses the input positive and negative sample judgment prompt words and adopts a one-time input strategy to perform positive and negative sample judgment on each candidate text segment in the candidate text segment set to obtain positive samples and negative samples: in, represents a positive sample; represents a negative sample; M() represents a large language model; P judge Indicates the positive and negative sample judgment prompt words; Each of the positive samples Include the positive sample candidate set and add each negative sample Include negative sample candidates.
10. A procurement data retrieval model fine-tuning system as claimed in claim 6, characterized in that: The positive and negative sample set construction module is specifically used for: The large language model refines the positive sample candidate set and the negative sample candidate set respectively through the input positive sample judgment prompt word and the negative sample judgment prompt word to construct a positive sample set and a negative sample set: C' pos = C pos {(q im , s p )} if M(q im , s p , P ifpos ) ≠ yes; C' neg =C neg \{(q im ,s n )}ifM(q im ,s n ,P ifneg )≠irrelevant; Among them, C' pos represents the positive sample set; C' neg represents the negative sample set; C pos represents the positive sample candidate set; C neg represents the negative sample candidate set; \ represents removing samples from the set; q im represents the retrieval problem of the mth text segment of the i-th procurement document; s p represents the positive sample in the positive sample candidate set; s n represents the negative samples in the negative sample candidate set; M() represents the large language model; P ifpos Indicates the positive sample judgment prompt word; P ifneg Indicates the negative sample judgment prompt word; The procurement data retrieval model fine-tuning module is specifically used for: A procurement data retrieval model is created, and a loss function of the procurement data retrieval model is set. For each retrieval question in the question text set, corresponding positive samples and negative samples are searched from the positive sample set and the negative sample set respectively. The procurement data retrieval model is fine-tuned through each of the retrieval questions, positive samples and negative samples. During the fine-tuning process, the loss function is used to maximize the similarity between the retrieval question and the positive sample, and minimize the similarity between the retrieval question and the negative sample.