Bid invitation document qualification examination processing method and system
By constructing a qualification performance vocabulary and query template, combined with the fine-tuning of the LLM model, the problems of low efficiency and poor accuracy in the qualification review of bidding documents are solved, and efficient and accurate qualification compliance review is achieved.
Patent Information
- Application Number
- CN202510442260.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the qualification review of bidding documents for electronic bidding mainly relies on manual processing, is inefficient and difficult to accurately deal with complex content. The LLM model lacks special high-quality data sets, is difficult to effectively train and reason, and has hallucination problems.
Construct a qualification performance vocabulary, design a query template to fit user input, and filter out the results with the highest similarity through cosine similarity comparison and vector database matching, and generate answers based on the LLM model, and create fine-tuning data sets to import the base model for fine-tuning to improve accuracy.
It significantly improves the accuracy of legal compliance review of bidding documents qualifications, alleviates the illusion problem of the LLM model, and improves the review efficiency and accuracy.
Smart Images

Figure CN120337890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for reviewing and processing the qualifications of tender documents. Background Art
[0002] Electronic tendering and bidding is a way to manage, publish, and process procurement and tendering processes using the Internet and digital technologies. Taking tender procurement as an example, tens of thousands of unstructured documents such as PDFs are accumulated every year in tender work. At present, tender review mainly relies on manual processing, and manual review faces challenges such as low efficiency and difficulty in accurately dealing with complex content.
[0003] In the prior art, large language models (LLM models) have provided many ideas for intelligent qualification review of tender documents with their powerful language understanding capabilities. However, the combination of RAG of large language models and tender review still faces formidable challenges. Currently, there is no research method specifically for tender documents, lacking a dedicated high-quality dataset, which is difficult to support the effective training and reasoning of LLM models; at the same time, LLM models also have the problem of hallucinations.
[0004] Therefore, there is an urgent need for an efficient intelligent review scheme to assist manual work in improving review efficiency. Summary of the Invention
[0005] In view of the defects existing in the prior art, the present invention provides a method and system for reviewing and processing the qualifications of tender documents.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: On the one hand, the present invention provides a method for reviewing and processing the qualifications of tender documents, including: Crawling the original tender legal document data and constructing a qualification and performance vocabulary table; Designing a query template to fit the user input; Vectorizing based on the original tender legal document data and the query template to construct a sample vector dataset; Filtering the user input based on the qualification and performance vocabulary table to obtain the keywords of the user input; Performing similarity comparison by calculating the cosine similarity between the keywords of the user input and the query templates in the sample vector dataset, taking the query template corresponding to the highest similarity, and combining it with the keywords of the user input to obtain a recombined query statement; The recombined query statement is initially matched in the vector database. Through vector similarity calculation, the 10 results with the highest similarity to the user input are selected; further screening is performed based on the presence or absence of the text additional field and the repetition rate of the content of the text additional field in the selected results to obtain the qualification and performance regulations corresponding to the user input; Combine the qualification and performance regulations corresponding to the user input with the keywords in the user input, and then input them into the LLM model to generate an answer.
[0007] Further, it also includes: Create a fine-tuning dataset, including instructions, inputs, and outputs; the instructions represent task instructions or questions, which are used to clarify the specific tasks that the LLM model needs to execute; the inputs are supplementary input information required for the tasks, which are used to provide context or additional data; the outputs are the target answers or outputs generated by the LLM model based on the instructions and inputs, and are used as the reference standard for training. Select an open-source large model as the base model, import the fine-tuning dataset for fine-tuning, and judge the importance of fine-tuning training according to the sample weights of the fine-tuning dataset.
[0008] Further, the cosine similarity is calculated according to the following formula:
[0009] Where is the i th generated vector and the j th reference vector the cosine similarity between them; is the function for calculating the cosine similarity between the generated vector and the reference vector between.
[0010] Further, the vector similarity is calculated according to the following formula:
[0011] Where represents the function for calculating the similarity between vectors A and B, and the result is a scalar value.
[0012] Further, the qualification and performance vocabulary is constructed according to the following steps: Based on the original tender legal document data, combined with the jieba word segmentation tool for coarse-grained and fine-grained word segmentation, extract the keyword set; Input the proper nouns of tender qualifications and performance in the keyword set to form the qualification and performance vocabulary.
[0013] Further, the query templates include single-target single-mapping matching, double-target single-mapping matching, double-target multi-mapping matching, multi-target multi-mapping matching, source-free error correction matching, single-source error correction matching, double-source single-mapping error correction matching, and double-source multi-mapping error correction matching.
[0014] Further, further screening based on the presence or absence of the text additional field in the selected results and the repetition rate of the content of the text additional field includes: Determine whether there are duplicate text additional fields among the top 10 results with the highest similarity. If not, directly output the result with the highest vector similarity as the qualification performance regulations corresponding to the user input; if so, determine whether the number of text additional fields in the judgment result is greater than the number of text additional fields required by the query template. If it is less, use all the text additional fields in the result as the qualification performance regulations corresponding to the user input; if it is greater, determine whether there are duplicate text additional fields. If there are, use the text additional field with the highest repetition rate as the qualification performance regulations corresponding to the user input; if not, use all the text additional fields in the result as the qualification performance regulations corresponding to the user input.
[0015] Further, the sample weights of the fine-tuning dataset are calculated according to the following formula:
[0016] where is the weight finally assigned to the sample ; is the total number of samples in the fine-tuning dataset; is the number of samples in the category ; is the smoothing parameter; is the balance parameter, and its value range is [0, 1]; is the smoothing exponent, and its value range is [0.5, 1]; is the importance score of the sample .
[0017] On the other hand, the present invention also provides a qualification review processing system for tender documents, including: The first module is used to crawl the original tender legal document data and construct a qualification performance vocabulary table; The second module is used to design a query template to fit the user input; The third module is used to vectorize based on the original tender legal document data and the query template to construct a sample vector dataset; The fourth module is used to filter the user input based on the qualification performance vocabulary table to obtain the keywords of the user input; The fifth module is used to perform similarity comparison by calculating the cosine similarity between the keywords of the user input and the query template in the sample vector dataset, take the query template corresponding to the highest similarity, and combine it with the keywords of the user input to obtain a recombined query statement. The sixth module is used to perform a preliminary match of the reorganized query statement in the vector database. Through vector similarity calculation, the 10 results with the highest similarity to the user input are screened out; further screening is performed based on the presence or absence of the text additional field in the screened results and the repetition rate of the content of the text additional field to obtain the qualification and performance regulations corresponding to the user input. The seventh module is used to combine the qualification and performance regulations corresponding to the user input with the keywords input by the user and then input them into the LLM model to generate an answer.
[0018] Furthermore, the system will save the entire process of generating an answer based on the user input.
[0019] Compared with the prior art, the beneficial technical effects of the present invention are as follows: The method and system for processing qualification review of tender documents provided by the present invention can effectively improve the filtering and keyword recognition efficiency of user question input by constructing a qualification and performance vocabulary table containing tender qualification professional terms and common vocabulary. Eight query templates (including 4 query query templates and 4 query error correction templates) are designed to fit the user input, so as to effectively meet the multi-type and special requirements of tender document qualification requirements.
[0020] The present invention obtains the user intention through a perception algorithm and matches it with a pre-designed intention template, combines word vector retrieval matching and vector data retrieval technology to generate an answer that not only meets the requirements of qualification and performance regulations but also can accurately respond to the user's original question. Moreover, the present invention makes a fine-tuning data set and imports it into the LLM model for fine-tuning, so that the fine-tuned LLM model can answer the user's questions more accurately and alleviate the problem of hallucinations generated by the LLM model.
[0021] The method and system for processing qualification review of tender documents provided by the present invention can significantly improve the accuracy of legal and compliant review of tender document qualifications. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a method for processing qualification review of tender documents provided in an embodiment; Figure 2 It is a schematic diagram of the process of reorganizing a query statement provided in an embodiment; Figure 3 Schematic diagram of double matching of vector database provided for an embodiment Figure 4 Schematic diagram of the process of further screening according to the text additional field in the result provided for an embodiment Figure 5 Schematic diagram of fine-tuning dataset provided for an embodiment Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] Referring to Figure 1 , an embodiment provides a method for qualification review and processing of bidding documents, including: Crawling the original bidding legal document data and constructing a qualification and performance vocabulary table Designing a query template to fit the user input Based on the original bidding legal document data and the query template vectorization, constructing a sample vector dataset Filtering the user input based on the qualification and performance vocabulary table to obtain the keywords of the user input Performing similarity comparison by calculating the cosine similarity between the keywords of the user input and the query templates in the sample vector dataset, taking the query template corresponding to the highest similarity, and combining with the keywords of the user input to obtain a recombined query statement The recombined query statement is initially matched in the vector database. Through vector similarity calculation, the 10 results with the highest similarity to the user input are screened out; further screening is performed according to the presence or absence of the text additional field in the screened results and the repetition rate of the content of the text additional field to obtain the qualification and performance regulations corresponding to the user input Combining the qualification and performance regulations corresponding to the user input with the keywords of the user input and inputting them into the LLM model to generate an answer
[0026] The vector database includes vector databases such as Milvus, Faiss, Pinecone, and Qdrant, which are used to efficiently store and retrieve large-scale vector data
[0027] The qualification and performance vocabulary table is constructed according to the following steps: Based on the original tender legal document data, combined with the jieba word segmentation tool for coarse-grained and fine-grained word segmentation, a keyword set is extracted; Input the proper nouns of tender qualification performance into the keyword set to form a qualification performance vocabulary.
[0028] By constructing a qualification performance vocabulary that includes tender qualification professional terms and common vocabulary, the filtering of user question input and the keyword recognition efficiency are improved. At the same time, it provides empowerment for the subsequent dual retrieval and the implementation of the user intention perception algorithm.
[0029] In one embodiment, the Milvus vector database is selected in this embodiment. The query templates include single-target single-mapping matching, double-target single-mapping matching, double-target multi-mapping matching, multi-target multi-mapping matching, passive error correction matching, single-source error correction matching, double-source single-mapping error correction matching, and double-source multi-mapping error correction matching; as shown in Table 1, taking the single-target single-mapping template as an example, the template sentence pattern is in the form of "B of A". The meaning of this query template is that "A" and "B" represent the elements of a certain row and a certain column in the table (at least one of the row and the column is a table header). From the perspective of the data storage method in the Milvus vector database, using the elements of the row and the column as "keys" and putting them into the vector database for matching can accurately obtain the target content "value". Therefore, as long as the row elements and column elements corresponding to the user input are accurately identified, the required purpose can be achieved. By designing 8 query templates to fit the user input, it can effectively cope with the multi-type and special requirements of the tender document qualification requirements.
[0030] Table 1 Example table of query templates
[0031] Refer to Figure 2 In one embodiment, based on the original tender legal document data, keyword sets of two excel tables of qualification performance - engineering and services, and qualification performance - materials are integrated, and the proper nouns of tender qualification performance are input into the keyword set to form a qualification performance vocabulary.
[0032] Construct a sample vector data set according to the keyword sets of the two excel tables of qualification performance - engineering and services, and qualification performance - materials in combination with the query templates; import various query templates of the sample vector data set into the paraphrase-MiniLM-L6-v2 model for vectorization; Filter the user input based on the qualification and performance glossary to obtain the keywords of the user input. Specifically, based on the qualification and performance glossary, segment the user input, and use a sliding window and diffilib to select the 5 keywords with the highest possible relevance at each position of the query template. Then, compare the cosine similarity between these five keywords and the query template in the sample vector dataset, select the query template corresponding to the highest similarity, and combine it with the keywords of the user input to obtain a recombined query statement through the word2vec model trained by skip-gram.
[0033] The cosine similarity is calculated according to the following formula:
[0034] Where, is the i th generated vector and the j th reference vector The cosine similarity between them; is the function to calculate the cosine similarity between the generated vector and the reference vector The cosine similarity is a measure of the similarity in direction between two vectors, and its value ranges from -1 to 1. Among them, 1 means the vectors are in exactly the same direction, -1 means exactly the opposite direction, and 0 means the vectors are orthogonal (i.e., independent of each other).
[0035] is the i th generated vector in the set of generated vectors, usually generated by a generation model or algorithm; is the j th reference vector in the set of reference vectors, usually the standard or target vector compared with .
[0036] Refer to Figure 3 , and the recombined query statement is initially matched in the Milvus vector database. Through vector similarity calculation, the 10 results with the highest similarity to the user input are screened out, that is, the Top10 match; on the basis of the initial match, further screening is carried out according to the presence or absence of the text additional field in the screened results and the repetition rate of the content of the text additional field to obtain the qualification and performance regulations corresponding to the user input.
[0037] The vector similarity is calculated according to the following formula:
[0038] Where, A function that calculates the similarity between vectors A and B, and the result is a scalar value. If (the preset value of the matching similarity limit for a specific template), then select this query template; if no query template can be matched, select the default template, that is, the statement after preliminary filtering is put into the Milvus vector database without changing the query template.
[0039] Refer to Table 2 and Table 3, which are the number of sample regulations of the query query template and the number of sample regulations of the query error correction template.
[0040] Table 2 Example table of the number of sample regulations of the query query template
[0041] Table 3 Example table of the number of sample regulations of the query error correction template
[0042] Refer to Figure 4 , and further filter according to the presence or absence of the text additional field in the selected results and the repetition rate of the content of the text additional field, including: Judge whether there are duplicate text additional fields among the 10 results with the highest similarity. If not, directly output the result with the highest vector similarity as the qualification and performance regulations corresponding to the user input; if so, judge whether the number of text additional fields in the judgment result is greater than the number of text additional fields required by the query template; If it is less, use all the text additional fields in the result as the qualification and performance regulations corresponding to the user input; if it is greater, judge whether there are duplicate text additional fields; If there are, use the text additional field with the highest repetition rate as the qualification and performance regulations corresponding to the user input; if not, use all the text additional fields in the result as the qualification and performance regulations corresponding to the user input.
[0043] Through the above settings, the accuracy and relevance of the qualification and performance regulations corresponding to the user input can be guaranteed.
[0044] Combine the qualification and performance regulations corresponding to the user input with the keywords input by the user through the prompt, and input the combined prompt into the LLM model to generate an answer. This combined prompt is input into the LLM model, and the LLM model uses its deep learning and context understanding capabilities to generate an answer that not only meets the requirements of the qualification and performance regulations but also can accurately answer the user input (i.e., the user's original question).
[0045] After the user's input has undergone the corresponding tender qualification performance retrieval, although the corresponding qualification performance provisions based on the user's input can be obtained, considering that in some cases when the user reviews the tender documents, a whole paragraph of the original tender document is copied and pasted into the system for questioning. At this time, if only the corresponding tender qualification regulations are available, the user's input cannot be reviewed. For example, when the user asks whether the qualification performance of advertising and publicity services is compliant, the original tender document usually states the following requirements for bidders first, and then enumerates the specific items of the qualification performance requirements corresponding to the specific provisions. At this time, relying solely on the retrieved provisions will not be able to assist the large model in making a good judgment. Therefore, the project also needs to fine-tune the large model to ensure that the large model can also make relatively accurate judgments under the specific circumstances described in some qualification performance regulations.
[0046] In one embodiment, the method further includes: Making a fine-tuning dataset, including instruction, input, and output; the instruction represents a task instruction or question, used to clarify the specific task that the LLM model needs to execute; the input is the supplementary input information required for the task, used to provide context or additional data, and is empty if the task does not depend on context; the output is the target answer or output generated by the LLM model according to the instruction and input, serving as the reference standard for training. In this embodiment, Qwen-7B-Struct is selected as the base model, the fine-tuning dataset is imported for fine-tuning, and the importance of fine-tuning training is judged according to the sample weights of the fine-tuning dataset.
[0047] The fine-tuning dataset can be made according to the user's usage scenario, so as to ensure the generality of the trained model. Refer to Figure 5 , as an example of the fine-tuning dataset. As shown in the figure, the fine-tuning datasets are tender qualification review cases (made according to the previous review cases), specific requirements for qualification performance (made according to the Provincial Company Material Procurement Strategy Guidance Manual for Material Tendering), and specific examples of qualification performance (made according to specific certificates, qualification condition certificates, etc.).
[0048] When training using the fine-tuning dataset, category weights are directly set in the loss function to solve the problem of different importance levels of fine-tuning training. The sample weights of the fine-tuning dataset are calculated according to the following formula:
[0049] Where is the weight finally assigned to the sample ; is the total number of samples in the fine-tuning dataset; is the category The number of samples; is the smoothing parameter; is the balance parameter, used to control the proportion of the class imbalance weight and the importance weight, and the value range is [0, 1]. Among them, approaching 1 means paying more attention to class imbalance, approaching 0 means paying more attention to sample importance. In this embodiment, ; is the smoothing exponent, used to control the sensitivity of class weight scaling, and the value range is [0.5, 1]. Among them, is no scaling, that is, completely according to the class ratio, is the weight reduced by the square root to reduce extreme weight differences. In this embodiment, since the number of relatively important samples is small, so ; is the sample importance score.
[0050] In one embodiment, a tender document qualification review processing system is provided, including: The first module is used to crawl the original tender legal document data and construct a qualification performance vocabulary table; The second module is used to design a query template to fit the user input; The third module is used to vectorize based on the original tender legal document data and the query template, and construct a sample vector dataset; The fourth module is used to filter the user input based on the qualification performance vocabulary table to obtain the keywords of the user input; The fifth module is used to perform similarity comparison by calculating the cosine similarity between the keywords of the user input and the query template in the sample vector dataset, take the query template corresponding to the highest similarity, and combine it with the keywords of the user input to obtain a reorganized query statement; The sixth module is used to perform a preliminary match of the reorganized query statement in the vector database, and through vector similarity calculation, screen out the 10 results with the highest similarity to the user input; further screen according to the presence or absence of the text additional field in the screened results and the repetition rate of the text additional field content to obtain the qualification performance regulations corresponding to the user input; The seventh module is used to combine the qualification performance regulations corresponding to the user input with the keywords of the user input and input them into the LLM model to generate an answer.
[0051] Moreover, the system will save the entire process of generating answers based on user input, including the user's original question, filtering operations, retrieved qualification and performance regulations, combined prompts, and answers generated by the LLM model. These historical records not only ensure the persistence of the Q&A process but also provide data resources for the system, which can be used for future analysis and optimization of the system. By analyzing these historical records, the system can continuously improve its prompt design, filtering algorithm, and answer generation quality, thereby providing more personalized and efficient services in future interactions. Through this cycle of iteration, the system can be continuously improved and adapted to the changing user needs.
[0052] To verify the similarity between the query statement and the filtered user input in the present invention, a query generation phase test was conducted. The BERTScore evaluation metric was used to quantitatively test the query generation phase; compared with traditional text evaluation metrics (such as BLEU, ROUGE), BERTScore can capture deeper semantic relationships by introducing word embedding vectors generated by pre-trained language models.
[0053] The calculation of the BERTScore evaluation metric is as follows: Given a reference text R and a generated text C, which are represented by word embedding vectors {r1, r2, …, rm} and {c1, c2, …, cn} respectively.
[0054] Calculate the cosine similarity between all reference word embeddings and candidate word embeddings:
[0055] where, ; For each generated word cj, find the most similar reference word:
[0056] The total precision is:
[0057] For each reference word ri, find the most similar generated word:
[0058] The total recall rate R is:
[0059] F1 is an index that comprehensively evaluates precision and recall rate and reflects the overall performance of the model. The F1 is calculated according to the following formula:
[0060] For this test, 100 pieces of user input data were prepared for each query template, and the BERTScore evaluation metric was used to evaluate the similarity between the user input and the retrieval statement. The results are shown in Tables 4 and 5.
[0061] Table 4 Evaluation Performance of Query Query Templates
[0062] Table 5 Evaluation Performance of Query Error Correction Templates
[0063] As can be seen from the table, taking the single-target single-mapping matching of the query query template as an example: The precision is 0.8339, indicating that 83.39% of the words (or semantic fragments) in the generated text can find highly similar matches in the reference text. A relatively high precision indicates a relatively high content quality of the generated text, but it may miss some information in the reference text.
[0064] The recall is 0.9155, indicating that 91.55% of the words in the reference text have relatively high matches in the generated text. A relatively high recall indicates that the generated text covers most of the information in the reference text.
[0065] F1 is the harmonic mean of precision and recall, used to balance the weights between the two; F1 is 0.8728, indicating that overall, the semantic similarity between the generated text and the reference text reaches 87.28%.
[0066] Matters not covered by this invention are well-known technologies.
[0067] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0068] The above-described embodiments merely represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
[0069] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Method for examining and processing qualifications in tender documents, characterized in that, Including: Crawl the data of original tender legal documents and construct a qualification and performance vocabulary table; Design a query template to fit the user input; Vectorize based on the original tender legal document data and the query template to construct a sample vector dataset; Filter the user input based on the qualification and performance vocabulary table to obtain the keywords of the user input; Perform similarity comparison by calculating the cosine similarity between the keywords of the user input and the query templates in the sample vector dataset, select the query template corresponding to the highest similarity, and combine it with the keywords of the user input to obtain a recombined query statement; The recombined query statement is preliminarily matched in the vector database. Through vector similarity calculation, the 10 results with the highest similarity to the user input are selected; further screening is performed based on the presence or absence of the text additional field and the repetition rate of the content of the text additional field in the selected results to obtain the qualification and performance regulations corresponding to the user input; Combine the qualification and performance regulations corresponding to the user input with the keywords of the user input and input them into the LLM model to generate an answer.
2. The method for reviewing and processing the qualifications of tender documents as described in claim 1, wherein It also includes: Make a fine-tuning dataset, including instructions, inputs, and outputs; the instructions represent task instructions or questions, which are used to clarify the specific tasks that the LLM model needs to perform; the inputs are supplementary input information required for the task, which are used to provide context or additional data; the outputs are the target answers or outputs generated by the LLM model according to the instructions and inputs, and are used as the reference standard for training; Select an open-source large model as the base model, import the fine-tuning dataset for fine-tuning, and judge the importance of fine-tuning training according to the sample weights of the fine-tuning dataset.
3. The method for reviewing and processing the qualifications of tender documents as described in claim 1, characterized in that, The cosine similarity is calculated according to the following formula: Among them, is the i th generation vector and the j th reference vector the cosine similarity between them; is a function for calculating the cosine similarity between the generation vector and the reference vector between them.
4. The method for reviewing and processing the qualifications in the tender document as described in claim 1, characterized in that, The vector similarity is calculated according to the following formula: Among them, represents a function for calculating the similarity between vectors A and B, and the result is a scalar value.
5. The method for reviewing and processing the qualifications of tender documents as described in claim 1, characterized in that, The qualification and performance vocabulary table is constructed according to the following steps: Based on the original tender legal document data, combined with the jieba word segmentation tool for coarse-grained word segmentation and fine-grained word segmentation, extract a keyword set; Input the proper nouns of tender qualifications and performance into the keyword set to form a qualification and performance vocabulary table.
6. The method for examining and approving the qualifications of tender documents as described in claim 1 is characterized in that, The query template includes single-target single-mapping matching, double-target single-mapping matching, double-target multi-mapping matching, multi-target multi-mapping matching, passive error correction matching, single-source error correction matching, double-source single-mapping error correction matching, and double-source multi-mapping error correction matching.
7. The method for reviewing and processing the qualifications of tender documents as described in claim 1 is characterized in that, Further screening based on the presence or absence of the text additional field and the repetition rate of the content of the text additional field in the selected results includes: Judge whether there are duplicate text additional fields among the 10 results with the highest similarity. If not, directly output the result with the highest vector similarity as the qualification and performance regulations corresponding to the user input; if so, judge whether the number of text additional fields in the judgment result is greater than the number of text additional fields required by the query template; If it is less, all text additional fields in the result are used as the qualification and performance regulations corresponding to the user input; if it is greater, judge whether there are duplicate text additional fields; If it exists, the text attachment field with the highest repetition rate is used as the qualification and performance regulations corresponding to the user input; if it does not exist, all text attachment fields in the result are used as the qualification and performance regulations corresponding to the user input.
8. The method for reviewing and processing the qualifications of tender documents as described in claim 2, characterized in that, The sample weights of the fine-tuning dataset are calculated according to the following formula: Among them, is the weight finally assigned to the sample ; is the total number of samples in the fine-tuning dataset; is the number of samples in the category ; is the smoothing parameter; is the balance parameter, and its value range is [0, 1]; is the smoothing exponent, and its value range is [0.5, 1]; is the importance score of the sample .
9. The qualification review processing system for tender documents is characterized in that, including: The first module is used to crawl the original tender legal document data and construct a qualification and performance vocabulary. The second module is used to design a query template to fit the user input. The third module is used to vectorize based on the original tender legal document data and the query template to construct an example vector dataset. The fourth module is used to filter the user input based on the qualification and performance vocabulary to obtain the keywords of the user input. The fifth module is used to perform a similarity comparison by calculating the cosine similarity between the keywords of the user input and the query templates in the example vector dataset, select the query template corresponding to the highest similarity, and combine it with the keywords of the user input to obtain a reorganized query statement. The sixth module is used to perform a preliminary match of the reorganized query statement in the vector database, calculate the vector similarity, and screen out the 10 results with the highest similarity to the user input; further screen according to the existence of the text attachment field in the screened results and the repetition rate of the text attachment field content to obtain the qualification and performance regulations corresponding to the user input. The seventh module is used to combine the qualification and performance regulations corresponding to the user input with the keywords of the user input and input them into the LLM model to generate an answer.
10. The tender document qualification review and processing system according to claim 9, characterized in that, The system will save the entire process of generating an answer according to the user input.
Citation Information
Cited By
Method and system for bid invitation document content extraction and qualification matching fused with RAG
CN121279270A
Rag-fused method and system for extracting contents of a tender document and matching qualifications
CN121279270B