Text review method, system, device and storage medium
By constructing a text library and automatically reviewing bid documents using similarity comparison and preset review dimensions, the problems of subjective deviation and inefficiency in manual review are solved, and more accurate and efficient review results are achieved.
Patent Information
- Application Number
- CN202410990742.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-07-23
AI Technical Summary
When manually reviewing bid documents, there are subjective deviations in the evaluation results and are not very accurate, and the evaluation efficiency is low.
By constructing a text library, the similarity comparison between the bidding documents and the benchmark text is compared, the target benchmark text is generated, and the similarity comparison between the bidding documents and the target benchmark text is compared, and the review operation is performed according to the preset text review dimension to generate the review results.
It reduces the subjective deviation of manual review, improves the accuracy and efficiency of review results, and ensures the comprehensiveness and accuracy of review results.
Smart Images

Figure CN118820401B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a text review method, system, device, and storage medium. Background Art
[0002] In the bidding process for engineering projects, the bidding documents serve as the basis for project implementation, describing the project overview and bidding requirements. The bid document, a response to the bidding documents, describes the bidder's basic information and the project implementation plan. Typically, after receiving the bid, the bid document is reviewed against the bidding documents, and the bidder's suitability for the project is determined based on the review results. For example, based on the bidding requirements in the bidding documents, the bidder's basic information in the bid document is reviewed to determine whether it meets the bidding requirements. This, in turn, determines the bidder's suitability for the project.
[0003] Currently, bid documents are usually manually reviewed by reviewers based on the bidding documents. This manual review method relies on the reviewer's professional level. Even if different reviewers review the same bid document, the review results may be inconsistent, which affects the accuracy of the review results.
[0004] Therefore, there is an urgent need for a method that can improve the accuracy of review results. Summary of the Invention
[0005] In view of this, the embodiments of the present disclosure provide a text review method, a text review system, an electronic device, and a computer-readable storage medium, which can reduce the subjective bias of manual review and improve the accuracy of review results through automatic review.
[0006] In one aspect, the present disclosure provides a text review method, the method comprising:
[0007] Performing a similarity comparison between the tender text in the tender document and a reference text in a text library used to represent the tender requirements, to obtain a target reference text that matches the tender text;
[0008] Comparing the bid text in the bid document with the target reference text for similarity to obtain a target bid text that matches the target reference text;
[0009] According to the preset text review dimensions, an evaluation operation is performed on the target bidding text to obtain an evaluation result for characterizing the text quality of the bidding document.
[0010] In the technical solutions of some embodiments of the present disclosure, by performing a similarity comparison between the tender text and the benchmark text in the text library, a target benchmark text corresponding to the tender document can be obtained. After performing a similarity comparison between the bid text and the target benchmark text, a target bid text that matches the target benchmark text in the bid document can be obtained. After performing a review operation on the target bid text according to the specified text review dimension, an review result that characterizes the quality of the bid document text can be generated. In this way, automatic review of the bid document is achieved, which can reduce the subjective bias of manual review and improve the accuracy of the review results.
[0011] In some embodiments, the similarity comparison of the tender text in the tender document with a reference text in a text library used to represent the tender requirements to obtain a target reference text that matches the tender text includes:
[0012] Dividing the bidding text into multiple bidding sub-texts;
[0013] For any reference text in the text library, performing a similarity comparison between the reference text and each of the tender sub-texts to select tender sub-texts whose similarity to the reference text is higher than a threshold;
[0014] Generate summary texts for each of the filtered tender sub-texts;
[0015] The summary text is compared with the reference text in the text library for similarity, and the reference text that matches the summary text is used as the target reference text.
[0016] By generating a summary text for the tender subtext and performing a similarity comparison between the summary text and the benchmark text, the text differences between the tender subtext and the benchmark text can be reduced and the accuracy of the similarity comparison results can be improved.
[0017] In some embodiments, the reference text is extracted from one or more files, and the text library further includes a source identifier of each reference text;
[0018] After obtaining the target reference text based on the bidding text screening, the method further includes:
[0019] If a reference text with a specified source identifier exists in the text library, the reference text with the specified source identifier is used as the target reference text.
[0020] By filtering the target benchmark text according to the specified source identifier, the benchmark text that cannot be filtered out based on the bidding documents can be used as the target benchmark text to ensure the evaluation accuracy of the bidding documents.
[0021] In some embodiments, the step of performing a similarity comparison between the bid text in the bid document and the target reference text to obtain a target bid text that matches the target reference text includes:
[0022] dividing the bidding text into a plurality of bidding sub-texts;
[0023] For any of the target benchmark texts, a similarity comparison is performed between the target benchmark text and each of the bid sub-texts, and a preset number of bid sub-texts corresponding to the target benchmark text are selected in descending order of similarity values as target bid texts that match the target benchmark text.
[0024] Compared with filtering the target bid text according to thresholds or other methods, the preset number of bid sub-texts corresponding to the target benchmark text are used as target bid texts that match the target benchmark text in order of similarity values from high to low. This can ensure that a matching target bid text is found for each target benchmark text, thereby improving the comprehensiveness of the review.
[0025] In some embodiments, before comparing the bid text with the target reference text for similarity, the bid text is extracted from the bid document based on the following method:
[0026] converting the bid document into a bid image, and identifying element regions of each element of the bid document from the bid image;
[0027] The text in the element area is extracted, and the extracted text is used as the bidding text.
[0028] After converting the bidding document into a bidding image, the bidding text is extracted based on the bidding image, which can improve the accuracy of text extraction.
[0029] In some embodiments, the elements in the bid document include a table;
[0030] The step of dividing the bid text into a plurality of bid sub-texts includes:
[0031] Merge the bid texts extracted from the table area into rows and treat the text in each row as a bid sub-text; or
[0032] Merge the bid texts extracted from the table area by columns, and treat the text in each column as a bid sub-text.
[0033] Since texts in the same row or column in a table are related, after the bid texts extracted from the table area are merged by rows or columns, the resulting bid sub-texts have richer and more accurate semantic information, which can improve the accuracy of the similarity comparison when the bid sub-texts are compared with the target benchmark texts.
[0034] In some embodiments, the target benchmark text is specifically used to represent the content that the bidding document needs to include;
[0035] The step of performing a review operation on the target bid text according to the preset text review dimension includes:
[0036] The target reference text and the target bid text are input into a first evaluation model, and based on an output result of the first evaluation model, it is determined whether the target bid text contains the content represented by the target reference text.
[0037] The first evaluation model is used to determine whether the target bid text contains the content represented by the target benchmark text. The result obtained is relatively accurate and the method is relatively simple to implement.
[0038] In some embodiments, if the target bid text contains content represented by the target reference text, performing a review operation on the target bid text according to a preset text review dimension further includes:
[0039] Inputting the target bid text into a second evaluation model, and judging the logic quality of the target bid text based on the output result of the second evaluation model; and / or
[0040] When the target bid text includes technology-related content, inputting the target bid text into a third evaluation model, and judging the degree of sophistication of the technology represented by the target bid text based on an output result of the third evaluation model; and / or
[0041] The target bid text is input into a fourth evaluation model, and the degree of matching between the target bid text and the project characteristics is determined based on the output result of the fourth evaluation model, wherein the project characteristics are the project corresponding to the bid document.
[0042] When the target bid text contains content represented by the target benchmark text, the target bid text will continue to be reviewed according to the text review dimensions of technological advancement, practical integration and logic. The review of the target bid text will be more comprehensive and the review results will be more reasonable.
[0043] In some embodiments, when there are multiple text review dimensions, performing a review operation on the target bid text according to the preset text review dimensions to obtain a review result for characterizing the text quality of the bid document includes:
[0044] Performing review operations on the target bid texts according to each of the text review dimensions to obtain dimensional review results of the bid document under each of the text review dimensions;
[0045] The review result is generated based on the dimension review result and the dimension weight of each text review dimension.
[0046] When generating the evaluation results of the bidding documents, since the dimension weights of the text evaluation dimensions are taken into consideration, the evaluation results obtained can be more in line with the actual situation and have higher accuracy.
[0047] In some embodiments, the text library further includes review items and review elements, wherein a review item includes one or more review elements, and a review element includes one or more benchmark texts;
[0048] The step of performing a review operation on the target bid text according to the preset text review dimension to obtain a review result for characterizing the text quality of the bid document includes:
[0049] The evaluation results of the bidding document are generated based on the dimension evaluation results, the dimension weights of each text evaluation dimension, the benchmark text weights of each target benchmark text, the evaluation element weights of the target evaluation elements to which each target benchmark text belongs, and the evaluation item weights of the target evaluation items to which each target evaluation element belongs.
[0050] When generating the evaluation results of the bidding documents, the accuracy of the evaluation results can be further improved by taking into account the benchmark text weight of each target benchmark text, the evaluation element weight of the target evaluation element to which each target benchmark text belongs, and the evaluation item weight of the target evaluation item to which each target evaluation element belongs.
[0051] Another aspect of the present disclosure provides a text review system, comprising:
[0052] A first comparison module is used to compare the similarity between the tender text in the tender document and the reference text used to represent the tender requirements in the text library to obtain a target reference text that matches the tender text;
[0053] A second comparison module is used to compare the bid text in the bid document with the target reference text to obtain a target bid text that matches the target reference text;
[0054] The review module is used to perform a review operation on the target bidding text according to a preset text review dimension, and obtain a review result for characterizing the text quality of the bidding document.
[0055] On the other hand, the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0056] On the other hand, the present disclosure further provides an electronic device, which includes a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method described above is implemented.
[0057] Another aspect of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the method described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The features and advantages of the present disclosure will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present disclosure in any way. In the accompanying drawings:
[0059] Figure 1 A schematic diagram showing a flow chart of a text review method provided by an embodiment of the present disclosure is shown;
[0060] Figure 2 A schematic diagram of modules of a text review system provided by one embodiment of the present disclosure is shown;
[0061] Figure 3 A schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0062] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0063] In some technologies, manual review of bidding documents by reviewers based on the bidding documents may have the following problems:
[0064] 1) The review results are subjective and not very accurate;
[0065] 2) Low review efficiency.
[0066] In order to solve the above problems, the present disclosure first constructs a text library similar to that shown in Table 1.
[0067] Table 1 Text Library
[0068]
[0069]
[0070] In Table 1, the text library includes evaluation items, evaluation elements, and benchmark text. The benchmark text is extracted from one or more documents and used to represent the bidding requirements. The one or more documents include, but are not limited to, historical bidding documents, national standards, and industry specifications. For example, suppose a historical bidding document contains the following original content:
[0071] The contractor shall establish a fixed temporary garbage storage point on site and set up necessary garbage bins on each floor or area; all garbage must be removed from the site on the same day and transported to the designated garbage disposal site in accordance with the regulations of the relevant administrative departments.
[0072] After sorting and refining the above original content, the benchmark text "temporary garbage storage point" in Table 1 can be extracted.
[0073] Review elements are used to identify the categories to which each benchmark document belongs, and review items are used to identify the categories to which each review element belongs. A review item can include one or more review elements, and a review element can include one or more benchmark documents.
[0074] Based on the above examples, we can see that the bidding requirements represented by the benchmark text can specifically refer to the content that the bid document must include. For example, suppose one of the benchmark texts in the text library is "Construction Wastewater Treatment Measures." Then, the bid document must provide a specific treatment plan for construction wastewater.
[0075] Furthermore, the text library can include the full set of benchmark texts extracted from national standards, industry norms, and historical bidding documents for multiple engineering projects. Since the benchmark texts for different engineering projects may vary, the bidding documents for a specific engineering project can include the content represented by some of the benchmark texts in the text library. For example, the bidding documents for Project A include the content represented by benchmark text a and benchmark text b, while the bidding documents for Project B include the content represented by benchmark text a and benchmark text c.
[0076] In some embodiments, the text library may also include source identifiers for each benchmark text. The source identifier indicates the category of the document used to extract the benchmark text. For example, assuming benchmark text A is extracted from the content of a national standard document, and benchmark text B is extracted from the content of a historical bidding document, the source identifier for benchmark text A may be 1, and the source identifier for benchmark text B may be 2. Based on the source identifier, the source of the benchmark text can be determined.
[0077] It should be noted that Table 1 is only an exemplary description of the text library. During the actual implementation of the program, the review items, review elements and benchmark texts can be determined based on actual conditions. Therefore, the content shown in Table 1 does not constitute a limitation to the present disclosure.
[0078] Based on a pre-built text library, the present disclosure provides a text review method that can reduce human errors and subjective biases of reviewers, improve the accuracy of review results, and at the same time, improve review efficiency. The text review method can be applied to industrial software or to electronic devices that run industrial software. Electronic devices include but are not limited to desktop computers, tablet computers, laptops, servers, etc. Figure 1 , which is a flow chart of a text review method provided in one embodiment of the present disclosure. Figure 1 In [1], the text review method includes the following steps:
[0079] Step S11 : performing a similarity comparison between the bidding text in the bidding document and the reference text in the text library used to represent the bidding requirements, and obtaining a target reference text that matches the bidding text.
[0080] In this embodiment, the reference text in the text library that has the greatest similarity to the tender text can be used as the target reference text. Specifically, the tender text can be divided into multiple tender subtexts (e.g., multiple tender subtexts at the sentence level). The multiple tender subtexts can be sentence-level texts. For any tender subtext, the tender subtext can be compared with each reference text in the text library for similarity to screen the reference text with the greatest similarity to the tender subtext. After summarizing all the reference texts screened, a target reference text that matches the tender text can be obtained.
[0081] Since bidding documents typically specify the content to be included in bid documents, and since the bidding text is similar to the target benchmark text, it can be assumed that the content represented by the target benchmark text is the content to be included in the bid document as specified by the bidding documents. For example, suppose the target benchmark text is "Construction wastewater treatment measures." Then, the bid document corresponding to the bidding document should provide a specific treatment plan for construction wastewater.
[0082] Step S12: performing a similarity comparison between the bid text in the bid document and the target reference text to obtain a target bid text that matches the target reference text.
[0083] Specifically, the bid text can be divided into multiple bid sub-texts (such as multiple bid sub-texts at the sentence level). For any target benchmark text, the target benchmark text can be compared with each bid sub-text for similarity, and the bid sub-text that meets the similarity requirement with the target benchmark text (such as the similarity is greater than a threshold) is used as the target bid text that matches the target benchmark text. In this step, one or more target bid texts that match each target benchmark text can be screened out. For example, the target bid texts that match the target benchmark text A are text a, text b, and text c, and the target bid texts that match the target benchmark text B are text a1, text b1, and text c1.
[0084] It is understandable that if a bid document includes content represented by a target benchmark text, then this content should include words similar to the target benchmark text, or have semantics similar to the target benchmark text. Therefore, a target bid text that meets the similarity requirements with the target benchmark text can be used as a candidate text in the bid document. This candidate text may or may not represent the content represented by the target benchmark text. For example, if target benchmark text A is "24-hour security and protection service," among texts a, b, and c in the bid document, text a describes the regional divisions for providing security and protection services, text b describes the number of people and shift system for providing security and protection services, and text c describes the benefits of providing security and protection services. Because texts a, b, and c all involve vocabulary related to security and protection services, all three texts can serve as target bid texts (i.e., candidate texts) that match target benchmark text A. However, the content of text c has little to do with how security and protection services are provided, so it can be considered that texts a and b represent the content represented by target benchmark text A, while text c does not.
[0085] Step S13: performing an evaluation operation on the target bid text according to the preset text evaluation dimension to obtain an evaluation result for characterizing the text quality of the bid document.
[0086] Understandably, one of the criteria for judging the quality of a bid document is whether it meets the requirements of the bidding documents, national standards, and industry regulations. The key to determining whether a bid document meets these requirements is whether it includes the content represented by the target benchmark text.
[0087] Based on the above description, the aforementioned text review dimension can include content integrity. When evaluating the target bid document based on content integrity, the target bid document, which matches the target base text, is evaluated to determine whether the target bid document includes the content represented by the target base text. If the target bid document includes the content represented by the target base text, the bid document is of good quality; if the target bid document does not include the content represented by the target base text, the bid document is of poor quality.
[0088] Specifically, the evaluation results of the bid documents can be represented by scores. The higher the score, the better the quality of the bid document; the higher the score, the better the quality of the bid document. For any target benchmark text, you can search in the target bid text that matches the target benchmark text to see whether the content represented by the target benchmark text exists. If it exists, the target benchmark text can be marked as 1; if it does not exist, the target benchmark text can be marked as 0. Among all the target benchmark texts, if the proportion of target benchmark texts marked as 1 is large, the score of the bid document can be higher (that is, better quality); if the proportion of target benchmark texts marked as 0 is large, the score of the bid document can be lower (that is, poorer quality).
[0089] In summary, in the technical solutions of some embodiments of the present disclosure, by performing a similarity comparison between the tender text and the benchmark text in the text library, a target benchmark text corresponding to the tender document can be obtained. After performing a similarity comparison between the tender text and the target benchmark text, a target tender text that matches the target benchmark text in the tender document can be obtained. After performing a review operation on the target tender text according to the specified text review dimension, an review result that characterizes the quality of the tender document text can be generated. In this way, automatic review of the tender document is achieved, which can reduce the subjective bias of manual review and improve the accuracy of the review results.
[0090] Secondly, compared to manual review, automated review is faster and can effectively improve the efficiency of bid review. This can significantly shorten the review time and reduce the human resources required when there are a large number of bids.
[0091] The technical solution of the present disclosure is further described below in conjunction with some other embodiments.
[0092] In some embodiments, under normal circumstances, the content represented by the benchmark text extracted from the national standard documents and industry specification documents is the content that all bidding documents need to include. However, in the above step S11, after the similarity comparison between the bidding text and the benchmark text in the text library is performed, the benchmark text extracted from the national standard documents and industry specification documents may not be used as the target benchmark text, which results in that in steps S12 and S13, the bidding documents cannot be reviewed based on these benchmark texts (that is, it is impossible to determine whether the bidding documents include the content represented by these benchmark texts). In view of this, in the above step S11, after the target benchmark text is obtained based on the bidding text screening, the method of the present disclosure may also include:
[0093] If there are reference texts with the specified source identifier in the text library, these reference texts with the specified source identifier are used as target reference texts.
[0094] Specifically, the designated source identifier may be used to identify national standard documents and industry specification documents. That is, for the benchmark texts extracted from national standard documents and industry specification documents, if these benchmark texts are not selected as target benchmark texts in step S11, these benchmark texts are added to the target benchmark text.
[0095] In the above embodiment, the target reference text is screened according to the designated source identifier, and the reference text that cannot be screened based on the bidding document can be used as the target reference text to ensure the evaluation accuracy of the bidding document.
[0096] In some embodiments, in steps S11 and S12, when performing a similarity comparison between the tender text in the tender document and a benchmark text in the text library, or when performing a similarity comparison between the bid text in the tender document and a target benchmark text, a text-based embedding model can be used to calculate the text's semantic vector, and the cosine similarity of the semantic vectors can be used as the text similarity comparison result. For ease of understanding, the similarity comparison between the tender text and the benchmark text is used as an example. For any tender subtext A and any benchmark text B in the tender document, tender subtext A and benchmark text B are respectively input into the text embedding model to obtain text vector A for tender subtext A and text vector B for benchmark text B. The cosine similarity between text vector A and text vector B can be used as the similarity between tender subtext A and benchmark text B. A greater cosine similarity indicates a greater semantic similarity between tender subtext A and benchmark text B; a smaller cosine similarity indicates a greater semantic difference between tender subtext A and benchmark text B.
[0097] In some embodiments, since the reference text in the text library is a relatively summary text extracted from historical bidding documents, national standard documents, and industry specification documents, and the bidding subtext is the unextracted text in the bidding document, directly performing a similarity comparison between the bidding subtext and the reference text may affect the accuracy of the similarity comparison result. In view of this, the above-mentioned step S11 of performing a similarity comparison between the bidding text in the bidding document and the reference text used to represent the bidding requirements in the text library to obtain a target reference text that matches the bidding text may include:
[0098] Divide the tender text into multiple tender sub-texts;
[0099] For any benchmark text in the text library, the benchmark text is compared with each tender sub-text for similarity, so as to select tender sub-texts whose similarity with the benchmark text is higher than a threshold;
[0100] Generate summary texts for each of the filtered tender sub-texts;
[0101] The summary text is compared with the benchmark text in the text library for similarity, and the benchmark text that matches the summary text is used as the target benchmark text.
[0102] The bidding sub-text division and similarity comparison can be found in the relevant description of step S11 and will not be repeated here.
[0103] The summary text may be a text obtained by refining the filtered tender sub-texts. The summary text may reflect the core semantic information of the tender sub-texts.
[0104] The filtered tender sub-texts can be input into the first language model. By constructing appropriate prompt words, the first language model can be guided to grasp the core semantics and remove noise. The first language model can then output the first summary text of each tender sub-text.
[0105] After obtaining the first summary text, the second language model can be used to perform grammatical correction and semantic optimization on the summary text to obtain an optimized second summary text. For each second summary text, a similarity comparison can be performed with each benchmark text in the text library to select the benchmark text with the greatest similarity to the second summary text. After all the benchmark texts obtained by screening are aggregated, the aggregated benchmark text can be used as the target benchmark text.
[0106] In the above embodiment, by generating a summary text for the tender subtext and performing a similarity comparison between the summary text and the reference text, the text difference between the tender subtext and the reference text can be reduced, thereby improving the accuracy of the similarity comparison result.
[0107] In some embodiments, in the above step S12, performing a similarity comparison between the bid text in the bid document and the target reference text to obtain a target bid text that matches the target reference text may include:
[0108] Divide the bidding text into multiple bidding sub-texts;
[0109] For any target benchmark text, the target benchmark text is compared with each bid sub-text for similarity, and the first preset number of bid sub-texts corresponding to the target benchmark text are selected in descending order of similarity values as target bid texts that match the target benchmark text.
[0110] For example, the similarity comparison between target benchmark text A and each bidding sub-text is performed. The similarity values between target benchmark text A and each bidding sub-text are as follows:
[0111] The similarity value with bid subtext a is 0.9;
[0112] The similarity value with bid subtext b is 0.1;
[0113] The similarity value with the bid subtext c is 0.3;
[0114] The similarity value with the bid subtext d is 0.7;
[0115] The similarity value with the bid subtext e is 0.6;
[0116] Sorting in descending order of similarity, the order of the bid subtexts is as follows: bid subtext a, bid subtext d, bid subtext e, bid subtext c, bid subtext b.
[0117] Assuming that the preset number is 3, the target bid texts that match the target reference text A are bid sub-text a, bid sub-text d, and bid sub-text e.
[0118] In the above embodiment, compared with filtering the target bid text according to thresholds or the like, a preset number of bid sub-texts corresponding to the target benchmark text are used as target bid texts that match the target benchmark text in order of similarity values from high to low. This ensures that a target bid text that matches each target benchmark text is found, thereby improving the comprehensiveness of the review.
[0119] In some embodiments, after obtaining target bid texts that match each target benchmark text, the step S13 of performing a review operation on the target bid text according to a preset text review dimension may include:
[0120] The target benchmark text and the target bid text are input into the first evaluation model, and based on the output result of the first evaluation model, it is determined whether the target bid text contains the content represented by the target benchmark text.
[0121] Among them, the first review model can be a trained third language model. For any target benchmark text, the target benchmark text and the target bid text that matches the target benchmark text can be input into the third language model. By constructing appropriate prompt words, the third language model can determine whether the target bid text includes the content represented by the target benchmark text. Specifically, the output result of the third language model can include identification information of each target bid text. If a target bid text is identified as a first value (such as 1), it means that the target bid text includes the content represented by the target benchmark text; if a target bid text is identified as a second value (such as 0), it means that the target bid text does not include the content represented by the target benchmark text.
[0122] For example, suppose that after the target base text A, the target bid subtext a, target bid subtext d, and target bid subtext c that match the target base text A are input into the third language model, the output of the third language model is as follows:
[0123] Target bid subtext a 1
[0124] Target bid subtext b 0
[0125] Target bid subtext c 1
[0126] Then it can be determined that the target bidding subtext a and the target bidding subtext c include the content represented by the target reference text A, and the target bidding subtext b does not include the content represented by the target reference text A.
[0127] In the above embodiment, the first evaluation model is used to determine whether the target bid text contains the content represented by the target reference text. The obtained result is relatively accurate and the method is relatively simple to implement.
[0128] In some embodiments, the text review dimensions may also include technological advancement, practical integration, and logic. Specifically, when evaluating the target bid text based on technological advancement, the degree of technological advancement represented by the target bid text is determined; when evaluating the target bid text based on practical integration, the degree of match between the target bid text and the project characteristics is determined, where the project characteristics are the project corresponding to the bid document; and when evaluating the target bid text based on logic, the degree of logic of the target bid text is determined.
[0129] Based on the above description, if the target bid text contains the content represented by the target benchmark text, the target bid text is reviewed according to the preset text review dimensions, further including:
[0130] Inputting the target bid text into the second evaluation model and judging the logic quality of the target bid text based on the output of the second evaluation model; and / or
[0131] When the target bid text includes technology-related content, input the target bid text into the third evaluation model and, based on the output of the third evaluation model, determine the level of sophistication of the technology represented by the target bid text; and / or
[0132] The target bid text is input into the fourth evaluation model, and the matching degree between the target bid text and the engineering project characteristics is determined based on the output result of the fourth evaluation model, wherein the engineering project characteristics are the engineering projects corresponding to the bid documents.
[0133] For example, assuming that the target bid sub-text a, target bid sub-text d and target bid sub-text c match the target benchmark text A, where the target bid sub-text a and target bid sub-text c include the content represented by the target benchmark text A, and the target bid sub-text b does not include the content represented by the target benchmark text A, then according to the text review dimensions of technological advancement, practical integration and logic, the review operation can continue to be performed on the target bid sub-text a and target bid sub-text c, and the review operation will no longer be performed on the target bid sub-text b.
[0134] Specifically, the second review model may be a trained fourth language model, the third review model may be a trained fifth language model, and the fourth review model may be a trained fifth language model.
[0135] In the above embodiment, when the target bid text contains content represented by the target benchmark text, the target bid text is continuously reviewed according to the text review dimensions of technological advancement, practical integration and logic. The review of the target bid text is more comprehensive and the review results are more reasonable.
[0136] Furthermore, when the target bid document is reviewed according to multiple text review dimensions, each text review dimension can have a corresponding dimension review result. After the dimension review results of the target bid document under each text review dimension are integrated, the resulting result can be used as the review result of the bid document.
[0137] Specifically, the dimensional review results of the target bid text under each text review dimension can be represented by a score. The higher the score of the target bid text under a text review dimension, the better the quality of the bid document under that text review dimension. The scores of the target bid text under each text review dimension are integrated (for example, added together), and the total score obtained can be used as the review result of the bid document. The higher the total score, the better the quality of the bid document when the bid document is reviewed according to the bidding documents; the lower the total score, the worse the quality of the bid document when the bid document is reviewed according to the bidding documents.
[0138] Furthermore, it is understandable that when performing an evaluation operation on the target bid text, the importance of different text review dimensions may not be exactly the same. For example, the importance of content integrity may generally be relatively high, but the importance of technological advancement, practical integration, and logic may generally be relatively low. When the scores of the target bid text under various text review dimensions are integrated, if the dimension review results under all text review dimensions are processed according to the same importance, the evaluation results of the obtained bid document will be biased. In view of this, in some embodiments, the evaluation operation is performed on the target bid text according to the preset text review dimensions in step S13 to obtain the evaluation results used to characterize the text quality of the bid document, which may include:
[0139] According to each text review dimension, the target bidding text is reviewed and the dimensional review results of the bidding document under each text review dimension are obtained;
[0140] The evaluation results of the bidding documents are generated based on the dimension evaluation results and the dimension weights of each text evaluation dimension.
[0141] Specifically, for text review dimensions with relatively high importance, the dimension weights may be relatively large; and for text review dimensions with relatively low importance, the dimension weights may be relatively small.
[0142] In the above embodiment, when generating the evaluation results of the bidding documents, since the dimension weights of the text evaluation dimensions are taken into consideration, the obtained evaluation results can be more consistent with the actual situation and have higher accuracy.
[0143] In some embodiments, the evaluation results of the bidding documents can also be generated based on the dimension evaluation results, the dimension weights of each text evaluation dimension, the benchmark text weights of each target benchmark text, the evaluation element weights of the target evaluation elements to which each target benchmark text belongs, and the evaluation item weights of the target evaluation items to which each target evaluation element belongs.
[0144] In the above embodiment, when generating the evaluation results of the bidding documents, the accuracy of the evaluation results can be further improved because the benchmark text weights of each target benchmark text, the evaluation element weights of the target evaluation elements to which each target benchmark text belongs, and the evaluation item weights of the target evaluation items to which each target evaluation element belongs are taken into consideration.
[0145] The following example illustrates how to determine the evaluation results of the bidding documents while taking weights into consideration.
[0146] 21) In the above step S11, target reference text A11, target reference text A12, target reference text B11, target reference text B12, target reference text C11 and target reference text C12 are obtained, wherein target reference text A11 and target reference text A12 belong to review element A1, target reference text B11 and target reference text B12 belong to review element B1, target reference text C11 and target reference text C12 belong to review element C1, review element A1 belongs to review item A, review element B1 and review element C1 belong to review item B; and,
[0147] The review item weight of review item A is 0.4, and the review item weight of review item B is 0.6;
[0148] The evaluation factor weight of evaluation factor A1 is 1, the evaluation factor weight of evaluation factor B1 is 0.6, and the evaluation factor weight of evaluation factor C1 is 0.4;
[0149] The weight of target reference text A11 is 0.6, the weight of target reference text A12 is 0.4, the weight of target reference text B11 is 0.3, the weight of target reference text B12 is 0.7, the weight of target reference text C11 is 0.2, and the weight of target reference text C12 is 0.8;
[0150] 22) The target bidding documents extracted from the bidding documents are evaluated for content completeness, technological advancement, practical integration and logic, with the dimension weight of content completeness being 0.5, the dimension weight of technological advancement being 0.2, the dimension weight of practical integration being 0.1 and the dimension weight of logic being 0.1.
[0151] Assuming the full score of the bid text is 100 points, then after allocating 100 minutes according to the weights:
[0152] The score of evaluation item A is 100*0.4=40 points;
[0153] The score of evaluation item B is 100*0.6=60 points;
[0154] Under evaluation item A, the score of evaluation factor A1 is 40*1=40 points;
[0155] Under evaluation item B, the score of evaluation factor B1 is 60*0.6=36 points, and the score of evaluation factor C1 is 60*0.4=24 points;
[0156] Under evaluation factor A1, the score of target benchmark document A11 is 40*0.6=24, and the score of target benchmark document A12 is 40*0.4=16;
[0157] Under evaluation element B1, the score for target benchmark document B11 is 36*0.3=10.8 points, and the score for target benchmark document B12 is 36*0.7=25.2 points.
[0158] Under evaluation factor C1, the score of target benchmark document C11 is 24*0.2=4.8, and the score of target benchmark document C12 is 24*0.8=19.2 points;
[0159] Under target benchmark text A11, the score for content completeness is 24*0.5=12 points, the score for technological advancement is 24*0.2=4.8 points, the score for practical integration is 24*0.1=2.4 points, and the score for logic is 24*0.1=2.4 points.
[0160] Under target benchmark text A12, the score for content completeness is 16*0.5=8 points, the score for technological advancement is 16*0.2=3.2 points, the score for practical integration is 16*0.1=1.6 points, and the score for logic is 16*0.1=1.6 points.
[0161] By analogy, the scores of each text review dimension under each target benchmark text can be calculated.
[0162] Assume that the bid document's initial score is 0. For target reference text A11, if target bid subtext 1 in the bid document, which matches target reference text A11, includes the content represented by target reference text A11 (i.e., satisfies content completeness), then the bid document's score is 12 points. If target bid subtext 1 satisfies technological advancement, the bid document's score is increased by 4.8 points, resulting in a score of 12 + 4.8 = 16.8 points. If target bid subtext 1 does not satisfy practical integration, the bid document's score remains at 16.8 points (i.e., no additional 2.4 points are added). If target bid subtext 1 satisfies logic requirements, the bid document's score is increased by 2.4 points, resulting in a score of 16.8 + 2.4 = 19.2 points.
[0163] And so on, by adding up the scores under each target benchmark text, we can finally get the score of the bidding document (i.e. the evaluation result).
[0164] This completes the instructions for calculating the review results.
[0165] In some embodiments, considering that bid documents are relatively complex (e.g., including tables, lists, etc.), extracting bid text from bid documents becomes a challenge. To this end, the present disclosure provides a solution. Specifically, before performing a similarity comparison between the bid text and the target reference text, the bid text can be extracted from the bid document based on the following steps 11) to 12).
[0166] 11) Converting the bidding document into a bidding image, and identifying element regions of various elements of the bidding document from the bidding image.
[0167] A bid image is an image file converted from a bid document. For example, a bid document in PDF format can be converted to a bid image using the pdf2image tool.
[0168] The elements of a bidding document refer to the content categories included in the bidding document, including but not limited to tables, titles, and paragraphs.
[0169] Based on a trained layout analysis model, a bid image can be analyzed to determine the element regions of each element in the bid document. Specifically, the bid image can be resized and normalized according to the input requirements of the layout analysis model. After the resized and normalized bid image is input into the layout analysis model, the layout analysis model outputs the bid image with a review frame and a confidence score. The review frame has a category label that corresponds to an element category. For example, a review frame with category label A identifies the title region; a review frame with category label B identifies the table region; and a review frame with category label C identifies the paragraph region. Each review frame has an associated confidence score, which represents the accuracy of the element category identified in the region where the review frame is located. For example, suppose a review frame with category label A identifies the title region. After region B of the bid image is annotated using the review frame with category label A, the corresponding confidence score is 0.7, indicating that there is a 70% probability that the content in region B is a title.
[0170] After obtaining the bid image with review frames and credibility, redundant review frames and those with low credibility can be removed. After the review frames are removed, the area defined by the remaining review frames is the element area. Based on the category labels of the review frames, the element category to which the text in the element area belongs can be identified.
[0171] 12) Extracting text from the element area and using the extracted text as the bid text.
[0172] Specifically, text may be extracted from each element region based on OCR (Optical Character Recognition) technology.
[0173] In the above embodiment, after the bidding document is converted into a bidding image, the bidding text is extracted based on the bidding image, which can improve the accuracy of text extraction.
[0174] For ease of understanding, assume that the review box with category label A is used to identify the area where the title is located; the review box with category label B is used to identify the area where the table is located; and the review box with category label C is used to identify the area where the paragraph is located.
[0175] First, title text can be extracted from the bid image. To extract title text from the bid image, the first review frame with the category label A can be searched for. Within the title area defined by the first review frame, title text can be extracted using optical character recognition (OCR) technology. However, since titles have a hierarchical relationship, this disclosure also provides a method for determining the hierarchical relationship of titles based on their location and the similarity between them. Specifically, if Title A and Title B are adjacent titles and their similarity exceeds a threshold, then Title A and Title B can be determined to be in a superior-subordinate relationship. Based on the positions of the two titles in the bid image, the superior and subordinate titles within the two titles can be determined. For example, if Title A is in front of Title B, then Title A is the superior title of Title B. Conversely, if Title A and Title B are adjacent titles and their similarity does not exceed the threshold, then Title A and Title B can be determined to be in a peer-level relationship. Based on their positions in the bid image, the order of the two titles can be determined. In this way, a title tree for the bid text can be constructed. Based on the title tree, you can extract the text under each title by title. For any title, you can extract the paragraph text and table text under the title as follows.
[0176] Specifically, when extracting paragraph text from the bidding image, the second review frame with the category label C may be searched, and the paragraph text may be extracted from the paragraph area defined by the second review frame based on the OCR technology.
[0177] When extracting table text from a bid image, a third review frame with a category label of B is searched for. Within the table area defined by the third review frame, a table line extraction algorithm based on image morphology is used to evaluate the horizontal and vertical lines of the table. Based on the extracted horizontal and vertical lines, the table area is segmented into cells, the intersections of the table lines are identified, and the boundaries of each cell are determined to obtain the cell area. Within each cell area, the table text is extracted using optical character recognition (OCR) technology.
[0178] In some embodiments, the bid texts extracted from the table area can be merged by rows, and the text in each row can be used as a bid sub-text; or the bid texts extracted from the table area can be merged by columns, and the text in each column can be used as a bid sub-text.
[0179] Since texts in the same row or column in a table are related, after the bid texts extracted from the table area are merged by rows or columns, the resulting bid sub-texts have richer and more accurate semantic information, which can improve the accuracy of the similarity comparison when the bid sub-texts are compared with the target benchmark texts.
[0180] This completes the relevant description of the disclosed solution.
[0181] Corresponding to the bidding method, the present disclosure also provides a text review system. Figure 2 , which is a module diagram of a text review system provided by an embodiment of the present disclosure.
[0182] A first comparison module is used to compare the similarity between the tender text in the tender document and the reference text used to represent the tender requirements in the text library to obtain a target reference text that matches the tender text;
[0183] The second comparison module is used to compare the bid text in the bid document with the target reference text to obtain a target bid text that matches the target reference text;
[0184] The review module is used to perform review operations on the target bidding text according to the preset text review dimensions, and obtain review results that are used to characterize the text quality of the bidding document.
[0185] In some embodiments, the first comparison module is specifically configured to:
[0186] Divide the tender text into multiple tender sub-texts;
[0187] For any benchmark text in the text library, the benchmark text is compared with each tender sub-text for similarity, so as to select tender sub-texts whose similarity with the benchmark text is higher than a threshold;
[0188] Generate summary texts for each of the filtered tender sub-texts;
[0189] The summary text is compared with the benchmark text in the text library for similarity, and the benchmark text that matches the summary text is used as the target benchmark text.
[0190] In some embodiments, the reference texts are extracted from one or more files, and the text library further includes a source identifier of each reference text;
[0191] After obtaining the target reference text based on the bidding text screening, the first comparison module is further used to:
[0192] If a base text with the specified source ID exists in the text library, the base text with the specified source ID is used as the target base text.
[0193] In some embodiments, the second comparison module is specifically configured to:
[0194] Divide the bidding text into multiple bidding sub-texts;
[0195] For any target benchmark text, the target benchmark text is compared with each bid sub-text for similarity, and the first preset number of bid sub-texts corresponding to the target benchmark text are selected in descending order of similarity values as target bid texts that match the target benchmark text.
[0196] In some embodiments, before comparing the bid text with the target reference text for similarity, the second comparison module further extracts the bid text from the bid document based on the following method:
[0197] converting the bidding document into a bidding image, and identifying element regions of each element of the bidding document from the bidding image;
[0198] Extract the text in the element area and use the extracted text as the bid text.
[0199] In some embodiments, the elements in the bid document include a table; and the second comparison module is specifically configured to:
[0200] Merge the bid texts extracted from the table area into rows and treat the text in each row as a bid sub-text; or
[0201] Merge the bid texts extracted from the table area by columns, and treat the text in each column as a bid sub-text.
[0202] In some embodiments, the target benchmark text is used to represent the content that the bid document needs to include; the review module is specifically used to:
[0203] The target benchmark text and the target bid text are input into the first evaluation model, and based on the output result of the first evaluation model, it is determined whether the target bid text contains the content represented by the target benchmark text.
[0204] In some embodiments, if the target bid text contains content represented by the target benchmark text, the review module is specifically configured to:
[0205] Inputting the target bid text into the second evaluation model and judging the logic quality of the target bid text based on the output of the second evaluation model; and / or
[0206] When the target bid text includes technology-related content, input the target bid text into the third evaluation model and, based on the output of the third evaluation model, determine the level of sophistication of the technology represented by the target bid text; and / or
[0207] The target bid text is input into the fourth evaluation model, and the matching degree between the target bid text and the engineering project characteristics is determined based on the output result of the fourth evaluation model, wherein the engineering project characteristics are the engineering projects corresponding to the bid documents.
[0208] In some embodiments, when there are multiple text review dimensions, the review module is specifically configured to:
[0209] According to each text review dimension, the target bidding text is reviewed and the dimensional review results of the bidding document under each text review dimension are obtained;
[0210] The review results are generated based on the dimension review results and the dimension weights of each text review dimension.
[0211] In some embodiments, the text library further includes review items and review elements, wherein a review item includes one or more review elements, and a review element includes one or more benchmark texts; the review module is specifically configured to:
[0212] The evaluation results of the bidding documents are generated based on the dimension evaluation results, the dimension weights of each text evaluation dimension, the benchmark text weights of each target benchmark text, the evaluation element weights of the target evaluation elements to which each target benchmark text belongs, and the evaluation item weights of the target evaluation items to which each target evaluation element belongs.
[0213] See also Figure 3 , is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. The electronic device includes a processor and a memory, wherein the memory is used to store a computer program. When the computer program is executed by the processor, the above method is implemented.
[0214] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0215] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods described in the embodiments of the present invention. The processor executes the non-transitory software programs, instructions, and modules stored in the memory to perform various processor functions and data processing, thereby implementing the methods described in the aforementioned method embodiments.
[0216] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0217] An embodiment of the present disclosure further provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, the above method is implemented.
[0218] The present disclosure also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0219] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A text review method, characterized in that: The method comprises: Performing a similarity comparison between the tender text in the tender document and a reference text in a text library used to represent the tender requirements, to obtain a target reference text that matches the tender text; Comparing the bid text in the bid document with the target reference text for similarity to obtain a target bid text that matches the target reference text; Performing an evaluation operation on the target bid text according to a preset text evaluation dimension to obtain an evaluation result for characterizing the text quality of the bid document; The step of performing a similarity comparison between the tender text in the tender document and a reference text in a text library used to represent the tender requirements to obtain a target reference text that matches the tender text includes: Dividing the bidding text into multiple bidding sub-texts; For any reference text in the text library, performing a similarity comparison between the reference text and each of the tender sub-texts to select tender sub-texts whose similarity to the reference text is higher than a threshold; Generate summary texts for each of the filtered tender sub-texts; The summary text is compared with the reference text in the text library for similarity, and the reference text that matches the summary text is used as the target reference text.
2. The method according to claim 1, wherein The reference text is extracted from one or more files, and the text library also includes a source identifier of each reference text; After obtaining the target reference text based on the bidding text screening, the method further includes: If a reference text with a specified source identifier exists in the text library, the reference text with the specified source identifier is used as the target reference text.
3. The method according to claim 1, wherein The step of performing a similarity comparison between the bid text in the bid document and the target reference text to obtain a target bid text that matches the target reference text includes: dividing the bidding text into a plurality of bidding sub-texts; For any of the target benchmark texts, a similarity comparison is performed between the target benchmark text and each of the bid sub-texts, and a preset number of bid sub-texts corresponding to the target benchmark text are selected in descending order of similarity values as target bid texts that match the target benchmark text.
4. The method according to claim 3, wherein Before comparing the similarity between the bid text and the target reference text, the bid text is extracted from the bid document based on the following method: converting the bid document into a bid image, and identifying element regions of each element of the bid document from the bid image; The text in the element area is extracted, and the extracted text is used as the bidding text.
5. The method according to claim 4, wherein The elements of the said tender documents include forms; The step of dividing the bid text into a plurality of bid sub-texts includes: Merge the bid texts extracted from the table area into rows and treat the text in each row as a bid sub-text; or Merge the bid texts extracted from the table area by columns, and treat the text in each column as a bid sub-text.
6. The method according to claim 1, wherein The target benchmark text is specifically used to represent the content that the bidding document needs to include; The step of performing a review operation on the target bid text according to the preset text review dimension includes: The target reference text and the target bid text are input into a first evaluation model, and based on an output result of the first evaluation model, it is determined whether the target bid text contains the content represented by the target reference text.
7. The method according to claim 6, wherein If the target bid text contains the content represented by the target reference text, performing the review operation on the target bid text according to the preset text review dimension further includes: Inputting the target bid text into a second evaluation model, and judging the logic quality of the target bid text based on the output result of the second evaluation model; and / or When the target bid text includes technology-related content, inputting the target bid text into a third evaluation model, and judging the degree of sophistication of the technology represented by the target bid text based on an output result of the third evaluation model; and / or The target bid text is input into a fourth evaluation model, and the degree of matching between the target bid text and the project characteristics is determined based on the output result of the fourth evaluation model, wherein the project characteristics are the project corresponding to the bid document.
8. The method according to claim 1, wherein In the case where there are multiple text review dimensions, performing a review operation on the target bid text according to the preset text review dimensions to obtain a review result for characterizing the text quality of the bid document includes: Performing review operations on the target bid texts according to each of the text review dimensions to obtain dimensional review results of the bid document under each of the text review dimensions; The review result is generated based on the dimension review result and the dimension weight of each text review dimension.
9. The method according to claim 8, wherein The text library further includes review items and review elements, wherein one review item includes one or more review elements, and one review element includes one or more benchmark texts; The step of performing a review operation on the target bid text according to the preset text review dimension to obtain a review result for characterizing the text quality of the bid document includes: The evaluation results of the bidding document are generated based on the dimension evaluation results, the dimension weights of each text evaluation dimension, the benchmark text weights of each target benchmark text, the evaluation element weights of the target evaluation elements to which each target benchmark text belongs, and the evaluation item weights of the target evaluation items to which each target evaluation element belongs.
10. A text review system, characterized in that: The system comprises: A first comparison module is configured to perform a similarity comparison between the tender text in the tender document and a reference text in a text library used to represent the tender requirements, to obtain a target reference text that matches the tender text. Specifically, the tender text is divided into a plurality of tender sub-texts; for any reference text in the text library, the reference text is compared with each of the tender sub-texts for similarity, so as to screen out tender sub-texts having a similarity with the reference text above a threshold; a summary text is generated for each of the screened tender sub-texts; the summary text is compared with the reference text in the text library for similarity, and the reference text that matches the summary text is used as the target reference text; A second comparison module is used to compare the bid text in the bid document with the target reference text to obtain a target bid text that matches the target reference text; The review module is used to perform a review operation on the target bidding text according to a preset text review dimension, and obtain a review result for characterizing the text quality of the bidding document.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
12. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method, device and medium for reviewing bidding document
CN114595661A
Electronic bidding document and clause matching method and device and medium
CN116303909A