A method and system for processing bill of quantities based on dual parsing correction
By adopting a dual analytical correction method in the processing of bill of quantities, combining projective geometry, Hough transformation, Bayesian algorithm and large language model, the inefficiency and inaccuracy caused by manual verification of bill of quantities in the existing technology are solved, and efficient and accurate bill of quantities are achieved.
Patent Information
- Application Number
- CN202510151815.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The preparation of the existing bill of quantities completely relies on manual verification, which has huge workload and extremely low efficiency, which is prone to artificial missed items and wrong items, resulting in untimely and inaccurate modification of the list.
The bill of quantities processing method based on double analytical correction is adopted. By combining projective geometry and Hough transformation algorithm, the paper-to-electronic version is converted, and the Bayesian algorithm and large language model are used for semantic analysis, multiple analysis and correction are realized to ensure the accuracy and consistency of the extraction results.
It significantly reduces the manual workload, improves the accuracy and efficiency of the identification of the bill of quantities, can update the bill of quantities in a timely manner, and improves the overall efficiency and safety of the project.
Smart Images

Figure CN119623476B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for processing bill of quantities based on dual parsing and correction. Background Art
[0002] The bill of quantities is an important component of the bidding documents and engineering contract terms, and is also the common basis for highway engineering projects to control project costs, prepare engineering base bids, adjust quantities, and serve as the tender offers of bidders.
[0003] Accurately processing the bill of quantities is crucial because it directly affects the cost control, resource allocation, and construction plan of engineering projects. Each engineering category and the corresponding quantities in the list are related to the project budget and final expenditure. If the data is incorrect or incomplete, it will lead to cost overruns, construction delays, and even quality problems. Therefore, precise processing of the bill of quantities helps ensure the economy and execution efficiency of the project and reduces project risks.
[0004] However, the existing bill of quantities compilation work completely relies on manual checking, which has a huge workload and extremely low efficiency. This is a common problem in the engineering industry. After the design unit modifies the drawings, it is necessary to manually recheck and adjust the engineering quantities. This cumbersome work is extremely prone to human omissions and errors during the compilation process, resulting in untimely and inaccurate list modifications. Summary of the Invention
[0005] In order to solve the technical problems existing in the prior art that the bill of quantities compilation work completely relies on manual checking, has a huge workload and extremely low efficiency, which is a common problem in the engineering industry. After the design unit modifies the drawings, it is necessary to manually recheck and adjust the engineering quantities. This cumbersome work is extremely prone to human omissions and errors during the compilation process, resulting in untimely and inaccurate list modifications, the present invention provides a method and system for processing bill of quantities based on dual parsing and correction.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First Aspect
[0008] A method for processing bill of quantities based on dual parsing and correction provided by an embodiment of the present invention includes:
[0009] S1: Obtain the bill of quantities in the engineering drawings, where the bill of quantities includes a paper-based bill of quantities and an electronic bill of quantities;
[0010] S2: Combine the parallelism preservation principle of projective geometry and the Hough transform algorithm to perform electronic conversion on the paper-based bill of quantities;
[0011] S3: Extract multiple identified project categories and the identified quantities of each identified project category from the target bill of quantities, where the target bill of quantities includes an electronic bill of quantities and a converted paper bill of quantities;
[0012] S4: Combine the Bayesian algorithm to extract the initial text of the target bill of quantities;
[0013] S5: Combine a large language model with dynamic context tokens as input data to perform semantic parsing on the initial text, and output the semantic text of the target bill of quantities, where the semantic text includes multiple parsed project categories and the parsed quantities corresponding to each parsed project category;
[0014] S6: Align each identified project category and each parsed project category to the target entity;
[0015] S7: Compare the identified quantities corresponding to the identified project categories and the parsed quantities corresponding to the parsed project categories that belong to the same target entity;
[0016] S8: When the comparison results are all consistent, output the semantic text as the processing result of the bill of quantities. Otherwise, mark the comparison difference results in the semantic text.
[0017] Second aspect
[0018] An engineering quantity list processing system based on dual parsing and correction provided by an embodiment of the present invention includes:
[0019] A processor;
[0020] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the engineering quantity list processing method based on dual parsing and correction as described in the first aspect is implemented.
[0021] Third aspect
[0022] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored. When the program is executed by a processor, the engineering quantity list processing method based on dual parsing and correction as described in the first aspect is implemented.
[0023] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:
[0024] In the present invention, the processing method of the bill of quantities is divided into two parallel methods: text extraction based on images and text extraction based on semantic recognition. Based on entity alignment technology, consistency verification is performed on the two extracted results on the same target entity, avoiding the problem that the effectiveness and accuracy of the extraction results cannot be ensured in single text extraction. Before text extraction, in combination with the principle of maintaining parallelism in projective geometry and the Hough transform algorithm, the paper bill of quantities is corrected for electronic conversion, avoiding the problem that the extraction results are affected by irrelevant factors and ensuring the accuracy in the entire text extraction process. During the text extraction process, the Bayesian algorithm is combined to extract the initial text of the target bill of quantities, correcting the semantic rationality problem of the extracted initial text. Moreover, a large language model with dynamic context tokens as input data performs semantic parsing on the initial text, which can reduce the problem of error accumulation in semantic recognition results caused by the sliding window length constraint and the problem of context semantic recognition discontinuity during text recognition, improving the accuracy of the semantic text extraction results. Finally, according to the comparison results, the accuracy of the extracted text can be quickly verified, and the difference results of the comparison are marked and output, reminding relevant personnel to quickly lock and check the relevant content. This processing method of the bill of quantities with multiple parsing and correction reduces the manual workload, greatly increases the recognition accuracy and recognition efficiency of the bill of quantities, can update the bill of quantities in a timely manner after the engineering drawings are modified, and improves the overall efficiency of the project and the safety of the engineering project. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 It is a schematic flowchart of a method for processing a bill of quantities based on dual parsing and correction provided by an embodiment of the present invention;
[0027] Figure 2 It is a schematic structural diagram of a system for processing a bill of quantities based on dual parsing and correction provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The following will describe the technical solutions in the present invention with reference to the drawings.
[0029] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0030] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same.
[0031] In the embodiments of the present invention, sometimes a subscript such as W1 may be miswritten as a non-subscript form such as W1. When the difference is not emphasized, their intended meanings are the same.
[0032] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0033] Refer to the attached Figure 1 , which shows a schematic flowchart of a method for processing a bill of quantities based on dual parsing and correction provided by an embodiment of the present invention.
[0034] The embodiments of the present invention provide a method for processing a bill of quantities based on dual parsing and correction. This method can be implemented by a device for processing a bill of quantities based on dual parsing and correction, and the device for processing a bill of quantities based on dual parsing and correction can be a terminal or a server. The processing flow of the method for processing a bill of quantities based on dual parsing and correction can include the following steps:
[0035] S1: Obtain the bill of quantities in the engineering drawings.
[0036] Among them, the bill of quantities includes a paper version of the bill of quantities and an electronic version of the bill of quantities.
[0037] It should be noted that by parsing the engineering drawings to obtain the bill of quantities, the bill of quantities includes two forms: paper version and electronic version. The purpose of this step is to extract and record the engineering quantity information in the drawings, providing a data basis for subsequent data conversion and automated processing.
[0038] S2: Combine the parallelism preservation principle of projective geometry and the Hough transform algorithm to perform electronic conversion on the paper version of the bill of quantities.
[0039] Among them, the parallelism preservation principle of projective geometry is that in projective geometry, parallel lines can maintain their directional relationship after projective transformation, ensuring that the shape of the object can be truly restored in the transformed image. The Hough transform algorithm is an image processing algorithm mainly used to detect shapes (such as lines or circles) in an image. By mapping the points in the image to the parameter space, the Hough transform can effectively detect and identify straight lines or other shapes in the image. The parallelism preservation principle of projective geometry is used to ensure the true restoration of the paper image, and at the same time, the Hough transform algorithm is combined to detect the lines in the image, converting the paper engineering quantity list into an accurate electronic version, realizing the digital conversion of data and improving the recognition accuracy of the subsequent engineering quantity list.
[0040] In a possible implementation manner, S2 specifically includes:
[0041] S201: Collect the engineering quantity list image of the paper engineering quantity list.
[0042] S202: Combine the parallelism preservation principle and the Hough transform algorithm to establish an image correction matrix:
[0043] Among them, represents the inclination angle of the engineering quantity list image, represents the offset parameter including the horizontal offset distance of the engineering quantity list image and the vertical offset distance of the engineering quantity list image , represents the pixel point coordinates of the engineering quantity list image, a and b respectively represent the perspective parameters caused by the horizontal fold of the engineering quantity list image and the vertical fold of the engineering quantity list image, represents the rotation matrix for correcting the inclination angle of the engineering quantity list image, represents the translation matrix for correcting the offset distance of the engineering quantity list image, represents the perspective transformation matrix for correcting the fold of the engineering quantity list image, and sin and cos respectively represent the sine function and the cosine function.
[0044] Among them, and are solved by substituting the four corner points of the engineering quantity list image into the following formula and using the least squares method:
[0045] Among them, represents the expected coordinates of the corrected engineering quantity list image at the th corner point, represents the uncorrected engineering quantity list image at the iThe coordinates of the corner points.
[0046] Among them, the expected coordinates are usually determined in advance. During the image correction process, the expected coordinates represent the positions of the four corner points of the image in the ideal state, and are usually set according to the true geometric ratio and alignment requirements of the image. For example, assuming the goal is to correct the image into a rectangle, these four expected coordinates may be the coordinates of the four corner vertices of the rectangle. These coordinates are determined before correction and are used as the standard positions that the corrected image should reach. In practical applications, by aligning the corner point coordinates of the uncorrected image with these expected coordinates, the least squares method can solve the transformation parameters of the image, thereby correcting the distorted image into an ideal and flat electronic image.
[0047] S203: Use the image correction matrix to correct the engineering quantity list image, obtain the corrected engineering quantity list image, and output the corrected engineering quantity list image as the electronic version conversion result of the paper-based engineering quantity list:
[0048] Among them, represents the pixel point coordinates of the corrected engineering quantity list image.
[0049] Specifically, in this process, first, the image of the paper-based engineering quantity list is collected. Then, by combining the principle of maintaining parallelism and the Hough transform algorithm, an image correction matrix is established. This matrix corrects the image according to parameters such as the tilt angle, offset distance, and wrinkles of the image. Specifically, the least squares method is used to determine parameters such as the offset and tilt angle to ensure that the parallel relationship and true ratio in the image are restored. Finally, the correction matrix is used to adjust the image to make the image content upright and distortion-free, obtaining a corrected image for electronic version conversion. This step ensures the high-quality digitization of the paper-based list content and provides a data basis for subsequent accurate recognition.
[0050] S3: Extract multiple recognized engineering categories of the target engineering quantity list and the recognized engineering quantities of each recognized engineering category based on the object detection algorithm.
[0051] Among them, the target engineering quantity list includes the electronic version engineering quantity list and the converted paper-based engineering quantity list.
[0052] Among them, the object detection algorithm is a computer vision technology used to identify and locate specific objects in images. Identifying engineering categories refers to automatically detecting and classifying different components or item types in engineering drawings. Identifying engineering quantities refers to extracting the specific quantities or dimensions of each identified engineering category. The object detection algorithm is used to automatically identify each engineering category and the corresponding engineering quantities from the bill of quantities, covering both electronic and converted paper-based list contents, ensuring the accurate extraction of detailed engineering data for each item. The extraction result serves as the data basis for subsequent comparison, jointly ensuring the accuracy of the preparation of the bill of quantities.
[0053] In a possible implementation manner, S3 specifically includes:
[0054] S301: Determine the detection regions through a pre-trained object detection model. Among them, the detection regions include the engineering category detection region and the engineering quantity detection region. Among them, the object detection model includes the YOLO object detection model, the Faster R-CNN object detection model, and the SSD object detection model.
[0055] Among them, the YOLO (You Only Look Once) object detection model, the Faster R-CNN (Region-based Convolutional Neural Network) object detection model, and the SSD (Single Shot MultiBox Detector) object detection model can all perform object detection quickly.
[0056] S302: Perform text recognition on the engineering category detection region and the engineering quantity detection region respectively, and output the identified engineering category and the identified engineering quantity of the identified engineering category.
[0057] Optionally, in step S302, OCR (Optical Character Recognition) can be directly used for semantic recognition. Or the OCR combined with the Bayesian algorithm in step S4 can be used for recognition.
[0058] S303: Repeat steps S301 to S303 until the target bill of quantities is extracted completely, and output multiple identified engineering categories and the identified engineering quantities of each identified engineering category.
[0059] It should be noted that the system first uses a pre-trained object detection model (such as YOLO, Faster R-CNN, SSD) to identify the detection regions in the engineering drawing, including the engineering category and the engineering quantity region. Subsequently, in S302, text recognition is performed on these regions to extract the specific engineering category names and quantities. This process is repeated until all engineering projects and corresponding engineering quantities in the list are extracted and output, and finally a complete structured data list is generated. This method realizes automated data extraction and effectively reduces manual intervention. This text extraction based on object detection can quickly locate the key regions in the engineering drawing, ensure that only the required text content is extracted, and avoid the resource consumption of full-image scanning. At the same time, the recognition efficiency and accuracy are improved, greatly reducing the workload of manual annotation and manual verification, applicable to the automatic processing of large-scale engineering data, providing accurate comparison data for subsequent result comparison, and completing effective and reliable result cross-verification.
[0060] S4: Combine the Bayesian algorithm to extract the initial text of the target engineering quantity list.
[0061] The probability optimization scheme improves the recognition accuracy by combining the image recognition output (the result of the OCR model) and the language model (prior probability). The prior probability helps the model correct the recognition result according to language common sense, especially for fuzzy or noisy picture content. This method can significantly reduce errors.
[0062] Among them, the Bayesian algorithm is based on Bayes' theorem and calculates the posterior probability by combining the prior probability and conditional probability of an event. In text recognition, the Bayesian algorithm helps correct recognition errors through context information. Using the Bayesian algorithm to extract the initial text of the engineering quantity list, the recognition result is corrected through the prior probability and context information, thereby improving the accuracy of text extraction and reducing recognition errors.
[0063] In a possible implementation manner, S4 specifically includes:
[0064] S401: Identify the first initial text of the target engineering quantity list through OCR.
[0065] S402: Combine the language model and Bayesian probability to establish the prior probability describing the correctness of the language rule:
[0066] Among them, represents the first initial text, represents the target engineering quantity list, represents given in the case of the posterior probability of occurrence, represents the probability calculated by OCR that given in the case of The probability of occurrence represents the prior probability based on the language model where represents the th token appearing after the th token and represents the total number of tokens in the first initial text.
[0067] Optionally, the language model can be an n-gram model. The n-gram model provides the probability of a certain word sequence by calculating the probability of one word following another. For example, it is more reasonable for the word "bill of quantities" to be followed by "list" than by "weather". We can use this language model to constrain the OCR results and improve the rationality of the OCR output.
[0068] S403: Optimize the first initial text according to the prior probability and output the initial text:
[0069] where represents the initial text, represents taking when taking the maximum value.
[0070] In the actual application process, this process is actually using prior knowledge to help OCR make more accurate guesses to improve the recognition accuracy. Prior knowledge can be understood as the context common sense and language rules of the text content. We combine this knowledge in the way of Bayesian probability, so that OCR not only depends on the image itself to recognize text, but also can refer to language habits to improve the rationality of the recognition result, which is also the data basis for the subsequent semantic text recognition accuracy.
[0071] Specifically, OCR simply recognizes each character or word in the image, but this process is isolated and without the help of "context". It is like "blind guessing" - judging what the characters might be based only on the shape of the image. For example, it may recognize "工宂量清且", but because the image is not clear, it may recognize "宂" as something similar to "程", and "且" as "单". However, in engineering-related documents, "工程量榜" is common rather than "工宂量清且". This language habit is the so-called "prior probability". In other words, OCR can use a "lexicon" or "language model" when recognizing, which will tell the system that combinations such as "工程量榜" are more common and more in line with language habits. When OCR gets the recognition result of "工宂量清且", it will not output it immediately, but will evaluate the rationality of this result again. OCR will think: "The possibility of 'bill of quantities' appearing in engineering documents is much higher than 'gong 宂量清且'." So it uses Bayes' theorem to make a judgment and believes that the overall possibility of the recognition result "bill of quantities" is higher. In the end, OCR combines image information and language habits to output a more reasonable "bill of quantities" result. This method avoids OCR directly outputting erroneous content in image recognition, but instead uses "prior knowledge" to correct the result. Significantly reduce OCR's incorrect recognition of initial text.
[0072] S5: Combined with a large language model that uses dynamic context words as input data, the initial text is semantically parsed to output the semantic text of the target bill of quantities.
[0073] The semantic text includes a plurality of parsing engineering categories and the parsing engineering quantities corresponding to each parsing engineering category.
[0074] Among them, dynamic context tokens are a technology that gradually adjusts input tokens in natural language processing, allowing the model to focus on key contextual information more flexibly in different rounds of input. Each input dynamically adjusts the context tokens according to the changes in the previous tokens and context to capture a wider range of text dependencies. This mechanism helps the model process longer texts within a limited context window and avoids the loss of contextual relevance due to window limitations. The dynamic context tokens are input into the large language model to perform semantic parsing on the initial text, ensuring that the parsed project categories and project quantities are accurately extracted under the context, and finally generating a complete list text with semantic associations.
[0075] Optionally, the large language model may be a BERT model, a RoBERTa model, an ERNIE model, or an XLNet model.
[0076] It is understandable that existing large models have good recognition effects on short texts. However, since the extracted semantic texts are continuous, it is impossible to manually split them again, which is time-consuming and laborious. If splitting is carried out, the workload is too large. However, the recognition of such long texts is usually limited by the size of the model's context window, resulting in the inability to input them all at once, which may lead to the incomplete capture of contextually related content, and then the problem of semantic truncation in the extracted semantic texts. This has a huge impact on the accounting of engineering bills of quantities that require rigorous processing. Using dynamic context tokens as input data can overcome the limitations of the context window. During each input process, consider the bias of each round of input, the attention score for the predicted token, and the continuity of the input tokens to improve the context tokens of each round of input, so as to make up for the shortcomings of semantic truncation caused by the limited context window and the recognition accuracy of semantic texts.
[0077] In a possible implementation manner, S5 specifically includes:
[0078] S501: Constrained by the context window length of the large language model, extract the first token sequence from the initial text in sequence, and input the first token sequence into the large language model to output the first predicted token.
[0079] S502: Generate the topic of the first token sequence through the LDA model:
[0080] Among them, represents the topic, represents the probability distribution of the first token sequence with respect to each alternative topic z , represents z the occurrence probability in the first token sequence, represents taking the alternative topic when takes the maximum value, LDA represents the LDA model, represents associated with the set of topic words, represents the th token in the first token sequence, represents the occurrence probability in the topic , represents the token screening probability threshold.
[0081] Among them, the LDA (Latent Dirichlet Allocation) model is a topic model used for topic extraction in text analysis.
[0082] S503: Calculate the attention scores of each token in the first token sequence relative to the first predicted token, and select multiple tokens from the first token sequence whose attention scores are greater than a preset attention score:
[0083] Among them, represents the token relative to the first predicted token 's attention score, represents 's query vector, S represents the selected tokens, represents 's key vector, represents the key vector dimension, represents the attention score threshold, respectively represent the length of the selected tokens, the length of the topic token, and the length of the context window, represents the adaptive proportion coefficient.
[0084] In a possible implementation, calculate the adaptive proportion coefficient based on information density and information entropy. The specific calculation method of the adaptive proportion coefficient is:
[0085] Among them, represents S 's information density, represents S 's information entropy, represents the basic proportion coefficient, represents 's relative importance in the set S , log represents the logarithmic function, represents the j th token belonging to S.
[0086] Among them, information density refers to the degree of concentration of effective information in an information sequence or dataset. It can be understood as how much "important" information is contained in a piece of content. Information entropy is a measure of the uncertainty in a piece of information. This adaptive proportion coefficient dynamically adjusts the retention ratio of tokens in the window according to information density and information entropy. In the case of high information complexity, increase the attention ratio to new tokens, so as to ensure that important information will not be ignored due to being restricted by the context window. This mechanism effectively optimizes information selection, making it take into account both the integrity and importance of information when processing long texts, and improving the overall parsing accuracy and coherence.
[0087] It is understandable that more space is reserved for new tokens when the information is complex (high information entropy). This is because in such cases, more new information needs to be captured to avoid excessive occupation of existing information. Among them, the maximum value of the basic proportion coefficient is 0.5, and 0.2 can also be taken. In this way, while avoiding the problem of large deviations caused by multi-round recognition and retaining the tokens with higher contribution degrees in the previous round, the continuity of the entire text recognition can be ensured to the greatest extent.
[0088] S504: Concatenate the tokens corresponding to the theme and the selected tokens, and store the concatenated token sequence in the placeholder.
[0089] S505: Sequentially extract the second token sequence from the initial text, where the length of the second token sequence is the difference between the length of the context window and the length of the placeholder.
[0090] S506: Concatenate the placeholder with the second token sequence, and input the concatenated second token sequence into the large language model to output the second predicted token.
[0091] S507: Use the second token sequence as the first token sequence, and return to step S502 until the initial text is completely input, obtaining multiple second predicted tokens.
[0092] S508: Concatenate the obtained first predicted tokens and second predicted tokens, and use the concatenated result as the semantic text output.
[0093] Specifically, the entire process first takes the length of the context window of the large language model as the limit, gradually extracts token sequences from the initial text, and sequentially inputs them into the model for prediction. The LDA model generates themes for each sequence, and filters out the high-probability tokens related to the theme. Then, calculate the attention scores of each token for the predicted token, and select the tokens with scores higher than the threshold to capture semantic associations. Next, concatenate the theme tokens and the tokens with high attention scores and store them in the placeholder. Then, extract the second token sequence from the initial text and concatenate it with the placeholder, and input it into the large language model to generate the next predicted token. Repeat this process until the initial text is parsed completely, and concatenate all the predicted tokens and output them as the semantic text. This process ensures semantic continuity and the accuracy of extracting semantic text by dynamically constructing the context.
[0094] It should be noted that by dynamically adjusting the context, the tokens with high attention scores are stored in the placeholder together with the theme-related tokens, and then concatenated with the subsequent second token sequence. In this way, each input not only retains the current semantic focus but also introduces new tokens, gradually enriching the context information and avoiding semantic truncation. This mechanism ensures the coherence and integrity of the long text parsing, enables the full expression of semantic information, reduces the information loss caused by window limitations, and improves the overall parsing accuracy.
[0095] S6: Align each identified project category and each parsed project category to the target entity.
[0096] Among them, the target entity refers to the final standardized project category in the project list. To ensure that the data remains consistent and clear at each stage. By aligning the identified project categories with the parsed project categories to the target entity, the consistency of each project category in the list is ensured, providing standardized basic data for subsequent comparison and calculation.
[0097] In a possible implementation manner, S6 specifically includes:
[0098] S601: Perform vector conversion on each identified project category and each parsed project category through a word embedding model.
[0099] S602: Calculate the cosine similarity between the converted identified project category and the parsed project category.
[0100] S603: Align the identified project category and the parsed project category with a cosine similarity greater than the preset cosine similarity to the target entity, where the target entity is the identified project category or the parsed project category.
[0101] Specifically, convert the identified and parsed project categories into vector representations through a word embedding model, and then calculate the cosine similarity between them to evaluate their semantic proximity. Categories pairs with a cosine similarity greater than the set threshold are considered similar and aligned, so as to unify the identified and parsed project categories to the same target entity. This method ensures the consistency of project categories from different sources in data, which helps with accurate list comparison and calculation.
[0102] It should be noted that those skilled in the art can set the size of the preset cosine similarity according to actual needs, and the present invention does not limit this here.
[0103] S7: Compare the identified project quantity corresponding to the identified project category belonging to the same target entity with the parsed project quantity corresponding to the parsed project category.
[0104] S8: When the comparison results are all consistent, output the semantic text as the processing result of the bill of quantities; otherwise, mark the comparison difference results in the semantic text.
[0105] In a possible implementation manner, after S8, it further includes:
[0106] Issue a warning when there are comparison difference results.
[0107] It should be noted that when there are differences between the recognized engineering quantity and the parsed engineering quantity in the comparison result, the added warning mechanism will automatically issue a warning reminder. Ensuring that users can promptly discover and handle data inconsistency issues helps improve the accuracy and reliability of the bill of quantities.
[0108] In the actual application process, first, the data in the engineering drawings are parsed and converted into electronic versions. The object detection algorithm is used to extract the engineering categories and quantities, and then the Bayesian algorithm is combined to correct the initial text to ensure the accuracy of the text. Subsequently, semantic parsing is performed through the dynamic context token input to the large language model to generate a list text with complete semantics. Then, the recognized and parsed categories are aligned to the target entity to ensure data consistency, and the engineering quantities are compared to verify the accuracy. Finally, if the data is consistent, the result is output; otherwise, the differences are marked for review. This method effectively improves the accuracy and efficiency of list processing and reduces manual operations.
[0109] The beneficial effects brought by the technical solution provided in the embodiments of the present invention at least include:
[0110] In the present invention, the processing method of the bill of quantities is divided into two parallel methods: text extraction based on images and text extraction based on semantic recognition. And based on the entity alignment technology, the consistency verification of the two extracted results is carried out on the same target entity, avoiding the problem that the effectiveness and accuracy of the extraction result cannot be ensured in single text extraction. Before text extraction, in combination with the parallelism preservation principle of projective geometry and the Hough transform algorithm, the correction of the conversion of the paper bill of quantities into an electronic version is carried out, avoiding the problem that the extraction result is affected by irrelevant factors and ensuring the accuracy in the whole text extraction process. During the text extraction process, the Bayesian algorithm is combined to extract the initial text of the target bill of quantities, correcting the semantic rationality problem of the extracted initial text. And the large language model with the dynamic context token as the input data performs semantic parsing on the initial text, which can reduce the problem of error accumulation of semantic recognition results caused by the sliding window length constraint and the problem of context semantic recognition break in the text recognition process, improving the accuracy of the semantic text extraction result. Finally, according to the comparison result, the accuracy of the extracted text can be quickly verified, and the comparison difference result is marked and output, reminding relevant personnel to quickly lock and check the relevant content. This processing method of the bill of quantities with multiple parsing and correction reduces the manual workload, greatly increases the recognition accuracy and recognition efficiency of the bill of quantities, can update the bill of quantities in time after the engineering drawings are modified, and improves the overall efficiency of the project and the safety of the engineering project.
[0111] Refer to the attached Figure 2 illustrates the structural schematic diagram of a bill of quantities processing system based on dual parsing and correction provided by the present invention.
[0112] The present invention also provides a bill of quantities processing system 20 based on dual parsing and correction, which is applied to the above-mentioned bill of quantities processing method based on dual parsing and correction, and includes:
[0113] A processor 201.
[0114] A memory 202, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor 201, the bill of quantities processing method based on dual parsing and correction as in the method embodiment is implemented.
[0115] The bill of quantities processing system 20 provided by the present invention can execute the above-mentioned bill of quantities processing method based on dual parsing and correction and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.
[0116] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0117] In the present invention, the processing method for the bill of quantities is divided into two parallel methods: text extraction based on images and text extraction based on semantic recognition. And based on the entity alignment technology, the consistency verification of the two extraction results is carried out on the same target entity, avoiding the problem that the effectiveness and accuracy of the extraction result cannot be ensured in single text extraction. Before text extraction, in combination with the parallelism preservation principle of projective geometry and the Hough transform algorithm, the correction of the conversion of the paper bill of quantities into an electronic version is carried out, avoiding the problem that the extraction result is affected by irrelevant factors and ensuring the accuracy in the whole text extraction process. During the text extraction process, the Bayesian algorithm is combined to extract the initial text of the target bill of quantities, correcting the semantic rationality problem of the extracted initial text, and the large language model with dynamic context tokens as input data is used to perform semantic parsing on the initial text, which can reduce the problem of error accumulation of semantic recognition results caused by the sliding window length constraint and the problem of context semantic recognition fault during the text recognition process, improving the accuracy of the semantic text extraction result. Finally, according to the comparison result, the accuracy of the extracted text can be quickly verified, and the difference results of the comparison are marked and output to remind relevant personnel to quickly lock and check the relevant content. This processing method for the bill of quantities with multiple parsing and correction reduces the manual workload, greatly increases the recognition accuracy and recognition efficiency of the bill of quantities, can update the bill of quantities in time after the engineering drawings are modified, and improves the overall efficiency and project safety of the project.
[0118] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0119] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0120] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0121] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0122] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0123] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0124] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0125] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0126] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0127] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0128] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0129] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0130] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the engineering quantity list processing method based on dual parsing and correction as in the method embodiment.
[0131] A computer-readable storage medium provided by the present invention can implement the steps and effects of the engineering quantity list processing method based on dual parsing and correction in the above method embodiment. To avoid repetition, the present invention will not elaborate further.
[0132] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0133] In the present invention, the processing method for the bill of quantities is divided into two parallel methods: text extraction based on images and text extraction based on semantic recognition. And based on the entity alignment technology, the two extracted results are verified for consistency on the same target entity, avoiding the problem that the effectiveness and accuracy of the extraction result cannot be ensured by a single text extraction. Before text extraction, in combination with the principle of maintaining parallelism in projective geometry and the Hough transform algorithm, the paper bill of quantities is corrected for electronic conversion, avoiding the problem of irrelevant factors affecting the extraction result and ensuring the accuracy throughout the text extraction process. During the text extraction process, the Bayesian algorithm is combined to extract the initial text of the target bill of quantities, correcting the semantic rationality problem of the extracted initial text. And a large language model with dynamic context tokens as input data performs semantic parsing on the initial text, which can reduce the problem of error accumulation in semantic recognition results caused by the sliding window length constraint and the problem of context semantic recognition discontinuity during text recognition, improving the accuracy of the semantic text extraction result. Finally, according to the comparison result, the accuracy of the extracted text can be quickly verified, and the difference results of the comparison are marked and output to remind relevant personnel to quickly lock and check the relevant content. This processing method for the bill of quantities with multiple parsing and correction reduces the manual workload, greatly increases the recognition accuracy and recognition efficiency of the bill of quantities, can update the bill of quantities in a timely manner after the engineering drawings are modified, and improves the overall efficiency and project safety of the project.
[0134] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
[0135] The following points need to be explained:
[0136] (1) The drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.
[0137] (2) For clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layer or region is enlarged or reduced, that is, these drawings are not drawn according to the actual ratio. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under another element or there can be an intermediate element.
[0138] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0139] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for processing bill of quantities based on double analytical correction, characterized in that: include: S1: Obtain a bill of quantities in an engineering drawing, wherein the bill of quantities includes a paper bill of quantities and an electronic bill of quantities; S2: Combining the parallelism preservation principle of projective geometry and the Hough transform algorithm, converting the paper version of the bill of quantities into an electronic version; S3: extracting multiple identified engineering categories and identified engineering quantities of each identified engineering category of a target engineering quantity list based on a target detection algorithm, wherein the target engineering quantity list includes the electronic engineering quantity list and the converted paper engineering quantity list; S4: extracting the initial text of the target bill of quantities in combination with a Bayesian algorithm; S5: performing semantic parsing on the initial text in combination with a large language model using dynamic context words as input data, and outputting a semantic text of the target bill of quantities, wherein the semantic text includes a plurality of parsed engineering categories and parsed engineering quantities corresponding to each parsed engineering category; S6: aligning each recognition engineering category and each parsing engineering category to the target entity; S7: comparing the identification engineering quantity corresponding to the identification engineering category and the analysis engineering quantity corresponding to the analysis engineering category belonging to the same target entity; S8: if the comparison results are consistent, output the semantic text as the result of bill of quantities processing; otherwise, mark the comparison difference results in the semantic text; Among them, the parallelism preservation principle of projective geometry is that in projective geometry, parallel lines can maintain their directional relationship after projective transformation, ensuring that the transformed image can truly restore the shape of the object; Among them, dynamic context tokens are gradually adjusted input tokens in natural language processing, allowing the model to focus more flexibly on key contextual information in different rounds of input. Each input will dynamically adjust the context tokens according to the changes in previous tokens and context to capture a wider range of text dependencies.
2. The method for processing bill of quantities based on double analytical correction according to claim 1 is characterized in that: The S2 specifically includes: S201: Acquire a bill of quantities image of the paper version of the bill of quantities; S202: Combining the parallelism preservation principle and the Hough transform algorithm, an image correction matrix is established: in, Indicates the tilt angle of the bill of quantities image. Indicates the horizontal offset distance of the bill of quantities image. Vertical offset distance from the bill of quantities image The offset parameter, Represents the pixel coordinates of the bill of quantities image, a and b denote the perspective parameters caused by the horizontal and vertical folds of the bill of quantities image, respectively. represents the rotation matrix used to correct the tilt angle of the bill of quantities image, represents the translation matrix used to correct the offset distance of the bill of quantities image, represents the perspective transformation matrix used to correct the wrinkles of the bill of quantities image, sin and cos represent the sine function and cosine function respectively; S203: performing image correction on the bill of quantities image using the image correction matrix to obtain a corrected bill of quantities image, and outputting the corrected bill of quantities image as the electronic version conversion result of the paper version of the bill of quantities: in, Represents the pixel coordinates of the corrected bill of quantities image.
3. The method for processing bill of quantities based on double analytical correction according to claim 1 is characterized in that: The S3 specifically includes: S301: determining a detection area through a pre-trained target detection model, wherein the detection area includes an engineering category detection area and an engineering quantity detection area, wherein the target detection model includes a YOLO target detection model, a Faster R-CNN target detection model, and an SSD target detection model; S302: performing text recognition on the engineering category detection area and the engineering quantity detection area respectively, and outputting the identified engineering category and the identified engineering quantity of the identified engineering category; S303: Repeat step S301 to step S302 until the target bill of quantities is extracted, and output a plurality of identified engineering categories and the identified engineering quantities of each identified engineering category.
4. The method for processing bill of quantities based on double analytical correction according to claim 1 is characterized in that: The S4 specifically includes: S401: recognizing a first initial text of the target bill of quantities by OCR; S402: Combine the language model and Bayesian probability to establish a priori probability describing the correctness of the language rules: in, represents the first initial text, represents the target bill of quantities, Indicates that in a given In the case The posterior probability of occurrence, Indicates the value obtained by OCR calculation in a given In the case The probability of occurrence, Represents the language model based on The prior probability of Indicates word In the word The probability of appearing later, represents the total number of tokens in the first initial text; S403: Optimize the first initial text according to the prior probability and output the initial text: in, Represents the initial text, Indicates that When the maximum value is taken .
5. The method for processing bill of quantities based on double analytical correction according to claim 1 is characterized in that: The S5 specifically includes: S501: extracting a first word-gram sequence from the initial text in order based on the context window length of the large language model as a constraint, inputting the first word-gram sequence into the large language model, and outputting a first predicted word-gram; S502: Generate the topic of the first word-gram sequence through the LDA model: in, Indicates the subject, Represents the first word sequence about each candidate topic z The probability distribution of express z The probability of occurrence in the first word sequence, Indicates that The alternative topic when taking the maximum value, LDA represents the LDA model, Representation and The associated subject headings, Indicates the first word in the sequence word, express In Theme The probability of occurrence in Indicates the probability threshold of word unit screening; S503: Calculate the attention score of each word in the first word-unit sequence relative to the first predicted word-unit, and select multiple words whose attention scores are greater than a preset attention score from the first word-unit sequence: in, Representation word Relative to the first predicted word The attention score, express The query vector is S Indicates the selected word. express The key vector of represents the key vector dimension, represents the attention score threshold, Respectively represent the selected word length, topic word length and context window length, Represents the adaptive proportion coefficient; S504: splicing the word-grams corresponding to the topic and the selected word-grams, and storing the spliced word-gram sequence into a placeholder; S505: extracting a second word sequence from the initial text in order, wherein the length of the second word sequence is the difference between the length of the context window and the length of the placeholder; S506: concatenating the placeholder with the second word-gram sequence, inputting the concatenated second word-gram sequence into the large language model, and outputting a second predicted word-gram; S507: taking the second word-gram sequence as the first word-gram sequence, and returning to step S502 until the initial text input is completed, and obtaining a plurality of second predicted word-grams; S508: Concatenate the obtained first predicted word-unit and the obtained second predicted word-unit, and output the concatenation result as the semantic text.
6. The method for processing bill of quantities based on double analytical correction according to claim 5 is characterized in that: The adaptive proportion coefficient is calculated based on information density and information entropy, and the adaptive proportion coefficient is calculated in the following manner: in, express S The information density, express S The information entropy of represents the basic proportion coefficient, express In the collection S The relative importance of , log represents the logarithmic function, Indicates the first j A word.
7. The method for processing bill of quantities based on double analytical correction according to claim 1 is characterized in that: The S6 specifically includes: S601: Convert each recognition engineering category and each parsing engineering category into a vector through a word embedding model; S602: Calculate the cosine similarity between the converted recognition engineering category and the parsing engineering category; S603: Align the recognition engineering categories and the analysis engineering categories whose cosine similarity is greater than a preset cosine similarity to the target entity, wherein the target entity is the recognition engineering category or the analysis engineering category.
8. The method for processing bill of quantities based on double analytical correction according to claim 1 is characterized in that: After S8, the method further includes: In the event of a difference in the comparison results, an early warning is issued.
9. A bill of quantities processing system based on dual analytical correction, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method for processing a bill of quantities based on double analysis correction as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for processing a bill of quantities based on double analysis correction as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Table semantic information extraction method, system and equipment based on cell coordinate optimization and medium
CN116543404A
Table information updating method and system based on intelligent recognition text
CN119206756A
Cited By
Multi-source data difference attribution system and method based on tensor decomposition
CN121743562A
System and method for multi-source data discrepancy attribution based on tensor decomposition
CN121743562B