A project information analysis and duplicate checking method

By combining deep learning and semantic understanding, optical character recognition technology has solved the problem of difficult paper document recognition, enabling efficient processing and accurate evaluation of diverse project information, and ensuring the quality of project information and the reliability of deduplication.

CN120162301BActive Publication Date: 2026-05-05SHANDONG SHENGLI PROJECT MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG SHENGLI PROJECT MANAGEMENT CO LTD
Filing Date
2025-02-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively process and analyze diverse project information, especially due to the low accuracy in recognizing handwritten and printed paper documents, and the lack of systematic project information quality assessment and deduplication methods.

Method used

We employ deep learning-based optical character recognition combined with semantic understanding for secondary recognition and verification to detect the flatness of handwritten paper documents. Layered optical character recognition processes printed paper documents, and we evaluate the quality of project information through multi-dimensional indicators. We also establish a project information database for deduplication.

Benefits of technology

It improves the accuracy of paper document recognition, ensures the standardization and consistency of project information, provides comprehensive and in-depth project evaluation and rapid deduplication capabilities, reduces subjective interference, and improves resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162301B_ABST
    Figure CN120162301B_ABST
Patent Text Reader

Abstract

This invention discloses a method for project information analysis and deduplication, comprising the following steps: Step 1: Import project information, including bidding documents and other project documents. Different processing methods are selected based on the type of project information to analyze it and convert it into a standard format. Step 2: After conversion to the standard format, a fill-in check is performed. If the fill-in check passes, semantic recognition is performed to obtain the analysis results of the filled-in content. If the fill-in check fails, a prompt message is generated, prompting a check of the original project information. Step 3: When the fill-in analysis results are normal, project information analysis is performed to obtain the project analysis results. Step 4: The project information is then analyzed again. Step 5: After the project information deduplication is completed, deduplication information is generated and sent to a preset receiving terminal. This invention provides more accurate project information analysis and deduplication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis, specifically to a method for project information analysis and deduplication. Background Technology

[0002] With the booming development of the smart construction field, various projects are characterized by large scale, high complexity, and numerous participants. From the construction of large-scale smart city infrastructure to the building of intelligent residential communities and smart factories, massive amounts of project information require efficient processing and precise management.

[0003] Project information from different sources comes in various formats, including traditional paper documents, electronic documents such as tender documents, various forms, handwritten annotations on construction drawings, and paper archives of various approval forms, as well as electronic documents such as progress reports generated by project management software and electronic drawings delivered by the design team.

[0004] To ensure the quality of project information, it is necessary to analyze and check for duplicates. Therefore, a method for analyzing and checking project information is proposed. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method for project information analysis and plagiarism detection, comprising the following steps:

[0006] Step 1: Import project information, which includes bidding documents and other project documents. Analyze the project information using different processing methods based on its type and convert it into a standard format.

[0007] Step 2: After converting to standard format, the form is checked. If the check passes, semantic recognition is performed to obtain the analysis results of the filled content. If the check fails, a prompt message is generated, prompting the user to check the original project information.

[0008] Step 3: If there are no abnormalities in the analysis results, proceed with the project information analysis to obtain the project analysis results;

[0009] Step 4: Next, analyze the project information, extract keywords, and then check for duplicates in the project information.

[0010] Step 5: After the project information is checked for plagiarism, the plagiarism information is generated and sent to the preset receiving terminal.

[0011] Furthermore, the specific process in step one is as follows:

[0012] Other project document types include handwritten paper project documents, printed paper project documents, and electronic project documents;

[0013] When other project files are paper-based handwritten project files:

[0014] First, a deep learning-based optical character recognition method is used to perform preliminary recognition, converting handwritten text into preliminary recognition results;

[0015] When the initial recognition result shows suspected errors or unclear characters, a second recognition verification is performed using an optical character recognition method based on semantic understanding. The system automatically captures the text content before and after the character, combines industry domain knowledge graphs and semantic analysis algorithms to make corrections, and marks the corrected content to obtain the second recognition result. If there are no suspected errors or unclear characters, the system proceeds directly to the next step.

[0016] Finally, the recognition results are manually verified through the manual verification module. Once the manual verification is passed, the final recognition results are exported.

[0017] When other project files are paper-printed project files:

[0018] After scanning paper documents and converting them into digital images, a layered optical character recognition (OCR) operation is performed.

[0019] The first layer uses a high-precision general optical character recognition engine to quickly extract text content and identify the main text and basic format information;

[0020] The second layer enables the intelligent optical character recognition module, which employs differentiated recognition strategies for different elements such as titles, paragraphs, tables, and charts.

[0021] Finally, a two-way format check is performed. On the one hand, from a visual presentation perspective, the layout of each element in the digital image is checked to see if it conforms to industry standard format specifications.

[0022] On the other hand, from the perspective of text content logic, compare whether the information between different elements is consistent;

[0023] Finally, export the recognition results that have passed the two-way format verification.

[0024] When other project files are in electronic document format, convert them directly to the standard format.

[0025] When the project information is a bidding document, the bidding document includes paper bidding documents and electronic bidding documents. For paper bidding documents, they are processed in the same way as paper-printed project documents. For electronic bidding documents, they are directly converted into the standard format.

[0026] Furthermore, when the project information is other project files, the types of other project files include paper handwritten project files and paper printed project files. Before recognizing paper handwritten project files, a flatness test is required. Once the flatness test is passed, the next step of recognition is allowed.

[0027] Furthermore, the specific process for conducting flatness testing is as follows:

[0028] First, image acquisition is performed, that is, using an image acquisition device, image information of paper handwritten project documents is acquired at a 45-degree angle downward from the left, a 45-degree angle downward from the right, a 45-degree angle downward from the top, and a 45-degree angle downward from the bottom, to obtain the first image information, the second image information, the third image information, and the fourth image information.

[0029] Then, the center point T1 of the left edge, the center point T2 of the right edge, the center point T3 of the top edge, and the center point T4 of the bottom edge are extracted from the first image information.

[0030] Then, using the plane where the paper-based handwritten project documents are placed as the reference plane, measure the distances q1 between T1 and the reference plane, m1 between T2 and the reference plane, f1 between T3 and the reference plane, and e1 between T4 and the reference plane.

[0031] Set q1, m1, f1 and e1 together to obtain the first evaluation parameter A1(q1, m1, f1, e1);

[0032] Then, the second image information, the third image information, and the fourth image information are processed in the same way as the first image information to obtain the second evaluation parameter A2 (q2, m2, f2, e2), the third evaluation parameter A3 (q3, m3, f3, e3), and the fourth evaluation parameter A4 (q4, m4, f4, e4).

[0033] Then, the difference Aa1 between the first evaluation parameter and the second evaluation parameter, and the difference Aa2 between the third evaluation parameter and the fourth evaluation parameter are calculated.

[0034] The difference between the first evaluation parameter and the third evaluation parameter is Aa3, the difference between the first evaluation parameter and the fourth evaluation parameter is Aa4, the difference between the second evaluation parameter and the fourth evaluation parameter is Aa5, and the difference between the second evaluation parameter and the third evaluation parameter is Aa6.

[0035] If at least four of Aa1, Aa2, Aa3, Aa4, Aa5, and Aa6 exceed the preset range, the flatness test fails.

[0036] Furthermore, the specific process of manual inspection through the manual proofreading module is as follows:

[0037] When a second verification result exists, at least two people must be arranged to confirm the suspected error or unclear characters at least twice. If the two confirmation results are the same, the export is allowed.

[0038] If the two confirmation results are different, a re-verification is required.

[0039] Furthermore, the specific process for completing the inspection is as follows:

[0040] First, perform text region recognition, collect image information after filling in the text, locate the position of the text box in the image, and obtain its upper left corner coordinates (x1, y1) and lower right corner coordinates (x2, y2);

[0041] Then, the range of the text within the input box is identified, and the coordinates of the top left corner (x3, y3) and bottom right corner (x4, y4) of the text area are obtained.

[0042] Next, calculate the center coordinates of the text region, x-coordinate. 中 = (x3 + x4) / 2, y-coordinate 中 = (y3+y4) / 2;

[0043] Next, calculate the center coordinates of the fill box, x 框中 = (x1+x2) / 2, y-coordinate 框中 = (y1+y2) / 2;

[0044] Calculate the deviation value of the abscissa Δx=|x 中 -x 框中 | and the deviation value of the ordinate Δy=|y 中 -y 框中 |;

[0045] When both the horizontal axis deviation value Δx and the vertical axis deviation value Δy are less than the set threshold, the text is considered to be centered, indicating that the filling check has passed.

[0046] Furthermore, the project analysis results include both projects with anomalies and projects without anomalies;

[0047] The process of obtaining the project analysis results is as follows:

[0048] Extract the on-time completion rate of key nodes, average daily construction volume, and progress coordination of each subsystem from the project information;

[0049] The on-time completion rate of key nodes, average daily construction volume and the coordination of progress of each subsystem are processed to obtain project progress indicators;

[0050] Then, extract the percentage of batches that failed material quality inspection, the rework rate, and the average failure interval of the intelligent system from the project information;

[0051] The percentage of batches failing material quality inspection, rework rate, and average failure interval of intelligent system are processed to obtain quality control indicators;

[0052] Then extract the cost-benefit ratio, budget change frequency and magnitude, and cost deviation rate from the project information;

[0053] The cost-benefit ratio, budget change frequency and magnitude, and cost deviation rate are processed to obtain cost budget indicators;

[0054] The project schedule indicators, quality control indicators, and cost budget indicators are processed to obtain a comprehensive evaluation indicator. When the comprehensive evaluation indicator is less than the preset value, it indicates that the project is abnormal; otherwise, it indicates that the project is normal.

[0055] Furthermore, the process of obtaining project progress indicators is as follows: a numerical mapping set of project progress scores is pre-established, which includes the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem. After obtaining the specific scores of the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem from the numerical mapping set of project progress scores, different weights are assigned to the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem. Then, the sum of the scores after assigning weights is calculated, which is the project progress indicator.

[0056] The process for obtaining quality control indicators is as follows:

[0057] A quality control score mapping set was pre-established, which included the percentage of batches failing material quality inspection, rework rate, and average fault interval of intelligent system. After obtaining the specific scores of the percentage of batches failing material quality inspection, rework rate, and average fault interval of intelligent system from the control score mapping set, different weights were assigned to the percentage of batches failing building material quality inspection, rework rate, and average fault interval of intelligent system. Then, the sum of the weighted scores was calculated, which yielded the quality control indicators.

[0058] The process of obtaining cost budget targets is as follows:

[0059] A cost budget scoring numerical mapping set of cost-benefit ratio, budget change frequency and magnitude and cost deviation rate was pre-established. After obtaining the specific scores of cost-benefit ratio, budget change frequency and magnitude and cost deviation rate from the cost budget scoring numerical mapping set, different weights were assigned to cost-benefit ratio, budget change frequency and magnitude and cost deviation rate. Then, the sum of the scores after assigning weights was calculated, which is the cost budget indicator.

[0060] The process for obtaining the comprehensive evaluation indicators is as follows:

[0061] Mark the project schedule indicator as G1, the quality control indicator as G2, and the cost budget indicator as G3;

[0062] Assign weight W1 to G1, weight W2 to G2, and weight W3 to G3, such that W1 + W2 + W3 = 1, and W3 > W2 = W1;

[0063] The comprehensive evaluation index Gg can be obtained by using the formula G1*W1+G2*W2+G3*W3=Gg.

[0064] Furthermore, the plagiarism detection process in step four is as follows:

[0065] First, keyword extraction is performed: the project information text is cleaned to remove noisy data and obtain pre-processed thick text;

[0066] Then, a statistical method was used: the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm was applied to calculate the importance score of each word in the project information text;

[0067] Then, a rule-based approach was used to combine domain knowledge and business rules related to the project information to formulate keyword extraction rules;

[0068] Next, part-of-speech tagging and filtering are performed on the preprocessed text to identify words of different parts of speech, such as nouns, verbs, and adjectives.

[0069] Keyword filtering and sorting: After obtaining candidate keywords through the above process, the keywords are filtered according to the set threshold (such as TF-IDF score threshold) to remove words with low scores;

[0070] At the same time, the selected keywords are sorted according to their importance scores, and the top-ranking keywords are selected as the final project information keywords.

[0071] Perform duplicate project information checks and establish a project information database: Store existing project information in the database, with each piece of project information having a unique identifier;

[0072] Create an index for the keywords extracted from each item in the database;

[0073] For the project information to be checked for plagiarism, keywords are extracted according to the set keyword extraction rules, and the keywords of the project information to be checked for plagiarism are matched with the keywords of the existing project information in the database;

[0074] The system determines whether there is duplicate project information based on the matching results. If a project information with keywords that highly match the project information to be checked for duplicates is found in the database, it is considered to be duplicated; otherwise, the project information is considered to be unique.

[0075] After all keywords have been matched, plagiarism detection information is generated.

[0076] Furthermore, in step four, when the imported project information is a bidding document, the specific plagiarism check process is as follows:

[0077] Import the bidding documents into the preset standard format library and read their content:

[0078] Then, the read content is stored in the standard library line by line, segment by segment, or according to specific fields, according to the preset standard library data entry specifications. At the same time, the basic information of the file is recorded, including the file name, upload time, and file source.

[0079] Then, a standard-compliant part is searched according to the pre-defined standard rules, which exist in the metadata area of ​​the standard library in the form of text descriptions, templates and keyword sets.

[0080] Then, using text matching and pattern recognition technology, paragraphs, clauses, and data items that conform to the standard rules are selected from the imported bidding documents and marked as conforming to the standard. At the same time, a search report is generated, recording the specific location and content summary of the conformity.

[0081] Next, extract the non-compliant parts, compare them with the retrieved compliant content, and use the difference comparison algorithm to find the parts in the bidding documents that are different from them. The different parts include text description, numerical range and format layout.

[0082] The different parts, i.e. the differences, are extracted separately and organized into independent datasets;

[0083] The extracted differences are input into a pre-trained deep learning model. The deep learning model analyzes the differences from the perspectives of semantic understanding, logical structure, and industry conventions to determine whether they conform to industry norms and outputs the analysis results, which include those that conform to and those that do not.

[0084] For discrepancies in the analysis results that do not conform to the specifications, the system automatically extracts the discrepancy features.

[0085] The differences are recorded in a structured form and stored in the violation feature library of the preset standard library;

[0086] When a deep learning model is unable to determine the content score, the content is manually judged. If the manual judgment finds that the content does not comply with the standards, the features of that part will be extracted and imported into the violation feature library of the preset standard library.

[0087] The beneficial effects of this invention are reflected in:

[0088] It can process different types of project information, such as handwritten paper, printed paper, and electronic documents, to meet the diverse information source needs in smart construction, ensuring that project information in various forms can enter the analysis process, and improving the versatility and applicability of the method.

[0089] For handwritten paper project documents, a combination of deep learning-based optical character recognition, semantic understanding-based secondary recognition and verification, and manual proofreading is used to effectively overcome the difficulties of handwritten text recognition and greatly improve the accuracy of recognition. For printed paper project documents, layered optical character recognition and two-way format verification ensure the quality of recognition results. Electronic document versions of project documents are directly converted into standard formats, which also facilitates subsequent unified processing.

[0090] Before recognizing handwritten project documents, the flatness detection of paper can be performed. By collecting image information from multiple angles and performing complex calculations and evaluations, paper unevenness problems can be detected in advance, avoiding the impact of paper deformation on recognition results and laying the foundation for accurate information extraction and analysis in the future.

[0091] By accurately identifying text regions and calculating coordinates, it can determine whether the filled content is centered, promptly identify problems with non-standard filling, and generate prompts, which helps improve the quality of project information entry and ensures the standardization and consistency of data.

[0092] By extracting multiple key indicators from three dimensions—project progress, quality control, and cost budget—a comprehensive evaluation of the project can fully and deeply reflect the actual situation of smart construction projects, providing project managers with rich and valuable decision-making basis.

[0093] Quantitative assessment improves scientific rigor: By pre-establishing a numerical mapping set of scores for each dimension of indicators and assigning different weights to them and calculating the total score, the project assessment indicators are quantified, making the assessment results more objective and scientific, reducing the interference of subjective factors, and improving the accuracy and credibility of the assessment.

[0094] A keyword extraction strategy that combines multiple methods can accurately and comprehensively extract representative keywords from project information text. These keywords can highly summarize the core content of the project, providing a reliable foundation for plagiarism detection.

[0095] By establishing a project information database and indexing keywords, the keywords of the project information to be checked for duplication can be matched with the existing information in the database. This allows for a quick and accurate determination of whether the project information is duplicated, which helps to avoid redundant construction, improve resource utilization efficiency, and ensure the uniqueness and innovation of smart construction projects. Attached Figure Description

[0096] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0097] Figure 1 This is an overall flowchart of the present invention;

[0098] Figure 2 This is a schematic diagram of the flatness detection image acquisition of the present invention. Detailed Implementation

[0099] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of the present invention and are therefore intended to limit the scope of protection of the present invention.

[0100] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0101] like Figures 1-2 As shown, a method for project information analysis and plagiarism detection includes the following steps:

[0102] Step 1: Import project information, which includes bidding documents and other project documents. Analyze the project information using different processing methods based on its type and convert it into a standard format.

[0103] Step 2: After converting to standard format, the form is checked. If the check passes, semantic recognition is performed to obtain the analysis results of the filled content. If the check fails, a prompt message is generated, prompting the user to check the original project information.

[0104] Step 3: If there are no abnormalities in the analysis results, proceed with the project information analysis to obtain the project analysis results;

[0105] Step 4: Next, analyze the project information, extract keywords, and then check for duplicates in the project information.

[0106] Step 5: After the project information is checked for plagiarism, the plagiarism information is generated and sent to the preset receiving terminal.

[0107] Furthermore, the specific process in step one is as follows:

[0108] Other project document types include handwritten paper project documents, printed paper project documents, and electronic project documents;

[0109] When other project files are paper-based handwritten project files:

[0110] First, a deep learning-based optical character recognition method is used to perform preliminary recognition, converting handwritten text into preliminary recognition results;

[0111] When the initial recognition result shows suspected errors or unclear characters, a second recognition verification is performed using an optical character recognition method based on semantic understanding. The system automatically captures the text content before and after the character, combines industry domain knowledge graphs and semantic analysis algorithms to make corrections, and marks the corrected content to obtain the second recognition result. If there are no suspected errors or unclear characters, the system proceeds directly to the next step.

[0112] Finally, the recognition results are manually verified through the manual verification module. Once the manual verification is passed, the final recognition results are exported.

[0113] When other project files are paper-printed project files:

[0114] After scanning paper documents and converting them into digital images, a layered optical character recognition (OCR) operation is performed.

[0115] The first layer uses a high-precision general optical character recognition engine to quickly extract text content and identify the main text and basic format information;

[0116] The second layer enables an intelligent optical character recognition module that focuses on complex typesetting analysis, employing differentiated recognition strategies for different elements such as titles, paragraphs, tables, and charts.

[0117] When identifying headings, an industry knowledge graph and a heading hierarchy model are used to determine their importance level within the overall project document architecture. For example, first-level headings typically cover the core project theme, while second-level headings detail key sections under the core theme. For paragraphs, in addition to analyzing hierarchy based on paragraph indentation and line spacing, text classification algorithms are used to categorize paragraph content into different types such as background information, technical solutions, and implementation plans. For tables, a semantic understanding model is constructed to ensure that row and column data are accurately transferred to the standard format table filling area and to interpret the logical relationships behind the table data. For example, in a project resource allocation table, the model analyzes whether the proportion of resources allocated to different departments is reasonable and in line with the overall project plan.

[0118] Finally, a two-way format check is performed. On the one hand, from a visual presentation perspective, the layout of each element in the digital image is checked to see if it conforms to industry standard format specifications, such as whether table lines are neat and whether chart axis labels are clear.

[0119] On the other hand, from the perspective of text content logic, the consistency of information between different elements is compared, such as whether the scope of the project mentioned in the title matches the scope reflected in the paragraph description and table data. Two-way verification ensures the accuracy and completeness of the information. After successful comparison, the precisely parsed project information is accurately filled into the corresponding fill-in boxes according to standard format requirements.

[0120] Finally, export the recognition results that have passed the two-way format verification.

[0121] When other project files are in electronic document format, convert them directly to the standard format.

[0122] When the project information is a bidding document, the bidding document includes paper bidding documents and electronic bidding documents. For paper bidding documents, they are processed in the same way as paper-printed project documents. For electronic bidding documents, they are directly converted into the standard format.

[0123] When the project information is other project files, the types of other project files include paper handwritten project files and paper printed project files. Before recognizing paper handwritten project files, a flatness test is required. Once the flatness test is passed, the next step of recognition is allowed.

[0124] The specific process for conducting a flatness test is as follows:

[0125] First, image acquisition is performed, that is, using an image acquisition device, image information of paper handwritten project documents is acquired at a 45-degree angle downward from the left, a 45-degree angle downward from the right, a 45-degree angle downward from the top, and a 45-degree angle downward from the bottom, to obtain the first image information, the second image information, the third image information, and the fourth image information.

[0126] Then, the center point T1 of the left edge, the center point T2 of the right edge, the center point T3 of the top edge, and the center point T4 of the bottom edge are extracted from the first image information.

[0127] Then, using the plane where the paper-based handwritten project documents are placed as the reference plane, measure the distances q1 between T1 and the reference plane, m1 between T2 and the reference plane, f1 between T3 and the reference plane, and e1 between T4 and the reference plane.

[0128] Set q1, m1, f1 and e1 together to obtain the first evaluation parameter A1(q1, m1, f1, e1);

[0129] Then, the second image information, the third image information, and the fourth image information are processed in the same way as the first image information to obtain the second evaluation parameter A2 (q2, m2, f2, e2), the third evaluation parameter A3 (q3, m3, f3, e3), and the fourth evaluation parameter A4 (q4, m4, f4, e4).

[0130] Then, the difference Aa1 between the first evaluation parameter and the second evaluation parameter, and the difference Aa2 between the third evaluation parameter and the fourth evaluation parameter are calculated.

[0131] The difference between the first evaluation parameter and the third evaluation parameter is Aa3, the difference between the first evaluation parameter and the fourth evaluation parameter is Aa4, the difference between the second evaluation parameter and the fourth evaluation parameter is Aa5, and the difference between the second evaluation parameter and the third evaluation parameter is Aa6.

[0132] If at least four of Aa1, Aa2, Aa3, Aa4, Aa5, and Aa6 exceed the preset range, the flatness test fails.

[0133] Furthermore, the specific process of manual inspection through the manual proofreading module is as follows:

[0134] When a second verification result exists, at least two people must be arranged to confirm the suspected error or unclear characters at least twice. If the two confirmation results are the same, the export is allowed.

[0135] If the two confirmation results are different, a re-verification is required.

[0136] Furthermore, the specific process for completing the inspection is as follows:

[0137] First, perform text region recognition, collect image information after filling in the text, locate the position of the text box in the image, and obtain its upper left corner coordinates (x1, y1) and lower right corner coordinates (x2, y2);

[0138] Then, the range of the text within the input box is identified, and the coordinates of the top left corner (x3, y3) and bottom right corner (x4, y4) of the text area are obtained.

[0139] Next, calculate the center coordinates of the text region, x-coordinate. 中 = (x3 + x4) / 2, y-coordinate 中 = (y3+y4) / 2;

[0140] Next, calculate the center coordinates of the fill box, x 框中 = (x1+x2) / 2, y-coordinate 框中 = (y1+y2) / 2;

[0141] Calculate the deviation value of the abscissa Δx=|x 中 -x 框中 | and the deviation value of the ordinate Δy=|y 中 -y 框中 |;

[0142] When both the horizontal axis deviation value Δx and the vertical axis deviation value Δy are less than the set threshold, the text is considered to be centered, indicating that the filling check has passed.

[0143] Project analysis results include projects with anomalies and projects without anomalies;

[0144] The process of obtaining the project analysis results is as follows:

[0145] Extract the on-time completion rate of key nodes, average daily construction volume, and progress coordination of each subsystem from the project information;

[0146] The on-time completion rate of key nodes, average daily construction volume and the coordination of progress of each subsystem are processed to obtain project progress indicators;

[0147] Then, extract the percentage of batches that failed material quality inspection, the rework rate, and the average failure interval of the intelligent system from the project information;

[0148] The percentage of batches failing material quality inspection, rework rate, and average failure interval of intelligent system are processed to obtain quality control indicators;

[0149] Then extract the cost-benefit ratio, budget change frequency and magnitude, and cost deviation rate from the project information;

[0150] The cost-benefit ratio, budget change frequency and magnitude, and cost deviation rate are processed to obtain cost budget indicators;

[0151] The project schedule indicators, quality control indicators, and cost budget indicators are processed to obtain a comprehensive evaluation indicator. When the comprehensive evaluation indicator is less than the preset value, it indicates that the project is abnormal; otherwise, it indicates that the project is normal.

[0152] The process of obtaining project progress indicators is as follows: First, a numerical mapping set of project progress scores is established in advance, which includes the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem. Then, the specific scores of the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem are obtained from the numerical mapping set of project progress scores. Different weights are assigned to the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem. Finally, the sum of the scores after assigning weights is calculated, which is the project progress indicator.

[0153] The process for obtaining quality control indicators is as follows:

[0154] A quality control score mapping set was pre-established, which included the percentage of batches failing material quality inspection, rework rate, and average fault interval of intelligent system. After obtaining the specific scores of the percentage of batches failing material quality inspection, rework rate, and average fault interval of intelligent system from the control score mapping set, different weights were assigned to the percentage of batches failing building material quality inspection, rework rate, and average fault interval of intelligent system. Then, the sum of the weighted scores was calculated, which yielded the quality control indicators.

[0155] The process of obtaining cost budget targets is as follows:

[0156] A cost budget scoring numerical mapping set of cost-benefit ratio, budget change frequency and magnitude and cost deviation rate was pre-established. After obtaining the specific scores of cost-benefit ratio, budget change frequency and magnitude and cost deviation rate from the cost budget scoring numerical mapping set, different weights were assigned to cost-benefit ratio, budget change frequency and magnitude and cost deviation rate. Then, the sum of the scores after assigning weights was calculated, which is the cost budget indicator.

[0157] The process for obtaining the comprehensive evaluation indicators is as follows:

[0158] Mark the project schedule indicator as G1, the quality control indicator as G2, and the cost budget indicator as G3;

[0159] Assign weight W1 to G1, weight W2 to G2, and weight W3 to G3, such that W1 + W2 + W3 = 1, and W3 > W2 = W1;

[0160] The comprehensive evaluation index Gg can be obtained by using the formula G1*W1+G2*W2+G3*W3=Gg.

[0161] Project progress indicators and scoring:

[0162] On-time completion rate of key milestones: This is calculated by dividing the number of milestones completed on time by the total number of key milestones, based on whether important milestones such as the completion of basic construction, the topping out of the main structure, and the initial debugging of intelligent systems are achieved on time.

[0163] Average daily construction volume: This is calculated by dividing the amount of work completed in each construction phase (such as building area, number of installed equipment, etc.) by the actual number of construction days, reflecting construction efficiency.

[0164] Schedule coordination of each subsystem: This measures the degree of matching between the schedules of subsystems such as main building construction, electromechanical installation, and intelligent system integration. The smaller the difference, the higher the coordination. It can be quantified by calculating the standard deviation of the schedule difference between each subsystem.

[0165] Scoring process:

[0166] On-time completion rate of key milestones: 20 points for completion rate of 95% or above; 15-20 points for completion rate of 85%-94%; 10-15 points for completion rate of 75%-84%; and 0-10 points for completion rate of less than 75%.

[0167] Average daily construction volume: Compared with the historical average daily construction volume of similar projects, 15 points are awarded if the volume reaches or exceeds the average by 10%; 10-15 points are awarded if the volume is within ±10% of the average; and 0-10 points are awarded if the volume is below -10% of the average.

[0168] Schedule coordination of each subsystem: 15 points for standard deviation less than or equal to 5%; 10-15 points for standard deviation between 5% and 10%; 0-10 points for standard deviation greater than 10%.

[0169] Quality control indicators and scores:

[0170] Percentage of batches failing material quality inspection: This refers to the percentage of batches of various building materials and intelligent components that fail inspection out of the total number of batches inspected.

[0171] Rework rate: Calculates the percentage of work that needs to be reworked due to quality issues, including rework caused by substandard construction techniques, equipment malfunctions, etc.

[0172] Intelligent system operation stability index - Mean Time Between Failures (MTBF): Records the average interval between two consecutive failures of an intelligent system within a certain operating period. The longer the interval, the higher the stability.

[0173] Scoring process:

[0174] Percentage of batches failing material quality inspection: 15 points for a percentage of 0%; deduct 1 point for every 1% increase, until all points are deducted.

[0175] Rework rate: 15 points for a rework rate below 3%; 10-15 points for a rework rate between 3% and 5%; 0-10 points for a rework rate above 5%.

[0176] Mean Time Between Failures (MTBF): 20 points are awarded if the MTBF value reaches or exceeds 1.2 times the average MTBF value of similar systems; 10-19 points are awarded if the value is between 1 and 1.2 times the average; and 0-9 points are awarded if the value is below the average.

[0177] Cost budget indicators and scoring:

[0178] Cost-benefit ratio: Calculated by dividing the expected revenue of a project by the total cost, it measures the return on investment.

[0179] Budget change frequency and magnitude: Statistics on the number of budget changes and the percentage of the change amount relative to the original budget reflect the accuracy of budget planning and the level of control.

[0180] Cost deviation rate: Calculate the ratio of the difference between the actual cost and the budget cost to the budget cost. A positive value indicates cost overruns, and a negative value indicates cost savings.

[0181] Scoring process:

[0182] Cost-benefit ratio: If it is 10% higher than the industry average, 20 points are given; for every 5% lower, 2 points are deducted.

[0183] Budget change frequency and amplitude: If the frequency is less than 2 times per year and the amplitude is less than 5%, 15 points are given; if the frequency is 2 - 3 times or the amplitude is 5% - 10%, a score within 10 - 15 is given; if the frequency exceeds 3 times or the amplitude exceeds 10%, a score within 0 - 10 is given.

[0184] Cost deviation rate: If the deviation rate is within ±3%, 15 points are given; if the deviation rate is between -3% and -10% or 3% and 10%, a score within 10 - 15 is given; if the deviation rate exceeds the above range, a score within 0 - 10 is given.

[0185] The duplicate checking process in Step 4 is as follows:

[0186] First, perform keyword extraction: Clean the project information text, remove the noise data in it, and obtain the preprocessed text.

[0187] Such as HTML tags, special characters, stop words (such as words like "de", "le", "shi" that have no actual meaning in the text), etc., and convert the text into a unified format, such as lowercase letter form, for subsequent processing.

[0188] After that, use the term frequency - inverse document frequency (TF-IDF) algorithm to calculate the importance score of each word in the project information text.

[0189] Term frequency (TF) represents the frequency of a word in the text, and inverse document frequency (IDF) measures the rarity of a word in the entire document collection. The higher the TF-IDF value, the higher the importance of the word in the text and the more likely it is to be a keyword. For example, in multiple project information about electronic products, the word "smartphone" may appear frequently in one article (high TF value), but rarely in most other documents (high IDF value), so its TF-IDF value will be relatively high and is very likely to be extracted as a keyword.

[0190] After that, use a rule-based method combined with the domain knowledge and business rules of the project information to formulate keyword extraction rules.

[0191] Next, part-of-speech tagging and filtering are performed on the preprocessed text to identify words of different parts of speech, such as nouns, verbs, and adjectives. Generally, nouns better represent the core content of the project information; therefore, nouns should be prioritized when extracting keywords. For example, in the sentence "This phone has a high-definition camera function," nouns like "phone" and "camera function" better reflect the key content of the project information and can be used as key keywords for extraction.

[0192] Keyword filtering and sorting: After obtaining candidate keywords through the above process, the keywords are filtered according to the set threshold (such as TF-IDF score threshold) to remove words with low scores;

[0193] At the same time, the selected keywords are sorted according to their importance scores, and the top-ranking keywords are selected as the final project information keywords.

[0194] Perform duplicate project information checks and establish a project information database: Store existing project information in the database, with each piece of project information having a unique identifier;

[0195] The database structure should include the basic content of the project information, extracted keywords, and other relevant attributes.

[0196] Create an index for the keywords extracted from each item in the database;

[0197] To facilitate quick searching and matching, an inverted index structure can be used, with keywords as index items, and each keyword corresponding to a list of items containing that keyword.

[0198] For the project information to be checked for plagiarism, keywords are extracted according to the set keyword extraction rules, and the keywords of the project information to be checked for plagiarism are matched with the keywords of the existing project information in the database;

[0199] Various matching strategies can be employed, such as exact matching and fuzzy matching. Exact matching refers to searching the database for items with keywords that are exactly the same as the item to be checked for plagiarism; fuzzy matching allows for a certain degree of difference, for example, by calculating the similarity between keywords (such as edit distance, cosine similarity, etc.) to determine whether similar item information exists.

[0200] The system determines whether there is duplicate project information based on the matching results. If a project is found in the database that highly matches the keywords of the project to be checked (exceeding the set similarity threshold), it is considered to be duplicated; otherwise, the project is considered not to be duplicated.

[0201] After all keywords have been matched, plagiarism detection information is generated.

[0202] In step four, when the imported project information is a bidding document, the specific plagiarism check process is as follows:

[0203] Import the bidding documents into the preset standard format library and read their content:

[0204] Then, the read content is stored in the standard library line by line, segment by segment, or according to specific fields, according to the preset standard library data entry specifications. At the same time, the basic information of the file is recorded, including the file name, upload time, and file source.

[0205] Then, a standard-compliant part is searched according to the pre-defined standard rules, which exist in the metadata area of ​​the standard library in the form of text descriptions, templates and keyword sets.

[0206] For example, it may stipulate that certain chapters must contain certain key terms, or specify formatting requirements (such as row and column specifications for tables).

[0207] Then, using text matching and pattern recognition technology, paragraphs, clauses, and data items that conform to the standard rules are selected from the imported bidding documents and marked as conforming to the standard. At the same time, a search report is generated, recording the specific location and content summary of the conformity.

[0208] Next, extract the non-compliant parts, compare them with the retrieved compliant content, and use the difference comparison algorithm to find the parts in the bidding documents that are different from them. The different parts include text description, numerical range and format layout.

[0209] The different parts, i.e. the differences, are extracted separately and organized into independent datasets;

[0210] The original file's related index is attached to facilitate tracing its origin and prepare for subsequent analysis.

[0211] The extracted differences are then input into a pre-trained deep learning model.

[0212] The architecture of deep learning models includes those built on convolutional neural networks (CNN), recurrent neural networks (RNN) and their variants, and are trained based on a large amount of historical bidding documents and manually labeled standardized data.

[0213] The deep learning model analyzes the differences in semantic understanding, logical structure and industry practices to determine whether they conform to the industry's common standards and outputs the analysis results, which include those that conform to the standards and those that do not.

[0214] For discrepancies in the analysis results that do not conform to the specifications, the system automatically extracts the discrepancy features.

[0215] Such as incorrect terminology, missing required information, and improper numerical settings.

[0216] The differences are recorded in a structured form and stored in the violation feature library of the preset standard library;

[0217] At the same time, it is linked with the corresponding bidding documents and the differences to facilitate subsequent statistical analysis and rule optimization.

[0218] When a deep learning model is unable to determine the content score, the content is manually judged. If the manual judgment finds that the content does not comply with the standards, the features of that part will be extracted and imported into the violation feature library of the preset standard library.

[0219] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for project information analysis and plagiarism detection, characterized in that: Includes the following steps: Step 1: Import project information, which includes bidding documents and other project documents. Analyze the project information using different processing methods based on its type and convert it into a standard format. Step 2: After converting to standard format, the form is checked. If the check passes, semantic recognition is performed to obtain the analysis results of the filled content. If the check fails, a prompt message is generated, prompting the user to check the original project information. Step 3: If there are no abnormalities in the analysis results, proceed with the project information analysis to obtain the project analysis results; Step 4: Next, analyze the project information, extract keywords, and then check for duplicates in the project information. Step 5: After the project information is checked for duplicates, the duplicate information is generated and sent to the preset receiving terminal; The specific process in step one is as follows: Other project document types include handwritten paper project documents, printed paper project documents, and electronic project documents; When other project files are paper-based handwritten project files: First, a deep learning-based optical character recognition method is used to perform preliminary recognition, converting handwritten text into preliminary recognition results; When the initial recognition result shows suspected errors or unclear characters, a second recognition verification is performed using an optical character recognition method based on semantic understanding. The system automatically captures the text content before and after the character, combines industry domain knowledge graphs and semantic analysis algorithms to make corrections, and marks the corrected content to obtain the second recognition result. If there are no suspected errors or unclear characters, the system proceeds directly to the next step. Finally, the recognition results are manually checked through the manual verification module. Once the manual verification is passed, the final recognition results are exported. When other project files are paper-printed project files: After scanning paper documents and converting them into digital images, a layered optical character recognition (OCR) operation is performed. The first layer uses a general optical character recognition engine to extract text content and identify the main text and basic format information. The second layer enables the intelligent optical character recognition module, which uses differentiated recognition strategies for different elements such as titles, paragraphs, tables, and charts. Finally, a two-way format check is performed. From a visual presentation perspective, the layout of each element in the digital image is checked to see if it conforms to industry standard format specifications. If it does not, it will fail. From the perspective of text content logic, compare whether the information between different elements is consistent, and finally export the recognition results that pass the two-way format verification. When other project files are in electronic document format, convert them directly to the standard format. When the project information is a bidding document, the bidding document includes paper bidding documents and electronic bidding documents. For paper bidding documents, they are processed in the same way as paper-printed project documents. For electronic bidding documents, they are directly converted into standard format. The specific process for filling out the checklist is as follows: First, perform text region recognition, collect image information after filling in the text, locate the position of the text box in the image, and obtain its upper left corner coordinates (x1, y1) and lower right corner coordinates (x2, y2). Then, the range of the text within the input box is identified, and the coordinates of the top left corner (x3, y3) and bottom right corner (x4, y4) of the text area are obtained. Next, calculate the center coordinates of the text region, x-coordinate. 中 = (x3 + x4) / 2, y-coordinate 中 = (y3+y4) / 2; Next, calculate the center coordinates of the fill box, x 框中 = (x1 + x2) / 2, y-coordinate 框中 = (y1+y2) / 2; Calculate the x-axis deviation value ∆x = |x 中 -x 框中 | and the deviation value of the ordinate ∆y=|y 中 -y 框中 |; When both the horizontal axis deviation value ∆x and the vertical axis deviation value ∆y are less than the set threshold, the text is considered to be centered, which means that the filling check has passed.

2. The project information analysis and deduplication method according to claim 1, characterized in that: When the project information is other project files, the types of other project files include paper handwritten project files and paper printed project files. Before recognizing paper handwritten project files, a flatness test is required. Once the flatness test is passed, the next step of recognition is allowed.

3. The project information analysis and deduplication method according to claim 2, characterized in that: The specific process of manual inspection using the manual verification module is as follows: When a second verification result exists, at least two people must be arranged to confirm the suspected error or unclear characters at least twice. If the two confirmation results are the same, the export is allowed. If the two confirmation results are different, a re-verification is required.

4. The project information analysis and deduplication method according to claim 3, characterized in that: Project analysis results include projects with anomalies and projects without anomalies; The process of obtaining the project analysis results is as follows: Extract the on-time completion rate of key nodes, average daily construction volume, and progress coordination of each subsystem from the project information; The on-time completion rate of key nodes, average daily construction volume and the coordination of progress of each subsystem are processed to obtain project progress indicators; Then, extract the percentage of batches that failed material quality inspection, the rework rate, and the average failure interval of the intelligent system from the project information; The percentage of batches failing material quality inspection, rework rate, and average failure interval of intelligent system are processed to obtain quality control indicators; Then extract the cost-benefit ratio, budget change frequency and magnitude, and cost deviation rate from the project information; The cost-benefit ratio, budget change frequency and magnitude, and cost deviation rate are processed to obtain cost budget indicators; The project schedule indicators, quality control indicators, and cost budget indicators are processed to obtain a comprehensive evaluation indicator. When the comprehensive evaluation indicator is less than the preset value, it indicates that the project is abnormal; otherwise, it indicates that the project is normal.

5. The project information analysis and deduplication method according to claim 4, characterized in that: The process of obtaining project progress indicators is as follows: First, a numerical mapping set of project progress scores is established in advance, which includes the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem. Then, the specific scores of the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem are obtained from the numerical mapping set of project progress scores. Different weights are assigned to the on-time completion rate of key nodes, the average daily construction volume, and the progress coordination of each subsystem. Finally, the sum of the scores after assigning weights is calculated, which is the project progress indicator. The process for obtaining quality control indicators is as follows: A quality control score mapping set was pre-established, which included the percentage of batches failing material quality inspection, rework rate, and average fault interval of intelligent system. After obtaining the specific scores of the percentage of batches failing material quality inspection, rework rate, and average fault interval of intelligent system from the control score mapping set, different weights were assigned to the percentage of batches failing building material quality inspection, rework rate, and average fault interval of intelligent system. Then, the sum of the weighted scores was calculated, which yielded the quality control indicators. The process of obtaining cost budget targets is as follows: A cost budget scoring numerical mapping set of cost-benefit ratio, budget change frequency and magnitude and cost deviation rate was pre-established. After obtaining the specific scores of cost-benefit ratio, budget change frequency and magnitude and cost deviation rate from the cost budget scoring numerical mapping set, different weights were assigned to cost-benefit ratio, budget change frequency and magnitude and cost deviation rate. Then, the sum of the scores after assigning weights was calculated, which is the cost budget indicator. The process for obtaining the comprehensive evaluation indicators is as follows: Mark the project schedule indicator as G1, the quality control indicator as G2, and the cost budget indicator as G3; Assign weight W1 to G1, weight W2 to G2, and weight W3 to G3. W1 + W2 + W3 = 1, and W3 > W2 = W1. The comprehensive evaluation index Gg can be obtained by using the formula G1×W1+G2×W2+G3×W3=Gg.

6. The project information analysis and deduplication method according to claim 1, characterized in that: The plagiarism detection process in step four is as follows: First, keyword extraction is performed: the project information text is cleaned to remove noisy data and obtain the preprocessed text; Then, the importance score of each word in the project information text was calculated using a statistical TF-IDF method; Then, a rule-based approach was used to combine domain knowledge and business rules related to the project information to formulate keyword extraction rules; Next, part-of-speech tagging and filtering are performed on the preprocessed text to identify words with different parts of speech and obtain candidate keywords; Keyword filtering and sorting: After obtaining candidate keywords according to the above process, the candidate keywords are filtered according to the set threshold to remove words with scores lower than the preset value, and the filtered keywords are obtained. At the same time, the selected keywords are sorted according to their importance scores, and the top-ranking keywords are selected as the final project information keywords. Perform duplicate project information checks and establish a project information database: Store existing project information in the database, with each piece of project information having a unique identifier; Create an index for the keywords extracted from each item in the database; For the project information to be checked for plagiarism, keywords are extracted according to the set keyword extraction rules, and the keywords of the project information to be checked for plagiarism are matched with the keywords of the existing project information in the database; The system determines whether there is duplicate project information based on the matching results. If a project with a keyword similarity greater than a preset value is found in the database, it is considered to be duplicated; otherwise, the project information is considered to be unique.

7. The project information analysis and deduplication method according to claim 1, characterized in that: In step four, when the imported project information is a bidding document, the specific plagiarism check process is as follows: Import the bidding documents into the preset standard format library and read their content: Then, the read content is stored in the standard library line by line, segment by segment, or according to specific fields, according to the preset standard library data entry specifications. At the same time, the basic information of the file is recorded, including the file name, upload time, and file source. Next, a standard-compliant section is searched. Based on pre-set standard rules, text matching and pattern recognition technologies are used to filter out paragraphs, clauses, and data items that conform to the standard rules from the imported bidding documents and mark them as conforming to the standard. At the same time, a search report is generated, recording the specific location and content summary of the conformity. Next, extract the non-compliant parts, compare them with the retrieved compliant content, and use a difference comparison algorithm to find the parts in the bidding documents that are different from them. The different parts, i.e. the differences, are extracted separately and organized into independent datasets; The extracted differences are input into a pre-trained deep learning model. The deep learning model analyzes the differences from the perspectives of semantic understanding, logical structure, and industry conventions to determine whether they conform to industry norms and outputs the analysis results, which include those that conform to and those that do not. For discrepancies in the analysis results that do not conform to the specifications, the system automatically extracts the discrepancy features. The differences are recorded in a structured form and stored in the violation feature library of the preset standard library; When there is content that the deep learning model cannot determine, the undetermined content is manually determined. When the manual determination finds that it does not conform to the standard, some features are extracted and imported into the violation feature library of the preset standard library.

Citation Information

Patent Citations

  • Picture recognition method and device, equipment and storage medium

    CN116311296A

  • Domain-oriented science and technology project duplicate checking method and system

    CN116431763A

  • File retrieval method and system based on OCR (Optical Character Recognition) technology

    CN117390214A