A method and system for automatic classification of construction labor contracts
By establishing a labor contract framework and using OCR technology and a pre-trained text classification model, labor contracts are automatically identified and classified, solving the problem of inconsistent agreements in different regions. This achieves unified contract formatting and efficient and accurate information extraction, ensuring the legality and compliance of contracts.
Patent Information
- Application Number
- CN202510206154.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Inconsistent labor agreements across regions make it difficult for the management system to automatically match relevant fields, resulting in cumbersome data analysis and statistics, and low efficiency and accuracy in information retrieval.
By establishing a labor contract framework for the corresponding region, using OCR technology and a pre-trained text classification model, the contract content is automatically identified and classified, integrated into a unified format, and abnormal contracts are manually reviewed to optimize the model.
It has achieved a unified format for labor contracts in different regions and the extraction of core information, which has improved the efficiency and accuracy of information retrieval and ensured the legality and compliance of contract content.
Smart Images

Figure CN120198061B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of contract classification and integration, and in particular to a method and system for automatically classifying construction site labor contracts. Background Technology
[0002] According to local government policies, the files of contract workers are uploaded to the corresponding internal and external monitoring systems. Staff are responsible for handling the scanning, detailed classification, merging, photographing, format conversion, entering information into the system, uploading merged files, and the final review process.
[0003] Inconsistent labor agreements across regions necessitate additional time for managers to understand and compare contract details from different areas when accessing or verifying contracts. Furthermore, the system integration process struggles to automatically match relevant fields, making data analysis and statistics cumbersome.
[0004] Therefore, there is a need for a method and system for automatically classifying construction site labor contracts to improve the efficiency and accuracy of information retrieval. Summary of the Invention
[0005] In view of this, it is necessary to provide a method and system for automatically classifying construction site labor contracts that improves the efficiency and accuracy of information retrieval, in order to solve the above problems.
[0006] The embodiments of this application provide a method for automatically classifying construction site labor contracts, retrieving labor agreements for the corresponding regions, and establishing construction site labor contracts for the corresponding regions;
[0007] Enter the relevant content of the construction site labor contract, and the document of the construction site labor contract for the corresponding region will be generated after the input is completed;
[0008] Scan documents and convert them into editable documents;
[0009] Differentiate editable documents and integrate them into documents of the same format. The differentiation of editable documents includes: contract type, project duration type, and contract period.
[0010] In at least one embodiment of this application, the method for automatically classifying construction site labor contracts further includes the step of:
[0011] Store editable documents, including either metadata storage or file storage.
[0012] Editable documents are transmitted to the labor supervision system and the labor management system, respectively.
[0013] In at least one embodiment of this application, the method of "inputting the relevant content of the construction site labor contract and generating a document of the construction site labor contract after input" includes the following steps:
[0014] Fill out the construction site labor contract and verify it;
[0015] Enter the information required for the construction site labor contract; input methods include: mobile phone input, computer input, and scanning of physical documents.
[0016] In at least one embodiment of this application, the method of "scanning a document and converting it into an editable document" further includes:
[0017] Using OCR light symbol recognition, extract content from entity documents, including: job content, work location, work hours, and remuneration;
[0018] Detect if the page is tilted; if so, introduce a page correction algorithm.
[0019] Use image preprocessing techniques to improve the quality of scanned images.
[0020] In at least one embodiment of this application, the method of "scanning a document and converting it into an editable document" further includes:
[0021] Generate keywords and text content related to the labor contract;
[0022] Use keywords and text content from labor law;
[0023] Use synonyms to integrate and match identical words.
[0024] In at least one embodiment of this application, the method of "distinguishing editable documents to form distinguishable content including: contract type, work period type, and contract term" includes:
[0025] A pre-trained text classification model is used for contract classification. Contract classification includes: text content, keyword distribution, and semantic feature vector.
[0026] Label the contract data and generate a model training set.
[0027] In at least one embodiment of this application, the method of "distinguishing editable documents to form distinguishable content including: contract type, project duration type, and contract period" further includes:
[0028] If a contract contains unmatched keywords or abnormal content, it will be marked for priority manual review.
[0029] Manually process contracts with abnormal classifications and confirm the modification of classification content;
[0030] Add the manually labeled results to the model training set.
[0031] In at least one embodiment of this application, the method of "classifying contracts using a pre-trained text classification model" includes:
[0032] Extracted keywords or text content that have already been entered;
[0033] Output the corresponding contract type label.
[0034] In at least one embodiment of this application, the method of "scanning a document and converting it into an editable document" further includes:
[0035] After integration, the original multi-regional labor agreements were re-established as labor contracts with the same format;
[0036] Keep the original document.
[0037] A system for automatically classifying construction site labor contracts, characterized in that the system includes any of the methods for automatically classifying construction site labor contracts described above.
[0038] The beneficial effects of the above-mentioned method and system for automatically classifying construction site labor contracts are as follows:
[0039] 1. Collect and integrate labor agreements from different regions, unify the format, and form standardized construction site labor contract documents.
[0040] 2. Extract core information from the contract (such as job description, location, and remuneration), and integrate similar expressions through synonym matching. Incorporate labor law-related keywords to ensure the comprehensiveness and accuracy of the extracted content.
[0041] 3. Contracts with unmatched keywords or abnormal content will be marked and prioritized for manual review. The review results will be used to optimize the classification model. Attached Figure Description
[0042] Figure 1 This is a flowchart of a method for automatically classifying construction site labor contracts. Detailed Implementation
[0043] The embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0044] It should be noted that when a component is considered to be "connected" to another component, it can be directly connected to the other component or may also have an intervening component. When a component is considered to be "placed" on another component, it can be directly placed on the other component or may also have an intervening component. The terms "top," "bottom," "upper," "lower," "left," "right," "front," "back," and similar expressions used in this article are for illustrative purposes only.
[0045] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0046] Please see Figure 1 This application provides a method for automatically classifying construction site labor contracts, the method comprising the following steps:
[0047] S1: Retrieve the labor agreements for the relevant regions and establish the corresponding construction site labor contracts.
[0048] S2: Enter the relevant content of the construction site labor contract. After the input is completed, a document of the construction site labor contract for the corresponding region will be generated.
[0049] S3: Scan the document and convert it into an editable document.
[0050] S4: Differentiate editable documents and integrate them into documents of the same format. The differentiation of editable documents includes: contract type, project duration type, and contract period.
[0051] Specifically, in step S1, firstly, based on the laws, regulations, and policies of the target construction site's location, labor agreement templates for that region are retrieved from relevant databases or official channels. These templates typically contain contract terms, format requirements, and legal constraints specific to that region. Based on these templates, a corresponding labor contract framework is established for the specific construction site. This framework will serve as the basis for subsequently inputting contract content, ensuring the legality and compliance of the contract.
[0052] Methods for obtaining a labor agreement template include:
[0053] The system can automatically retrieve labor agreement templates for specific regions through API interfaces or database queries.
[0054] Based on automated retrieval, necessary adjustments and supplements are made by professionals to adapt to the actual needs of specific construction sites.
[0055] In step S2, within the established labor contract framework, specific contract details are entered based on the actual conditions of the construction site and the results of negotiations between both parties. These details include, but are not limited to, information about both parties, job duties, work location, working hours, wages and benefits, and liability for breach of contract. After the data is entered, the system automatically generates a complete and standardized construction site labor contract document.
[0056] The methods for generating construction site labor contract documents include: First, managers can input the contract content using a keyboard or handwriting tablet. Second, if some contract data already exists, it can be imported into the system via file upload or database import. Third, using an established contract framework, contract content can be quickly generated by selecting or filling in predefined options and fields. Fourth, the document can be scanned and converted into an editable document.
[0057] For paper contracts or existing non-editable electronic documents (such as PDFs, images, etc.), use a scanner or OCR (Optical Character Recognition) technology to convert them into editable electronic documents (such as Word, Excel, etc.). This process ensures the digitization and editability of the contract content, facilitating subsequent processing.
[0058] To improve the accuracy of scanned paper contracts, a high-resolution scanner is used to scan the contracts into image files. OCR software is then used to recognize characters in the scanned image files, converting them into editable text files. The format of the OCR-recognized documents is then adjusted to ensure consistency with the target format.
[0059] Editable documents are categorized and consolidated into documents with the same format. The distinctions include: contract type, project duration type, and contract term. The converted editable documents are further processed to differentiate between contract types (e.g., probationary contracts, formal contracts, renewal contracts), project duration types (e.g., long-term contracts, short-term contracts, temporary contracts), and contract terms (e.g., fixed-term contracts, indefinite-term contracts). Based on these distinctions, the documents are consolidated into documents with the same format. This process ensures clear categorization and consistent formatting of contract data, facilitating subsequent management and analysis.
[0060] Based on predefined rules or pattern matching algorithms, key information in contracts is automatically identified and categorized.
[0061] Contracts that cannot be accurately identified by rule matching are manually categorized and adjusted by professionals.
[0062] Use document processing software or tools to convert documents of different formats into the same target format.
[0063] In one specific embodiment, the method further includes the step of:
[0064] S5: Stores editable documents, including either metadata storage or file storage.
[0065] S6: Transfer editable documents to the labor supervision system and the labor management system respectively.
[0066] Specifically, the differentiated and categorized contract documents are stored in designated storage media (such as databases, file systems, or cloud storage) and then transmitted to the labor supervision system and labor management system. This process ensures the security and availability of contract data while promoting cross-departmental data sharing.
[0067] Choose the appropriate storage medium and storage strategy based on the type, size, and access frequency of the data.
[0068] Contract data can be transferred via API interfaces, file transfer protocols (such as FTP and SFTP), or data synchronization tools.
[0069] Assign appropriate access permissions to different users or systems to ensure the security and privacy protection of contract data.
[0070] In one specific embodiment, step S2 includes the following steps:
[0071] S21. Fill out the labor service agreement and verify it; the verification includes: identity information, salary details and work period type, etc.
[0072] S22. Enter the content to be filled in for the labor agreement; the input methods include: mobile phone input, computer input, and scanning of physical documents.
[0073] Specifically, during the completion of the employment contract, ensure that all necessary information is filled in accurately and completely. This includes key clauses such as the identity information of both parties, job duties, and salary and benefits. After completion, the contract should be checked by professionals or automatically by the system to ensure its accuracy and consistency.
[0074] Use predefined contract templates to fill out the forms, ensuring consistency in format and completeness of content.
[0075] Use data validation rules or algorithms to automatically check the entered content, such as ID card number verification and telephone number format verification.
[0076] For content that cannot be fully covered by automatic verification, it is manually checked and confirmed by professionals.
[0077] During the completion of the employment contract, ensure that all necessary information is filled in accurately and completely. This includes key clauses such as the identity information of both parties, job duties, and salary and benefits. After completion, the contract should be checked by professionals or automatically by the system to ensure its accuracy and consistency.
[0078] Use predefined contract templates to fill out the forms, ensuring consistency in format and completeness of content.
[0079] Use data validation rules or algorithms to automatically check the entered content, such as ID card number verification and telephone number format verification.
[0080] For content that cannot be fully covered by automatic verification, it is manually checked and confirmed by professionals.
[0081] In one specific embodiment, step S3 includes the following steps:
[0082] S31: Use OCR light symbol recognition to extract content from entity documents. The extracted content includes: job content, work location, work hours, and remuneration.
[0083] S32: Detect whether the page is tilted. If so, introduce a page correction algorithm.
[0084] S33: Use image preprocessing techniques to improve the quality of scanned images.
[0085] Specifically, for paper contracts or existing image documents, OCR technology is used for character recognition. During the recognition process, page tilt is detected and corrected using a page correction algorithm. Simultaneously, image preprocessing techniques (such as noise reduction and contrast enhancement) are used to improve the quality of the scanned image, thereby increasing the accuracy and efficiency of OCR recognition.
[0086] Relevant keywords and text content are generated based on the contract content. These keywords and text content will be used for subsequent classification, retrieval, and management. Simultaneously, labor law keywords and text content are used to conduct compliance checks on the labor contracts, ensuring the legality and compliance of the contract content.
[0087] Keywords are extracted from the content of labor contracts using natural language processing techniques or keyword extraction algorithms.
[0088] Generate a concise and clear text description or summary based on the content of the labor contract.
[0089] The content of the labor contract is compared with the relevant provisions of the labor law to ensure the legality and compliance of the contract.
[0090] Synonym integration and matching refers to identifying synonyms or near-synonyms appearing in multiple labor contracts and integrating them into a unified vocabulary or tag for subsequent processing and analysis.
[0091] In one specific embodiment, step S4 includes the following steps:
[0092] S41: Contract classification is performed using a pre-trained text classification model. Contract classification includes: text content, keyword distribution, and semantic feature vector.
[0093] S42: Label contract data and generate a model training set. Contract data includes: probationary contracts, formal contracts, renewal contracts, etc.
[0094] S43: Detect unmatched keywords or abnormal content in the contract and mark them for priority manual review.
[0095] S431: Manually process contracts with abnormal classifications and confirm the modification of classification content.
[0096] S432: Add the manually labeled results to the model training set.
[0097] Specifically, before using a pre-trained text classification model to classify contracts, text preprocessing is performed. Text preprocessing mainly includes removing irrelevant information such as stop words, punctuation marks, and numbers, as well as performing operations such as word segmentation, stemming, or word form restoration to improve the effectiveness of subsequent steps.
[0098] Text preprocessing can be achieved using tools such as regular expressions and natural language processing libraries (e.g., NLTK, spaCy).
[0099] Text vectorization is the process of converting text into numerical vectors so that machine learning models can process it. Common methods include the Bag of Words model, TF-IDF, and word embeddings (such as Word2Vec and BERT).
[0100] Taking TF-IDF as an example, its formula is as follows:
[0101] TF (Term Frequency): The number of times a word appears in a document divided by the total number of words in the document.
[0102]
[0103] IDF (Inverse Document Frequency): The logarithm of the total number of documents divided by the number of documents containing the word.
[0104]
[0105] TF-IDF value: The product of TF and IDF.
[0106] TF-IDF(t,d)=TF(t,d)×IDF(t).
[0107] The above example illustrates this: Consider a document set containing two contracts, Contract 1: "This contract sells goods," and Contract 2: "This contract concludes the contract." For the word "sell," the TF in Contract 1 is 1 / 3 (because Contract 1 contains 3 words), and the IDF in the document set is log(2 / 1) = 0 (because only one document contains "sell," but we usually add a small smoothing term to avoid dividing by 0). Therefore, the TF-IDF value for "sell" is (1 / 3)*0 (actually, it will be a very small positive number because of the smoothing term).
[0108] During the model training phase, preprocessed and vectorized data are needed to train a classifier, such as logistic regression, support vector machine (SVM), Naive Bayes, or deep learning models (such as convolutional neural network CNN, recurrent neural network RNN, BERT, etc.).
[0109] Taking logistic regression as an example, its hypothesis function is:
[0110]
[0111] Here, x is the input feature vector (i.e., the result of text vectorization), and θ is the model parameters. The training process of the model is to find an optimal set of θ that minimizes the prediction error of the model on the training set.
[0112] The contracts are categorized into two types: sales contracts and service contracts. After text preprocessing and vectorization, we obtain a feature matrix X (each row represents a contract, and each column represents a feature) and a label vector y (each element represents the corresponding contract category). We can then use a logistic regression algorithm to train a classifier that learns how to classify new contract texts as either sales contracts or service contracts.
[0113] A text classification model was trained using a large amount of contract data to improve classification accuracy and generalization ability. The classification criteria included text content, keyword distribution, and semantic feature vectors.
[0114] The contract data is manually annotated to clearly identify the type of each contract (e.g., trial period contract, formal contract, renewal contract, etc.), the type of work period, and the contract duration. This annotated data will be used to generate the model training set.
[0115] The labeled contract data is divided into training, validation, and test sets according to a certain ratio (e.g., 70% training set, 15% validation set, and 15% test set). The training set is used to train the text classification model, the validation set is used to adjust the model parameters to improve performance, and the test set is used to evaluate the model's classification performance.
[0116] During the classification process, algorithms are used to detect unmatched keywords or content anomalies in the contracts. These anomalous contracts are marked for priority manual review to ensure classification accuracy.
[0117] Contracts marked as abnormal are manually reviewed and processed by professionals. Based on the review results, the classification information is confirmed and modified.
[0118] Manually processed contract data (including modified classification content and manual annotation results) is added to the model training set. This helps to continuously update and improve the text classification model, enhancing its classification performance and accuracy.
[0119] In one specific embodiment, step S41 includes the following steps:
[0120] Step S411: Extract the entered keywords or text content.
[0121] Step S412: Output the corresponding contract type label.
[0122] Specifically, before classification, the input keywords or text content are extracted from the contract. These keywords or text content will serve as input features for the text classification model.
[0123] The text classification model outputs corresponding contract type labels based on the input keywords or text content. These labels will be used for subsequent classification and management.
[0124] In one specific embodiment, step S3 further includes the step:
[0125] S34: After integration, the original multi-regional labor agreements will be re-established as labor contracts with the same format.
[0126] S35: Retain the original document.
[0127] The categorized and processed contract data will be integrated into a standardized labor contract document. Original documents will be retained for future review and auditing. The integrated contract document will be used for subsequent contract management and data analysis.
[0128] The categorized and processed contract data will be integrated and stored in a unified format.
[0129] Store the original documents in a specified storage medium (such as a database, file system, or cloud storage) and set appropriate access permissions and backup policies.
[0130] The contract management system is used to manage and maintain the integrated contract documents, including operations such as contract querying, modification, renewal and termination.
[0131] Therefore, the beneficial effects of the above-mentioned method and system for automatically classifying construction site labor contracts are as follows:
[0132] We retrieved and integrated labor agreements from different regions, standardized the format, and created standardized construction site labor contract documents.
[0133] Extract core information from the contract (such as job description, location, and remuneration), and integrate similar expressions through synonym matching. Incorporate labor law-related keywords to ensure the comprehensiveness and accuracy of the extracted content.
[0134] Contracts with mismatched keywords or abnormal content are marked and prioritized for manual review. The review results are then used to optimize the classification model.
[0135] The above description is merely an embodiment of this application. It should be noted that those skilled in the art can make improvements without departing from the inventive concept of this application, but these improvements all fall within the protection scope of this application.
Claims
1. A method for automatic classification of construction service contracts, characterized by, The method comprises the steps of: calling the labor agreement of the corresponding region, establishing the construction site labor contract of the corresponding region; inputting the relevant content of the construction site labor contract, and forming the document of the construction site labor contract of the corresponding region after inputting; scanning the document and converting the document into an editable document; the specific steps are as follows: generating keywords and text content related to the construction site labor contract; using the keywords and text content of the labor law; using the same words to integrate and match the same words using the same words; distinguishing the editable document and integrating it into the same format document, the distinguishing content of the editable document includes: contract type, construction period type, contract period; the specific steps are as follows: using a pre-trained text classification model to classify the contract, the contract classification includes: text content, keyword distribution, semantic feature vector; annotating the contract data and generating a model training set; detecting unmatched keywords or content abnormalities in the contract, and marking them for manual review and priority processing; manually processing the classified abnormal contract and confirming the modification of the classification content; adding the manual annotation result to the model training set.
2. The method of automatically classifying construction services contracts of claim 1, wherein, The method of automatic classification of the construction site labor contract further comprises the steps of: storing the editable document, the storage method includes: metadata database storage or file storage; transmitting the editable document to the labor supervision system and the labor management system respectively.
3. The method of automatically classifying construction services contracts of claim 1, wherein, The method of "inputting the relevant content of the construction site labor contract, and forming the document of the construction site labor contract after inputting" comprises the steps of: filling out the construction site labor contract and checking the construction site labor contract; inputting the content filled out in the construction site labor contract; the input method includes: mobile phone input, computer filling and scanning of physical documents.
4. The method of automatically classifying construction services contracts of claim 1, wherein, The method of "scanning the document and converting the document into an editable document" further comprises: using OCR optical symbol recognition to extract the content in the physical document, the extracted content includes: work content, work location, work time and work remuneration; detecting whether the page is inclined, if so, introducing a page correction algorithm; using image preprocessing technology to improve the quality of the scanned picture.
5. The method of automatically classifying construction service contracts of claim 1, wherein, The method of "using a pre-trained text classification model to classify the contract" comprises: extracting the inputted keywords or text content; outputting the corresponding contract type label.
6. The method of automatically classifying construction services contracts of claim 1, wherein, The method of "scanning the document and converting the document into an editable document" further comprises: after integration, the original multi-region labor agreement is newly established as a labor contract of the same format; keeping the original document.
7. A system for automatic classification of construction service contracts, characterized by The system comprises the method of automatic classification of the construction site labor contract according to any one of claims 1-6.
Citation Information
Patent Citations
System and method for digitally and intelligently identifying contract
CN119251861A