An intelligent document recognition and classification method in power project management

By setting key nodes in power project management and dynamically adjusting document priorities, combining frequency analysis and classification models, the processing lag caused by static settings of document priorities is solved, and intelligent and efficient management of document processing is realized.

CN120029982BActive Publication Date: 2025-07-22FUJIAN YIRONG INFORMATION TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510511997.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the existing power project management, document priorities cannot be adjusted dynamically, resulting in lagging processing of key documents, affecting project progress and quality.

Method used

By setting key nodes of the project, the project progress is obtained in real time, the initial priority is determined based on the adaptability of the document type in the project stage, and the project progress pressure function is used for dynamic adjustment, the document processing queue is updated in combination with the frequency weighting and attenuation functions, the text, semantics and structured features are extracted into vectors, and the random forest model is used for multi-category classification, and the document priority is adaptively adjusted.

Benefits of technology

It realizes intelligent sorting and resource optimization of document processing, ensures timely processing of high-priority documents, and improves document processing efficiency and accuracy in power project management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029982B_ABST
    Figure CN120029982B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent document recognition and classification method in power project management, which relates to the technical field of document management in power projects and is used to solve the problem of poor document recognition and classification management in power projects. By setting key project nodes, obtaining the project progress in real time, determining the initial priority according to the adaptability of document types and project stages, and dynamically adjusting using the project progress pressure function, the present invention realizes the intelligent sorting of document processing and resource optimization. Through the frequency weighting function and the frequency decay function, the document processing queue is updated in real time to ensure that high-priority documents are processed in a timely manner. Then, the text, semantic, and structured features of the documents are extracted as vectors and classified into multiple categories through a random forest model. According to the classification results, the document priorities are adaptively adjusted, realizing the dynamic adjustment and precise management of document priorities, effectively improving the document processing efficiency and intelligent level in power project management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document management for power projects. More specifically, the present invention relates to an intelligent document recognition and classification method in power project management. Background Art

[0002] In power project management, the quantity and types of documents are numerous, including engineering design drawings, construction plans, contracts, technical specifications, equipment maintenance manuals, etc. These documents cover all stages of the project from planning, design, construction to operation and maintenance, with complex content and diverse formats. At the same time, as the project progresses, the documents will be continuously updated, and different versions of the documents may be used alternately.

[0003] Deficiencies of the prior art: In the existing document management of power project management, the document priority is usually set based on static rules and cannot be dynamically adjusted as the project stage progresses. Especially when the project enters the construction and acceptance stages from the planning and design stages, the requirements and importance of the documents change significantly, and it is impossible to timely perceive the changes in the project progress and automatically adjust the document priority. The limitations of this static priority setting often lead to delays in the processing of key documents such as construction plans, equipment lists, and acceptance reports in the later stage of the project, and even the resources are occupied by low-priority documents and cannot be processed and concerned in a timely manner, thus affecting the overall progress and quality of the project. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the following solutions are provided to solve the problem of poor document management efficiency in the power project in the above-mentioned background art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] An intelligent document recognition and classification method in power project management, comprising the following steps:

[0007] Set key project nodes, obtain and update the project progress in real time, determine the initial priority of the documents according to the adaptability of the document types in the project stage, then adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority;

[0008] Conduct multi-dimensional analysis on the sorted document priorities, determine the dynamic priority of the documents according to the frequency weighting function and the frequency decay function, and update the document processing queue in real time according to the dynamic priority of the documents;

[0009] Extract and transform the text, semantics, and structured features of the documents into vectors, and input the vectors into a random forest model for multi-class classification of the documents, and adaptively adjust the document priority according to the document category classification results;

[0010] Analyze the process of adapting document priorities, obtain and analyze the priority adjustment information during the scheduling optimization process, and adjust the document management strategy according to different signals generated by the analysis.

[0011] In a preferred embodiment, set project key nodes, obtain and update the project progress in real time, and determine the initial priority of the document according to the adaptability of the document type in the project phase. The specific steps are as follows:

[0012] Setting project key nodes includes the starting point of the design phase, the starting point of the construction phase, and the starting point of the acceptance phase;

[0013] When the project progress reaches each phase, automatically update the current phase and assign a calibration value;

[0014] Define document types. First, define the types of documents, and establish a matching function between document types and project phases based on an empirical scoring mechanism;

[0015] Calculate the adaptability feedback function based on the actual usage frequency and modification times of the document during project operation. Combine the project phase matching function and the adaptability feedback function to calculate the document adaptability coefficient, and adjust the initial priority of the document according to the document adaptability coefficient.

[0016] In a preferred embodiment, further adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority. The specific steps include:

[0017] Determine the growth trend of project pressure based on the planned completion time and the current time of the project;

[0018] Construct a document priority function based on the initial priority of the document, the document adaptability coefficient, and the growth trend of project pressure;

[0019] After the document priority function is constructed and calculated, sort all the documents to be processed according to the calculated priority value, arrange the documents in reverse order according to the priority value, and store the sorted documents in the project's document management database;

[0020] The database structure includes document ID, priority value, document type, and project phase information.

[0021] In a preferred embodiment, conduct a multi-dimensional analysis of the sorted document priorities, determine the dynamic priority of the document according to the frequency weighting function and the frequency decay function, and update the document processing queue in real time according to the dynamic priority of the document. The specific steps are as follows:

[0022] Collect data on access behavior, editing behavior, continuous usage duration, and user operation patterns;

[0023] The access behavior, editing behavior, continuous usage time, and user operation mode data are used to construct a usage frequency weighted function to determine the usage frequency of the document. A usage frequency decay mechanism is introduced to control the frequency decay of the document and determine the dynamic priority adjustment function of the document in each time period.

[0024] The dynamic priority value of the document is updated according to the dynamic priority adjustment function, the dynamic priority values of all documents are recalculated in a fixed time period, and the document queue is reordered according to the new dynamic priority value.

[0025] In a preferred embodiment, the text, semantic and structural features of the document are extracted and converted into vectors, and the vectors are input into a random forest model for multi-category classification of the document, and the document priority is adaptively adjusted according to the document category classification results, including the following steps:

[0026] Extract text features from documents based on regular expressions and perform word frequency statistics using word frequency inverse document frequency, and use the extracted features as text features;

[0027] Use the BERT model for semantic extraction and semantic features;

[0028] Use TableNet to extract table data and YOLO image processing algorithm to extract images, and use the extracted features as structured features;

[0029] The text features, semantic features, and structural features are input into the random forest model as document feature vectors, and the document category is obtained through majority voting;

[0030] After the document category classification is completed, the priority of each document is dynamically adjusted according to the matching degree between the document category and the project stage.

[0031] In a preferred embodiment, analyzing the adaptive adjustment process of document priority, obtaining and analyzing priority adjustment information in the scheduling optimization process, includes the following steps:

[0032] Obtaining priority adjustment information during the scheduling optimization process, including document processing information and resource usage information;

[0033] The fluctuation trend information includes the document processing urgency index, and the load response information includes the document complexity index;

[0034] The acquired document processing urgency index and document complexity index are combined to generate a document management coefficient;

[0035] The urgency index of document processing, the complexity index of documents, and the document management coefficient are in a direct proportional relationship.

[0036] In a preferred embodiment, according to the different signals generated by the analysis, the document management strategy is adjusted, including the following steps:

[0037] Compare the generated prediction adjustment coefficient with the set management control threshold;

[0038] If the prediction adjustment coefficient is greater than or equal to the management control threshold, a document management fluctuation signal is generated, indicating that the processing urgency and complexity of the document have reached a critical level. Additional measures need to be taken to address the unstable situation of document management, and the frequency of monitoring and inspection of the document is increased;

[0039] If the prediction adjustment coefficient is less than the management control threshold, a document management stability signal is generated, and there is no need to allocate more resources or take emergency measures for the document, and no additional intervention is required.

[0040] The technical effects and advantages of an intelligent document recognition and classification method in power project management of the present invention:

[0041] By setting the key nodes of the project, the progress of the project is obtained in real time, and according to the adaptability of the document type in the project stage, the priority of the document is initially determined. Then, the dynamic adjustment is carried out by using the project progress pressure function, realizing the intelligent sorting and optimal allocation of resources for document processing. Combining the frequency weighting function and the frequency decay function, the document processing queue is updated in real time according to the usage frequency of the document to ensure that high-priority documents are processed in a timely manner. Then, by extracting text, semantic, and structured features, the document is converted into a vector, and a random forest model is used for multi-category classification. According to the classification results, the priority of the document is further adaptively adjusted to ensure the dynamic matching of the document processing order and the project requirements, so as to realize the dynamic adjustment of the document priority, improve the intelligence and accuracy of document processing, and effectively improve the document processing efficiency in power project management. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic flowchart of an intelligent document recognition and classification method in power project management of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] To achieve the above object, Figure 1 The structural schematic diagram of an intelligent document recognition and classification method in a power project management of the present invention is given, and the specific steps are as follows;

[0045] Set the key project nodes, obtain and update the project progress in real time, determine the initial priority of the document according to the adaptability of the document type in the project stage, then adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority;

[0046] Conduct multi-dimensional analysis on the sorted document priorities, determine the dynamic priority of the documents according to the frequency weighting function and the frequency decay function, and update the document processing queue in real time according to the dynamic priority of the documents;

[0047] Extract and transform the text, semantics and structural features of the document into vectors, and input the vectors into the random forest model for multi-category classification of the document, and adaptively adjust the document priority according to the document category classification result;

[0048] Analyze the process of adaptively adjusting the document priority, obtain and analyze the priority adjustment information in the scheduling optimization process, and adjust the document management strategy according to different signals generated by the analysis.

[0049] Step 1: Conduct project progress tracking and document priority initialization, that is, in the process of intelligent document recognition and classification in power project management, the tracking of project progress is the basis for dynamically adjusting the document priority. To ensure that the document priority can accurately reflect the current project stage, it is first necessary to obtain the key nodes of the project in real time through project progress management and associate them with the document type and usage requirements, and then assign appropriate priorities to the documents through initial calculation. The specific steps are as follows:

[0050] The project calibrates the start and end of each stage through pre-set key nodes (such as planning completion, design review, construction start, project acceptance, etc.), and sets the time points of these key nodes as , where is the starting point of the design stage, and the stage calibration value is 1 (such as the planning completion time); is the starting point of the construction stage (such as the design completion review time), and the stage calibration value is 2; is the starting point of the acceptance stage (such as the construction completion time), and the stage calibration value is 3. Through the start time of the project and the stage starting point of the project, the project stage at the current time point t is determined and the stage calibration value is set. For example, when , the calibration value is 2, and the current time belongs to the design stage;

[0051] Track the progress of the project in real time. When the project progresses to each stage, automatically update the current stage and assign a calibration value, which will be used for the dynamic adjustment of the document priority in the subsequent steps. This information can be automatically extracted from the project management platform. Through the synchronization mechanism with the project plan and schedule, ensure the accuracy of the project progress tracking data. Through stage calibration, the current stage of the project (such as planning, design, construction, etc.) can be determined at any time point, providing a basis for document processing in each stage. This step is achieved through automated project management data synchronization and plays a fundamental role in the adjustment of document priority;

[0052] In power project management, different types of documents have different importance at different stages. For example, planning documents are more important in the initial stage of the project, while during the construction stage, documents such as construction plans and equipment lists are the main ones. Therefore, a computer mechanism is needed to measure the adaptability of a specific document type in the current project stage and adjust the initial priority of the document accordingly. The specific steps are as follows:

[0053] Define the document type. First, define the document type d, such as design documents, construction documents, acceptance documents, etc. Each document type d corresponds to its adaptability in different project stages;

[0054] Establish a matching function between the document type and the project stage to measure the importance of the document type d in the current project stage. The project stage matching function adopts an experience-based scoring mechanism, and the value range is set as [0, 1]. The higher the value, the more suitable the document is for the current stage, as follows: ; ;

[0055] To enable the document priority to dynamically respond to the actual situation of the project, introduce an adaptability feedback function, which is calculated based on factors such as the actual usage frequency and modification times of the document during the project operation. For example, if a document is frequently used in a certain stage, its priority will be further improved. The expression of the adaptability feedback function is as follows: , where is the inherent usage frequency of the document (calculated based on historical projects), is the modification frequency of the document in the current stage (such as the change rate of the number of edits and access times);

[0056] The adaptive feedback function is used to reflect the dynamic performance of a document in a specific phase of a project. Based on the usage of the document in the current project phase and combined with its historical performance, it dynamically adjusts its adaptability. The application of this function focuses more on project progress management and the matching of document types. The adaptive feedback function is mainly used to capture the immediate usage frequency and modification status of a document in the current project phase. For example, if a document is frequently cited or modified during the planning phase, the function will correspondingly increase the priority of the document in this phase. By combining the project phase and document type for immediate priority adjustment, it ensures that the document priority matches the current needs of the project;

[0057] Similarly, the document adaptability coefficient is calculated. Finally, the document adaptability coefficient is calculated by combining the project phase matching function and the adaptive feedback function: , where, and are preset weight parameters that control the proportion of the static matching degree and dynamic feedback of the document, and adjust the initial priority of the document according to the static adaptability and actual usage of the document type;

[0058] After the project progress tracking and the calculation of the document adaptability coefficient are completed, the document priority needs to be further quantified by comprehensively considering multiple factors. The document priority function not only depends on the adaptability between the document type and the project phase, but also needs to combine the urgency of the project, the progress pressure, and the inherent importance of the document. By constructing the priority function, it ensures the accurate adjustment of the document priority in different phases of the project and provides a basis for subsequent document processing. The specific steps are as follows:

[0059] Construction of the project progress pressure function h(t): The closer the project progress is to the deadline, the higher the urgency of document processing. Therefore, it is necessary to construct a progress pressure function h(t) to describe the growth trend of the project pressure at the current time t, which is defined as follows: , where, is the planned completion time of the project, t is the current time, and k is the adjustment coefficient that controls the pressure growth, which reflects the sensitivity of the pressure change when the project is approaching completion. This function increases as time t approaches so as to ensure that the priority of the document will be significantly improved when the project is approaching completion;

[0060] Construct the document priority function, , where, is the initial priority of the document, which can be determined according to the importance of the aforementioned document type. This formula ensures that the document priority is gradually adjusted as the project progresses based on the adaptability between the document type and the project phase, and the priority of urgent documents is significantly improved when approaching the completion of the project;

[0061] After the document priority function is calculated, all pending documents will be sorted according to the priority value. Documents with higher priority values will be ranked at the top. The sorted documents and their priority data will be stored in the project's document management database. The database structure can include the following fields: document ID (uniquely identifying the document), priority value, document type, and project phase information.

[0062] Step 2: Analysis of document usage frequency and dynamic priority adjustment. After completing the project progress tracking and initial priority setting of documents, the document priority needs to be dynamically adjusted according to its usage frequency in the actual project. The access times, edit times, and actual usage of the document reflect the project's dependence on the document. The specific steps are as follows:

[0063] Collect access behavior data. Automatically monitor and record the access times of each document through the document management system, which reflects the total number of times a certain document is opened by team members within the set time period. Whenever a user accesses a document, the system will automatically increase the access count of the document and record the precise timestamp of the access.

[0064] Collect edit behavior data, that is, monitor the edit behavior of the document and collect the edit times, which means recording the number of times the document is modified within the set time period. Whenever the content of the document changes (such as text addition, deletion, version update, etc.), increase the edit count of the document and mark the detailed modification records of each edit, including the modified content and the modifying user. The data collection of edit behavior is not limited to single modifications, but also includes the counting of batch operations.

[0065] Collect continuous usage duration data. After accessing and editing the document, the usage duration of the document is another important dimension. Record the duration that the document is continuously opened after each access or edit. By monitoring the user's interaction activities with the document (such as mouse movement, scroll bar operation, etc.), judge whether the document is still in use. If it is detected that the document has not been operated for a long time (exceeding the set timeout threshold), record the current usage duration and classify it into the corresponding time period.

[0066] Collect user operation mode data. To more accurately reflect the usage frequency of the document, monitor the operation mode of the user when accessing the document. For example, judge the activity level of document usage by the frequency of operation activities such as mouse clicks, scrolling, copy-pasting, etc. The operation activities are quantified as operation counts and combined with the usage frequency to ensure that documents that are opened for a long time but have no operations are not given unreasonably high priorities.

[0067] All the above collected data are recorded and stored in the document management database in real time.

[0068] The usage frequency of a document does not solely depend on the number of accesses or edits. Instead, it is the result of a comprehensive set of factors. The construction of the usage frequency weighting function requires considering multiple dimensions, such as the number of accesses, the number of edits, the usage duration, and the user operation activity level, to ensure that each dimension has a reasonable impact on the document's priority. To avoid bias caused by excessive inflation of data in a single dimension, appropriate weighting and dynamic adjustment are required for different dimensions. The specific definition of the usage frequency weighting function is as follows: , where is the number of accesses, reflecting the frequency of document access, is the number of edits, representing the frequency of document modification. The weight coefficient reflects the impact of the edit behavior on the document's priority; is the continuous usage duration of the document. The logarithmic function is used to smooth the usage duration over a long period to avoid excessive inflation of its impact on the priority. The weight coefficient reflects the impact of the document's continuous usage duration on the document's priority; is the user operation activity level, which reflects the user's activity during document usage through the operation frequency. The weight coefficient is used to regulate the impact of the operation activity level on the priority; is the weight coefficient for regulating the duration inflation, ensuring that the priority does not increase unreasonably when the document is opened for a long time but not actually used. The settings of each weight coefficient are adjusted according to actual needs;

[0069] The usage frequency weighting function is used to analyze the actual usage frequency of a document throughout the project. It is not limited to a specific project stage but comprehensively measures the usage frequency of the document by long-term monitoring of multiple dimensions such as the number of accesses, the number of edits, and the usage duration, and adjusts the priority through this function. The usage frequency weighting function is used to dynamically track the actual usage of a document, especially those with high-frequency access and high-frequency modification. This requirement may occur at any stage of the project and is used to measure the degree of continuous usage of the document;

[0070] For some documents that maintain an overly high priority in later stages due to high historical usage frequency, a usage frequency decay mechanism needs to be introduced to make the historical usage frequency of the document gradually decay over time. The definition of the frequency decay function is as follows: , where is the decay rate, which controls the speed of frequency decay. The larger the value, the faster the decay; is the time point of the last high-frequency usage of the document, which is used to record the time of historical high-frequency usage to ensure that the priority of the document will gradually decrease within a certain future time period;

[0071] After combining the document usage frequency weighting function and the usage frequency decay function, the dynamic priority adjustment function is defined as follows: , where is the initial document priority function (calculated by steps 1 and 2). The dynamic priority function ensures that within each time period t, the priority of the document can be adjusted according to its usage frequency, and the priority of the documents that have been frequently used in history will gradually decay to prevent affecting the later priority allocation;

[0072] The dynamic priority value of the document is updated in real time through the document management platform , and it is passed to the document processing queue. Every fixed time period (such as every hour or every day), the dynamic priority values of all documents are recalculated, and the document queue is reordered according to the new dynamic priority values to ensure that high-priority documents can be processed in a timely manner, avoiding the long-term occupation of resources by historical high-frequency documents. Finally, the accuracy and real-time nature of the priority calculation are ensured through the real-time feedback mechanism.

[0073] Step 3: Perform document content category recognition and priority adaptive adjustment. In power project management, different types of documents have different priority requirements at different stages of the project. By analyzing the content characteristics of the documents, the category of the document (such as planning, design, construction, etc.) can be automatically judged, and the priority of the document can be dynamically adjusted according to the matching degree between the document type and the project stage to ensure that the processing order of the documents matches the project requirements and realize document management. The specific steps are as follows:

[0074] Perform text feature extraction, that is, word segmentation and cleaning. Use a word segmentation algorithm (such as a word segmenter based on regular expressions or natural language processing toolkits) to segment the text content of the document. After word segmentation, remove non-informative content such as stop words, punctuation marks, and irrelevant characters;

[0075] Perform word frequency statistics. Use TF-IDF (term frequency-inverse document frequency) or other text vectorization methods to convert the word segmentation results into numerical vectors. The weight of each word is calculated according to its importance in the entire document set. Common words (such as "of", "and") have lower weights, and rare technical terms have higher weights. Convert the text content of each document into a feature vector , where each element represents the TF-IDF value of a word in the document. The vectorized text features are used for subsequent classification tasks;

[0076] Adopt a pre-trained deep learning model (such as BERT, GPT) to perform semantic encoding on the sentences in the document. Models such as BERT can perform more accurate semantic understanding of words according to the context and extract the deep semantic information of the sentences. The specific steps are as follows:

[0077] Input the document text into the BERT model to obtain the context vectors of each sentence;

[0078] For each sentence in the document, generate semantic vectors , where each vector represents the semantic features of a sentence in the document, , where, is the semantic vector of the i-th sentence in the document;

[0079] Aggregate the semantic vectors of all sentences (e.g., through weighted average or max pooling) to obtain a document-level semantic feature vector for capturing the semantic integrity of the entire document. The expression is as follows: , where, represents the overall semantic feature vector of the document, and m is the total number of sentences in the document;

[0080] Semantic feature extraction encodes each sentence in the document through a deep learning model to capture the deep semantic information of the document, thereby enabling the understanding of the actual meaning and context relationship of the text;

[0081] For documents containing tables or structured data, use image recognition or table parsing algorithms (such as TableNet, Tabula) to extract the table data. First, identify the table positions in the document through the image recognition algorithm, and then parse the row and column data in the table to extract the key numerical values or information in the table;

[0082] Example: A table in a power equipment list document contains equipment names, specifications, quantities, etc., and these fields need to be automatically parsed and extracted from the table; for example, , where, represents the table data feature vector extracted from document type d;

[0083] For technical drawings or documents containing images, use image processing algorithms (such as OpenCV or YOLO) to extract the key features in the drawings, such as equipment symbols, connection lines, etc. in the drawings, to ensure that the important graphic elements in technical documents can be accurately parsed; for example, , where, represents the table data feature vector extracted from document type d;

[0084] Convert the extracted table and drawing information into vector form and input it as part of the document content into the subsequent classification model. The purpose of structured information extraction is to ensure that the non-text information in the document can also participate in the document classification and priority adjustment process: , Structured information extraction ensures the ability to identify tabular and image data in documents, especially the key equipment lists or drawing information in technical documents, and this data plays a role in document classification and priority adjustment.

[0085] After the content feature extraction of the documents is completed, each document needs to be automatically classified into categories. The identification of document categories is a multi-classification task based on the text, semantics, and structured features of the documents through a machine learning model. The classified document categories will be used to guide subsequent priority adjustments to ensure that different types of documents have a reasonable processing order at different stages of the project. The specific steps are as follows:

[0086] Construct a training dataset that contains documents of known categories (such as planning documents, design documents, construction documents, etc.), and each document has been manually classified and labeled with category tags according to its content. , meanwhile, the text features , semantic features , structured features are used as the input of the random forest model, and the input vector is the document feature vector: ;

[0087] During the training process, the random forest constructs multiple decision trees by randomly selecting features and samples to avoid over-reliance of the model on a single feature. Each decision tree uses a different subset of features to generate classification rules, and finally determines the category of the document through majority voting. The formulaic expression of the model is: , where represents the classification result or prediction value of the h-th decision tree in the random forest for the input sample, and N is the total number of decision trees in the random forest;

[0088] After the input document feature vector , each decision tree in the random forest classifies the document independently, and finally determines the final category of the document through majority voting. ;

[0089] Tune the hyperparameters of the random forest to obtain the best classification effect. The key hyperparameters include: the number of decision trees. More trees can improve the classification accuracy but also increase the computational cost, and the typical value is between 100 and 500;

[0090] The maximum tree depth, which is used to control the complexity of the decision tree and prevent overfitting. When the depth is relatively shallow, the model tends to be simple and has good generalization; when the depth is relatively large, the model has stronger fitting ability;

[0091] The minimum number of samples for splitting, which is used to control the minimum number of samples when each tree node splits. A larger value can reduce the complexity of the model;

[0092] Using the preprocessed document dataset (including feature vectors and class labels ), train the random forest model. Use the document dataset in historical projects and adjust the model parameters through supervised learning until the model can accurately classify documents. The training process optimizes the model by minimizing the classification error, and the loss function is defined as the classification error: , where is an indicator function indicating whether the classification result of each tree matches the true class ;

[0093] Use the cross-validation method (such as K-fold cross-validation) to evaluate the model. During the training process, divide the dataset into multiple subsets and alternately use them as the training set and the validation set to evaluate the performance of the model on different data and ensure that the model has good classification ability.

[0094] Conduct document type adaptability adjustment. The goal of document type adaptability adjustment is to dynamically adjust the priority of each document according to the matching degree between the document category and the current project stage after the document category classification is completed. Different categories of documents have different priorities at different project stages, so evaluate their matching degree to ensure that important documents are given priority at the correct time;

[0095] Analyze the process of document priority adaptability adjustment, dynamically adjust the priority of documents to ensure that the priority adjustment process is more reasonable, ensure that key documents are given priority at the correct stage, and obtain the priority adjustment information during the document priority adaptability adjustment process. The priority adjustment information includes document processing information and resource usage information;

[0096] The document processing information includes the document processing urgency index calibrated as WDC, and the resource usage information includes the document complexity index calibrated as FZX;

[0097] The document processing urgency index is used to measure the urgency of document processing in the current time period, that is, the time pressure of the document from its last processing deadline (such as the task deadline), and considers the relative position of the document in the current processing queue. By evaluating the time constraints of the document, ensure that the system can give priority to processing those documents that are close to the processing deadline and are more urgent;

[0098] The main role of the document processing urgency index is to help the management system dynamically adjust the priority according to the time pressure of the document. As the document processing deadline approaches, the urgency index of the document will increase, and the system will raise the priority of the document to ensure that it can be processed in a timely manner:

[0099] In power project management, different types of documents (such as planning documents, design documents, and construction documents) usually have clear task time limits. The processing urgency index provides data support for task management and time scheduling, ensuring that in case of time urgency, resources and manpower can be preferentially allocated to tasks that need to be processed immediately, avoiding the situation where key documents affect the project progress due to delayed processing;

[0100] By quantifying the time pressure of documents, the urgency index can help the system preferentially allocate resources and time to documents that need to be processed as soon as possible, thereby improving the overall document processing efficiency and avoiding delays in the processing of important documents due to low-priority documents occupying resources;

[0101] For example, during the construction stage, some temporarily added design documents or approval documents may have very short processing deadlines. At this time, the document processing urgency index will quickly elevate the priority of the documents by considering the absolute time urgency and the relative position in the queue, avoiding these documents from affecting the urgent tasks of the project;

[0102] The acquisition logic of the document processing urgency index is as follows:

[0103] Obtain the final deadline for the current document to be completed or submitted , obtain the total number of all pending documents in the current set processing queue and the current time t, the relative processing order of the document in the current set processing queue ;

[0104] Obtain the time buffer for the current document processing , calculate the non-linear time buffer difference according to the current time t and the final deadline of the current document, and the calculation expression is as follows: , calculate the document processing urgency index: .

[0105] It should be noted that the final deadline is usually provided by the document attributes in the project plan or task assignment system; the documents are sorted in the processing queue according to multiple factors such as the initial priority and due date, and all pending documents are sorted by priority, and a sequence number is generated for each document , and the sequence number represents the relative processing order of the document in the current processing queue, indicating the processing priority order of the document in the queue. The smaller the value, the earlier the document is and the closer it is to being processed.

[0106] The Document Complexity Index (FZX) is used to measure the complexity of a document in multiple dimensions, such as its content structure, the number of technical terms, charts, technical drawings in the document, and the complexity of cross - departmental collaboration. It is used to comprehensively evaluate the difficulty of document processing. This index reflects the processing resources, time required for the document, and the degree of dependence on professional knowledge, helping the system better allocate processing priorities and reasonably allocate resources when resources are limited;

[0107] The main function of the Document Complexity Index is to help relevant management systems reasonably allocate resources when resources are limited. For documents with higher complexity, the system will allocate more processing time and resources to ensure that the documents can be fully processed. For simpler documents, the system can give priority to processing when resources are scarce, avoiding wasting time and energy;

[0108] For complex documents, the risks during processing are often higher (such as processing errors, understanding errors, etc.). The Document Complexity Index helps identify these potential risks and arranges more detailed reviews or more experienced personnel to handle these documents to reduce the risk of document processing errors;

[0109] By adjusting the document processing order through the Document Complexity Index, documents with lower complexity (such as regular reports or notifications) can be processed first at the initial stage of the project or when resources are sufficient, while documents with higher complexity may need to be delayed until there is enough time and resources for processing;

[0110] The acquisition logic of the Document Complexity Index is as follows:

[0111] Obtain the character length of the p - th technical term in the document , obtain the correlation degree between the p - th term and other terms, and the calculation expression is as follows: , where represents the term and the term 's co - occurrence times; obtain the number of departments involved in cross - departmental collaboration related to the p - th term , calculate the cross - departmental collaboration complexity value, and the calculation expression is: ;

[0112] Obtain the error feedback generated during document processing, the time when the error occurred and the total processing duration of the document , calculate the document error accumulation value, and the calculation expression is: , in the formula, represents the q - th error, calculate the Document Complexity Index, and the calculation expression is: , in the formula, is the total number of terms.

[0113] It should be noted that the length of a term is usually related to its complexity. The longer the term, the more difficult it is to interpret and process. After identifying the technical terms in the document, the system calculates the number of characters of each term through a string processing function, and the length of each term is recorded and stored. For different types of documents or terms, basic values ​​can be set for the association of different terms through historical data to better reflect the complex relationship between terms.

[0114] The document processing information and resource usage information are combined to generate the document management coefficient, that is, the acquired document processing urgency index and document complexity index are combined to generate the document management coefficient. The expression is: , where is the preset proportionality factor of the document processing urgency index and document complexity index, and Both are greater than 0.

[0115] The specific method of jointly generating a document management coefficient may involve a variety of algorithms and models, which depends on the actual situation and application requirements. This embodiment can use a weighted summation method to combine the document processing urgency index and the document complexity index to generate a comprehensive document management coefficient. This document management coefficient can be used as an input for evaluating the adjustment of document priority in the power project management process to determine whether the document priority needs to be adjusted.

[0116] It should be noted that the size of the preset proportionality coefficient is a specific value obtained by quantifying each parameter. In order to facilitate subsequent comparison, the size of the coefficient depends on the amount of sample data and the preset proportionality coefficient initially set by technical personnel in this field for each group of sample data. It is not unique, as long as it does not affect the proportional relationship between the parameter and the quantized value. For example, the document processing urgency index is proportional to the document management coefficient. The document processing urgency index and the document complexity index are normalized to have the same dimension and range. This can be achieved by subtracting the mean from the original data and dividing it by the standard deviation, or mapping the data to the range of [0, 1].

[0117] The greater the document processing urgency index and the document complexity index, the greater the jointly generated document management coefficient, indicating that the processing urgency and complexity of the document are at a high level. The time, resources, personnel and technical support required to process the document must increase. More professional manpower, equipment and time need to be allocated to this document to ensure that the complex document can be processed on time and with high quality. This means that the document has a very high priority in the current project or task process and must be processed first to avoid delays or errors in the project progress due to improper processing;

[0118] The smaller the document processing urgency index and the smaller the document complexity index, the smaller the jointly generated document management coefficient, indicating that the urgency and complexity of the document in the current project phase are relatively low. This means that there is relatively ample time for document processing and the processing difficulty is small. Therefore, other more urgent or complex documents can be processed first, and such documents can be processed later without threatening the overall project progress;

[0119] Compare the generated document management coefficient with the pre-set management regulation threshold to generate a document management fluctuation signal and a document management stability signal;

[0120] After obtaining the document management coefficient, compare the document management coefficient with the management regulation threshold;

[0121] If the document management coefficient is greater than or equal to the management regulation threshold, a document management fluctuation signal is generated at this time, indicating that the processing urgency and complexity of the document have reached a critical level. At this time, the priority of the document may suddenly rise, and the relevant document management system needs to respond immediately and take additional measures to deal with the unstable situation of document management. Such documents need to be monitored and inspected more frequently. It is possible to ensure that the document is processed within the specified time by starting real-time monitoring, strengthening reminders, notifying relevant responsible personnel, etc. For example, project managers should trigger emergency mechanisms, including accelerating the approval process, increasing the processing frequency, and even carrying out cross-departmental collaboration to ensure that the document is completed before the critical node;

[0122] If the document management coefficient is less than the management regulation threshold, a document management stability signal is generated, indicating that the processing urgency and complexity of the document are within the controllable range. This document can be processed at the normal pace with the existing resources, and there is no need to allocate more resources or take emergency measures for it. The current resources and processing capabilities can meet the processing requirements of this document without additional intervention, indicating that the document management work is progressing steadily.

[0123] When a document management fluctuation signal is generated, it means that the management coefficient of the document has exceeded the pre-set management regulation threshold, indicating that the processing urgency and complexity of the document have reached a relatively high level, which may affect the project progress or resource allocation. In this case, a series of countermeasures need to be taken to ensure that the document can be processed in a timely manner and avoid negative impacts on the overall project progress. The following are the specific processing steps:

[0124] Increase the document priority: The priority of this document should be immediately re-evaluated and moved to the front of the queue for processing to avoid delaying the project progress due to delayed processing. Prioritized processing can prevent the document from exceeding the deadline or the complex problems from intensifying;

[0125] Resource Allocation and Task Assignment: According to the changes in the document management coefficient, more processing resources should be considered for this document, which may include adding personnel, mobilizing more technical support, or extending the processing time to cope with the complexity of the document. Analyze the complexity of the document processing task, determine whether senior employees or experts are needed to handle the document, and optimize the task assignment according to the document requirements to ensure that suitable personnel are involved in the processing;

[0126] At the same time, monitor the progress of processing this document at a high frequency, track each step of the processing flow in real time to ensure that the document is completed on time. An automatic reminder or progress reporting mechanism can be set up to detect and solve potential problems in a timely manner.

[0127] In summary, when a document management fluctuation signal is generated, project managers need to take emergency handling measures, including raising the priority, allocating more resources, monitoring the processing progress in real time, cross-departmental coordination, and adjusting the project schedule in a timely manner. These steps ensure that the system can respond efficiently in case of document emergencies, avoid adverse effects on the project caused by lagging document processing, and record and analyze the reasons for the fluctuation signal to optimize future document management processes.

[0128] It should be noted that the threshold information related to this embodiment is pre-set by professionals and will not be explained in detail here. There are some cases where the English letters of some parameters in the embodiment are the same, but different meanings are explained when used, and they will not be explained one by one here.

[0129] The present invention realizes the intelligent sorting and optimal allocation of document processing by setting project key nodes, obtaining the project progress in real time, initially determining the priority of the document according to the adaptability of the document type in the project stage, and then dynamically adjusting it using the project progress pressure function. Combining the frequency weighting function and the frequency decay function, the document processing queue is updated in real time according to the usage frequency of the document to ensure that high-priority documents are processed in a timely manner. Then, by extracting text, semantic, and structured features, the document is transformed into a vector, and a random forest model is used for multi-class classification. According to the classification results, the priority of the document is further adaptively adjusted to ensure the dynamic matching of the document processing order and project requirements, so as to realize the dynamic adjustment of document priority, improve the intelligence and accuracy of document processing, and effectively improve the document processing efficiency in power project management.

[0130] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0131] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0132] Those of ordinary skill in the art will realize that the modules and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0133] In addition, in each of the embodiments of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0134] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0135] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should all be included in the protection scope of the present invention.

Claims

1. An intelligent document recognition and classification method in power project management, characterized in that: The steps are as follows: Set the key project nodes, obtain and update the project progress in real time, determine the initial priority of the document according to the adaptability of the document type in the project stage, then adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority; Conduct multi-dimensional analysis on the sorted document priorities, determine the dynamic priorities of the documents according to the frequency weighting function and the frequency decay function, and update the document processing queue in real time according to the dynamic priorities of the documents; Extract and transform the text, semantics and structural features of the document into vectors, and input the vectors into the random forest model for multi-class classification of the document, and adaptively adjust the document priority according to the document category classification results; Analyze the process of adaptive adjustment of the document priority, obtain and analyze the priority adjustment information in the scheduling optimization process, and adjust the document management strategy according to the different signals generated by the analysis; Then adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority. The specific steps include: Determine the growth trend of the project pressure according to the planned completion time and the current time of the project; Construct a document priority function according to the initial priority of the document, the document adaptability coefficient, and the growth trend of the project pressure; After the document priority function is constructed and calculated, sort all the documents to be processed according to the calculated priority value, arrange the documents in reverse order according to the priority value, and store the sorted documents in the project document management database; The database structure includes document ID, priority value, document type and project stage information; Conduct multi-dimensional analysis on the sorted document priorities, determine the dynamic priorities of the documents according to the frequency weighting function and the frequency decay function, and update the document processing queue in real time according to the dynamic priorities of the documents. The specific steps are as follows: Collect data on access behavior, editing behavior, continuous usage duration, and user operation mode; Construct a usage frequency weighting function for the access behavior, editing behavior, continuous usage duration, and user operation mode data, determine the usage frequency of the document, and introduce a usage frequency decay mechanism to control the frequency decay of the document, and determine the dynamic priority adjustment function of the document in each time period; Update the dynamic priority value of the document according to the dynamic priority adjustment function, set a fixed time period to recalculate the dynamic priority value of all documents, and re-sort the document queue according to the new dynamic priority value.

2. The intelligent document recognition and classification method in power project management according to claim 1, characterized in that: Set the key project nodes, obtain and update the project progress in real time, and determine the initial priority of the document according to the adaptability of the document type in the project stage. The specific steps are as follows: Setting the key project nodes includes the starting point of the design stage, the starting point of the construction stage, and the starting point of the acceptance stage; When the project progress reaches each stage, automatically update the current stage and assign a calibration value; Define the document type, first define the type of the document, and establish a matching function between the document type and the project stage based on an empirical scoring mechanism; The adaptability feedback function is calculated based on the actual frequency of use and number of modifications of the document during project operation. The document adaptability coefficient is calculated by combining the project stage matching function with the adaptability feedback function, and the initial priority of the document is adjusted based on the document adaptability coefficient.

3. The intelligent document recognition and classification method in power project management according to claim 1, characterized in that: The text, semantic and structural features of the document are extracted and converted into vectors, and the vectors are input into the random forest model for multi-category classification of the document. The document priority is adaptively adjusted according to the document category classification results, including the following steps: Extract text features from documents based on regular expressions and perform word frequency statistics using word frequency inverse document frequency, and use the extracted features as text features; Use the BERT model for semantic extraction and semantic features; Use TableNet to extract table data and YOLO image processing algorithm to extract images, and use the extracted features as structured features; The text features, semantic features, and structural features are input into the random forest model as document feature vectors, and the document category is obtained through majority voting; After the document category classification is completed, the priority of each document is dynamically adjusted according to the matching degree between the document category and the project stage.

4. An intelligent document recognition and classification method in power project management according to claim 3, characterized in that: Analyze the adaptive adjustment process of document priority, obtain and analyze the priority adjustment information in the scheduling optimization process, including the following steps: Obtaining priority adjustment information during the scheduling optimization process, including document processing information and resource usage information; The fluctuation trend information includes the document processing urgency index, and the load response information includes the document complexity index; The acquired document processing urgency index and document complexity index are combined to generate a document management coefficient; The document processing urgency index and document complexity index are directly proportional to the document management coefficient.

5. The intelligent document recognition and classification method in a power project management according to claim 1, characterized in that: The document management strategy is adjusted based on the different signals generated by the analysis, including the following steps: Compare the generated forecast adjustment factor with the set management control threshold; If the forecast adjustment coefficient is greater than or equal to the management control threshold, a document management fluctuation signal is generated, indicating that the urgency and complexity of document processing have reached a critical level, and additional measures need to be taken to deal with the instability of document management, and the frequency of monitoring and checking of documents should be increased; If the predicted adjustment coefficient is less than the management control threshold, a document management stability signal is generated. There is no need to allocate more resources to the document or take emergency measures, and no additional intervention is required.

Citation Information

Patent Citations

  • Electronic pregnant woman health record management method and system

    CN117854664A