Intelligent document identification and classification method in electric power project management

By obtaining project progress and document adaptability in real time in power project management, combining technical means such as project progress pressure function and frequency weighting function, dynamically adjusting document priority, solving the problem of document processing lag in the existing technology, and achieving efficient document processing and resource optimization.

CN120029982AActive Publication Date: 2025-05-23FUJIAN YIRONG INFORMATION TECH

Patent Information

Application Number
CN202510511997.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The document management system in the existing power project management cannot dynamically adjust document priorities, resulting in lagging key document processing at different stages of the project, affecting project progress and quality.

Method used

By setting key nodes of the project, the project progress is obtained in real time, the initial priority is determined based on the adaptability of the document type in the project stage, and the project progress pressure function, frequency weighting function and frequency attenuation function are used for dynamic adjustment, and the document priority is adaptively adjusted based on the extraction of text, semantics and structured features and the multi-category classification of the random forest model.

Benefits of technology

It realizes intelligent sorting and resource optimization allocation of document processing, ensures that high-priority documents are processed in a timely manner, improves the intelligence and accuracy of document processing, and effectively improves document processing efficiency in power project management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029982A_ABST
    Figure CN120029982A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent document identification and classification method in electric power project management, relates to the technical field of document management of electric power projects, and is used for solving the problem of poor document identification and classification management in the electric power projects. According to the method, key nodes of a project are set, project progress is obtained in real time, the initial priority is determined according to the adaptability of document types and project stages, dynamic adjustment is carried out through a project progress pressure function, intelligent sorting and resource optimization of document processing are achieved, and through a frequency weighting function and a frequency attenuation function, document processing efficiency is improved. According to the method, a document processing queue is updated in real time, it is ensured that high-priority documents are processed in time, then text, semantic and structured features of the documents are extracted as vectors, multi-category classification is conducted through a random forest model, the document priorities are adjusted in a self-adaptive mode according to classification results, dynamic adjustment and accurate management of the document priorities are achieved, and the document processing efficiency is improved. The document processing efficiency and the intelligent level in electric power project management are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document management of electric power projects, and more specifically, to an intelligent document recognition and classification method in electric power project management. Background Art

[0002] In power project management, there are many documents in terms of quantity and types, including engineering design drawings, construction plans, contracts, technical specifications, equipment maintenance manuals, etc. These documents cover all stages of the project from planning, design, construction to operation and maintenance, with complex content and various formats. At the same time, as the project progresses, the documents will be continuously updated, and different versions of the documents may be used alternately.

[0003] Deficiencies in existing technologies: In the existing document management of power project management, document priority is usually set based on static rules and cannot be dynamically adjusted as the project stages progress. Especially when the project enters the construction and acceptance stages from the planning and design stages, the demand and importance of documents change significantly, and it is impossible to perceive changes in project progress in a timely manner and automatically adjust the priority of documents. The limitations of this static priority setting often lead to delays in the processing of key documents such as construction plans, equipment lists, acceptance reports, etc. in the later stages of the project, and even resources are occupied by low-priority documents, which cannot be processed and paid attention to in a timely manner, thereby affecting the overall progress and quality of the project. Summary of the invention

[0004] In order to overcome the above-mentioned defects of the prior art, there is a solution as follows to solve the problem of poor document management efficiency of power projects in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions: An intelligent document recognition and classification method in power project management includes the following steps: Set key project nodes, obtain and update project progress in real time, determine the initial priority of documents based on the adaptability of document types in the project stage, adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority; Perform multi-dimensional analysis on the priorities of the sorted documents, determine the dynamic priorities of the documents based on the frequency weighting function and the frequency decay function, and update the document processing queue in real time based on the dynamic priorities of the documents; The text, semantic and structural features of the document are extracted and converted into vectors, and the vectors are input into the random forest model for multi-category classification of the document. The document priority is adaptively adjusted according to the document category classification results. Analyze the adaptive adjustment process of document priority, obtain and analyze the priority adjustment information in the scheduling optimization process, and adjust the document management strategy based on the different signals generated by the analysis.

[0006] In a preferred implementation, key project nodes are set, project progress is acquired and updated in real time, and the initial priority of documents is determined according to the adaptability of document types at the project stage. The specific steps are as follows: The key nodes of the project are set, including the starting point of the design phase, the starting point of the construction phase, and the starting point of the acceptance phase; When the project progresses to each stage, the current stage is automatically updated and the calibration value is assigned; Define the document type. First, define the document type and establish a matching function between the document type and the project stage based on an experience-based scoring mechanism. The adaptability feedback function is calculated based on the actual frequency of use and number of modifications of the document during project operation. The document adaptability coefficient is calculated by combining the project stage matching function with the adaptability feedback function, and the initial priority of the document is adjusted based on the document adaptability coefficient.

[0007] In a preferred implementation, the document priority is adjusted by the project progress pressure function, and the documents are sorted according to the adjusted priority. The specific steps include: Determine the growth trend of project pressure based on the project's planned completion time and current time; Construct a document priority function based on the document's initial priority, document adaptability coefficient, and the growth trend of project pressure; When the document priority function is constructed and calculated, all documents to be processed are sorted according to the calculated priority values, and the documents are sorted in reverse order according to the priority values, and the sorted documents are stored in the project's document management database; The database structure contains document ID, priority value, document type, and project phase information.

[0008] In a preferred embodiment, a multi-dimensional analysis is performed on the sorted document priorities, the dynamic priorities of the documents are determined according to the frequency weighting function and the frequency decay function, and the document processing queue is updated in real time according to the dynamic priorities of the documents. The specific steps are as follows: Collect data on access behavior, editing behavior, continuous usage duration, and user operation patterns; The access behavior, editing behavior, continuous usage time, and user operation mode data are used to construct a usage frequency weighted function to determine the usage frequency of the document. A usage frequency decay mechanism is introduced to control the frequency decay of the document and determine the dynamic priority adjustment function of the document in each time period. The dynamic priority value of the document is updated according to the dynamic priority adjustment function, the dynamic priority values ​​of all documents are recalculated in a fixed time period, and the document queue is reordered according to the new dynamic priority value.

[0009] In a preferred embodiment, the text, semantic and structural features of the document are extracted and converted into vectors, and the vectors are input into a random forest model for multi-category classification of the document, and the document priority is adaptively adjusted according to the document category classification results, including the following steps: Extract text features from documents based on regular expressions and perform word frequency statistics using word frequency inverse document frequency, and use the extracted features as text features; Use the BERT model for semantic extraction and semantic features; Use TableNet to extract table data and YOLO image processing algorithm to extract images, and use the extracted features as structured features; The text features, semantic features, and structural features are input into the random forest model as document feature vectors, and the document category is obtained through majority voting; After the document category classification is completed, the priority of each document is dynamically adjusted according to the matching degree between the document category and the project stage.

[0010] In a preferred embodiment, analyzing the adaptive adjustment process of document priority, obtaining and analyzing priority adjustment information in the scheduling optimization process, includes the following steps: Obtaining priority adjustment information during the scheduling optimization process, including document processing information and resource usage information; The fluctuation trend information includes the document processing urgency index, and the load response information includes the document complexity index; The acquired document processing urgency index and document complexity index are combined to generate a document management coefficient; The document processing urgency index and document complexity index are directly proportional to the document management coefficient.

[0011] In a preferred embodiment, adjusting the document management strategy according to the different signals generated by the analysis includes the following steps: Compare the generated forecast adjustment factor with the set management control threshold; If the forecast adjustment coefficient is greater than or equal to the management control threshold, a document management fluctuation signal is generated, indicating that the urgency and complexity of document processing have reached a critical level, and additional measures need to be taken to deal with the instability of document management, and the frequency of monitoring and checking of documents should be increased; If the predicted adjustment coefficient is less than the management control threshold, a document management stability signal is generated. There is no need to allocate more resources to the document or take emergency measures, and no additional intervention is required.

[0012] Technical effects and advantages of an intelligent document recognition and classification method in power project management of the present invention: The present invention sets key project nodes, obtains project progress in real time, preliminarily determines the priority of documents according to the adaptability of document types in the project stage, and then uses the project progress pressure function for dynamic adjustment, thereby realizing intelligent sorting of document processing and optimal resource allocation. In combination with the frequency weighting function and the frequency attenuation function, the document processing queue is updated in real time according to the frequency of document use to ensure that high-priority documents are processed in a timely manner. Then, the documents are converted into vectors through the extraction of text, semantic and structured features, and multi-category classification is performed using the random forest model. The priority of the documents is further adaptively adjusted according to the classification results to ensure that the document processing order is dynamically matched with the project requirements, thereby realizing dynamic adjustment of the document priority, improving the intelligence and accuracy of document processing, and effectively improving the document processing efficiency in power project management. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 The present invention is a flowchart of an intelligent document recognition and classification method in power project management. DETAILED DESCRIPTION

[0014] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0015] In order to achieve the above objectives, Figure 1 A structural schematic diagram of an intelligent document recognition and classification method in power project management of the present invention is given, which specifically includes the following steps: Set key project nodes, obtain and update project progress in real time, determine the initial priority of documents based on the adaptability of document types in the project stage, adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority; Perform multi-dimensional analysis on the priorities of the sorted documents, determine the dynamic priorities of the documents based on the frequency weighting function and the frequency decay function, and update the document processing queue in real time based on the dynamic priorities of the documents; The text, semantic and structural features of the document are extracted and converted into vectors, and the vectors are input into the random forest model for multi-category classification of the document. The document priority is adaptively adjusted according to the document category classification results. Analyze the adaptive adjustment process of document priority, obtain and analyze the priority adjustment information in the scheduling optimization process, and adjust the document management strategy based on the different signals generated by the analysis.

[0016] Step 1: Track project progress and initialize document priority. That is, in the process of intelligent document recognition and classification in power project management, tracking project progress is the basis for dynamically adjusting document priority. To ensure that the document priority can accurately reflect the current project stage, it is necessary to first obtain the key nodes of the project in real time through project progress management, and associate them with document types and usage requirements, and then assign appropriate priorities to documents through initial calculations. The specific steps are as follows: The project uses pre-set key nodes (such as planning completion, design review, construction start, project acceptance, etc.) to mark the start and end of each stage. The time points of these key nodes are set as ,in, It is the starting point of the design phase, and the phase calibration value is 1 (such as the planned completion time); is the starting point of the construction phase (e.g., the time of design completion review), and the phase calibration value is 2; is the starting point of the acceptance phase (such as the completion time of construction), the phase calibration value is 3, and the phase of the project at the current time point t is determined by the start time of the project and the starting point of the project phase, and the phase calibration value is set. For example, when , the calibration value is 2, and the current time belongs to the design stage; Track the progress of the project in real time. When the project progress reaches each stage, the current stage is automatically updated and a calibration value is assigned. The calibration value will be used to dynamically adjust the document priority in subsequent steps. This information can be automatically extracted from the project management platform. Through the synchronization mechanism with the schedule and calendar, the project progress tracking data is ensured to be accurate. Through stage calibration, the current stage of the project (such as planning, design, construction, etc.) can be determined at any time point, providing a basis for document processing at each stage. This step is achieved through automated project management data synchronization and plays a fundamental role in adjusting the document priority. In power project management, different types of documents have different importance at different stages. For example, planning documents are more important at the beginning of the project, while the construction stage is mainly based on documents such as construction plans and equipment lists. Therefore, a calculation mechanism is needed to measure the adaptability of specific document types at the current project stage and adjust the initial priority of the documents accordingly. The specific steps are as follows: Define the document type. First, define the document type d, such as design document, construction document, acceptance document, etc. Each document type d corresponds to its adaptability in different project stages. Create document type and project phase matching function , used to measure the current project stage The importance of document type d is shown below. The project stage matching function adopts an experience-based scoring mechanism with a value range of [0, 1]. The higher the value, the more suitable the document is for the current stage, as follows: ; In order to enable the document priority to dynamically respond to the actual situation of the project, an adaptive feedback function is introduced to calculate the document priority based on the actual frequency of use and number of modifications of the document during the project operation. For example, if a document is frequently used at a certain stage, its priority will be further improved. The adaptive feedback function expression is as follows: ,in, is the intrinsic usage frequency of the document (calculated based on historical projects), is the modification frequency of the document at the current stage (e.g., the rate of change of the number of edits and accesses); The adaptive feedback function is used to reflect the dynamic performance of a document in a specific stage of a project. The adaptability of a document is dynamically adjusted based on its usage in the current project stage and its historical performance. The application of this function is more focused on project schedule management and document type matching. The adaptive feedback function is mainly used to capture the instant usage frequency and modification of a document in the current project stage. For example, if a document is frequently referenced or modified in the planning stage, the function will increase the priority of the document in that stage accordingly. By combining the project stage and document type for instant priority adjustment, it ensures that the document priority matches the current needs of the project. Similarly, the document adaptability coefficient is calculated. Finally, the document adaptability coefficient is calculated by combining the project stage matching function with the adaptability feedback function: ,in, and It is a preset weight parameter that controls the proportion of static matching degree and dynamic feedback of the document, and adjusts the initial priority of the document according to the static adaptability and actual usage of the document type; After the project progress tracking and document adaptability coefficient calculation are completed, the document priority needs to be further quantified by comprehensively considering multiple factors. The document priority function not only depends on the adaptability of the document type and the project stage, but also needs to be combined with the urgency of the project, the schedule pressure and the inherent importance of the document. By building a priority function, we can ensure the accurate adjustment of the document priority at different stages of the project and provide a basis for subsequent document processing. The specific steps are as follows: Construction of project progress pressure function h(t): The closer the project progress is to the deadline, the higher the urgency of document processing. Therefore, it is necessary to construct a progress pressure function h(t) to describe the growth trend of project pressure at the current time t. It is defined as follows: ,in, is the planned completion time of the project, t is the current time, and k is the adjustment coefficient for controlling the pressure growth, which reflects the sensitivity of pressure changes when the project is close to completion. Increases in size, ensuring that the document's priority is significantly increased as the project nears completion; Construct the document priority function. ,in, It is the initial priority of the document, which can be determined based on the importance of the aforementioned document types. This formula ensures that the document priority is gradually adjusted as the project progresses based on the adaptability of the document type and the project stage. The priority of urgent documents is significantly increased when the project is close to completion. When the document priority function is calculated, all documents to be processed will be sorted according to the priority value. Documents with higher priority values ​​will be ranked at the front. The sorted documents and their priority data will be stored in the project's document management database. The database structure can include the following fields: document ID (uniquely identifies the document), priority value, document type, and project phase information.

[0017] Step 2: Document usage frequency analysis and dynamic priority adjustment. After completing project progress tracking and initial document priority setting, the document priority needs to be dynamically adjusted based on its usage frequency in the actual project. The number of document accesses, edits, and actual usage reflect the project's reliance on the document. The specific steps are as follows: Collect access behavior data. The document management system automatically monitors and records the number of times each document is accessed, reflecting the total number of times a document is opened by team members within a set time period. Every time a user accesses a document, the system automatically increases the access count of the document and records the exact timestamp of the access. Editing behavior data collection, that is, monitoring the editing behavior of the document, collecting the number of edits, that is, recording the number of times the document is modified within a set time period. Whenever the document content changes (such as text addition, deletion, version update, etc.), the edit count of the document is increased, and the detailed modification record of each edit is marked, including the modified content and the modifying user. The data collection of editing behavior is not limited to a single modification, but also includes the count of batch operations; Continuous usage duration data collection: After accessing and editing a document, the duration of document usage is another important dimension. The duration of a document being opened after each access or edit is recorded. By monitoring the user's interaction with the document (such as mouse movement, scroll bar operation, etc.), it is determined whether the document is still in use. If it is detected that the document has not been operated for a long time (exceeding the set timeout threshold), the duration of this usage is recorded and classified into the corresponding time period. User operation pattern data collection, that is, to more accurately reflect the frequency of document use, monitor the user's operation mode when accessing the document. For example, the frequency of mouse clicks, scrolling, copying and pasting operations is used to judge the activeness of document use. Operation activities are quantified as operation counts and combined with the frequency of use to ensure that documents that have been opened for a long time but not operated on are not given unreasonably high priority; All the data collected above are recorded in real time and stored in the document management database; The frequency of document usage does not depend solely on the number of visits or edits, but is the result of a combination of factors. The construction of the frequency weighted function needs to combine multiple dimensions, such as the number of visits, the number of edits, the length of time used, and the user's operational activity, to ensure that each dimension has a reasonable impact on the priority of the document. In order to avoid deviations caused by excessive expansion of data in a certain dimension, different dimensions need to be appropriately weighted and dynamically adjusted. The specific definition of the frequency weighted function is as follows: ,in, is the number of visits, reflecting the frequency with which the document is visited. is the number of edits, representing how frequently the document is modified, and the weight coefficient Reflect the impact of editing actions on document priority; is the duration of continuous use of the document, using a logarithmic function To smooth out the usage time of a long period of time, avoid excessive expansion of the priority, and the weight coefficient Reflects the impact of the duration of document use on document priority; It is the user operation activity, which reflects the user's activeness in document use through the operation frequency. The weight coefficient Used to regulate the impact of operation activity on priority; It is the weight coefficient that regulates the expansion of the duration, ensuring that when a document is opened for a long time but not actually used, the priority will not be increased unreasonably. The settings of each weight coefficient are adjusted according to actual needs; The frequency-weighted function is used to analyze the frequency of actual use of documents throughout the project. It is not limited to a specific project stage, but through long-term monitoring of document access times, edit times, usage time and other dimensions, the frequency of use of documents is comprehensively measured, and the priority is adjusted through this function. The frequency-weighted function is used to dynamically track the actual use of documents, especially those documents that are frequently accessed and frequently modified. This requirement may occur at any stage of the project, and is used to measure the degree of continuous use of documents. Some documents still have too high a priority in the later stage due to their high historical usage. It is necessary to introduce a frequency decay mechanism so that the historical usage frequency of the document gradually decays over time. The frequency decay function is defined as follows: ,in, The decay rate controls the speed of frequency decay. The larger the value, the faster the decay. The time point when the document was last used frequently. It is used to record the time when the document was used frequently in history, ensuring that the priority of the document will gradually decrease in a certain period of time in the future. After combining the document usage frequency weighting function and usage frequency decay function, the dynamic priority adjustment function is defined as follows: ,in, is the initial document priority function (calculated by steps 1 and 2). The dynamic priority function ensures that within each time period t, the priority of a document can be adjusted according to its usage frequency, and that documents with a high historical frequency of use will gradually decay their priority to prevent affecting the subsequent priority allocation; Update the dynamic priority value of documents in real time through the document management platform , and pass it to the document processing queue. The dynamic priority values ​​of all documents are recalculated at fixed time intervals (such as every hour or every day), and the document queue is reordered according to the new dynamic priority values ​​to ensure that high-priority documents can be processed in a timely manner, avoiding historical high-frequency documents from occupying resources for a long time, and ultimately ensuring the accuracy and real-time nature of the priority calculation through a real-time feedback mechanism.

[0018] Step 3: Identify the document content category and adaptively adjust the priority. In power project management, different types of documents have different priority requirements at different stages of the project. By analyzing the content characteristics of the document, the document category (such as planning, design, construction, etc.) can be automatically determined, and the document priority can be dynamically adjusted according to the matching degree between the document type and the project stage to ensure that the processing order of the document matches the project requirements and realize document management. The specific steps are as follows: Perform text feature extraction, i.e., tokenization and cleaning. Use a tokenization algorithm (such as a tokenizer based on regular expressions or a natural language processing toolkit) to tokenize the text content of the document. After tokenization, remove non-informative content such as stop words, punctuation marks, and irrelevant characters; Perform word frequency statistics. Use TF-IDF (Term Frequency-Inverse Document Frequency) or other text vectorization methods to convert the tokenization results into numerical vectors. The weight of each word is calculated based on its importance in the entire document collection. Common words (such as "de", "he") have lower weights, while rare technical terms have higher weights. Convert the text content of each document into a feature vector , where each element represents the TF-IDF value of a word in the document. The vectorized text features are used for subsequent classification tasks; Adopt a pre-trained deep learning model (such as BERT, GPT) to perform semantic encoding on the sentences in the document. Models such as BERT can perform more accurate semantic understanding of words based on the context and extract the deep semantic information of the sentences. The specific steps are as follows: Input the document text into the BERT model to obtain the context vector of each sentence; For each sentence in the document, generate a semantic vector , where each vector represents the semantic features of a certain sentence in the document, , where, is the semantic vector of the i-th sentence in the document; Aggregate the semantic vectors of all sentences (such as through weighted average or max pooling) to obtain a document-level semantic feature vector for capturing the semantic integrity of the entire document. The expression is as follows: , where, represents the overall semantic feature vector of the document, and m is the total number of sentences in the document; Semantic feature extraction encodes each sentence in the document through a deep learning model to capture the deep semantic information of the document, so as to be able to understand the actual meaning and context relationship of the text; For documents containing tables or structured data, use image recognition or table parsing algorithms (such as TableNet, Tabula) to extract the table data. First, use the image recognition algorithm to identify the position of the table in the document, and then parse the row and column data in the table to extract the key numerical values or information in the table; Example: A table in a power equipment list document contains equipment names, specifications, quantities, etc. It is necessary to automatically parse the table and extract these fields; for example, , where, represents the table data feature vector extracted from document type d; For technical drawings or documents containing images, use image processing algorithms (such as OpenCV or YOLO) to extract key features in the drawings, such as equipment symbols, connection lines, and other information in the drawings, to ensure that important graphic elements in technical documents can be accurately parsed; for example, ,in, represents the feature vector of tabular data extracted from document type d; The extracted table and drawing information is converted into vector form and input into the subsequent classification model as part of the document content. The purpose of structured information extraction is to ensure that non-text information in the document can also participate in the document classification and priority adjustment process: ,Structured information extraction ensures that table and image data in documents can be identified, especially key equipment lists or drawing information in technical documents, which are useful for document classification and priority adjustment.

[0019] After the content feature extraction of the document is completed, each document needs to be automatically classified. The identification of document categories is a task of multi-category classification based on the text, semantic and structural features of the document through a machine learning model. The classified document categories will be used to guide subsequent priority adjustments to ensure that different types of documents have a reasonable processing order at different stages of the project. The specific steps are as follows: Build a training dataset that contains documents of known categories (such as planning documents, design documents, construction documents, etc.), each of which has been manually classified and labeled according to its content At the same time, the text features in the feature extraction step , semantic features , structural features As input to the random forest model, the input vector is the document feature vector: ; During the training process, random forest constructs multiple decision trees by randomly selecting features and samples to avoid excessive reliance on a single feature on the model. Each decision tree uses a different subset of features to generate classification rules, and ultimately the category of the document is determined by majority voting. The model is formulated as follows: ,in, It represents the classification result or prediction value of the hth decision tree in the random forest for the input sample, and N is the total number of decision trees in the random forest; In the input document feature vector After that, each decision tree in the random forest classifies the document independently, and finally determines the final category of the document by majority voting. ; Tune the hyperparameters of the random forest to achieve the best classification results. The key hyperparameters include: the number of decision trees. More trees can improve the accuracy of classification, but also increase the computational cost. The typical value is between 100 and 500. Maximum tree depth is used to control the complexity of the decision tree and prevent overfitting. When the depth is shallow, the model tends to be simple and generalizes well; when the depth is large, the model has stronger fitting ability; Minimum sample split number, used to control the minimum number of samples when each tree node is split. A larger value can reduce the complexity of the model. Using the preprocessed document dataset (including feature vectors and category labels ), train the random forest model, use the document dataset from the historical project, adjust the model parameters through supervised learning, until the model can accurately classify documents, the training process optimizes the model by minimizing the classification error, and the loss function is defined as the classification error: ,in, Is an indicator function, indicating whether the classification result of each tree is consistent with the true category match; Use cross-validation methods (such as K-fold cross-validation) to evaluate the model. During the training process, divide the data set into multiple subsets, which are alternately used as training sets and validation sets to evaluate the performance of the model on different data and ensure that the model has good classification capabilities.

[0020] Perform adaptive adjustment of document types. The goal of adaptive adjustment of document types is to dynamically adjust the priority of each document based on the matching degree between the document category and the current project stage after the document category classification is completed. Documents of different categories have different priorities at different project stages, so their matching degree is evaluated to ensure that important documents are prioritized at the right time. Analyze the adaptive adjustment process of document priority, dynamically adjust the priority of documents, ensure that the priority adjustment process is more reasonable, ensure that key documents are given priority at the correct stage, and obtain the priority adjustment information in the adaptive adjustment process of document priority, which includes document processing information and resource usage information; Document processing information includes the document processing urgency index and is calibrated as WDC, and resource usage information includes the document complexity index and is calibrated as FZX; The document processing urgency index is used to measure the urgency of document processing in the current time period, that is, the time pressure of the document from its final processing deadline (such as the task deadline), and takes into account the relative position of the document in the current processing queue. By evaluating the time constraints of the document, it ensures that the system can give priority to those documents that are close to the processing deadline and are more urgent; The main function of the document processing urgency index is to help the management system dynamically adjust the priority according to the time pressure of the document. As the document processing deadline approaches, the document's urgency index will increase, and the system will increase the priority of the document to ensure that it can be processed in time: In power project management, different types of documents (such as planning documents, design documents, and construction documents) usually have clear task deadlines. The processing urgency index provides data support for task management and time scheduling, ensuring that when time is tight, resources and manpower can be allocated to tasks that need to be processed immediately first, avoiding the delay of key documents that affects project progress; By quantifying the time pressure of documents, the urgency index can help the system prioritize resources and time to documents that need to be processed as quickly as possible, thereby improving the overall document processing efficiency and avoiding delays in the processing of important documents due to low-priority documents occupying resources; For example, during the construction phase, some temporarily added design documents or approval documents may have a very short processing period. At this time, the document processing urgency index will quickly increase the priority of the document by considering the absolute time urgency and the relative position in the queue to prevent these documents from affecting the urgent tasks of the project. The logic for obtaining the document processing urgency index is as follows: Get the final deadline for the current document to be completed or submitted , get the total number of all pending documents in the currently set processing queue and the current time t, the relative processing order of the document in the currently set processing queue ; Get the time buffer for the current document processing , calculate the nonlinear time buffer difference based on the current time t and the final deadline of the current document. The calculation expression is as follows: , calculate the document processing urgency index: .

[0021] It should be noted that the final deadline is usually provided by the document properties in the project planning or task allocation system; documents are sorted in the processing queue based on multiple factors such as initial priority and expiration time, and all pending documents are sorted by priority, and a sequential number is generated for each document The sequence number indicates the relative processing order of the document in the current processing queue, indicating the processing priority order of the document in the queue. The smaller the value, the closer the document is to processing.

[0022] The Document Complexity Index (FZX) is used to measure the content structure of the document, the number of technical terms, the diagrams and technical drawings in the document, and the complexity of cross-departmental collaboration. It is used to comprehensively evaluate the difficulty of document processing. This index reflects the processing resources, time and reliance on professional knowledge required for the document, helping the system to better allocate processing priorities and reasonably allocate resources when resources are limited. The main function of the document complexity index is to help the relevant management system to reasonably allocate resources when resources are limited. For documents with higher complexity, the system will allocate more processing time and resources to ensure that the documents can be fully processed. For simpler documents, the system can give priority to processing when resources are tight to avoid wasting time and energy. For complex documents, the risks in processing are often higher (such as processing errors, misunderstandings, etc.). The document complexity index helps to identify these potential risks and arrange more detailed reviews or more experienced personnel to process these documents to reduce the risk of document processing errors; Adjust the order of document processing by using the document complexity index. Documents with lower complexity (such as routine reports or notices) can be processed first at the beginning of a project or when resources are sufficient, while documents with higher complexity may need to be delayed until there is enough time and resources to process them. The logic for obtaining the document complexity index is as follows: Get the character length of the pth technical term in the document , get the correlation between the pth term and other terms, the calculation expression is as follows: ,in, Representation term and terminology The number of co-occurrences of the p-th term; Get the number of departments that have cross-departmental collaboration related to the p-th term , the complex value of cross-departmental collaboration is calculated, and the calculation expression is: ; Get error feedback and error time during document processing and the total processing time of the document , calculate the document error accumulation value, the calculation expression is: , where Indicates the qth error, and the document complexity index is calculated. The calculation expression is: , where is the total number of terms.

[0023] It should be noted that the length of a term is usually related to its complexity. The longer the term, the more difficult it is to interpret and process. After identifying the technical terms in the document, the system calculates the number of characters of each term through a string processing function, and the length of each term is recorded and stored. For different types of documents or terms, basic values ​​​​can be set for the association of different terms through historical data to better reflect the complex relationship between terms.

[0024] The document processing information and resource usage information are combined to generate the document management coefficient, that is, the acquired document processing urgency index and document complexity index are combined to generate the document management coefficient. The expression is: , where is the preset proportionality factor of the document processing urgency index and document complexity index, and Both are greater than 0.

[0025] The specific method of jointly generating a document management coefficient may involve a variety of algorithms and models, which depends on the actual situation and application requirements. This embodiment can use a weighted summation method to combine the document processing urgency index and the document complexity index to generate a comprehensive document management coefficient. This document management coefficient can be used as an input for evaluating the adjustment of document priority in the power project management process to determine whether the document priority needs to be adjusted.

[0026] It should be noted that the size of the preset proportionality coefficient is a specific value obtained by quantifying each parameter. In order to facilitate subsequent comparison, the size of the coefficient depends on the amount of sample data and the preset proportionality coefficient initially set by technical personnel in this field for each group of sample data. It is not unique, as long as it does not affect the proportional relationship between the parameter and the quantized value. For example, the document processing urgency index is proportional to the document management coefficient. The document processing urgency index and the document complexity index are normalized to have the same dimension and range. This can be achieved by subtracting the mean from the original data and dividing it by the standard deviation, or mapping the data to the range of [0, 1].

[0027] The greater the document processing urgency index and the document complexity index, the greater the jointly generated document management coefficient, indicating that the processing urgency and complexity of the document are at a high level. The time, resources, personnel and technical support required to process the document must increase. More professional manpower, equipment and time need to be allocated to this document to ensure that the complex document can be processed on time and with high quality. This means that the document has a very high priority in the current project or task process and must be processed first to avoid delays or errors in the project progress due to improper processing; The smaller the document processing urgency index and the document complexity index, the smaller the jointly generated document management coefficient, indicating that the urgency and complexity of the document in the current project phase are low, which means that there is relatively ample time for document processing and the processing difficulty is relatively low. Therefore, other more urgent or complex documents can be processed first. Such documents can be processed later and will not pose a threat to the overall progress of the project. Compare the generated document management coefficient with the preset management control threshold to generate a document management fluctuation signal and a document management stability signal; After obtaining the document management coefficient, compare the document management coefficient with the management control threshold; If the document management coefficient is greater than or equal to the management control threshold, a document management fluctuation signal is generated, indicating that the urgency and complexity of document processing have reached a critical level. At this time, the priority of the document may suddenly rise, requiring the relevant management system of the document to respond immediately and take additional measures to deal with the instability of document management, and monitor and check such documents more frequently. You can ensure that the document is processed within the specified time by starting real-time monitoring, strengthening reminders, and notifying relevant responsible personnel. For example, project managers should trigger emergency mechanisms, including speeding up the approval process, increasing the processing frequency, and even cross-departmental collaboration to ensure that the document is completed before the key node; If the document management coefficient is less than the management control threshold, a document management stability signal is generated, indicating that the urgency and complexity of document processing are within a controllable range, and the document can be processed at a normal pace with existing resources. There is no need to allocate more resources or take emergency measures. The current resources and processing capabilities can cope with the processing needs of the document without additional intervention, indicating that the document management work is progressing stably.

[0028] When a document management fluctuation signal is generated, it means that the document management coefficient has exceeded the preset management control threshold, indicating that the urgency and complexity of document processing have reached a high level, which may affect the project progress or resource allocation. In this case, a series of countermeasures need to be taken to ensure that the document can be processed in time and avoid negative impact on the overall progress of the project. The following are the specific processing steps: Improve document priority: The document should be re-evaluated immediately and moved to the front of the queue to avoid delays that could hinder project progress. Prioritization can prevent documents from exceeding deadlines or increasing complexity; Resource allocation and task assignment: Based on the changes in the document management coefficient, more processing resources should be allocated to the document, which may include adding staff, mobilizing more technical support or extending the processing time to cope with the complexity of the document. The complexity of the document processing task should be analyzed to determine whether senior staff or experts are needed to process the document, and task assignment should be optimized according to the document requirements to ensure that the appropriate personnel are involved in the processing; At the same time, the progress of the document processing is monitored at a high frequency, and each step of the processing process is tracked in real time to ensure that the document is completed on time. An automatic reminder or progress reporting mechanism can be set up to promptly identify and resolve potential problems.

[0029] In summary, when a document management fluctuation signal is generated, project managers need to take emergency measures, including raising priority, allocating more resources, real-time monitoring of processing progress, cross-departmental coordination, and timely adjustment of project schedules. These steps ensure that the system can respond efficiently in document emergencies and avoid adverse effects on the project caused by delayed document processing. At the same time, the causes of the fluctuation signals are recorded and analyzed to optimize future document management processes.

[0030] It should be noted that the threshold information in this embodiment is pre-set by professionals and will not be explained in detail here. Some parameter English letters in the embodiments have the same situation, but different meanings are explained when used, which will not be explained one by one here.

[0031] The present invention sets key project nodes, obtains project progress in real time, preliminarily determines the priority of documents according to the adaptability of document types in the project stage, and then uses the project progress pressure function for dynamic adjustment, thereby realizing intelligent sorting of document processing and optimal resource allocation. In combination with the frequency weighting function and the frequency attenuation function, the document processing queue is updated in real time according to the frequency of document use to ensure that high-priority documents are processed in a timely manner. Then, the documents are converted into vectors through the extraction of text, semantic and structured features, and multi-category classification is performed using the random forest model. The priority of the documents is further adaptively adjusted according to the classification results to ensure that the document processing order is dynamically matched with the project requirements, thereby realizing dynamic adjustment of the document priority, improving the intelligence and accuracy of document processing, and effectively improving the document processing efficiency in power project management.

[0032] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0033] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0034] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0035] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0036] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0037] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An intelligent document recognition and classification method in power project management, characterized by: The steps include: Set key project nodes, obtain and update project progress in real time, determine the initial priority of documents based on the adaptability of document types in the project stage, adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority; Perform multi-dimensional analysis on the priorities of the sorted documents, determine the dynamic priorities of the documents based on the frequency weighting function and the frequency decay function, and update the document processing queue in real time based on the dynamic priorities of the documents; The text, semantic and structural features of the document are extracted and converted into vectors, and the vectors are input into the random forest model for multi-category classification of the document. The document priority is adaptively adjusted according to the document category classification results. Analyze the adaptive adjustment process of document priority, obtain and analyze the priority adjustment information in the scheduling optimization process, and adjust the document management strategy based on the different signals generated by the analysis.

2. The intelligent document recognition and classification method in power project management according to claim 1 is characterized by: Set key project nodes, obtain and update project progress in real time, and determine the initial priority of documents based on the adaptability of document types at the project stage. The specific steps are as follows: The key nodes of the project are set, including the starting point of the design phase, the starting point of the construction phase, and the starting point of the acceptance phase; When the project progresses to each stage, the current stage is automatically updated and the calibration value is assigned; Define the document type. First, define the document type and establish a matching function between the document type and the project stage based on an experience-based scoring mechanism. The adaptability feedback function is calculated based on the actual frequency of use and number of modifications of the document during project operation. The document adaptability coefficient is calculated by combining the project stage matching function with the adaptability feedback function, and the initial priority of the document is adjusted based on the document adaptability coefficient.

3. The intelligent document recognition and classification method in power project management according to claim 2 is characterized by: Then adjust the document priority through the project progress pressure function, and sort the documents according to the adjusted priority. The specific steps include: Determine the growth trend of project pressure based on the project's planned completion time and current time; Construct a document priority function based on the document's initial priority, document adaptability coefficient, and the growth trend of project pressure; When the document priority function is constructed and calculated, all documents to be processed are sorted according to the calculated priority values, and the documents are sorted in reverse order according to the priority values, and the sorted documents are stored in the project's document management database; The database structure contains document ID, priority value, document type, and project phase information.

4. The intelligent document recognition and classification method in power project management according to claim 3 is characterized by: Perform multi-dimensional analysis on the priorities of the sorted documents, determine the dynamic priorities of the documents based on the frequency weighting function and the frequency attenuation function, and update the document processing queue in real time based on the dynamic priorities of the documents. The specific steps are as follows: Collect data on access behavior, editing behavior, continuous usage duration, and user operation patterns; The access behavior, editing behavior, continuous usage time, and user operation mode data are used to construct a usage frequency weighted function to determine the usage frequency of the document. A usage frequency decay mechanism is introduced to control the frequency decay of the document and determine the dynamic priority adjustment function of the document in each time period. The dynamic priority value of the document is updated according to the dynamic priority adjustment function, the dynamic priority values ​​of all documents are recalculated in a fixed time period, and the document queue is reordered according to the new dynamic priority value.

5. The intelligent document recognition and classification method in power project management according to claim 4 is characterized by: The text, semantic and structural features of the document are extracted and converted into vectors, and the vectors are input into the random forest model for multi-category classification of the document. The document priority is adaptively adjusted according to the document category classification results, including the following steps: Extract text features from documents based on regular expressions and perform word frequency statistics using word frequency inverse document frequency, and use the extracted features as text features; Use the BERT model for semantic extraction and semantic features; Use TableNet to extract table data and YOLO image processing algorithm to extract images, and use the extracted features as structured features; The text features, semantic features, and structural features are input into the random forest model as document feature vectors, and the document category is obtained through majority voting; After the document category classification is completed, the priority of each document is dynamically adjusted according to the matching degree between the document category and the project stage.

6. The intelligent document recognition and classification method in power project management according to claim 5 is characterized by: Analyze the adaptive adjustment process of document priority, obtain and analyze the priority adjustment information in the scheduling optimization process, including the following steps: Obtaining priority adjustment information during the scheduling optimization process, including document processing information and resource usage information; The fluctuation trend information includes the document processing urgency index, and the load response information includes the document complexity index; The acquired document processing urgency index and document complexity index are combined to generate a document management coefficient; The document processing urgency index and document complexity index are directly proportional to the document management coefficient.

7. The intelligent document recognition and classification method in power project management according to claim 4 is characterized by: The document management strategy is adjusted based on the different signals generated by the analysis, including the following steps: Compare the generated forecast adjustment factor with the set management control threshold; If the forecast adjustment coefficient is greater than or equal to the management control threshold, a document management fluctuation signal is generated, indicating that the urgency and complexity of document processing have reached a critical level, and additional measures need to be taken to deal with the instability of document management, and the frequency of monitoring and checking of documents should be increased; If the predicted adjustment coefficient is less than the management control threshold, a document management stability signal is generated. There is no need to allocate more resources to the document or take emergency measures, and no additional intervention is required.

Citation Information

Patent Citations

  • Electronic pregnant woman health record management method and system

    CN117854664A

  • Science and technology project document key information extraction method based on document elements

    CN118349674A

  • Medical data document classification and marking system

    CN119621972A

  • Intelligent archive management system based on multi-type analysis

    CN119646278A

  • Cloud document management method and platform based on block chain technology

    CN119807157A

Cited By

  • Enterprise data asset intelligent management system for server research, development and manufacturing

    CN120765199A

  • An enterprise data asset intelligent management system for server research and development

    CN120765199B