Auxiliary filing method based on offshore wind power data
By constructing a dual-core knowledge base and automated processes, the problems of low efficiency, inconsistent standards, and low digitization in the archiving and management of offshore wind power data have been solved, achieving intelligent and standardized archiving management and improving archiving quality and retrieval efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-04-14
AI Technical Summary
The existing offshore wind power data archiving management suffers from problems such as low efficiency, reliance on manual labor, inconsistent standards, difficulty in traceability, and low level of digitization, making it difficult to improve the quality and efficiency of archiving.
A dual-core knowledge base is constructed. Standardized metadata is collected through a combination of automated process extraction and manual supplementation. This enables intelligent identification and classification, automatic volume assembly by volume/item rules, generation of standardized archive catalogs and multi-dimensional indexes, and verification of completeness and consistency, outputting standardized data packages.
This has enabled the transformation of offshore wind power data management from manual to intelligent, shortening the processing cycle, ensuring consistency of archived results, improving retrieval efficiency, reducing cross-project coordination and organization costs, and activating the value of data reuse.
Smart Images

Figure CN121858797A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data-assisted archiving, and specifically to a method for data-assisted archiving of offshore wind power. Background Technology
[0002] With the rapid development of the offshore wind power industry and the continuous expansion of project scale, the amount of documents and materials spanning the entire project lifecycle (covering construction phase survey and design, equipment procurement, civil construction, installation and commissioning, and production phase operation records, maintenance, inspection and monitoring, and technological upgrades) is growing exponentially. The current document archiving management methods commonly used in the industry have significant pain points: First, they are inefficient and highly dependent on manual labor. Document collection, organization, classification, volume (item) assembly, numbering, and catalog generation all require manual verification of massive amounts of documents, relying on personnel experience to judge categories and archiving rules, resulting in heavy workload and extremely long processing times. Second, inconsistent standards are prone to errors. Offshore wind power industry archiving standards (such as the "Table of Archiving Scope and Retention Period for Wind Power Enterprise Documents and Materials" and the "Regulations on Archive Management of Power Construction Projects") are complex, requiring differentiation between "item-based archiving" and "volume-based archiving," etc. Manual processing is prone to errors due to misunderstandings or oversights in the standards. Errors in class, association, and file numbering affect the quality of archiving; third, poor consistency and difficulty in traceability, with archiving results processed by different projects and personnel having difficulty in unifying formats and standards, making it difficult to find specific files during later audits or retrievals, and significantly reducing the traceability and utilization value of archives; fourth, low level of digitization, with existing document management systems mostly only realizing electronic storage of files, lacking intelligent processing capabilities for archiving business scenarios, and unable to complete the automated transformation from original files to standardized archiving results. There is an urgent need for an intelligent, efficient, and accurate method to solve the above-mentioned standardized archiving problems and improve the quality and efficiency of digital delivery results. Summary of the Invention
[0003] The purpose of this invention is to solve the technical problems mentioned above and to propose a method for auxiliary archiving of offshore wind power data, comprising the following steps: S1. Construct a dual-core knowledge base, digitally model the offshore wind power industry archiving standards into structured archiving rules, and configure explicit association rules and store them as database tables. S2. Obtain documents and capture metadata. Support batch uploading of multi-format offshore wind power project documents. Collect standardized metadata and complete verification through a combination of automatic extraction and manual supplementation. S3. Based on a dual-core knowledge base, intelligent identification and classification are achieved through document feature extraction and rule matching to determine the file's archiving category, retention period, archiving format, and associated tags. S4. Automatically assemble volumes according to volume / item rules and solidify the association relationships; S5. Generate a standardized archive directory and unique file number; S6. Automatically generate multi-dimensional indexes based on relationships and standardized metadata; S7. Verify human-computer interaction and optimize indexes; the system automatically completes integrity and consistency checks. S8 outputs standardized data packages, adaptable to offline archiving and digital system import.
[0004] In the preferred embodiment, the structured archiving rules of the dual-core knowledge base in step S1 include an archive classification table, a retention period table, a file number rule table, and a file type comparison table. The association rules include cross-stage technology association rules, device binding rules, hierarchical association rules, aggregated volume rules, and metadata cross-association rules. The rules can be added, deleted, modified, and queried through a graphical interface.
[0005] In the preferred embodiment, the standardized metadata in step S2 includes project code / unit project code, equipment number, technical topic keywords, professional category, responsible person, project stage, and retention period. Among them, the equipment number corresponds to the equipment lifecycle index, the technical topic keywords correspond to the technical process index, and the project code / unit project code and professional category correspond to the hierarchical classification index and the volume aggregation index.
[0006] In the preferred embodiment, the document feature extraction in step S3 includes extracting the file name, content keywords and standardized metadata. The rule matching achieves multi-dimensional matching through a pre-trained text classification model, and supports previewing the classification results and manual fine-tuning on the human-computer interaction interface.
[0007] In the preferred embodiment, the file grouping rule for archiving by volume in step S4 is "same unit project + same specialty + same business type", and the documents within the volume are arranged in the order of construction or time. The solidification of the association relationship is achieved by recording the association IDs of the documents with the case files, equipment, and technical topics, and the system automatically checks the integrity of the association and generates a verification report.
[0008] In the preferred embodiment, the multi-dimensional index in step S6 includes a technical process index, a device lifecycle index, a hierarchical classification index, a volume aggregation index, a responsible entity index, a time dimension index, and a retention period index, wherein the retention period index marks files that are about to expire.
[0009] In the preferred embodiment, the verification in step S7 includes the classification, volume assembly, and file number compliance verification of archived results, as well as the integrity and accuracy of the index and the associated logic verification. A verification report containing problem items and rectification suggestions is generated, and incremental index updates are triggered when new files are added or the associated rules change.
[0010] In the preferred embodiment, the output format of the standardized data packet in step S8 includes Excel, PDF, and JSON. The JSON format is compatible with the digital archive retrieval system, and the data packet can be directly used for physical binding and archiving, or imported into the digital archive, smart wind power platform, and digital twin power plant system with one click.
[0011] In the preferred solution, a digital archiving assistance platform based on a B / S architecture is used. The digital archiving assistance platform includes a user interaction layer, a business logic layer, a rule engine layer, and a data storage layer. The rule engine layer stores a dual-core knowledge base, and the data storage layer stores documents, metadata, relationships, and index results.
[0012] In the preferred embodiment, step 2 involves extracting text information from unstructured documents using OCR technology and automatically identifying basic metadata using an NLP model. In step S6, automatic index aggregation is achieved through keyword matching, ID association, and sorting algorithms, and the user interaction layer provides a visual display interface for the index.
[0013] Compared with existing technologies, the beneficial effects of this invention are as follows: This method replaces repetitive manual labor with automated processes and unifies management standards through a standardized framework. In the archiving stage, it significantly shortens the processing cycle, reduces reliance on professional experience, avoids human error and rework costs, and ensures consistent archiving results across different projects and personnel, while also reducing investment in new employee training. In the retrieval and usage stage, multi-dimensional indexing breaks the limitations of scattered and fragmented data connections, enabling precise positioning of target files across stages and types, providing efficient data support for troubleshooting and technical upgrade demonstrations. In full-cycle management, the standardized system reduces the cost of cross-project coordination and organization, adapts to digital platforms without secondary processing, and avoids redundant investment in "data reconstruction." In terms of compliance and traceability, it follows industry standards throughout the process, shortens the audit and inspection cycle, clarifies the responsible parties and time nodes for traceability, and improves traceability efficiency. Ultimately, it realizes the transformation of offshore wind power data management from "manual-led" to "intelligent-driven," fully activating the value of data reuse, laying a solid data foundation for the industry's digital upgrade, and the larger the project scale and the more complex the data, the more significant the overall improvement effect. Attached Figure Description
[0014] Figure 1 This is a flowchart of a method for auxiliary archiving of offshore wind power data. Detailed Implementation
[0015] Example 1 A method for auxiliary archiving of offshore wind power data, such as Figure 1 This includes the following steps: S1. Construct a dual-core knowledge base of "archiving rules + association relationships" to digitally model the archiving standards of the offshore wind power industry into structured archiving rules, while configuring explicit association rules and storing them as database tables; S2. Document Acquisition and "Index-Oriented" Metadata Capture: Supports batch uploading of multi-format offshore wind power project documents, and collects standardized metadata and completes verification through a combination of automatic extraction and manual supplementation. S3. Based on a dual-core knowledge base, intelligent identification and classification are achieved through document feature extraction and rule matching to determine the file's archiving category, retention period, archiving format, and associated tags. S4. Automatically assemble volumes according to volume / item rules, and solidify cross-stage technical associations, device-document bindings, hierarchical associations, and other related relationships; S5. Generate a standardized archiving catalog and a unique file number. The archiving catalog includes an "Archived List", a "Case File Catalog", and a "File Internal Directory". The file number follows a preset encoding format. S6. Automatically generate multi-dimensional indexes based on relationships and standardized metadata; S7, Human-computer interaction verification and index optimization, supports manual adjustment of archive and index results, and the system automatically completes integrity and consistency verification; S8 outputs standardized data packages, including electronic document sets, archive directories, multi-dimensional index results, and metadata and related relationship files, which are compatible with offline archiving and digital system import.
[0016] Preferably, the structured archiving rules of the dual-core knowledge base in step S1 include an archive classification table, a retention period table, a file number rule table, and a file type comparison table. The association rules include cross-stage technology association rules, device binding rules, hierarchical association rules, aggregated volume rules, and metadata cross-association rules. The rules can be added, deleted, modified, and queried through a graphical interface.
[0017] Preferably, the standardized metadata in step S2 includes project code / unit project code, equipment number, technical topic keywords, professional category, responsible person, project stage, and retention period. Among them, the equipment number corresponds to the equipment life cycle index, the technical topic keywords correspond to the technical process index, and the project code / unit project code and professional category correspond to the hierarchical classification index and the volume aggregation index.
[0018] Preferably, the document feature extraction in step S3 includes extracting file name, content keywords and standardized metadata. Rule matching achieves multi-dimensional matching through a pre-trained text classification model, and supports previewing the classification results and manual fine-tuning on the human-computer interaction interface.
[0019] Preferably, the file grouping rule for archiving by volume in step S4 is "same unit project + same specialty + same business type", and the files in the volume are arranged in the order of construction or time. The association relationship is solidified by recording the association IDs of files, files, equipment and technical topics, and the system automatically checks the integrity of the association and generates a verification report.
[0020] Preferably, the multi-dimensional index in step S6 includes a technical process index, a device lifecycle index, a hierarchical classification index, a volume aggregation index, a responsible entity index, a time dimension index, and a retention period index, wherein the retention period index marks files that are about to expire.
[0021] Preferably, the verification in step S7 includes verification of the classification, volume assembly, and file number compliance of archived results, as well as verification of the integrity of the index and the accuracy of the association logic. A verification report containing problem items and rectification suggestions is generated, and incremental index updates are triggered when new files are added or association rules change.
[0022] Preferably, the output format of the standardized data packet in step S8 includes Excel, PDF, and JSON. The JSON format is compatible with the digital archive retrieval system, and the data packet can be directly used for physical binding and archiving, or imported into the digital archive, smart wind power platform, and digital twin power plant system with one click.
[0023] Preferably, the method is implemented through a B / S architecture digital archiving auxiliary platform, which includes a user interaction layer, a business logic layer, a rule engine layer, and a data storage layer. The rule engine layer stores a dual-core knowledge base, and the data storage layer stores documents, metadata, relationships, and index results.
[0024] Preferably, in step S2, text information of unstructured documents is extracted using OCR technology, and basic metadata is automatically identified using an NLP model; in step S6, automatic index aggregation is achieved through "keyword matching + ID association + sorting algorithm", and the user interaction layer provides an index visualization display interface.
[0025] Example 2 This embodiment uses a 300MW offshore wind farm project as an application scenario. The project includes 25 6MW wind turbines and one offshore substation, requiring the archiving and management of over 100,000 documents, including survey and design records, equipment procurement records, civil construction records, installation and commissioning records, and operation and maintenance records, from the construction and production phases. This embodiment utilizes a B / S architecture-based digital archiving support platform (deployed on the project's private cloud, supporting 20 simultaneous online operations), fully implementing the core technical solutions of claims 1-10. The specific implementation process is as follows: I. Implementation Prerequisites: System Architecture and Core Support The digital archiving support platform upon which this method is based consists of four layers, and the functions and technical support of each layer are as follows: User interaction layer: Provides a web-based visual interface, including a document upload entry, metadata entry form, rule configuration panel, and index preview window (supporting tree / timeline views). Business logic layer: integrates Tesseract OCR engine (extracts text from scanned documents), BERT text classification model (document classification), volume grouping algorithm (aggregates files according to rules), and verification engine (integrity check); Rule engine layer: Stores a dual-core knowledge base of "archived rules + relationships", and uses a MySQL database to implement dynamic rule invocation; Data storage layer: MySQL is used to store metadata / relationships, MinIO is used to store electronic files (supporting PDF / DOC / DWG / XLS formats, etc.), and Elasticsearch is used to store index data (supporting fast retrieval).
[0026] II. Detailed Implementation Steps Step 1: Construct a dual-core knowledge base consisting of "archiving rules + association relationships" 1.1 Digital Modeling of Archiving Rules The "Scope of Filing and Retention Period of Wind Power Enterprise Documents and Materials" and the "Regulations on the Management of Archives for Power Construction Projects" are broken down into structured database tables. The core table design and examples are as follows: Table 1. Example of core table design for a structured database.
[0027] 1.2 Explicit Configuration of Association Rules A new "Association Rules Table" has been added, defining 5 types of core association logic, as shown in the following example: Cross-phase technology association rules: Keyword "wind turbine foundation" → mapping to "design drawings (survey period) + pouring records (construction period) + settlement monitoring data (operation and maintenance period)"; Equipment binding rules: Equipment number "WT-XX-08" (No. 8 fan) → associated with "Purchase contract (equipment) + Installation report (construction) + Operation log (maintenance) + Technical modification plan (update)"; Hierarchical association rules: Project code "XXFD-2023" → Unit project code "WT-08" (No. 8 wind turbine project) → Professional category "Civil Engineering" → Case file number "TS-008" → Document number "TS-008-001"; Aggregated document assembly rules: The document assembly conditions are "same unit project (WT-08) + same specialty (civil engineering) + same business type (acceptance)", and the document sorting is "time order (acceptance records first, then test reports)". Metadata cross-association rules: Responsible party "XX Construction Group" → aggregate "Construction Records, Acceptance Documents" type files.
[0028] 1.3 Visualized Rule Management After the "Regulations on the Management of Archives for Power Construction Projects" were updated, the administrator completed the modification of the "Equipment Ledger Retention Period from '20 Years' to '30 Years'" in the "Retention Period Table" in just 15 minutes through the platform's "Rule Configuration Interface". The system automatically synchronized the modification to the entire process rule call stage (without needing to restart the service).
[0029] Step 2: Document Acquisition and Index-Oriented Metadata Capture 2.1 Batch document upload The project team uploaded 52 documents related to the wind turbine foundation construction dated March 8, 2024, in batches through the platform's "Document Upload" module. These documents included: Structured documents: DOC version of "WT-08 Wind Turbine Foundation Pouring Construction Log", XLS version of "Concrete Test Block Test Report"; Unstructured documents: DWG version of "WT-08 Wind Turbine Foundation Reinforcement Drawing" (scanned copy), PDF version of "Supervisor's Acceptance Signature Scanned Copy".
[0030] 2.2 Standardized Collection of Metadata Automatic extraction: The scanned document was extracted using Tesseract OCR, showing "Document Date: 2024.03.15" and "Pages: 8". The file name / content was identified using the BERT-NLP model, showing "Equipment Number: WT-XX-08" and "Technical Topic Keywords: Wind Turbine Foundation Pouring". Manual completion: The platform pops up a metadata form, prompting the user to fill in "Responsible party: XX Construction Group Civil Engineering Division" and "Project stage: Construction period - Civil construction" (the system automatically matches "Professional category: Civil Engineering" and "Retention period: Permanent", and the user submits after confirming that it is correct). Metadata verification: The system automatically verifies the "equipment number format (must conform to 'WT-XX-XX')" and the "retention period and document type matching". It found that a "Cable Test Report" was incorrectly filled with "retention period: permanent", which triggered the prompt "Cable Test Report belongs to '30-year retention', do you want to correct it?". After the user confirmed, it was corrected.
[0031] Table 2. Example of final metadata collection results
[0032] Step 3: Intelligent Recognition and Classification (Association Rule Driven) 3.1 Document Feature Extraction The system extracts core features from the "WT-08 Wind Turbine Foundation Pouring Construction Log": File name characteristics: "WT-08", "Wind Turbine Foundation Pouring", "Construction Log"; Content characteristics: "C30 concrete", "pouring volume 520m³", "pouring time 2024.03.15-2024.03.16"; Metadata characteristics: Equipment number "WT-XX-08", professional category "civil engineering".
[0033] 3.2 Rule Matching Classification The BERT text classification model is matched with a dual-core knowledge base, and the following results are output: Archive category: Construction period - Civil construction - Construction records; Storage period: Permanent; Archiving format: By volume (meeting the criteria of "same unit project + same specialty + same business type"); Related tags: #windmill foundation #WT-08 #civil engineering construction #permanent storage
[0034] 3.3 Preview and Fine-tuning of Classification Results The platform displays the matching results on the "Categorization Results" page. The user found that one document, "WT-08 Wind Turbine Foundation Temporary Power Supply Plan," was misclassified as "Installation and Commissioning." The user manually adjusted it to "Civil Construction - Safety Plan." The system automatically updated the associated tags to #TemporaryPowerSupply#WT-08 #CivilConstruction.
[0035] Step 4: Automatic volume assembly and solidification of association relationships 4.1 File archives by volume According to the "Aggregation and File Combination Rules", the system aggregates the following 6 files into 1 virtual file: 1. WT-08 Wind Turbine Foundation Reinforcement Drawing (DWG); 2. WT-08 Wind Turbine Foundation Pouring Construction Log (DOC); 3. Concrete Specimen Test Report (XLS); 4. Scanned copy of the supervisor's acceptance signature (PDF); 5. WT-08 Temporary Power Supply Plan for Wind Turbine Foundations (DOC); 6. "WT-08 Wind Turbine Foundation Reinforcement Acceptance Record" (DOC); The volume title is automatically generated: "Civil Construction and Acceptance File of Wind Turbine Foundation No. 8 Project". The documents in the volume are ordered in the logical order of "Design → Construction → Acceptance" (i.e., 1→5→2→3→6→4 above).
[0036] 4.2 Solidification of Associations Cross-stage technology linkage: By using the keyword "wind turbine foundation", the "WT-08 Wind Turbine Foundation Geological Survey Report" (September 2023) during the survey period and the "WT-08 Wind Turbine Foundation Settlement Monitoring Monthly Report (July 2024)" during the operation and maintenance period are linked to form a technology chain of "survey → construction → operation and maintenance". Equipment-Document Binding: Using "WT-XX-08" as the core key, it links the equipment procurement period "No. 8 Fan Procurement Contract" (October 2023), the installation period "No. 8 Fan Lifting Report" (February 2024), the above-mentioned files during the construction period, and the operation and maintenance period "No. 8 Fan Operation Log (July 2024)", forming a complete equipment lifecycle archive package; Hierarchical association: Assign hierarchical identifiers to case files as “XXFD-2023 (project) → WT-08 (unit project) → civil engineering (specialty) → TS-008 (case file)”, and each document corresponds to a unique document identifier “TS-008-001~TS-008-006”.
[0037] 4.3 Validation of Association Results The system automatically generated an "Association Integrity Verification Report," which indicated that "the No. 8 fan equipment file package is missing the 'Equipment Factory Certificate of Conformity.'" After the user uploaded the missing information, the system automatically updated the equipment-document association and the verification passed again.
[0038] Step 5: Generate archive directory and unique file number 5.1 Archive Directory Generation Based on the paper compilation results, the system automatically generates three sets of core directories: Case File Catalog (Excel / PDF format): Includes case file number "TS-008", case file title "XX Project No. 8 Wind Turbine Foundation Civil Construction Acceptance Case File", professional category "Civil Engineering", retention period "Permanent", number of documents in the file "6", and file number "QZ-XXFD-2023-TS-008" (QZ is the project record number); "Document Directory" (Excel / PDF format): Records the document sequence number, title, responsible person, document date, page number, and document identifier within the volume. For example, sequence number 1 corresponds to "WT-08 Wind Turbine Foundation Reinforcement Drawing", XX Design Institute, 2023.09.10, 12 pages, TS-008-001; Archived List (applicable to documents archived individually): such as "Feasibility Study Report for Project XX" (archived separately), record the document number, title, responsible person, retention period, and file number "QZ-XXFD-2023-KY-001".
[0039] 5.2 Unique File Numbering File numbers strictly follow the format of "archive number-project code-classification number-file number / item number", for example: File number: QZ-XXFD-2023-TS-008; File number for archived items: QZ-XXFD-2023-KY-001 (KY is the classification number for "feasibility study").
[0040] Step 6: Automatic generation of multi-dimensional indexes Based on relationships and metadata, the system generates seven types of indexes using a "keyword matching + ID association + time sorting algorithm." A core example is shown below: Table 3 Core Examples of Multi-Dimensional Indexes
[0041] Step 7: Human-computer interaction verification and index optimization 7.1 Result Preview and Manual Adjustment Users discovered the following in the "Index Preview" interface: In the equipment lifecycle index, the "No. 8 Fan Installation Report" was mistakenly associated with "WT-09". After manually deleting it, it was reassigned to "WT-08". The technical process index is missing the "2024.04 Wind Turbine Foundation Lightning Protection Test Report". Manually supplement and upload it to trigger an incremental index update (only update the technical process index and equipment index, no need to regenerate the whole index).
[0042] 7.2 Integrity and Consistency Verification The system generates an "Archives and Index Verification Report," which states: Issue 1: The file number "QZ-XXFD-2023-TS-009" is duplicated (conflicting with the file number of the civil engineering case file for wind turbine No. 9). The user has changed it to "QZ-XXFD-2023-TS-009-01". Question 2: The "2024.03 Construction Documents" in the time dimension index was not sorted by date. The system automatically re-sorted it and the verification passed.
[0043] Step 8: Standardized Output and Multi-dimensional Delivery 8.1 Packaging of Output Materials The system integrates the following deliverables into a "Civil Construction Archive Data Package for Project XX, March 2024": 1. Electronic document set: sorted by "archive category → file / document", such as "construction period / civil construction / TS-008 file / TS-008-001 reinforcement drawing.dwg"; 2. Archive Catalog: Case File Catalog (Excel + PDF), File Contents Catalog (Excel + PDF), Archive List (Excel); 3. Multi-dimensional index: Excel version of 7 types of indexes (for offline viewing) and JSON version (adapted to digital archive search interface); 4. Metadata and Relationship Files: MySQL backup files (containing all file metadata and relational ID mappings).
[0044] 8.2 Multi-dimensional delivery Offline archiving: After downloading the data package, the project team prints the "Case File Catalog" and "Catalog of Documents in the File", binds them with the paper documents, and stores them in the project archives. Digital delivery: Through the platform's "one-click import" function, the JSON version of the index and electronic document set are imported into the "XX Province Wind Power Digital Archives System", and the archives retrieval system directly calls the multi-dimensional index; Cross-system linkage: The entire lifecycle index of the equipment is synchronized to the "XX Project Smart Operation and Maintenance Platform". When troubleshooting the No. 8 wind turbine, the operation and maintenance personnel can directly access the associated construction records and acceptance reports.
[0045] III. Implementation Results Verification This embodiment completes the archiving and indexing of 52 civil construction documents for Project XX in just 45 minutes (compared to 3 people / 2 days using traditional manual methods), with 100% archiving accuracy (no classification or file number errors). The multi-dimensional index improves subsequent retrieval efficiency by 70%. For example, if maintenance personnel are looking for "documents related to the foundation of Wind Turbine No. 8", they only need to enter "WT-08" in the equipment index to obtain data from the entire process of surveying, construction, and maintenance within 10 seconds, completely solving the pain points of traditional manual archiving, which is time-consuming, error-prone, and difficult to retrieve.
[0046] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for auxiliary archiving of offshore wind power data, characterized in that, Includes the following steps: S1. Construct a dual-core knowledge base, digitally model the offshore wind power industry archiving standards into structured archiving rules, and configure explicit association rules and store them as database tables. S2. Obtain documents and capture metadata. Support batch uploading of multi-format offshore wind power project documents. Collect standardized metadata and complete verification through a combination of automatic extraction and manual supplementation. S3. Based on a dual-core knowledge base, intelligent identification and classification are achieved through document feature extraction and rule matching to determine the file's archiving category, retention period, archiving format, and associated tags. S4. Automatically assemble volumes according to volume / item rules and solidify the association relationships; S5. Generate a standardized archive directory and unique file number; S6. Automatically generate multi-dimensional indexes based on relationships and standardized metadata; S7. Verify human-computer interaction and optimize indexes; the system automatically completes integrity and consistency checks. S8 outputs standardized data packages, adaptable to offline archiving and digital system import.
2. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, The structured archiving rules of the dual-core knowledge base in step S1 include an archive classification table, a retention period table, a file number rule table, and a file type comparison table. The association rules include cross-stage technology association rules, device binding rules, hierarchical association rules, aggregated volume rules, and metadata cross-association rules. The rules can be added, deleted, modified, and queried through a graphical interface.
3. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, The standardized metadata in step S2 includes project code / unit project code, equipment number, technical topic keywords, professional category, responsible party, project stage, and retention period. Among them, the equipment number corresponds to the equipment lifecycle index, the technical topic keywords correspond to the technical process index, and the project code / unit project code and professional category correspond to the hierarchical classification index and the volume aggregation index.
4. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, The document feature extraction in step S3 includes extracting file name, content keywords and standardized metadata. Rule matching achieves multi-dimensional matching through a pre-trained text classification model, and supports previewing the classification results and manual fine-tuning on the human-computer interaction interface.
5. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, The file grouping rule for archiving by volume in step S4 is "same unit project + same specialty + same business type". The files in the volume are arranged in the order of construction or time. The association relationship is solidified by recording the association ID of the file with the case file, equipment and technical subject. The system automatically checks the integrity of the association and generates a verification report.
6. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, The multi-dimensional index in step S6 includes a technical process index, a device lifecycle index, a hierarchical classification index, a volume aggregation index, a responsible entity index, a time dimension index, and a retention period index, among which the retention period index marks files that are close to their expiration date.
7. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, The verification in step S7 includes the classification, volume assembly, and file number compliance verification of archived results, as well as the integrity and accuracy of the index and related logic verification. It generates a verification report containing problem items and rectification suggestions, and supports triggering incremental index updates when new files are added or related rules change.
8. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, The output formats of the standardized data packets in step S8 include Excel, PDF, and JSON. The JSON format is compatible with digital archive retrieval systems. The data packets can be directly used for physical binding and archiving, or imported into digital archives, smart wind power platforms, and digital twin power plant systems with one click.
9. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, This is achieved through a B / S architecture-based digital archiving support platform, which includes a user interaction layer, a business logic layer, a rule engine layer, and a data storage layer. The rule engine layer stores a dual-core knowledge base, while the data storage layer stores documents, metadata, relationships, and index results.
10. The method for auxiliary archiving of offshore wind power data according to claim 1, characterized in that, In step 2, text information of unstructured documents is extracted using OCR technology, and basic metadata is automatically identified by combining it with an NLP model. In step S6, automatic index aggregation is achieved through keyword matching, ID association, and sorting algorithms, and the user interaction layer provides a visual display interface for the index.
Citation Information
Cited By
Meteorological archive filing system and method
CN122220306A