Computer data storage method and system

Through multimodal feature extraction and glyph topology analysis, combined with tree-like namespace and blockchain index table, the problem of lack of active perception of data semantics in the power system is solved, efficient isolated storage and accurate retrieval of power drawings are achieved, and the stability of data management and query efficiency are improved.

CN120653711AInactive Publication Date: 2025-09-16YANTAI NANSHAN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510472188.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing storage technologies lack the ability to actively perceive data semantics in power systems, resulting in deep contradictions in data organization logic and poor retrieval system performance. When mixed storage of unstructured and structured data occurs, the logical associations between the contents are ignored, causing a data swamp effect and affecting data extraction efficiency.

Method used

Multimodal feature extraction and glyph topology analysis are adopted, and convolutional neural networks are used to identify vector graphic elements and text symbols in power drawings to generate feature vectors. A tree-like namespace structure and dynamic term mapping module are used for data routing and verification. Combined with blockchain index tables and hybrid storage media, efficient isolated storage and retrieval of data are achieved.

Benefits of technology

It realizes the detailed modeling and efficient retrieval of power drawings, improves the accuracy of data storage and query efficiency, ensures the stability of the system and the security of data traceability, and meets the power system's real-time management needs for large-scale unstructured data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653711A_ABST
    Figure CN120653711A_ABST
Patent Text Reader

Abstract

The invention discloses a computer data storage method and system, and relates to the technical field of computer data storage.The computer data storage method comprises the steps that variant samples are generated through multi-modal feature extraction, font topology analysis and adversarial training, and fine modeling of Chinese character connection and pixel-level interference in an electric power drawing is achieved; an isolated storage pool path is dynamically generated through a tree-shaped namespace, and a physical storage address and version information can efficiently and uniquely correspond and retrieve in combination with bidirectional mapping of version identifier construction and block chain indexes; according to a confidence score obtained through weighted calculation of the topological similarity and the term library matching degree, once the confidence score is lower than a preset threshold value, the system pushes a manual recheck mark to a specified terminal through an asynchronous message queue, and meanwhile, read-only data locking is implemented to ensure that the writing process is not interrupted; the consensus management mechanism is based on an improved practical Byzantine fault-tolerant algorithm, and block chain index consistency and data traceability security are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data storage, and in particular to a computer data storage method and system. Background Art

[0002] In the digital transformation of power systems, data lakes are becoming a core storage architecture, integrating diverse data such as meter readings, equipment logs, and design drawings. While existing storage technologies address the aggregation of massive amounts of data, they pose risks in data association and semantic consistency. In power grid data lakes, operations and maintenance personnel often need to retrieve substation design drawings and real-time monitoring logs simultaneously for fault analysis. However, during actual retrieval, a CAD drawing labeled "10kV switchgear" may be completely out of context due to subtle semantic deviations in the storage process. This problem stems not from storage capacity or read / write speeds, but rather from underlying contradictions within the data organization logic. Traditional data storage methods rely on physical features such as file extensions and directory structures to manage data. For example, the design department might store CAD drawings by voltage level directory, and archive maintenance logs by timestamp. While this classification appears clear and reasonable, when annotation text within a CAD drawing, such as an allowable deviation of ±0.5%, is mistakenly recognized by an OCR tool as ±0.5%, the actual file content becomes disconnected from the directory label. Worse still, these errors are not intercepted during storage, but instead flow through the data, contaminating downstream analysis systems and easily misidentifying equipment as out-of-tolerance, leading to unnecessary downtime and maintenance. Some existing solutions use strengthened rule constraints to alleviate the problem, such as requiring manual addition of metadata tags such as voltage level and accuracy range when uploading drawings. However, in actual operation, the subjectivity of manual labeling becomes a new problem. Different engineers mark the same drawing as 10kV, 10KV or ten kilovolts. This non-standardized expression makes the keyword-based retrieval system ineffective. The deeper problem is that traditional storage architecture lacks the ability to actively perceive data semantics. When unstructured data, such as drawing scans, is mixed with structured data such as sensor logs, the system only focuses on the byte storage location, but ignores the logical relationship between the content. This ultimately leads to the data swamp effect. The larger the amount of data, the more difficult it is to extract effective information. Summary of the Invention

[0003] In view of the above existing problems, the present invention is proposed.

[0004] The present invention provides a computer data storage method and system to solve the problem in the power grid data lake that current storage technology focuses on physical storage efficiency but sacrifices data semantic integrity.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, an embodiment of the present invention provides a computer data storage method, which includes: Step S1: Receive a mixed input stream of CAD files containing power drawings and structured logs through the heterogeneous data access module, and trigger the multimodal classification module to perform feature extraction on the input data; Step S2: The graphic feature extraction unit identifies vector graphic elements in the input data based on a pre-trained convolutional neural network model, and the text feature extraction unit parses text symbols in the data using a glyph topology analysis model to generate a feature vector containing the graphic type, text semantics, and version identifier. Step S3: Routing the power drawings to an isolated storage pool based on the graphic type identifier in the feature vector. The isolated storage pool adopts a tree-like namespace structure, and its node path is dynamically generated by the voltage level code and the equipment type code. Step S4: Deploy a dynamic terminology mapping module at the entrance of the isolated storage pool. This module has a built-in power standard terminology library and is associated with a real-time updated on-site common terminology library to perform terminology consistency verification on written power drawings. Step S5: Add a version identifier to the verified power drawings. The version identifier is generated by combining the design tool type, timestamp, and hash value, and is synchronously recorded in the blockchain index table.

[0006] As a preferred solution of the computer data storage method described in the present invention, wherein: in step S2, the glyph topology analysis model is generated through adversarial training, and the training data includes multiple variant samples of electric power symbols, and the variant samples cover letter case confusion, handwriting and printed mixed input scenarios, and the model output includes the symbol's stroke connection relationship matrix and topological similarity score.

[0007] As a preferred solution of the computer data storage method of the present invention, the method for generating a node path of the tree-like namespace structure in step S3 includes: According to the voltage level coding rules, voltage levels including 10kV, 35kV, and 110kV are mapped into three-level digital codes; The equipment type code adopts a composite coding form, which includes the equipment function code and the safety level code. The safety level code is dynamically calculated based on the topological position of the equipment in the power system.

[0008] As a preferred solution of the computer data storage method of the present invention, the term consistency check in step S4 includes: Normalize the symbols of text annotations in unstructured power drawings and map unconventional expressions to standard terminology libraries; Generate a correction log and append it to the data metadata segment, wherein the correction log records the original expression, standard terminology, and confidence score; When the confidence score is lower than the threshold, a manual review mark is triggered without interrupting the data writing process.

[0009] As a preferred embodiment of the computer data storage method described in the present invention, the blockchain index table in step S5 adopts a lightweight consortium chain structure, and its consensus mechanism is based on an improved practical Byzantine fault-tolerant algorithm. Participating nodes include design department terminals, operation and maintenance servers, and a data lake management platform. The version identifier is written into the block as transaction content and a bidirectional index is established with the physical storage address in the isolated storage pool. The improved practical Byzantine fault-tolerant algorithm introduces a time window sharding mechanism. Each consensus cycle is divided into a preparation phase, a proposal phase, and a confirmation phase. In the proposal phase, only nodes with the latest version identifier are allowed to initiate proposals. The confirmation phase adopts the majority rule and requires verification of the hash consistency of the version identifier.

[0010] As a preferred solution of the computer data storage method described in the present invention, the physical storage structure of the isolated storage pool adopts a hybrid storage medium, wherein the high-frequency access power drawings are stored in the phase change memory partition, and the low-frequency data is stored in the three-dimensional stacked flash memory partition. The phase change memory partition realizes storage unit mode switching through laser-induced phase change technology and supports SLC / MLC hybrid encoding.

[0011] In a second aspect, the present invention provides a computer data storage system comprising: The heterogeneous data access module is used to receive mixed input streams including CAD files of electrical drawings and structured logs, and transmit the input data to the multimodal classification module; A multimodal classification module, connected to the heterogeneous data access module, comprising a graphic feature extraction unit and a text feature extraction unit; The graphic feature extraction unit is configured to extract vector graphic elements from input data through a pre-trained convolutional neural network model and identify the boundary between the device graphic element and the annotation text; The text feature extraction unit is configured to analyze the stroke connection relationship of the text symbol through the glyph topology analysis model to generate a semantic feature including a symbol topology similarity score; a storage routing control module, connected to the multimodal classification module, dynamically generating a tree-like namespace path based on the device type code and voltage level code output by the graphic feature extraction unit, and routing the power drawings to the isolated storage pool; A dynamic term mapping module, deployed on the write interface of the isolated storage pool, having a built-in power standard terminology library and a field common name library, configured to perform term normalization processing on input power drawings, generate a correction log, and write it into the data metadata segment; a version control module, connected to the isolated storage pool and the blockchain index table, configured to attach a version identifier to the verified power drawings, where the version identifier is generated by combining the design tool type, timestamp, and hash value, and synchronously write the identifier to the lightweight consortium chain node of the blockchain index table; A hybrid storage medium module, including a phase-change memory partition and a three-dimensional stacked flash memory partition. The phase-change memory partition uses laser-induced phase change technology to support SLC / MLC hybrid encoding, while the three-dimensional stacked flash memory partition uses QLC mode to store low-frequency data. The laser-induced phase change technology in the hybrid storage medium module specifically includes: a laser parameter control unit configured to dynamically adjust the laser wavelength and pulse width according to the data access frequency, wherein the high-frequency data region uses a 1064nm wavelength and a 10ns pulse, and the low-frequency data region uses a 1550nm wavelength and a 50ns pulse; a phase change mode switching unit configured to control the crystalline state of the storage unit via an energy density threshold, the threshold being dynamically calibrated according to a temperature resistance characteristic curve of the phase change material; The consensus management module is connected to the blockchain index table and adopts an improved practical Byzantine fault tolerance algorithm with a time window sharding mechanism. It is configured to screen nodes with the latest version identifier to initiate proposals in the proposal phase and verify hash consistency in the confirmation phase.

[0012] As a preferred solution of the computer data storage system of the present invention, the dynamic term mapping module further includes: A term normalization submodule is configured to map non-standard expressions in power drawings to a standard terminology library, wherein the mapping rule is based on a weighted calculation of the glyph topology similarity score and the terminology library matching degree; The correction log generation submodule is configured to record the original expression, standard terminology, and confidence score. When the confidence score falls below a dynamic threshold, it triggers an asynchronous message queue to push a manual review request to a designated terminal. The data locking submodule is configured to lock the relevant data in read-only mode during manual review and append a temporary locking identifier to the meta-information segment.

[0013] As a preferred solution of the computer data storage system of the present invention, the generation rules of the tree-like namespace path include: The voltage level coding adopts three-level digital coding, where 10kV is mapped to 001, 35kV is mapped to 002, and 110kV is mapped to 003; The equipment type code adopts the composite form of function code and safety level code. The safety level code is generated based on the calculation result of the betweenness centrality of the equipment in the power grid topology diagram, and the betweenness threshold is dynamically optimized according to the historical failure rate.

[0014] As a preferred solution of the computer data storage system described in the present invention, the time window sharding mechanism of the consensus management module specifically includes: Preparation phase: Each node synchronizes the hash chain of the latest version identifier, and the timestamp deviation does not exceed 1 consensus cycle; Proposal stage: Only nodes that have submitted valid version identifiers within three consecutive blocks are allowed to initiate proposals; Confirmation phase: The majority rule is used to verify the consistency between the version identifier hash of the proposal node and the locally stored blockchain index table.

[0015] The beneficial effects of the present invention are as follows: the present invention utilizes multimodal feature extraction, glyph topology analysis, and adversarial training to generate variant samples, thereby realizing fine modeling of the ligatures and pixel-level interference of Chinese characters in electric power drawings; dynamically generates isolated storage pool paths through a tree-like namespace, and combines version identifier construction with bidirectional mapping of blockchain indexes, so that physical storage addresses and version information can be efficiently and uniquely mapped and retrieved; the confidence score obtained by weighted calculation based on topological similarity and terminology library matching, once it is lower than a preset threshold, the system will push a manual review mark to the designated terminal through an asynchronous message queue, and implement read-only data locking at the same time to ensure that the writing process is not interrupted; the consensus management mechanism is based on an improved practical Byzantine fault-tolerant algorithm to ensure blockchain index consistency and data traceability security.

[0016] The present invention overcomes the limitations of traditional storage technologies in processing diversified power data, improves the accuracy of data storage, query efficiency and system stability, and meets the power system's needs for efficient and real-time management of large-scale unstructured data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 Schematic diagram of the process of computer data storage method in Example 1.

[0019] Figure 2 This is a schematic diagram of the framework of the computer data storage system in Example 1. DETAILED DESCRIPTION

[0020] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0022] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0023] Example 1, with reference to Figure 1 and Figure 2 , this embodiment provides a computer data storage method, comprising the following steps: Step S1: Receive a mixed input stream of CAD files containing power drawings and structured logs through the heterogeneous data access module, and trigger the multimodal classification module to perform feature extraction on the input data; Step S2: The graphic feature extraction unit identifies vector graphic elements in the input data based on a pre-trained convolutional neural network model, and the text feature extraction unit parses text symbols in the data using a glyph topology analysis model to generate a feature vector containing the graphic type, text semantics, and version identifier. In step S2, a glyph topology analysis model is generated through adversarial training. The training data includes multiple variant samples of the power symbol. The variant samples cover scenarios where uppercase and lowercase letters are mixed and handwritten and printed characters are mixed. The model output includes the symbol's stroke connection relationship matrix and topological similarity score. In step S2, variant samples are generated: Suppose we construct a transformation function for the connecting effect of letters k and K, and smoothly fuse the original glyph matrix. The formula is: ,in, Represents the glyph image matrix output after connecting the strokes. is the degree of connection coefficient, which is used to control the ratio of the two glyphs. is the graphic matrix of the original lowercase letter k, Represents the translation transformation operator, which is used to simulate stroke connection. The parameters are Determine is the graphic matrix of the original capital letter K; Translation amount parameter Adopt the calculation of the starting and ending positions of key strokes. The calculation formula is: , where represents the translation amount in the translation operation, is the calibration coefficient used to adjust the translation sensitivity, is the ending abscissa of the key stroke, is the starting abscissa of the key stroke; Adopt a weighted fusion model to simulate the pixel-level superposition interference between the symbol ± and the Chinese character 土. The model formula is: , where represents the pixel matrix of the generated variant sample after fusion, is the superposition weight, which controls the contribution ratio of the pixel information of the two, is the pixel matrix of the symbol ±, is the pixel matrix of the Chinese character 土; Adopt a normalized mapping function to describe the superposition weight , and the formula is: , where represents the normalization function that maps the input value to the weight range, is the original pixel difference index, which reflects the difference between the two graphics at the pixel level; Specifically, in the simulation of connected strokes, a weighted translation fusion method is adopted. Through translation transformation, the capital letter and the lowercase letter are smoothly connected in strokes, thus presenting a connected stroke effect. In the pixel-level superposition interference, the pixel matrices of two different graphics are weighted and fused through weights, introducing local interference to simulate the mixed input scenario; Step S3: According to the graphic type identifier in the feature vector, route the power drawing to the isolation storage pool. The isolation storage pool adopts a tree-like namespace structure, and its node path is dynamically generated by the voltage level code and the device type code; The method for generating the node path of the tree-like namespace structure in step S3 includes: According to the voltage level coding rule, map the voltage levels including 10kV, 35kV, and 110kV to three-level digital codes; The device type code adopts a composite coding form, including the device function code and the safety level code, where the safety level code is dynamically calculated according to the topological position of the device in the power system; Step S4: Deploy a dynamic terminology mapping module at the entrance of the isolated storage pool. This module has a built-in power standard terminology library and is associated with a real-time updated on-site common terminology library to perform terminology consistency verification on written power drawings. The term consistency check in step S4 includes: Normalize the symbols of text annotations in unstructured power drawings and map unconventional expressions to standard terminology libraries; Generate a correction log and append it to the data metadata segment. The correction log records the original expression, standard terminology, and confidence score. When the confidence score is lower than the threshold, a manual review mark is triggered without interrupting the data writing process; In step S4, the method of triggering the manual review mark without interrupting the data writing process is: The confidence score is calculated using a weighted formula described as: ,in, represents the confidence score, which is used to quantify the consistency of the current data symbol mapping, is the weight coefficient, whose value depends on the type of power equipment and balances the contribution of topological similarity and terminology matching. Represents the topological similarity score, reflecting the structural consistency of the glyph symbols, Indicates the terminology matching degree, reflecting the matching between the drawing text and the standard terminology database; Weight coefficient Dynamic adjustment is performed based on the device type, and the calculation formula is defined as: , in, Numerical code indicating the type of electrical equipment, used to distinguish the characteristics of different equipment. is a scaling factor used to adjust the sensitivity of the device type to the weight adjustment, ensuring Always within the interval (0,1), so that the weight distribution has smoothness and dynamic responsiveness; In the comparison threshold judgment part, if the confidence score is lower than the preset threshold, the manual review mark is triggered. The judgment condition is expressed as: like ,but , like ,but ; in, It is a trigger indicator variable, and its value is 1, indicating that manual review is required. Indicates the preset confidence score threshold, which is used to determine whether the data meets the automatic writing standard; In satisfaction Under these conditions, the review mark information is pushed to the designated terminal through the asynchronous message queue module, and the corresponding data in the data lake is locked in read-only mode. The push operation is expressed as: ,in, Represents an asynchronous message queue, used to send review notifications. Indicates the target terminal identifier, indicating the receiving device of the review request. Indicates review mark information, including relevant prompts for data anomalies; The data locking operations are: ,in, Indicates data lock status, Designated as read-only mode to ensure that data is not overwritten during manual review, thereby achieving the requirement of uninterrupted data writing process; Specifically, by constructing a confidence score The mathematical model uses the dynamic weight coefficient of topological similarity and terminology matching based on the device type. Perform weighted fusion to accurately reflect the consistency of data semantics. When the score is lower than the threshold When a request is received, the asynchronous message queue is triggered by conditional judgment to send a review mark to the designated terminal. At the same time, the corresponding data in the data lake is locked for read-only, ensuring that abnormal data can be manually reviewed in a timely manner. At the same time, the continuity of the data writing process and the stability of the system are maintained, effectively balancing automation and manual intervention. Step S5: Add a version identifier to the verified power drawings. The version identifier is generated by combining the design tool type, timestamp, and hash value, and is synchronously recorded in the blockchain index table. In step S5, the blockchain index table adopts a lightweight consortium chain structure, and its consensus mechanism is based on an improved practical Byzantine fault-tolerant algorithm. Participating nodes include the design department terminal, operation and maintenance server, and data lake management platform. The version identifier is written into the block as transaction content and a bidirectional index is established with the physical storage address in the isolated storage pool. The improved practical Byzantine fault tolerance algorithm introduces a time window sharding mechanism. Each consensus cycle is divided into a preparation phase, a proposal phase, and a confirmation phase. In the proposal phase, only nodes with the latest version identifier are allowed to initiate proposals. The confirmation phase adopts the majority rule and requires verification of the hash consistency of the version identifier. The process of establishing a bidirectional index in step S5 is as follows: First, define blockchain transaction records as: ,in, Represents a transaction record, whose content is written into the blockchain. Represents the version identifier, which is generated by combining the design tool type, timestamp, and hash value. Indicates the physical storage address in the isolated storage pool. Represents a reverse hash pointer, used to point from the physical storage address to the version identifier; Define the reverse hash pointer, set ,in, represents a reverse hash pointer, Represents the reverse hash function, used to map the version identifier to reverse pointer information, Represents a version identifier; Use Bloom filter to construct index, assuming that Bloom filter is: ,in, Represents a Bloom filter, whose length is Bit, Indicates the length of the bit array of the Bloom filter; use A hash function inserts the version identifier into the Bloom filter, and the generated hash set is: ,in, Indicates the A hash function is used to generate the corresponding index bits, represents the number of hash functions, Represents a version identifier; When querying, calculate the version identifier Hash value, and check whether the corresponding bit of the Bloom filter is 1. The judgment formula is: , in, Represents the version identifier If all corresponding bits are 1, it means may exist in the index, It is an indicator function that returns 1 if the input condition is met, otherwise it returns 0. Indicates the Hash function pair The calculation results are: represents the number of hash functions, Represents a version identifier; In the blockchain transaction records, Directly mapped to physical storage addresses in the isolated storage pool , while the metadata field of the physical storage address records the reverse hash pointer , thereby achieving arrive With Reverse Lookup Bidirectional index of ; Specifically, by constructing blockchain transaction records, the version identifier, physical storage address, and reverse hash pointer are organically combined to achieve bidirectional indexing. Using Bloom filters to perform hash insertion and query on the version identifier, the index search process is greatly accelerated. Bloom filters can also quickly determine the existence of data in large-scale data with minimal space consumption. Reverse hash pointers are embedded in the metadata of the physical storage address, allowing for reverse lookup of the corresponding version information when accessing physical data. The physical storage structure of the isolated storage pool uses hybrid storage media, in which high-frequency access power drawings are stored in phase change memory partitions, and low-frequency data are stored in three-dimensional stacked flash memory partitions. The phase change memory partition uses laser-induced phase change technology to achieve storage unit mode switching and supports SLC / MLC hybrid encoding.

[0024] This embodiment also provides an application system of the above-mentioned computer data storage method, including the following modules: The heterogeneous data access module is used to receive mixed input streams including CAD files of electrical drawings and structured logs, and transmit the input data to the multimodal classification module; The multimodal classification module is connected to the heterogeneous data access module and includes a graphic feature extraction unit and a text feature extraction unit; The graphic feature extraction unit is configured to extract vector graphic elements from input data through a pre-trained convolutional neural network model and identify boundaries between device graphics elements and annotation text; The text feature extraction unit is configured to analyze the stroke connection relationship of the text symbol through the glyph topology analysis model to generate a semantic feature including a symbol topology similarity score; The storage routing control module is connected to the multimodal classification module. It dynamically generates a tree-like namespace path based on the device type code and voltage level code output by the graphic feature extraction unit, and routes the power drawings to the isolated storage pool. The dynamic term mapping module is deployed in the write interface of the isolated storage pool. It has a built-in power standard terminology library and a field common name library. It is configured to perform terminology normalization on input power drawings, generate correction logs, and write them into the data metadata segment. The dynamic term mapping module further includes: A term normalization submodule is configured to map non-standard expressions in power drawings to a standard terminology library, wherein the mapping rule is based on a weighted calculation of the glyph topology similarity score and the terminology library matching degree; The correction log generation submodule is configured to record the original expression, standard terminology, and confidence score. When the confidence score falls below a dynamic threshold, it triggers an asynchronous message queue to push a manual review request to a designated terminal. A data locking submodule configured to lock relevant data in read-only mode during manual review and append a temporary lock flag to the meta-information segment; The rules for generating a tree namespace path include: The voltage level coding adopts three-level digital coding, where 10kV is mapped to 001, 35kV is mapped to 002, and 110kV is mapped to 003; The equipment type code is a composite of a function code and a safety level code. The safety level code is generated based on the betweenness centrality calculation results of the equipment in the power grid topology diagram. The betweenness threshold is dynamically optimized based on the historical failure rate. A version control module, connected to the isolated storage pool and the blockchain index table, is configured to attach a version identifier to the verified power drawings. The version identifier is generated by combining the design tool type, timestamp, and hash value, and is synchronously written to the lightweight consortium chain node of the blockchain index table. A hybrid storage medium module, including a phase-change memory partition and a three-dimensional stacked flash memory partition. The phase-change memory partition uses laser-induced phase change technology to support SLC / MLC hybrid encoding, while the three-dimensional stacked flash memory partition uses QLC mode to store low-frequency data. The laser-induced phase change technology in the hybrid storage medium module specifically includes: a laser parameter control unit configured to dynamically adjust the laser wavelength and pulse width according to the data access frequency, wherein the high-frequency data region uses a 1064nm wavelength and a 10ns pulse, and the low-frequency data region uses a 1550nm wavelength and a 50ns pulse; a phase change mode switching unit configured to control the crystalline state of the storage unit via an energy density threshold, the threshold being dynamically calibrated based on a temperature resistance characteristic curve of the phase change material; The consensus management module is connected to the blockchain index table and adopts an improved practical Byzantine fault tolerance algorithm with a time window sharding mechanism. It is configured to filter nodes with the latest version identifier to initiate proposals during the proposal phase and verify hash consistency during the confirmation phase. The time window sharding mechanism of the consensus management module specifically includes: Preparation phase: Each node synchronizes the hash chain of the latest version identifier, and the timestamp deviation does not exceed 1 consensus cycle; Proposal stage: Only nodes that have submitted valid version identifiers within three consecutive blocks are allowed to initiate proposals; Confirmation phase: The majority rule is used to verify the consistency between the version identifier hash of the proposal node and the locally stored blockchain index table.

[0025] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A computer data storage method, characterized in that: include, Step S1: Receive a mixed input stream of CAD files containing power drawings and structured logs through the heterogeneous data access module, and trigger the multimodal classification module to perform feature extraction on the input data; Step S2: The graphic feature extraction unit identifies vector graphic elements in the input data based on a pre-trained convolutional neural network model, and the text feature extraction unit parses text symbols in the data using a glyph topology analysis model to generate a feature vector containing the graphic type, text semantics, and version identifier. Step S3: Routing the power drawings to an isolated storage pool based on the graphic type identifier in the feature vector. The isolated storage pool adopts a tree-like namespace structure, and its node path is dynamically generated by the voltage level code and the equipment type code. Step S4: Deploy a dynamic terminology mapping module at the entrance of the isolated storage pool. This module has a built-in power standard terminology library and is associated with a real-time updated on-site common terminology library to perform terminology consistency verification on written power drawings. Step S5: Add a version identifier to the verified power drawings. The version identifier is generated by combining the design tool type, timestamp, and hash value, and is synchronously recorded in the blockchain index table.

2. A computer data storage method according to claim 1, characterized in that: In step S2, the glyph topology analysis model is generated through adversarial training. The training data includes multiple variant samples of electric power symbols. The variant samples cover letter case confusion and mixed input scenarios of handwriting and print. The model output includes the symbol's stroke connection relationship matrix and topological similarity score.

3. A computer data storage method according to claim 1, characterized in that: The node path generation method of the tree-like namespace structure in step S3 includes: According to the voltage level coding rules, voltage levels including 10kV, 35kV, and 110kV are mapped into three-level digital codes; The equipment type code adopts a composite coding form, which includes the equipment function code and the safety level code. The safety level code is dynamically calculated based on the topological position of the equipment in the power system.

4. A computer data storage method according to claim 1, characterized in that: The term consistency check in step S4 includes: Normalize the symbols of text annotations in unstructured power drawings and map unconventional expressions to standard terminology libraries; Generate a correction log and append it to the data metadata segment, wherein the correction log records the original expression, standard terminology, and confidence score; When the confidence score is lower than the threshold, a manual review mark is triggered without interrupting the data writing process.

5. A computer data storage method according to claim 1, characterized in that: The blockchain index table in step S5 adopts a lightweight consortium chain structure, and its consensus mechanism is based on an improved practical Byzantine fault-tolerant algorithm. Participating nodes include design department terminals, operation and maintenance servers, and the data lake management platform. The version identifier is written into the block as transaction content and establishes a bidirectional index with the physical storage address in the isolated storage pool; The improved practical Byzantine fault-tolerant algorithm introduces a time window sharding mechanism. Each consensus cycle is divided into a preparation phase, a proposal phase, and a confirmation phase. In the proposal phase, only nodes with the latest version identifier are allowed to initiate proposals. The confirmation phase adopts the majority rule and requires verification of the hash consistency of the version identifier.

6. A computer data storage method according to claim 5, characterized in that: The physical storage structure of the isolated storage pool adopts a hybrid storage medium, in which high-frequency access power drawings are stored in phase change memory partitions, and low-frequency data are stored in three-dimensional stacked flash memory partitions. The phase change memory partitions realize storage unit mode switching through laser-induced phase change technology and support SLC / MLC hybrid encoding.

7. A computer data storage system, based on a computer data storage method according to any one of claims 1 to 6, characterized in that: include: The heterogeneous data access module is used to receive mixed input streams including CAD files of electrical drawings and structured logs, and transmit the input data to the multimodal classification module; A multimodal classification module, connected to the heterogeneous data access module, comprising a graphic feature extraction unit and a text feature extraction unit; The graphic feature extraction unit is configured to extract vector graphic elements from input data through a pre-trained convolutional neural network model and identify the boundary between the device graphic element and the annotation text; The text feature extraction unit is configured to analyze the stroke connection relationship of the text symbol through the glyph topology analysis model to generate a semantic feature including a symbol topology similarity score; a storage routing control module, connected to the multimodal classification module, dynamically generating a tree-like namespace path based on the device type code and voltage level code output by the graphic feature extraction unit, and routing the power drawings to the isolated storage pool; A dynamic term mapping module, deployed on the write interface of the isolated storage pool, having a built-in power standard terminology library and a field common name library, configured to perform term normalization processing on input power drawings, generate a correction log, and write it into the data metadata segment; a version control module, connected to the isolated storage pool and the blockchain index table, configured to attach a version identifier to the verified power drawings, where the version identifier is generated by combining the design tool type, timestamp, and hash value, and synchronously write the identifier to the lightweight consortium chain node of the blockchain index table; A hybrid storage medium module, including a phase-change memory partition and a three-dimensional stacked flash memory partition. The phase-change memory partition uses laser-induced phase change technology to support SLC / MLC hybrid encoding, while the three-dimensional stacked flash memory partition uses QLC mode to store low-frequency data. The laser-induced phase change technology in the hybrid storage medium module specifically includes: a laser parameter control unit configured to dynamically adjust the laser wavelength and pulse width according to the data access frequency, wherein the high-frequency data region uses a 1064nm wavelength and a 10ns pulse, and the low-frequency data region uses a 1550nm wavelength and a 50ns pulse; a phase change mode switching unit configured to control the crystalline state of the storage unit via an energy density threshold, the threshold being dynamically calibrated according to a temperature resistance characteristic curve of the phase change material; The consensus management module is connected to the blockchain index table and adopts an improved practical Byzantine fault tolerance algorithm with a time window sharding mechanism. It is configured to screen nodes with the latest version identifier to initiate proposals in the proposal phase and verify hash consistency in the confirmation phase.

8. A computer data storage system according to claim 7, characterized in that: The dynamic term mapping module further includes: A term normalization submodule is configured to map non-standard expressions in power drawings to a standard terminology library, wherein the mapping rule is based on a weighted calculation of the glyph topology similarity score and the terminology library matching degree; The correction log generation submodule is configured to record the original expression, standard terminology, and confidence score. When the confidence score falls below a dynamic threshold, it triggers an asynchronous message queue to push a manual review request to a designated terminal. The data locking submodule is configured to lock the relevant data in read-only mode during manual review and append a temporary locking identifier to the meta-information segment.

9. A computer data storage system according to claim 7, characterized in that: The generation rules of the tree-like namespace path include: The voltage level coding adopts three-level digital coding, where 10kV is mapped to 001, 35kV is mapped to 002, and 110kV is mapped to 003; The equipment type code adopts the composite form of function code and safety level code. The safety level code is generated based on the calculation result of the betweenness centrality of the equipment in the power grid topology diagram, and the betweenness threshold is dynamically optimized according to the historical failure rate.

10. A computer data storage system according to claim 7, characterized in that: The time window sharding mechanism of the consensus management module specifically includes: Preparation phase: Each node synchronizes the hash chain of the latest version identifier, and the timestamp deviation does not exceed 1 consensus cycle; Proposal stage: Only nodes that have submitted valid version identifiers within three consecutive blocks are allowed to initiate proposals; Confirmation phase: The majority rule is used to verify the consistency between the version identifier hash of the proposal node and the locally stored blockchain index table.

Citation Information

Cited By

  • Method and system for intercepting portable device registry creation in operating system

    CN121234361A

  • Digital base system

    CN121689559A

  • Intelligent verification method for consistency of electric power design engineering drawings

    CN121811446A

  • An electric power design engineering drawing consistency intelligent checking method

    CN121811446B