Multi-source heterogeneous data security storage method for smart engineering
Patent Information
- Application Number
- CN202610733719.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]针对现有技术的不足,本发明提供了面向智慧工程的多源异构数据安全存储方法,解决了多源异构数据兼容性适配不足以及存储策略同质化,安全与效率失衡的问题
本发明针对结构化数据,采用样本统计+正则匹配双机制自动识别字段类型,解决数据库差异、字段命名混乱导致的兼容问题,针对半结构化数据,通过深度优先遍历+领域知识图谱联动自动提取核心字段,提升适配效率,针对非结构化数据,精准提取业务关联元数据,生成统一结构化索引表,提高数据孤岛破解率。
Smart Images

Figure CN122839416A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage technology, specifically to a method for secure storage of multi-source heterogeneous data for smart engineering projects. Background Technology
[0002] With the deep integration of new-generation information technology and engineering construction, smart engineering has become the mainstream of development in the engineering construction industry. The entire life cycle of smart engineering will generate massive amounts of multi-source heterogeneous data, and its secure storage and processing are directly related to the safety, economy and compliance of engineering construction.
[0003] Existing technologies mostly adopt generalized data processing solutions, lacking customized adaptation mechanisms for smart engineering scenarios. This results in severe data silos, making it impossible to achieve cross-type and cross-system data correlation analysis. Secondly, the use of a single storage architecture to process all levels of data without designing differentiated storage solutions for data security levels leads to resource waste. Finally, encryption technologies suffer from single strategies, coarse granularity, and chaotic key management, making it difficult to balance encryption security and data usage efficiency. Furthermore, encryption, storage, and classification / grading strategies lack linkage, failing to form a full-link security protection system. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a secure storage method for multi-source heterogeneous data in smart engineering, which solves the problems of insufficient compatibility and adaptation of multi-source heterogeneous data, homogenization of storage strategies, and imbalance between security and efficiency.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for secure storage of multi-source heterogeneous data in smart engineering, which specifically includes the following steps: Step 1: Collect multi-source heterogeneous data from smart engineering projects, classify them into structured data, semi-structured data, and unstructured data according to data type, and perform compatibility adaptation processing for each. Step 2: Classify and grade the multi-source heterogeneous data after compatibility processing. Combine the business scenarios in the whole life cycle of smart engineering, and conduct quantitative analysis from the aspects of security risks and business value. Calculate the total score of the grade and divide it into core level, important level and general level to generate data grading results. Step 3: Based on the data classification results, adopt a differentiated storage strategy and configure corresponding storage architectures and backup mechanisms for core-level, important-level, and general-level data respectively. Step 4: Encrypt the data based on the data storage information, extract sensitive fields and generate encryption keys in a specific way to encrypt the data.
[0006] As a further aspect of the present invention, the compatibility adaptation process specifically includes: For structured data, it is compatible with mainstream relational databases and tabular files. It extracts the first 100 data samples of each field, determines the basic type by data length, numerical range and character features, and completes the field format standardization by combining preset rules. For semi-structured data, we traverse all data samples to count the frequency of each field, combine the knowledge graph of the smart engineering field to screen the business-relevant fields and remove redundant fields, clarify the field names, data types and constraints based on the extracted core fields, generate a standardized architecture, and design fault tolerance and adaptation logic. For unstructured data, metadata is extracted using a multimodal parsing model, and a structured index table in a unified format is generated based on the metadata.
[0007] As a further aspect of the present invention, the metadata extraction of unstructured data specifically includes: Video data: Frame features are extracted using CNN, and shooting time, duration, and frame rate are obtained by combining time series analysis. Shooting location is obtained by correlating GPS data, and device ID and video resolution information are extracted simultaneously. Engineering drawing data: CNN is used to extract component outlines, dimensions, and text information. The Transformer model parses the drawing title, scale, design unit, and version number. Combined with the BIM domain dictionary, the component type and component number are identified. Document data: The OCR function recognizes the text in scanned documents, and the Transformer model extracts the document title, author, publication time, associated project number, and keywords, automatically classifying the document type. Image data: Extract metadata such as shooting time, resolution, geographic range, shooting height, and image quality, and link it to the grid partition information of the GIS system through georeferencing.
[0008] As a further aspect of the present invention, the classification and grading process specifically includes: Business scenarios are divided based on the entire lifecycle of smart engineering, and quantitative analysis is conducted from both security risk and business value perspectives. To address security risks, we analyze the risks of sensitive data types in business scenarios, the scope of impact of leaks, and the degree of harm caused by tampering, and assign values to each. We then assign different weights to each risk factor and sum them up to obtain the security risk factor. To address business value, we analyze the business coreness, irreplaceability, and timeliness value of data in business scenarios, and assign values to each. We then assign different weights to each value factor and sum them up to obtain the business value factor. Different weights are assigned to safety risk factors and business value factors, and the weighted sum is used to obtain the total score for each level. The levels are divided according to the total score: a score of ≥8 is the core level, a score of 5 ≤ score <8 is the important level, and a score <5 is the general level.
[0009] As a further aspect of the present invention, the assignment rules for security risk factors are as follows: for sensitive types of risks, personal information is assigned 6 points, trade secrets are assigned 10 points, and core engineering data is assigned 9 points; for the scope of leakage impact, internal use only is assigned 3 points, cross-enterprise collaboration is assigned 7 points, and public accessibility is assigned 10 points; for the degree of tampering harm, no impact on business is assigned 2 points, impact on local processes is assigned 6 points, and causing project shutdown / accident is assigned 10 points. The rules for assigning business value factors are as follows: In terms of business coreness, non-core businesses are assigned 3 points, important businesses are assigned 7 points, and core businesses are assigned 10 points; in terms of data irreplaceability, replaceable businesses are assigned 2 points, difficult-to-replace businesses are assigned 6 points, and irreplaceable businesses are assigned 10 points; in terms of data timeliness value, long-term effectiveness is assigned 5 points, short-term effectiveness is assigned 8 points, and real-time effectiveness is assigned 10 points.
[0010] As a further aspect of the present invention, the differentiated storage strategy specifically includes: Core-level data: It adopts an enterprise-level distributed storage system combined with local backup, and uses a 3-replica storage mechanism of production replica + local hot backup replica + off-site cold backup replica. The backup strategy is real-time synchronization + scheduled full backup + incremental backup. Real-time synchronization realizes real-time synchronization between production replica and local hot backup replica. Full backup is performed once every Sunday at midnight. Incremental backup is performed every day at midnight. For critical data: a combination of enterprise-grade centralized storage and cloud backup is used, with a backup strategy of one full backup per month and one incremental backup every 3 days. General-level data: Uses ordinary disk array storage, adopts local full backup strategy, performs full backup once a month, stores backup data on local spare hard drive, and regularly checks and maintains disk array.
[0011] As a further aspect of the present invention, the method for generating the encryption key specifically includes: Extract sensitive fields from the data and label them with i, where i = 1, 2, ..., j, and j is the type of sensitive field. Obtain the length of each sensitive field and calculate the average length. Use the average length as the standard to divide the core-level data into equally divided data segments. The sensitive fields are converted to binary. Four consecutive identical binary numbers in the converted sensitive fields are identified and replaced with a single binary number. The replaced sensitive fields are then bundled in pairs according to their positional order to obtain a bundle. Convert the evenly divided data segments into binary. If the number of binary segments is odd, pad the end with 0. Divide the binary numbers according to the average length and generate a combined binary stream based on the division results. The binary data segments of the evenly divided data are XORed with the bundled data to generate encrypted data segments. A hash function is used to generate a unique hash value for each encrypted data segment. The final encryption key is generated by combining the hash value with a preset random salt value and using a key generation algorithm.
[0012] As a further aspect of the present invention, the generation rule for the combined binary stream is as follows: If the binary number can be divided equally, the divided binary streams are combined sequentially, with the first group combined with the last group, the second group combined with the second-to-last group, and so on. If the binary number cannot be divided equally, the remaining binary number is marked as the remaining binary stream. The divided binary streams are then combined sequentially to generate the basic combined binary stream. The remaining binary stream is then combined with the basic combined binary stream in turn.
[0013] As a further aspect of the present invention, the fault-tolerant adaptation logic for semi-structured data is as follows: When a field is missing, it is filled with an invalid value and a missing marker is recorded; when there is a structural variation, the variation node is automatically identified and the newly added hierarchical field is mapped to the corresponding position in the standardized architecture.
[0014] This invention provides a secure storage method for multi-source heterogeneous data in smart engineering. Compared with existing technologies, it has the following advantages: This invention addresses compatibility issues caused by database differences and inconsistent field naming for structured data by employing a dual mechanism of sample statistics and regular expression matching to automatically identify field types. For semi-structured data, it automatically extracts core fields through depth-first traversal and domain knowledge graph linkage to improve adaptation efficiency. For unstructured data, it accurately extracts business-related metadata and generates a unified structured index table to increase the rate of breaking down data silos.
[0015] This invention employs a tiered encryption strength adaptation mechanism. Core-level data is encrypted using triple encryption at the field level, file block level, and metadata level. Important-level data is encrypted using dual encryption at the file level and sensitive field level. General-level data is encrypted using basic file level encryption. This avoids the efficiency loss or security deficiencies caused by a one-size-fits-all approach to encryption. It also innovates a key generation mechanism based on data characteristics, combining sensitive field length and binary structure characteristics to generate customized keys. Key security is enhanced through bitwise XOR, hash value, and random salt value, increasing the difficulty of key cracking. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the steps of the multi-source heterogeneous data secure storage method of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figure 1 This application provides a method for secure storage of multi-source heterogeneous data for smart engineering, which specifically includes the following steps: Step 1: Collect multi-source heterogeneous data for smart engineering projects. This multi-source heterogeneous data includes various data sources such as sensors, monitoring equipment, BIM models, and GIS systems. Based on the data type of the multi-source heterogeneous data, it is classified into structured data, semi-structured data, and unstructured data. Structured data includes MySQL / Oracle data tables, Excel progress tables, and BIM model parameter tables. Semi-structured data includes construction logs (JSON / XML), sensor time-series data, and GIS geographic information data. Unstructured data includes monitoring videos, engineering drawings, and drone aerial images. At the same time, compatibility adaptation processing is performed according to different data types. For structured data, it is compatible with mainstream relational databases. It extracts the first 100 data samples of each field, determines the basic type by data length, numerical range and character features, and completes field format standardization based on preset rules. The preset rules include format unification rules, unit unification rules and outlier handling rules. The outlier handling rules specifically mean that it automatically identifies null values, duplicate values and data that exceeds the reasonable range. Duplicate values are retained with the latest one, and outliers are marked and stored in a separate outlier data table. For semi-structured data, all data samples are traversed, the frequency of occurrence of each field is counted, and knowledge graphs in the field of smart engineering are combined to select fields that are strongly related to business, eliminate redundant fields, and based on the extracted core fields, clarify the field names, data types, and constraints to generate a standardized architecture. At the same time, fault tolerance and adaptation logic is designed. When a field is missing, it is filled with an invalid value and a missing mark is recorded. When the structure changes, the mutation node is automatically identified and the newly added hierarchical fields are mapped to the corresponding positions in the standardized architecture. For unstructured data, metadata is extracted using a multimodal parsing model. For video data, frame features are extracted using CNN, and shooting time, duration, and frame rate are obtained by combining time series analysis. The shooting location is obtained by using GPS correlation data, and basic information such as device ID and video resolution are extracted simultaneously. Engineering drawing data is processed by CNN to extract component outlines, dimensions, and text information from the drawings. The Transformer model parses the drawing title, scale, design unit, and version number, and combines the BIM domain dictionary to identify core business metadata such as component type and component number. For document data, OCR is used to recognize the text in scanned documents, and the Transformer model extracts the document title, author, publication time, associated project number, and keywords, and automatically classifies the document type. Metadata such as shooting time, resolution, geographical range, shooting height, and image quality are extracted from image data. This data is then linked to the grid partition information of the GIS system through georeferencing. Based on the extracted metadata, a structured index table with a unified format is generated.
[0019] Step 2: Obtain multi-source heterogeneous data after compatibility processing, classify and grade it, and combine it with the entire life cycle of smart engineering to divide business scenarios and obtain multiple sets of business scenario data, such as engineering design, construction execution, and bidding and contract. Then, conduct quantitative analysis from the perspectives of security risks and business value. To address security risks, we analyzed the sensitive types of data in business scenarios, the scope of leakage impact, and the degree of tampering harm, and assigned values to each. Sensitive types of risks were assigned 6 points for personal information, 10 points for trade secrets, and 9 points for core engineering data. The scope of leakage impact was assigned 3 points for internal use only, 7 points for cross-enterprise collaboration, and 10 points for public accessibility. The degree of tampering harm was assigned 2 points for no impact on business, 6 points for impact on local processes, and 10 points for causing project shutdown / accidents. We also assigned different weights to sensitive types of risks, the scope of leakage impact, and the degree of tampering harm, and then calculated the weighted sum to obtain the security risk factor. To assess business value, the analysis considers the core business nature, irreplaceability, and timeliness of data within business scenarios, assigning values accordingly. Core business nature is scored as follows: 3 points for non-core business, 7 points for important business, and 10 points for core business. Irreplaceability is scored as follows: 2 points for replaceable, 6 points for difficult to replace, and 10 points for irreplaceable. Timeliness is scored as follows: 5 points for long-term validity, 8 points for short-term validity, and 10 points for real-time validity. Furthermore, different weights are assigned to core business nature, irreplaceability, and timeliness, and the resulting weighted sum is used to calculate the business value factor. Based on the obtained security risk factors and business value factors, different weights are assigned to the two, and a weighted sum is calculated to obtain the total score for the grade. Those with a total score greater than 8 points are classified as core level, those with a total score greater than 5 points but less than 8 points are classified as important level, and those with a total score less than 5 points are classified as general level. This process is repeated to categorize all business scenario data and generate data classification results.
[0020] Step 3: Based on the obtained data classification results, data storage management is carried out. Differentiated storage strategies are adopted for data of different classifications. For core data, a combination of distributed storage system and local backup is used. The distributed storage system is an enterprise-level distributed storage system with a 3-replica storage mechanism of production replica + local hot backup replica + off-site cold backup replica. At the same time, a combination strategy of real-time synchronization + scheduled full backup + incremental backup is adopted. Specifically, real-time synchronization means that the production replica and the local hot backup replica are synchronized in real time. Full backup means that a full encrypted backup is performed once every Sunday at midnight. Incremental backup means that an incremental backup is performed every day at midnight to reduce backup time and storage usage. For critical data, a combination of centralized storage and cloud backup is used. The centralized storage system is an enterprise-grade centralized storage system, and backup is performed using full backup and incremental backup. Full backup means a full backup is performed once a month and stored in local centralized storage and the cloud. Incremental backup means an incremental backup is performed every 3 days, uploading only changed data to reduce bandwidth consumption. At the same time, critical data is backed up to the cloud storage platform, taking advantage of the elastic expansion and disaster recovery capabilities of cloud storage to ensure data security and availability. For general-level data, a standard disk array is used for storage, employing a "local full backup" strategy. A full backup is performed once a month, with the backup data stored on a local spare hard drive, eliminating the need for off-site backups. At the same time, the data in the disk array is regularly checked and maintained to ensure data integrity and readability.
[0021] Step 4: Based on the obtained data storage information, encrypt the data, extract the sensitive fields, and generate a separate key. The specific generation method is as follows: The extracted sensitive fields are obtained and labeled as i, where i = 1, 2, ..., j, and j represents the type of sensitive field. The field length corresponding to sensitive field i is obtained. Then, the average length of all sensitive field lengths is calculated. At the same time, the core-level data is segmented based on the average length to obtain evenly divided data segments. Next, the sensitive fields are converted to binary. Consecutive identical binary numbers in the converted sensitive fields are replaced. The specific replacement rule is to identify four consecutive identical binary numbers and replace them with a single binary number, for example, 0000 is replaced with 0, 1111 is replaced with 1. At the same time, all the replaced sensitive fields are obtained and bundled in pairs according to their position order to obtain bundled packages. For evenly divided data segments, first convert them to binary and obtain the number of binary segments. If the number of binary segments is odd, pad with 0s at the end to make it even. If the number of binary segments is even, no padding is needed. Then, divide the binary segments in order based on the average length. If the binary segments can be evenly divided, obtain all binary streams and combine them in the order of the segments. Specifically, prevent the binary streams labeled with the first group from being combined with the binary streams labeled with the last group, combine the second group with the second to last group, and so on, to generate combined binary streams. If the binary segments cannot be evenly divided, obtain the remaining binary segments and mark them as the remaining binary streams. At the same time, combine the binary streams obtained after the segments in the order of the segments to generate combined binary streams. Using the time series as the standard, specifically the time series [1, 2, ..., 12], combine the remaining binary streams with the combined binary streams in sequence. When the corresponding time is 2, combine the remaining binary streams with the combined binary streams labeled with the first group. When the corresponding time is 2, combine the remaining binary streams with the combined binary streams labeled with the second group, and so on. Next, the binary data segments of the evenly divided data segments are XORed with the bundled data segments to generate encrypted data segments. At the same time, a hash function is used to generate a unique hash value for each encrypted data segment. The hash value is used as part of the key, and combined with a preset random salt value, the final encryption key is generated through a key generation algorithm.
[0022] Some of the data in the above formulas are numerical calculations with dimensions removed, and the contents not described in detail in this specification are all prior art known to those skilled in the art.
[0023] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A method for secure storage of multi-source heterogeneous data for smart engineering, characterized in that, The method specifically includes the following steps: Step 1: Collect multi-source heterogeneous data from smart engineering projects, classify them into structured data, semi-structured data, and unstructured data according to data type, and perform compatibility adaptation processing for each. Step 2: Classify and grade the multi-source heterogeneous data after compatibility processing. Combine the business scenarios in the whole life cycle of smart engineering, and conduct quantitative analysis from the aspects of security risks and business value. Calculate the total score of the grade and divide it into core level, important level and general level to generate data grading results. Step 3: Based on the data classification results, adopt a differentiated storage strategy and configure corresponding storage architectures and backup mechanisms for core-level, important-level, and general-level data respectively. Step 4: Encrypt the data based on the data storage information, extract sensitive fields and generate encryption keys in a specific way to encrypt the data.
2. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 1, characterized in that, Compatibility adaptation handling specifically includes: For structured data, it is compatible with mainstream relational databases and tabular files. It extracts the first 100 data samples of each field, determines the basic type by data length, numerical range and character features, and completes the field format standardization by combining preset rules. For semi-structured data, we traverse all data samples to count the frequency of each field, combine the knowledge graph of the smart engineering field to screen the business-relevant fields and remove redundant fields, clarify the field names, data types and constraints based on the extracted core fields, generate a standardized architecture, and design fault tolerance and adaptation logic. For unstructured data, metadata is extracted using a multimodal parsing model, and a structured index table in a unified format is generated based on the metadata.
3. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 1, characterized in that, Metadata extraction from unstructured data specifically includes: Video data: Frame features are extracted using CNN, and shooting time, duration, and frame rate are obtained by combining time series analysis. Shooting location is obtained by correlating GPS data, and device ID and video resolution information are extracted simultaneously. Engineering drawing data: CNN is used to extract component outlines, dimensions, and text information. The Transformer model parses the drawing title, scale, design unit, and version number. Combined with the BIM domain dictionary, the component type and component number are identified. Document data: The OCR function recognizes the text in scanned documents, and the Transformer model extracts the document title, author, publication time, associated project number, and keywords, automatically classifying the document type. Image data: Extract metadata such as shooting time, resolution, geographic range, shooting height, and image quality, and link it to the grid partition information of the GIS system through georeferencing.
4. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 1, characterized in that, The classification and grading process specifically includes: Business scenarios are divided based on the entire lifecycle of smart engineering, and quantitative analysis is conducted from both security risk and business value perspectives. To address security risks, we analyze the risks of sensitive data types in business scenarios, the scope of impact of leaks, and the degree of harm caused by tampering, and assign values to each. We then assign different weights to each risk factor and sum them up to obtain the security risk factor. To address business value, we analyze the business coreness, irreplaceability, and timeliness value of data in business scenarios, and assign values to each. We then assign different weights to each value factor and sum them up to obtain the business value factor. Different weights are assigned to safety risk factors and business value factors, and the weighted sum is used to obtain the total score for each level. The levels are divided according to the total score: a score of ≥8 is the core level, a score of 5 ≤ score <8 is the important level, and a score <5 is the general level.
5. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 4, characterized in that, The scoring rules for security risk factors are as follows: For sensitive types of risks, personal information is scored 6 points, trade secrets are scored 10 points, and core engineering data is scored 9 points; for the scope of impact of leakage, internal use only is scored 3 points, cross-enterprise collaboration is scored 7 points, and public accessibility is scored 10 points; for the degree of harm caused by tampering, no impact on business is scored 2 points, impact on local processes is scored 6 points, and causing project shutdown / accident is scored 10 points. The rules for assigning business value factors are as follows: non-core businesses are assigned 3 points, important businesses are assigned 7 points, and core businesses are assigned 10 points. In the assessment of data non-substitutability, substitutable assignments are worth 2 points, difficult-to-substitut assignments are worth 6 points, and non-substitutable assignments are worth 10 points. Data timeliness value is assigned as follows: long-term validity is assigned 5 points, short-term validity is assigned 8 points, and real-time validity is assigned 10 points.
6. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 1, characterized in that, Differentiated storage strategies specifically include: Core-level data: It adopts an enterprise-level distributed storage system combined with local backup, and uses a 3-replica storage mechanism of production replica + local hot backup replica + off-site cold backup replica. The backup strategy is real-time synchronization + scheduled full backup + incremental backup. Real-time synchronization realizes real-time synchronization between production replica and local hot backup replica. Full backup is performed once every Sunday at midnight. Incremental backup is performed every day at midnight. For critical data: a combination of enterprise-grade centralized storage and cloud backup is used, with a backup strategy of one full backup per month and one incremental backup every 3 days. General-level data: Uses ordinary disk array storage, adopts local full backup strategy, performs full backup once a month, stores backup data on local spare hard drive, and regularly checks and maintains disk array.
7. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 1, characterized in that, The specific methods for generating encryption keys include: Extract sensitive fields from the data and label them with i, where i = 1, 2, ..., j, and j is the type of sensitive field. Obtain the length of each sensitive field and calculate the average length. Use the average length as the standard to divide the core-level data into equally divided data segments. The sensitive fields are converted to binary. Four consecutive identical binary numbers in the converted sensitive fields are identified and replaced with a single binary number. The replaced sensitive fields are then bundled in pairs according to their positional order to obtain a bundle. Convert the evenly divided data segments into binary. If the number of binary segments is odd, pad the end with 0. Divide the binary numbers according to the average length and generate a combined binary stream based on the division results. The binary data segments of the evenly divided data are XORed with the bundled data to generate encrypted data segments. A hash function is used to generate a unique hash value for each encrypted data segment. The final encryption key is generated by combining the hash value with a preset random salt value and using a key generation algorithm.
8. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 7, characterized in that, The rules for generating combined binary streams are as follows: If the binary number can be divided equally, then the divided binary streams are combined sequentially, with the first group combined with the last group, the second group combined with the second to last group, and so on. If the binary number cannot be divided equally, the remaining binary number is marked as the remaining binary stream. The binary streams after equal division are combined end to end in order to generate the basic combined binary stream. The remaining binary stream is then combined with the basic combined binary stream in turn.
9. The method for secure storage of multi-source heterogeneous data for smart engineering according to claim 2, characterized in that, The fault tolerance adaptation logic for semi-structured data is as follows: When a field is missing, it is filled with an invalid value and a missing marker is recorded; when there is a structural variation, the variation node is automatically identified and the newly added hierarchical field is mapped to the corresponding position in the standardized architecture.