One-way database data acquisition method and system based on file ferry
By using a one-way database data acquisition method in the file transfer zone, and leveraging unified data middleware and optical one-way transmission equipment, the problems of data silos and security in data transmission in steel and metallurgical enterprises have been solved, achieving efficient and secure data transmission and analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BERIS ENG & RES CORP
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-28
AI Technical Summary
Data transmission in steel and metallurgical enterprises suffers from data silos, inconsistencies, poor security, and the risk of data leakage, impacting data analysis and security during digital transformation.
A one-way database data acquisition method based on file transfer is adopted. The data is mapped, converted and encrypted through a unified data middleware, and physical isolation and security verification are achieved by using optical one-way transmission equipment to ensure one-way data transmission between the source database and the target database.
It improves the security and efficiency of data transmission, reduces the risk of data leakage, ensures the reliability and consistency of data transmission, and supports the digital transformation of steel and metallurgical enterprises.
Smart Images

Figure CN121935121A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a one-way database data acquisition method and system based on file transfer. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] Currently, the steel industry is experiencing a continuous downturn, and many steel companies are making adjustments to their high-quality development strategies to cope with the possibility of a long-term "winter" in the industry. Digital transformation projects such as management and control centers are being implemented in full swing. In the process of digital transformation of steel and metallurgical enterprises, eliminating data silos and building digital assets are very important aspects.
[0004] The informatization of iron and steel metallurgical enterprises has been gradually improved. Major process steps, such as blast furnaces and raw materials, have their own secondary systems. This leads to inconsistencies in database types, table formats, and data dictionaries across these secondary systems, causing numerous problems for data acquisition during digital transformation. The lack of effective verification during data transmission increases the risk of data tampering, creating bottlenecks in secure transmission. Furthermore, the complexity and diversity of data in iron and steel metallurgical processes, coupled with poor consistency across systems, severely impacts data analysis throughout the entire process. Summary of the Invention
[0005] To overcome the shortcomings of the prior art, this invention provides a one-way database data acquisition method and system based on file transfer, used to acquire database data from the process machine system of an iron and steel metallurgical enterprise. A file transfer area is set up between the source database and the target database, which effectively improves the security and efficiency of data transmission and reduces the risk of data leakage.
[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: In a first aspect, the present invention provides a one-way database data acquisition method based on file transfer, comprising: Data is collected from multiple source databases to obtain multi-source data; The multi-source data are respectively input into the unified data middleware for mapping and conversion to obtain data in a unified intermediate format. The unified intermediate format data is transmitted to the intranet side of the file transfer area. The unified intermediate format data is then subjected to security verification and encryption in the file transfer area. The encrypted data is then transmitted unidirectionally to the extranet side of the file transfer area. The encrypted data is transmitted from the external network side of the file transfer area to the target database.
[0007] In a further technical solution, the unified data middleware sequentially includes an access layer, an intelligent mapping engine, an execution engine, and a monitoring and management layer.
[0008] In a further technical solution, the access layer is used to receive multi-source data, parse it, and extract the corresponding data types; the intelligent mapping engine is used to automatically identify and match similar data in different systems based on a word vector model; and the execution engine is used to perform field mapping and transformation on similar data based on a mapping library to obtain data in a unified intermediate format.
[0009] A further technical solution is that the intelligent mapping engine is used to automatically identify and match similar data in different systems based on a word vector model, specifically as follows: The multi-source data is preprocessed to obtain preprocessed multi-source data. Construct multi-dimensional features based on preprocessed multi-source data; The multi-dimensional features are input into the word vector model to obtain the vectorization of the multi-dimensional features, and the similarity of the multi-dimensional features is calculated respectively. Data with a similarity greater than a set threshold will be matched as similar data.
[0010] A further technical solution is that the mapping library includes a data dictionary and mapping rules, and the mapping rules include field mapping rules, data type conversion rules, and data format conversion rules.
[0011] In a further technical solution, the file transfer area is physically isolated from the source database and the target database respectively through a one-way optical transmission device.
[0012] In a further technical solution, the security verification uses hash values to verify the integrity of the data.
[0013] Secondly, the present invention provides a one-way database data acquisition system based on file transfer, comprising: The data acquisition module is configured to collect data from multiple source databases to obtain multi-source data. The intermediate conversion module is configured to input the multi-source data into the unified data middleware for mapping and conversion to obtain data in a unified intermediate format. The intermediate isolation module is configured to: transmit the data of the unified intermediate format to the intranet side of the file transfer area, perform security verification and encryption on the data of the unified intermediate format in the file transfer area, and transmit the encrypted data unidirectionally to the extranet side of the file transfer area; The target security module is configured to: obtain the encrypted data from the external network side of the file transfer area and transmit it to the target database.
[0014] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a one-way database data acquisition method based on file transfer as described in the first aspect.
[0015] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a one-way database data acquisition method based on file transfer as described in the first aspect.
[0016] The above one or more technical solutions have the following beneficial effects: This invention is applied to iron and steel metallurgical enterprises. A file transfer zone is established between the source and target databases, isolated into an internal network and an external network. Data from the source database is collected in real-time, exported as a JSON file, and securely transmitted to the internal network of the transfer zone. After security verification and encryption, the data is transmitted unidirectionally to the external network and finally imported into the target database. This effectively improves the security and efficiency of data transmission and reduces the risk of data leakage.
[0017] This invention enables unidirectional data flow at the hardware level through a unidirectional optical transmission device, physically blocking reverse data transmission and using hash values for data security verification to prevent data tampering and improve the reliability of collected data.
[0018] This invention introduces a unified data middleware based on intelligent mapping and transformation before unidirectional data transmission, so as to realize the automatic adaptation and unification of data between different systems. Attached Figure Description
[0019] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0020] Figure 1 This is a flowchart of a one-way database data acquisition method based on file transfer according to an embodiment of the present invention. Detailed Implementation
[0021] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0022] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0023] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0024] Example 1 like Figure 1 As shown in the figure, this embodiment discloses a one-way database data acquisition method based on file transfer, which includes the following steps: S1: Collect data from multiple source databases to obtain multi-source data; In this embodiment, data is collected from the secondary system databases of each process segment in an iron and steel metallurgical enterprise. Each secondary system database is considered a source database, and data is collected from each source database to obtain multi-source data. During the data collection process, the source databases act as data senders, and their locations have a lower security level, constituting a low-security zone.
[0025] On the source database side, Flink CDC data collection tools are deployed to monitor the database's runtime logs in real time, such as MySQL's binlog, undolog, and redolog, to obtain relevant data, including metadata changes, table changes, and transaction changes. This allows for real-time data monitoring, distribution, and synchronization. The data collected by CDC is then exported as JSON files for subsequent processing and transmission. The export process ensures data integrity and accuracy through checksum verification.
[0026] During file transfer and data processing, the JSON file name needs to be processed. Therefore, the naming rules for the file are agreed upon. The file name consists of 5 parts connected by the # symbol. The meaning of the connected parts is shown in Table 1.
[0027] Table 1 Data File Naming Rules
[0028] S2: Input the multi-source data into the unified data middleware for mapping and conversion to obtain data in a unified intermediate format; In this embodiment, a unified data middleware based on intelligent mapping and transformation enables automatic adaptation and unification of data between different systems. The unified data middleware adopts a layered architecture design, including an access layer, a metadata management layer, an intelligent mapping engine, an execution engine, and a monitoring management layer.
[0029] The access layer receives data from multiple sources, automatically identifies metadata, parses it, and extracts the corresponding data types. Metadata refers to information describing the data structure obtained from parsing multi-source data, such as field names, field types, formats, lengths, process segments, and equipment identifiers. Metadata is an abstract description of the actual business data, used for subsequent semantic recognition, mapping, and transformation.
[0030] The process of automatically identifying metadata is as follows: (1) Parse the JSON files or table structures of multi-source databases; (2) Infer the semantic category of the field based on the data naming conventions (such as field naming, process section identification), data type, and format; (3) The above results are encapsulated into a unified metadata structure for subsequent mapping engine calls. This process automatically parses structured information, rather than manually defining metadata types. Identifying metadata is the first step in building a unified semantic space, providing basic input for subsequent word vector similarity calculations and field mapping transformations.
[0031] The metadata management layer is used to manage the mapping library. Primarily responsible for centralized management of the mapping library, it is the core management module in the unified data middleware. This layer does not directly process business data, but rather organizes, maintains, and schedules the metadata extracted from multi-source data and its corresponding relationships to support the work of the intelligent mapping engine and execution engine. The functions of the metadata management layer include: managing the data dictionary: storing basic metadata such as standard field names, field meanings, data types, unit specifications, and value ranges for each process segment, providing a unified, standardized semantic space; managing mapping rules: maintaining field mapping rules, data type conversion rules, data format conversion rules, and synonym mapping relationships for use by the intelligent mapping engine and execution engine; providing metadata query and access services: providing a standardized benchmark for the mapping engine when performing semantic similarity calculations, ensuring that the field matching process between different systems is controllable, reusable, and auditable; and ensuring the consistency of the mapping library: through centralized management, ensuring the consistent use of mapping rules throughout the entire conversion chain, avoiding field conflicts or type inconsistencies caused by differences in multi-source systems. In summary, the role of the metadata management layer is to provide standard basis and rule support for the entire data unification process, so that intelligent identification and unified conversion have a reliable standard foundation, and ensure that the final generated unified intermediate format data is accurate, consistent and standardized.
[0032] The intelligent mapping engine is used to automatically identify and match similar data in different systems based on word vector models.
[0033] The execution engine is used to map and transform similar data based on the mapping library to obtain data in a unified intermediate format. The execution engine converts similar fields from different source databases into a unified intermediate format, and the specific process includes the following steps: (1) Read the standard field definitions in the mapping library The mapping library pre-stores standard field names, field types, data formats, and synonym mapping rules for each process segment. The execution engine first locates the corresponding standard field item in the mapping library based on the identified fields of the same type.
[0034] (2) Field name mapping (unified names) The execution engine uses field mapping rules to uniformly replace field names that differ across systems with standard field names. For example, fields such as "steelTemp", "TSTEEL", and "temp_steel" are uniformly mapped to "steel_temperature".
[0035] (3) Data type conversion Based on the type conversion rules in the mapping library, the data type of the source field is converted to the target unified type. For example, a string field for "speed" or an int field for "temperature" is converted to a float or numeric type to ensure consistency in subsequent processing.
[0036] (4) Data format conversion Standardize and transform data in different formats, for example: Convert the time field to a unified format (such as ISO 8601 time format). Numerical fields are standardized in units according to rules (e.g., m / min is converted to m / s). JSON internal objects are flattened into a uniform key-value structure.
[0037] (5) Generate data in a unified intermediate format according to the mapping rules. The execution engine reorganizes the processed fields into a standard intermediate format, such as a standardized JSON structure, according to a unified data structure template. This format includes uniform field names, types, and formats, facilitating subsequent validation and data transfer.
[0038] (6) Optional manual confirmation mechanism In cases where semantic similarity is high but ambiguity exists, the execution engine can provide a manual verification step to ensure the accuracy of the mapping and avoid data errors caused by mismatches.
[0039] Therefore, field mapping and transformation is a multi-step process based on the mapping library, including field name replacement, type standardization, and format unification, which ultimately ensures that similar data from different systems are presented in a consistent structure, forming a unified intermediate format data.
[0040] The monitoring and management layer is used to validate and repair data. The purpose of validation and repair is to ensure the correctness of the unified intermediate format data after mapping and transformation in terms of basic syntax, type, and structure. The specific process includes: (1) Syntax and format verification The monitoring management team performs basic syntax checks on the converted data, including: whether the field data types are consistent with the standard types in the mapping library (e.g., numeric fields should not appear as strings); whether the data format conforms to predefined standards (e.g., whether the time field format is valid and the numeric field format is correct); and whether the JSON structure is complete and whether there are any missing fields, redundant fields, or structure parsing errors.
[0041] (2) Mapping rule consistency verification The monitoring management layer checks the mapping library to ensure that: field names have been successfully converted to standard field names according to the mapping rules; data types have been correctly converted according to the mapping rules; and format conversion rules have been successfully executed, such as type conversion and format unification. If a conversion is found to be inconsistent with the mapping rules, the monitoring management layer will mark the field as abnormal.
[0042] (3) Data content legality verification The monitoring management checks: whether the field value is empty but cannot be empty according to the rules; whether the field value has format errors, such as illegal characters or abnormal symbols; and whether the numeric field meets the format requirements (such as it must be an integer or floating-point type).
[0043] (4) Automatic repair mechanism For formatting issues detected during the above validation, the monitoring management layer will repair them according to predefined rules. For example, missing fields will be filled with default values (such as empty strings or 0) to ensure structural integrity; invalid data will be corrected, such as removing illegal characters or converting string-based numeric values to valid numeric types; and mismatched field types will be re-converted (such as restoring the string "123" to a numeric type). Data that cannot be automatically repaired will be marked as an anomaly for system logging or manual processing.
[0044] (5) Verification and repair recording mechanism The monitoring management retains all verification and repair results as logs for subsequent auditing, investigation, and data tracing.
[0045] Furthermore, the mapping library in the metadata management layer includes data dictionaries and mapping rules for the secondary systems of each process segment. The mapping library defines the correspondence between similar data in different systems, including field names, data types, data formats, etc.
[0046] The data dictionary originates from the database structure description files of the secondary systems (such as ironmaking, steelmaking, and rolling) of various process sections in the steel enterprise. After parsing the multi-source database structure at the access layer, the structural information of each field is organized into dictionary items, specifically including: the source field name (such as temp, speed_val, ts, etc.); the original data type of the field (such as varchar, double, int, datetime); the original representation format of the field (such as "2024 / 01 / 01 12:00", strings with units such as "1200℃", "12.50 t / h", etc.); a description of the business meaning of the field (field comments or configuration descriptions from the secondary systems); and the process section and system source to which the field belongs (for subsequent semantic judgment).
[0047] The data dictionary includes process entity identifiers (such as material codes, unique equipment identifiers, etc.), process parameters, and metadata specifications. Process parameters include composition parameters (such as carbon content in molten steel, alloy element ratios), speed parameters (such as mill line speed, continuous casting speed), temperature parameters (such as rolling temperature, molten steel temperature, molten iron temperature), pressure parameters (such as rolling force, blast furnace top pressure), and time parameters (such as smelting cycle). The metadata specifications provide field naming conventions, namely
Process Segment
Equipment
Process Parameter
Type
[0048] Mapping rules describe the conversion method from "source system field to unified standard field". Mapping rules include field mapping rules, data type conversion rules, and data format conversion rules. Field mapping rules set up a synonym mapping table to convert synonym fields into unified standard terms. Data type conversion rules convert the source type into the target type. Data format conversion rules convert data of different formats into a unified intermediate format.
[0049] Field mapping rules specify that fields from multiple different systems should be standardized to the same standard field. As shown in Table 1, for example: temp, temperature_value → standard field temperature; ts, time_stamp, collect_time → standard field timestamp. Data type conversion rules standardize how the original types of fields from different systems are converted to intermediate formats. For example: converting the varchar type "1250" to float; converting the int type timestamp to a unified datetime string; parsing the string "12.5 t / h" with units as the numerical value 12.5. Data format conversion rules ensure that data is output in a unified format. For example: all time fields are unified as "YYYY-MM-DD HH:mm:ss"; all numeric fields are unified as "floating-point decimal form, without retaining units".
[0050] The construction process of the mapping library is as follows: 1. Parse the multi-source database structure; the access layer obtains the data table structure and field metadata of each secondary system. 2. Automatically generate an initial data dictionary; by parsing the table structure, field names, types, formats, field comments, etc., are solidified into standardized dictionary items. 3. Perform semantic clustering based on a word vector model; the intelligent mapping engine generates vectors for field names and field comments, and identifies "same type fields" through similarity calculation. 4. Solidify mapping rules after manual review; the field correspondences identified by the model are written into the mapping library after manual confirmation, forming rules for field name mapping, type conversion, format conversion, etc. 5. Dynamically supplement and correct during subsequent operation; when a new table structure appears, the system automatically generates new dictionary items and adds rules. An example of the mapping library is shown in Table 2 below.
[0051] Table 2 Example of Mapped Library Contents
[0052] Furthermore, the intelligent mapping engine uses pre-trained word vector models (such as the Word2Vec model) to calculate the semantic similarity in multi-source data, and automatically identify and match similar data in different systems.
[0053] The specific identification steps are as follows: (1) Preprocess the multi-source data to obtain the preprocessed multi-source data. The preprocessing includes field name standardization, term splitting and stop word filtering. Special characters such as "_" and "#" are removed by field name standardization, and the case is unified. Long fields are divided into multiple short word fields by term splitting. Meaningless words are removed by stop word filtering based on a predefined stop word list.
[0054] (2) Construct multi-dimensional features based on preprocessed multi-source data. Multi-dimensional features include semantic features, structural features and physical features. The word segmentation results of field names are used as semantic features, the process segment and equipment to which the data belongs are used as structural features, and the data type is used as physical features.
[0055] (3) Input the multi-dimensional features into the word vector model to obtain the vectorization of the multi-dimensional features; calculate the similarity of the multi-dimensional features respectively.
[0056] (4) Match data with similarity greater than the set threshold as data of the same type.
[0057] Semantic similarity calculations are performed on field-level metadata, not on entire data entries. This field metadata originates from the table structure of the multi-source process segment's secondary system, and the access layer automatically parses it to obtain field names, data types, field formats, field comments, and information about the process segment and equipment to which it belongs. Based on this, the intelligent mapping engine compares and calculates similarity between "fields" to identify fields with the same semantic meaning in different systems.
[0058] Before calculating similarity, the system first filters comparable objects based on the structural characteristics of the fields, including: 1. Process consistency filtering, for example, fields related to the ironmaking system are only compared with fields related to the ironmaking system to avoid incorrect matching across processes. 2. Equipment type consistency filtering, such as temperature-related fields being compared only with data fields of the same type of equipment. 3. Data type compatibility filtering, such as temperature fields all being numeric; time fields all being time-type, etc. After the above filtering, the set of comparable fields is limited to "fields with similar business backgrounds".
[0059] Semantic similarity is determined by three types of features: 1) Field semantic features The intelligent mapping engine processes field names and field annotations as follows: field name standardization (removing underscores and unifying capitalization), word segmentation and term splitting, stop word filtering, and vectorization of word segmentation results through a pre-trained word vector model. Multiple word vectors are averaged or weighted to form field semantic vectors.
[0060] 2) Field structure characteristics Encode the process segment, system origin, and equipment type of the field into a dense vector (such as embedding / unique heat encoding). This is used to reflect the industrial context of the field.
[0061] 3) Field physical characteristics The data types of fields (int, float, varchar, datetime, etc.) are uniformly encoded to ensure physical comparability.
[0062] 4) Multi-dimensional similarity calculation Semantic vectors, structural vectors, and physical vectors are combined into field feature vectors, and similarity is calculated using the following methods: cosine similarity (primarily for semantic vectors); Euclidean distance or Manhattan distance (for structural / physical vectors); and the final similarity is combined in a weighted manner. This ensures that fields with similar names, consistent process segments, and matching types receive higher similarity scores.
[0063] The specific steps for matching fields of the same type are as follows: 1) Parse field metadata: The access layer extracts information such as field name, type, format, comments, process section, and equipment from different systems to form a set of fields to be processed.
[0064] 2) Field preprocessing and feature construction: Generate semantic features, structural features, and physical features for each field, and then quantize them.
[0065] 3) Limit the candidate comparison set: Only compare fields that are consistent in process segment, equipment type, and data type.
[0066] 4) Calculate similarity based on word vectors: Perform multi-dimensional similarity calculations on each pair of candidate fields to obtain a field similarity matrix.
[0067] 5) Threshold filtering for candidates of the same type of field; if the similarity of a field pair is greater than the set threshold (e.g., 0.75), it will be automatically marked as a candidate of the same type of field.
[0068] 6) Manual confirmation and writing to the mapping library: The operator confirms the automatic matching results and writes the final matching relationship into the "field mapping rules" as part of the mapping library.
[0069] 7) Unified field conversion by the execution engine: The execution engine calls the above mapping rules during data conversion to convert the field name, data type, format, etc. of similar fields and generate a unified intermediate format.
[0070] Furthermore, before performing the conversion of similar data, operators can first confirm the matched similar data to determine if they are indeed similar. If so, the conversion is performed; otherwise, the matched data is adjusted to avoid data errors caused by conversion mistakes.
[0071] The execution engine uses a data format converter to convert data of different formats in the same type into a unified intermediate format according to the mapping rules in the mapping library, and converts field names into unified standard terms and data types into target types.
[0072] The execution engine actually performs three types of conversion tasks according to the three categories of rules in the mapping library: field name conversion, data type conversion, and data format conversion. Field name conversion is performed by the execution engine, which replaces the source system field names with standard field names based on the "field name mapping rules" in the mapping library. This process is completed by the field name conversion submodule. Data type conversion is performed by the execution engine, which converts the original type of the source field (such as varchar, int) to the target type (such as float, datetime) required by the unified intermediate format, based on the "type conversion rules." This process is completed by the data type conversion submodule. The data format converter is responsible for applying the "format conversion rules" in the mapping library, such as unifying time formats, removing units, and standardizing string formats. This process is completed by the data format conversion submodule. The execution engine completes the unified conversion of field names, data types, and data formats by sequentially calling the above three conversion submodules according to the corresponding rules in the mapping library.
[0073] Furthermore, within the monitoring and management layer, the converted data is received and validated to ensure its integrity and consistency. Data validation includes syntax verification, process rule verification, and production logic verification. Syntax verification includes validating data types (e.g., temperature values must be floating-point numbers) and verifying data format compliance (e.g., timestamps must conform to ISO 8601 standards). Process rule verification involves validating different process parameters according to corresponding rules (e.g., molten steel temperature should be between [1500℃, 1700℃]). Production logic verification addresses process time constraints, such as the steelmaking start time being less than the tapping time, and the tapping time being less than the continuous casting start time.
[0074] S3: The data in the unified intermediate format is transmitted to the intranet side of the file transfer area. The data in the unified intermediate format is subjected to security verification and encryption in the file transfer area. The encrypted data is transmitted unidirectionally to the extranet side of the file transfer area. In this embodiment, a multi-layered one-way data transmission mechanism based on logical isolation and protocol control is constructed, and a multi-layered one-way transmission system based on "physical isolation + protocol control" is established.
[0075] The file transfer zone is a one-way transmission intermediate layer, serving as a buffer for logical isolation and protocol control. Located between the source and target databases, it isolates data on the internal network and external network sides. Data from the source database is securely transmitted to the internal network side after processing by a unified data middleware, then undergoes security verification and encryption before being transmitted unidirectionally to the external network side, and finally imported into the target database.
[0076] The internal network side and external network side refer to the physical isolation of the dual-network environment upon which the file transfer zone relies. The file transfer zone is deployed in corresponding directories or storage areas on both sides of a unidirectional transmission device (such as a unidirectional gateway). This unidirectional transmission device provides unidirectional physical isolation between the internal and external networks and unidirectional data transmission capabilities. Specifically: Isolation is physical: The internal storage area (internal network side) and the external storage area (external network side) of the file transfer zone, relying on the unidirectional transmission device, are naturally physically isolated, and the networks cannot access each other; Isolation is achieved through the unidirectional transmission device: The unidirectional transmission device provides an internal-to-outside unidirectional transmission channel, achieving the isolation effect of "the internal network can write to the external network unidirectionally, but the external network cannot access it in reverse"; The role of the file transfer zone is to utilize these two physically isolated areas to achieve secure relay: The unified data middleware writes the processed files to the internal network side directory, automatically transmits them unidirectionally to the external network side directory via the unidirectional device, and then imports them into the target database from the external network side. Therefore, the "internal network side / external network side" of the file transfer area relies on the physical isolation and unidirectional transmission capability provided by the unidirectional transmission device. This application achieves logical isolation and security control in this way.
[0077] The file transfer area serves as a secure, isolated zone between the source and target databases, enabling unidirectional data flow. The specific settings for the file transfer area are as follows: 1) Physical Isolation: First, the file transfer area is physically isolated from both the source and target database systems, ensuring that direct connections at the network layer are severed to prevent potential network attacks or data leaks. This is achieved by installing a hardware-level isolation mechanism—a unidirectional optical transmission device—between the networks of the source and target databases. This hardware-level isolation mechanism employs physical optical gate technology, implemented using unidirectional fiber optic devices. The transmitting end is configured with a laser transmitter, while the receiving end does not physically have an optical receiver module; data transmission is achieved through photoelectric conversion to achieve unidirectional physical signal transmission. All routing protocols are disabled on the network devices, and a unidirectional access control list (ACL) is configured, allowing only outbound connections. Physical interfaces are managed separately, with sending and receiving using independent network devices.
[0078] 2) Access Control: Strict access control is implemented for the file transfer area. Only authorized users and hosts can access this area, and access behavior must be audited and logged for traceability and monitoring. Access to the file transfer area from both internal and external networks requires strong password verification and multi-factor authentication before access to the target folder is granted. All access is logged.
[0079] 3) Storage Security: Storage devices within the file transfer area should possess high reliability and security, supporting encrypted data storage. A verification mechanism should be set up on the host to periodically verify the data on the storage devices, ensuring their normal operation and preventing data loss or damage during transmission.
[0080] 4) File Processing Rules: Within the file transfer area, explicit file processing rules are established, and a message queue system is used to handle file transfers. Message queues provide asynchronous transmission capabilities, reducing the complexity of direct communication between services and improving system stability and reliability. Data files within the intranet transfer area are sorted by transfer time and file name, and sequentially transferred to the external network area. The target database consumes data files by accessing the external network area. If a file fails to transfer or be consumed during intranet-extranet file transfer, a .bak file containing the corresponding data content is generated, and a data transfer failure alarm is sent to the data platform.
[0081] The original protocol header and information are stripped away at the intermediate layer of transmission, and the payload undergoes deep content inspection, including detection of malicious code and sensitive information. It is then re-encapsulated using a dedicated one-way protocol. Only the predefined simple SFTP protocol is allowed; the protocol implementation has been specially modified to remove all response mechanisms, and each session uses an independent encrypted channel.
[0082] Furthermore, the stripping of the original protocol header and deep content inspection are performed by the protocol processing module (transmission intermediate layer) of the one-way transmission device, which is a built-in function of the one-way transmission device. The processing object is the file transfer data stream uploaded to the one-way device's intranet via SFTP. After receiving the data stream, the protocol processing module automatically removes the SFTP / TCP protocol header, handshake information, and other control fields, retaining only the actual data portion of the file. This process does not process the file content; it only removes the encapsulation information of the transmission protocol, resulting in the file body content, i.e., the payload. The payload refers to the file content itself retained after stripping the protocol header, such as CSV, JSON, TXT, and other data files. Deep content inspection is performed by the same module to ensure the security of data to be transmitted across the network. This mainly includes: parsing the file content; scanning for malicious code characteristics (such as script fragments, executable instructions, etc.); checking sensitive information (such as accounts, passwords, IP addresses, etc.); and performing format integrity verification (such as whether it has been truncated or embedded with abnormal characters). Only after passing the detection does the subsequent one-way protocol re-encapsulation and transmission process begin.
[0083] Furthermore, the special modification involves a one-way simplification of the SFTP communication method. This is automatically completed at the protocol stack level by the protocol processing module of the one-way transmission device. The purpose is to ensure that the cross-network link only has one-way transmission capability and no form of reverse response capability. The protocol modification and response mechanism removal are executed by the SFTP protocol processing module built into the one-way transmission device, rather than being manually implemented by the application system or user. The modification process mainly masks session control and response messages in the SFTP protocol, including: ACK / OK and other acknowledgment messages; responses related to session negotiation and handshake; and reverse control information such as error receipts (e.g., STATUS messages). Before data enters the one-way link, the protocol processing module automatically: filters and discards all response messages from the outside; prohibits the generation of any form of return response to the source; and retains only the "minimum set of messages necessary for one-way transmission," forming a simplified transmission protocol. The modified SFTP is only responsible for file sending on a one-way transmission link and no longer supports: session status confirmation, transmission feedback, error receipt, or any reverse interaction process. In other words, it completely removes the "two-way interaction capability" and retains the "one-way push capability".
[0084] Data in a unified intermediate format is transferred to the file transfer area via SFTP (SSH File Transfer Protocol). Within the file transfer area, received files undergo security verification and necessary encryption. Verification includes checking file integrity (data integrity verification) and the legitimacy of the source; encryption involves selecting an appropriate encryption algorithm based on actual needs to encrypt the file, ensuring its security during storage and subsequent transmission.
[0085] Furthermore, a strong hashing algorithm, SHA-256, is used in the source database to calculate the hash value of the file to be transmitted. This calculation process generates a fixed-length hash value string, which serves as the file's unique "fingerprint." This hash value is recorded or transmitted along with the file before it is sent, serving as the base value for subsequent verification.
[0086] After a JSON file is transferred to the file transfer area via the SFTP encryption protocol, the file receiving system in the file transfer area immediately triggers a file verification process. This process first performs the same hash algorithm on the received file to generate a new hash value, which is then compared with the baseline value recorded before sending. If the two hash values match exactly, it means the file has not been tampered with during transmission and has maintained its integrity; the verification passes, and the file can safely enter the file transfer process. If the hash values do not match, it indicates that the file may have been corrupted or maliciously altered during transmission. The system should immediately issue an alert and may trigger error handling mechanisms (such as requesting a new file, logging errors, etc.).
[0087] The output data processed by the unified data middleware is saved and transmitted in the form of JSON files. These JSON files, transmitted via SFTP encryption to the file transfer area, include not only the source data but also the unified intermediate format file generated by the middleware after completing field mapping, type conversion, and format unification. The JSON format is characterized by its clear structure and highly readable fields, clearly expressing the standard field names, unified types, and unified formats generated by the middleware, making it very suitable as an intermediate file format for cross-network transmission.
[0088] S4: Obtain the encrypted data from the external network side of the file transfer area and transmit it to the target database.
[0089] In this embodiment, the data receiving end is a high-security zone, and the data flow strictly follows a unidirectional path from the low-security zone → intermediate layer → high-security zone. Any reverse transmission channel is blocked at both the physical and logical levels.
[0090] The target database iterates through all JSON files in the specified directory from the external network side of the file transfer area via SFTP, sorts them according to the file's sequence, processes the JSON file with the smallest sequence first, updates the corresponding content in the file to the corresponding database, and deletes the corresponding file in the file transfer area after the operation is completed.
[0091] A dual-buffer structure is adopted, physically separating read and write operations. Access timing is controlled through a semaphore mechanism, and new data is only allowed to be written after the buffer is empty. The low-security zone can only add tasks to the queue, while the high-security zone can only retrieve tasks from the queue. The intermediate scheduler implements unidirectional task transfer, and integrity check codes are added during transmission.
[0092] Semaphores are primarily used to establish a controlled access sequence between readers and writers, avoiding read / write contention. Specifically, two semaphores, "writable" and "readable," are set up for the dual buffers. The writing end must acquire the "writable" semaphore before writing data; this semaphore is only released after the buffer has been marked as read and reset. Similarly, the reading end needs to acquire the "readable" semaphore before reading; this semaphore is only released after the writing end completes the write and commits the buffer. Through the correspondence between these two semaphores, a fixed sequence of "write → commit → read → clear → write again" can be formed, ensuring deterministic access timing and preventing simultaneous or out-of-order access to the same buffer.
[0093] The intermediate scheduler is not a standalone add-on module, but rather a necessary control link for task interaction between the two security zones. Located between the low-security zone and the high-security zone, it receives write requests from the low-security zone and, after integrity verification, places the data into a readable buffer in the high-security zone according to a predetermined sequence. This scheduler is needed because the two security zones have different permissions and access scopes, preventing direct communication. By implementing unidirectional task transfer through the scheduler, it ensures that data can only flow from the low-security zone to the high-security zone, forming a unidirectional channel consistent with the security link, while preventing the high-security zone from being directly exposed to the low-security zone. Therefore, the intermediate scheduler's role is not only task transfer itself, but also a core node for performing access isolation, verification control, and timing management, forming a complete security link in the overall solution together with the dual buffer and semaphore mechanism.
[0094] The entire data transmission process is monitored, recording information such as who, when, where, why, and how all transmission operations are performed, and saving original data snapshots for post-event auditing and real-time monitoring of transmission traffic.
[0095] In summary, the data acquisition method of this invention demonstrates significant advantages in the digital transformation of the steel industry. First, it can be compatible with diverse heterogeneous production databases, acquiring real-time database changes, including equipment operating status, raw material consumption, and product quality parameters, providing a solid foundation for precise management. Second, this data acquisition method achieves automation and intelligence, significantly reducing manual intervention, improving data acquisition efficiency, and effectively reducing human error rates. Third, through big data analytics, enterprises can deeply mine data value, quickly identify production bottlenecks, optimize processes, and achieve efficient resource allocation and precise cost control. Finally, based on real-time data feedback, enterprises can flexibly adjust production strategies, quickly respond to market changes, enhance competitiveness, and drive the steel industry towards a green, intelligent, and efficient direction. In conclusion, the data acquisition method in the steel industry's digital transformation, with its efficient, accurate, and intelligent characteristics, injects strong momentum into the industry's transformation and upgrading.
[0096] Example 2 This embodiment discloses a one-way database data acquisition system based on file transfer, including: The data acquisition module is configured to collect data from multiple source databases to obtain multi-source data. The intermediate conversion module is configured to input the multi-source data into the unified data middleware for mapping and conversion to obtain data in a unified intermediate format. The intermediate isolation module is configured to: transmit the data of the unified intermediate format to the intranet side of the file transfer area, perform security verification and encryption on the data of the unified intermediate format in the file transfer area, and transmit the encrypted data unidirectionally to the extranet side of the file transfer area; The target security module is configured to: obtain the encrypted data from the external network side of the file transfer area and transmit it to the target database.
[0097] Example 3 The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method of Embodiment 1.
[0098] Example 4 The purpose of this embodiment is to provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the method of Embodiment 1.
[0099] The steps and methods involved in the apparatuses of Embodiments 3 and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0100] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0102] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A one-way database data acquisition method based on file transfer, characterized in that, include: Data is collected from multiple source databases to obtain multi-source data; The multi-source data are respectively input into the unified data middleware for mapping and conversion to obtain data in a unified intermediate format. The unified intermediate format data is transmitted to the intranet side of the file transfer area. The unified intermediate format data is then subjected to security verification and encryption in the file transfer area. The encrypted data is then transmitted unidirectionally to the extranet side of the file transfer area. The encrypted data is transmitted from the external network side of the file transfer area to the target database.
2. The one-way database data acquisition method based on file transfer as described in claim 1, characterized in that, The unified data middleware includes, in sequence, an access layer, an intelligent mapping engine, an execution engine, and a monitoring and management layer.
3. The one-way database data acquisition method based on file transfer as described in claim 2, characterized in that, The access layer is used to receive multi-source data, parse it, and extract the corresponding data types; the intelligent mapping engine is used to automatically identify and match similar data in different systems based on word vector models; the execution engine is used to perform field mapping and transformation on similar data based on the mapping library to obtain data in a unified intermediate format.
4. The one-way database data acquisition method based on file transfer as described in claim 3, characterized in that, The intelligent mapping engine is used to automatically identify and match similar data from different systems based on a word vector model. The multi-source data is preprocessed to obtain preprocessed multi-source data. Construct multi-dimensional features based on preprocessed multi-source data; The multi-dimensional features are input into the word vector model to obtain the vectorization of the multi-dimensional features, and the similarity of the multi-dimensional features is calculated respectively. Data with a similarity greater than a set threshold will be matched as similar data.
5. The one-way database data acquisition method based on file transfer as described in claim 3, characterized in that, The mapping library includes a data dictionary and mapping rules, which include field mapping rules, data type conversion rules, and data format conversion rules.
6. The one-way database data acquisition method based on file transfer as described in claim 1, characterized in that, The file transfer area is physically isolated from the source database and the target database respectively through optical unidirectional transmission equipment.
7. The one-way database data acquisition method based on file transfer as described in claim 1, characterized in that, The security verification uses hash values to verify the integrity of the data.
8. A one-way database data acquisition system based on file transfer, characterized in that, include: The data acquisition module is configured to collect data from multiple source databases to obtain multi-source data. The intermediate conversion module is configured to input the multi-source data into the unified data middleware for mapping and conversion to obtain data in a unified intermediate format. The intermediate isolation module is configured to: transmit the data of the unified intermediate format to the intranet side of the file transfer area, perform security verification and encryption on the data of the unified intermediate format in the file transfer area, and transmit the encrypted data unidirectionally to the extranet side of the file transfer area; The target security module is configured to: obtain the encrypted data from the external network side of the file transfer area and transmit it to the target database.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the one-way database data acquisition method based on file transfer as described in any one of claims 1-7.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the one-way database data acquisition method based on file transfer as described in any one of claims 1-7.