Cross-network heterogeneous data fusion processing method, device, equipment and medium

By constructing a cross-network data transfer channel, a protocol adaptation module, a multimodal parsing engine, and a semantic governance center, the problem of fusion processing of multi-source heterogeneous traffic data in a multi-network environment was solved, realizing secure and efficient data fusion and unified services, and improving system operating efficiency and data utilization capabilities.

CN121934780APending Publication Date: 2026-04-28PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-15
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing traffic data systems, due to isolated network environments, highly heterogeneous data structures, and a lack of unified data governance and hierarchical storage service mechanisms, struggle to achieve secure, efficient fusion processing and unified external services for multi-source heterogeneous traffic data in multi-network environments.

Method used

A cross-network data transfer channel is constructed for one-way isolated transmission. A protocol adaptation module and a multimodal parsing engine are used to structure the data and align it with the spatiotemporal reference. Metadata registration and semantic tagging are performed through a semantic governance center to generate standardized data after governance. The data is written into a hierarchical storage architecture based on time attributes and access frequency characteristics. Finally, data access services are provided to the outside world through a unified data service gateway.

Benefits of technology

It enables secure migration and unified parsing of multi-source heterogeneous data in an isolated network environment, reduces the complexity of data processing and access, and improves the overall operating efficiency and data utilization capability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121934780A_ABST
    Figure CN121934780A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hierarchical storage, can be applied to a multi-source heterogeneous data fusion processing scene in a multi-network isolation environment such as smart traffic, and discloses a cross-network heterogeneous data fusion processing method, device and equipment and a medium. Comprising the following steps: realizing safe migration of multi-source heterogeneous original data in an isolated network environment by constructing a cross-network data ferry channel, and completing protocol adaptation, multi-modal analysis, structured processing and space-time alignment of the data in an internal safe area; and performing semantic governance on the processed data, writing the processed data into a hierarchical storage architecture according to time attributes and access popularity, and finally providing a data calling service to the outside through a unified data service gateway. Through secure cross-network transmission, multi-source data unified analysis and semantic management, and in combination with hierarchical storage and unified service output, efficient fusion and ordered management of heterogeneous data are realized, the data processing and access complexity is reduced, and the overall operation efficiency and data utilization capability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hierarchical storage technology, and in particular to a method, apparatus, device and medium for cross-network heterogeneous data fusion processing. Background Technology

[0002] With the deepening of smart city construction, traffic management is increasingly reliant on multi-source data. Traffic operation monitoring, command and dispatch, and situation analysis require the simultaneous use of data resources from different network environments, including public security networks, government networks, industry networks, and the internet. However, existing traffic data systems generally suffer from clear network boundaries and strict isolation mechanisms. Different network areas are typically protected by physical or logical isolation, resulting in complex cross-network data transfer paths and low transmission efficiency. Existing cross-network data exchange methods largely rely on manual export, timed transfers, or offline approval processes, leading to long data migration cycles and making it difficult to support the real-time and continuous data requirements of traffic management scenarios.

[0003] Meanwhile, existing traffic data is highly heterogeneous in terms of source systems, data structures, and representation formats. Data generated by different systems lacks unified standards in terms of format, field semantics, time reference, and spatial coordinates. Existing technologies typically employ different processing links for structured, semi-structured, and unstructured data, lacking a unified access and parsing mechanism. This results in data requiring extensive customized processing after entering the platform before it can participate in fusion analysis. This heterogeneous processing approach not only increases system complexity but also makes it difficult to process cross-source and cross-type data consistently within the same process, thus hindering the comprehensive utilization of multi-source traffic data.

[0004] Furthermore, existing traffic data platforms still fall short in data governance, storage organization, and external services. Data from different sources lacks unified metadata management and semantic standards, resulting in inconsistent data naming, unquantifiable quality, and difficulty in tracking data evolution, impacting data credibility and usability. With the continuous growth of data volume, the existing storage system has failed to effectively stratify data based on temporal and access characteristics, leading to a mismatch between storage costs and access performance due to the mixed storage of hot and cold data. Simultaneously, data is provided externally in a fragmented manner, lacking a unified data service entry point. Data access requires adaptation to different storage locations and interface types, increasing system integration and maintenance burdens and limiting the efficient sharing and application of traffic data resources. Summary of the Invention

[0005] The main objective of this invention is to provide a cross-network heterogeneous data fusion processing method, apparatus, device, and storage medium, aiming to solve the technical problems in existing technologies for traffic data management, which are difficult to achieve secure and efficient fusion processing and unified external service of multi-source heterogeneous traffic data in a multi-network environment due to network environment isolation, highly heterogeneous data structure, and lack of unified data governance and hierarchical storage service mechanisms.

[0006] To achieve the above objectives, the present invention provides a cross-network heterogeneous data fusion processing method, comprising: Construct a cross-network data transfer channel between the external network area and the internal security area, and use a one-way isolation transmission mechanism to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area; The multi-source heterogeneous raw data is received through the protocol adaptation module in the internal security area, and the multi-modal parsing engine is called to perform structured feature extraction and spatiotemporal benchmark alignment on the multi-source heterogeneous raw data to generate standardized data packets. The semantic governance center performs metadata registration, semantic tagging, and synonym mapping operations on the standardized data packets to generate standardized data after governance. The time attributes and access popularity characteristics of the standardized data after governance are analyzed, and the standardized data after governance is written into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. The unified data service gateway responds to data call requests and outputs the standardized, governed data stored in the hierarchical storage architecture.

[0007] Furthermore, to achieve the above objectives, the present invention provides a cross-network heterogeneous data fusion processing apparatus, comprising: The cross-network data transfer module is used to build a cross-network data transfer channel between the external network area and the internal security area, and to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area using a one-way isolation transmission mechanism. The multimodal parsing module is used to receive the multi-source heterogeneous raw data through the protocol adaptation module in the internal security area, and call the multimodal parsing engine to perform structured feature extraction and spatiotemporal benchmark alignment processing on the multi-source heterogeneous raw data to generate standardized data packets; The semantic governance module is used to perform metadata registration, semantic tag annotation, and synonym mapping operations on the standardized data packets through the semantic governance center to generate standardized data after governance. The hierarchical storage scheduling module is used to analyze the time attributes and access popularity characteristics of the standardized data after governance, and write the standardized data after governance into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. The unified data service module is used to respond to data call requests through the unified data service gateway to output the standardized data after governance stored in the hierarchical storage architecture.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a cross-network heterogeneous data fusion processing program stored in the memory and executable on the processor, wherein when the cross-network heterogeneous data fusion processing program is executed by the processor, it implements the steps of the cross-network heterogeneous data fusion processing method as described above.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a cross-network heterogeneous data fusion processing program, wherein when the cross-network heterogeneous data fusion processing program is executed by a processor, it implements the steps of the cross-network heterogeneous data fusion processing method as described above.

[0010] Beneficial Effects: This invention relates to the field of hierarchical storage technology and can be applied to multi-source heterogeneous data fusion processing scenarios in multi-network isolated environments such as intelligent transportation. It discloses a cross-network heterogeneous data fusion processing method, apparatus, device, and medium, including: constructing a cross-network data transfer channel to achieve secure migration of multi-source heterogeneous raw data in an isolated network environment; completing data protocol adaptation, multimodal parsing, structured processing, and spatiotemporal alignment within an internal secure area; performing semantic governance on the processed data and writing it into a hierarchical storage architecture based on time attributes and access frequency; and finally providing data access services externally through a unified data service gateway. This invention achieves efficient fusion and orderly management of heterogeneous data through secure cross-network transmission, unified parsing and semantic governance of multi-source data, combined with hierarchical storage and unified service output, reducing data processing and access complexity, and improving the overall system operating efficiency and data utilization capabilities. Attached Figure Description

[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a cross-network heterogeneous data fusion processing method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the cross-network heterogeneous data fusion processing method of the present invention; Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the cross-network heterogeneous data fusion processing device of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The cross-network heterogeneous data fusion processing method provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can use the client to build a cross-network data transfer channel to achieve secure migration of multi-source heterogeneous raw data in an isolated network environment. Within its internal secure area, it completes protocol adaptation, multimodal parsing, structured processing, and spatiotemporal alignment of the data. Semantic governance is performed on the processed data, and it is written into a hierarchical storage architecture based on time attributes and access frequency. Finally, data access services are provided externally through a unified data service gateway. This invention achieves efficient integration and orderly management of heterogeneous data through secure cross-network transmission, unified parsing and semantic governance of multi-source data, combined with hierarchical storage and unified service output. This reduces data processing and access complexity and improves overall system efficiency and data utilization. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The following detailed description of specific embodiments further illustrates this invention.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the cross-network heterogeneous data fusion processing method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the cross-network heterogeneous data fusion processing method proposed in this invention includes the following steps: S10, construct a cross-network data transfer channel between the external network area and the internal security area, and use a one-way isolation transmission mechanism to migrate the multi-source heterogeneous raw data generated in the external network area to the internal security area; In this embodiment, in a system environment operating with multiple isolated networks, there is typically no direct communication path between the external network area and the internal security area. Data generated in the external network area cannot directly enter the internal security area for processing. To address this issue, a cross-network data transfer channel is constructed between the two network areas. This channel is deployed at the network boundary and establishes independent connections with both the external network area and the internal security area, serving to carry out data transfer processes across isolation boundaries. The purpose of the cross-network data transfer channel is to provide a controlled data migration path, rather than establishing a conventional network communication link.

[0016] Multi-source heterogeneous raw data refers to data sets generated by different systems, data formats, and generation mechanisms. These data exist in their raw form within the external network area and have not yet undergone unified modeling or security processing. A one-way isolation transmission mechanism is used to limit the direction of data flow, allowing data to be transmitted only from the external network area to the internal secure area. This mechanism eliminates reverse communication capabilities through physical or logical means, ensuring that the internal secure area does not expose any external access channels when receiving data. Through this one-way isolation transmission mechanism, multi-source heterogeneous raw data migrates from its source to the internal secure area in a controlled manner, ensuring that data migration behavior is consistent with the network isolation strategy.

[0017] This embodiment constructs a cross-network data transfer channel between the external network area and the internal security area, and uses a one-way isolation transmission mechanism to complete the migration of multi-source heterogeneous raw data. This allows data to enter the internal security area without violating the network isolation boundary, avoiding the security risks caused by two-way communication and improving the controllability and stability of the cross-network data migration process.

[0018] S20, in the internal security area, the multi-source heterogeneous raw data is received through the protocol adaptation module, and the multi-modal parsing engine is called to perform structured feature extraction and spatiotemporal benchmark alignment processing on the multi-source heterogeneous raw data to generate standardized data packets; In this embodiment, within the internal security zone, the multi-source heterogeneous raw data has already completed cross-network migration. However, the source systems differ significantly, with inconsistent communication methods, data encapsulation formats, and field organization rules, making them uncomprehending to be directly understood by subsequent processing units. To achieve unified access, a protocol adaptation module is set up within the internal security zone. This module identifies the communication protocol type used when data accesses the data and performs protocol-level parsing and conversion based on the identification results, enabling data from different protocol environments to enter the internal processing flow in a unified manner. The protocol adaptation module focuses on the communication and transport layers, without altering the data content itself, only eliminating access barriers caused by protocol differences.

[0019] After receiving the data, the multimodal parsing engine performs content-level parsing processing on the multi-source heterogeneous raw data. The engine decomposes the raw data according to its organizational structure, transforming unstructured or semi-structured expressions into structured feature expressions, thus converting the data from its raw form into a feature set with fields, attributes, and semantic meaning. During structured feature extraction, different types of data are processed through corresponding parsing paths to ensure consistency in the extraction results at the expression level. Subsequently, the multimodal parsing engine performs spatiotemporal benchmark alignment processing on the parsed feature data. By unifying the time benchmark and spatial reference system, it eliminates differences in temporal precision and spatial coordinates between different data sources, enabling the multi-source data to have the basic conditions for alignment and comparison. After completing the above processing, the parsing results are encapsulated into a unified format data set, forming a standardized data package.

[0020] This embodiment introduces a protocol adaptation module and a multimodal parsing engine within the internal security zone, enabling the elimination of protocol differences in multi-source heterogeneous raw data during the receiving phase and achieving structured expression and spatiotemporal benchmark unification during the parsing phase. This provides a consistent and alignable data foundation for subsequent data processing and reduces the complexity of heterogeneous data mixing processing.

[0021] S30, The semantic governance center performs metadata registration, semantic tag annotation and synonym mapping operations on the standardized data packets to generate standardized data after governance; In this embodiment, although standardized data packets have a unified form after being structured and spatiotemporally aligned, they still suffer from semantic differences, naming inconsistencies, and implied business meanings, which can easily lead to misunderstandings if used directly. To address this issue, a semantic governance center is introduced to centrally process standardized data packets. The semantic governance center is responsible for the semantic unification and governance control of data, and its processing objects are data content that has already achieved standardized formatting.

[0022] In the semantic governance center, metadata registration is first performed on data packets. By establishing a unique identifier for each packet and recording its source path, generation process, and basic attributes, the data becomes traceable and manageable in subsequent use. The result of metadata registration forms the basic descriptive information of the data, providing contextual support for semantic processing. Subsequently, the semantic governance center semantically tags the data content, associating abstract data content with explicit business semantics based on the meaning of data fields, business attributes, and contextual relationships. This transforms data from technical expression into a semantic expression with business relevance.

[0023] After completing semantic tagging, the Semantic Governance Center further performs synonym mapping on the data fields. By uniformly mapping fields with the same meaning but different names in different systems, semantic ambiguity and naming issues are eliminated. This process uses a standard semantic system as a reference to merge multi-source fields under a unified semantic expression, ultimately forming semantically consistent and named data content, thus generating standardized data after governance.

[0024] This embodiment uses a semantic governance center to register metadata, label semantic tags, and map synonyms for standardized data packets. This enables data to achieve semantic consistency and traceability management while maintaining structural uniformity, reducing the risk of ambiguity in the understanding and use of multi-source data, and improving data availability and governance consistency.

[0025] S40, Analyze the time attributes and access popularity characteristics of the standardized data after governance, and write the standardized data after governance into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. In this embodiment, after the standardized data has undergone semantic unification, its access value will vary over time and with changes in usage frequency. If a unified storage method is adopted, it can easily lead to a decrease in access efficiency for high-value data or high-cost resource consumption for low-value data. To solve this problem, the standardized data after governance is analyzed for time attributes and access popularity characteristics, which serve as the basis for storage hierarchy division.

[0026] The time attribute characterizes the timing of data generation or updates. By parsing the timestamp information carried in the data packet, the validity period of the data relative to the current time is determined. This attribute reflects the freshness and timeliness requirements of the data and is an important basis for judging whether the data requires high-performance storage. Access frequency characteristics describe the frequency with which data is accessed within a certain time window. By statistically analyzing the number of requests and access distribution for the data in historical access records, its actual usage intensity is quantified. The time attribute and access frequency characteristics together constitute the basic dimensions for judging the value of data.

[0027] After completing the attribute analysis, the data is written into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer, and a cold data storage layer, based on the analysis results. The hot data storage layer is used to hold data with high timeliness and high access frequency, the warm data storage layer is used to hold data with medium access demand or data in transitional phases, and the cold data storage layer is used to hold data with low access frequency and long retention periods, thereby realizing the tiered utilization of storage resources.

[0028] This embodiment analyzes the time attributes and access frequency characteristics of the standardized data after governance and performs hierarchical writing, so that the allocation of storage resources matches the actual use value of the data, improves the access efficiency of high-frequency data while reducing the storage cost of low-frequency data, and achieves a balance between storage performance and resource utilization.

[0029] S50, through a unified data service gateway, responds to data call requests to output the standardized data after governance stored in the hierarchical storage architecture.

[0030] In this embodiment, after tiered storage is completed, if the standardized data after governance lacks a unified external service exit, it will still be directly accessed by different systems in their own ways, which can easily lead to problems such as fragmented access paths, inconsistent access control, and untraceable call behavior. To solve this problem, a unified data service gateway is used to centrally carry out data output and access control capabilities, thereby decoupling data calls from the storage layer.

[0031] The unified data service gateway receives data access requests, parses the request content, and determines the identifier information of the required data. Data access requests reflect the access needs of external systems or internal applications for specific data content; their request format can be manifested as query conditions, data ranges, or return format constraints. By unifying the access of these requests, all access behaviors converge at a single logical entry point, avoiding direct penetration through the hierarchical storage architecture.

[0032] After parsing the data retrieval request, the gateway retrieves the corresponding standardized data from the tiered storage architecture based on the data identifier indicated in the request. At this stage, the tiered storage architecture only handles data transport; the gateway centrally manages the data location, read order, and return organization, thus shielding the underlying storage differences. The retrieved data is then formatted and output, allowing the caller to obtain results in a unified manner.

[0033] This embodiment uses a unified data service gateway to centrally respond to data call requests and output standardized data after governance, thereby unifying the data access path, making the storage layer transparent to the outside world, reducing the coupling between systems, and improving the controllability and consistency of data calls.

[0034] In one embodiment, step S10 includes: S101, Deploy an external isolation buffer pool in the external network area to verify the identity and credibility of the data source that generates multi-source heterogeneous raw data, and write the verified multi-source heterogeneous raw data into the external isolation buffer pool for temporary storage. S102, activate the physical unidirectional optical gate component in the unidirectional isolation transmission mechanism, establish a unidirectional optical signal transmission link from the external isolation buffer pool to the internal security area, read the multi-source heterogeneous raw data in the external isolation buffer pool through the unidirectional optical signal transmission link, and deliver the multi-source heterogeneous raw data to the internal isolation buffer pool deployed in the internal security area in the form of a non-network protocol physical signal. S103, trigger a policy-driven automatic approval process in the internal isolation buffer pool, and determine whether to allow the multi-source heterogeneous original data based on the data usage attributes and security level attributes of the multi-source heterogeneous original data. S104, if the automatic approval process determines to allow the data to pass, the national cryptographic algorithm encryption module is called to identify sensitive entity fields in the multi-source heterogeneous original data and perform dynamic desensitization and encryption processing, and generate a data flow audit log.

[0035] In this embodiment, a cross-network data transfer channel is deployed between the external network area and the internal security area. This channel is used to migrate multi-source, heterogeneous raw data generated in the external network area to the internal security area while maintaining network isolation boundaries. The distinction between the external network area and the internal security area is based on network boundaries and security control strength. The external network area focuses on supporting business systems and data sources with open access and broad connectivity, while the internal security area focuses on supporting business systems and data resources with sensitive data processing and controlled services. The external network area is typically located in a position where it can interconnect with the government extranet, video conferencing network, or the internet. It features external interface exposure, cross-domain access, and third-party system access. Access control primarily relies on account systems, interface authentication, and boundary protection, emphasizing access scale and availability. The internal security area is typically located within a dedicated security domain, employing stricter network segmentation and access whitelist control. It emphasizes minimal data exposure, auditable operations, minimal permissions, and prevention of external connections. Communication paths, port sets, and maintenance entry points between internal systems are centrally managed, and data import / export typically requires a controlled channel.

[0036] Taking the transportation sector as an example, the external network area can include the network environment where data sources such as real-time traffic conditions from internet-based navigation platforms, ride-hailing platform trajectories, and shared bicycle parking heat maps reside, as well as the network environment where data services such as municipal construction plans, public transportation operation announcements, and weather warnings reside from the government extranet. These environments share the common characteristic of needing to provide data interfaces to multiple systems or receive external calls, with network links involving cross-unit interconnections, cross-system calls, and cross-protocol access. The internal security area can include the network environment where sensitive data and core applications within the public security traffic management network reside, such as checkpoint vehicle passage records, electronic police violation information, accident alarm logs, and command and dispatch system analysis data. These systems typically require isolation from the internet, demanding finer-grained control and traceability over access subjects, access times, and access content. Furthermore, data services provided externally must undergo approval, anonymization, encryption, and auditing before being output.

[0037] In practical deployments, the boundary between the two is typically defined by security isolation devices and policies. For example, the system on the external network side will first allow the multi-source heterogeneous raw data it generates to fall into the access area or buffer near the boundary. The internal security area side only receives the incoming data payload through a controlled cross-network data transfer channel. After entering the system, the protocol adaptation module and multimodal parsing engine complete the structured and spatiotemporal alignment processing, preventing the internal security area from directly exposing network ports or service interfaces that can be accessed externally. This distinction allows the external network area to bear the broad connectivity attribute of the data source, while the internal security area bears the strong control attribute of sensitive data processing and unified services, thus forming a secure boundary that can be implemented when traffic data is integrated across domains.

[0038] A one-way isolation transmission mechanism is used to limit the transmission direction, ensuring that data only enters the internal secure area along a predetermined path, preventing backflow communication from the internal secure area from reaching the external network area. Multi-source heterogeneous raw data encompasses data payloads from different sources and in different formats. Before entering the internal secure area, the migration process requires source trust verification, isolation delivery, policy approval, and sensitive field processing to ensure that data received in subsequent processing stages has consistent constraints in terms of source, compliance, and traceability.

[0039] An external isolation buffer pool is deployed in the external network area for data temporary storage and batch transfer. The deployment location and access boundary of the external isolation buffer pool are used to absorb fluctuations from upstream multi-source inputs and to uniformly place multi-source heterogeneous raw data into the same buffer medium before entering the unidirectional isolation transmission mechanism. Identity credibility verification is used to constrain the access qualifications of the data source side. The verification object is the data source that generated the multi-source heterogeneous raw data. The verification content can establish consistent judgment rules based on access credentials, device identification, certificate chain, or signature verification. The verification result is used to form a binary decision of granting or denying. The act of writing to the external isolation buffer pool not only realizes temporary storage but also serves to solidify the input snapshot, avoiding disturbances caused by changes and retransmissions of upstream data sources in the subsequent delivery stage. The temporarily stored data units can carry necessary attributes such as data source identification, generation time, and purpose tags, providing data basis for subsequent approval and judgment.

[0040] The physical unidirectional optical gate component is a hardware implementation of the unidirectional isolation transmission mechanism, used to establish a unidirectional optical signal transmission link from the external isolation buffer pool to the internal secure area. This unidirectional optical signal transmission link replaces network protocol handshakes with physical direction constraints, ensuring that the transmission channel does not provide reverse backhaul capability at the electrical and protocol levels. Reading multi-source heterogeneous raw data from the external isolation buffer pool requires converting the data units in the buffer pool into a sequence adapted for optical signal transmission according to a preset encoding method, forming a strict unidirectional pipeline between the reading and delivery actions. Multi-source heterogeneous raw data is delivered to the internal isolation buffer pool in the form of physical signals that are not part of the network protocol. The delivery process emphasizes that it does not rely on traditional network protocol stacks for session establishment, confirmation, and response, thus avoiding the formation of exploitable bidirectional interaction surfaces at the isolation boundary. The internal isolation buffer pool is deployed on the internal secure area side to receive the delivered data and perform re-writing or re-encapsulation. The internal isolation buffer pool also performs traffic shaping at the isolated receiving end, ensuring that subsequent approval processes are based on a stable data set for judgment.

[0041] The strategy-driven automated approval process is triggered within an internal isolation buffer. Triggering conditions can be based on internal states such as data arrival events or buffer level thresholds. Strategy-driven approaches emphasize that approval rules are expressed through configurable strategies, enabling approval logic to combine and determine data usage and security classification attributes. The data usage attribute expresses the intended use of multi-source, heterogeneous raw data entering the internal secure area, while the security classification attribute expresses the hierarchical constraints of the data within the security system. The automated approval process maps these two types of attributes to release rules and outputs a decision on whether to release the data. The decision process can include attribute consistency checks, whitelist matching, cross-constraint verification of usage and security classification, and rejection rules for abnormal attributes, ensuring that release decisions have an interpretable rule path.

[0042] When the automated approval process determines to release the data, the national cryptographic algorithm encryption module identifies sensitive entity fields in the multi-source heterogeneous original data and performs dynamic desensitization and encryption processing. Sensitive entity field identification is used to locate data elements requiring protection. Identification rules can be established based on field name, field type, content mode, or a predefined set of sensitive entities. The identification results are used to determine the scope of desensitization and encryption. Dynamic desensitization is used to mask, truncate, replace, or generalize sensitive entity fields without compromising the usability of the data structure. Its dynamic nature is reflected in the fact that desensitization rules can be adjusted according to changes in data usage attributes and security level attributes. Encryption processing applies national cryptographic algorithm protection to sensitive entity fields that still need to maintain the recoverability of their original values, ensuring that sensitive information remains in encrypted form during internal transfer and storage. Data transfer audit logs are generated during the release and processing. The log content records external isolation buffer writes, unidirectional optical signal delivery, internal isolation buffer reception, automated approval determination, sensitive entity field identification result summaries, desensitization and encryption action identifiers, and associated time information and responsible entity identifiers, forming traceable evidence of cross-network migration links.

[0043] This embodiment verifies and temporarily stores the identity credibility of multi-source heterogeneous raw data through an external isolation buffer pool, enabling source screening and forming stable input batches before entering the isolation link. This reduces the impact of untrusted data sources and input jitter on cross-network migration. By using a physical unidirectional optical gate component and a unidirectional optical signal transmission link, data delivery is limited to unidirectional transmission of physical signals without network protocols, reducing the bidirectional interaction surface at the isolation boundary and improving the security of data migration between the external network area and the internal security area. Through a policy-driven automatic approval process triggered by the internal isolation buffer pool and release decisions based on data usage attributes and security level attributes, the migration process has a configurable compliance control path. After release, combined with sensitive entity field identification, dynamic desensitization and encryption processing, and data flow audit log generation by the national cryptographic algorithm encryption module, data entering the internal security area has auditable closed-loop constraints in terms of sensitive information protection and flow traceability.

[0044] In one embodiment, step S20 above includes: S201, the protocol adaptation module is activated in the internal security area to automatically identify the communication protocol type of the data source and establish a data transmission connection according to the communication protocol type to receive the multi-source heterogeneous raw data; S202, invoke the multimodal parsing engine to identify the data format of the multi-source heterogeneous raw data; S203, based on the identified data format, the multimodal parsing engine is invoked to perform differentiated parsing; when the data format is structured data, a standardized data model mapping is performed; when the data format is semi-structured data, the meaning of the fields is inferred using a pattern inference algorithm; when the data format is unstructured data, an artificial intelligence recognition model is invoked to perform structured feature extraction. S204, The multimodal parsing engine is used to perform spatiotemporal reference alignment processing on the parsed data generated after differential parsing to generate spatiotemporally aligned data; S205, the spatiotemporally aligned data is encapsulated in a unified format to generate a standardized data packet.

[0045] In this embodiment, the internal security zone is used to carry out controlled access and parsing processing of multi-source heterogeneous raw data. The protocol adaptation module, as the access entry point, is responsible for identifying the communication protocol type and establishing a data transmission connection without pre-defining the data source form, enabling multi-source heterogeneous raw data to enter the internal security zone in an acceptable connection form. Starting the protocol adaptation module includes loading protocol identification rules, initializing connection parameters, and enabling the connection management state machine. Protocol identification is performed based on the communication protocol type of the data source. The communication protocol type can be identified based on port characteristics, handshake sequences, header fields, connection keep-alive characteristics, or pre-registered data source configurations. The identification result is used to select a matching connection driver and decoder within the protocol adaptation module. A data transmission connection is established according to the communication protocol type. Connection establishment includes session parameter negotiation, connection authentication, transmission channel creation, and receive buffer configuration, enabling the protocol adaptation module to continuously receive multi-source heterogeneous raw data and form an input stream or batch that can be parsed subsequently. During the reception process, basic processing can be performed on fragmentation, retransmission, out-of-order delivery, and packet loss to ensure that the data boundaries delivered to the multimodal parsing engine can be determined.

[0046] The multimodal parsing engine is used to perform format recognition, differential parsing, and structured feature extraction on multi-source heterogeneous raw data. The actions of calling the multimodal parsing engine to identify data formats include reading sample fragments, extracting structural fingerprints, and outputting the data format category. Data formats cover three categories: structured data, semi-structured data, and unstructured data. Data format recognition can determine structured data based on structural signals such as field separators and fixed column widths; semi-structured data based on key-value pair hierarchy, tag structure, and pattern drift features; and unstructured data based on binary container headers, media encoding identifiers, natural language text features, or image pixel array features. The recognition results serve as routing conditions for differential parsing, avoiding the loss of structure or parsing failure caused by using the same parsing path for different data formats.

[0047] Differential parsing is performed based on the identified data format. This parsing converts different data formats into uniformly processable parsed data and provides a standardized set of fields and temporal and spatial elements for subsequent spatiotemporal benchmark alignment. When the data format is structured, a standardized data model mapping is performed. This mapping aligns column names, data types, enumerated values, and unit expressions in the structured data based on the target field set and field type constraints. The mapping process may include field renaming, type conversion, application of missing value imputation rules, and primary key candidate generation, ensuring a consistent structure at the field level in the output parsed data. When the data format is semi-structured, a pattern inference algorithm is used to infer field meanings. This algorithm establishes candidate field semantics based on the key set, hierarchical path, value distribution, and contextual co-occurrence relationships. The inference results are used to map dynamic key paths to a stable set of fields and determine field meanings. Determining field meanings may include field role determination, numerical dimension inference, and temporal and spatial field identification, enabling stable parsed data to be generated even when the structure of semi-structured data is not fixed. When the data format is unstructured, the AI ​​recognition model is invoked to perform structured feature extraction. The AI ​​recognition model is used to extract structured feature elements from images, videos, audio, or text. Structured feature extraction includes object detection and attribute extraction, event fragment recognition, text entity extraction and relation extraction, or acoustic fragment annotation. The output feature elements are entered into the parsed data in a field-based form, enabling the unstructured data to participate in subsequent alignment and encapsulation processes.

[0048] The multimodal parsing engine performs spatiotemporal benchmark alignment on the parsed data. This alignment unifies temporal and spatial representations, enabling parsed data from different data sources to be correlated under the same temporal and spatial reference. Spatiotemporal benchmark alignment includes timestamp normalization and spatial coordinate normalization. Timestamp normalization converts time fields of different precisions to a uniform precision and handles time zone offsets, leap second corrections, and acquisition delay compensation. Spatial coordinate normalization converts different coordinate systems or spatial index representations to a uniform coordinate benchmark and performs projection transformations, map matching, or grid index mapping. The alignment process requires preserving the mapping relationship before and after alignment for traceability. The generated spatiotemporally aligned data includes a unified time field and a unified spatial field in its field set, allowing subsequent fusion and storage stages to directly rely on these fields for consistent processing.

[0049] After spatiotemporal alignment, the data is encapsulated in a unified format to form standardized data packets. This unified format encapsulation involves defining the encapsulation structure, filling in necessary metadata, and outputting a transmittable and storable data packet carrier. The encapsulation structure of the standardized data packet can include a data payload area and a metadata area. The data payload area carries the field-based records of the spatiotemporally aligned data, while the metadata area carries the data source identifier, data format identifier, parsing path identifier, alignment benchmark identifier, and verification information, making the standardized data packet verifiable and reusable. The unified format encapsulation can also consistently configure field order, field encoding, compression methods, and fragmentation strategies, ensuring that the standardized data packet maintains clear boundaries and consistent parsing during cross-module transmission. This provides a stable input object for subsequent metadata registration, semantic tag annotation, and synonym mapping operations in the semantic governance center.

[0050] This embodiment uses a protocol adaptation module to automatically identify communication protocol types and establish data transmission connections, enabling multi-source heterogeneous raw data to form a unified and accessible receiving format within an internal secure area, reducing reliance on a single protocol or access method. Through a multimodal parsing engine, it identifies and differentiates data formats, integrating standardized data model mapping for structured data, pattern inference of field meanings for semi-structured data, and extraction of structured features from unstructured data into the same processing chain. This ensures that parsed data can be processed consistently in terms of field structure and semantic carrying methods. By performing spatiotemporal benchmark alignment processing on the parsed data and generating spatiotemporally aligned data, and then encapsulating the spatiotemporally aligned data in a unified format to generate standardized data packets, cross-source data can be directly aligned in time and space and output using a unified carrier. This provides stable input for subsequent governance and hierarchical storage, reducing the compensation costs for heterogeneous formats and spatiotemporal differences in subsequent processing.

[0051] In one embodiment, step S30 above includes: S301, The semantic governance center generates a unique identifier for the standardized data packet and records the source path, processing process and usage scenario of the standardized data packet to form a data lineage graph; S302, The semantic governance center calls the natural language processing model and expert policy library to annotate the standardized data packets with business semantic tags; S303, retrieve a preset standard terminology dictionary through the semantic governance center and map the non-standard field names in the standardized data packet to standard names; S304, The semantic governance center uses a dynamic data quality scoring model to analyze the quality score of the standardized data packet, and triggers an alarm when the quality score is lower than a preset threshold. S305, Based on the data lineage map, the business semantic tags, the standard names, and the quality scores, generate standardized data after governance.

[0052] In this embodiment, the semantic governance center is responsible for the structured registration, semantic normalization, and quality monitoring of standardized data packets after they enter the unified governance domain. The input object is the standardized data packet, and the output object is the standardized data after governance. The output result needs to be directly usable by subsequent storage tiering and data service calls. Therefore, traceable information, searchable information, and measurable information are formed simultaneously in the processing link. A unique identifier is generated for the standardized data packet to establish a stable reference anchor. The source of the unique identifier can be an encoded string generated by the semantic governance center according to preset generation rules, or it can be calculated by combining the key fields of the standardized data packet and introducing random perturbation or sequence number expansion in case of conflict. After generation, the unique identifier is written into the metadata registration item and associated with the internal record of the standardized data packet, so that any subsequent retrieval, verification, alarm, and distribution can locate the same object through the unique identifier. Recording the source path is used to clarify which data source the standardized data packet comes from and which access channels and relay nodes it passes through. The source path can be obtained from fields such as source identifier, link identifier, and collection task identifier carried by the access side, or it can be generated by the semantic governance center based on the receiving port, authentication subject, access protocol, and file landing point. Recording the processing process is used to depict the set of processing actions that the standardized data packet undergoes before entering the semantic governance center or within the governance center. The processing process can be solidified into the metadata registration item in the form of processing operator name, version number, key parameter summary, and changes in input and output fields. Recording the usage scenario is used to bind the standardized data packet with the consumer's intent. The usage scenario can come from the purpose tag, data subject, and call domain identifier carried by the business system when delivering data, or it can be inferred and backfilled by the semantic governance center based on business semantic tags. Source path, processing process, and usage scenario are used together to form a data lineage graph. The data lineage graph can be organized using a node and edge structure. Nodes are used to represent entities such as data sources, datasets, standardized data packages, and standardized data after governance. Edges are used to represent relationship types such as migration, processing, derivation, and aggregation. The edge attributes record the time, responsible party, and processing action, so that when quality alarms, disputes over standards, or field tracing occur, the upstream source and intermediate processing links can be located along the graph.

[0053] Business semantic tags are used to describe the category, theme, entity type, or event type of standardized data packets in terms of business meaning. The semantic governance center calls a natural language processing model and an expert policy library to complete the annotation. The input of the natural language processing model can be text or semi-structured information such as field names, field comments, value examples, data theme descriptions, and usage scenarios. The output is a set of candidate semantic tags and a confidence distribution. The expert policy library is used to provide interpretable supplementary constraints and correction rules. The policy content can be expressed as keyword matching rules, regular expression patterns, field combination judgment rules, numerical range and unit verification rules, regional coding or industry coding mapping rules, etc. In terms of the calling relationship, the semantic governance center can first have the natural language processing model provide candidate tags, and then have the expert policy library filter, rearrange, or cover the candidate tags. Alternatively, the expert policy library can first directly provide tags based on strong rules, and then the natural language processing model can complete the tags for uncovered fields. The annotation results are written back to the metadata registration item and associated with a unique identifier, so that the business semantic tags can be directly used for subsequent retrieval, hierarchical storage, and data call request parsing.

[0054] Synonym mapping is used to eliminate the fragmentation of fusion and query caused by differences in field naming. The semantic governance center retrieves a pre-defined standard terminology dictionary and maps non-standard field names in the standardized data package to standard names. The standard terminology dictionary needs to include standard names, a set of synonyms, a set of abbreviations and aliases, field semantic descriptions, scope of application, and conflict resolution rules. The retrieval process can use a combination of exact matching, normalized matching, and similarity matching. Normalized matching can include case normalization, delimiter normalization, camelCase splitting, and stemming. Similarity matching can generate candidate mappings based on edit distance, vector similarity, or pinyin similarity. When there are polysemous conflicts in candidate mappings, the semantic governance center performs disambiguation based on business semantic tags, usage scenarios, and field value distribution. After disambiguation is successful, the mapping relationship is solidified into a field mapping record, and non-standard field names are replaced with standard names or a table of standard names to original fields is established in the standardized data package to ensure that the query scope is unified and the source field can be traced back. Standard names serve as the basis for the field definitions of standardized data after subsequent governance, avoiding statistical discrepancies caused by fields with the same meaning appearing with different names in different data sources.

[0055] The quality score quantifies the overall status of standardized data packets across dimensions such as completeness, consistency, accuracy, and timeliness. The Semantic Governance Center utilizes a dynamic data quality scoring model to perform analysis and trigger alerts. The input to the dynamic data quality scoring model includes measurable indicators such as the field missing rate, enumeration validity, cross-field constraint consistency, timestamp reasonableness, spatial coordinate range reasonableness, synonym mapping coverage, and business semantic tag confidence of the standardized data packets. Different indicator weights are set based on the usage scenario to form the quality score. The dynamic nature is reflected in the fact that weights or thresholds can be adjusted according to business cycles, data source fluctuations, and historical alert distribution, making the score sensitive to data source drift and business changes. A preset threshold is used to define the acceptable lower limit. The Semantic Governance Center compares the quality score with the preset threshold. When the quality score falls below the preset threshold, an alert is triggered. The alert content can include elements such as a unique identifier, source path, processing process, triggering indicator, the difference between the quality score and the threshold, and suggested handling strategies to locate problematic data and trace the responsibility chain. Ultimately, the semantic governance center generates standardized data after governance based on the data lineage graph, business semantic tags, standard names, and quality scores. The generation process can be represented by merging and encapsulating the standardized data package and governance results and outputting them as a unified structure. The unified structure simultaneously includes a standardized set of fields, field mapping records, a set of semantic tags, and pointers to quality scores and lineage associations, so that the standardized data after governance has a traceable, understandable, and controllable quality.

[0056] This embodiment uses a semantic governance center to generate unique identifiers, record source paths, record processing procedures, and record usage scenarios for standardized data packets, forming a data lineage graph. It then combines a natural language processing model and an expert strategy library to complete business semantic labeling. A standard terminology dictionary is used to normalize the mapping of non-standard field names to standard names. A dynamic data quality scoring model is employed to analyze quality scores, triggering alarms when the quality score falls below a preset threshold. This ensures that the output standardized data after governance possesses stable referencing capabilities, end-to-end traceability, semantic interpretability, unified field definitions, and measurable quality. This reduces the uncertainty in usage caused by inconsistent definitions and uncontrollable quality during heterogeneous data fusion.

[0057] In one embodiment, step S40 above includes: S401, Analyze the standardized data after governance, extract the timestamp of the standardized data after governance to determine the time attribute, and analyze the historical access records of the standardized data after governance to determine the access popularity characteristics; S402, if the time attribute is within the first preset time period and the access popularity feature indicates high-frequency access, then the standardized data after treatment is written into the hot data storage layer. S403, if the time attribute is within the first preset time period and the access popularity feature indicates non-high frequency access, or the time attribute exceeds the first preset time period but does not exceed the second preset time period, then the standardized data after treatment is written into the warm data storage layer. S404, if the time attribute exceeds the second preset time period, then the standardized data after treatment is written into the cold data storage layer. S405, establish a unified data index among the hot data storage layer, the warm data storage layer and the cold data storage layer.

[0058] In this embodiment, before the standardized data enters the storage stage, it needs to form time attributes and access popularity characteristics that can be used for hierarchical judgment. The standardized data is then written into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer, and a cold data storage layer, so that the storage location matches the access pattern of subsequent data retrieval requests. Analyzing the standardized data includes reading the structural load and governance result fields of the standardized data. The source of the structural load is the set of business fields encapsulated in the standardized data. The source of the governance result fields is the business semantic tags, standard names, quality scores, and other information encapsulated in the standardized data. On the storage side, the analysis action is manifested as generating a judgment input set for hierarchical routing. The judgment input set includes at least a timestamp, historical access records, a first preset time period, a second preset time period, and a judgment threshold or judgment rule for access popularity characteristics. The timestamps of the standardized data after governance are extracted to determine the time attributes. The timestamps can originate from the business time field or collection time field of the standardized data, or from the storage time field written during the generation of the standardized data. After extraction, the timestamps are format-normalized to make them comparable to the first and second preset time periods. Format normalization can include time zone unification, precision unification, and missing data completion. Missing data completion can be estimated from the access time recorded in the source path or the processing time during the processing. The time attributes are used to express the relative relationship between the standardized data after governance and the current time or the time period to which it belongs. The time attributes can be represented as the time difference from the current time, the time window identifier, or the time period identifier mapped according to the business cycle. The generation of time attributes requires comparing the timestamps with the time boundaries corresponding to the first and second preset time periods. The comparison results form the time condition input for hierarchical judgment, making subsequent branch judgments deterministic.

[0059] Historical access records of the standardized data after governance are analyzed to determine access frequency characteristics. These records originate from interface call statistics, query condition hit information, and data read logs from the unified data service gateway or storage access layer. Historical access records need to be associated with the standardized data after governance; the association key can be a unique identifier for the standardized data, a dataset identifier in the standard name set, or an index key in the unified data index. After association, access frequency, access concurrency, access interval distribution, and data read volume are aggregated within a preset statistical window. The aggregation result forms the input for calculating the access frequency characteristics. Access frequency characteristics indicate the frequency of access to the standardized data after governance over a subsequent period. These characteristics can be represented as a discrete state of high-frequency or low-frequency access, or as a popularity score mapped to a discrete state through a threshold. The popularity score can be calculated using sliding window counting, exponential decay weighted counting, or business cycle weighted counting, ensuring that recent accesses contribute more to the results and reflect periodic fluctuations. The criteria for determining high-frequency access can be that the number of accesses exceeds a preset threshold, the access concurrency exceeds a preset threshold, or the access interval is lower than a preset threshold. Non-high-frequency access corresponds to the state that does not meet the criteria for determining high-frequency access, so that the access popularity feature can be directly referenced by subsequent hierarchical writing rules.

[0060] When the time attribute falls within a first preset time period and the access frequency characteristic indicates high-frequency access, the standardized data after governance is written to the hot data storage layer. The first preset time period is used to define the range of recent data and, together with the high-frequency access condition, forms the hot layer writing condition. The writing action includes converting the standardized data after governance into a writing carrier for the hot data storage layer and submitting the write transaction. The writing carrier can be a row record, a key-value entry, or a query-oriented wide table record, and uses a unique identifier or a unified data index key as the primary key to support fast location. The hot data storage layer is used to hold data sets with high requirements for low-latency reading. Therefore, during writing, index entries or cache pages oriented towards query conditions can be generated simultaneously to shorten the subsequent retrieval path of the unified data service gateway. At the same time, the writing result is linked to the index entries of the unified data index, enabling the unified data index to indicate the physical location or logical partition in the hot data storage layer.

[0061] When the time attribute is within the first preset time period and the access popularity characteristic indicates low-frequency access, or when the time attribute exceeds the first preset time period but does not exceed the second preset time period, the standardized data after governance is written to the warm data storage layer. This branch aggregates near-term but low-frequency data and data within the medium-term time window into the warm data storage layer, avoiding low-access data occupying the resources of the hot data storage layer. The second preset time period is used to define the range of data that can be served online but whose access frequency is relatively low. When writing to the warm data storage layer, the organization method of the standardized data after governance can be adjusted for batch scanning and aggregation. The organization method adjustment can include column layout, partition key generation, and file organization grouped by standard name or business semantic tag, making the warm data storage layer adaptable to multi-condition filtering and range query. Low-frequency access data within the first preset time period is written using the warm layer, so that the writing rules cover all near-term data combinations, avoiding the problem of data with time attributes within the first preset time period but low access popularity characteristics having no writing destination.

[0062] When the time attribute exceeds the second preset time period, the standardized data after governance is written to the cold data storage layer. This branch is used to handle long-term data. During the write operation, the standardized data after governance can be encapsulated for long-term storage. The encapsulation can include segmented archiving, generation of archive buckets by timestamp, and generation of verification information. The verification information can be a hash digest or checksum to support subsequent integrity verification. The cold data storage layer focuses on capacity and cost, so the write operation can adopt append write and reduce random updates. At the same time, the mapping relationship between the unified data index and the cold data storage layer object identifier or archive bucket identifier is maintained, so that the unified data service gateway can locate the corresponding standardized data after governance when backtracking is required.

[0063] A unified data index is established between the hot, warm, and cold data storage layers to provide a unified entry point for cross-layer location. The index key of the unified data index can be a unique identifier of the standardized data after governance, a standard name combination key, or a composite key containing business semantic tags. The index value must include at least a storage layer identifier, a storage location pointer, a time attribute summary, and an access frequency feature summary, enabling the unified data service gateway to quickly determine which storage layer to read from and locate the specific position based on query conditions. The unified data index can be created by inserting or updating the index after each write operation, or by refreshing the index in batches according to a time window. Index updates must be consistent with the write results to avoid the index pointing to invalid locations. If necessary, a version number or effective time can be recorded in the index to support index rollback and consistency verification. The coupling relationship between the unified data index and the hierarchical storage architecture is that the unified data index provides the logical entry point, while the hot, warm, and cold data storage layers provide the physical carrier, thus forming a storage organization structure that allows cross-layer retrieval and controllable hierarchical layering.

[0064] This embodiment determines the time attribute by extracting the timestamp of the standardized data after governance and determines the access popularity characteristics based on historical access records. It constructs branch writing rules using a first preset time period and a second preset time period, writes the recently governed standardized data with high-frequency access to the hot data storage layer, writes the governed standardized data with low-frequency access or in the middle window within the first preset time period to the warm data storage layer, and writes the governed standardized data with high-frequency access beyond the second preset time period to the cold data storage layer. A unified data index is established between the hot data storage layer, the warm data storage layer, and the cold data storage layer, so that the storage location of the governed standardized data corresponds to the access pattern and has cross-layer unified positioning capability, thereby reducing the resource consumption and retrieval path uncertainty caused by the co-storage of high-frequency access data and low-frequency access data.

[0065] In one embodiment, step S50 above includes: S501, Configure a variety of standard service interfaces in the unified data service gateway. The standard service interfaces include a structured query language interface for custom queries, a representation state transition interface for calling specific datasets, a streaming subscription interface for real-time data push, and an embedded component interface for visualization integration. S502, receive data call requests through the standard service interface, and perform identity and permission verification and traffic rate limiting checks on the requester initiating the data call request; S503, when the identity and permission verification is passed and the traffic rate limit check is passed, the query conditions of the data call request are parsed, and the corresponding standardized data after governance is retrieved from the hierarchical storage architecture according to the query conditions; S504, the retrieved standardized data after governance is output through the unified data service gateway, and the interface call statistics for the data call request are recorded.

[0066] In this embodiment, a unified data service gateway responds to data call requests to output standardized, governed data from a tiered storage architecture. This involves interface capability orchestration, request access verification, query condition parsing, tiered storage architecture retrieval, and a closed-loop output and statistics mechanism. The unified data service gateway serves as the entry point for external or internal data services. The "unified" aspect refers to the consistent access path used by different requesters to reach the tiered storage architecture. The data service gateway and the tiered storage architecture communicate via a unified data index or storage layer access adapter for location and retrieval. The response process consists of receiving, verifying, parsing, retrieving, and outputting. The output is limited to standardized, governed data stored within the tiered storage architecture, thus clearly defining the boundaries between the gateway's processing scope and the data source scope.

[0067] The unified data service gateway is configured with various standard service interfaces. Configuration actions include interface type registration, interface routing rule binding, interface parameter contract definition, and interface return structure definition, enabling different interfaces to share a unified identity and permission verification and traffic rate limiting check process. Standard service interfaces include a Structured Query Language (SCL) interface, a representation state transition interface, a streaming subscription interface, and an embedded component interface. The SCL interface is used for custom queries, originating from structured query statements for analytical scenarios. Interface input includes the query text and execution parameters, which may include timeout thresholds, maximum number of returned rows, and pagination identifiers. On the gateway side, interface processing involves syntax validation of the query statement, extraction of query conditions, and generation of an executable request. The representation state transition interface is used for specific dataset calls, originating from resource-location-based interfaces for business calls. Interface input includes dataset identifiers, filtering conditions, and time ranges. On the gateway side, the dataset identifier is mapped to a unified data index key or standard name set, and the filtering conditions are standardized into searchable conditions, thereby reducing the complexity of custom queries for callers and stabilizing request formats. The streaming subscription interface is used for real-time data push. It originates from subscription-based interactions that require continuous data updates. The interface input includes the subscription topic, update frequency, or incremental conditions. The gateway establishes a subscription session for each subscription and maintains a cursor or offset, linking subsequent pushes to the incremental data location within the tiered storage architecture. The embedded component interface is used for visualization integration. It originates from interactive methods that embed data services into visualization pages or components. The interface input includes the component identifier, the set of fields required for rendering, and the aggregation method. The gateway maps the field set to the projection set in the query conditions, forming a data structure adapted to the visualization component, avoiding repeated field assembly and structure transformation on the front end for the caller.

[0068] The system receives data call requests through a standard service interface and performs identity and permission verification and traffic rate limiting checks on the requester. The receiving actions include request message parsing, requester identification, and request context establishment. The request context includes at least the requester identifier, interface type, request time, and request parameters. Identity and permission verification determines whether the requester has the authorized scope to access the standardized data after governance. The identity source can be an access token, digital certificate, or signature field. The verification process may include token validity verification, signature verification, and requester identifier mapping. Permission determination may include interface-level permissions, dataset-level permissions, and field-level permissions, creating a constraint relationship between the scope of standardized data accessible to the requester and the query conditions. Traffic rate limiting checks control concurrency and throughput. Rate limiting can be based on the requester's concurrency limit, the interface's rate limit, or the overall system resource level. The check can be implemented using a counter or token bucket, mapping each data call request to an event that consumes a certain amount of data and comparing it with a preset threshold. If the threshold is exceeded, the request is rejected or downgraded, thus preventing the tiered storage architecture from being overwhelmed by sudden traffic surges. The order of identity and permission verification and traffic rate limiting checks is fixed. Verification before rate limiting or rate limiting before verification can both be implemented. However, in this step chain, both are placed before parsing query conditions and retrieval, so that subsequent retrieval is always performed under authorization and rate limiting constraints.

[0069] When identity and permission verification and traffic rate limiting checks are passed, the query conditions of the data call request are parsed, and the corresponding standardized data after governance is retrieved from the hierarchical storage architecture based on the query conditions. The query conditions originate from the input differences of different standard service interfaces. In the Structured Query Language interface, the query conditions come from the predicate, projection, and aggregation parts of the query statement; in the representation state transition interface, the query conditions come from parameterized filtering conditions and dataset identifiers; in the streaming subscription interface, the query conditions come from the subscription topic and incremental conditions; and in the embedded component interface, the query conditions come from the field set and aggregation method. The parsing process includes condition extraction, condition normalization, and condition constraint injection. Condition extraction converts the interface input into a unified filtering expression; condition normalization converts time ranges, spatial ranges, field names, etc., into expressions consistent with standard names and aligned with the unified data index key; and condition constraint injection converts the permission determination results into additional filtering conditions or field masking rules to ensure that the retrieval scope does not exceed the authorized scope. Based on the query criteria, the gateway retrieves the corresponding standardized data from the hierarchical storage architecture. The retrieval process includes locating the storage layer and performing a read operation. Location can be based on a unified data index, mapping the query criteria to location pointers in hot, warm, or cold data storage layers. Location pointers can be partition keys, object identifiers, or file paths. The read operation can employ different access methods in different storage layers, but presents a unified set of read results to the gateway. To reduce cross-layer retrieval overhead, the gateway can first extract the index key based on the query criteria and complete the candidate set filtering in the unified data index before performing a hierarchical read on the candidate set. This reduces the probability of a full scan in the cold data storage layer and shortens the response path.

[0070] The retrieved standardized data after governance is output through a unified data service gateway, and the interface call statistics for data call requests are recorded. Output actions include result encapsulation, result return, and exception feedback. Result encapsulation generates a response payload according to the return contract of the standard service interface. The Structured Query Language interface can return tabular results and field metadata. The representation state transition interface can return structured objects oriented towards the dataset. The streaming subscription interface can segment the standardized data after governance into incremental messages with cursor update information. The embedded component interface can return aggregated results and dimensional field sets directly consumed by visualization components. Interface call statistics describe the running status and resource consumption of a data call request. Statistical fields are derived from the request context and execution results on the gateway side, and at least include call time, processing time, returned data volume, storage layer hit identifier, and result status code. Interface call statistics can be aggregated with historical access records for subsequent access popularity feature analysis, thereby enabling data service behavior to support the write strategy and index optimization of the hierarchical storage architecture. The recording of interface call statistics needs to be bound to the output action within the same request lifecycle; both successful and failed outputs are recorded to form a traceable call loop and support exception localization.

[0071] This embodiment configures a structured query language interface, a representation state transition interface, a streaming subscription interface, and an embedded component interface in a unified data service gateway. After receiving a data call request, it performs identity and permission verification and traffic rate limiting checks. After successful verification, it parses the query conditions of the data call request and retrieves standardized data after governance from the hierarchical storage architecture based on the query conditions. Then, it outputs the standardized data after governance and records interface call statistics. This allows different types of data call requests to obtain a consistent verification and retrieval link through a unified entry point. At the same time, it precipitates the call behavior into usable statistical information, thereby improving the controllability and traceability of data output and reducing the inconsistency of retrieval paths and the complexity of operation and maintenance caused by the dispersion of interface forms.

[0072] In one embodiment, after step S50 above, the method further includes: S601 records and analyzes the access logs of historical data call requests to obtain user access habit characteristics and business cycle pattern characteristics. S602, Based on the user access habit characteristics and the business cycle pattern characteristics, the prediction module is used to predict the target dataset that will be frequently accessed in the next time period. S603, Locate the target dataset from the hierarchical storage architecture and extract the target dataset into the memory cache in advance; S604, establish an access link between the memory cache and the unified data service gateway, so as to prioritize responding with data from the memory cache when a data call request for the target dataset is received.

[0073] In this embodiment, after responding to data call requests through the unified data service gateway and outputting the standardized data after governance in the hierarchical storage architecture, a preloading link based on access behavior is introduced to shorten the retrieval path of subsequent similar requests. The overall action consists of access log accumulation, feature extraction, prediction module inference, target dataset location, memory cache prefetching, and access link activation. The constraints are as follows: user access habit characteristics and business cycle regularity characteristics come from the access logs of historical data call requests; the target dataset comes from the prediction results of the prediction module for high-frequency access objects in the next time period; the memory cache comes from the pre-extraction results of the target dataset; and the access link is used to prioritize the memory cache for subsequent data call requests for the target dataset.

[0074] Access logs, which record and analyze historical data call requests, form the basis for computable behavioral data. These historical data call requests originate from request instances already received and processed by the unified data service gateway. The access log entries must include at least the requester's identifier, request time, request interface type, query condition summary, hit dataset identifier, returned data volume, and processing time, ensuring the logs reflect both access frequency and time period, as well as access cost. Recording actions can be completed within the request lifecycle of the unified data service gateway. Access log entries are formed by expanding the fields of interface call statistics and organized using a time window aggregation method. Time windows can be segmented by minute, hour, or day granularity for subsequent statistics. Analysis actions use access log entries as input to construct user access habit characteristics and business cycle pattern characteristics. User access habit characteristics characterize the requester's preferences in the time and content dimensions. The time dimension reflects peak access periods, differences between weekdays and non-weekdays, and the distribution of consecutive access intervals. The content dimension reflects the probability of repeated access to the target dataset, interface type preferences, and fixed patterns in query conditions. Business cycle pattern features are used to characterize periodic fluctuations related to business rhythms. The source can be the frequency sequence and time sequence aggregated by calendar cycle in the access log. The cycle characterization can cover daily cycle, weekly cycle or monthly cycle, so that the subsequent prediction module can distinguish between occasional sudden increases and periodic high frequency.

[0075] Based on user access habits and business cycle patterns, the prediction module forecasts the target dataset to be frequently accessed in the next time period. This is an inference process that maps behavioral features to a preloaded object set. The prediction module's input consists of user access habits and business cycle patterns, and its output is the target dataset and its corresponding high-frequency access confidence score or ranking result. The next time period is used to limit the effective prediction interval and align with the preload window. The time period length can be consistent with the granularity of access log aggregation for closed-loop verification. The prediction module can employ a multi-factor scoring method, combining factors such as access frequency trends, time period matching, cycle phase matching, and the interval between the most recent accesses to form a high-frequency access score. The target dataset is then determined based on a scoring threshold or a Top set, thus ensuring the target dataset has an interpretable source and a controllable upper limit on size. To avoid including excessively large datasets in the preload, the prediction module can introduce data volume constraints or cache capacity constraints, jointly filtering high-frequency access scores and resource constraints to ensure the target dataset set matches the capacity of the memory cache.

[0076] Locating the target dataset within a hierarchical storage architecture and pre-fetching it into a memory cache is an execution process that transforms prediction results into a directly serviceable data-resident format. The location action takes the target dataset as input and utilizes the unified data index or dataset identifier mapping rules of the hierarchical storage architecture to determine the target dataset's location within the hot, warm, or cold data storage layer. Location representation can include partition key ranges, object identifier sets, or file path sets, providing the extraction action with a set of executable read instructions. The pre-fetching action initiates a read operation on the corresponding storage layer based on the location result, writes the read result to the memory cache, and establishes a mapping relationship between cache entries and the target dataset. Cache entries can contain data block sequences, columnar fragments, or key-value segments, and the mapping relationship can include the target dataset identifier, version identifier, or time range identifier to support subsequent cache hits based on the target dataset. The memory cache provides shorter read paths and lower access latency. Its internal organization can employ a key-value structure, segmented structure, or page structure. Key generation can be based on a combination of the target dataset identifier and query condition summary, ensuring that different commonly used query formats for the same target dataset can hit the corresponding cache entries. To reduce the data consistency risks introduced by preloading, a generation timestamp and expiration date can be written to the memory cache entries, and an update can be triggered when the expiration date expires or the hit rate decreases. The update source is still the target dataset location and extraction process in the hierarchical storage architecture.

[0077] Establishing an access link between the memory cache and the unified data service gateway is used to implement a priority hit strategy and complete the path switching from cache to output on the gateway side. The establishment of the access link includes inserting a cache determination node into the request processing chain of the unified data service gateway, configuring the mapping rules from the target dataset to the cache key, and configuring the return encapsulation rules after a cache hit. The cache determination node is used to identify whether a request belongs to the target dataset when a data call request for the target dataset is received. This identification can be based on the matching relationship between the dataset identifier parameters, query conditions, and the target dataset list, or the matching relationship between the unified data index key. If a match is successful, the response data is preferentially read from the memory cache and generated. If a match fails or the cache is not hit, the process falls back to the tiered storage architecture retrieval path, forming a controllable priority and fallback logic. To ensure the consistency of the unified data service gateway's output, the access link needs to reuse the gateway's existing output encapsulation logic. This ensures that cache hit output and tiered storage architecture retrieval output are consistent in terms of field sets, sorting rules, and return structures. At the same time, cache hit and rollback events are recorded in the access log. This allows the subsequent construction of user access habit characteristics and business cycle pattern characteristics to reflect the impact of caching strategies on access behavior and provides a closed-loop correction basis for the prediction module.

[0078] Example Description: Taking the cross-network integration construction of an urban traffic management platform as an example, the external network area covers data generation and circulation environments such as the Internet, government extranets, and video private networks, while the internal security area corresponds to the data processing and command application environment within the public security private network. The platform deploys a cross-network data transfer channel between the external network area and the internal security area, and enables a one-way isolation transmission mechanism, ensuring that multi-source heterogeneous raw data can only enter the internal security area along a predetermined direction. In implementation, an external isolation buffer pool is set up in the external network area. Before accessing multi-source heterogeneous raw data such as real-time traffic conditions from navigation platforms, shared bicycle parking heat maps, government construction plans, weather warnings, and structured event results from video private networks, the credibility of the data source is verified. After successful verification, the multi-source heterogeneous raw data is temporarily stored in the external isolation buffer pool. Subsequently, the physical unidirectional optical gate component in the unidirectional isolation transmission mechanism is activated, establishing a unidirectional optical signal transmission link from the external isolation buffer pool to the internal secure area. This link reads the multi-source heterogeneous raw data from the external isolation buffer pool and delivers it to the internal isolation buffer pool in the internal secure area using physical signals that are not part of the network protocol. Within the internal isolation buffer pool, a policy-driven automatic approval process is triggered. This process reads the data purpose and security classification attributes of the multi-source heterogeneous raw data and determines whether to release it. For example, internet navigation traffic conditions are labeled as congestion analysis and signal timing linkage based on purpose, while government construction plans are labeled as internally collaboratively visible based on security classification. After the automatic approval process determines that the data is released, the national cryptographic algorithm encryption module is invoked to perform sensitive entity field identification on the multi-source heterogeneous raw data. Sensitive entity fields such as license plate numbers, facial features, and precise trajectory points are dynamically desensitized and encrypted. Simultaneously, a data flow audit log is generated to record the data flow time, data source identifier, approval conclusion, and processing actions, facilitating subsequent audit traceability.

[0079] After multi-source heterogeneous raw data enters the internal secure area, the protocol adaptation module establishes a unified access point on the access side. Upon startup, the protocol adaptation module automatically identifies the communication protocol type of the data source and establishes a data transmission connection based on the protocol type to receive the multi-source heterogeneous raw data. For example, HTTPS or SFTP connections are used for exchanging documents on the government extranet, message queues or streaming channels are used for video private network event pushes, and REST pull or subscription methods are used for open internet data. After data reception, the multimodal parsing engine identifies the data format of the multi-source heterogeneous raw data and performs differentiated parsing based on the data format to form a manageable structured representation. When the data format is structured, the multimodal parsing engine performs standardized data model mapping, such as aligning fields in checkpoint vehicle records to a unified set of fields for vehicle, time, and location. When the data format is semi-structured, the multimodal parsing engine uses pattern inference algorithms to infer the meaning of fields, such as inferring from JSON logs that both `plate_no` and `vehicle_id` point to the same type of license plate field semantics and extracting them as unified fields. When the data format is unstructured, the multimodal parsing engine calls an artificial intelligence recognition model to perform structured feature extraction, such as extracting license plate text, vehicle type, and lane occupancy status from images, and extracting structured features such as lane occupancy, wrong-way driving, and congestion queue length from video events. After generating parsed data through differential parsing, the multimodal parsing engine performs spatiotemporal benchmark alignment processing on the parsed data to generate spatiotemporally aligned data. The alignment process unifies the differences in timestamp precision from different sources to the same timing benchmark and unifies different coordinate systems to the same spatial benchmark, enabling a unified expression of the same event that is comparable and correlated in time and space. Finally, the spatiotemporally aligned data is encapsulated in a unified format to generate standardized data packets, so that each piece of data entering the governance chain carries a set of fields, aligned time and location expressions, main elements and resolution confidence information in a unified carrier format.

[0080] After standardized data packets enter the semantic governance center, the center performs metadata registration, semantic tagging, and synonym mapping operations on the standardized data packets and outputs the standardized data after governance. The metadata registration stage generates a unique identifier for each standardized data packet and records its source path, processing process, and usage scenario to form a data lineage graph. For example, the source path records the navigation platform interface name or government system directory identifier; the processing process records the parsing type and alignment rule version; and the usage scenario records congestion assessment, construction coordination, or accident handling. The semantic tagging stage calls upon natural language processing models and expert policy libraries to tag standardized data packets with business semantic tags. For example, morning rush hour traffic events on main roads are tagged as related to commuter vehicles, accident-prone road segment events are tagged as accident black spots, and key vehicle passage events are tagged as related to key regulatory targets. Subsequently, the semantic governance center retrieves a pre-set standard terminology dictionary and maps non-standard field names in the standardized data packets to standard names. For example, it maps differing fields such as plate_no, license_num, and vehicle_id to the same standard name to eliminate naming fragmentation. In the data quality analysis phase, the Semantic Governance Center uses a dynamic data quality scoring model to analyze the quality score of standardized data packets. The quality score is obtained by comprehensively considering dimensions such as completeness, consistency, accuracy, and timeliness. When the quality score falls below a preset threshold, an alarm is triggered to indicate data source anomalies, parsing anomalies, or alignment anomalies. Finally, the Semantic Governance Center generates standardized data after governance based on the data lineage graph, business semantic tags, standard names, and quality scores, ensuring that the data achieves a unified state of manageability and usability in terms of semantics, fields, and quality.

[0081] Before the standardized data is incorporated into the tiered storage architecture, the platform analyzes its time attributes and access frequency characteristics. Based on the analysis results, it allocates writes to the hot data storage layer, warm data storage layer, and cold data storage layer. The time attributes are extracted from the timestamps of the standardized data, and the access frequency characteristics are obtained by analyzing the historical access records of the standardized data. These historical access records can be derived from the access trajectories formed by the statistical information accumulated from the interface calls of the unified data service gateway. During the write allocation process, when the time attribute is within the first preset time period and the access popularity characteristic indicates high-frequency access, the standardized data after governance is written to the hot data storage layer. For example, real-time traffic within the past seven days and high-frequency query results of violations at key intersections are entered into the hot data storage layer to support millisecond-level queries. When the time attribute is within the first preset time period and the access popularity characteristic indicates low-frequency access, or when the time attribute exceeds the first preset time period but does not exceed the second preset time period, the standardized data after governance is written to the warm data storage layer. For example, medium-frequency trend analysis data within the past thirty days is entered into the warm data storage layer to support reports and retrospective analysis. When the time attribute exceeds the second preset time period, the standardized data after governance is written to the cold data storage layer. For example, long-term archived video-related events, historical logs, and sets of atomic events with low-frequency access are entered into the cold data storage layer to control costs. After completing the layered write, a unified data index is established between the hot data storage layer, the warm data storage layer, and the cold data storage layer. The unified data index records the mapping relationship between the standardized data after governance and the storage layer location, partition range, and dataset identifier, enabling subsequent retrieval to locate across layers and support consistent access across multiple interfaces.

[0082] When business systems, analysts, or visualization dashboards need to access data, the unified data service gateway is responsible for responding to data request requests and outputting standardized, governed data stored in a tiered storage architecture. The unified data service gateway is configured with various standard service interfaces to adapt to different consumption patterns. These standard service interfaces include a structured query language interface, a representation state transition interface, a streaming subscription interface, and an embedded component interface. For example, the structured query language interface allows analysts to perform combined queries by intersection, time period, and violation type; the representation state transition interface allows business systems to access data based on fixed datasets, such as the top 10 violations at a certain intersection yesterday; the streaming subscription interface is used to subscribe to event push notifications triggered by congestion thresholds; and the embedded component interface is used to embed statistical layers and event lists into command dashboards or mobile applications. After receiving data request requests through the standard service interfaces, the unified data service gateway performs identity and permission verification and traffic rate limiting checks on the requester. Identity and permission verification binds the requester's identity to the range of accessible data, while traffic rate limiting controls concurrency and frequency within thresholds to avoid resource overload. Once identity and permission verification is successful and the traffic rate limiting check passes, the unified data service gateway parses the query conditions of the data call request and retrieves the corresponding standardized data from the hierarchical storage architecture based on the query conditions and the unified data index. The retrieval process can directly locate the hot data storage layer or cross-layer to locate the warm and cold data storage layers and summarize the results. After the retrieval is complete, the standardized data is output through the unified data service gateway, and interface call statistics for the data call request are recorded. These statistics include interface type, query condition summary, hit storage layer, returned data volume, and latency, for subsequent operational analysis and popularity calculation.

[0083] After the service chain of the unified data service gateway is running stably, the platform further introduces a pre-loading closed loop of memory caching after data output to improve the response efficiency of high-frequency access scenarios. The system records and analyzes access logs of historical data call requests to obtain user access habit characteristics and business cycle pattern characteristics. The access logs are continuously accumulated by the unified data service gateway when processing data call requests. User access habit characteristics depict the requester's preferences for road segments, intersections, event types, and interface forms, while business cycle pattern characteristics depict periodic access fluctuations such as morning and evening peak hours, holidays, and construction windows. Based on user access habit characteristics and business cycle pattern characteristics, the prediction module predicts the target datasets that will be frequently accessed in the next time period. For example, before the morning peak, it predicts that target datasets related to main road traffic flow and queue length at key intersections will be frequently accessed. Subsequently, the target dataset is located from the hierarchical storage architecture and pre-extracted into the memory cache. The location relies on the unified data index to map the target dataset to the specific location of the hot data storage layer, warm data storage layer, or cold data storage layer. Pre-extraction keeps the target dataset in the memory cache in a form suitable for fast reading. Finally, an access link between the memory cache and the unified data service gateway is established. When the unified data service gateway receives a data call request for the target dataset, it first responds with data from the memory cache. When the memory cache is not hit or the target dataset is not in the preloaded set, it falls back to the retrieval path of the hierarchical storage architecture. Thus, without changing the external interface form, high-frequency access requests are transformed from cross-layer retrieval to memory-level reading and continuously rely on access logs and interface call statistics to form a self-correcting preload closed loop.

[0084] This embodiment records and analyzes access logs of historical data call requests to form user access habit characteristics and business cycle patterns. It then uses a prediction module to predict the target dataset that will be frequently accessed in the next time period. The target dataset is then located from the hierarchical storage architecture and extracted to the memory cache in advance. By establishing an access link between the memory cache and the unified data service gateway, when a data call request for the target dataset is received, the data is responded to from the memory cache first. This switches the retrieval path of the hierarchical storage architecture to the short path response of the memory cache, reducing the time consumption caused by repeated retrieval and cross-layer reading. At the same time, the hit and rollback information is recorded as access logs to support the subsequent prediction module correction, thereby improving the response efficiency and stability of data call requests for the target dataset.

[0085] In one embodiment, a cross-network heterogeneous data fusion processing apparatus is provided, which corresponds one-to-one with the cross-network heterogeneous data fusion processing method described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the cross-network heterogeneous data fusion processing device of the present invention. The modules include a cross-network data transfer module 10, a multimodal parsing module 20, a semantic governance module 30, a hierarchical storage scheduling module 40, and a unified data service module 50. Detailed descriptions of each functional module are as follows: Cross-network data transfer module 10 is used to construct a cross-network data transfer channel between the external network area and the internal security area, and to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area using a one-way isolation transmission mechanism. The multimodal parsing module 20 is used to receive the multi-source heterogeneous raw data through the protocol adaptation module in the internal security area, and call the multimodal parsing engine to perform structured feature extraction and spatiotemporal benchmark alignment processing on the multi-source heterogeneous raw data to generate standardized data packets. The semantic governance module 30 is used to perform metadata registration, semantic tag annotation and synonym mapping operations on the standardized data packets through the semantic governance center to generate standardized data after governance. The hierarchical storage scheduling module 40 is used to analyze the time attributes and access popularity characteristics of the standardized data after governance, and write the standardized data after governance into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. The unified data service module 50 is used to respond to data call requests through the unified data service gateway to output the standardized data after governance stored in the hierarchical storage architecture.

[0086] Specific limitations regarding the cross-network heterogeneous data fusion processing device can be found in the aforementioned limitations regarding the cross-network heterogeneous data fusion processing method, and will not be repeated here. Each module in the aforementioned cross-network heterogeneous data fusion processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0087] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a cross-network heterogeneous data fusion processing method on the server side.

[0088] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a cross-network heterogeneous data fusion processing method.

[0089] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Construct a cross-network data transfer channel between the external network area and the internal security area, and use a one-way isolation transmission mechanism to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area; The multi-source heterogeneous raw data is received through the protocol adaptation module in the internal security area, and the multi-modal parsing engine is called to perform structured feature extraction and spatiotemporal benchmark alignment on the multi-source heterogeneous raw data to generate standardized data packets. The semantic governance center performs metadata registration, semantic tagging, and synonym mapping operations on the standardized data packets to generate standardized data after governance. The time attributes and access popularity characteristics of the standardized data after governance are analyzed, and the standardized data after governance is written into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. The unified data service gateway responds to data call requests and outputs the standardized, governed data stored in the hierarchical storage architecture.

[0090] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Construct a cross-network data transfer channel between the external network area and the internal security area, and use a one-way isolation transmission mechanism to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area; The multi-source heterogeneous raw data is received through the protocol adaptation module in the internal security area, and the multi-modal parsing engine is called to perform structured feature extraction and spatiotemporal benchmark alignment on the multi-source heterogeneous raw data to generate standardized data packets. The semantic governance center performs metadata registration, semantic tagging, and synonym mapping operations on the standardized data packets to generate standardized data after governance. The time attributes and access popularity characteristics of the standardized data after governance are analyzed, and the standardized data after governance is written into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. The unified data service gateway responds to data call requests and outputs the standardized, governed data stored in the hierarchical storage architecture.

[0091] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0092] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0093] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

[0094] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.

Claims

1. A method for cross-network heterogeneous data fusion processing, characterized in that, Includes the following steps: Construct a cross-network data transfer channel between the external network area and the internal security area, and use a one-way isolation transmission mechanism to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area; The multi-source heterogeneous raw data is received through the protocol adaptation module in the internal security area, and the multi-modal parsing engine is called to perform structured feature extraction and spatiotemporal benchmark alignment on the multi-source heterogeneous raw data to generate standardized data packets. The semantic governance center performs metadata registration, semantic tagging, and synonym mapping operations on the standardized data packets to generate standardized data after governance. The time attributes and access popularity characteristics of the standardized data after governance are analyzed, and the standardized data after governance is written into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. The unified data service gateway responds to data call requests and outputs the standardized, governed data stored in the hierarchical storage architecture.

2. The cross-network heterogeneous data fusion processing method as described in claim 1, characterized in that, Constructing a cross-network data transfer channel deployed between an external network area and an internal security area, and using a one-way isolation transmission mechanism to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area, including: Deploy an external isolation buffer pool in the external network area to verify the identity and credibility of the data source that generates multi-source heterogeneous raw data, and write the verified multi-source heterogeneous raw data into the external isolation buffer pool for temporary storage. The physical unidirectional optical gate component in the unidirectional isolation transmission mechanism is activated to establish a unidirectional optical signal transmission link from the external isolation buffer pool to the internal security area. The multi-source heterogeneous raw data in the external isolation buffer pool is read through the unidirectional optical signal transmission link, and the multi-source heterogeneous raw data is delivered to the internal isolation buffer pool deployed in the internal security area in the form of a non-network protocol physical signal. In the internal isolation buffer pool, a policy-driven automatic approval process is triggered to determine whether to allow the multi-source heterogeneous raw data based on the data usage attributes and security level attributes of the raw data. If the automatic approval process determines that the data should be approved, the national cryptographic algorithm encryption module is invoked to identify sensitive entity fields in the multi-source heterogeneous original data and perform dynamic desensitization and encryption processing, and generate a data flow audit log.

3. The cross-network heterogeneous data fusion processing method as described in claim 1, characterized in that, Within the internal security zone, the multi-source heterogeneous raw data is received via a protocol adaptation module, and a multimodal parsing engine is invoked to perform structured feature extraction and spatiotemporal benchmark alignment on the multi-source heterogeneous raw data, generating standardized data packets, including: The protocol adaptation module is activated in the internal security zone to automatically identify the communication protocol type of the data source and establish a data transmission connection according to the communication protocol type to receive the multi-source heterogeneous raw data. The multimodal parsing engine is invoked to identify the data format of the multi-source heterogeneous raw data; Based on the identified data format, the multimodal parsing engine is invoked to perform differential parsing; when the data format is structured data, a standardized data model mapping is performed; when the data format is semi-structured data, a pattern inference algorithm is used to infer the meaning of the fields; when the data format is unstructured data, an artificial intelligence recognition model is invoked to perform structured feature extraction. The multimodal parsing engine is used to perform spatiotemporal reference alignment processing on the parsed data generated after differential parsing to generate spatiotemporally aligned data; The spatiotemporally aligned data is then encapsulated in a unified format to generate standardized data packets.

4. The cross-network heterogeneous data fusion processing method as described in claim 1, characterized in that, The semantic governance center performs metadata registration, semantic tagging, and synonym mapping operations on the standardized data packets to generate standardized data after governance, including: The semantic governance center generates a unique identifier for the standardized data packets and records the source path, processing process, and usage scenario of the standardized data packets to form a data lineage graph. The semantic governance center invokes natural language processing models and expert policy libraries to annotate the standardized data packets with business semantic tags. By retrieving a pre-defined standard terminology dictionary through the semantic governance center, the non-standard field names in the standardized data packet are mapped to standard names; The semantic governance center uses a dynamic data quality scoring model to analyze the quality score of the standardized data packets and triggers an alarm when the quality score is lower than a preset threshold. Based on the data lineage map, the business semantic tags, the standard names, and the quality scores, standardized data after governance is generated.

5. The cross-network heterogeneous data fusion processing method as described in claim 1, characterized in that, Analyze the time attributes and access frequency characteristics of the standardized data after governance, and write the standardized data after governance into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer, and a cold data storage layer according to the time attributes and access frequency characteristics, including: Analyze the standardized data after governance, extract the timestamps of the standardized data after governance to determine the time attributes, and analyze the historical access records of the standardized data after governance to determine the access popularity characteristics; If the time attribute is within the first preset time period and the access popularity feature indicates high-frequency access, then the standardized data after treatment will be written into the hot data storage layer. If the time attribute is within the first preset time period and the access popularity feature indicates non-high frequency access, or if the time attribute exceeds the first preset time period but does not exceed the second preset time period, then the standardized data after treatment will be written into the warm data storage layer. If the time attribute exceeds the second preset time period, the standardized data after treatment will be written into the cold data storage layer. A unified data index is established between the hot data storage layer, the warm data storage layer, and the cold data storage layer.

6. The cross-network heterogeneous data fusion processing method as described in claim 1, characterized in that, Through a unified data service gateway, in response to data access requests, the standardized, governed data stored in the tiered storage architecture is output, including: The unified data service gateway is configured with a variety of standard service interfaces, including a structured query language interface for custom queries, a representation state transition interface for calling specific datasets, a streaming subscription interface for real-time data push, and an embedded component interface for visualization integration. The system receives data call requests through the standard service interface and performs identity and permission verification and traffic rate limiting checks on the requester initiating the data call request. When the identity and permission verification is successful and the traffic rate limit check is passed, the query conditions of the data call request are parsed, and the corresponding standardized data after governance is retrieved from the hierarchical storage architecture according to the query conditions; The retrieved standardized data after governance is output through the unified data service gateway, and the interface call statistics for the data call requests are recorded.

7. The cross-network heterogeneous data fusion processing method as described in claim 1, characterized in that, After responding to data request through the unified data service gateway and outputting the standardized, governed data stored in the tiered storage architecture, the system further includes: Record and analyze access logs of historical data call requests to obtain user access habit characteristics and business cycle pattern characteristics; Based on the user access habit characteristics and the business cycle pattern characteristics, the prediction module is used to predict the target dataset that will be frequently accessed in the next time period. Locate the target dataset from the hierarchical storage architecture and pre-extract the target dataset into the memory cache; Establish an access link between the memory cache and the unified data service gateway so that when a data call request for the target dataset is received, data is responded to from the memory cache first.

8. A cross-network heterogeneous data fusion processing device, characterized in that, The cross-network heterogeneous data fusion processing device includes: The cross-network data transfer module is used to build a cross-network data transfer channel between the external network area and the internal security area, and to migrate multi-source heterogeneous raw data generated in the external network area to the internal security area using a one-way isolation transmission mechanism. The multimodal parsing module is used to receive the multi-source heterogeneous raw data through the protocol adaptation module in the internal security area, and call the multimodal parsing engine to perform structured feature extraction and spatiotemporal benchmark alignment processing on the multi-source heterogeneous raw data to generate standardized data packets; The semantic governance module is used to perform metadata registration, semantic tag annotation, and synonym mapping operations on the standardized data packets through the semantic governance center to generate standardized data after governance. The hierarchical storage scheduling module is used to analyze the time attributes and access popularity characteristics of the standardized data after governance, and write the standardized data after governance into a hierarchical storage architecture consisting of a hot data storage layer, a warm data storage layer and a cold data storage layer according to the time attributes and access popularity characteristics. The unified data service module is used to respond to data call requests through the unified data service gateway to output the standardized data after governance stored in the hierarchical storage architecture.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a cross-network heterogeneous data fusion processing program stored in the memory and executable on the processor. When the cross-network heterogeneous data fusion processing program is executed by the processor, it implements the steps of the cross-network heterogeneous data fusion processing method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a cross-network heterogeneous data fusion processing program, which, when executed by a processor, implements the steps of the cross-network heterogeneous data fusion processing method as described in any one of claims 1-7.