Real-time data loading method and system for digital twin channel
By acquiring and preprocessing multi-source heterogeneous data in real time, and using dynamic loading strategies and spatiotemporal semantic maps to generate a compact hash code collection in the digital twin waterway system, the problem of insufficient data loading efficiency and stability in the existing technology is solved, and efficient data management and query is realized, which is suitable for fine query requirements in complex scenarios.
Patent Information
- Application Number
- CN202510472359.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
AI Technical Summary
The existing digital twin channel technology has challenges in data loading efficiency and stability. The data transmission delay is high, the index structure is difficult to meet the needs of fast and complex query, and the data formats from different sources are different, and the standardization and interoperability are insufficient.
By acquiring and preprocessing multi-source heterogeneous data in real time, the data is loaded into the digital twin waterway system using a dynamic loading strategy, and a spatiotemporal semantic map is built to generate a compact hash code collection to achieve efficient data management and query.
It significantly improves the channel operation efficiency and system response speed, reduces system overhead, improves query efficiency and matching accuracy, is suitable for fine query requirements in complex scenarios, and provides solid technical support for real-time perception and intelligent decision-making of the channel.
Smart Images

Figure CN119988693A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital twin waterways, and in particular relates to a real-time data loading method and system for a digital twin waterway. Background Art
[0002] The core of digital twin waterway technology is to build a virtual digital model corresponding to the real waterway environment, collect multi-source heterogeneous data in the waterway environment in real time, such as meteorology, hydrology, ship dynamics, etc., and load and index them efficiently in the digital model, so as to achieve a comprehensive and accurate perception of the waterway environment. This will not only help improve shipping efficiency and reduce operating costs, but also significantly improve the safety of waterway passage and reduce accident risks.
[0003] In the existing technology, data loading for the digital twin waterway system is realized through advanced sensor networks and adaptive data acquisition and transmission algorithms, and the multi-dimensional index structure of distributed storage and the semantic web-based method are used to organize and manage massive multi-source heterogeneous data to support fast query and complex analysis needs. However, these methods still have certain shortcomings: with the increase in the number of sensors and the sharp increase in the amount of data, the existing technology faces challenges in data loading efficiency and stability, and the data transmission delay is high; for the efficient management of massive heterogeneous data, the existing index structure is difficult to meet the fast and complex query requirements; in addition, the data formats from different sources are different, and the realization of standardization and interoperability still needs to be strengthened.
[0004] Therefore, there is an urgent need to develop a real-time data loading method and system for digital twin waterways, which can reduce system overhead, improve query efficiency, ensure the accuracy of results, and be suitable for the fine query needs in complex scenarios of digital twin waterways. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a real-time data loading method and system for a digital twin waterway, which can reduce system overhead, improve query efficiency, ensure the accuracy of results, and is suitable for the fine query needs in complex scenarios of digital twin waterways.
[0006] The present invention provides a real-time data loading method for a digital twin waterway, the method comprising the following steps: S1, real-time acquisition of multi-source heterogeneous data of actual waterways and pre-processing; S2, loading the pre-processed multi-source heterogeneous data into the digital twin waterway system through a dynamic loading strategy; S3. Construct a spatiotemporal semantic graph in the digital twin waterway system, and obtain a compact hash code set based on the spatiotemporal semantic graph; S4, converting the target query data into a hash code and matching it with the compact hash code set to obtain an index result; S5. Generate a waterway operation strategy based on the index results.
[0007] Furthermore, in S1, real-time acquisition of multi-source heterogeneous data of the actual waterway and preprocessing include: S11, real-time acquisition of multi-source heterogeneous data of actual waterways; S12, perform spatiotemporal alignment and missing value interpolation processing on multi-source heterogeneous data of actual waterways; S13. Calculate the timeliness weight, importance weight and usage frequency weight of the processed multi-source heterogeneous data.
[0008] Furthermore, in S13, the timeliness weight, importance weight, and usage frequency weight of the multi-source heterogeneous data after calculation and processing include: The calculation formula of timeliness weight is as follows: ; Where T(d) represents the timeliness weight, t represents the current time, and t d Indicates the generation time of data, T w Indicates the valid time window of the data; The calculation formula of importance weight is as follows: ; Among them, I(d) represents the importance weight, R d Indicates the influence range of the data, max (R d ) represents the maximum value of the data impact range; The frequency weight is calculated as follows: ; ; Among them, U(d) represents the usage frequency weight, F d Indicates the access frequency of data, max (F d ) represents the maximum value of data access frequency, FT w Indicates that in the effective time window T w The number of times the data is accessed.
[0009] Furthermore, in S2, the pre-processed multi-source heterogeneous data is loaded into the digital twin waterway system through a dynamic loading strategy, including: S21, calculating the data loading priority of the pre-processed multi-source heterogeneous data according to the timeliness weight, importance weight and usage frequency weight; S22, allocating the preprocessed multi-source heterogeneous data to the memory database or disk database of the digital twin waterway system according to the data loading priority and the priority threshold; If the data loading priority is greater than the priority threshold, the corresponding preprocessed multi-source heterogeneous data will be allocated to the in-memory database of the digital twin waterway system; If the data loading priority is less than or equal to the priority threshold, the corresponding pre-processed multi-source heterogeneous data will be allocated to the disk database of the digital twin waterway system; S23. Dynamically adjust the storage location of the pre-processed multi-source heterogeneous data through a periodic update mechanism.
[0010] Furthermore, in S3, a spatiotemporal semantic graph is constructed in the digital twin waterway system, and a compact hash code set is obtained based on the spatiotemporal semantic graph, including: S31. Construct a spatiotemporal semantic graph in the digital twin waterway system; define the nodes of the spatiotemporal semantic graph to represent ships, navigation marks, waterway sections and ship attributes, and define the edges of the spatiotemporal semantic graph to represent the spatiotemporal relationship between nodes; S32, input the spatiotemporal semantic graph into the graph attention network to obtain the semantic embedding vector of each node; The calculation formula is as follows: ; Among them, G represents the spatiotemporal semantic graph, X represents the initial feature matrix of the node, and h v Represents the semantic embedding vector of node v; S33, inputting the semantic embedding vector of each node into the neural hash model to obtain a compact hash code of each node to form a compact hash code set; The calculation formula is as follows: ; Among them, b v represents a set of compact hash codes for node v, and Hash represents a hash function.
[0011] Furthermore, in S4, the target query data is converted into a hash code and matched with the compact hash code set, and the index results obtained include: S41, constructing a target query feature vector according to the target query data, and obtaining a hash code of the target query data through a hash function; S42, calculating the similarity between the hash code of the target query data and each node in the compact hash code set, and taking the nodes whose similarity is greater than or equal to a preset threshold as a candidate set; S43, constructing a target query graph according to the target query feature vector; S44. Execute a subgraph matching strategy in the candidate set according to the target query graph, and use the matching result with the highest similarity score as the index result.
[0012] Furthermore, in S42, the similarity between the hash code of the target query data and each node in the compact hash code set is calculated, and the calculation formula is as follows: ; Among them, S(b q ,b v ) represents the similarity between the hash code of the target query data and node v, b q represents the hash code of the target query data, and ||·|| represents the modulus length of the vector.
[0013] Further, in S43, a target query graph is constructed according to the target query feature vector, including: S431, mapping the target query data into n query nodes, each query node including a corresponding feature vector; S432, constructing query edges between query nodes according to the relationship between nodes, and assigning features to the edges; S433. Construct a target query graph according to the query nodes and query edges.
[0014] The present invention also provides a real-time data loading system for a digital twin waterway, which is used to execute the real-time data loading method for a digital twin waterway, and is characterized in that the system includes the following modules: Data acquisition module, used to obtain multi-source heterogeneous data of actual waterways in real time and perform preprocessing; The dynamic loading module is connected to the data acquisition module and is used to load the pre-processed multi-source heterogeneous data into the digital twin waterway system through a dynamic loading strategy; The hash code conversion module is connected with the dynamic loading module and is used to construct a spatiotemporal semantic graph in the digital twin waterway system and obtain a compact hash code set according to the spatiotemporal semantic graph; An indexing module, connected to the hash code conversion module, is used to convert the target query data into a hash code and match it with a compact hash code set to obtain an index result; An output module is connected to the index module and is used to generate a waterway operation strategy according to the index result.
[0015] The embodiments of the present invention have the following technical effects: This solution realizes real-time perception, intelligent analysis and efficient decision support of waterway operation status through the digital twin waterway system combined with dynamic loading strategy of multi-source heterogeneous data, construction of spatiotemporal semantic graph and efficient subgraph matching method; the dynamic loading strategy prioritizes loading key data into the memory database through comprehensive calculation of timeliness weight, importance weight and usage frequency weight, which significantly improves the waterway operation efficiency and system response speed, and avoids the performance bottleneck caused by traditional static loading method; the spatiotemporal semantic graph converts multi-source heterogeneous data into structured graph form, and uses graph attention network to generate compact hash code, which realizes efficient representation and storage of complex semantic information and reduces storage and computing costs. At the same time, the fast screening and subgraph matching algorithm based on hash code can accurately locate the subgraph related to the target query, improving query efficiency and matching accuracy; the optimized storage mechanism performs hierarchical storage according to the importance and usage frequency of data, effectively reducing system overhead and improving overall operation efficiency; the compact hash code greatly reduces the amount of data that needs to be compared during query, speeds up retrieval speed, and significantly reduces computational complexity. Performing subgraph matching after screening out the candidate set not only improves query efficiency, but also ensures the accuracy of the results. It is particularly suitable for sophisticated query needs in complex scenarios, provides solid technical support for real-time perception and intelligent decision-making of waterways, and comprehensively improves the intelligence and efficiency of waterway management. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 It is a flow chart of a real-time data loading method for a digital twin waterway provided in an embodiment of the present invention; Figure 2 It is a logic diagram of a dynamic data loading method for a digital twin waterway provided in an embodiment of the present invention; Figure 3 This is a loading time comparison diagram of a loading technology of the present solution and a traditional loading technology provided by an embodiment of the present invention; Figure 4 This is a query response time comparison diagram of an indexing technology of the present solution and a traditional indexing technology provided by an embodiment of the present invention; Figure 5 It is a structural schematic diagram of a real-time data loading system for a digital twin waterway provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.
[0019] The embodiment of the present invention provides a real-time data loading method for a digital twin waterway, which is implemented based on a digital twin waterway system. The construction method of the digital twin waterway system is as follows: A1. First, data collection is carried out, including: Terrain data collection: The real-life data of the terrain along the waterway is collected through oblique photogrammetry technology, and high-resolution images are obtained using drones and other equipment, which are then processed by software to generate a three-dimensional real-life model of the terrain. At the same time, data such as topographic maps and digital elevation models (DEMs) can also be obtained through satellite remote sensing, geographic information systems (GIS), and other means to provide a basis for the scene base model.
[0020] Underwater terrain data collection: Use multi-beam bathymetric systems and other equipment to measure the waterway depth, obtain water depth point and isobath data, and after processing, obtain underwater terrain data in raster form for constructing underwater terrain models.
[0021] Hydrometeorological data collection: Use hydrometeorological observation stations, buoys and other equipment to collect hydrometeorological data such as water level, flow velocity, flow direction, wind speed and wind direction. These data will provide real-time environmental information for the digital twin waterway system and be used to simulate the dynamic changes of water bodies.
[0022] Video surveillance data collection: Video surveillance cameras are installed along the waterway and at important nodes to obtain real-time video data of the waterway, provide intuitive visual information for the digital twin waterway, and assist in ship monitoring, waterway condition monitoring, etc.
[0023] Ship data collection: Receive the ship’s position, speed, heading and other information through the ship’s automatic identification system (AIS), realize real-time tracking and monitoring of the ship, and provide data support for the dynamic simulation of the ship in the digital twin channel.
[0024] A2. Perform data processing and fusion based on the collected data, including: Data cleaning and preprocessing: Clean the collected data to remove noise, errors and redundant information to ensure the accuracy and reliability of the data. For example, perform radiation correction and geometric correction on the oblique photogrammetry data to improve the image quality; perform filtering and interpolation on the water depth data to make it smoother and more complete.
[0025] Data fusion: Fusion of data from different sources and in different formats to build a unified data model. For example, terrain data can be fused with underwater terrain data to form an integrated terrain model above and below the water; hydrological and meteorological data can be associated and fused with video surveillance data, ship data, etc. to achieve data interconnection and collaborative analysis.
[0026] A3. Perform 3D modeling based on the processed data, including: Terrain modeling along the route: Based on the processed terrain data, use 3D modeling software to construct a 3D model of the terrain along the route, including mountains, valleys, plains and other terrain features, to provide a real geographical background for the digital twin waterway.
[0027] Underwater terrain modeling: Based on underwater terrain data, professional underwater terrain modeling technology is used to construct a three-dimensional model of the underwater terrain, simulating underwater landforms such as riverbeds, reefs, underwater obstacles, etc., to provide a reference for ship navigation safety.
[0028] Modeling of navigation-related buildings: For navigation-related buildings on the waterway, such as locks, bridges, docks, navigation marks, etc., detailed modeling is carried out using 3D modeling software based on their design drawings, appearance photos and other information to ensure the accuracy and authenticity of the model so that it can be effectively managed and monitored in the digital twin waterway.
[0029] Ship modeling: Based on the type, size and other parameters of the ship, a three-dimensional model library of the ship is constructed using parametric modeling methods. When it is necessary to display the ship in the digital twin channel, the corresponding model can be selected from the model library based on the ship's AIS data for real-time drawing and display.
[0030] A4. Build a platform based on the constructed 3D model and collected data, including: Build a 3D simulation module: Select a suitable digital twin engine, such as Unity, Unreal Engine, etc., build a 3D simulation module, import the constructed 3D model into the engine, render and display the scene, and realize 3D visualization of the waterway.
[0031] Construction of scene roaming module: Develop the scene roaming function to allow users to roam freely in the digital twin waterway, view detailed information of the waterway from different angles and positions, and improve users' intuitive feeling and understanding of the waterway.
[0032] Data-driven module construction: Establish a data-driven mechanism to associate and drive the collected real-time data with the three-dimensional model, so that the model can be updated and dynamically responded to data changes in real time. For example, the position and navigation state of the ship model in the digital twin channel can be driven by the ship's AIS data; the dynamic changes of the water model can be driven by hydrological and meteorological data, etc.
[0033] Data loading and indexing module construction: including the Figure 5 As shown in the structure, a real-time data loading method for a digital twin waterway provided in an embodiment of the present invention is implemented based on a data loading and indexing module in the digital twin waterway system.
[0034] A5. Integrate the built platform to obtain a digital twin waterway system, including: The three-dimensional simulation module, scene roaming module, data driving module and data loading and indexing module are integrated to form a complete digital twin waterway system. Data is exchanged between modules. For example, the three-dimensional simulation module relies on the dynamic loading module in the data loading and indexing module to provide real-time data flow, the scene roaming module can achieve fast spatial retrieval through the hash code conversion module and indexing module in the data loading and indexing module, and the data driving module forms feedback control through the retrieval results output by the output module in the data loading and indexing module and each three-dimensional model.
[0035] In addition, the digital twin waterway system is also connected with relevant business systems of waterway management, such as ship dispatching systems, waterway maintenance systems, etc., to achieve data sharing and business collaboration, and provide comprehensive support for the operation and management of the waterway.
[0036] Figure 1 is a flowchart of a real-time data loading method for a digital twin waterway provided by an embodiment of the present invention, see Figure 1 , the method comprises the following steps: S1. Acquire multi-source heterogeneous data of the actual waterway in real time and perform preprocessing.
[0037] S11. Obtain multi-source heterogeneous data of the actual waterway in real time.
[0038] In some embodiments, the multi-source heterogeneous data of the actual waterway include, but are not limited to, waterway perception data, including ship position, speed, heading, navigation mark position, water depth, flow rate, etc.; waterway environmental information, such as meteorological data (wind speed, wind direction, temperature, humidity), hydrological data (tides, water level changes), etc.; historical operation data of the waterway, including ship trajectories, event records, traffic flow, etc.
[0039] S12. Perform spatiotemporal alignment and missing value interpolation processing on multi-source heterogeneous data of actual waterways.
[0040] In some embodiments, performing spatiotemporal alignment on multi-source heterogeneous data of an actual waterway may include: The timestamps collected by different devices are unified into UTC time, and the location coordinates are converted into the WGS84 standard.
[0041] In some embodiments, performing missing value interpolation processing on multi-source heterogeneous data of actual waterways may include: Use the space-time interpolation method to complete the missing points of the ship trajectory: ; ; Where x(t) represents the position of the ship at time t, t i Indicates t i At the same time, v represents the average speed of the ship at adjacent sampling points.
[0042] S13. Calculate the timeliness weight, importance weight and usage frequency weight of the processed multi-source heterogeneous data.
[0043] Among them, the timeliness weight represents the real-time demand of data, the importance weight represents the impact of data on waterway operation, and the frequency weight represents the access frequency of data.
[0044] In some embodiments, the calculation formula of the timeliness weight is as follows: ; Where T(d) represents the timeliness weight, t represents the current time, and t d Indicates the generation time of data, T w Represents the effective time window of the data; the formula shows that the greater the difference between the data generation time and the current time, the lower the timeliness weight; The calculation formula of importance weight is as follows: ; Among them, I(d) represents the importance weight, R d Indicates the influence range of the data, max (R d ) represents the maximum value of the data impact range; The frequency weight is calculated as follows: ; ; Among them, U(d) represents the usage frequency weight, F d Indicates the access frequency of data, max (F d ) represents the maximum value of data access frequency, FT w Indicates that in the effective time window T w The number of times the data is accessed.
[0045] S2. Load the preprocessed multi-source heterogeneous data into the digital twin waterway system through a dynamic loading strategy.
[0046] Figure 2This is a logic diagram of a dynamic data loading method for a digital twin waterway provided by an embodiment of the present invention, see Figure 2 , S2 includes the following sub-steps: S21. Calculate the data loading priority of the preprocessed multi-source heterogeneous data according to the timeliness weight, importance weight and usage frequency weight.
[0047] In some embodiments, a data loading priority function may be defined as P(d) to calculate the data loading priority of multi-source heterogeneous data: P(d)=w1×T(d)+w2×I(d)+w3×U(d); Among them, w1, w2, and w3 represent weight coefficients respectively, and satisfy w1+w2+w3=1.
[0048] The priority of data is dynamically calculated by comprehensively considering timeliness, importance, and frequency of use.
[0049] S22. Allocate the pre-processed multi-source heterogeneous data to the memory database or disk database of the digital twin waterway system according to the data loading priority and priority threshold.
[0050] In some embodiments, the priority threshold may be set as needed, for example, based on an empirical value or by setting different priority thresholds and testing system performance, and selecting the value with the best performance as the priority threshold.
[0051] In some embodiments, if the data loading priority is greater than a priority threshold, the corresponding pre-processed multi-source heterogeneous data is allocated to the in-memory database of the digital twin waterway system for quick access; If the data loading priority is less than or equal to the priority threshold, the corresponding preprocessed multi-source heterogeneous data will be allocated to the disk database of the digital twin waterway system and loaded on demand.
[0052] Through dynamic loading strategies, the use of memory and storage resources is optimized to improve system performance.
[0053] S23. Dynamically adjust the storage location of the pre-processed multi-source heterogeneous data through a periodic update mechanism.
[0054] In some embodiments, a periodic update mechanism can also be set, such as recalculating data priority every 5 seconds, dynamically adjusting the storage location of preprocessed multi-source heterogeneous data, and transferring data from the memory database to the disk database if the data priority in the memory database drops below the threshold; if the data priority in the disk database rises above the threshold, it is loaded into the memory database. This dynamic loading strategy can achieve a balance between memory and storage resources, ensuring the efficient operation of the digital twin waterway system.
[0055] S3. Construct a spatiotemporal semantic graph in the digital twin waterway system, and obtain a compact hash code set based on the spatiotemporal semantic graph.
[0056] S31. Construct a spatiotemporal semantic map in the digital twin waterway system.
[0057] The nodes that define the spatiotemporal semantic graph represent ships, navigation marks, channel segments, and ship attributes (such as heading, speed, timestamp, etc.), and the edges that define the spatiotemporal semantic graph represent the spatiotemporal relationships between nodes (such as the distance between nodes, relative orientation, etc.).
[0058] The spatiotemporal semantic graph is used to describe the dynamic characteristics and temporal and spatial relationships of the waterway.
[0059] S32. Input the spatiotemporal semantic graph into the graph attention network to obtain the semantic embedding vector of each node.
[0060] The semantic embedding vector is used to capture the complex semantic information of a node. The calculation formula of the semantic embedding vector is as follows: ; Where G represents the spatiotemporal semantic graph, X represents the initial feature matrix of the node, X is a V×D matrix, V represents the number of nodes in the graph, D represents the feature dimension of each node, the vth row element represents the initial feature vector of the vth node, and h v Represents the semantic embedding vector of node v.
[0061] S33, inputting the semantic embedding vector of each node into the neural hash model to obtain a compact hash code of each node to form a compact hash code set; Compact hash codes are used for efficient retrieval and matching. The calculation formula for compact hash codes is as follows: ; Among them, b v represents a compact hash code set of node v, and Hash represents a hash function used to map a high-dimensional embedding vector to a low-dimensional hash code space.
[0062] In this way, complex spatiotemporal semantic graphs can be converted into efficient hash code representations to provide support for subsequent fast retrieval.
[0063] S4. Convert the target query data into a hash code and match it with the compact hash code set to obtain the index result.
[0064] S41. Construct a target query feature vector according to the target query data, and obtain a hash code of the target query data through a hash function.
[0065] In some embodiments, the target query data is represented as a target query feature vector; The high-dimensional feature vector is mapped to the low-dimensional hash code space through random projection to obtain the hash code of the target query data: b q =sign(R×q); Among them, represents, sign represents the sign function, R represents the randomly generated projection matrix, and q represents the target query feature vector.
[0066] S42, calculating the similarity between the hash code of the target query data and each node in the compact hash code set, and taking the nodes whose similarity is greater than or equal to a preset threshold as a candidate set.
[0067] In some embodiments, the calculation formula for the similarity between the hash code of the target query data and each node in the compact hash code set is as follows: ; Among them, S(b q ,b v ) represents the similarity between the hash code of the target query data and node v, b q represents the hash code of the target query data, and ||·|| represents the modulus length of the vector.
[0068] The candidate set C is filtered according to the similarity, and the graph subgraph matching operation is performed only on the nodes in the candidate set C. The candidate set is filtered by hash code, which can reduce the amount of calculation and improve the query efficiency.
[0069] In some embodiments, the Hamming distance between the hash code of the target query data and each node in the compact hash code set may be calculated, and nodes whose Hamming distance is less than or equal to a preset byte may be retained as a candidate set; illustratively, the preset byte may be 3.
[0070] S43. Construct a target query graph according to the target query feature vector.
[0071] S431. Map the target query data into n query nodes, each query node including a corresponding feature vector.
[0072] S432: construct a query edge between query nodes according to the relationship between the nodes, and assign features to the edge (such as the distance between two nodes, relative orientation, etc.).
[0073] S433, constructing a target query graph G based on query nodes and query edges q =(V q ,E q ), where V q Represents the query node set, E q Represents the query edge set.
[0074] S44. Execute a subgraph matching strategy in the candidate set according to the target query graph, and use the matching result with the highest similarity score as the index result.
[0075] In some embodiments, the subgraph matching algorithm may adopt a VF2 algorithm, an approximate matching algorithm, etc. Taking the VF2 algorithm as an example, the process of performing the subgraph matching strategy is as follows: Sa, for the target query graph G q For each query node o, find the most similar node v in the candidate set C through the similarity scoring function: ; Among them, S(o,v) represents the similarity score between the query node o and the target node v, φ(o) represents the feature vector of the query node o, and φ(v) represents the feature vector of the target node v.
[0076] Sb, for the target query graph G q For each query edge, the corresponding edge is found in the candidate set C through the similarity scoring function. The principle is the same as Sa.
[0077] Sc, sum the similarity scores of nodes and edges to get the comprehensive similarity score, Sd. During the search process, the matching result with the highest comprehensive similarity score is recorded as the index result.
[0078] S5. Generate a waterway operation strategy based on the index results.
[0079] Generate channel control instructions based on the index results and send them to channel equipment and ships.
[0080] This embodiment compares the loading technology and indexing technology of this solution with the traditional loading technology and indexing technology. Figure 3 This is a loading time comparison diagram of a loading technology provided by an embodiment of the present invention and a traditional loading technology. Figure 4 This is a comparison chart of query response time between the indexing technology of the present solution and the traditional indexing technology provided by an embodiment of the present invention, see Figure 3 , Figure 4 And Table 1, Table 2.
[0081] Table 1 Comparison of loading technologies
[0082] Table 2 Index technology comparison table
[0083] In summary, the traditional technology uses a relational database to insert one entry at a time, and the daily loading time is 30 minutes (1800 seconds). The single-threaded processing mode results in a CPU utilization rate of only 20%, and there is a significant phenomenon of resource idleness. The linear growth of processing time is positively correlated with the amount of data, and the system throughput is about 0.55GB / min (1TB / 30min). This solution uses a multi-source heterogeneous data dynamic loading strategy to achieve parallel processing and pipeline optimization, and the daily loading time is shortened to 10 minutes. The use of multi-core parallel computing technology increases the CPU utilization rate to 80%, and the system throughput reaches 20GB / min (200GB / 10min), which is 36 times more efficient. The resource utilization of traditional technologies presents an unbalanced state of "low CPU-high memory". This solution can achieve coordinated optimization of computing and storage resources.
[0084] Traditional technology takes an average of 2 seconds (2000ms) to query regional ships (5km² range) during indexing, with a standard deviation of ±300ms for response delay. It uses a single B+ tree index structure, which requires traversing multiple layers of tree nodes for range queries, with I / O access times as high as 120 times / query, and lacks a spatial retrieval optimization mechanism. This solution technology compresses the response time of the same query to 500ms through spatiotemporal semantic graph reconstruction, improving efficiency by 75%. It uses a composite structure of spatiotemporal semantic graph subgraph matching and compact hash code sets, optimizing I / O access times to 20 times / query.
[0085] In summary, this solution has been verified by stress testing and has a significant performance improvement over traditional technologies.
[0086] This solution realizes real-time perception, intelligent analysis and efficient decision support of waterway operation status through the digital twin waterway system combined with dynamic loading strategy of multi-source heterogeneous data, construction of spatiotemporal semantic graph and efficient subgraph matching method; the dynamic loading strategy prioritizes loading key data into the memory database through comprehensive calculation of timeliness weight, importance weight and usage frequency weight, which significantly improves the waterway operation efficiency and system response speed, and avoids the performance bottleneck caused by traditional static loading method; the spatiotemporal semantic graph converts multi-source heterogeneous data into structured graph form, and uses graph attention network to generate compact hash code, which realizes efficient representation and storage of complex semantic information and reduces storage and computing costs. At the same time, the fast screening and subgraph matching algorithm based on hash code can accurately locate the subgraph related to the target query, improving query efficiency and matching accuracy; the optimized storage mechanism performs hierarchical storage according to the importance and usage frequency of data, effectively reducing system overhead and improving overall operation efficiency; the compact hash code greatly reduces the amount of data that needs to be compared during query, speeds up retrieval speed, and significantly reduces computational complexity. Performing subgraph matching after screening out the candidate set not only improves query efficiency, but also ensures the accuracy of the results. It is particularly suitable for sophisticated query needs in complex scenarios, provides solid technical support for real-time perception and intelligent decision-making of waterways, and comprehensively improves the intelligence and efficiency of waterway management.
[0087] The embodiment of the present invention also provides a real-time data loading system for a digital twin waterway, which is used to execute the real-time data loading method for a digital twin waterway. Figure 5 is a structural diagram of a real-time data loading system for a digital twin waterway provided by an embodiment of the present invention, see Figure 5 , the system includes the following modules: Data acquisition module, used to obtain multi-source heterogeneous data of actual waterways in real time and perform preprocessing; The dynamic loading module is connected to the data acquisition module and is used to load the pre-processed multi-source heterogeneous data into the digital twin waterway system through a dynamic loading strategy; The hash code conversion module is connected with the dynamic loading module and is used to construct a spatiotemporal semantic graph in the digital twin waterway system and obtain a compact hash code set according to the spatiotemporal semantic graph; An indexing module, connected to the hash code conversion module, is used to convert the target query data into a hash code and match it with a compact hash code set to obtain an index result; An output module is connected to the index module and is used to generate a waterway operation strategy according to the index result.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A real-time data loading method for a digital twin waterway, characterized in that: The method comprises the following steps: S1, real-time acquisition of multi-source heterogeneous data of actual waterways and pre-processing; S2, loading the pre-processed multi-source heterogeneous data into the digital twin waterway system through a dynamic loading strategy; S3, constructing a spatiotemporal semantic graph in the digital twin waterway system, and obtaining a compact hash code set according to the spatiotemporal semantic graph; S4, converting the target query data into a hash code, and matching it with the compact hash code set to obtain an index result; S5. Generate a waterway operation strategy according to the index result.
2. A real-time data loading method for a digital twin waterway according to claim 1, characterized in that: In S1, real-time acquisition of multi-source heterogeneous data of the actual waterway and preprocessing include: S11, real-time acquisition of multi-source heterogeneous data of actual waterways; S12, performing spatiotemporal alignment and missing value interpolation processing on the multi-source heterogeneous data of the actual waterway; S13. Calculate the timeliness weight, importance weight and usage frequency weight of the processed multi-source heterogeneous data.
3. A real-time data loading method for a digital twin waterway according to claim 2, characterized in that: In S13, the timeliness weight, importance weight and usage frequency weight of the processed multi-source heterogeneous data are calculated and processed, including: The calculation formula of timeliness weight is as follows: ; Where T(d) represents the timeliness weight, t represents the current time, and t d Indicates the generation time of data, T w Indicates the valid time window of the data; The calculation formula of importance weight is as follows: ; Among them, I(d) represents the importance weight, R d Indicates the influence range of the data, max (R d ) represents the maximum value of the data impact range; The frequency weight is calculated as follows: ; ; Among them, U(d) represents the usage frequency weight, F d Indicates the access frequency of data, max (F d ) represents the maximum value of data access frequency, FT w Indicates that in the effective time window T w The number of times the data is accessed.
4. A real-time data loading method for a digital twin waterway according to claim 2, characterized in that: In S2, loading the pre-processed multi-source heterogeneous data into the digital twin waterway system through a dynamic loading strategy includes: S21, calculating the data loading priority of the pre-processed multi-source heterogeneous data according to the timeliness weight, importance weight and usage frequency weight; S22, allocating the preprocessed multi-source heterogeneous data to the memory database or disk database of the digital twin waterway system according to the data loading priority and priority threshold; If the data loading priority is greater than the priority threshold, the corresponding pre-processed multi-source heterogeneous data is allocated to the in-memory database of the digital twin waterway system; If the data loading priority is less than or equal to the priority threshold, the corresponding pre-processed multi-source heterogeneous data is allocated to the disk database of the digital twin waterway system; S23. Dynamically adjust the storage location of the pre-processed multi-source heterogeneous data through a periodic update mechanism.
5. The real-time data loading method for a digital twin waterway according to claim 1 is characterized in that: In S3, a spatiotemporal semantic graph is constructed in the digital twin waterway system, and a compact hash code set is obtained according to the spatiotemporal semantic graph, including: S31, constructing a spatiotemporal semantic graph in the digital twin waterway system; defining the nodes of the spatiotemporal semantic graph to represent ships, navigation marks, waterway sections and ship attributes, and defining the edges of the spatiotemporal semantic graph to represent the spatiotemporal relationship between nodes; S32, inputting the spatiotemporal semantic graph into a graph attention network to obtain a semantic embedding vector for each node; The calculation formula is as follows: ; Among them, G represents the spatiotemporal semantic graph, X represents the initial feature matrix of the node, and h v Represents the semantic embedding vector of node v; S33, inputting the semantic embedding vector of each node into the neural hash model to obtain a compact hash code of each node to form a compact hash code set; The calculation formula is as follows: ; Among them, b v represents a set of compact hash codes for node v, and Hash represents a hash function.
6. A real-time data loading method for a digital twin waterway according to claim 5, characterized in that: In S4, the target query data is converted into a hash code and matched with the compact hash code set to obtain an index result including: S41, constructing a target query feature vector according to the target query data, and obtaining a hash code of the target query data through a hash function; S42, calculating the similarity between the hash code of the target query data and each node in the compact hash code set, and taking the nodes whose similarity is greater than or equal to a preset threshold as a candidate set; S43, constructing a target query graph according to the target query feature vector; S44. Execute a subgraph matching strategy in the candidate set according to the target query graph, and use the matching result with the highest similarity score as the index result.
7. A real-time data loading method for a digital twin waterway according to claim 6, characterized in that: In S42, the similarity between the hash code of the target query data and each node in the compact hash code set is calculated, and the calculation formula is as follows: ; Among them, S(b q ,b v ) represents the similarity between the hash code of the target query data and node v, b q represents the hash code of the target query data, and ||·|| represents the modulus length of the vector.
8. A real-time data loading method for a digital twin waterway according to claim 6, characterized in that: In S43, constructing a target query graph according to the target query feature vector includes: S431, mapping the target query data into n query nodes, each query node including a corresponding feature vector; S432, constructing query edges between query nodes according to the relationship between nodes, and assigning features to the edges; S433. Construct a target query graph according to the query nodes and query edges.
9. A real-time data loading system for a digital twin waterway, used to execute a real-time data loading method for a digital twin waterway according to any one of claims 1 to 8, characterized in that: The system includes the following modules: Data acquisition module, used to obtain multi-source heterogeneous data of actual waterways in real time and perform preprocessing; A dynamic loading module, connected to the data acquisition module, is used to load the pre-processed multi-source heterogeneous data into the digital twin waterway system through a dynamic loading strategy; A hash code conversion module, connected to the dynamic loading module, is used to construct a spatiotemporal semantic graph in the digital twin waterway system and obtain a compact hash code set according to the spatiotemporal semantic graph; An indexing module, connected to the hash code conversion module, for converting target query data into hash codes and matching them with the compact hash code set to obtain indexing results; An output module is connected to the index module and is used to generate a waterway operation strategy according to the index result.
Citation Information
Patent Citations
Similar information retrieval method and system based on isoproton map neural network
CN114168804A
Operation and maintenance fault diagnosis and analysis method based on subgraph matching and distributed query
CN114186073A
Intelligent water conservancy inspection method, device and equipment and storage medium
CN119106880A
Quick starting and content preloading method and device of network high-definition player
CN119835486A
Cited By
Channel model dynamic association method based on multi-source data index
CN120372786A