Intelligent port and navigation AI application system under digital twinning enabling
Through the collaborative design of an inland waterway-specific AI model, high-noise voice interaction, and digital twin fusion modules, the problems of model adaptability, voice recognition accuracy, and interaction latency in inland waterway smart port and shipping have been solved, creating an efficient digital twin scenario and improving the efficiency and safety of inland waterway shipping management.
Patent Information
- Application Number
- CN202511822011.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies for smart port and shipping applications in inland waterways suffer from problems such as poor AI model adaptability, low voice recognition accuracy, high interaction latency, insufficient waterway modeling precision, and lack of integration of meteorological factors, making it difficult to achieve accurate decision-making and efficient interaction.
By employing an inland waterway-specific AI model module, a high-noise voice interaction module, a command-scene mapping interaction module, and a digital twin fusion module, and through technologies such as scene feature extraction, multimodal fusion, knowledge graph constraints, model lightweighting, voice enhancement and semantic understanding, command mapping, and high-precision modeling, an efficient smart port and shipping AI application system empowered by digital twins is constructed.
It has achieved real-time perception and accurate decision-making of waterway status, effective voice interaction in noisy environments, accurate mapping of voice commands and digital twin scenarios, and real-time fusion of high-precision waterway models and meteorology, which has improved the efficiency of inland waterway transportation, ensured navigation safety, and reduced operating costs.
Smart Images

Figure CN121527341A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence application systems, and relates to a smart port and navigation AI application system under digital twin empowerment. BACKGROUND
[0002] With the continuous advancement of smart port and navigation construction, digital twin technology, as a key technology for realizing the deep integration of the physical world and the digital world, has been widely applied in inland waterway management, ship lock control and ship scheduling fields. Digital twin technology realizes real-time monitoring, analysis and prediction of entity state by constructing virtual mapping of physical entities, providing a new technical path for smart port and navigation. At present, the construction of inland waterway digital twin scene has made certain progress. CN113223162B discloses a construction method and device of an inland waterway digital twin scene, which includes multi-scale waterway scene construction, multi-channel video three-dimensional real-time fusion, digital asset three-dimensional labeling and Internet of Things sensing data aggregation fusion, and is used to solve the problems of video fragmentation, multi-source Internet of Things data separation and lack of correlation in inland waterway supervision. CN114529680A proposes a digital twin waterway construction method and system, which establishes a digital twin scene through a commercial or open-source engine, monitors changes in scene conditions by accessing sensing data, and applies it to waterway management. In the field of ship lock control, CN113722806A introduces a ship lock control method based on a digital twin system, which acquires ship lock operation data and transmits it to the digital twin system, simulates the ship lock operation state in the digital twin system, and then realizes precise control of the actual ship lock. CN116416393B proposes a construction method of interactive real scene three-dimensional waterway, which realizes real scene simulation of inland waterway based on digital twin technology, realizes real-time display of digital waterway, and achieves the purpose of viewing ship, navigation mark and water target information through interaction. In the aspect of ship navigation control, CN116824912B discloses a digital twin inland ship navigation control system, which perceives ship information in the target water area through artificial intelligence machine vision technology, fuses radar data and AIS data to support dynamic digital twin visualization of inland waterway, and realizes real-time monitoring of regional ship operation safety risk through ship collision risk calculation method.
[0003] However, the existing technology still has the following deficiencies in the application of inland smart port and navigation: Firstly, the existing AI model lacks adaptation to the particularity of inland waterway. Inland waterway has the characteristics of narrow water area, many curves and complex navigation conditions compared with ocean waterway, and general AI model is difficult to accurately perceive the situation of inland waterway and make accurate decisions. At the same time, the real-time performance and accuracy of the model are difficult to balance, especially when deploying on edge devices, the computing resources are limited, and it is difficult to meet the real-time requirements. Secondly, the inland waterway shipping environment is noisy, with engine noise, water flow noise, and wind noise, which seriously affect the accuracy of speech recognition. Existing voice interaction systems have a significantly lower recognition rate in noisy environments and lack the semantic understanding of inland waterway technical terms and crew's colloquial instructions, making it difficult to support efficient voice interaction. Third, the mapping between voice commands and digital twin scene operations in existing technologies is not accurate enough, resulting in problems such as ambiguous commands and incomplete parameters; at the same time, the interaction latency is generally high, making it difficult to achieve the latency requirement of less than 100ms, which affects user experience and operational efficiency. Fourth, the existing 3D modeling accuracy of inland waterways is insufficient, especially the underwater topography modeling has a large error. In addition, meteorological factors have a significant impact on inland waterway navigation, but the existing system lacks effective integration of meteorological simulation and digital twin scenarios. The combination of real-time monitoring images and twin scenarios is also not seamless enough, making it difficult to achieve an intuitive display that combines the virtual and real worlds.
[0004] In summary, there is an urgent need for a smart port and shipping AI application system empowered by digital twins that can solve the above-mentioned technical problems, so as to improve the situational awareness and decision-making capabilities of inland waterways, improve the voice interaction experience in high-noise environments, realize the accurate mapping and low-latency interaction between voice commands and digital twin scenarios, and build a fusion of high-precision waterway 3D models and real-time monitoring images. Summary of the Invention
[0005] The purpose of this invention is to address the above-mentioned problems by providing a smart port and shipping AI application system empowered by digital twins.
[0006] To achieve the above objectives, the present invention adopts the following technical solutions: A smart port and shipping AI application system empowered by digital twins includes an inland waterway-specific AI model module, a high-noise voice interaction module, a command-scene mapping interaction module, and a digital twin fusion module. The inland waterway-specific AI model module is used to realize real-time perception and accurate decision-making of waterway status. The high-noise voice interaction module is used to improve the robustness of voice recognition and the accuracy of semantic understanding in high-noise environments. The command-scene mapping interaction module is used to realize accurate mapping and low-latency interaction between voice commands and digital twin scene operations. The digital twin fusion module is used to construct a high-precision three-dimensional model of the waterway and realize the fusion of meteorological simulation and real-time monitoring images.
[0007] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the inland waterway-specific AI model module includes a scene feature extraction unit, a multimodal fusion architecture unit, a knowledge graph constraint unit, a model lightweighting unit, and an edge deployment unit. The scene feature extraction unit is used to collect physical parameters and shipping rules of inland waterways and construct a feature library. The multimodal fusion architecture unit adds a physical feature encoding layer on top of CNN and fuses physical parameter vectors, image features, and ship dynamic features through an attention mechanism. The knowledge graph constraint unit constructs an inland waterway shipping knowledge graph and sets a knowledge verification module in the model output layer. The model lightweighting unit optimizes the model using L1 regularization pruning and quantization compression. The edge deployment unit deploys the compressed model to waterway edge nodes.
[0008] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the model lightweighting unit converts 32-bit floating-point parameters into 16-bit or 8-bit integers through quantization compression, and implements quantization inference in the TensorRT framework; the edge nodes deployed by the edge deployment unit are NVIDIA Jetson series chips for shore monitoring terminals.
[0009] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the high-noise speech interaction module includes a high-noise speech enhancement unit and a semantic precision understanding unit. The high-noise speech enhancement unit uses microphone array beamforming, deep learning noise reduction, and endpoint detection optimization techniques to process speech signals. The semantic precision understanding unit improves semantic understanding accuracy through domain word vector training, BERT domain fine-tuning, and colloquial parsing rules.
[0010] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the microphone array in the strong noise speech enhancement unit adopts a 4-microphone linear array, and the noise direction is calculated by the MVDR algorithm; the deep learning noise reduction adopts the LSTM noise reduction model, inputs the Mel spectrum of noisy speech, outputs the clean speech spectrum, and then reconstructs the speech waveform by the Griffin-Lim algorithm; the endpoint detection optimization accurately locates speech segments in strong noise by combining energy threshold and spectral entropy analysis.
[0011] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the semantic precision understanding unit uses inland waterway shipping regulations and crew dialogue records as training data for domain word vector training, and trains domain word vectors through Word2Vec; BERT domain fine-tuning uses inland waterway instructions to fine-tune the pre-trained BERT model; and colloquial parsing rules match colloquial expressions through a regular expression library.
[0012] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the instruction-scene mapping interaction module includes a semantic parameter extraction unit, an action mapping table construction unit, an ambiguity resolution unit, an edge interaction engine deployment unit, a rendering resource preloading unit, and an asynchronous processing link unit. The semantic parameter extraction unit extracts key parameters from voice instructions using a named entity recognition model. The action mapping table construction unit establishes instruction-API mapping relationships. The ambiguity resolution unit completes parameters based on context or confirms them through voice questioning. The edge interaction engine deployment unit deploys a lightweight interaction engine locally on the digital twin server. The rendering resource preloading unit performs rendering caching on frequently accessed virtual scenes and pre-generates texture data from a 360° perspective. The asynchronous processing link unit uses multi-threaded parallel processing for voice parsing by the semantic parameter extraction unit, parameter mapping by the action mapping table construction unit, and scene rendering by the action mapping table construction unit.
[0013] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the digital twin fusion module includes a high-precision 3D modeling unit, a meteorological simulation fusion unit, and a monitoring screen fusion unit. The high-precision 3D modeling unit collects underwater data using multibeam sonar and stitches point clouds together using the ICPT algorithm. It then combines this data with waterway maps and laser scanning data to construct a 3D model of waterway facilities and build a parametric model library for ships. The meteorological simulation fusion unit connects to the meteorological bureau's API to obtain meteorological data and maps it into dynamic effects in the digital twin scenario, simulating the impact of meteorology on waterways based on a fluid dynamics model. The monitoring screen fusion unit uses the SIFT feature matching algorithm to calibrate the coordinates of the monitoring camera images and the virtual scene, achieving a superimposed display of virtual and real images.
[0014] In the aforementioned AI application system for smart ports and waterways empowered by digital twins, the high-precision 3D modeling unit uses inverse distance weighted interpolation to supplement the point cloud density in the shoal area; the 3D model of waterway facilities is constructed using Blender software, which assigns material properties to the model; the ship parametric model library is constructed using the size parameters of typical inland waterway vessels, and supports the automatic generation of corresponding virtual vessels based on the length and draft of the vessel.
[0015] In the aforementioned AI application system for smart ports and shipping empowered by digital twins, the monitoring screen fusion unit uses a virtual-real overlay display to enable clicking on the monitoring icon in the virtual scene to bring up the real-time image of the corresponding camera, while simultaneously marking the location of the ship identified in the monitoring on the corresponding coordinates in the virtual scene.
[0016] Compared with existing technologies, the advantages of this invention are: 1. This system achieves real-time perception and accurate decision-making of waterway status, effective voice interaction in noisy environments, accurate mapping of voice commands to virtual scenes, and fusion of high-precision waterway models with meteorology and monitoring through the coordinated operation of an inland waterway-specific AI model module, a high-noise voice interaction module, a command-scene mapping interaction module, and a digital twin fusion module. This design effectively solves the problems of low efficiency, inconvenient interaction, and poor scene reproduction in traditional inland waterway shipping management, thereby improving inland waterway shipping efficiency, ensuring navigation safety, and reducing operating costs.
[0017] 2. The inland waterway-specific AI model module constructs a feature library through a scene feature extraction unit, integrates multi-source features through a multi-modal fusion architecture unit, ensures compliant decision-making through a knowledge graph constraint unit, optimizes model performance through a lightweight model unit, and achieves low-latency inference through an edge deployment unit. This structure solves the problems of traditional AI models' poor adaptability to inland waterway scenarios and the difficulty in balancing real-time inference with accuracy. It enables the model to accurately perceive the waterway situation and provide compliant decision-making suggestions, improving the efficiency of waterway anomaly identification and the accuracy of scheduling decisions.
[0018] 3. In the noisy voice interaction module, the noisy voice enhancement unit uses a variety of technologies to process the voice signal, and the semantic accurate understanding unit improves the understanding accuracy through domain adaptation technology. This structure solves the problems of poor robustness of traditional speech recognition in noisy environments and insufficient understanding of professional terms and colloquial instructions, realizes effective voice communication in noisy environments, and improves the accuracy of speech recognition and semantic understanding.
[0019] 4. The instruction-scene mapping interaction module extracts key information through the semantic parameter extraction unit, establishes the mapping between instructions and APIs through the action mapping table construction unit, eliminates instruction ambiguity through the ambiguity resolution unit, reduces response latency through the edge interaction engine deployment unit, speeds up scene loading through the rendering resource preloading unit, and processes tasks in parallel through the asynchronous processing link unit. This structure solves the problems of inaccurate mapping between traditional voice instructions and virtual scenes and high interaction latency, and realizes accurate mapping and low-latency interaction between voice instructions and digital twin scene operations, thereby improving interaction efficiency. 5. The digital twin fusion module constructs waterway, facility, and ship models through high-precision 3D modeling units, maps the impact of weather on waterways through meteorological simulation fusion units, and realizes virtual-real overlay display through monitoring screen fusion units. This structure solves the problems of low accuracy in traditional waterway modeling, difficulty in simulating meteorological impact, and poor integration of monitoring and virtual scenes. It constructs a high-fidelity digital twin scene, realizes real-time linkage between the physical world and the virtual scene, and improves the overall control capability.
[0020] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the present invention; Figure 2 This is a detailed flowchart of the instruction-scene mapping interaction module; Figure 3 This is a detailed flowchart illustrating the constraint units of a knowledge graph. Figure 4 This is a detailed flowchart of the edge deployment unit.
[0022] The diagram shows the following modules: Inland Waterway Dedicated AI Model Module 1, High-Noise Voice Interaction Module 2, Command-Scene Mapping Interaction Module 3, Digital Twin Fusion Module 4, Scene Feature Extraction Unit 11, Multimodal Fusion Architecture Unit 12, Knowledge Graph Constraint Unit 13, Model Lightweighting Unit 14, Edge Deployment Unit 15, High-Noise Voice Enhancement Unit 21, Semantic Precision Understanding Unit 22, Semantic Parameter Extraction Unit 31, Action Mapping Table Construction Unit 32, Ambiguity Resolution Unit 33, Edge Interaction Engine Deployment Unit 34, Rendering Resource Preloading Unit 35, Asynchronous Processing Link Unit 36, High-Precision 3D Modeling Unit 41, Meteorological Simulation Fusion Unit 42, Monitoring Screen Fusion Unit 43, Knowledge Verification Module 131, Riverside Monitoring Terminal 151, and NVIDIA Jetson Series Chip 152. Detailed Implementation
[0023] like Figures 1-4 As shown, a smart port and shipping AI application system empowered by digital twins includes an inland waterway-specific AI model module 1, a high-noise voice interaction module 2, a command-scene mapping interaction module 3, and a digital twin fusion module 4. The inland waterway-specific AI model module 1 is used to realize real-time perception and accurate decision-making of waterway status. The high-noise voice interaction module 2 is used to improve the robustness of voice recognition and the accuracy of semantic understanding in high-noise environments. The command-scene mapping interaction module 3 is used to realize accurate mapping and low-latency interaction between voice commands and digital twin scene operations. The digital twin fusion module 4 is used to construct a high-precision three-dimensional model of the waterway and realize the fusion of meteorological simulation and real-time monitoring images.
[0024] This system, through the coordinated operation of an inland waterway-specific AI model module, a high-noise voice interaction module, a command-scene mapping interaction module, and a digital twin fusion module, achieves real-time perception and accurate decision-making of waterway status, effective voice interaction in high-noise environments, accurate mapping of voice commands to virtual scenes, and fusion of high-precision waterway models with meteorological and monitoring systems. This design effectively solves the problems of low efficiency, inconvenient interaction, and poor scene reproduction in traditional inland waterway shipping management, thereby improving inland waterway shipping efficiency, ensuring navigation safety, and reducing operating costs.
[0025] Specifically, the inland waterway-specific AI model module 1 includes a scene feature extraction unit 11, a multimodal fusion architecture unit 12, a knowledge graph constraint unit 13, a model lightweighting unit 14, and an edge deployment unit 15. The scene feature extraction unit 11 is used to collect physical parameters and shipping rules of inland waterways and build a feature library. The multimodal fusion architecture unit 12 adds a physical feature encoding layer on the basis of CNN and fuses physical parameter vectors, image features, and ship dynamic features through an attention mechanism. The knowledge graph constraint unit 13 constructs an inland waterway shipping knowledge graph and sets a knowledge verification module 131 in the model output layer. The model lightweighting unit 14 optimizes the model using L1 regularization pruning and quantization compression. The edge deployment unit 15 deploys the compressed model to the edge nodes of the waterway.
[0026] The inland waterway-specific AI model module constructs a feature library through a scene feature extraction unit, integrates features from multiple sources through a multimodal fusion architecture unit, ensures compliant decision-making through a knowledge graph constraint unit, optimizes model performance through a lightweight model unit, and achieves low-latency inference through an edge deployment unit. This structure solves the problems of traditional AI models' poor adaptability to inland waterway scenarios and the difficulty in balancing real-time inference with accuracy. It enables the model to accurately perceive the waterway situation and provide compliant decision-making suggestions, improving the efficiency of waterway anomaly identification and the accuracy of scheduling decisions.
[0027] Specifically, the scene feature extraction unit 11 collects physical parameters of the inland waterway, including but not limited to water flow speed, water flow direction, waterway width, and bridge clearance height. The shipping rules refer to the "Inland Waterway Collision Avoidance Rules". The water flow speed parameter is collected by sensors. Sensors are set up along the river section to collect water flow speed and direction data of different river sections. The image features in the multimodal fusion architecture unit 12 are the monitoring screen, and the ship dynamic features are the ship's GPS trajectory. Those skilled in the art should understand that adding a physical feature encoding layer on top of CNN is a method to integrate prior physical knowledge into a deep learning model. It aims to improve the model's performance, interpretability, and generalization ability by utilizing physical laws. It is commonly used to handle physics-related tasks. After adding a physical feature encoding layer on top of CNN, the physical parameter vector, image features, and ship dynamic features can be fused through an attention mechanism to enable the model to automatically focus on ship avoidance priorities in scenarios such as "narrow channel + against the current". The inland waterway shipping knowledge graph constructed by the knowledge graph constraint unit 13 includes entities such as "restricted navigation area", "bridge area" and "vehicle type", and relationships such as "restricted navigation period" and "avoidance relationship". The knowledge verification module 131 can filter scheduling suggestions that violate the rules, such as prohibiting the recommendation of passing instructions in the bridge area.
[0028] Specifically, those skilled in the art should understand that L1 regularization pruning is a model lightweighting method that combines L1 regularization technology with model pruning strategies. It aims to simplify the model structure by removing redundant parameters or neurons while maintaining the core performance of the model. The model lightweighting unit 14 uses L1 regularization to sparsify the model weights, removes connections below the threshold, and reduces the number of parameters. It can filter and remove redundant or unimportant parameters and neurons. Quantization compression performs low-bit-width representation transformation on the model parameters to achieve model structure simplification and optimization.
[0029] Furthermore, in the model lightweighting unit 14, quantization compression converts the 32-bit floating-point parameters into 16-bit or 8-bit integers, enabling quantized inference within the TensorRT framework. The edge deployment unit 15 deploys edge nodes on the NVIDIA Jetson series chips 152 of the shore-based monitoring terminal 151. The model lightweighting unit reduces the number of parameter bits through quantization compression, enabling efficient inference within the TensorRT framework. The edge deployment unit deploys the model on the NVIDIA Jetson series chips of the shore-based monitoring terminal. This design reduces model computation and size, improves inference speed, avoids the high latency of cloud deployment, and keeps model inference latency at a low level, meeting the millisecond-level decision-making requirements for dynamic ship adjustments.
[0030] Specifically, those skilled in the art should understand that TensorRT is a high-performance deep learning inference optimization framework launched by NVIDIA, which focuses on optimizing, deploying and accelerating trained deep learning models. It is widely used in inference tasks in fields such as computer vision, natural language processing, and speech recognition. Experimental results show that quantization compression can reduce the model size by 75% and increase the inference speed by 3 times after converting 32-bit floating-point parameters to 8 bits, while keeping the accuracy loss within 2%. NVIDIA Jetson series chips are embedded processors designed by NVIDIA for edge computing and artificial intelligence applications, and are widely used in robotics, drones, intelligent monitoring and other fields. After the edge deployment unit 15 deploys the compressed model to the channel edge node of the shore monitoring terminal 151 (which is the NVIDIA Jetson series chip 152), the ship anomaly identification model can be optimized from a latency of 100ms+ transmitted from the cloud to a response time of less than 50ms.
[0031] Specifically, the noisy speech interaction module 2 includes a noisy speech enhancement unit 21 and a semantically accurate understanding unit 22. The noisy speech enhancement unit 21 uses microphone array beamforming, deep learning noise reduction, and endpoint detection optimization techniques to process the speech signal. The semantically accurate understanding unit 22 improves semantic understanding accuracy through domain word vector training, BERT domain fine-tuning, and colloquial parsing rules. In the noisy speech interaction module, the noisy speech enhancement unit uses multiple techniques to process the speech signal, and the semantically accurate understanding unit improves understanding accuracy through domain adaptation technology. This structure solves the problems of poor robustness of traditional speech recognition in noisy environments and insufficient understanding of professional terms and colloquial instructions, realizing effective speech communication in noisy environments and improving speech recognition accuracy and semantic understanding accuracy.
[0032] Furthermore, in the strong noise speech enhancement unit 21, the microphone array adopts a 4-microphone linear array, and the noise direction is calculated through the MVDR algorithm; the deep learning denoising adopts the LSTM denoising model, inputs the Mel spectrum of the noisy speech, outputs the clean speech spectrum, and then reconstructs the speech waveform through the Griffin-Lim algorithm; the endpoint detection optimization accurately locates speech segments in strong noise by combining energy threshold and spectral entropy analysis. After the strong noise speech enhancement unit uses a 4-microphone linear array combined with the MVDR algorithm to calculate the noise direction, it can form a directional beam focusing speech signal acquisition direction within ±30° to suppress engine roar, which usually comes from the stern of the ship and water flow noise; The LSTM noise reduction model constructs a training pair by collecting noisy speech from the bridge of an inland waterway vessel and recording clean speech in a recording studio. It takes the Mel spectrum of the noisy speech as input and outputs the spectrum of the clean speech. Combined with the Griffin-Lim algorithm, the speech waveform is reconstructed. This design utilizes the ability of LSTM to capture temporal features, effectively separates speech and noise components, significantly improves speech clarity, and can improve the signal-to-noise ratio by more than 15dB. It solves the problem of noise in the ship's bridge interfering with the speech signal and provides high-quality speech input for scenarios such as voice communication and command recognition. Endpoint detection is performed by combining energy thresholding and spectral entropy analysis. Energy thresholding initially filters speech segments by differentiating the energy of speech from noise; spectral entropy analysis utilizes the differences in spectral distribution characteristics between speech and pure noise, such as engine noise. Speech spectral entropy is lower and fluctuates more, while noise spectral entropy is higher and more stable, thus accurately identifying valid speech. This design can accurately locate speech segments in strong noise, effectively eliminating invalid periods such as pure engine noise, reducing invalid recognition, and improving the efficiency and accuracy of subsequent speech processing such as noise reduction and recognition.
[0033] Those skilled in the art should understand that the MVDRMinimum Variance Distortionless Response algorithm is a classic adaptive beamforming technique widely used in the field of array signal processing. Its core objective is to suppress interference and noise while preserving the target signal without distortion, thereby improving the signal-to-noise ratio of the array output. LSTM (Long Short-Term Memory) network denoising model is a speech enhancement technique based on temporal deep learning, specifically designed to separate and restore clean speech from noisy speech, especially suitable for processing speech signals with temporal dependencies. Its core advantage lies in leveraging LSTM's ability to capture long-term dependencies, effectively distinguishing the temporal features of speech from noise, and achieving high-quality denoising. The Griffin-Lim algorithm is a classic speech signal reconstruction technique, primarily used to recover complete speech waveforms from amplitude spectra lacking phase information. In speech processing tasks such as denoising and speech synthesis, models often only predict the amplitude spectrum of the signal, while phase information is difficult to estimate accurately. The Griffin-Lim algorithm, through an iterative optimization strategy, can reconstruct high-quality speech waveforms from the amplitude spectrum, solving the waveform reconstruction problem caused by phase loss.
[0034] Specifically, in the semantic precision understanding unit 22, domain word vector training uses inland waterway regulations and crew dialogue records as training data, and trains domain word vectors through Word2Vec; BERT domain fine-tuning uses inland waterway instructions to fine-tune the pre-trained BERT model; and colloquial parsing rules match colloquial expressions through a regular expression library. The semantic precision understanding unit trains domain word vectors through Word2Vec, performs BERT domain fine-tuning to adapt to professional instructions, and uses a regular expression library to parse colloquial expressions. This design solves the problem of low parsing accuracy of traditional semantic understanding for riverway professional terminology and colloquial instructions, significantly improving the model's accuracy in recognizing professional instructions and colloquial expressions, covering most daily instruction scenarios for crew members. Furthermore, crew dialogue records include, but are not limited to, VHF call recordings transcribed into text; BERT domain fine-tuning: Based on the pre-trained BERT model, fine-tuning was performed using 20,000 inland waterway commands such as "Query XX shoal water level" and "Is the channel ahead congested?", adjusting the weights of the fully connected layers to adapt to professional terminology. After fine-tuning, the domain command recognition accuracy improved from 72% to 91%; The colloquial parsing rules use a regular expression library to match colloquial expressions, including but not limited to mapping "Is it congested ahead?" to "Query the congestion status of the channel ahead", and mapping "How long do I have to wait to pass through the lock?" to "Query the lock waiting time", covering more than 80% of the colloquial instructions of crew members.
[0035] Those skilled in the art should understand that Word2Vec training is the process of converting text corpora into low-dimensional word vectors. Its core is learning the co-occurrence patterns of words in context through neural networks, ultimately generating vector representations that reflect semantic relationships. The training process balances efficiency and semantic capture capabilities, making it a crucial step in the practical application of word embedding technology. BERT domain fine-tuning refers to retraining a pre-trained BERT model on domain-specific corpora to adapt it to the language features and task requirements of that domain, thereby improving the model's performance in downstream tasks within that domain. While pre-trained BERT models possess general language understanding capabilities, their performance is limited in specialized domains such as law, medicine, and shipping, particularly in terms of terminology, sentence structure, and semantic logic. Domain fine-tuning compensates for this gap by injecting domain knowledge.
[0036] Specifically, the instruction-scene mapping interaction module 3 includes a semantic parameter extraction unit 31, an action mapping table construction unit 32, an ambiguity resolution unit 33, an edge interaction engine deployment unit 34, a rendering resource preloading unit 35, and an asynchronous processing link unit 36. The semantic parameter extraction unit 31 extracts key parameters from voice instructions using a named entity recognition model. The action mapping table construction unit 32 establishes an instruction-API mapping relationship. The ambiguity resolution unit 33 completes parameters based on context or confirms them through voice questioning. The edge interaction engine deployment unit 34 deploys a lightweight interaction engine locally on the digital twin server. The rendering resource preloading unit 35 performs rendering caching on frequently accessed virtual scenes and pre-generates texture data from a 360° perspective. The asynchronous processing link unit 36 uses multi-threaded parallel processing to handle the voice parsing of the semantic parameter extraction unit 31, the parameter mapping of the action mapping table construction unit 32, and the scene rendering of the action mapping table construction unit 32. The instruction-scene mapping interaction module extracts key information through the semantic parameter extraction unit, establishes the mapping between instructions and APIs through the action mapping table construction unit, eliminates instruction ambiguity through the ambiguity resolution unit, reduces response latency through the edge interaction engine deployment unit, speeds up scene loading through the rendering resource preloading unit, and processes tasks in parallel through the asynchronous processing link unit. This structure solves the problems of inaccurate mapping between traditional voice instructions and virtual scenes and high interaction latency, and realizes accurate mapping and low-latency interaction between voice instructions and digital twin scene operations, thereby improving interaction efficiency. Furthermore, the semantic parameter extraction unit 31 extracts key parameters from the voice command through the named entity recognition model, including but not limited to extracting the entity "virtual ship", the action "turn left", and the parameter "30 degrees" in "control the virtual ship to turn left 30 degrees", corresponding to the ship ID and rotation angle in the digital twin scenario; Establish command-API mapping relationships, such as "view facilities in XX area" corresponding to calling the digital twin platform interface, predefine mapping rules for 100+ high-frequency commands, and the parsing time is <20ms; Ambiguity resolution unit 33 completes parameters by combining context or confirms them by asking a question in voice. For example, if "view upstream" does not specify a river segment, it will automatically complete the parameters by combining context, such as the crew's current location, or confirm them by asking a question in voice, such as "Which upstream channel are you looking for?". The edge interaction engine deployment unit 34 deploys a lightweight interaction engine locally on the digital twin server, such as a real-time communication module developed based on C++, to reduce cross-network call latency. Actual tests show that edge deployment can reduce command response time from 200ms in the cloud to less than 50ms. The rendering resource preloading unit 35 performs rendering caching on frequently accessed virtual scenes and pre-generates texture data for a 360° view. This data can be directly loaded when called, avoiding real-time rendering time consumption. Actual tests show that the latency of switching views is reduced from 80ms to 15ms after preloading. The asynchronous processing link unit 36 can process the scene rendering module in parallel while parsing instructions, prepare the basic scene model, and keep the total latency within 80ms.
[0037] Specifically, the digital twin fusion module 4 includes a high-precision 3D modeling unit 41, a meteorological simulation fusion unit 42, and a monitoring screen fusion unit 43. The high-precision 3D modeling unit 41 collects underwater data using multi-beam sonar and stitches point clouds together using the ICPT algorithm. It then combines this data with waterway maps and laser scanning data to construct a 3D model of waterway facilities and build a parametric model library for ships. The meteorological simulation fusion unit 42 connects to the meteorological bureau's API to obtain meteorological data and maps it as dynamic effects in the digital twin scene, simulating the impact of weather on the waterway based on a fluid dynamics model. The monitoring screen fusion unit 43 uses the SIFT feature matching algorithm to calibrate the coordinates of the monitoring camera images and the virtual scene, achieving a superimposed display of the virtual and real worlds. The digital twin fusion module constructs waterway, facility, and ship models through the high-precision 3D modeling unit, maps the impact of weather on the waterway through the meteorological simulation fusion unit, and achieves a superimposed display of the virtual and real worlds through the monitoring screen fusion unit. This structure solves the problems of low accuracy in traditional waterway modeling, difficulty in simulating meteorological impacts, and poor integration of monitoring and virtual scenes. It constructs a high-fidelity digital twin scene, realizes real-time linkage between the physical world and the virtual scene, and improves overall control capabilities. Specifically, underwater data is collected using multibeam sonar, and point clouds measured by multiple ships are stitched together using the ICPT iterative nearest point algorithm, which can eliminate stitching errors and control them within 0.5 meters. Waterway drawings include, but are not limited to, CAD drawings; Data obtained from the meteorological bureau's API includes, but is not limited to, wind speed, precipitation, and visibility data. In the digital twin scenario, this data is mapped to dynamic effects—for example, when the wind speed is greater than 10 m / s, a wave animation with the corresponding wave height is generated on the virtual water surface; when the visibility is less than 1 km, a fogging effect is added to the scene. The simulation of the impact of meteorology on waterways is based on fluid dynamics models, including but not limited to simulating the impact of water level rise caused by heavy rain on waterway depth, such as a water level rise of 0.5 meters after 24 hours of heavy rain, and automatically updating the safe navigation draft value of shallow shoals. The SIFT feature matching algorithm can be used to geometrically calibrate the surveillance camera footage with the corresponding area in the virtual scene, such as the virtual shoreline corresponding to shoreline monitoring, with the error controlled within 10 pixels.
[0038] Those skilled in the art should understand that ICP-T Iterative Closest Point is a classic iterative optimization algorithm for point cloud registration. Its core objective is to find the optimal transformation, translation, rotation, or scaling between two point sets so that they coincide to the greatest extent in space. It is widely used in fields such as 3D reconstruction, object recognition, and robot navigation. SIFT (Scale-Invariant Feature Transform) is an image feature extraction and matching algorithm with scale and rotation invariance. It can stably extract key features from images and achieve accurate matching under different scales, rotations, lighting changes, and even partial occlusion. It is widely used in image stitching, target recognition, stereo vision and other fields.
[0039] Furthermore, in the high-precision 3D modeling unit 41, inverse distance weighted interpolation is used to supplement the point cloud density in the shoal area; the 3D model of waterway facilities is constructed using Blender software, which assigns material properties to the model; the parametric ship model library is constructed using the size parameters of typical inland waterway vessels, and supports the automatic generation of corresponding virtual vessels based on ship length and draft. For detailed areas such as shoals, inverse distance weighted IDW interpolation is used to supplement the point cloud density, which can improve the modeling accuracy to the 0.1-meter level; the 3D model of waterway facilities includes, but is not limited to, bridge cores, navigation mark cores, etc., which are assigned material properties using Blender software, including but not limited to steel reflectivity, concrete texture, etc., which can achieve seamless splicing of cores and terrain models; the parametric ship model library collects the size parameters of 100+ typical inland waterway vessels, including cargo ships, passenger ships, and tugboats.
[0040] Specifically, in the monitoring image fusion unit 43, the virtual and real superimposed display enables clicking the monitoring icon in the virtual scene to pop up the real-time image of the corresponding camera, while marking the position of the ship identified in the monitoring on the corresponding coordinates in the virtual scene.
[0041] The monitoring screen fusion unit calibrates coordinates using the SIFT feature matching algorithm, enabling the function of clicking on the virtual monitoring icon to pop up real-time images and ship position markings. This design solves the problem of loose integration between the monitoring screen and the virtual scene, ensuring small coordinate calibration errors between the virtual and real scenes, and achieving seamless linkage display of "virtual scene + real entity". This shortens the time for managers to grasp the overall waterway situation and improves regulatory efficiency.
[0042] Specifically, in the digital twin platform, clicking the monitoring icon of the virtual scene will automatically bring up the real-time image of the corresponding camera; at the same time, the position of the ship identified in the monitoring will be marked on the corresponding coordinates of the virtual scene through target detection, so as to realize the linkage display of "virtual scene + real-time entity".
[0043] Through a dedicated AI model module for inland waterways, the accuracy rate for identifying abnormal events such as ship collisions and groundings reaches 98.5%, improving efficiency by 50 times compared to traditional video surveillance and manual inspections, while reducing the false negative rate to below 0.3%. The scheduling suggestions, combined with knowledge graph constraints, achieve 100% compliance. In narrow channel encounter scenarios, the recommended avoidance routes reduce the risk of ship collisions by 90% and shorten the time for a single encounter by 15 minutes. After optimization, the model's inference latency on edge devices is controlled within 50ms, a 60% speedup compared to cloud-based deployment solutions. Through a high-noise voice interaction module, microphone array + LSTM noise reduction technology achieves a 95% accuracy rate for voice recognition in high-noise environments in the cockpit, a 40 percentage point improvement over traditional single-microphone solutions. The domain-adjusted BERT model achieves a 96% accuracy rate in parsing professional instructions and a 92% accuracy rate in recognizing colloquial expressions, covering 98% of the daily instruction scenarios for crew members. The response speed of voice-controlled virtual scenes is 3 times faster than mouse operation, reducing the average time for crew members to query waterway information from 20 seconds to 2 seconds. Through the digital twin fusion module, the 3D reconstruction error of the waterway topography is ≤0.1 meters, and the model details of bridges, navigation marks and other facilities are 99% consistent with the actual objects; the digital twin scene can map the impact of wind speed and precipitation on the waterway in real time, providing ships with a prediction of the feasibility of waterway passage in the next 12 hours, and improving the navigation safety factor by 60% under extreme weather conditions; the coordinate calibration error between the monitoring screen and the twin scene is ≤10 pixels, and the time for managers to grasp the status of the entire waterway is shortened from 30 minutes to 1 minute.
[0044] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
[0045] Although this paper extensively uses terms such as Inland Waterway Dedicated AI Model Module 1, Noisy Voice Interaction Module 2, Command-Scene Mapping Interaction Module 3, Digital Twin Fusion Module 4, Scene Feature Extraction Unit 11, Multimodal Fusion Architecture Unit 12, Knowledge Graph Constraint Unit 13, Model Lightweighting Unit 14, Edge Deployment Unit 15, Noisy Voice Enhancement Unit 21, Semantic Precision Understanding Unit 22, Semantic Parameter Extraction Unit 31, Action Mapping Table Construction Unit 32, Ambiguity Resolution Unit 33, Edge Interaction Engine Deployment Unit 34, Rendering Resource Preloading Unit 35, Asynchronous Processing Link Unit 36, High-Precision 3D Modeling Unit 41, Meteorological Simulation Fusion Unit 42, Monitoring Screen Fusion Unit 43, Knowledge Verification Module 131, Shoreline Monitoring Terminal 151, and NVIDIA Jetson Series Chip 152, these terms are used merely to more conveniently describe and explain the essence of this invention; interpreting them as any additional limitation would contradict the spirit of this invention.
Claims
1. A smart port and shipping AI application system empowered by digital twins, characterized in that, It includes an inland waterway-specific AI model module (1), a high-noise voice interaction module (2), a command-scene mapping interaction module (3), and a digital twin fusion module (4). The inland waterway-specific AI model module (1) is used to realize real-time perception and accurate decision-making of waterway status. The high-noise voice interaction module (2) is used to improve the robustness of voice recognition and the accuracy of semantic understanding in a high-noise environment. The command-scene mapping interaction module (3) is used to realize accurate mapping and low-latency interaction between voice commands and digital twin scene operations. The digital twin fusion module (4) is used to construct a high-precision three-dimensional waterway model and realize the fusion of meteorological simulation and real-time monitoring images.
2. The smart port and shipping AI application system empowered by digital twins as described in claim 1, characterized in that, The inland waterway-specific AI model module (1) includes a scene feature extraction unit (11), a multimodal fusion architecture unit (12), a knowledge graph constraint unit (13), a model lightweighting unit (14), and an edge deployment unit (15). The scene feature extraction unit (11) is used to collect physical parameters and shipping rules of inland waterways and build a feature library. The multimodal fusion architecture unit (12) adds a physical feature encoding layer on the basis of CNN and fuses physical parameter vectors, image features, and ship dynamic features through an attention mechanism. The knowledge graph constraint unit (13) constructs an inland waterway shipping knowledge graph and sets a knowledge verification module (131) in the model output layer. The model lightweighting unit (14) optimizes the model using L1 regularization pruning and quantization compression. The edge deployment unit (15) deploys the compressed model to the edge nodes of the waterway.
3. The smart port and shipping AI application system empowered by digital twins according to claim 2, characterized in that, In the model lightweight unit (14), quantization compression converts 32-bit floating-point parameters into 16-bit or 8-bit integers, and quantization inference is implemented in the TensorRT framework; the edge nodes deployed by the edge deployment unit (15) are NVIDIA Jetson series chips (152) of the shore monitoring terminal (151).
4. The smart port and shipping AI application system empowered by digital twins according to claim 1, characterized in that, The strong noise speech interaction module (2) includes a strong noise speech enhancement unit (21) and a semantic accurate understanding unit (22). The strong noise speech enhancement unit (21) uses microphone array beamforming, deep learning noise reduction and endpoint detection optimization technology to process speech signals. The semantic accurate understanding unit (22) improves the semantic understanding accuracy through domain word vector training, BERT domain fine-tuning and colloquial parsing rules.
5. The smart port and shipping AI application system empowered by digital twins according to claim 4, characterized in that, In the strong noise speech enhancement unit (21), the microphone array adopts a 4-microphone linear array, and the noise direction is calculated by the MVDR algorithm; the deep learning noise reduction adopts the LSTM noise reduction model, inputs the Mel spectrum of the noisy speech, outputs the clean speech spectrum, and then reconstructs the speech waveform by the Griffin-Lim algorithm; the endpoint detection optimization accurately locates the speech segment in strong noise by combining energy threshold and spectral entropy analysis.
6. The smart port and shipping AI application system empowered by digital twins according to claim 4, characterized in that, In the semantic precision understanding unit (22), the domain word vector training uses inland waterway shipping regulations and crew dialogue records as training data, and trains domain word vectors through Word2Vec; the BERT domain fine-tuning uses inland waterway instructions to fine-tune the pre-trained BERT model; the colloquial parsing rules match colloquial expressions through a regular expression library.
7. The smart port and shipping AI application system empowered by digital twins according to claim 1, characterized in that, The instruction-scene mapping interaction module (3) includes a semantic parameter extraction unit (31), an action mapping table construction unit (32), an ambiguity resolution unit (33), an edge interaction engine deployment unit (34), a rendering resource preloading unit (35), and an asynchronous processing link unit (36). The semantic parameter extraction unit (31) extracts key parameters from voice instructions through a named entity recognition model. The action mapping table construction unit (32) establishes an instruction-API mapping relationship. The ambiguity resolution unit (33) completes parameters in conjunction with the context or confirms them through voice questioning. The edge interaction engine deployment unit (34) deploys a lightweight interaction engine locally on the digital twin server. The rendering resource preloading unit (35) performs rendering caching on frequently accessed virtual scenes and pre-generates texture data from a 360° perspective. The asynchronous processing link unit (36) uses multi-threaded parallel processing to perform voice parsing of the semantic parameter extraction unit (31), parameter mapping of the action mapping table construction unit (32), and scene rendering of the action mapping table construction unit (32).
8. The smart port and shipping AI application system empowered by digital twins according to claim 1, characterized in that, The digital twin fusion module (4) includes a high-precision 3D modeling unit (41), a meteorological simulation fusion unit (42), and a monitoring screen fusion unit (43). The high-precision 3D modeling unit (41) collects underwater data through multi-beam sonar and stitches point clouds using the ICPT algorithm. It then combines waterway maps and laser scanning data to construct a 3D model of waterway facilities and builds a parametric model library for ships. The meteorological simulation fusion unit (42) connects to the meteorological bureau's API to obtain meteorological data and maps it into dynamic effects in the digital twin scene. It simulates the impact of meteorology on waterways based on a fluid dynamics model. The monitoring screen fusion unit (43) uses the SIFT feature matching algorithm to calibrate the coordinates of the monitoring camera images and the virtual scene to achieve virtual-real overlay display.
9. The smart port and shipping AI application system empowered by digital twins according to claim 8, characterized in that, In the high-precision three-dimensional modeling unit (41), the point cloud density is supplemented by inverse distance weighted interpolation in the shoal area; the three-dimensional model of the waterway facility is constructed using Blender software, which assigns material properties to the model; the ship parametric model library is constructed using the size parameters of typical inland waterway vessels, and supports the automatic generation of corresponding virtual vessels based on the length and draft.
10. The smart port and shipping AI application system empowered by digital twins according to claim 8, characterized in that, In the monitoring screen fusion unit (43), the virtual and real superposition display realizes that clicking the monitoring icon of the virtual scene will pop up the real-time screen of the corresponding camera, and at the same time, the position of the ship identified in the monitoring is marked on the corresponding coordinates of the virtual scene.
Citation Information
Patent Citations
A method and apparatus for constructing a digital twin scenario for inland waterways
CN113223162B
Digital twin channel construction method and system
CN114529680A
A construction method for an interactive real - scene three - dimensional waterway
CN116416393B
Cited By
Ship draft intelligent measurement method and system based on marker template matching
CN122335948A