Multi-modal perception and multi-agent collaboration method and system for smart port
Patent Information
- Application Number
- CN202610713813.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-25
AI Technical Summary
这种架构存在明显的局限性:视频数据传输带宽需求大,中心处理压力高,实时性难以保障;单一视觉感知在恶劣天气条件下性能显著下降;缺乏与其他感知手段的深度融合机制
[0015]本发明与现有技术相比所具有的优点或效果,有益效果应注意以下几点:
Smart Images

Figure CN122820150A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a multimodal perception and multi-agent collaborative decision-making method and system for smart ports. Background Technology
[0002] 1. Current technological status of smart port systems: Currently, the main technological solutions adopted for smart port construction both domestically and internationally are as follows: The first type is a sensing and data acquisition system based on traditional Internet of Things (IoT). This system uses a distributed sensor network for data collection, including single-function sensors such as inductive loops, RFID readers, and GPS positioning devices. The system architecture is mostly based on a single-center data acquisition and centralized processing approach, lacking collaborative sensing capabilities between sensing nodes and failing to achieve real-time fusion of multi-source data. Data processing uses simple threshold judgments or preset rule engines, making it difficult to adapt to the complex and ever-changing port business scenarios.
[0003] The second type is an intelligent analysis system based on video surveillance. This system uses video surveillance as its core, combining it with computer vision technology to achieve target detection and recognition. The system typically employs a centralized video processing architecture, transmitting surveillance video back to a central server for AI analysis. This architecture has significant limitations: high bandwidth requirements for video data transmission, high central processing pressure, and difficulty in guaranteeing real-time performance; performance of single visual perception degrades significantly under adverse weather conditions; and it lacks a deep integration mechanism with other sensing methods.
[0004] The third type is a decision support system based on rule engines. Some advanced ports have introduced intelligent decision-making systems based on rule engines or simple machine learning models, with a single central system performing data analysis and decision-making. This approach has the following shortcomings: the centralized architecture struggles to handle the processing needs of large-scale real-time data, resulting in poor system scalability; single decision nodes are prone to single-point failures and insufficient reliability; there is a lack of effective collaboration mechanisms between various business systems, easily leading to information silos; and it cannot adapt to the complexity and dynamic changes of port business scenarios.
[0005] The fourth type is a simple linkage-based equipment collaboration system. Some ports have implemented simple linkages between systems such as customs inspection and quarantine, border inspection, and port scheduling, triggering equipment actions through preset rules. This solution has the following shortcomings: the rule engine cannot handle complex scenarios and boundary conditions, requiring a large amount of manual maintenance; it lacks an evaluation and optimization mechanism for decision results, making continuous improvement impossible; each subsystem operates independently, lacking a global perspective for coordinated optimization; and it cannot handle emergencies and unconventional situations, exhibiting poor flexibility.
[0006] 2. Existing technical deficiencies: A comprehensive analysis of existing smart port system technical solutions reveals the following common technical deficiencies: First, the system lacks the capability to fuse multi-source heterogeneous data. Existing systems typically can only process single-type or a few types of sensing data, making it difficult to effectively fuse multi-source heterogeneous data such as video, radar, sensors, and electronic tags. Different sensors exhibit significant differences in data format, time reference, and spatial coordinates. Current technologies lack a unified fusion framework and efficient fusion algorithms, resulting in blind spots and data gaps in port sensing, impacting the comprehensiveness and accuracy of decision-making.
[0007] Second, the coordination between real-time perception and intelligent decision-making is poor. Existing systems often have their perception and decision-making systems built independently, lacking a deep integration mechanism. Perceived data must undergo lengthy collection, transmission, storage, and processing processes before reaching the decision-making system, resulting in significant response delays from perception to decision-making, making it difficult to meet the high-throughput and high-real-time requirements of ports. Simultaneously, the decision-making system struggles to dynamically adjust its decision-making strategies based on real-time perception data.
[0008] Third, the collaborative decision-making capabilities of multiple business systems are limited. Smart ports involve multiple business systems such as customs, inspection and quarantine, border control, port affairs, and logistics, each with its own independent business logic and decision-making mechanisms. Existing technical solutions struggle to achieve intelligent collaborative decision-making across systems, frequently resulting in business conflicts and decision-making contradictions. The lack of effective information sharing and coordination mechanisms between systems prevents the formation of a unified decision-making view.
[0009] Fourth, the system lacks sufficient adaptability and fault tolerance. Existing systems typically employ static rules or fixed models, making them unable to adaptively adjust to changes in port operations and the external environment. When business rules change, manual reconfiguration is required, resulting in high maintenance costs and a high risk of errors. Furthermore, the system lacks robust fault tolerance and degradation mechanisms, meaning a single point of failure could paralyze the entire system.
[0010] Fifth, the decision-making process suffers from poor interpretability and traceability. Most existing intelligent decision-making systems operate as "black box" models, making the decision-making process difficult to explain and the causes of problems hard to trace. Decision-makers cannot understand why the intelligent system made a certain decision, making it difficult to review and intervene in the decisions. This hinders regulatory auditing and continuous optimization, and also affects users' trust in the intelligent system. Summary of the Invention
[0011] In view of the above-mentioned technical defects and deficiencies in the existing technology, the technical problem to be solved by the present invention is: how to provide a smart port system method and architecture that can integrate multimodal perception information and support multi-agent collaborative decision-making, so as to realize comprehensive perception, intelligent understanding and collaborative decision-making of the complex port environment and business, and improve port operation efficiency and decision quality.
[0012] The multimodal perception and multi-agent collaboration method for smart ports includes the following steps: S1. Multimodal sensing data acquisition: Collect heterogeneous data of the port environment through video surveillance, lidar, millimeter-wave radar, RFID readers, electronic weighbridges and environmental monitoring equipment, and perform spatiotemporal synchronization and data cleaning; S2. Multi-source data fusion and feature extraction: Using a feature fusion network with a cross-modal attention mechanism, the multi-source data processed in step S1 is aligned and fused to generate a unified target representation; S3. Digital Twin Status Update: Based on the fused data output in step S2, update the location and status attributes of vehicles, goods, and equipment in the port digital twin model, and perform anomaly detection and simulation. S4. Distributed Decision Making: Based on twin state attributes, local decision-making schemes are generated in parallel for multiple business agents in security monitoring, scheduling optimization, emergency response, and customer service; each agent adopts a policy network based on multi-agent reinforcement learning. S5. Collaborative Optimization: Conflict detection is performed on the local decision schemes generated in step S4; a global decision scheme is generated based on a multi-criteria decision-making method. S6. Decision Execution and Feedback: Distribute the global decision scheme described in S5 to the execution device and collect execution feedback data to update the digital twin model.
[0013] It also includes step S7: performing causal reasoning analysis on the global decision-making scheme based on the causal graph model, and recording the decision chain for post-event tracing.
[0014] The Smart Port's multimodal perception and multi-agent collaborative system includes: a multimodal perception layer, a data fusion layer, a digital twin layer, a multi-agent decision-making layer, a collaborative control layer, and an interactive presentation layer; The multimodal perception layer includes a video surveillance module, a lidar module, a millimeter-wave radar module, an RFID reader / writer module, an electronic weighbridge module, and an environmental monitoring module. The data fusion layer, connected to the multimodal perception layer, includes a spatiotemporal synchronization module, a target association module, and a feature fusion module based on a cross-modal attention mechanism; The digital twin layer, connected to the data fusion layer, is used to construct and update the port's 3D model, including a real-time mapping module and a simulation and deduction module; A multi-agent decision-making layer, connected to the digital twin layer, includes multiple business agents based on reinforcement learning; The collaborative control layer, connected to the multi-agent decision-making layer, includes a conflict arbitration module and a consensus formation module; An interactive presentation layer, connected to the collaborative control layer, is used to visualize the global decision-making scheme.
[0015] The advantages or effects of this invention compared to the prior art, and the beneficial effects, should be noted in the following points: 1. Enhanced Perception Capabilities Through Multimodal Perception Fusion. This invention utilizes multimodal perception fusion technology to simultaneously acquire multi-dimensional information about the port environment, including visual, location, identity, and weight information, thus solving the problem of insufficient information from single perception methods. Different modalities complement each other, providing supplementary information when the performance of a single modality declines. Even under adverse weather conditions such as heavy fog, rain, and snow, the system maintains a high level of perception accuracy.
[0016] 2. Improved System Reliability Through Multi-Agent Architecture. This invention achieves decentralized and redundant decision-making functions through a layered distributed multi-agent architecture. The failure of a single agent will not paralyze the entire system; other agents can take over its functions. The system employs a multi-replica backup mechanism, with key agents running replicas on different physical nodes. Automatic failover is achieved through heartbeat detection, effectively ensuring the continuity of port operations.
[0017] 3. Improved operational efficiency through collaborative decision-making. This invention achieves coordinated decision-making across business systems through a multi-agent collaborative optimization mechanism, reducing business conflicts and redundant operations. Each agent gains a global perspective through information sharing, enabling more rational local decisions. The collaborative control layer resolves decision-making contradictions through conflict arbitration and consensus-building mechanisms, generating the globally optimal solution.
[0018] 4. Continuously Improved Decision Quality Through Adaptive Learning. This invention utilizes a multi-agent reinforcement learning mechanism, enabling the system to continuously learn and optimize from historical decisions. The agents continuously improve their decision-making strategies through interaction with the environment, adapting to changes in business operations and the environment. The system employs a combination of offline and online learning mechanisms. Offline learning involves large-scale training in a simulation environment, while online learning fine-tunes strategies based on real-world feedback.
[0019] 5. Enhanced Compliance and Trust Through Explainable Decision-Making. This invention utilizes a causal reasoning explanation mechanism, enabling the traceability of the basis and process of each decision. Decision-makers can understand why the intelligent system made a certain decision and can review and intervene in those decisions. When a decision deviates from its intended course, the cause can be quickly identified and adjustments made. Stakeholder trust and acceptance of the intelligent decision-making system are significantly increased, creating conditions for its widespread application. Attached Figure Description
[0020] Figure 1 A schematic diagram of the framework for a multimodal perception and multi-agent collaborative decision-making system for smart ports; Figure 2 This is a schematic diagram of multimodal perception and multi-agent collaborative decision-making technology for smart ports. Detailed Implementation
[0021] The present invention is further described below in conjunction with embodiments and accompanying drawings: The multimodal perception and multi-agent collaboration method for smart ports includes the following steps: S1. Multimodal sensing data acquisition: Collect heterogeneous data of the port environment through video surveillance, lidar, millimeter-wave radar, RFID readers, electronic weighbridges and environmental monitoring equipment, and perform spatiotemporal synchronization and data cleaning; S2. Multi-source data fusion and feature extraction: Using a feature fusion network with a cross-modal attention mechanism, the multi-source data processed in step S1 is aligned and fused to generate a unified target representation; S3. Digital Twin Status Update: Based on the fused data output in step S2, update the location and status attributes of vehicles, goods, and equipment in the port digital twin model, and perform anomaly detection and simulation. S4. Distributed Decision Making: Based on twin state attributes, local decision-making schemes are generated in parallel for multiple business agents in security monitoring, scheduling optimization, emergency response, and customer service; each agent adopts a policy network based on multi-agent reinforcement learning. S5. Collaborative Optimization: Conflict detection is performed on the local decision schemes generated in step S4; a global decision scheme is generated based on a multi-criteria decision-making method. S6. Decision Execution and Feedback: Distribute the global decision scheme described in S5 to the execution device and collect execution feedback data to update the digital twin model.
[0022] The feature fusion network of the cross-modal attention mechanism in step S2 includes: calculating the correlation weights between different modal features through a multi-head self-attention mechanism, and adaptively combining the modal features using a gated fusion network.
[0023] In step S4, the generation of local decision-making schemes is divided into strategic-level decision-making schemes, tactical-level decision-making schemes, and execution-level decision-making schemes for three levels: long-term resource planning, medium-term task coordination, and real-time control. The decision-making cycle of the strategic-level decision-making scheme is days or weeks, the decision-making cycle of the tactical-level decision-making scheme is hours, and the decision-making cycle of the execution-level decision-making scheme is seconds.
[0024] The collaborative optimization in step S5 includes: using the Contract Network protocol or auction algorithm for task allocation, and using a practical Byzantine fault-tolerant algorithm to achieve multi-agent consensus processing.
[0025] It also includes step S7: performing causal reasoning analysis on the global decision-making scheme based on the causal graph model, and recording the decision chain for post-event tracing.
[0026] Step S3 involves simulation and deduction, which includes predicting the future multi-step state sequence based on the current state and pre-evaluating the effectiveness of the decision-making scheme in the digital twin environment based on the prediction results.
[0027] The Smart Port's multimodal perception and multi-agent collaborative system includes: a multimodal perception layer, a data fusion layer, a digital twin layer, a multi-agent decision-making layer, a collaborative control layer, and an interactive presentation layer; The multimodal perception layer includes a video surveillance module, a lidar module, a millimeter-wave radar module, an RFID reader / writer module, an electronic weighbridge module, and an environmental monitoring module. The data fusion layer, connected to the multimodal perception layer, includes a spatiotemporal synchronization module, a target association module, and a feature fusion module based on a cross-modal attention mechanism; The digital twin layer, connected to the data fusion layer, is used to construct and update the port's 3D model, including a real-time mapping module and a simulation and deduction module; A multi-agent decision-making layer, connected to the digital twin layer, includes multiple business agents based on reinforcement learning; The collaborative control layer, connected to the multi-agent decision-making layer, includes a conflict arbitration module and a consensus formation module; An interactive presentation layer, connected to the collaborative control layer, is used to visualize the global decision-making scheme.
[0028] The agents in the multi-agent decision-making layer adopt a centralized training and decentralized execution architecture. Each agent shares global information during the training phase and makes decisions based on local observations during the execution phase.
[0029] The collaborative control layer also includes a decision synthesis module, which is used to synthesize local decisions into global decisions based on a multi-objective optimization algorithm.
[0030] The system also includes a causal reasoning module for constructing a decision causal graph and generating an interpretable decision report.
[0031] The purpose of this invention is to address the technical shortcomings of existing smart port systems in areas such as multi-source data fusion, real-time perception and decision-making, multi-service collaboration, adaptive fault tolerance, and decision interpretability. It provides a multimodal perception and multi-agent collaborative decision-making method and system for smart ports. This system can: First, achieve deep fusion of multi-source heterogeneous sensing data. By constructing a multimodal sensing network covering the entire port area, deep learning technology is used to align, fuse, and enhance multi-source heterogeneous data such as video, radar, sensors, and electronic tags, breaking through the limitations of single sensing methods and achieving comprehensive perception of the port environment.
[0032] Second, a multi-agent collaborative decision-making architecture is constructed. Port operations are divided into multiple agents based on functional domains, such as security monitoring, scheduling optimization, emergency response, and customer service. Each agent independently perceives the environment, makes local decisions, and achieves global optimization through a collaborative mechanism. This distributed architecture improves the system's scalability and reliability.
[0033] Third, establish an efficient collaboration mechanism among intelligent agents. This involves using a message bus to achieve information sharing, a task allocation protocol to optimize resource allocation, a conflict arbitration mechanism to resolve decision-making conflicts, and a consensus algorithm to achieve consistent decision-making. These mechanisms ensure that multiple agents can work in a coordinated and consistent manner in complex environments.
[0034] Fourth, it provides full traceability of the decision-making process. Through technologies such as decision log recording, decision tree visualization, and causal reasoning analysis, it achieves a complete record and interpretable presentation of the decision-making process. This makes the basis and process of each decision traceable, facilitating regulatory auditing and continuous optimization.
[0035] The technical solution is the core part of the patent application materials, and the applicant should describe this part in as much detail as possible. The technical solution of this invention mainly includes the following core components: 1. System Overall Architecture This invention provides a multimodal perception and multi-agent collaborative decision-making system for smart ports. Its overall architecture comprises six layers: a multimodal perception layer, a data fusion layer, a digital twin layer, a multi-agent decision-making layer, a collaborative control layer, and an interactive presentation layer. These layers are connected and exchange data through standardized interfaces, forming a complete closed loop for smart port perception and decision-making.
[0036] The multimodal perception layer is responsible for collecting various types of perception data about the port environment. This layer includes multiple functional modules: a video surveillance module that uses deep learning networks to achieve real-time target detection and identification, including vehicle detection, personnel detection, and cargo detection; a lidar module that uses 3D point cloud data to achieve spatial positioning and shape reconstruction of targets; a millimeter-wave radar module that detects and tracks moving targets under various weather conditions; an RFID reader / writer module that enables rapid identification and information reading of electronic tags; an electronic weighbridge module that collects weight data of vehicles and cargo; an environmental monitoring module that collects environmental parameters such as temperature, humidity, wind speed, and visibility; and a vessel traffic management module that connects to the port's VTS system for vessel navigation data. All modules use a unified real-time data bus for data transmission to ensure data synchronization.
[0037] The data fusion layer preprocesses and aligns multimodal sensing data. This layer includes the following core modules: a spatiotemporal synchronization module, which uses a unified time protocol to unify data from different sensing devices to the same time base and aligns spatial coordinates through coordinate transformation; a target association module, which uses a multi-hypothesis tracking algorithm to achieve target association across sensors and achieves continuous target tracking through a data association matrix and Kalman filtering; a data cleaning module, which uses statistical methods and robust estimation techniques to filter and correct abnormal data and noise; and a feature extraction module, which extracts high-level semantic features such as vehicle features, cargo features, and environmental features from the raw sensing data and uses deep neural networks to achieve automatic feature learning.
[0038] The digital twin layer constructs a digital mirror of the port's physical space. This layer includes the following functional modules: a 3D modeling module, which uses BIM technology and point cloud fusion technology to construct a detailed 3D model of port facilities, including wharves, storage yards, roads, warehouses, and other facilities; a real-time mapping module, which maps perceived data into the digital twin model in real time, achieving synchronization between the physical and digital worlds; a state estimation module, which estimates the hidden states of various port elements based on Bayesian filtering algorithms and multi-source information fusion technology; and a simulation and deduction module, which supports business simulation and scheme evaluation based on the digital twin, providing predictive analysis capabilities for decision-making.
[0039] The multi-agent decision-making layer comprises multiple business agents with autonomous decision-making capabilities. Each agent employs an independent knowledge base and decision-making model, enabling it to make local decisions based on its own observations and received information. The agents include: a safety monitoring agent, responsible for identifying safety hazards, abnormal behaviors, and safety events, employing deep learning anomaly detection algorithms and time-series analysis techniques; a scheduling optimization agent, responsible for optimizing decisions such as vehicle guidance, yard allocation, and operation scheduling, employing reinforcement learning and operations research algorithms; an emergency response agent, responsible for detecting, assessing, and responding to emergencies, employing event reasoning and emergency plan matching techniques; a customer service agent, responsible for responding to customer inquiries, handling complaints, and providing service requests, employing natural language processing and knowledge graph technologies; and an environmental management agent, responsible for energy management, equipment maintenance, and green port initiatives, employing predictive maintenance and energy consumption optimization algorithms.
[0040] The collaborative control layer is responsible for coordinating the decision-making behavior of each agent. This layer is the core coordination hub of the system and includes the following key modules: an information distribution module, which distributes decision-related information among agents based on a subscription mechanism; a task coordination module, which decomposes and allocates cross-agent collaborative tasks, using contract net protocol and auction algorithm to optimize task allocation; a conflict arbitration module, which detects and resolves conflicts between multi-agent decisions, using multi-criteria decision-making and game theory methods for conflict resolution; a consensus formation module, which achieves consensus among multiple agents on the global decision through a distributed consensus algorithm, using Byzantine fault tolerance and practical Byzantine fault tolerance algorithms; and a decision synthesis module, which synthesizes the local decisions of each agent to generate the optimal global decision scheme, using multi-objective optimization and preference aggregation techniques.
[0041] The interactive presentation layer is responsible for the interaction between the system and the user. This layer provides multiple interaction methods: a monitoring screen module, which displays the panoramic situation of the port in a 3D visualization in the command center, supporting multi-view switching and detailed drilling; a mobile terminal module, which supports mobile inspection and remote command, and provides AR augmented reality navigation functions; a business system interface module, which conducts standardized data exchange with business systems such as customs and port authorities; and an alarm notification module, which pushes alarm information to relevant personnel through sound, light, SMS, and APP, supporting hierarchical and categorized notification strategies.
[0042] 2. Connection relationships between various modules of the system The connections and data flow between the modules are as follows: The outputs of each sensing module in the multimodal perception layer are connected to the corresponding preprocessing modules in the data fusion layer via a real-time data bus, enabling real-time data transmission and synchronization. The output of the data fusion layer is also connected to both the digital twin layer and the multi-agent decision-making layer, providing fused sensing data to both layers. The digital twin layer provides the multi-agent decision-making layer with global situational information and historical state query services. Each agent module communicates bidirectionally with the collaborative control layer, receiving collaborative instructions and reporting local decisions. The interactive presentation layer connects to both the digital twin layer and the collaborative control layer, acquiring situational display data and decision execution feedback.
[0043] The specific connection relationships are described as follows: The video surveillance module, lidar module, millimeter-wave radar module, RFID reader / writer module, electronic weighbridge module, environmental monitoring module, and ship traffic management module are respectively connected to the video processing module, point cloud processing module, radar processing module, tag processing module, weighbridge processing module, meteorological processing module, and VTS processing module of the data fusion layer. The spatiotemporal synchronization module of the data fusion layer simultaneously outputs to the target association module and the real-time mapping module of the digital twin layer. The 3D modeling module of the digital twin layer is communicatively connected to the simulation and deduction module, and the simulation and deduction module is communicatively connected to the task coordination module of the collaborative control layer. Each intelligent agent module is connected to the corresponding coordination submodule of the collaborative control layer, including the information distribution submodule, task coordination submodule, conflict arbitration submodule, consensus formation submodule, and decision synthesis submodule. The output of the decision synthesis module of the collaborative control layer is connected to the monitoring screen module of the interactive presentation layer.
[0044] Communication between modules uses a unified message format, including a message header, message body, and timestamp. The message header contains information such as message type, source module, destination module, and priority. The message body uses the Protobuf serialization format, supporting flexible data structure expansion. Connections between modules employ a redundant design, with backup channels deployed on critical links to ensure system reliability.
[0045] 3. Method and Flow Based on the above system architecture, this invention also provides a core method for multi-agent cooperative decision-making, including the following steps: Step 1: Multimodal Sensing Data Acquisition. Each sensing device collects port environmental data according to preset frequencies and operating modes. The video surveillance module collects high-definition video streams from key areas such as intersections, storage yards, and gates, using H.265 encoding for compressed transmission; the lidar module collects 3D point cloud data within its field of view, reducing transmission bandwidth through point cloud compression algorithms; the millimeter-wave radar module collects distance, speed, and angle information of moving targets, extracting target echoes using a constant false alarm rate (CFAR) detection algorithm; RFID readers collect information from electronic tags passing through the reading / writing area, handling simultaneous readings of multiple tags using anti-collision algorithms; electronic weighbridges collect vehicle load data, eliminating vehicle vibration interference through dynamic weighing algorithms; environmental monitoring equipment collects meteorological and environmental parameter data, including temperature, humidity, wind speed, wind direction, and visibility. All sensing data carries a unified timestamp and device identifier for easy subsequent data fusion processing.
[0046] Step Two: Multi-Source Data Fusion Processing. The data fusion layer fuses the collected multi-source data. First, spatiotemporal synchronization is performed, unifying the data from various sensing devices to the same time reference and spatial coordinate system. Bilinear interpolation and coordinate transformation are used to achieve spatiotemporal alignment of the data. Then, target association is performed, using a multi-hypothesis tracking algorithm to associate observations from different sensors with the same target, and the optimal association matching is solved using the Hungarian algorithm. Next, data cleaning is performed, using Mahalanobis distance and statistical tests to filter outliers and noise. Finally, feature extraction is performed, using deep convolutional neural networks and recurrent neural networks to extract high-level features such as target category, location, velocity, and identity from the original sensing data to obtain a unified target representation.
[0047] Step 3: Digital Twin Status Update. The digital twin layer updates the status of the port's digital twin model based on the fused perception data. Detected elements such as vehicles, goods, and equipment are added to the digital twin model, updating the position, status, and other attributes of each element; hidden states such as target position, velocity, and acceleration are estimated based on the extended Kalman filter algorithm; anomalies and equipment malfunctions are detected using a time-series model-based anomaly detection algorithm; and the display status of the 3D rendering model is updated to reflect changes in the physical world in real time. The digital twin layer also maintains a historical status database, supporting decision backtracking and trend analysis.
[0048] Step Four: Multi-Agent Distributed Decision Making. Each business agent makes local decisions based on locally observed and received global information. Each agent maintains its own knowledge base and decision model. The knowledge base includes business rules, historical cases, and domain expert experience, while the decision model employs a deep reinforcement learning network. Agents use attention mechanisms to filter relevant content from global information, reducing information overload. After generating a local decision plan, agents exchange information with other agents through a collaborative control layer, or they can negotiate point-to-point with other agents. During the decision-making process, each agent considers multiple dimensions such as security constraints, efficiency goals, and cost limitations.
[0049] Step 5: Multi-Agent Collaborative Optimization. The collaborative control layer performs collaborative optimization on the local decisions of each agent. The information distribution module distributes decision-related information to the agents that need it, using a publish-subscribe mechanism based on subscription. The task coordination module decomposes and allocates cross-agent collaborative tasks, using auction algorithms and contract net protocols to optimize task allocation. The conflict arbitration module detects and resolves conflicts between multi-agent decisions, such as the trade-off between security constraints and scheduling efficiency, using multi-criteria decision-making methods for conflict resolution. The consensus formation module achieves consensus among multiple agents on the global decision through a practical Byzantine fault-tolerant algorithm. The decision synthesis module synthesizes the local decisions of each agent to generate the optimal global decision scheme, using weighted summation and multi-objective optimization methods. The collaborative optimization process is iterative, continuing until all agents reach a consensus or the maximum number of iterations is reached.
[0050] Step Six: Decision Execution and Feedback. The collaboratively optimized decision plan is issued for execution, and the execution results are fed back through the interactive presentation layer. The decision execution status is displayed on a large monitoring screen, showing the real-time execution status of decision instructions; mobile terminals guide on-site personnel, providing task navigation and execution confirmation functions; relevant equipment actions are triggered through business system interfaces, including gate control, signal light switching, and broadcast playback; execution results and perception feedback are collected to update the digital twin model status, providing a basis for the next round of decision-making. Execution feedback information is also transmitted to the collaborative control layer for evaluating the decision's effectiveness and triggering adjustments. The entire decision-making process runs continuously at fixed intervals, achieving real-time optimization of port operations.
[0051] 4. Key Technological Innovations This invention has the following key technological innovations compared to the prior art: Technical Innovation Point 1: A Multimodal Perception Fusion Method Based on Attention Mechanism. This invention designs a multi-source data fusion method based on a cross-modal attention mechanism, which can effectively integrate multimodal perception data such as video, radar, and sensors. The method first extracts deep feature representations of each modality through independent perceptual encoder networks, employing a neural network architecture optimized for different modalities. Second, a cross-modal attention module calculates the correlation weights between features of different modalities. The attention weights are calculated through a multi-head self-attention mechanism, which can capture complex relationships between modalities. Finally, a unified multimodal representation is obtained through a feature fusion layer, and a gated fusion network adaptively combines the contributions of different modalities. This method improves target detection accuracy by approximately 15% compared to single-modal methods and significantly enhances robustness under complex weather conditions.
[0052] Technical Innovation Point Two: Layered Distributed Multi-Agent Collaborative Decision-Making Architecture. This invention designs a layered distributed multi-agent collaborative decision-making architecture, dividing the port operation agents into three layers: a strategic layer, a tactical layer, and an execution layer. The strategic layer agents are responsible for long-term planning and resource allocation, such as annual plans and facility construction, with decision-making cycles measured in days or weeks. The tactical layer agents are responsible for mid-term decisions and task coordination, such as work scheduling and personnel shift work, with decision-making cycles measured in hours. The execution layer agents are responsible for real-time control and action execution, such as equipment operation and security protection, with decision-making cycles measured in seconds. Agents at each layer interact through a standardized communication protocol. Upper-layer agents provide goals and constraints to lower-layer agents, while lower-layer agents provide feedback on their execution status to upper-layer agents. This layered architecture realizes a complete decision-making chain from strategy to execution, improving the system's flexibility and scalability.
[0053] Technical Innovation Point 3: Collaborative Decision-Making Optimization Method Based on Multi-Agent Reinforcement Learning. This invention designs a collaborative decision-making optimization method based on multi-agent reinforcement learning, enabling each agent to learn the optimal collaborative strategy through interaction. This method models the port collaborative decision-making problem as a multi-agent partially observable Markov decision process, employing a centralized training and decentralized execution framework. During the training phase, each agent can acquire global information and learn the global value function through a critic network; during the execution phase, it makes decisions based solely on local observations, outputting actions through an actor network. By combining value function decomposition and attention mechanisms, efficient multi-agent collaborative learning is achieved. Agents share key information through a communication network, accelerating training convergence. This method is trained in a simulation environment and fine-tuned through online learning after being transferred to a real environment, continuously adapting to business changes.
[0054] Technical Innovation Point Four: Explainable Decision-Making Mechanism Based on Causal Reasoning. This invention designs an explainable decision-making mechanism based on causal reasoning, making the decision-making process transparent and traceable. This mechanism first constructs a causal graph model of port operations to describe the causal relationships between business elements; secondly, it records the inputs and outputs of key decision nodes during the decision-making process, forming a decision chain; then, it analyzes the influencing factors of each decision based on causal reasoning algorithms to explain why that decision was made; finally, it provides a decision visualization interface to display the decision basis and confidence level. This mechanism supports post-event retrospective analysis, enabling rapid identification of causes when problems arise in decision-making, providing a basis for system optimization.
[0055] Technical Innovation Point Five: Predictive Decision-Making Method Based on Digital Twins. This invention designs a predictive decision-making method based on digital twins, utilizing digital twin models to predict future states and pre-evaluate solutions. The method first establishes a digital twin model of port operations, including a facility model, a business process model, and a resource model; secondly, it loads the current state and predicted future inputs into the digital twin model; then, it generates future state sequences under various scenarios through simulation; finally, it optimizes the solutions based on the prediction results and selects the optimal decision-making solution. This method can identify potential problems in advance, optimize resource allocation, and improve the foresight of decision-making.
[0056] like Figure 1 As shown, the system framework includes a first-layer multimodal perception layer, a second-layer data fusion layer, a third-layer digital twin layer, a fourth-layer multi-agent decision-making layer, a fifth-layer collaborative control layer, and a sixth-layer interactive presentation layer. Specifically, the multimodal perception layer collects multi-source perception data in the smart port scenario; the data fusion layer aligns, fuses, and extracts features from multi-source heterogeneous data; the digital twin layer constructs a synchronous mapping between the port's physical and digital spaces, providing state estimation, simulation, and historical state support; the multi-agent decision-making layer generates local decisions based on the global situation, historical states, and simulation results provided by the digital twin layer; the collaborative control layer distributes information, coordinates tasks, arbitrates conflicts, forms consensus, and synthesizes decisions reported by each agent; and the interactive presentation layer displays, distributes, and provides feedback on global decision schemes, execution instructions, and alarm information. Figure 1 Solid arrows in the diagram represent the main data flow or control flow, while dashed arrows represent feedback or support relationships. Through these data flows, control flows, and feedback relationships, the system forms a closed-loop processing framework encompassing multimodal perception, data fusion, digital twin modeling, multi-agent local decision-making, collaborative optimization, and interactive presentation. Figure 2In this system, the multimodal perception input section receives data from video surveillance, LiDAR, millimeter-wave radar, RFID, electronic weighbridges, environmental monitoring, and VTS. The multimodal perception fusion technology module performs perceptual encoding, cross-modal attention fusion, gating fusion, and unified target representation on multi-source heterogeneous data. The digital twin modeling and prediction technology module performs 3D modeling, real-time mapping, state estimation, and simulation, outputting global situational awareness and prediction results. The hierarchical distributed multi-agent architecture module includes strategic, tactical, and executive agents, enabling hierarchical decision-making and execution feedback. The multi-agent reinforcement learning collaborative optimization module generates optimized decisions through centralized training / decentralized execution, communication collaboration, task allocation, and conflict arbitration / consensus formation. The causal interpretable decision-making technology module outputs interpretable decision results through causal graph models, decision chain records, influencing factor analysis, and traceability visualization. The application output section outputs global decision-making solutions, security warnings, scheduling optimization, emergency response, customer service, and operation and maintenance management results. Figure 1 Solid arrows indicate the main technical processes, while dashed arrows indicate feedback, learning, or model update processes.
[0057] The preceding technical solution section provides a general description of the structural features of the device or apparatus of the present invention, while the detailed implementation section provides a specific description of the structural features of the device or apparatus. Specifically: Example 1: Intelligent Container Scheduling Scenario. A large container port needs to optimize the allocation of import and export containers in the yard and the scheduling of container trucks. The system's multimodal perception layer collects the following information: import and export vessel plans, 3D laser scan data of the container yard, container electronic tag information, container truck GPS positioning data, and road traffic flow data. These data are aligned and correlated through the data fusion layer to form a unified target view.
[0058] The system's digital twin layer constructs a 3D model of the container yard, updating the status of each container location in real time, including location occupancy, stacking layers, and owner information. A scheduling optimization agent analyzes the task sequence and yard status to generate an initial yard allocation plan—placing export containers near the gate and grouping import containers according to destination port and owner. This agent employs a deep reinforcement learning algorithm, considering multiple objectives such as operational efficiency, container turnover rate, and equipment energy consumption.
[0059] The safety monitoring agent detected an abnormal travel trajectory of a truck, deviating from the recommended route. The collaborative control layer initiated a conflict arbitration mechanism, coordinating the scheduling optimization agent and the safety monitoring agent. Considering that safety takes precedence over efficiency, the system decided to temporarily restrict the truck's travel area and simultaneously notified backend personnel to follow up. This conflict arbitration mechanism ensures a balance between safety and efficiency.
[0060] The system generates final dispatch instructions: pushing yard guidance information and optimal driving routes to truck drivers via mobile terminals, and updating navigation paths in real time. Execution results show that the average waiting time for trucks decreased from forty minutes to twenty minutes, and the yard container turnover rate decreased by approximately 30%. This scenario verifies the effectiveness of multi-agent collaborative decision-making in container dispatching.
[0061] Example 2: Customs Cargo Inspection Collaboration Scenario. A customs office at a port needs to issue an inspection order for a batch of imported goods. The multimodal perception layer collects X-ray scan images of the goods, customs declaration data, cargo weighing data, historical inspection records, etc. These data come from different business systems and are integrated through the data fusion layer.
[0062] The risk analysis agent assesses the risk level of goods based on multi-source data: X-ray images show abnormal density distribution of goods, discrepancies exist between the declared product name and image features, historical records show that the cargo owner has a history of violations, and the comprehensive risk score reaches 75 points, triggering a level-two inspection order. This agent employs multimodal feature fusion technology, integrating image features, text features, and statistical features to obtain a comprehensive risk assessment result.
[0063] The collaborative control layer assigns inspection tasks to the on-site inspection agent and the customer service agent. The on-site inspection agent generates an inspection plan, determines inspection locations and sampling requirements, and considers inspection efficiency and coverage. The customer service agent generates an inspection notification, prepares the inspection site and equipment, and coordinates the inspection schedule. The two agents share information and coordinate tasks through the collaborative control layer.
[0064] During the inspection process, the video monitoring module records the inspection footage in real time, providing evidence for post-inspection traceability. Inspection results are uploaded to the system in real time via mobile terminals, and the system automatically generates an inspection report. The system then feeds the inspection results back to the risk analysis agent, updating the cargo risk model and forming a closed-loop management system. This scenario demonstrates the effectiveness of multi-agent collaboration in inspection.
[0065] Example 3: Port Emergency Evacuation Scenario. A hazardous materials leak occurs at a port, requiring rapid personnel evacuation and area control. The environmental monitoring module detects that the concentration of harmful gases exceeds the standard, triggering the alarm threshold. The safety monitoring intelligent agent immediately identifies it as a high-risk event and triggers the emergency response process.
[0066] The emergency response agent quickly activates the emergency plan: it analyzes the affected area and personnel distribution through a digital twin layer, and calculates the best evacuation route based on a 3D model; it calculates the optimal evacuation route through a scheduling optimization agent, taking into account road capacity and personnel density; and it sends early warning notifications to relevant personnel through a customer service agent, ensuring information dissemination through multiple channels.
[0067] The collaborative control layer organizes multi-agent collaborative responses: the scheduling and optimization agent controls the gate to close, preventing vehicles from entering the danger zone, updates the traffic light timing scheme at the intersection, and guides vehicles to detour; the customer service agent initiates voice broadcasts to guide on-site personnel to evacuate, and publishes evacuation instructions through LED screens; the security monitoring agent continuously monitors the leakage situation, updates the boundary of the security area in real time, and dynamically adjusts the control scope.
[0068] The system records the entire emergency response process, including timestamps and key parameters for each stage, such as event detection, decision generation, command issuance, and execution feedback. These records are used for post-event review and analysis to evaluate the effectiveness of the emergency response and identify areas for improvement. This scenario validates the system's rapid response and multi-agent collaborative capabilities in emergency situations.
[0069] Example 4: Nighttime Unmanned Operation Scenario. An automated terminal needs to achieve unmanned horizontal transport by container trucks at night, requiring reliable perception and decision-making under low-light conditions. The system switches to nighttime operation mode, and the multimodal perception layer is adjusted to a nighttime perception configuration: an infrared thermal imager replaces the visible light camera, utilizing thermal radiation information for target detection; millimeter-wave radar improves detection sensitivity, employing a high-resolution mode to enhance small target detection capabilities; and the high-precision positioning system uses RTK differential positioning to achieve centimeter-level positioning accuracy.
[0070] The scheduling optimization agent generates nighttime operation plans, comprehensively considering factors such as ship berthing schedules, yard layout, and traffic flow, and allocates transportation tasks to each unmanned truck. The agent employs a time-series prediction-based optimization algorithm to identify potential traffic conflicts in advance. The digital twin layer updates the location of the unmanned trucks in real time and displays it on a large monitoring screen. The collaborative control layer monitors the operational status, detects anomalies, and issues early warnings.
[0071] The safety monitoring agent continuously detects obstacles and anomalies, employing a multi-sensor fusion obstacle detection algorithm to effectively distinguish between personnel, vehicles, and static obstacles. Upon detecting intrusion, it immediately triggers emergency braking and notifies the monitoring center. Simultaneously, the system activates audible and visual alarms to alert on-site personnel. During nighttime operations, the system completed unmanned transport of 300 standard containers without any safety incidents. Decision logs show smooth collaboration among the agents, with an average decision response time of less than 100 milliseconds. This scenario validates the system's adaptability under special conditions.
Claims
1. A multimodal perception and multi-agent collaborative method for smart ports, characterized in that... Includes the following steps: S1. Multimodal sensing data acquisition: Collect heterogeneous data of the port environment through video surveillance, lidar, millimeter-wave radar, RFID readers, electronic weighbridges and environmental monitoring equipment, and perform spatiotemporal synchronization and data cleaning; S2. Multi-source data fusion and feature extraction: Using a feature fusion network with a cross-modal attention mechanism, the multi-source data processed in step S1 is aligned and fused to generate a unified target representation; S3. Digital Twin Status Update: Based on the fused data output in step S2, update the location and status attributes of vehicles, goods, and equipment in the port digital twin model, and perform anomaly detection and simulation. S4. Distributed Decision Making: Based on twin state attributes, local decision-making schemes are generated in parallel for multiple business agents in security monitoring, scheduling optimization, emergency response, and customer service; each agent adopts a policy network based on multi-agent reinforcement learning. S5. Collaborative Optimization: Conflict detection is performed on the local decision schemes generated in step S4; a global decision scheme is generated based on a multi-criteria decision-making method. S6. Decision Execution and Feedback: The global decision mentioned in S5 is sent to the execution device, and execution feedback data is collected to update the digital twin model.
2. The multimodal perception and multi-agent collaboration method for smart ports according to claim 1, characterized in that... The feature fusion network of the cross-modal attention mechanism in step S2 includes: calculating the correlation weights between different modal features through a multi-head self-attention mechanism, and adaptively combining the modal features using a gated fusion network.
3. The multimodal perception and multi-agent collaboration method for smart ports according to claim 1, characterized in that... In step S4, the generation of local decision-making schemes is divided into strategic-level decision-making schemes, tactical-level decision-making schemes, and execution-level decision-making schemes for three levels: long-term resource planning, medium-term task coordination, and real-time control. The decision-making cycle of the strategic-level decision-making scheme is days or weeks, the decision-making cycle of the tactical-level decision-making scheme is hours, and the decision-making cycle of the execution-level decision-making scheme is seconds.
4. The multimodal perception and multi-agent collaboration method for smart ports according to claim 1, characterized in that... The collaborative optimization in step S5 includes: using the contract network protocol or auction algorithm for task allocation, and using a practical Byzantine fault-tolerant algorithm to achieve multi-agent consensus processing.
5. The multimodal perception and multi-agent collaboration method for smart ports according to claim 1, characterized in that, It also includes step S7: performing causal reasoning analysis on the global decision-making scheme based on the causal graph model, and recording the decision chain for post-event tracing.
6. The multimodal perception and multi-agent collaboration method for smart ports according to claim 1, characterized in that... Step S3 involves simulation and deduction, which includes predicting the future multi-step state sequence based on the current state and pre-evaluating the effectiveness of the decision-making scheme in the digital twin environment based on the prediction results.
7. A multimodal perception and multi-agent collaborative system for smart ports, used to implement the method described in any one of claims 1-6, characterized in that... include: Multimodal perception layer, data fusion layer, digital twin layer, multi-agent decision-making layer, collaborative control layer, and interactive presentation layer; The multimodal perception layer includes a video surveillance module, a lidar module, a millimeter-wave radar module, an RFID reader / writer module, an electronic weighbridge module, and an environmental monitoring module. The data fusion layer, connected to the multimodal perception layer, includes a spatiotemporal synchronization module, a target association module, and a feature fusion module based on a cross-modal attention mechanism; The digital twin layer, connected to the data fusion layer, is used to construct and update the port's 3D model, including a real-time mapping module and a simulation module; A multi-agent decision-making layer, connected to the digital twin layer, includes multiple business agents based on reinforcement learning; The collaborative control layer, connected to the multi-agent decision-making layer, includes a conflict arbitration module and a consensus formation module; An interactive presentation layer, connected to the collaborative control layer, is used to visualize the global decision-making scheme.
8. The multimodal perception and multi-agent collaborative system for smart ports according to claim 7, characterized in that... The agents in the multi-agent decision-making layer adopt a centralized training and decentralized execution architecture. Each agent shares global information during the training phase and makes decisions based on local observations during the execution phase.
9. The multimodal perception and multi-agent collaborative system for smart ports according to claim 1, characterized in that... The collaborative control layer also includes a decision synthesis module, which is used to synthesize local decisions into global decisions based on a multi-objective optimization algorithm.
10. The multimodal perception and multi-agent collaborative system for smart ports according to claim 1, characterized in that... The system also includes a causal reasoning module for constructing a decision causal graph and generating an interpretable decision report.