Intelligent processing method and system for aviation logistics
The intelligent processing system for air logistics, which integrates multimodal fusion and knowledge graphs, solves the problems of manual reliance and low efficiency in air logistics processing. It achieves automated data governance and verification, improves data accuracy and security, and adapts to changes in the business environment.
Patent Information
- Application Number
- CN202610090428.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-23
- Publication Date
- 2026-02-24
AI Technical Summary
The existing air logistics processing flow relies on manual operation, which is inefficient, prone to errors, has inconsistent quality of multi-source heterogeneous data, and the various processing links are isolated from each other. It also lacks data governance and verification mechanisms, resulting in insufficient security for security checks.
An intelligent processing system for air logistics employing multimodal fusion and knowledge graphs receives multimodal data, performs data governance and multimodal verification, utilizes the cross-attention model of the Transformer architecture and graph neural network for data verification and decision-making, and introduces a reinforcement learning mechanism to optimize the system.
It achieves full-process automation and collaborative optimization, improves data accuracy and security, reduces error rate, enhances system flexibility and robustness, and adapts to changes in the business environment.
Smart Images

Figure CN121563353A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technology, and in particular to an intelligent processing method and system for air logistics based on multimodal fusion and knowledge graph. Background Technology
[0002] With the digital transformation of the air logistics industry, massive amounts of business data are constantly being generated. How to efficiently and accurately process this data and empower business decisions has become crucial for the industry's development. In traditional air logistics operations, key processes such as security checkpoint allocation and cargo arrival sorting rely heavily on manual experience. Operators need to make decisions based on complex allocation rules, such as the priority of different airlines, the type of special cargo, and cargo size restrictions. This manual approach is not only inefficient but also prone to errors in allocation or misclassification due to human negligence or misjudgment, thus affecting the timeliness and reliability of the entire logistics chain.
[0003] Furthermore, air logistics data typically originates from multiple different entities, including airport cargo terminals, airlines, and freight forwarders, and takes various forms, such as electronic waybills, manifests, and cargo status reports. The quality of data reported from different sources varies, frequently resulting in information conflicts or inconsistencies. For example, the names of goods for the same shipment reported by different agents may differ. Current technologies often lack a unified and effective data governance and verification mechanism to resolve these data conflict issues, leading to inaccurate basic data upon which business systems rely and severely impacting the reliability of subsequent data analysis and business decisions.
[0004] Currently, while some information systems have attempted to automate logistics processes, such as automating customs clearance or sorting of goods according to preset rules, these systems typically treat security checks, distribution, and data governance as independent modules, creating information silos and lacking integration and collaborative optimization of the entire process. More importantly, these systems often assume the accuracy of the data they receive, failing to effectively address the inconsistencies caused by the aforementioned multi-source heterogeneous data. Furthermore, for the crucial security verification step of "document-cargo consistency"—verifying whether the cargo information declared on the waybill matches the actual cargo—current technology still primarily relies on manual interpretation at X-ray security scanners. There is a lack of intelligent verification methods that can automatically compare the text information of the shipment with the image information of the cargo. This is not only an efficiency bottleneck in the security check process but also introduces potential security risks. Summary of the Invention
[0005] The purpose of this application is to provide an intelligent processing system and method for air logistics based on multimodal fusion and knowledge graph, which aims to solve the technical problems in the existing air logistics processing process, such as reliance on manual labor, low efficiency, susceptibility to errors, inconsistent quality of multi-source heterogeneous data, and isolation between processing links.
[0006] To achieve the above objectives, this application provides an intelligent processing method for air logistics, which includes the following steps: Receive multimodal air logistics data, which includes at least business message data and cargo image data; Perform data governance, which includes: configuring trust weights for different data sources, and when new data that conflicts with existing data is received, executing a preset data update strategy by comparing the trust weights of the sources of the new and old data; Multimodal data verification is performed by using a cross-attention model based on the Transformer architecture to interactively analyze the text description in the business message data and the visual features in the cargo image data to verify the consistency between the order and the cargo. Based on the governed and verified data, execute automated decisions and allocations for at least one logistics task.
[0007] Optionally, configuring trust weights for different data sources specifically includes configuring corresponding trust weights for the data sources based on the enterprise type, data type, or data item of the data source.
[0008] Optionally, the multimodal data verification further includes: calculating a consistency score between the text and image content based on the results of the interaction analysis, so as to quantify the consistency of the goods.
[0009] Optionally, the logistics operations include at least one of automatic optimized allocation of security checkpoints and intelligent warehouse sorting after goods arrive at the port.
[0010] Optionally, the automatic optimization allocation of the security checkpoints specifically includes: based on preset priority rules, availability rules, and queuing order rules, filtering and weighting the goods to be processed to generate a security checkpoint allocation instruction.
[0011] Optionally, the intelligent warehouse division and diversion of goods after they enter the port specifically includes: assigning a destination warehouse area to the goods to be processed based on the preset mapping relationship and priority between special cargo codes and special cargo operation codes.
[0012] Optionally, the automated decision-making and allocation are implemented through a graph neural network running on a knowledge graph; the knowledge graph defines entity nodes such as goods, special goods codes, operation codes, and warehouse areas, as well as the relationship edges between entity nodes, and the graph neural network realizes reasoning and decision-making on business rules by performing message passing and aggregation on the knowledge graph.
[0013] Optionally, the method further includes: introducing a reinforcement learning mechanism, using the decision-making effect of logistics business as a reward signal, and dynamically adjusting the trust weight or the business rule weight on which the decision allocation is based.
[0014] To achieve the above objectives, this application also provides an intelligent processing system for air logistics, characterized by comprising: a data receiving module for receiving multimodal air logistics data, wherein the multimodal air logistics data includes at least business message data and cargo image data; a data governance module for performing data governance, wherein the data governance module is configured to: configure trust weights for different data sources, and when new data conflicting with existing data is received, execute a preset data update strategy by comparing the trust weights of the sources of the new and old data; a multimodal verification module for performing multimodal data verification, wherein the multimodal verification module incorporates a cross-attention model based on the Transformer architecture for interactive analysis of the text description in the business message data and the visual features in the cargo image data to verify the consistency between the shipment and the goods; and a decision execution module for performing automated decision-making and allocation of at least one logistics business based on the governed and verified data.
[0015] Compared with the prior art, this application has the following beneficial effects: First, this application integrates data governance, multimodal verification, and intelligent decision-making into a unified system and methodology, streamlining key processes such as data processing, security checks, and data routing. This achieves end-to-end automation and collaborative optimization, improving overall efficiency and synergy. Second, based on precise business rules and algorithmic models, this application automates decision-making and automatically verifies "item-goods consistency," significantly reducing allocation errors that may result from manual operations and effectively identifying security risks such as false reporting and misreporting of product names, thereby lowering the error rate and enhancing security. Third, by establishing a weight-based data governance mechanism, this application automatically resolves conflicts between multi-source heterogeneous data, ensuring the accuracy, consistency, and authority of business data, providing a reliable data foundation for upper-level intelligent decision-making. Finally, by introducing advanced algorithmic models such as graph neural networks and reinforcement learning, this application makes business rules and system weights learnable and adaptive, enhancing the system's flexibility and robustness in responding to changes in business processes and the data environment, enabling continuous self-optimization. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Furthermore, these drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments.
[0017] Figure 1 A schematic diagram of the overall architecture of an intelligent air logistics processing system provided in this application embodiment; Figure 2 A flowchart illustrating an intelligent air logistics processing method provided in this application embodiment; Figure 3 A schematic diagram of the security check and ticket queuing process provided in this application embodiment; Figure 4 A schematic diagram of the intelligent cargo diversion sub-process provided in the embodiments of this application; Figure 5 A schematic diagram illustrating the principle of multimodal "document-to-goods consistency" verification provided in this application embodiment; Figure 6 A schematic diagram of flexible rule-based reasoning based on GNN provided for an embodiment of this application; Figure 7 This is a schematic diagram of a dynamic weight optimization loop based on reinforcement learning, provided as an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0019] Example 1 This embodiment provides a basic implementation scheme for an intelligent processing system and method for air logistics. The scheme aims to solve problems such as low efficiency, data inconsistency, and reliance on manual security verification in traditional air logistics processing flows, and fully demonstrates the entire process from receiving multimodal data, performing data governance, conducting multimodal verification, to finally achieving intelligent decision-making and allocation.
[0020] Please see Figure 1This diagram illustrates the overall architecture of an intelligent air logistics processing system according to an embodiment of this application. The system can be deployed on a cloud server, a local server cluster, or a hybrid cloud environment, and its core logic consists of a series of functional modules. Specifically, the system includes a data input source 10, a data processing engine 20, a rule engine 30, a decision engine 40, a data storage layer 50, and a result output terminal 60.
[0021] Data input source 10 serves as the entry point for the system to interact with the external world, receiving air logistics data from various business entities such as airport cargo terminal systems, airline systems, freight forwarder systems, and customs systems. This data can be accessed in various ways, such as real-time push via application programming interfaces (APIs) or retrieval from external databases or file servers via scheduled batch processing tasks. It should be noted that the received data is multimodal, including at least structured business message data (such as electronic air waybill messages and cargo manifest messages recommended by the International Air Transport Association) and unstructured cargo image data (such as images captured by X-ray security scanners, high-speed scanners, etc., during cargo inspection or warehousing).
[0022] The data processing engine 20, as the core data processing hub of the system, receives raw data from the data input source 10 and performs a series of cleaning, transformation, governance, and verification operations. In one embodiment of this application, the data processing engine 20 can be further divided into a data receiving module, a data governance module, and a multimodal verification module. The data receiving module is responsible for receiving and parsing input data (multimodal air logistics data) in different formats and unifying it into an internal standard format. The data governance module is responsible for resolving data conflicts and inconsistencies, i.e., performing data governance, including configuring trust weights for different data sources, and when new data conflicting with existing data is received, executing a preset data update strategy by comparing the trust weights of the sources of the new and old data. The multimodal verification module is responsible for cross-validating data from different modalities, and incorporates a cross-attention model based on the Transformer architecture to interactively analyze the text descriptions in business message data and the visual features in cargo image data to verify the consistency between the shipment and the goods.
[0023] Rule Engine 30 is a configurable business rule management center used to store and manage the decision logic for various logistics operations. Business experts can configure and modify these rules (such as security checkpoint allocation rules, warehouse zoning rules, etc.) through a graphical interface without changing the system code.
[0024] As the decision-making core of the system, the decision engine 40 is based on the high-quality data provided by the data processing engine 20, which has been governed and verified, and combined with the business rules defined in the rule engine 30 and the historical data or knowledge graph information obtained from the data storage layer 50, to execute specific automated decision-making and allocation tasks for logistics business.
[0025] The data storage layer 50 provides persistent data storage capabilities for the entire system. Understandably, this layer can be composed of a combination of various database technologies, such as using relational databases to store structured business data, document databases or key-value stores to store semi-structured message data, object storage to store large amounts of image files, and graph databases to build and store aviation logistics knowledge graphs.
[0026] Output terminal 60 is responsible for sending the decision instructions generated by decision engine 40 to the corresponding execution terminals. These terminals can be web interfaces used by on-site operators, handheld mobile device applications, or control systems integrated with automated equipment such as sorting line controllers and robotic arms.
[0027] The following will combine Figure 2 This application describes the flow of an intelligent air logistics processing method provided in its embodiments. This method can be... Figure 1 The system execution shown includes the following steps: Step S100: Receive multimodal data. When a shipment enters the logistics process (e.g., arrives at an airport cargo terminal for security inspection), the system's data receiving module receives multimodal data related to the shipment from different data input sources 10. For example, it receives the electronic waybill message for the shipment from the freight forwarder's system. Simultaneously, when the shipment passes through an X-ray security scanner, the security equipment transmits the collected X-ray image data of the shipment to the system in real time.
[0028] Step S200: Perform data governance. Since information reported from different data sources may conflict, the data governance module processes the received data to ensure its consistency and authority. The core mechanism of this module is to configure trust weights for different data sources.
[0029] As a specific implementation method, this weight can be configured with fine granularity based on the type of enterprise, data type, or specific data item from which the data originates. For example, the system administrator can pre-set the following in the configuration table of data storage layer 50: data directly reported by airlines (as carriers) has a weight of 0.9, data reported by large, reputable core agents has a weight of 0.8, and data reported by small or newly cooperating agents has a weight of 0.6. The weight value can be a floating-point number between 0 and 1, with higher values representing higher trust levels.
[0030] When the system receives new data that conflicts with existing data in the database, the data governance module will initiate a data update strategy. For example, a newly received electronic waybill message lists the product name as "lithium battery," while a record in the database reported by another agent shows the same shipment as "electronic equipment." This strategy decides based on the trust weight of the sources of the new and old data. In this example, suppose the data "lithium battery" comes from a core agent with a weight of 0.8, while the data "electronic equipment" comes from a regular agent with a weight of 0.7. Since 0.8 > 0.7, the system will adopt the new data from the data source with the higher trust weight, update the product name information in the database to "lithium battery," and log this update operation.
[0031] Step S300: Perform multimodal data verification. To further ensure the accuracy of cargo information, especially to prevent security risks such as misreporting or falsely reporting product names, the system performs automated verification of "document and cargo consistency." This step is completed by the multimodal verification module. Please refer to [link / reference]. Figure 5 It demonstrates the principle of multimodal "single-item consistency" verification. The module has a built-in cross-attention model 130 based on the Transformer architecture, which is pre-trained on a large number of data pairs (business messages, X-ray images, and matching / non-matching labels).
[0032] During runtime, the model receives two inputs: a text description in the processed business message 110, such as the product name "lithium battery"; and the corresponding X-ray image 120 of the goods. Internally, the text encoder converts the text entity 111 of "lithium battery" into a high-dimensional semantic vector. Simultaneously, the image encoder divides the X-ray image 120 into multiple image patches and extracts the visual feature vector for each patch. The key cross-attention mechanism works by using the semantic vector of the text entity 111 as the query and the visual feature vectors of all image patches as keys and values. The model calculates the similarity between the query vector and each key vector, generating an attention weight distribution. This distribution highlights the visual feature regions 121 in the image most relevant to the semantics of "lithium battery." The model then aggregates the features of these highlighted regions and deeply fuses them with the text features. Finally, the model's output layer provides a consistency score (e.g., a probability value between 0 and 1) to quantify the degree of consistency between the text and image content. If the score is higher than a preset threshold (e.g., 0.9), the single-goods consistency check is considered passed.
[0033] Step S400: Execute intelligent decision-making and allocation. Based on the high-quality, processed, and validated data generated in the previous steps, the decision engine 40 executes at least one specific automated logistics business decision. In this embodiment, the automatic optimization allocation of security checkpoints is used as an example for illustration, namely... Figure 2 Security check and ticket queuing steps S410.
[0034] Please see Figure 3 It details the security check and ticketing process. The process begins at step S411, where the decision engine 40 receives the verified cargo information. This information is multi-dimensional, including details such as the cargo's master order number, product name, quantity, weight, dimensions, special cargo code, airline, and agent information.
[0035] Next, rule classification is performed in step S412. Decision engine 40 loads the rule set related to security checkpoint ticketing from rule engine 30 and classifies these rules. These rules can generally be divided into three categories: 1. Priority rules: These define which goods must or have priority access to specific channels, such as a dedicated channel contracted by an airline, or a special channel for handling dangerous goods or live animals.
[0036] 2. Available Class Rules: Defines the applicability restrictions of the passage, such as the physical size restrictions of the passage (oversized or overweight goods cannot pass through), equipment capacity (some passages do not have refrigeration equipment), or current status (whether it is open or not).
[0037] 3. Queuing order rules: Define the queuing priority of goods when multiple channels are available. For example, a "agent consistency" rule can be set to allocate multiple shipments from the same agent to the same channel for centralized processing; or a "number of items" rule can be set to guide small items to idle channels to balance the load.
[0038] Subsequently, channel filtering and verification are performed in step S413. Decision engine 40 first applies priority class rules to determine whether there is a required channel. If not, it applies availability class rules to traverse all currently available security check channels and filter out a list of candidate channels that meet the attributes of the goods (such as size, special cargo type).
[0039] Then, in step S414, weighted sorting is performed. For each channel in the candidate channel list, the decision engine 40 calculates a comprehensive score based on queuing order rules. For example, for a shipment containing the attributes "SF Express agent," "refrigerated goods," and "oversized," the system filters out channels C01 and C02, which can handle refrigerated goods and meet the size requirements. In this case, if the rule sets "SF Express agent" to have a high priority, and the goods currently being processed by channel C01 are mostly SF Express agents, then according to the "agent consistency" rule, C01 will receive a higher weighted score. The system integrates the scores of all queuing order rules and performs weighted sorting of the candidate channels.
[0040] Finally, a ticket number is generated in step S415. The decision engine 40 selects the highest-ranked channel (e.g., C01) as the final allocation result and generates a unique security ticket number for the cargo. This ticket number typically contains channel information and a serial number, such as "C01-0005".
[0041] Step S500: Output decision instructions. The security checkpoint allocation instructions (including ticket number, destination channel, cargo information, etc.) generated by the decision engine are pushed to the airport cargo terminal's on-site execution system in a structured data format (e.g., JSON) through the result output terminal 60, and displayed on the operator's computer screen or handheld device to guide them to send the cargo to the designated C01 channel.
[0042] Through the above process, this embodiment achieves fully automated processing of a single shipment, from data reception, cleaning, and verification to final security check and allocation. This not only ensures the accuracy of basic data but also enhances the security of the security check process through intelligent multimodal verification and improves the utilization efficiency of security check resources through rule-based refined allocation.
[0043] Example 2 This embodiment, based on Embodiment 1, provides a more advanced and flexible variant of the implementation of the decision engine 40. Traditionally, the decision logic in step S400 may consist of a large number of hard-coded conditional statements or decision tables. When business rules are complex and changeable, this approach suffers from limitations such as high maintenance costs and poor scalability. This embodiment introduces graph neural network technology to perform reasoning on the aviation logistics knowledge graph, thereby achieving flexible reasoning and adaptive decision-making for complex and composite business rules.
[0044] In this embodiment, an aviation logistics knowledge graph based on dynamic weight rules is constructed and maintained in the data storage layer 50. This knowledge graph is stored in the form of a graph database, which defines various types of entity nodes and relational edges. Entity nodes include, but are not limited to: cargo nodes, special cargo code nodes (such as COL node 221 representing refrigerated goods, DGR node 222 representing dangerous goods), operation code nodes (such as DCOL node 230 representing refrigerated dangerous goods operations), warehouse area nodes 240 (such as refrigerated dangerous goods areas), airline nodes, agent nodes, etc. Relational edges define the connections between these entities, such as "contains," "mapped to," "higher priority," "stored in," etc. Accordingly, relational edges can also have weights to represent the strength or priority of the relationship. For example, when multiple data sources have different weight data for the same shipment, the system will prioritize adopting the node information corresponding to the data of the enterprise with higher weight, realizing knowledge-based intelligent data governance.
[0045] When constructing an aviation logistics knowledge graph based on dynamic weight rules, in addition to defining traditional entities (such as enterprises, waybills, cargo, airports, and flights) and relationships (such as receiving and carrying), two special types of entities are introduced: Rule entities: For example, airline priority rules and data weighting rules. These entities have attributes such as rule effective time and rule priority.
[0046] Metric entities: For example, data accuracy, data timeliness, and data authenticity. These entities are linked to enterprise entities, and their values can be updated dynamically.
[0047] At the same time, define a new relation type to connect them: hasRule: Connects airport and airline priority rules.
[0048] hasWeight: Connects enterprises with data weight rules; the relationship itself can carry a weight value attribute.
[0049] Furthermore, graph neural networks supporting relational types, such as R-GCN, are used for knowledge representation learning. The key logic of the learning is that during message passing in R-GCN, when information is transmitted from an enterprise node to the waybill node it carries, the amount of information transmitted is modulated (e.g., multiplied by the dynamic weight value of the hasWeight relation edge). Therefore, the node corresponding to data uploaded by a high-weight enterprise will have a stronger signal influence in the graph, giving it a more dominant position in subsequent graph inference. This means that the business concept of "data weight" is directly encoded into the propagation function of the graph neural network.
[0050] In this way, the knowledge graph internalizes the dynamic business rules and data quality weights that were originally external to the knowledge graph into the knowledge graph itself, enabling the graph to carry and express dynamic business knowledge, thereby achieving a unified representation of perception (data) and cognition (rules).
[0051] The following describes the intelligent warehouse sorting task after goods arrive at the port (i.e.) Figure 2 Taking the library separation step S420 as an example, this section illustrates the decision-making process based on a graph neural network. Please refer to [link / reference]. Figure 4 The diagram illustrates the sub-processes of intelligent cargo diversion.
[0052] Step S421: Parse the Electronic Waybill (FWB) message. After the system receives and processes the data of an inbound cargo shipment, the decision engine 40 first parses the message data and extracts all relevant special cargo codes. Suppose a shipment is declared with both "COL" (refrigerated) and "DGR" (dangerous goods, a certain category) special cargo codes.
[0053] Steps S422 and S423: Matching special cargo operation codes and determining priority. In traditional rule engines, the system needs to query the rule base to find matching rules, such as "IF a.has('COL') AND a.has('DGR') THENassign_to('DCOL_AREA')". This method becomes extremely cumbersome when there are many rule combinations.
[0054] In this embodiment, the decision engine 40 uses a graph neural network inference engine to complete this task, replacing the traditional, rigid "if-else" rule chain. Through the message passing mechanism of the graph neural network (GNN) on the graph, it simulates the expert-level comprehensive reasoning process and realizes the automatic and flexible processing of complex and composite business rules.
[0055] Specifically, taking waybill information - cargo type identification as an example, a subgraph is constructed from each identification request. This subgraph includes: Central node: Current waybill.
[0056] First-degree neighbor nodes: the special cargo code (such as COL, DGR) and customs code of the goods for this waybill.
[0057] Second-degree neighbor nodes: Special cargo operation codes (such as COL, DCOL), product names (such as refrigerated goods, dangerous goods) connected to the above neighbors, and priority relationships connecting these entities.
[0058] Next, a priority-based message passing mechanism is designed, which includes the following processes: 1. Initialization: All special cargo codes and customs codes for goods are assigned an initial activation value based on their existence.
[0059] 2. First round of message passing: The special cargo code node transmits the activation signal to the special cargo operation code nodes connected to it. The signal strength is affected by the priority value on the connection edge (the higher the priority, the stronger the signal).
[0060] 3. Competition and Aggregation: A special operation code node (such as DCOL) may receive signals from multiple special operation code nodes (COL and DGR). The aggregation function of GNN (such as max pooling) will selectively retain the strongest input signal.
[0061] 4. Second round of message passing: The activated special cargo operation code node transmits the signal to the product name node. Similarly, the cargo customs code node also directly transmits the signal to its bound product name node.
[0062] 5. Decision Readout: Ultimately, all product name nodes connected to the current waybill will receive an activation score. The system selects the product name with the highest activation score as the target.
[0063] This mechanism naturally handles rule conflicts and complex rules. For example, even if a cargo has multiple special cargo codes, only the one with the highest priority will ultimately "win" through competition and determine its destination. Furthermore, the entire process involves no "IF-THEN" statements; instead, the decision emerges naturally through the flow of signals and competition on the graph structure.
[0064] For a better understanding, please refer to Figure 6 This illustrates the flexible rule-based reasoning process based on graph neural networks. The system first locates the cargo node 210 representing the shipment in the knowledge graph, as well as the COL node 221 and DGR node 222 connected to it.
[0065] Next, the graph neural network inference engine initiates the message passing process. In the graph, COL node 221 may connect to the "Ordinary Cold Storage Operation" node, and DGR node 222 may connect to the "Ordinary Hazardous Goods Operation" node. However, there is also a rule in the knowledge graph that COL node 221 and DGR node 222 both point to a more specific operation code node, namely DCOL node 230. Furthermore, the priority weight of this composite relationship from (COL, DGR) to DCOL is set higher than the weight of their respective connections to ordinary operation nodes.
[0066] Graph neural network models perform inference by propagating and aggregating node information (or "messages") along the edges of the graph. This process can be understood as follows: an "activation" signal from cargo node 210 flows to COL node 221 and DGR node 222. In the next iteration, these two nodes continue propagating the signal along their out-degree edges. During the aggregation phase, DCOL node 230 receives signals from both COL node 221 and DGR node 222. Because the path leading to it has the highest total weight, this node obtains the strongest activation score after calculation using aggregation functions (such as weighted summation) and activation functions (such as ReLU).
[0067] Step S424: Assign destination storage area. The decision engine 40 selects the DCOL node 230 with the highest activation score as the inference result, and further queries the graph to find the storage area node 240 connected to DCOL node 230 through the "stored in" relationship, namely the "refrigerated hazardous goods area". Finally, the system generates an instruction to assign the shipment to the refrigerated hazardous goods area.
[0068] Understandably, the entire process requires no explicit "IF...AND...THEN..." rules; business rules are embedded within the structure and weights of the knowledge graph. When rules need to be added or modified (e.g., adding a new combination of special products), only new nodes and relationships need to be added to the graph, and the system can automatically learn and adapt. This approach greatly enhances the system's ability to handle complex combinations of rules that are not predefined, reduces reliance on rigid rule engines, and improves the flexibility and robustness of decision-making.
[0069] Example 3 This embodiment, building upon Embodiments 1 and 2, further introduces a reinforcement learning mechanism, aiming to enable the system's core parameters to self-optimize and dynamically adjust, thereby adapting to the ever-changing business environment. Specifically, this embodiment improves the trust weight system in the data governance module and the decision model weight system in the decision execution module.
[0070] Please see Figure 7 This diagram illustrates a dynamic weight optimization loop based on reinforcement learning. At the heart of this mechanism is a closed-loop interaction between an agent 310 and the environment 320.
[0071] In the context of this application, Environment 320 refers to the entire intelligent air logistics processing system and its actual business operation environment.
[0072] An agent 310 is an independent computing unit, typically consisting of one or more neural networks (such as policy networks and value networks), designed to learn optimal behavioral policies.
[0073] State 330, which is a snapshot of the environment 320 observed by agent 310 at a certain moment, is usually represented as a numerical vector or tensor, containing key indicators that describe the current operating status of the environment. For example, the state may include: the real-time queue length of each security checkpoint, the average waiting time of each checkpoint in the past hour, the frequency of data inconsistency events, the success rate of data verification submitted by different agents, and the CPU and memory usage of the system server, etc.
[0074] Action 340 is the operation that agent 310 can perform based on the currently observed state 330 and its own policy. In this embodiment, the action space is defined as the adjustment of key weights within the system. Specifically, it may include: 1. Adjust the trust weight of the data source: For example, agent 310 can decide to increase or decrease the trust weight of a certain agent by 0.05.
[0075] 2. Adjusting feature weights in the decision model: For example, in the security check and ticketing weight ranking step S414 of Example 1, if the ranking model is an interpretable linear model or a ranking learning algorithm model such as LambdaMART, the agent 310 can adjust the weights of various features (such as "agent consistency", "number of goods", "airline priority").
[0076] Reward 350 is a scalar signal from the environment 320 to the agent 310 after the agent performs an action 340, used to evaluate the merits of that action. The design of the reward function is crucial; it needs to be positively correlated with the ultimate business objective. For example, the reward function can be defined as a linear combination of metrics related to the business objective (such as increasing total security throughput or reducing average waiting time). A simple reward function could be: ,in, This is the original average waiting time. The average waiting time after executing action 340. This refers to the reduction in average waiting time. If an action leads to a decrease in average waiting time, the reward is positive; otherwise, it is negative.
[0077] To better understand this, the following will illustrate the optimization loop through a specific workflow: 1. Through the monitoring system, agent 310 observes the current state 330: security checkpoint C03 is severely congested due to a large influx of small items in a short period of time, and its average waiting time exceeds the preset warning threshold.
[0078] 2. The policy network of agent 310 receives the state vector as input and outputs a probability distribution of action 340. Based on sampling of this distribution, the agent decides to perform an action: fine-tuning the internal feature weights of the security check ticket sorting model (step S414), temporarily increasing the weight of the feature "few items" by 20%, while slightly decreasing the weight of "agent consistency".
[0079] 3. In the next time window (e.g., 15 minutes), decision engine 40 uses the adjusted model to schedule tickets. As the weight of "few items" is increased, smaller batches of goods arriving later are more likely to be directed to other relatively empty channels, thus diverting goods heading to C03 and alleviating congestion at C03.
[0080] 4. After the time window ends, agent 310 observes the new state and finds that the overall average waiting time of the system has decreased compared to before. Based on the reward function calculation, it receives a positive reward of 350.
[0081] 5. This positive reward signal is used to update the policy network of agent 310. Through algorithms such as policy gradient (e.g., proximal policy optimization algorithm), the agent will adjust its network parameters so that in the future, when encountering a situation such as "a channel is congested due to small items", it is more likely to take the advantageous action of "increasing the diversion weight of small items".
[0082] Through continuous learning and trial-and-error processes, Agent 310 can learn a complex strategy for dynamically adjusting system parameters, enabling the system to automatically adapt to changes in business traffic, sudden events, etc., and continuously optimize resource allocation efficiency and data governance accuracy without human intervention, thereby realizing the online evolution and continuous improvement of system decision-making strategies.
[0083] Example 4 This embodiment focuses on the training method of the artificial intelligence model in the system, aiming to ensure that the learning objectives of each model component within the system are highly aligned with the final business objectives, thereby improving the end-to-end performance of the system in real business scenarios. As an optional implementation, this embodiment adopts a business-oriented end-to-end joint optimization training strategy.
[0084] In Examples 1 and 2, the system includes multiple models based on machine learning or deep learning, such as a multimodal verification module (with a built-in cross-attention model) for "single-item matching" verification, and a graph neural network model that may be used for decision-making. The traditional approach is to train these models in stages and independently. For example, the multimodal verification model might be trained on a dataset labeled "match / mismatch," with the optimization objective being to maximize classification accuracy; while the graph neural network model might be trained on another dataset labeled "correct storage area," with the optimization objective being to minimize the cross-entropy loss of storage area predictions. It should be noted that this approach has certain limitations, namely that each module only focuses on its own local optimum, and the superposition of local optima does not necessarily achieve a global optimum. For example, a model with an accuracy of 99% on the verification task might have a 1% error that leads to a subsequent catastrophic storage area allocation error.
[0085] This embodiment proposes a different training paradigm, which constructs a unified, fully differentiable deep learning architecture from raw input to final decision output. This architecture connects the aforementioned modules (such as text encoder, image encoder, cross-attention layer in multimodal verification module, graph neural network inference layer, etc.) as different layers of the network.
[0086] The training process is as follows: 1. Prepare training data: The training sample is a tuple containing all the information required for the end-to-end process, such as (business message data, cargo X-ray image data, and real business decision labels). Among them, the "real business decision label" is the final business result. For example, for the warehouse allocation task, it is the one-hot encoded vector of the correct warehouse area that the cargo should be assigned to.
[0087] 2. Constructing a unified model and loss function: All model components are integrated into a large computational graph. The loss function is directly defined on the final business output. For example, for the warehouse area distribution task, the unified loss function is the cross-entropy loss between the warehouse area allocation probability vector and the true warehouse area label vector in the final output of the model. :
[0088] in, It is the total number of reservoir areas. It is the first One true label (one-hot vector) It is the probability predicted by the model.
[0089] Joint Backpropagation: In each step of training, a training sample is input into the unified model for a complete forward propagation. The data flows through all stages, including text processing, image processing, multimodal fusion, and graph inference, ultimately yielding a prediction result at the output layer. Then, the loss value is calculated according to the end-to-end loss function defined above. Crucially, the gradient generated by this single loss value propagates backward from the output layer, traversing the entire computational graph, while simultaneously updating all trainable parameters of the graph neural network inference layer, multimodal validation module, text encoder, and image encoder. During joint backpropagation, gradient masking and selective backpropagation are implemented. The core logic is: instead of simply calculating the total loss and averaging it backpropagation, a gradient masking mechanism is introduced: 1. The model performs forward computation to obtain the final prediction and loss. 2. During backpropagation, the system analyzes which intermediate module's "judgment" played a crucial role in the final correct decision. 3. For intermediate modules whose outputs have errors but do not affect the final correct decision, a smaller gradient penalty or even gradient masking is applied.
[0090] For example, the text recognition module might identify "COL" as "CEL" (90% confidence), but the image module provides strong evidence of a "cold storage" entry. Meanwhile, the knowledge graph inference module fails to find a match based on the incorrect code "CEL," ultimately correctly predicting the "cold storage" entry by combining all the information. In this case, the loss is minimal. During backpropagation, the strong image evidence and the error-correcting capabilities of the knowledge graph "protect" the system. Gradients primarily flow to the image encoder and knowledge graph to reinforce their capabilities. For the text recognition module, because its errors are corrected by subsequent modules and do not lead to a final error, the gradient signal it receives is weak. This is equivalent to telling the text model: "You made a mistake this time, but it's okay, your teammates are there to cover for you. However, you should still learn in the right direction." This end-to-end joint training mechanism yields significant benefits. It enables all modules in the system to work collaboratively, sharing responsibility for the ultimate business objective. For example, during backpropagation, if a small modification to the multimodal verification module (even if it performs worse on a local verification task) provides more discriminative features to the subsequent graph neural network, thereby reducing the final library allocation error rate, then this modification will be encouraged by gradient descent. Conversely, even if a module performs perfectly on its local task, its parameters will not be effectively optimized if the features it extracts are unhelpful or even misleading to the final decision.
[0091] Therefore, by employing this end-to-end joint optimization strategy oriented towards the final business decision loss for unified training, it can be ensured that all AI models within the system learn the most discriminative features for the final business decision, rather than merely optimizing their own intermediate task metrics. This makes the entire system act as a unified intelligent agent responsible for the final business outcome, greatly improving the model's generalization ability and final decision accuracy in real-world, complex business scenarios. The model trained according to this embodiment, when deployed into the system framework described in Embodiment 1 or Embodiment 2, inherently possesses characteristics oriented towards global business optimization.
[0092] In summary, compared with the prior art, this application has the following beneficial effects: First, this application integrates data governance, multimodal verification, and intelligent decision-making into a unified system and methodology, streamlining key processes such as data processing, security checks, and data routing. This achieves end-to-end automation and collaborative optimization, improving overall efficiency and synergy. Second, based on precise business rules and algorithmic models, this application automates decision-making and automatically verifies "item-goods consistency," significantly reducing allocation errors that may result from manual operations and effectively identifying security risks such as false reporting and misreporting of product names, thereby lowering the error rate and enhancing security. Third, by establishing a weight-based data governance mechanism, this application automatically resolves conflicts between multi-source heterogeneous data, ensuring the accuracy, consistency, and authority of business data, providing a reliable data foundation for upper-level intelligent decision-making. Finally, by introducing advanced algorithmic models such as graph neural networks and reinforcement learning, this application makes business rules and system weights learnable and adaptive, enhancing the system's flexibility and robustness in responding to changes in business processes and the data environment, enabling continuous self-optimization.
[0093] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.
[0094] It should be noted that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means at least two.
[0095] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0096] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0097] Those skilled in the art will understand that all or part of the steps of the methods implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0098] Furthermore, the functional units in the various embodiments of this invention can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0099] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0100] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. An intelligent processing method for air logistics, characterized in that, Includes the following steps: Receive multimodal air logistics data, which includes at least business message data and cargo image data; Perform data governance, which includes: configuring trust weights for different data sources, and when new data that conflicts with existing data is received, executing a preset data update strategy by comparing the trust weights of the sources of the new and old data; Multimodal data verification is performed by using a cross-attention model based on the Transformer architecture to interactively analyze the text description in the business message data and the visual features in the cargo image data to verify the consistency between the order and the cargo. Based on the governed and verified data, execute automated decisions and allocations for at least one logistics task.
2. The method according to claim 1, characterized in that, The configuration of trust weights for different data sources specifically includes: Configure corresponding trust weights for the data source based on the enterprise type, data type, or data item from which the data originates.
3. The method according to claim 1, characterized in that, The multimodal data verification process also includes: Based on the results of the interaction analysis, a consistency score is calculated between the text and image content to quantify the consistency of the goods.
4. The method according to claim 1, characterized in that, The logistics operations include at least one of the following: automatic optimization allocation of security checkpoints and intelligent warehouse sorting after goods arrive at the port.
5. The method according to claim 4, characterized in that, The automatic optimization allocation of the security checkpoints specifically includes: Based on preset priority rules, availability rules, and queuing order rules, the goods to be processed are filtered and weighted to generate security check lane allocation instructions.
6. The method according to claim 4, characterized in that, The intelligent warehouse sorting and distribution of goods after they arrive at the port specifically includes: Based on the preset mapping relationship and priority between special cargo codes and special cargo operation codes, destination warehouse areas are assigned to the goods to be processed.
7. The method according to claim 4, characterized in that, The automated decision-making and allocation are implemented through a graph neural network running on a knowledge graph; The knowledge graph defines entity nodes, including goods, special goods codes, operation codes, and warehouse areas, as well as the relationship edges between entity nodes. The graph neural network realizes reasoning and decision-making on business rules by performing message passing and aggregation on the knowledge graph.
8. The method according to claim 1, characterized in that, The method further includes: A reinforcement learning mechanism is introduced, using the decision-making effect of logistics operations as a reward signal to dynamically adjust the trust weight or the business rule weight on which the decision allocation is based.
9. An intelligent processing system for air logistics, characterized in that, include: The data receiving module is used to receive multimodal air logistics data, which includes at least business message data and cargo image data; The data governance module is used to perform data governance. The data governance module is configured to: configure trust weights for different data sources, and when new data that conflicts with existing data is received, execute a preset data update strategy by comparing the trust weights of the sources of the new and old data. A multimodal verification module is used to perform multimodal data verification. The multimodal verification module has a built-in cross-attention model based on the Transformer architecture, which is used to perform interactive analysis on the text description in the business message data and the visual features in the cargo image data to verify the consistency between the order and the cargo. The decision execution module is used to automate the decision-making and allocation of at least one logistics operation based on the governed and verified data.
10. The system according to claim 9, characterized in that, The cross-attention model in the multimodal verification module and other models involved in decision-making in the decision execution module are uniformly trained using an end-to-end joint optimization strategy oriented towards the final business decision loss.
Citation Information
Patent Citations
Oil unloading and transportation method and oil unloading and transportation device
CN119848297A
Commercial and trade circulation supply chain optimization method based on multi-modal data fusion
CN120218364A
Ship identity multi-modal verification method and system based on computer vision and deep learning
CN120372543A
Multi-modal hierarchical feature fusion and decision-making method, device, equipment and medium
CN120951246A
Method and device for identifying logistics attributes, electronic equipment and storage medium
CN121051571A