A freight service verification method and system fusing multi-dimensional data
By creating a freight business risk identification model on the online freight platform and using ETL tools and rule engines for automated verification of multi-dimensional data, the problems of data silos and lag in traditional verification methods are solved, achieving efficient and accurate risk identification and management.
Patent Information
- Application Number
- CN202511029670.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Online freight platforms face risks such as business fraud, false tracking, and mismatched fund flows. Traditional manual verification methods suffer from data silos, lag, and limited verification dimensions, resulting in low verification efficiency and insufficient accuracy.
A freight business verification method integrating multi-dimensional data is adopted. By creating a freight business risk identification model, using ETL tools to obtain multi-dimensional data, and combining a rule engine and a pre-trained model to perform pre-event, in-event, and post-event verification, an end-to-end risk management closed loop is constructed to achieve automated verification.
It significantly improves the accuracy and timeliness of freight business verification, reduces manual intervention, enhances adaptability to changes in the business environment, reduces risk exposure, and improves operational efficiency and security.
Smart Images

Figure CN120542940B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and freight information verification, and in particular to a freight business verification method and system that integrates multi-dimensional data. Background Technology
[0002] With the rapid development of information technology, especially the deep integration and popularization of technologies such as the Internet, big data, the Internet of Things (IoT), and cloud computing, the traditional road freight transportation industry is undergoing a profound digital and intelligent transformation. Against this backdrop, network freight platforms have emerged and rapidly developed into a core component of the modern logistics system.
[0003] Online freight platforms leverage the powerful information integration and resource allocation capabilities of internet platforms to efficiently connect and optimize the allocation of scattered transportation resources (such as carriers and drivers). Acting as the legal entity of the "carrier," these platforms directly sign transportation service contracts with shippers and assume overall transportation responsibility. However, they do not directly participate in actual transportation operations; instead, they delegate specific cargo transportation tasks to qualified actual carriers connected to the platform. This innovative model significantly improves the transparency of the logistics chain, the efficiency of transportation resource utilization, and reduces transaction costs caused by information asymmetry, providing crucial support for promoting cost reduction, efficiency improvement, and high-quality development in the logistics industry.
[0004] However, while achieving efficient operation and value creation, the unique "asset-light, platform-based" operation model of online freight platforms also introduces significant challenges and risks. Among them, the "risk of business fraud" is particularly prominent, directly threatening the platform's stable operation, integrity system construction, data value mining, and compliance requirements. This risk manifests itself in the following ways: 1. Fictitious waybill risk: Platform users (such as some shippers or actual carriers) may generate false waybill data by forging non-existent cargo transportation needs or transaction information; their motives may include, but are not limited to: obtaining platform subsidies or freight differences, artificially inflating transaction volume to improve platform ranking or obtain financial support (such as waybill financing), and evading tax supervision. 2. Fake trajectory risk: In order to cover up non-compliant operations (such as detours, abnormal stops, transporting prohibited goods), or to cooperate with the need for fictitious waybills, actual carriers may use technical means such as GPS signal simulators, deliberately interfering with vehicle positioning equipment, and tampering with positioning data in the background to forge or modify the vehicle's transportation trajectory; fake trajectories seriously undermine the visualization and traceability of the transportation process, rendering supervision and risk control measures ineffective. 3. Risk of mismatch between fund flow records: There are inexplicable discrepancies between the transaction amounts and payment settlement information recorded by the platform and the actual fund flow (such as bank statements and third-party payment statements); this may be due to falsified waybill information, offline transactions outside the platform's supervision, off-book fund circulation, etc.
[0005] The aforementioned behaviors not only erode the trust foundation and market fairness of online freight platforms, leading to platform losses and reduced user stickiness, but also potentially trigger serious compliance risks, impacting the construction of a healthy industry ecosystem. Therefore, verification of freight operations is necessary. Traditionally, verification of freight operations on online freight platforms has been conducted manually, but this method has the following drawbacks:
[0006] 1. Data silos and barriers to collaborative verification:
[0007] The freight business process of online freight platforms involves multiple stages and a large amount of heterogeneous data, such as: transportation contracts and cargo information on the shipper's side; order information and user identity information on the platform's side; vehicle / driver qualifications and waybill status on the actual carrier's side; real-time vehicle location and trajectory on the IoT side; and fund receipt and payment records and invoice information on the payment and settlement side. Under the traditional architecture, this key data is usually scattered and stored in different subsystems within the platform (such as contract management system, order system, trajectory system, and financial settlement system) or in external third-party systems (such as payment gateways and tax systems). These systems often lack unified data interface standards, real-time and efficient data sharing mechanisms, and strong access control. The efficiency of manually pulling, comparing, and analyzing cross-system data is extremely low, and there are even technical obstacles. This physical dispersion and logical isolation of data creates a serious "data silo" effect, making it difficult for verification personnel to implement joint verification and cross-verification. As a result, the partial information recorded by a single system can be easily used by counterfeiters, making it impossible to identify complex fraudulent behaviors involving cross-stage and cross-system collaboration.
[0008] 2. Severe lag leads to risk control failure:
[0009] Manual verification processes are typically complex and lengthy, relying on sampling inspections, report analysis, or post-event reporting, resulting in significant delays. Verification often occurs after problematic waybills are completed and transactions are settled, sometimes even days or weeks later. This "post-event audit" model implies that risks have already materialized: problematic waybills have been completed and included in business volume, and mismatched fund flows may have been transferred or concealed. Online freight platforms respond passively after problems occur, akin to "closing the barn door after the horse has bolted." While they can hold those responsible accountable and punish those who have discovered problems, they cannot conduct real-time monitoring and proactive interception at critical stages of risk (such as waybill generation, abnormal tracking points, and the eve of payment). This provides ample "time windows" for fraudulent activities to evade inspection, significantly reducing the preventative effectiveness of risk control measures.
[0010] 3. Insufficient verification depth due to a single verification dimension:
[0011] Constrained by the limitations of manual verification and data silos, traditional methods tend to focus on verifying the authenticity of single-point data. For example, a verifier may only be able to check whether the uploaded electronic waybill format is standardized (whether there are obvious signs of forgery), or to perform a simple location comparison of individual vehicle trajectory points. This verification model based on a single data stream has obvious loopholes: it can verify the superficial authenticity of a certain link (such as the waybill), but it cannot effectively assess the correlation and consistency of this data with other core data in the entire business process (such as the signing status of the corresponding contract, the continuity and speed of the waybill trajectory points, and the matching payment amount and time). Forgers can take advantage of this weakness of "point-based verification" to carefully forge data in specific links so that it "passes" under a single-dimensional check, resulting in superficial verification results that cannot penetrate to reveal the deep, multi-linked chain of forgery.
[0012] In summary, while improving logistics efficiency, online freight platforms also face severe and increasingly complex risks of business fraud. Traditional verification methods, which rely on manual operation, are limited to post-event audits, and are constrained by data silos and single verification dimensions, are inherently inefficient, lagging, and superficial, and can no longer meet the high governance requirements and massive business volume processing needs of online freight platforms. Therefore, how to provide a freight business verification method and system that integrates multi-dimensional data to improve the accuracy and timeliness of freight business verification has become an urgent technical problem to be solved. Summary of the Invention
[0013] The technical problem to be solved by the present invention is to provide a method and system for verifying freight business by integrating multi-dimensional data, so as to improve the accuracy and timeliness of freight business verification.
[0014] In a first aspect, the present invention provides a method for verifying freight business by integrating multi-dimensional data, comprising the following steps:
[0015] Step S1: Create a freight business risk identification model and set the loss function of the freight business risk identification model;
[0016] Step S2: Obtain a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data; preprocess and label each of the historical freight business data to construct a dataset.
[0017] Step S3: Train the freight business risk identification model using the dataset and loss function, and deploy the trained freight business risk identification model to the server;
[0018] Step S4: The server sets up a business verification rule set containing several verification rules;
[0019] Step S5: The server uses an ETL tool to obtain real-time freight business data, including real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time bill data, from the contract system, business system, tracking system, funds system, and bill system.
[0020] Step S6: The server uses the rule engine to call the business verification rule set to perform pre-verification of the real-time freight business data;
[0021] Step S7: The server performs in-process verification of the real-time freight business data using the deployed freight business risk identification model.
[0022] Step S8: The server constructs an incremental dataset based on the real-time freight business data, and performs post-event optimization of the freight business risk identification model using the incremental dataset.
[0023] Secondly, this invention provides a freight business verification system that integrates multi-dimensional data, comprising the following modules:
[0024] The freight business risk identification model creation module is used to create a freight business risk identification model and set the loss function of the freight business risk identification model.
[0025] The dataset construction module is used to acquire a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data. The dataset is constructed after preprocessing and labeling each of the historical freight business data.
[0026] The freight business risk identification model training module is used to train the freight business risk identification model using the dataset and loss function, and to deploy the trained freight business risk identification model to the server.
[0027] The business verification rule set setting module is used by the server to set a business verification rule set containing several verification rules.
[0028] The multidimensional data acquisition module is used by the server to acquire real-time freight business data, including real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time invoice data, from the contract system, business system, trajectory system, financial system, and invoice system through ETL tools.
[0029] The pre-verification module is used by the server to call the business verification rule set through the rule engine to perform pre-verification on the real-time freight business data.
[0030] The in-process verification module is used by the server to perform in-process verification on the real-time freight business data through the deployed freight business risk identification model;
[0031] The post-event optimization module is used by the server to construct an incremental dataset based on the real-time freight business data, and to perform post-event optimization on the freight business risk identification model using the incremental dataset.
[0032] The advantages of this invention are:
[0033] 1. A freight business risk identification model is created, and its loss function is defined. Then, a large dataset of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical invoice data, is acquired to construct the dataset. The freight business risk identification model is trained using this dataset and the loss function. The trained model is then deployed to a server, which is configured with a business verification rule set containing several verification rules. Next, the server uses an ETL tool to obtain real-time freight business data, including real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time invoice data, from the contract system, business system, trajectory system, financial system, and invoice system. The rule engine then calls the business verification rule set to process the real-time freight business data. The system performs pre-verification of freight business data, in-process verification of real-time freight business data using a deployed freight business risk identification model, and post-optimization of the freight business risk identification model based on real-time freight business data. Specifically, it uses ETL tools to acquire multi-dimensional data (real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time invoice data) from different systems for freight business verification, overcoming traditional data silos and collaborative verification barriers, as well as the problem of traditionally singular verification dimensions. Pre-verification is performed by calling the business verification rule set through a rule engine, and in-process verification is performed through a pre-trained freight business risk identification model, overcoming the severe lag that traditionally leads to risk control failure, ultimately greatly improving the accuracy and timeliness of freight business verification.
[0034] 2. By integrating multi-dimensional data (including contract data, business data, trajectory data, financial data, and invoice data), the limitations of a single data source are avoided. This enables the freight business risk identification model to capture more comprehensive risk characteristics (such as contract fraud, abnormal funds, or trajectory deviation) during risk identification, thereby improving the accuracy of verification. For example, the fusion of trajectory data and financial data can effectively identify fraudulent transportation activities.
[0035] 3. Automated verification is achieved through a dual mechanism of rule engine and freight business risk identification model, which greatly reduces manual intervention. Specifically, pre-verification (based on rule engine) quickly screens common violations, while in-process verification (based on trained model) handles complex risk scenarios (such as abnormal fund flow) in real time, shortening the verification response time. At the same time, real-time freight business data is automatically collected using ETL tools, which greatly improves data processing efficiency.
[0036] 4. By constructing an incremental dataset and optimizing the freight business risk identification model, post-event closed-loop optimization was achieved. The freight business risk identification model can be automatically updated based on new data, enhancing its adaptability to changes in the business environment (such as new fraud methods). Combined with the defined loss function, the freight business risk identification model is continuously optimized, improving the long-term robustness of verification. This self-learning mechanism can resist "concept drift" (changes in data distribution), making the system more reliable and durable, unlike static solutions.
[0037] 5. An end-to-end risk management closed loop is constructed through three stages: pre-event, during-event, and post-event. Pre-event rule verification provides basic screening, during-event model verification handles complex logic, and post-event optimization ensures continuous improvement. It not only covers the entire lifecycle of freight business (such as order execution to fund settlement), but also supports real-time decision-making (such as during-event verification can automatically trigger risk alarms), effectively reducing overall risk exposure and improving business security.
[0038] 6. Large-scale data is processed through ETL tools and dataset construction methods, ensuring efficient data utilization; historical freight business data is used to train the freight business risk identification model, real-time freight business data supports verification, and combined with a configurable business verification rule set (verification rules can be added at any time), it is easy to extend to new scenarios (such as multimodal freight), optimize data resources, reduce redundant storage and computing overhead, effectively improve hardware resource utilization, and support multi-business type adaptation.
[0039] 7. By integrating multi-dimensional real-time and historical data (contracts, business, trajectory, funds, and invoices), and combining pre-verification by the rule engine with in-process verification by the AI model (freight business risk identification model), a dynamic risk control closed-loop system covering the entire freight business process has been constructed. Its core advantages are significantly improved in terms of the accuracy (avoiding blind spots from a single perspective by utilizing multi-dimensional data) and timeliness (real-time verification for rapid response). At the same time, through layered processing (preliminary rule screening + precise model judgment), computing resources are optimized and labor costs are reduced. Furthermore, relying on an incremental learning mechanism, the system continuously self-optimizes, enabling the model to dynamically evolve to adapt to business changes. Ultimately, the system achieves the goal of automated, highly robust, and scalable risk prevention and control, effectively reducing the risk of fraud and violations.
[0040] 8. The freight business risk identification model integrates multi-source heterogeneous data, including contract data, business data, trajectory data, financial data, and invoice data, through a feature extraction layer. This enables comprehensive monitoring of freight business risks, avoiding the limitations of a single data source and capturing a wider range of risk factors (such as contract defaults, abnormal funds, trajectory deviations, etc.). Dedicated modules process different types of data (e.g., Transformer for text, CNN / GRU for sequences), enhancing the practicality and market adaptability of the solution and meeting the freight industry's needs for comprehensive risk assessment.
[0041] 9. Contract data is processed through a Transformer bidirectional encoder to capture key semantic dependencies and improve the accuracy of text understanding; business and invoice data are processed through a multi-layer fully connected network to model linear and non-linear relationships and enhance the ability to analyze complex business indicators; trajectory and funding data are processed by combining GRU and CNN, with GRU capturing temporal dependencies and CNN extracting local features (such as transaction patterns) to ensure efficient identification of dynamic risks; this hybrid architecture (combining Transformer, CNN, GRU, etc.) can efficiently process structured, unstructured, and temporal data, reduce feature loss, and improve model robustness.
[0042] 10. The feature fusion layer dynamically calculates the feature weights of each feature through a multi-head attention module and performs weighted fusion. Then, it uses LSTM to model temporal dependencies and global average pooling to reduce dimensionality, ensuring that key risk features are given priority. At the same time, it processes the time series characteristics in the data (such as capital flow trends) to improve the fusion effect. The combination of multi-head attention and LSTM enables adaptive feature fusion, reduces redundant information, and improves computational efficiency.
[0043] 11. The prediction output layer adopts a dual-branch structure (risk item prediction and risk level prediction). It outputs independent risk item probabilities and risk level probabilities through Sigmoid and Softmax activation functions respectively, and jointly outputs a comprehensive report. This allows the freight business risk identification model to handle multiple related tasks simultaneously (such as identifying specific risk items and assessing the overall risk level), improving the fineness and accuracy of prediction. The multi-task learning framework reduces model complexity by sharing features (whole graph feature vectors), and the loss function design (weighted binary cross-entropy and classification cross-entropy) balances the weights of different tasks, optimizes the training process, effectively improves the generalization ability of the freight business risk identification model, and makes the output business risk identification report more comprehensive.
[0044] 12. The loss function is based on a weighted sum of risk term loss (binary cross-entropy) and risk level loss (classification cross-entropy). This allows for dynamic adjustment of the priority of different tasks during training (such as focusing more on high-risk levels or specific risk terms), improving the convergence speed and stability of the freight business risk identification model. This weighted loss mechanism is innovative and can effectively handle the imbalance problem in multi-task learning (such as the low frequency of some risk terms), reduce the risk of overfitting, and improve the reliability of the freight business risk identification model in real-world scenarios.
[0045] 13. Modular design (such as modular data processing in the feature extraction layer) and feature integration (global average pooling) reduce computational complexity and improve inference speed. At the same time, the model structure is easy to extend, for example, by adding new data sources (such as weather data) or adjusting modules (such as replacing GRU with Transformer) to adapt to different freight scenarios. While maintaining high performance, it reduces hardware resource requirements, is suitable for edge computing or cloud deployment, and its scalability enhances long-term competitiveness.
[0046] 14. By integrating five types of heterogeneous data—contracts, business, trajectory, funds, and invoices—and using targeted neural network modules (such as Transformer for extracting contract semantics, CNN+GRU for capturing fund sequences, and fully connected layers for modeling business indicators) to deeply mine multi-dimensional risk features, and utilizing a multi-head attention mechanism for adaptive weighted feature fusion, combined with LSTM for modeling temporal dependencies and global average pooling for compressing key information; on this basis, a comprehensive risk report is generated through a dual-task collaborative prediction mechanism (Sigmoid outputs the probability of independent risk items + Softmax classifies risk levels). Its multi-task weighted loss function (binary cross-entropy + classification cross-entropy) effectively balances the training objectives, ensuring high-precision risk identification while significantly improving the model's generalization ability and computational efficiency in complex freight scenarios, ultimately achieving end-to-end automated risk control decision support.
[0047] 15. By employing multi-step preprocessing (including deduplication, missing value handling, error correction, data standardization, data normalization, and data encoding), the original data is systematically cleaned and transformed, reducing the impact of noise, inconsistency, and bias, and ensuring the consistency and operability of the dataset. For example, standardization and normalization processes enable data from different sources (such as speed information and monetary information) to be compared, thereby improving the accuracy of subsequent risk labeling and model training.
[0048] 16. By integrating multi-dimensional historical freight data (including contract, business, trajectory, fund and bill data) and adopting systematic preprocessing (such as deduplication and standardization) to improve data quality, and combining risk labeling to build a high-quality dataset, and then innovatively applying adversarial networks to expand the sample size, the comprehensiveness, accuracy and robustness of freight risk identification are significantly improved. This effectively solves the prediction bias problem caused by single data or insufficient samples in traditional models, while realizing automated and efficient processing, providing scalable technical support for logistics risk management.
[0049] 17. By using stratified sampling to divide the dataset into training, validation, and test sets in an 8:1:1 ratio, the data distribution was balanced and representative, sampling bias was reduced, and the model's generalization ability (i.e., the model's performance on unseen data) was improved, thereby reducing the misjudgment rate in freight business risk identification. The 8:1:1 ratio also optimized resource utilization, avoided overfitting or underfitting problems, and enhanced the reliability of the model.
[0050] 18. During the training phase, by continuously optimizing hyperparameters (such as learning rate and batch size) until the loss value of the loss function is less than the preset threshold, the model achieves rapid convergence and efficient training, reducing training time and computational resource consumption, while ensuring that the model achieves high accuracy in the freight risk identification task; the iterative optimization mechanism (such as dynamic adjustment based on loss value) improves the adaptability of training, enabling the model to better handle noise and complexity in freight data.
[0051] 19. The validation set is used to independently evaluate the model's accuracy and sets an accuracy threshold as a pass standard. If the threshold is not met, the training set is expanded to continue training, introducing a feedback loop. This improves the model by increasing data diversity (such as adding more freight cases), avoiding "stagnation" in the validation phase, improving the model's robustness, ensuring high identification accuracy in variable freight environments (such as risk patterns in different regions or seasons), and reducing missed or false alarms in business risks.
[0052] 20. The test set is used for the final evaluation of the model, and a confidence threshold check (such as the confidence level of the model's prediction) is introduced, which increases the credibility of risk identification. If the confidence level is not up to standard, the training set is expanded and the model is retrained to form an iterative improvement, ensuring the rigorous verification of the model before real deployment, reducing the probability of misjudging high-risk events in freight business (such as cargo loss or fraud), and providing quantitative evaluation through the confidence index, which is convenient for monitoring and optimization.
[0053] 21. By optimizing dataset partitioning through stratified sampling and combining iterative training with a dual threshold verification mechanism, the accuracy, robustness, and reliability of the freight business risk identification model were significantly improved. Containerization technology enabled efficient deployment and elastic scaling, greatly reducing operation and maintenance costs. Simultaneously, through terminal authentication and structured storage of rule sets, the customizability and real-time decision-making capabilities of business rules were enhanced while ensuring system security. This formed a closed-loop optimization system from data preprocessing to model deployment, effectively solving the core pain points of low risk identification efficiency, high misjudgment rate, and poor system scalability in freight scenarios.
[0054] 22. By using ETL tools to automatically integrate real-time freight business data from five major systems—contracts, business, tracking, funds, and invoices—and through standardized processing such as cleaning, formatting, and aggregation, a unified, high-quality data warehouse was built. This completely broke down information silos, providing real-time, accurate, and multi-dimensional data support for business decisions, and significantly improving operational efficiency, risk management capabilities, and the company's core competitiveness.
[0055] 23. By automatically calling the rule engine through the API interface, the processing of freight business events is triggered in real time, avoiding the delays of traditional manual verification. The rule engine reads the verification rule set from the specified path and performs multi-dimensional checks on the data (such as data format, integrity, logical consistency, vehicle qualifications, personnel qualifications, vehicle trajectory, waybill timeliness and invoice information), ensuring that the verification process is efficient and automated, greatly reducing business downtime and improving the overall throughput of the freight process (for example, it can quickly generate reports when failure occurs), thereby improving operational efficiency and reducing labor costs.
[0056] 24. A multi-layered encryption mechanism (including SM9 algorithm, RC6 algorithm, character shifting, and splitting / swapping) provides enhanced protection for the data. Specifically: hash values are used to ensure data integrity and prevent verification reports and timestamps from being tampered with; the SM9 algorithm provides post-quantum level asymmetric encryption, with first-level encryption ensuring the confidentiality of sensitive information; splitting / swapping (splitting and swapping in a 7:3 ratio) increases the difficulty of cracking; the RC6 algorithm further encrypts the data and introduces perturbations through a 5-bit rightward cyclic shift of characters, making the encryption process more complex; this multi-step process significantly improves the data's resistance to attacks and is suitable for freight data environments with high security requirements; finally, the data is uploaded to the blockchain, utilizing its distributed ledger characteristics to ensure data immutability and traceability, enhancing the overall system's non-repudiation capabilities.
[0057] 25. By using a rules engine to automatically verify freight business data in real time (covering key dimensions such as data format, integrity, logical consistency, qualification and timeliness), the efficiency and accuracy of verification are greatly improved, avoiding human delays. At the same time, a multi-layer dynamic encryption mechanism (SM9 algorithm and RC6 algorithm combined with data segmentation, swapping and displacement operations) and blockchain notarization are adopted to ensure that the data is tamper-proof and traceable throughout the process. The verification failure report is pushed in real time through the TLS protocol to realize the immediate warning and closed-loop management of business risks. Ultimately, while ensuring freight safety and compliance, the system significantly reduces operating costs and enhances system reliability.
[0058] 26. By using hardware acceleration technologies (such as GPUs or FPGAs) for reasoning in the freight business risk identification model, the computational latency of real-time data processing is significantly reduced, enabling in-process verification and avoiding the time lag of traditional batch processing methods. It has a high response speed (millisecond-level reasoning) and can promptly identify and handle high-risk freight events (such as cargo damage or fraud), thereby reducing business losses. This real-time capability enhances the practicality of the system, making it particularly suitable for high-frequency, dynamic freight scenarios.
[0059] 27. By adopting multi-layered encryption (including XTEA, IDEA algorithms and custom character swapping) and combining hash value verification (such as the second hash value based on the report and verification time), a strong data security guarantee is provided: (1) Algorithm fusion enhances security: XTEA algorithm is used for initial encryption to ensure basic data security; then, by swapping characters (such as swapping 7 with B, 8 with 9, and 10 with A in hexadecimal data), a confusion layer is introduced to increase the difficulty of reverse engineering and resist pattern recognition attacks; IDEA algorithm performs deep encryption to form a "three-layer" protection; (2) Data integrity guarantee: Hash values are embedded in the hash value calculation and encryption process to ensure that the data is not tampered with during transmission and storage, effectively responding to the risk of man-in-the-middle attacks or data tampering; (3) TLS protocol ensures real-time communication security: when the risk level is higher than the preset value, a report is pushed through the TLS protocol to prevent the data from being eavesdropped or intercepted during transmission; (4) Blockchain technology guarantees permanent immutability: finally, the encrypted report is uploaded to the blockchain, and the characteristics of the distributed ledger are used to provide traceability and auditing capabilities.
[0060] 28. Through hardware acceleration, millisecond-level real-time verification of freight business risks is achieved. Combined with a unique multi-layer encryption mechanism (integrating XTEA / IDEA algorithms and character swapping obfuscation technology) and blockchain evidence storage, while ensuring the security of real-time business risk report transmission and the immutability of data, the closed loop is optimized by using an incremental dataset-driven freight business risk identification model. This continuously improves the accuracy of risk identification and forms a closed-loop intelligent risk control system from real-time early warning to dynamic evolution, significantly reducing the operational risks and manual intervention costs of freight business.
[0061] 29. By integrating multi-dimensional heterogeneous data (contracts, business, trajectory, funds, and invoices), a freight business risk identification model is constructed. Combined with a rule engine, it achieves dual protection of pre-verification and in-process prediction. Incremental data is used to optimize the freight business risk identification model in real time, forming a closed-loop risk control system throughout the entire process. The innovation is reflected in the use of Transformer, GRU, multi-head attention and other modules to achieve deep feature extraction and fusion, improve the accuracy of risk identification, and ensure system efficiency, scalability and data security through hardware acceleration, containerized deployment and blockchain encryption, which significantly reduces freight business risks and improves verification efficiency. Attached Figure Description
[0062] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0063] Figure 1 This is a flowchart of a freight business verification method that integrates multi-dimensional data according to the present invention.
[0064] Figure 2 This is a schematic diagram of the structure of a freight business verification system that integrates multi-dimensional data according to the present invention. Detailed Implementation
[0065] The overall approach of the technical solution in this application is as follows: Multi-dimensional data (real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time invoice data) is obtained from different systems using ETL tools for freight business verification, overcoming traditional data silos and collaborative verification barriers, and overcoming the problem of traditionally singular verification dimensions; Pre-verification is performed by calling the business verification rule set through a rule engine, and in-process verification is performed through a pre-trained freight business risk identification model, overcoming the traditional serious lag that leads to risk control failure, thereby improving the accuracy and timeliness of freight business verification.
[0066] Please refer to Figures 1 to 2 As shown, a preferred embodiment of the freight business verification method integrating multi-dimensional data according to the present invention includes the following steps:
[0067] Step S1: Create a freight business risk identification model and set the loss function of the freight business risk identification model;
[0068] Step S2: Obtain a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data; preprocess and label each of the historical freight business data to construct a dataset.
[0069] By integrating multidimensional data (including contract data, business data, trajectory data, financial data, and invoice data), the limitations of a single data source are avoided. This enables the freight business risk identification model to capture more comprehensive risk characteristics (such as contract fraud, abnormal funds, or trajectory deviation) during risk identification, thereby improving the accuracy of verification. For example, the fusion of trajectory data and financial data can effectively identify fraudulent transportation activities.
[0070] Step S3: Train the freight business risk identification model using the dataset and loss function, and deploy the trained freight business risk identification model to the server;
[0071] Step S4: The server sets up a business verification rule set containing several verification rules;
[0072] Step S5: The server uses an ETL tool to obtain real-time freight business data, including real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time bill data, from the contract system, business system, tracking system, funds system, and bill system.
[0073] Step S6: The server uses the rule engine to call the business verification rule set to perform pre-verification of the real-time freight business data;
[0074] Step S7: The server performs in-process verification of the real-time freight business data using the deployed freight business risk identification model.
[0075] The automated verification is achieved through a dual mechanism of rule engine and freight business risk identification model, which greatly reduces manual intervention. Specifically, pre-verification (based on rule engine) quickly screens common violations, while in-process verification (based on trained model) handles complex risk scenarios (such as abnormal fund flow) in real time, shortening the verification response time. At the same time, real-time freight business data is automatically collected using ETL tools, which greatly improves data processing efficiency.
[0076] Step S8: The server constructs an incremental dataset based on the real-time freight business data, and performs post-event optimization of the freight business risk identification model using the incremental dataset;
[0077] By constructing an incremental dataset and optimizing the freight business risk identification model, post-event closed-loop optimization was achieved. The freight business risk identification model can be automatically updated based on new data, enhancing its adaptability to changes in the business environment (such as new fraud methods). Combined with the defined loss function, the freight business risk identification model is continuously optimized, improving the long-term robustness of verification. This self-learning mechanism can resist "concept drift" (changes in data distribution), making the system more reliable and durable, unlike static solutions.
[0078] By constructing an end-to-end risk management closed loop through three stages: pre-event, during-event, and post-event, the pre-event rule verification provides basic screening, the during-event model verification handles complex logic, and the post-event optimization ensures continuous improvement. It not only covers the entire lifecycle of freight business (such as order execution to fund settlement), but also supports real-time decision-making (such as during-event verification can automatically trigger risk alarms), effectively reducing overall risk exposure and improving business security.
[0079] In step S1, the freight business risk identification model is constructed based on a feature extraction layer, a feature fusion layer, and a prediction output layer.
[0080] The feature extraction layer is constructed based on a contract data processing module, a business data processing module, a trajectory data processing module, a capital data processing module, and a bill data processing module. The contract data processing module extracts key semantic dependencies from contract data using a bidirectional encoder built with a Transformer to obtain contract risk features. The business data processing module models linear and nonlinear relationships between business indicators in business data using a multi-layered first fully connected network to obtain business operation features. The trajectory data processing module extracts trajectory risk features from trajectory data using a first gated recurrent unit. The capital data processing module captures local transaction features from capital data using a one-dimensional convolutional neural network and captures sequence dependency features of capital data using a second gated recurrent unit, outputting capital flow features based on the local transaction features and sequence dependency features. The bill data processing module extracts bill integrity features from bill data using a multi-layered second fully connected network.
[0081] The freight business risk identification model integrates multi-source heterogeneous data, including contract data, business data, trajectory data, financial data, and invoice data, through a feature extraction layer. This enables comprehensive monitoring of freight business risks, avoiding the limitations of a single data source and capturing a wider range of risk factors (such as contract defaults, abnormal funding, and trajectory deviations). Dedicated modules process different types of data (e.g., Transformer for text, CNN / GRU for sequences), enhancing the practicality and market adaptability of the solution and meeting the freight industry's needs for comprehensive risk assessment.
[0082] By processing contract data through a Transformer bidirectional encoder, key semantic dependencies are captured, improving the accuracy of text understanding. By processing business and invoice data through a multi-layer fully connected network, linear and non-linear relationships are modeled, enhancing the ability to analyze complex business indicators. By combining GRU and CNN to process trajectory and funding data, GRU captures temporal dependencies, and CNN extracts local features (such as transaction patterns), ensuring efficient identification of dynamic risks. This hybrid architecture (combining Transformer, CNN, GRU, etc.) can efficiently process structured, unstructured, and temporal data, reduce feature loss, and improve model robustness.
[0083] The feature fusion layer is constructed based on a multi-head attention module, a long short-term memory module, and a feature integration module. The multi-head attention module is used to calculate the feature weights of contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features through multi-head self-attention units. Based on the feature weights, the contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features are fused to obtain fused features. The long short-term memory module is used to model the temporal dependencies of the fused features through LSTM units to obtain serialized features. The feature integration module is used to perform global average pooling on the serialized features to obtain the whole graph feature vector.
[0084] The feature fusion layer dynamically calculates the feature weights of each feature through a multi-head attention module and performs weighted fusion. Then, it uses LSTM to model temporal dependencies and global average pooling to reduce dimensionality, ensuring that key risk features are given priority. At the same time, it processes the time series characteristics in the data (such as capital flow trends) to improve the fusion effect. The combination of multi-head attention and LSTM enables adaptive feature fusion, reduces redundant information, and improves computational efficiency.
[0085] The prediction output layer is constructed based on a risk item prediction branch, a risk level prediction branch, and a joint output module. The risk item prediction branch is used to infer the feature vector of the entire graph through a first fully connected layer and a sigmoid activation function to obtain the independent probabilities of different risk items (such as fraud, delay, default, etc.), supporting multi-label prediction. The risk level prediction branch is used to infer the feature vector of the entire graph through a second fully connected layer and a Softmax activation function to obtain the category probabilities of risk levels (low, medium, high). The joint output module is used to output a business risk identification report including risk items and risk levels based on the independent probabilities and category probabilities.
[0086] The prediction output layer adopts a dual-branch structure (risk item prediction and risk level prediction), using Sigmoid and Softmax activation functions to output independent risk item probabilities and risk level probabilities respectively, and jointly outputting a comprehensive report. This allows the freight business risk identification model to handle multiple related tasks simultaneously (such as identifying specific risk items and assessing the overall risk level), improving the fineness and accuracy of prediction. The multi-task learning framework reduces model complexity by sharing features (whole graph feature vectors), and the loss function design (weighted binary cross-entropy and classification cross-entropy) balances the weights of different tasks, optimizes the training process, effectively improves the generalization ability of the freight business risk identification model, and makes the output business risk identification report more comprehensive.
[0087] In step S1, the loss function is constructed based on the weighted sum of risk item loss and risk level loss; the risk item loss adopts binary cross-entropy loss; the risk level loss adopts classification cross-entropy loss.
[0088] The loss function is based on a weighted sum of risk term loss (binary cross-entropy) and risk level loss (classification cross-entropy), which allows for dynamic adjustment of the priority of different tasks during training (such as focusing more on high-risk levels or specific risk terms), improving the convergence speed and stability of the freight business risk identification model. This weighted loss mechanism is innovative and can effectively handle the imbalance problem in multi-task learning (such as the low frequency of some risk terms), reduce the risk of overfitting, and improve the reliability of the freight business risk identification model in real-world scenarios.
[0089] Step S2 specifically involves:
[0090] Acquire a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data;
[0091] The historical contract data includes at least the basic contract information, information on the cooperating parties, information on the goods, transportation service requirements, fee terms, liability for breach of contract, and dispute resolution methods;
[0092] The basic contract information includes at least the contract number, signing date, and contract validity period, used for contract classification, retrieval, and management to ensure timeliness and traceability. The information on the cooperating parties includes at least the shipper's name, carrier's name, address, contact information, and business license number, used to clarify the identities and responsibilities of both parties, providing a basis for subsequent business operations and dispute resolution. The cargo information includes at least the cargo's name, specifications, model, weight, volume, quantity, and value, used for rationally arranging transport vehicles, calculating freight costs, and ensuring cargo safety. The transportation service requirements include at least the origin, destination, transportation time requirements (e.g., latest pick-up time, latest delivery time), and mode of transport (road, rail, etc.). The terms and conditions for handling goods include: (1) waterway transportation and (2) loading and unloading requirements, to ensure that the transportation service can meet the specific needs of the customer; (2) fee terms, to include at least the method of calculating freight (e.g., by weight, volume or mileage), payment method (e.g., prepayment, cash on delivery, monthly settlement), payment time, and invoice issuance requirements, to clarify the financial responsibilities and settlement process of both parties; (3) liability for breach of contract, to stipulate corresponding liability and compensation methods for possible breaches of contract during the performance of the contract, such as delayed delivery, damage or loss of goods, to protect the legitimate rights and interests of both parties; and (4) dispute resolution, to stipulate the means of resolving contract disputes, such as negotiation, mediation, arbitration or litigation, as well as the applicable law and arbitration institution or court of jurisdiction, to provide a mechanism for resolving possible disputes.
[0093] The business data includes at least waybill information, order logs, online transaction logs, customer information, vehicle and driver information, and logistics information;
[0094] The waybill information includes at least the waybill number, consignor information, consignee information, cargo description (consistent with the cargo information in the contract), transportation route, transportation mileage, estimated transportation time, actual transportation time, and waybill status (e.g., ordered, picked up, in transit, delivered), used to track and manage the entire process of each transportation transaction; the order log records the time, operator, and reason for operations such as order generation, modification, and cancellation, facilitating the tracing of order changes and ensuring the transparency and auditability of business operations; the online transaction log includes at least records of user online freight platform activities such as login, browsing, querying, order placement, quotation, and transaction completion, used to reflect user transaction habits and platform transaction activity, providing a basis for data analysis and business optimization; the customer information includes, in addition to the consignor information involved in the contract... In addition to information on people and carriers, the system also includes customer registration information, credit ratings, historical transaction records, and preference settings. This helps online freight platforms better understand customer needs and provide personalized services, while also facilitating customer relationship management and risk assessment. The vehicle and driver information includes at least license plate number, vehicle type, approved load capacity, vehicle status (e.g., idle, in transit, under maintenance), driver's name, driver's contact information, driver's license information, professional qualification certificate information, driving experience, and credit record. This information is used to rationally allocate vehicle and driver resources and ensure transportation safety. The logistics information includes at least the real-time location of goods, transportation status (e.g., whether transportation is normal, whether any abnormalities have occurred), loading and unloading times, and transit information. By combining this information with trajectory data, the system enables full-process monitoring and visual management of the logistics process, improving logistics efficiency and customer satisfaction.
[0095] The trajectory data includes at least time information, geographical location information, driving route information, speed information, mileage information, and stop point information;
[0096] The time information refers to the timestamps of the vehicle at various locations, including departure time, time of passing through each node, and time of arrival at the destination. This time information allows for the calculation of vehicle speed, dwell time, and other parameters, thereby analyzing transportation efficiency and identifying any abnormal stops. The geographical location information refers to the real-time latitude and longitude coordinates of the vehicle during transportation, obtained through positioning devices (such as GPS and BeiDou). This information is used to determine the vehicle's precise location, map its trajectory, and enable dynamic tracking and monitoring of the transportation process. The driving path information refers to the actual route the vehicle travels, including road segments, intersections, and highway entrances / exits. Comparing this information with the planned transportation route allows for the determination of whether the vehicle is following the predetermined route and whether there are any abnormalities such as detours or deviations. The speed information... The vehicle speed information refers to the vehicle's speed over different time periods. Monitoring speed allows analysis of whether the vehicle is speeding and whether there are abnormal speed changes due to road conditions, traffic control, or other factors, providing a reference for assessing transportation safety and efficiency. The mileage information refers to the distance the vehicle travels during transportation, which can be used to calculate costs such as fuel consumption and vehicle wear and tear. Combined with information such as cargo weight, it can also assess transportation cost-effectiveness and provide data support for optimizing transportation plans. The stop point information refers to the vehicle's stopping locations and duration during transportation, including loading and unloading points, refueling stations, rest areas, etc. Analyzing stop points can help understand the vehicle's operating mode and the driver's driving habits, and also help identify potential risks, such as the risk of cargo theft due to prolonged stops.
[0097] The financial data includes at least fund flow information, account balance information, and settlement information;
[0098] The fund flow information includes records of every income and expenditure, such as freight charges, freight payments to the actual carrier, and platform service fees. It details the flow of funds, amount, transaction time, and transaction counterparty to ensure transparency and traceability. The account balance information reflects the fund account balances of both the shipper and carrier on the platform, indicating their financial status and available funds, providing a basis for fund settlement and financial management. The settlement information includes at least the settlement cycle, settlement method, settlement time, and settlement amount, clarifying the settlement rules and procedures to ensure timely and accurate settlement and protect the legitimate rights and interests of all parties.
[0099] The invoice data includes at least the invoice number, invoice date, invoice amount, and invoice type;
[0100] The historical freight business data are preprocessed, including at least deduplication, missing value handling, error correction, data standardization, data normalization, and data encoding. Risk items and risk levels are labeled on the preprocessed historical freight business data. A dataset is constructed using the labeled historical freight business data. The sample size of the dataset is expanded using an adversarial network.
[0101] By employing multi-step preprocessing (including deduplication, missing value handling, error correction, data standardization, data normalization, and data encoding), the raw data is systematically cleaned and transformed, reducing the impact of noise, inconsistencies, and biases, and ensuring the consistency and operability of the dataset. For example, standardization and normalization processes enable data from different sources (such as speed information and monetary information) to be compared, thereby improving the accuracy of subsequent risk labeling and model training.
[0102] By integrating multi-dimensional historical freight data (including contract, business, trajectory, financial and invoice data) and adopting systematic preprocessing (such as deduplication and standardization) to improve data quality, and combining risk labeling to build a high-quality dataset, and then innovatively applying adversarial networks to expand the sample size, the comprehensiveness, accuracy and robustness of freight risk identification are significantly improved. This effectively solves the prediction bias problem caused by single data or insufficient samples in traditional models, while realizing automated and efficient processing, providing scalable technical support for logistics risk management.
[0103] Step S3 specifically involves:
[0104] The dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1 using stratified sampling. The freight business risk identification model is trained using the training set. During the training process, the hyperparameters of the freight business risk identification model are continuously optimized until the loss value of the loss function is less than the preset loss threshold.
[0105] The trained freight business risk identification model is validated using the validation set to determine whether the identification accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded for further training; if yes, the validation passes, and:
[0106] The validated freight business risk identification model is tested using the test set to determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded for further training. If yes, the test passes, and the validated freight business risk identification model is deployed to the server using containerization technology.
[0107] By dividing the dataset into training, validation, and test sets in an 8:1:1 ratio using stratified sampling, the data distribution was balanced and representative, sampling bias was reduced, and the model's generalization ability (i.e., the model's performance on unseen data) was improved, thereby reducing the misjudgment rate in freight business risk identification. The 8:1:1 ratio also optimized resource utilization, avoided overfitting or underfitting problems, and enhanced the reliability of the model.
[0108] During the training phase, by continuously optimizing hyperparameters (such as learning rate and batch size) until the loss value of the loss function is less than a preset threshold, the model achieves rapid convergence and efficient training, reducing training time and computational resource consumption, while ensuring that the model achieves high accuracy in the freight risk identification task; the iterative optimization mechanism (such as dynamic adjustment based on loss value) improves the adaptability of training, enabling the model to better handle noise and complexity in freight data.
[0109] By optimizing dataset partitioning through stratified sampling and combining iterative training with a dual threshold verification mechanism, the accuracy, robustness, and reliability of the freight business risk identification model were significantly improved. Containerization technology enabled efficient deployment and elastic scaling, greatly reducing operation and maintenance costs. Simultaneously, through terminal authentication and structured storage of rule sets, the customizability and real-time decision-making capabilities of business rules were enhanced while ensuring system security. This formed a closed-loop optimization system from data preprocessing to model deployment, effectively solving the core pain points of low risk identification efficiency, high false positive rate, and poor system scalability in freight scenarios.
[0110] Step S4 specifically involves:
[0111] After authenticating the accessing mobile terminal, the server obtains several verification rules set by the mobile terminal, constructs a business verification rule set based on each of the verification rules, and stores the business verification rule set in a structured manner to a specified path.
[0112] Step S5 specifically involves:
[0113] The server uses an ETL tool to extract real-time freight business data from the contract system, business system, trajectory system, fund system, and bill system based on a preset extraction period. This data includes real-time contract data, real-time business data, real-time trajectory data, real-time fund data, and real-time bill data. The server then performs data transformation operations on each of the real-time freight business data, including at least data cleaning, data formatting, and data aggregation. Finally, the server loads the transformed real-time freight business data into a preset data warehouse.
[0114] Step S6 specifically involves:
[0115] The server triggers events based on freight business and calls the rule engine through the API interface. The rule engine reads the business verification rule set from the specified path and performs pre-verification on real-time freight business data based on the business verification rule set, including at least data format, data integrity, logical consistency, vehicle qualification, personnel qualification, vehicle trajectory, waybill timeliness, and invoice information, and generates a pre-verification report carrying whether the pre-verification was successful or failed.
[0116] When the pre-verification report indicates that the pre-verification has failed, the pre-verification report will be pushed to the pre-associated management terminal in real time via the TLS protocol.
[0117] The server obtains the pre-verification time, calculates the first hash value of the pre-verification report and the pre-verification time, encrypts the pre-verification report, the pre-verification time, and the first hash value into first-level encrypted data using the SM9 algorithm, divides the first-level encrypted data in a 7:3 ratio and swaps the order to obtain second-level encrypted data, encrypts the second-level encrypted data into third-level encrypted data using the RC6 algorithm, and shifts each character of the third-level encrypted data cyclically 5 bits to the right to obtain the pre-encrypted report, which is then uploaded to the blockchain.
[0118] By automatically calling the rule engine through the API interface, the processing of freight business events is triggered in real time, avoiding the delays of traditional manual verification. The rule engine reads the verification rule set from the specified path and performs multi-dimensional checks on the data (such as data format, integrity, logical consistency, vehicle qualifications, personnel qualifications, vehicle trajectory, waybill timeliness, and invoice information), ensuring that the verification process is efficient and automated, greatly reducing business downtime, improving the overall throughput of the freight process (for example, it can quickly generate reports when failure occurs), thereby improving operational efficiency and reducing labor costs.
[0119] A multi-layered encryption mechanism (including the SM9 algorithm, RC6 algorithm, character shifting, and splitting / swapping) provides enhanced protection for the data. Specifically: hash values are used to ensure data integrity and prevent verification reports and timestamps from being tampered with; the SM9 algorithm provides post-quantum level asymmetric encryption, with first-level encryption ensuring the confidentiality of sensitive information; splitting / swapping (splitting and swapping in a 7:3 ratio) increases the difficulty of cracking; the RC6 algorithm further encrypts the data and introduces perturbations through a 5-bit rightward cyclic shift of characters, making the encryption process more complex; this multi-step process significantly improves the data's resistance to attacks and is suitable for freight data environments with high security requirements; finally, the data is uploaded to the blockchain, leveraging its distributed ledger characteristics to ensure data immutability and traceability, enhancing the overall system's non-repudiation capabilities.
[0120] Step S7 specifically involves:
[0121] The server inputs the real-time freight business data into the deployed freight business risk identification model. The freight business risk identification model performs inference through hardware acceleration technology to obtain a real-time business risk identification report that includes real-time risk items and real-time risk levels, so as to verify the real-time freight business data in real time.
[0122] When the risk level is higher than the preset level, the real-time business risk identification report will be pushed to the pre-associated management terminal in real time via the TLS protocol;
[0123] The server obtains the in-process verification time, calculates the second hash value of the real-time business risk identification report and the in-process verification time, and encrypts the real-time business risk identification report, the in-process verification time, and the second hash value into a first layer of encrypted data using the XTEA algorithm. The first layer of encrypted data is converted into hexadecimal data, and the number 7 is swapped with the letter B, the number 8 is swapped with the number 9, and the number 10 is swapped with the letter A to obtain a second layer of encrypted data. The second layer of encrypted data is encrypted into a third layer of encrypted data using the IDEA algorithm. A random string of a specified length is added to a specified position in the third layer of encrypted data to obtain an in-process encrypted report, and the in-process encrypted report is uploaded to the blockchain.
[0124] By using hardware acceleration technologies (such as GPUs or FPGAs) for inference of the freight business risk identification model, the computational latency of real-time data processing is significantly reduced, enabling in-process verification and avoiding the time lag of traditional batch processing methods. It has a high response speed (millisecond-level inference) and can promptly identify and handle high-risk freight events (such as cargo damage or fraud), thereby reducing business losses. This real-time capability enhances the practicality of the system, making it particularly suitable for high-frequency, dynamic freight scenarios.
[0125] Step S8 specifically involves:
[0126] The server constructs an incremental dataset based on the real-time freight business data. This incremental dataset is labeled with real-time risk items, real-time risk levels, actual risk items, and actual risk levels. When the data volume of the incremental dataset exceeds a preset threshold, the freight business risk identification model is post-optimized using the incremental dataset. The real-time risk items and real-time risk levels are predicted data, while the actual risk items and actual risk levels are actual data.
[0127] A preferred embodiment of the freight business verification system integrating multi-dimensional data according to the present invention includes the following modules:
[0128] The freight business risk identification model creation module is used to create a freight business risk identification model and set the loss function of the freight business risk identification model.
[0129] The dataset construction module is used to acquire a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data. The dataset is constructed after preprocessing and labeling each of the historical freight business data.
[0130] By integrating multidimensional data (including contract data, business data, trajectory data, financial data, and invoice data), the limitations of a single data source are avoided. This enables the freight business risk identification model to capture more comprehensive risk characteristics (such as contract fraud, abnormal funds, or trajectory deviation) during risk identification, thereby improving the accuracy of verification. For example, the fusion of trajectory data and financial data can effectively identify fraudulent transportation activities.
[0131] The freight business risk identification model training module is used to train the freight business risk identification model using the dataset and loss function, and to deploy the trained freight business risk identification model to the server.
[0132] The business verification rule set setting module is used by the server to set a business verification rule set containing several verification rules.
[0133] The multidimensional data acquisition module is used by the server to acquire real-time freight business data, including real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time invoice data, from the contract system, business system, trajectory system, financial system, and invoice system through ETL tools.
[0134] The pre-verification module is used by the server to call the business verification rule set through the rule engine to perform pre-verification on the real-time freight business data.
[0135] The in-process verification module is used by the server to perform in-process verification on the real-time freight business data through the deployed freight business risk identification model;
[0136] The automated verification is achieved through a dual mechanism of rule engine and freight business risk identification model, which greatly reduces manual intervention. Specifically, pre-verification (based on rule engine) quickly screens common violations, while in-process verification (based on trained model) handles complex risk scenarios (such as abnormal fund flow) in real time, shortening the verification response time. At the same time, real-time freight business data is automatically collected using ETL tools, which greatly improves data processing efficiency.
[0137] The post-event optimization module is used by the server to construct an incremental dataset based on the real-time freight business data, and to perform post-event optimization on the freight business risk identification model using the incremental dataset.
[0138] By constructing an incremental dataset and optimizing the freight business risk identification model, post-event closed-loop optimization was achieved. The freight business risk identification model can be automatically updated based on new data, enhancing its adaptability to changes in the business environment (such as new fraud methods). Combined with the defined loss function, the freight business risk identification model is continuously optimized, improving the long-term robustness of verification. This self-learning mechanism can resist "concept drift" (changes in data distribution), making the system more reliable and durable, unlike static solutions.
[0139] By constructing an end-to-end risk management closed loop through three stages: pre-event, during-event, and post-event, the pre-event rule verification provides basic screening, the during-event model verification handles complex logic, and the post-event optimization ensures continuous improvement. It not only covers the entire lifecycle of freight business (such as order execution to fund settlement), but also supports real-time decision-making (such as during-event verification can automatically trigger risk alarms), effectively reducing overall risk exposure and improving business security.
[0140] In the freight business risk identification model creation module, the freight business risk identification model is constructed based on a feature extraction layer, a feature fusion layer, and a prediction output layer;
[0141] The feature extraction layer is constructed based on a contract data processing module, a business data processing module, a trajectory data processing module, a capital data processing module, and a bill data processing module. The contract data processing module extracts key semantic dependencies from contract data using a bidirectional encoder built with a Transformer to obtain contract risk features. The business data processing module models linear and nonlinear relationships between business indicators in business data using a multi-layered first fully connected network to obtain business operation features. The trajectory data processing module extracts trajectory risk features from trajectory data using a first gated recurrent unit. The capital data processing module captures local transaction features from capital data using a one-dimensional convolutional neural network and captures sequence dependency features of capital data using a second gated recurrent unit, outputting capital flow features based on the local transaction features and sequence dependency features. The bill data processing module extracts bill integrity features from bill data using a multi-layered second fully connected network.
[0142] The freight business risk identification model integrates multi-source heterogeneous data, including contract data, business data, trajectory data, financial data, and invoice data, through a feature extraction layer. This enables comprehensive monitoring of freight business risks, avoiding the limitations of a single data source and capturing a wider range of risk factors (such as contract defaults, abnormal funding, and trajectory deviations). Dedicated modules process different types of data (e.g., Transformer for text, CNN / GRU for sequences), enhancing the practicality and market adaptability of the solution and meeting the freight industry's needs for comprehensive risk assessment.
[0143] By processing contract data through a Transformer bidirectional encoder, key semantic dependencies are captured, improving the accuracy of text understanding. By processing business and invoice data through a multi-layer fully connected network, linear and non-linear relationships are modeled, enhancing the ability to analyze complex business indicators. By combining GRU and CNN to process trajectory and funding data, GRU captures temporal dependencies, and CNN extracts local features (such as transaction patterns), ensuring efficient identification of dynamic risks. This hybrid architecture (combining Transformer, CNN, GRU, etc.) can efficiently process structured, unstructured, and temporal data, reduce feature loss, and improve model robustness.
[0144] The feature fusion layer is constructed based on a multi-head attention module, a long short-term memory module, and a feature integration module. The multi-head attention module is used to calculate the feature weights of contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features through multi-head self-attention units. Based on the feature weights, the contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features are fused to obtain fused features. The long short-term memory module is used to model the temporal dependencies of the fused features through LSTM units to obtain serialized features. The feature integration module is used to perform global average pooling on the serialized features to obtain the whole graph feature vector.
[0145] The feature fusion layer dynamically calculates the feature weights of each feature through a multi-head attention module and performs weighted fusion. Then, it uses LSTM to model temporal dependencies and global average pooling to reduce dimensionality, ensuring that key risk features are given priority. At the same time, it processes the time series characteristics in the data (such as capital flow trends) to improve the fusion effect. The combination of multi-head attention and LSTM enables adaptive feature fusion, reduces redundant information, and improves computational efficiency.
[0146] The prediction output layer is constructed based on a risk item prediction branch, a risk level prediction branch, and a joint output module. The risk item prediction branch is used to infer the feature vector of the entire graph through a first fully connected layer and a sigmoid activation function to obtain the independent probabilities of different risk items (such as fraud, delay, default, etc.), supporting multi-label prediction. The risk level prediction branch is used to infer the feature vector of the entire graph through a second fully connected layer and a Softmax activation function to obtain the category probabilities of risk levels (low, medium, high). The joint output module is used to output a business risk identification report including risk items and risk levels based on the independent probabilities and category probabilities.
[0147] The prediction output layer adopts a dual-branch structure (risk item prediction and risk level prediction), using Sigmoid and Softmax activation functions to output independent risk item probabilities and risk level probabilities respectively, and jointly outputting a comprehensive report. This allows the freight business risk identification model to handle multiple related tasks simultaneously (such as identifying specific risk items and assessing the overall risk level), improving the fineness and accuracy of prediction. The multi-task learning framework reduces model complexity by sharing features (whole graph feature vectors), and the loss function design (weighted binary cross-entropy and classification cross-entropy) balances the weights of different tasks, optimizes the training process, effectively improves the generalization ability of the freight business risk identification model, and makes the output business risk identification report more comprehensive.
[0148] In the freight business risk identification model creation module, the loss function is constructed based on the weighted sum of risk item loss and risk level loss; the risk item loss adopts binary cross-entropy loss; the risk level loss adopts classification cross-entropy loss.
[0149] The loss function is based on a weighted sum of risk term loss (binary cross-entropy) and risk level loss (classification cross-entropy), which allows for dynamic adjustment of the priority of different tasks during training (such as focusing more on high-risk levels or specific risk terms), improving the convergence speed and stability of the freight business risk identification model. This weighted loss mechanism is innovative and can effectively handle the imbalance problem in multi-task learning (such as the low frequency of some risk terms), reduce the risk of overfitting, and improve the reliability of the freight business risk identification model in real-world scenarios.
[0150] The dataset construction module is specifically used for:
[0151] Acquire a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data;
[0152] The historical contract data includes at least the basic contract information, information on the cooperating parties, information on the goods, transportation service requirements, fee terms, liability for breach of contract, and dispute resolution methods;
[0153] The basic contract information includes at least the contract number, signing date, and contract validity period, used for contract classification, retrieval, and management to ensure timeliness and traceability. The information on the cooperating parties includes at least the shipper's name, carrier's name, address, contact information, and business license number, used to clarify the identities and responsibilities of both parties, providing a basis for subsequent business operations and dispute resolution. The cargo information includes at least the cargo's name, specifications, model, weight, volume, quantity, and value, used for rationally arranging transport vehicles, calculating freight costs, and ensuring cargo safety. The transportation service requirements include at least the origin, destination, transportation time requirements (e.g., latest pick-up time, latest delivery time), and mode of transport (road, rail, etc.). The terms and conditions for handling goods include: (1) waterway transportation and (2) loading and unloading requirements, to ensure that the transportation service can meet the specific needs of the customer; (2) fee terms, to include at least the method of calculating freight (e.g., by weight, volume or mileage), payment method (e.g., prepayment, cash on delivery, monthly settlement), payment time, and invoice issuance requirements, to clarify the financial responsibilities and settlement process of both parties; (3) liability for breach of contract, to stipulate corresponding liability and compensation methods for possible breaches of contract during the performance of the contract, such as delayed delivery, damage or loss of goods, to protect the legitimate rights and interests of both parties; and (4) dispute resolution, to stipulate the means of resolving contract disputes, such as negotiation, mediation, arbitration or litigation, as well as the applicable law and arbitration institution or court of jurisdiction, to provide a mechanism for resolving possible disputes.
[0154] The business data includes at least waybill information, order logs, online transaction logs, customer information, vehicle and driver information, and logistics information;
[0155] The waybill information includes at least the waybill number, consignor information, consignee information, cargo description (consistent with the cargo information in the contract), transportation route, transportation mileage, estimated transportation time, actual transportation time, and waybill status (e.g., ordered, picked up, in transit, delivered), used to track and manage the entire process of each transportation transaction; the order log records the time, operator, and reason for operations such as order generation, modification, and cancellation, facilitating the tracing of order changes and ensuring the transparency and auditability of business operations; the online transaction log includes at least records of user online freight platform activities such as login, browsing, querying, order placement, quotation, and transaction completion, used to reflect user transaction habits and platform transaction activity, providing a basis for data analysis and business optimization; the customer information includes, in addition to the consignor information involved in the contract... In addition to information on people and carriers, the system also includes customer registration information, credit ratings, historical transaction records, and preference settings. This helps online freight platforms better understand customer needs and provide personalized services, while also facilitating customer relationship management and risk assessment. The vehicle and driver information includes at least license plate number, vehicle type, approved load capacity, vehicle status (e.g., idle, in transit, under maintenance), driver's name, driver's contact information, driver's license information, professional qualification certificate information, driving experience, and credit record. This information is used to rationally allocate vehicle and driver resources and ensure transportation safety. The logistics information includes at least the real-time location of goods, transportation status (e.g., whether transportation is normal, whether any abnormalities have occurred), loading and unloading times, and transit information. By combining this information with trajectory data, the system enables full-process monitoring and visual management of the logistics process, improving logistics efficiency and customer satisfaction.
[0156] The trajectory data includes at least time information, geographical location information, driving route information, speed information, mileage information, and stop point information;
[0157] The time information refers to the timestamps of the vehicle at various locations, including departure time, time of passing through each node, and time of arrival at the destination. This time information allows for the calculation of vehicle speed, dwell time, and other parameters, thereby analyzing transportation efficiency and identifying any abnormal stops. The geographical location information refers to the real-time latitude and longitude coordinates of the vehicle during transportation, obtained through positioning devices (such as GPS and BeiDou). This information is used to determine the vehicle's precise location, map its trajectory, and enable dynamic tracking and monitoring of the transportation process. The driving path information refers to the actual route the vehicle travels, including road segments, intersections, and highway entrances / exits. Comparing this information with the planned transportation route allows for the determination of whether the vehicle is following the predetermined route and whether there are any abnormalities such as detours or deviations. The speed information... The vehicle speed information refers to the vehicle's speed over different time periods. Monitoring speed allows analysis of whether the vehicle is speeding and whether there are abnormal speed changes due to road conditions, traffic control, or other factors, providing a reference for assessing transportation safety and efficiency. The mileage information refers to the distance the vehicle travels during transportation, which can be used to calculate costs such as fuel consumption and vehicle wear and tear. Combined with information such as cargo weight, it can also assess transportation cost-effectiveness and provide data support for optimizing transportation plans. The stop point information refers to the vehicle's stopping locations and duration during transportation, including loading and unloading points, refueling stations, rest areas, etc. Analyzing stop points can help understand the vehicle's operating mode and the driver's driving habits, and also help identify potential risks, such as the risk of cargo theft due to prolonged stops.
[0158] The financial data includes at least fund flow information, account balance information, and settlement information;
[0159] The fund flow information includes records of every income and expenditure, such as freight charges, freight payments to the actual carrier, and platform service fees. It details the flow of funds, amount, transaction time, and transaction counterparty to ensure transparency and traceability. The account balance information reflects the fund account balances of both the shipper and carrier on the platform, indicating their financial status and available funds, providing a basis for fund settlement and financial management. The settlement information includes at least the settlement cycle, settlement method, settlement time, and settlement amount, clarifying the settlement rules and procedures to ensure timely and accurate settlement and protect the legitimate rights and interests of all parties.
[0160] The invoice data includes at least the invoice number, invoice date, invoice amount, and invoice type;
[0161] The historical freight business data are preprocessed, including at least deduplication, missing value handling, error correction, data standardization, data normalization, and data encoding. Risk items and risk levels are labeled on the preprocessed historical freight business data. A dataset is constructed using the labeled historical freight business data. The sample size of the dataset is expanded using an adversarial network.
[0162] By employing multi-step preprocessing (including deduplication, missing value handling, error correction, data standardization, data normalization, and data encoding), the raw data is systematically cleaned and transformed, reducing the impact of noise, inconsistencies, and biases, and ensuring the consistency and operability of the dataset. For example, standardization and normalization processes enable data from different sources (such as speed information and monetary information) to be compared, thereby improving the accuracy of subsequent risk labeling and model training.
[0163] By integrating multi-dimensional historical freight data (including contract, business, trajectory, financial and invoice data) and adopting systematic preprocessing (such as deduplication and standardization) to improve data quality, and combining risk labeling to build a high-quality dataset, and then innovatively applying adversarial networks to expand the sample size, the comprehensiveness, accuracy and robustness of freight risk identification are significantly improved. This effectively solves the prediction bias problem caused by single data or insufficient samples in traditional models, while realizing automated and efficient processing, providing scalable technical support for logistics risk management.
[0164] The freight business risk identification model training module is specifically used for:
[0165] The dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1 using stratified sampling. The freight business risk identification model is trained using the training set. During the training process, the hyperparameters of the freight business risk identification model are continuously optimized until the loss value of the loss function is less than the preset loss threshold.
[0166] The trained freight business risk identification model is validated using the validation set to determine whether the identification accuracy is greater than a preset accuracy threshold. If not, the validation fails, and the training set is expanded for further training; if yes, the validation passes, and:
[0167] The validated freight business risk identification model is tested using the test set to determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded for further training. If yes, the test passes, and the validated freight business risk identification model is deployed to the server using containerization technology.
[0168] By dividing the dataset into training, validation, and test sets in an 8:1:1 ratio using stratified sampling, the data distribution was balanced and representative, sampling bias was reduced, and the model's generalization ability (i.e., the model's performance on unseen data) was improved, thereby reducing the misjudgment rate in freight business risk identification. The 8:1:1 ratio also optimized resource utilization, avoided overfitting or underfitting problems, and enhanced the reliability of the model.
[0169] During the training phase, by continuously optimizing hyperparameters (such as learning rate and batch size) until the loss value of the loss function is less than a preset threshold, the model achieves rapid convergence and efficient training, reducing training time and computational resource consumption, while ensuring that the model achieves high accuracy in the freight risk identification task; the iterative optimization mechanism (such as dynamic adjustment based on loss value) improves the adaptability of training, enabling the model to better handle noise and complexity in freight data.
[0170] By optimizing dataset partitioning through stratified sampling and combining iterative training with a dual threshold verification mechanism, the accuracy, robustness, and reliability of the freight business risk identification model were significantly improved. Containerization technology enabled efficient deployment and elastic scaling, greatly reducing operation and maintenance costs. Simultaneously, through terminal authentication and structured storage of rule sets, the customizability and real-time decision-making capabilities of business rules were enhanced while ensuring system security. This formed a closed-loop optimization system from data preprocessing to model deployment, effectively solving the core pain points of low risk identification efficiency, high false positive rate, and poor system scalability in freight scenarios.
[0171] The business verification rule set setting module is specifically used for:
[0172] After authenticating the accessing mobile terminal, the server obtains several verification rules set by the mobile terminal, constructs a business verification rule set based on each of the verification rules, and stores the business verification rule set in a structured manner to a specified path.
[0173] The multidimensional data acquisition module is specifically used for:
[0174] The server uses an ETL tool to extract real-time freight business data from the contract system, business system, trajectory system, fund system, and bill system based on a preset extraction period. This data includes real-time contract data, real-time business data, real-time trajectory data, real-time fund data, and real-time bill data. The server then performs data transformation operations on each of the real-time freight business data, including at least data cleaning, data formatting, and data aggregation. Finally, the server loads the transformed real-time freight business data into a preset data warehouse.
[0175] The pre-verification module is specifically used for:
[0176] The server triggers events based on freight business and calls the rule engine through the API interface. The rule engine reads the business verification rule set from the specified path and performs pre-verification on real-time freight business data based on the business verification rule set, including at least data format, data integrity, logical consistency, vehicle qualification, personnel qualification, vehicle trajectory, waybill timeliness, and invoice information, and generates a pre-verification report carrying whether the pre-verification was successful or failed.
[0177] When the pre-verification report indicates that the pre-verification has failed, the pre-verification report will be pushed to the pre-associated management terminal in real time via the TLS protocol.
[0178] The server obtains the pre-verification time, calculates the first hash value of the pre-verification report and the pre-verification time, encrypts the pre-verification report, the pre-verification time, and the first hash value into first-level encrypted data using the SM9 algorithm, divides the first-level encrypted data in a 7:3 ratio and swaps the order to obtain second-level encrypted data, encrypts the second-level encrypted data into third-level encrypted data using the RC6 algorithm, and shifts each character of the third-level encrypted data cyclically 5 bits to the right to obtain the pre-encrypted report, which is then uploaded to the blockchain.
[0179] By automatically calling the rule engine through the API interface, the processing of freight business events is triggered in real time, avoiding the delays of traditional manual verification. The rule engine reads the verification rule set from the specified path and performs multi-dimensional checks on the data (such as data format, integrity, logical consistency, vehicle qualifications, personnel qualifications, vehicle trajectory, waybill timeliness, and invoice information), ensuring that the verification process is efficient and automated, greatly reducing business downtime, improving the overall throughput of the freight process (for example, it can quickly generate reports when failure occurs), thereby improving operational efficiency and reducing labor costs.
[0180] A multi-layered encryption mechanism (including the SM9 algorithm, RC6 algorithm, character shifting, and splitting / swapping) provides enhanced protection for the data. Specifically: hash values are used to ensure data integrity and prevent verification reports and timestamps from being tampered with; the SM9 algorithm provides post-quantum level asymmetric encryption, with first-level encryption ensuring the confidentiality of sensitive information; splitting / swapping (splitting and swapping in a 7:3 ratio) increases the difficulty of cracking; the RC6 algorithm further encrypts the data and introduces perturbations through a 5-bit rightward cyclic shift of characters, making the encryption process more complex; this multi-step process significantly improves the data's resistance to attacks and is suitable for freight data environments with high security requirements; finally, the data is uploaded to the blockchain, leveraging its distributed ledger characteristics to ensure data immutability and traceability, enhancing the overall system's non-repudiation capabilities.
[0181] The in-process verification module is specifically used for:
[0182] The server inputs the real-time freight business data into the deployed freight business risk identification model. The freight business risk identification model performs inference through hardware acceleration technology to obtain a real-time business risk identification report that includes real-time risk items and real-time risk levels, so as to verify the real-time freight business data in real time.
[0183] When the risk level is higher than the preset level, the real-time business risk identification report will be pushed to the pre-associated management terminal in real time via the TLS protocol;
[0184] The server obtains the in-process verification time, calculates the second hash value of the real-time business risk identification report and the in-process verification time, and encrypts the real-time business risk identification report, the in-process verification time, and the second hash value into a first layer of encrypted data using the XTEA algorithm. The first layer of encrypted data is converted into hexadecimal data, and the number 7 is swapped with the letter B, the number 8 is swapped with the number 9, and the number 10 is swapped with the letter A to obtain a second layer of encrypted data. The second layer of encrypted data is encrypted into a third layer of encrypted data using the IDEA algorithm. A random string of a specified length is added to a specified position in the third layer of encrypted data to obtain an in-process encrypted report, and the in-process encrypted report is uploaded to the blockchain.
[0185] By using hardware acceleration technologies (such as GPUs or FPGAs) for inference of the freight business risk identification model, the computational latency of real-time data processing is significantly reduced, enabling in-process verification and avoiding the time lag of traditional batch processing methods. It has a high response speed (millisecond-level inference) and can promptly identify and handle high-risk freight events (such as cargo damage or fraud), thereby reducing business losses. This real-time capability enhances the practicality of the system, making it particularly suitable for high-frequency, dynamic freight scenarios.
[0186] The post-optimization module is specifically used for:
[0187] The server constructs an incremental dataset based on the real-time freight business data. This incremental dataset is labeled with real-time risk items, real-time risk levels, actual risk items, and actual risk levels. When the data volume of the incremental dataset exceeds a preset threshold, the freight business risk identification model is post-optimized using the incremental dataset. The real-time risk items and real-time risk levels are predicted data, while the actual risk items and actual risk levels are actual data.
[0188] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for verifying freight operations by integrating multidimensional data, characterized in that: Includes the following steps: Step S1: Create a freight business risk identification model and set the loss function of the freight business risk identification model; the freight business risk identification model is constructed based on a feature extraction layer, a feature fusion layer and a prediction output layer; The feature extraction layer is constructed based on a contract data processing module, a business data processing module, a trajectory data processing module, a capital data processing module, and a bill data processing module. The contract data processing module extracts key semantic dependencies from contract data using a bidirectional encoder built with a Transformer to obtain contract risk features. The business data processing module models linear and nonlinear relationships between business indicators in business data using a multi-layered first fully connected network to obtain business operation features. The trajectory data processing module extracts trajectory risk features from trajectory data using a first gated recurrent unit. The capital data processing module captures local transaction features from capital data using a one-dimensional convolutional neural network and captures sequence dependency features of capital data using a second gated recurrent unit, outputting capital flow features based on the local transaction features and sequence dependency features. The bill data processing module extracts bill integrity features from bill data using a multi-layered second fully connected network. The feature fusion layer is constructed based on a multi-head attention module, a long short-term memory module, and a feature integration module. The multi-head attention module is used to calculate the feature weights of contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features through multi-head self-attention units. Based on the feature weights, the contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features are fused to obtain fused features. The long short-term memory module is used to model the temporal dependency of the fused features through LSTM units to obtain serialized features; the feature integration module is used to perform global average pooling on the serialized features to obtain the whole image feature vector. The prediction output layer is constructed based on a risk item prediction branch, a risk level prediction branch, and a joint output module. The risk item prediction branch is used to infer the feature vector of the entire graph through a first fully connected layer and a sigmoid activation function to obtain the independent probabilities of different risk items. The risk level prediction branch is used to infer the feature vector of the entire graph through a second fully connected layer and a softmax activation function to obtain the category probability of the risk level. The joint output module is used to output a business risk identification report including risk items and risk levels based on the independent probabilities and category probabilities. The loss function is constructed based on a weighted sum of risk term loss and risk level loss; the risk term loss uses binary cross-entropy loss; the risk level loss uses classification cross-entropy loss. Step S2: Obtain a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data; preprocess and label each of the historical freight business data to construct a dataset. Step S3: Train the freight business risk identification model using the dataset and loss function, and deploy the trained freight business risk identification model to the server; Step S4: The server sets up a business verification rule set containing several verification rules; Step S5: The server uses an ETL tool to obtain real-time freight business data, including real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time bill data, from the contract system, business system, tracking system, funds system, and bill system. Step S6: Based on the freight business trigger event, the server calls the rule engine through the API interface. The rule engine reads the business verification rule set from the specified path and performs pre-verification on the real-time freight business data through the business verification rule set, including at least data format, data integrity, logical consistency, vehicle qualification, personnel qualification, vehicle trajectory, waybill timeliness and invoice information, and generates a pre-verification report carrying the pre-verification success or failure. When the pre-verification report indicates that the pre-verification has failed, the pre-verification report will be pushed to the pre-associated management terminal in real time via the TLS protocol. The server obtains the pre-verification time, calculates the first hash value of the pre-verification report and the pre-verification time, encrypts the pre-verification report, the pre-verification time, and the first hash value into first-level encrypted data using the SM9 algorithm, divides the first-level encrypted data in a 7:3 ratio and swaps the order to obtain second-level encrypted data, encrypts the second-level encrypted data into third-level encrypted data using the RC6 algorithm, shifts each character of the third-level encrypted data 5 bits to the right to obtain the pre-encrypted report, and uploads the pre-encrypted report to the blockchain. Step S7: The server performs in-process verification of the real-time freight business data using the deployed freight business risk identification model. Step S8: The server constructs an incremental dataset based on the real-time freight business data, and performs post-event optimization of the freight business risk identification model using the incremental dataset.
2. The freight business verification method integrating multi-dimensional data as described in claim 1, characterized in that: Step S2 specifically involves: Acquire a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data; The historical contract data includes at least the basic contract information, information on the cooperating parties, information on the goods, transportation service requirements, fee terms, liability for breach of contract, and dispute resolution methods; The business data includes at least waybill information, order logs, online transaction logs, customer information, vehicle and driver information, and logistics information; The trajectory data includes at least time information, geographical location information, driving route information, speed information, mileage information, and stop point information; The financial data includes at least fund flow information, account balance information, and settlement information; The invoice data includes at least the invoice number, invoice date, invoice amount, and invoice type; The historical freight business data are preprocessed, including at least deduplication, missing value handling, error correction, data standardization, data normalization, and data encoding. Risk items and risk levels are labeled on the preprocessed historical freight business data. A dataset is constructed using the labeled historical freight business data. The sample size of the dataset is expanded using an adversarial network.
3. The freight business verification method integrating multi-dimensional data as described in claim 1, characterized in that: Step S3 specifically involves: The dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1 using stratified sampling. The freight business risk identification model is trained using the training set. During the training process, the hyperparameters of the freight business risk identification model are continuously optimized until the loss value of the loss function is less than the preset loss threshold. The trained freight business risk identification model is validated using the validation set to determine whether the identification accuracy is greater than the preset accuracy threshold. If not, the validation fails, and the training set is expanded for further training. If so, the verification passes, and: The validated freight business risk identification model is tested using the test set to determine whether the confidence level is greater than a preset confidence level threshold. If not, the test fails, and the training set is expanded for further training. If yes, the test passes, and the validated freight business risk identification model is deployed to the server using containerization technology. Step S4 specifically involves: After authenticating the accessing mobile terminal, the server obtains several verification rules set by the mobile terminal, constructs a business verification rule set based on each of the verification rules, and stores the business verification rule set in a structured manner to a specified path.
4. The freight business verification method integrating multi-dimensional data as described in claim 1, characterized in that: Step S5 specifically involves: The server uses an ETL tool to extract real-time freight business data from the contract system, business system, trajectory system, fund system, and bill system based on a preset extraction period. This data includes real-time contract data, real-time business data, real-time trajectory data, real-time fund data, and real-time bill data. The server then performs data transformation operations on each of these real-time freight business data, including at least data cleaning, data formatting, and data aggregation. Finally, the server loads the transformed real-time freight business data into a preset data warehouse.
5. The freight business verification method integrating multi-dimensional data as described in claim 1, characterized in that: Step S7 specifically involves: The server inputs the real-time freight business data into the deployed freight business risk identification model. The freight business risk identification model performs inference through hardware acceleration technology to obtain a real-time business risk identification report that includes real-time risk items and real-time risk levels, so as to verify the real-time freight business data in real time. When the risk level is higher than the preset level, the real-time business risk identification report will be pushed to the pre-associated management terminal in real time via the TLS protocol; The server obtains the in-process verification time, calculates the second hash value of the real-time business risk identification report and the in-process verification time, and encrypts the real-time business risk identification report, the in-process verification time, and the second hash value into a first layer of encrypted data using the XTEA algorithm. The first layer of encrypted data is converted into hexadecimal data, and the number 7 is swapped with the letter B, the number 8 is swapped with the number 9, and the number 10 is swapped with the letter A to obtain a second layer of encrypted data. The second layer of encrypted data is encrypted into a third layer of encrypted data using the IDEA algorithm. A random string of a specified length is added to a specified position in the third layer of encrypted data to obtain an in-process encrypted report, and the in-process encrypted report is uploaded to the blockchain. Step S8 specifically involves: The server constructs an incremental dataset based on the real-time freight business data, and labels the incremental dataset with real-time risk items, real-time risk levels, actual risk items, and actual risk levels. When the amount of data in the incremental dataset exceeds a preset threshold, the freight business risk identification model is optimized ex-post using the incremental dataset.
6. A freight business verification system integrating multi-dimensional data, characterized in that: Includes the following modules: The freight business risk identification model creation module is used to create a freight business risk identification model and set the loss function of the freight business risk identification model; the freight business risk identification model is constructed based on a feature extraction layer, a feature fusion layer and a prediction output layer; The feature extraction layer is constructed based on a contract data processing module, a business data processing module, a trajectory data processing module, a capital data processing module, and a bill data processing module. The contract data processing module extracts key semantic dependencies from contract data using a bidirectional encoder built with a Transformer to obtain contract risk features. The business data processing module models linear and nonlinear relationships between business indicators in business data using a multi-layered first fully connected network to obtain business operation features. The trajectory data processing module extracts trajectory risk features from trajectory data using a first gated recurrent unit. The capital data processing module captures local transaction features from capital data using a one-dimensional convolutional neural network and captures sequence dependency features of capital data using a second gated recurrent unit, outputting capital flow features based on the local transaction features and sequence dependency features. The bill data processing module extracts bill integrity features from bill data using a multi-layered second fully connected network. The feature fusion layer is constructed based on a multi-head attention module, a long short-term memory module, and a feature integration module. The multi-head attention module is used to calculate the feature weights of contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features through multi-head self-attention units. Based on the feature weights, the contract risk features, business operation features, trajectory risk features, capital flow features, and bill integrity features are fused to obtain fused features. The long short-term memory module is used to model the temporal dependency of the fused features through LSTM units to obtain serialized features; the feature integration module is used to perform global average pooling on the serialized features to obtain the whole image feature vector. The prediction output layer is constructed based on a risk item prediction branch, a risk level prediction branch, and a joint output module. The risk item prediction branch is used to infer the feature vector of the entire graph through a first fully connected layer and a sigmoid activation function to obtain the independent probabilities of different risk items. The risk level prediction branch is used to infer the feature vector of the entire graph through a second fully connected layer and a softmax activation function to obtain the category probability of the risk level. The joint output module is used to output a business risk identification report including risk items and risk levels based on the independent probabilities and category probabilities. The loss function is constructed based on a weighted sum of risk term loss and risk level loss; the risk term loss uses binary cross-entropy loss; the risk level loss uses classification cross-entropy loss. The dataset construction module is used to acquire a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data. The dataset is constructed after preprocessing and labeling each of the historical freight business data. The freight business risk identification model training module is used to train the freight business risk identification model using the dataset and loss function, and to deploy the trained freight business risk identification model to the server. The business verification rule set setting module is used by the server to set a business verification rule set containing several verification rules. The multidimensional data acquisition module is used by the server to acquire real-time freight business data, including real-time contract data, real-time business data, real-time trajectory data, real-time financial data, and real-time invoice data, from the contract system, business system, trajectory system, financial system, and invoice system through ETL tools. The pre-verification module is used by the server to call the rule engine through the API interface based on the event triggered by the freight business. The rule engine reads the business verification rule set from the specified path and performs pre-verification on the real-time freight business data through the business verification rule set, including at least data format, data integrity, logical consistency, vehicle qualification, personnel qualification, vehicle trajectory, waybill timeliness and invoice information, and generates a pre-verification report with the pre-verification success or failure. When the pre-verification report indicates that the pre-verification has failed, the pre-verification report will be pushed to the pre-associated management terminal in real time via the TLS protocol. The server obtains the pre-verification time, calculates the first hash value of the pre-verification report and the pre-verification time, encrypts the pre-verification report, the pre-verification time, and the first hash value into first-level encrypted data using the SM9 algorithm, divides the first-level encrypted data in a 7:3 ratio and swaps the order to obtain second-level encrypted data, encrypts the second-level encrypted data into third-level encrypted data using the RC6 algorithm, shifts each character of the third-level encrypted data 5 bits to the right to obtain the pre-encrypted report, and uploads the pre-encrypted report to the blockchain. The in-process verification module is used by the server to perform in-process verification on the real-time freight business data through the deployed freight business risk identification model; The post-event optimization module is used by the server to construct an incremental dataset based on the real-time freight business data, and to perform post-event optimization on the freight business risk identification model using the incremental dataset.
7. The freight business verification system integrating multi-dimensional data as described in claim 6, characterized in that: The dataset construction module is specifically used for: Acquire a large amount of historical freight business data, including historical contract data, historical business data, historical trajectory data, historical financial data, and historical bill data; The historical contract data includes at least the basic contract information, information on the cooperating parties, information on the goods, transportation service requirements, fee terms, liability for breach of contract, and dispute resolution methods; The business data includes at least waybill information, order logs, online transaction logs, customer information, vehicle and driver information, and logistics information; The trajectory data includes at least time information, geographical location information, driving route information, speed information, mileage information, and stop point information; The financial data includes at least fund flow information, account balance information, and settlement information; The invoice data includes at least the invoice number, invoice date, invoice amount, and invoice type; The historical freight business data are preprocessed, including at least deduplication, missing value handling, error correction, data standardization, data normalization, and data encoding. Risk items and risk levels are labeled on the preprocessed historical freight business data. A dataset is constructed using the labeled historical freight business data. The sample size of the dataset is expanded using an adversarial network.
8. The freight business verification system integrating multi-dimensional data as described in claim 6, characterized in that: The freight business risk identification model training module is specifically used for: The dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1 using stratified sampling. The freight business risk identification model is trained using the training set. During the training process, the hyperparameters of the freight business risk identification model are continuously optimized until the loss value of the loss function is less than the preset loss threshold. The trained freight business risk identification model is validated using the validation set to determine whether the identification accuracy is greater than the preset accuracy threshold. If not, the validation fails, and the training set is expanded for further training. If so, the verification passes, and: The validated freight business risk identification model is tested using the test set to determine whether the confidence level is greater than the preset confidence level threshold. If not, the test fails, and the training set is expanded for further training. If so, the test is passed, and the freight business risk identification model that has passed the test will be deployed to the server using containerization technology; The business verification rule set setting module is specifically used for: After authenticating the accessing mobile terminal, the server obtains several verification rules set by the mobile terminal, constructs a business verification rule set based on each of the verification rules, and stores the business verification rule set in a structured manner to a specified path.
9. A freight business verification system integrating multi-dimensional data as described in claim 6, characterized in that: The multidimensional data acquisition module is specifically used for: The server uses an ETL tool to extract real-time freight business data from the contract system, business system, trajectory system, fund system, and bill system based on a preset extraction period. This data includes real-time contract data, real-time business data, real-time trajectory data, real-time fund data, and real-time bill data. The server then performs data transformation operations on each of these real-time freight business data, including at least data cleaning, data formatting, and data aggregation. Finally, the server loads the transformed real-time freight business data into a preset data warehouse.
10. A freight business verification system integrating multi-dimensional data as described in claim 6, characterized in that: The in-process verification module is specifically used for: The server inputs the real-time freight business data into the deployed freight business risk identification model. The freight business risk identification model performs inference through hardware acceleration technology to obtain a real-time business risk identification report that includes real-time risk items and real-time risk levels, so as to verify the real-time freight business data in real time. When the risk level is higher than the preset level, the real-time business risk identification report will be pushed to the pre-associated management terminal in real time via the TLS protocol; The server obtains the in-process verification time, calculates the second hash value of the real-time business risk identification report and the in-process verification time, and encrypts the real-time business risk identification report, the in-process verification time, and the second hash value into a first layer of encrypted data using the XTEA algorithm. The first layer of encrypted data is converted into hexadecimal data, and the number 7 is swapped with the letter B, the number 8 is swapped with the number 9, and the number 10 is swapped with the letter A to obtain a second layer of encrypted data. The second layer of encrypted data is encrypted into a third layer of encrypted data using the IDEA algorithm. A random string of a specified length is added to a specified position in the third layer of encrypted data to obtain an in-process encrypted report, and the in-process encrypted report is uploaded to the blockchain. The post-optimization module is specifically used for: The server constructs an incremental dataset based on the real-time freight business data, and labels the incremental dataset with real-time risk items, real-time risk levels, actual risk items, and actual risk levels. When the amount of data in the incremental dataset exceeds a preset threshold, the freight business risk identification model is optimized ex-post using the incremental dataset.
Citation Information
Patent Citations
Network freight service abnormal state identification method based on multi-source data fusion
CN118071229A
Dangerous cargo risk early warning method and system and computer program
CN119359057A