Cross-border logistics data collaboration optimization method and device, electronic equipment and storage medium

By dynamically adjusting the anonymization level and building a global model in cross-border logistics, the problem of data silos in cross-border logistics has been solved, the data value has been maximized and the reliability of the optimization scheme has been improved, and the resource allocation efficiency and data collaborative optimization performance of cross-border logistics have been enhanced.

CN120851300BActive Publication Date: 2026-02-13SHENZHEN MINGXIN DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511340336.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-02-13
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

In cross-border logistics, the inability of logistics companies in different countries to share raw data due to trade secrets has led to a serious data silo phenomenon, which affects the accuracy of regional freight demand forecasting and the efficiency of logistics resource allocation. Existing static desensitization methods lack dynamic adjustment capabilities, resulting in poor performance of cross-border logistics data collaborative optimization.

Method used

By inputting feature data into a preset prediction network for initial desensitization, calculating the prediction error to obtain the desensitization level, and dynamically adjusting the data according to the preset desensitization strategy, an optimization scheme is generated by combining graph attention network and global model, and the model gradient parameters and feature contribution are aggregated to build a cross-border logistics data collaborative optimization system.

Benefits of technology

While ensuring data privacy, it maximizes the retention and dynamic adjustment of data value, improves the flexibility and accuracy of cross-border logistics data collaborative optimization, reduces the risk of data leakage, and enhances the reliability of logistics resource allocation and optimization schemes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851300B_ABST
    Figure CN120851300B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of cross-border logistics data cooperation optimization, and discloses a cross-border logistics data cooperation optimization method and device, electronic equipment and a storage medium, wherein the method comprises the following steps: inputting feature data into a preset prediction network to obtain a first prediction value, and preliminarily desensitizing the feature data, and inputting desensitized feature data in a desensitized regional data set into the preset prediction network to obtain a second prediction value; calculating the prediction error of each feature data, obtaining a desensitization level, and desensitizing the feature data in the regional data set to obtain a target regional data set after desensitization; and inputting target desensitized feature data into a preset global model to obtain an optimization scheme. The application has the beneficial technical effect that the ability of dynamically adjusting the desensitization granularity is realized, and the deficiency of a traditional static desensitization method in data value adjustment is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cross-border logistics data collaboration optimization, and in particular to a cross-border logistics data collaboration optimization method and device, an electronic device and a storage medium. BACKGROUND

[0002] In the modern logistics industry, the data island dilemma has become an important factor restricting business development and improving operational efficiency. Due to the protection of commercial secrets, logistics companies in different countries cannot share raw data such as customer information and freight rate strategies, and the overall information flow in the industry is greatly limited. This phenomenon makes the error rate of regional freight demand forecasting as high as 18%-25%, leading to inefficiency and uncertainty in resource allocation, transportation planning and decision-making processes for logistics companies, further affecting service quality and customer satisfaction. In order to solve this problem, logistics companies have gradually turned to sharing data that has been desensitized to protect sensitive information while achieving data sharing. The necessity of this process lies in the fact that by desensitizing raw data, enterprises can retain valuable information and reduce the risk of information leakage, thereby breaking through the shackles of data islands.

[0003] However, traditional static desensitization methods such as generalization and substitution often rely on fixed rules for processing and lack the ability to dynamically adjust the value of data, resulting in poor performance in cross-border logistics data collaboration optimization. SUMMARY

[0004] Therefore, it is necessary to propose a cross-border logistics data collaboration optimization method, device, electronic device and storage medium for the existing cross-border logistics data collaboration optimization problem.

[0005] A cross-border logistics data collaboration optimization method, the method comprising:

[0006] inputting feature data in a region data set of a specified region to be optimized into a preset prediction network to obtain a first prediction value corresponding to each feature data, and preliminarily desensitizing the feature data in the region data set to obtain a desensitized region data set after desensitization; wherein the region data set comprises a plurality of feature data;

[0007] inputting desensitized feature data in the desensitized region data set into the preset prediction network to obtain a second prediction value corresponding to each desensitized feature data;

[0008] calculating a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to each feature data;

[0009] obtaining a desensitization level of each feature data according to the prediction error;

[0010] According to a preset desensitization level and desensitization strategy corresponding table, the feature data in the region data set is desensitized to obtain a desensitized target region data set;

[0011] The target desensitized feature data in the target region data set is input into a preset global model to obtain an optimization scheme for the region to be optimized.

[0012] Further, before the step of inputting the target desensitized feature data in the target region data set into the preset global model to obtain the optimization scheme for the region to be optimized, the step further comprises:

[0013] Obtain the model gradient parameters and feature contribution degree uploaded by a plurality of preset local models;

[0014] Adjust the aggregation weight of the corresponding model gradient parameter based on the feature contribution degree of each preset local model;

[0015] Aggregate the model gradient parameters and the corresponding aggregation weights of each preset local model to obtain the preset global model.

[0016] Further, before the step of aggregating the model gradient parameters and the corresponding aggregation weights of each preset local model to obtain the preset global model, the step further comprises:

[0017] Calculate the offset value corresponding to each model gradient parameter based on all model gradient parameters;

[0018] According to the offset value corresponding to each model gradient parameter, set the corresponding model gradient parameter noise to obtain the target model gradient parameter of each preset local model.

[0019] Further, before the step of inputting the feature data in the region data set of the specified region to be optimized into the preset prediction network to obtain the first prediction value corresponding to each feature data, and preliminarily desensitizing the feature data in the region data set to obtain the desensitized desensitized region data set, the step further comprises:

[0020] Obtain a plurality of order text data of the specified region to be optimized;

[0021] Extract the feature data in each order text data through a preset graph attention network to form the region data set.

[0022] Further, the step of obtaining a plurality of order text data of the specified region to be optimized comprises:

[0023] Obtain a plurality of order documents of the specified region to be optimized;

[0024] Each order document is identified by a preset XLM-RoBERTa model to obtain corresponding temporary text data of each order document;

[0025] Each temporary text data is subjected to entity recognition by a preset BiLSTM-CRF model to obtain corresponding order text data of each temporary text data.

[0026] Further, before the step of performing desensitization processing on the feature data in the region data set according to the preset desensitization level and desensitization strategy correspondence table to obtain the target region data set after desensitization, the method further comprises:

[0027] Obtaining scene information corresponding to the region data set of the specified region to be optimized;

[0028] Calculating the relevance value of each feature data in the region data set and the scene information;

[0029] The feature data with a relevance value greater than a relevance threshold value is recorded as relevant feature data;

[0030] Setting a desensitization level for the relevant feature data.

[0031] Further, the step of performing desensitization processing on the feature data in the region data set according to the preset desensitization level and desensitization strategy correspondence table to obtain the target region data set after desensitization, comprises:

[0032] According to the preset desensitization level and desensitization strategy correspondence table, the feature data in the region data set is desensitized to obtain a temporary region data set;

[0033] Adding Gaussian noise to the feature data in the temporary region data set with a preset desensitization level lower than a preset level to obtain the target region data set.

[0034] A cross-border logistics data collaborative optimization device, the device comprises:

[0035] An optimization region data input module is configured to input feature data in a region data set of a specified region to be optimized into a preset prediction network to obtain a first prediction value corresponding to each feature data, and to perform preliminary desensitization on the feature data in the region data set to obtain a desensitized region data set after desensitization; wherein the region data set comprises a plurality of feature data;

[0036] A desensitized feature data input module is configured to input desensitized feature data in the desensitized region data set into the preset prediction network to obtain a second prediction value corresponding to each desensitized feature data;

[0037] a prediction error calculation module configured to calculate a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to the feature data;

[0038] a desensitization level acquisition module configured to acquire a desensitization level of each feature data according to the prediction error;

[0039] a desensitization processing module configured to perform desensitization processing on the feature data in the region data set according to a preset desensitization level and desensitization strategy correspondence table, to obtain a target region data set after desensitization;

[0040] an optimization scheme acquisition module configured to input target desensitization feature data in the target region data set into a preset global model, to obtain an optimization scheme of the region to be optimized.

[0041] An electronic device includes a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to cause the processor to perform the following steps:

[0042] input feature data in a region data set of a region to be optimized into a preset prediction network, to obtain a first prediction value corresponding to each feature data, and perform preliminary desensitization on the feature data in the region data set, to obtain a desensitized region data set after desensitization; wherein the region data set includes a plurality of feature data;

[0043] input desensitized feature data in the desensitized region data set into the preset prediction network, to obtain a second prediction value corresponding to each desensitized feature data;

[0044] calculate a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to the feature data;

[0045] acquire a desensitization level of each feature data according to the prediction error;

[0046] perform desensitization processing on the feature data in the region data set according to a preset desensitization level and desensitization strategy correspondence table, to obtain a target region data set after desensitization;

[0047] input target desensitization feature data in the target region data set into a preset global model, to obtain an optimization scheme of the region to be optimized.

[0048] A computer readable storage medium stores a computer program, the computer program being executed by a processor to cause the processor to perform the following steps:

[0049] Input feature data in a region data set specifying a region to be optimized into a preset prediction network to obtain a first prediction value corresponding to each feature data, and preliminarily desensitize the feature data in the region data set to obtain a desensitized region data set after desensitization; wherein the region data set comprises a plurality of feature data;

[0050] Input desensitized feature data in the desensitized region data set into the preset prediction network to obtain a second prediction value corresponding to each desensitized feature data;

[0051] Calculate a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to each feature data;

[0052] Obtain a desensitization level of each feature data according to the prediction error;

[0053] Desensitize the feature data in the region data set according to a preset desensitization level and desensitization strategy correspondence table to obtain a target region data set after desensitization;

[0054] Input target desensitized feature data in the target region data set into a preset global model to obtain an optimization scheme of the region to be optimized.

[0055] The beneficial effects of the present application are: generating a first prediction value based on feature data in a region data set and a second prediction value after preliminary desensitization, obtaining a desensitization level of each feature data by calculating a prediction error, and formulating a corresponding desensitization strategy for different feature data according to a preset desensitization level and desensitization strategy correspondence table, so as to ensure that data validity is maximally retained on the premise of guaranteeing data privacy, realize the ability of dynamically adjusting desensitization granularity, and effectively solve the deficiency of traditional static desensitization method in data value adjustment. BRIEF DESCRIPTION OF DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0057] Among them:

[0058] Figure 1 It is an application environment diagram of the cross-border logistics data collaborative optimization method in one embodiment;

[0059] Figure 2 It is a flowchart of the cross-border logistics data collaborative optimization method in one embodiment;

[0060] Figure 3 A structural block diagram of a cross-border logistics data collaborative optimization device in an embodiment;

[0061] Figure 4 A structural block diagram of an electronic device in an embodiment. DETAILED DESCRIPTION

[0062] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0063] Figure 1 A cross-border logistics data collaborative optimization application environment diagram in an embodiment. Referring to Figure 1 The cross-border logistics data collaborative optimization method is applied to a cross-border logistics data collaborative optimization system. The cross-border logistics data collaborative optimization system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 can be a desktop terminal or a mobile terminal. The mobile terminal can be at least one of a mobile phone, a tablet computer, a notebook computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to obtain a regional data set, and the server 120 is used to generate an optimization scheme.

[0064] As shown in Figure 2 In an embodiment, a cross-border logistics data collaborative optimization method is provided. The method can be applied to a terminal or a server. The embodiment is exemplified by application to a terminal. The cross-border logistics data collaborative optimization method specifically includes the following steps:

[0065] S1: inputting feature data in a regional data set of a specified region to be optimized into a preset prediction network to obtain a first prediction value corresponding to each feature data, and performing preliminary desensitization on the feature data in the regional data set to obtain a desensitized regional data set after desensitization; wherein the regional data set includes multiple feature data;

[0066] S2: inputting desensitized feature data in the desensitized regional data set into the preset prediction network to obtain a second prediction value corresponding to each desensitized feature data;

[0067] S3: calculating a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to each feature data;

[0068] S4: obtaining a desensitization level of each feature data according to the prediction error;

[0069] S5: According to the preset desensitization level and desensitization strategy corresponding table, the feature data in the region data set is desensitized to obtain a target region data set after desensitization.

[0070] S6: The target desensitization feature data in the target region data set is input into a preset global model to obtain an optimization scheme of the region to be optimized.

[0071] As described in step S1, the feature data in the region data set of the specified region to be optimized is input into a preset prediction network to obtain a first prediction value corresponding to each feature data, and the feature data in the region data set is preliminarily desensitized to obtain a desensitized region data set. The region data set includes a plurality of feature data. The feature data in the region data set to be optimized is input into a preset prediction network. The feature data includes important factors affecting cross-border logistics, such as transportation time, cost, cargo type, transportation mode, etc. At the same time, the region data set is preliminarily desensitized. In the data desensitization stage, various technologies can be used, such as data encryption, random noise addition or aggregation processing, etc. so that the data can retain certain meaning and characteristics while ensuring that individual information cannot be restored. The ultimate goal is to generate a desensitized region data set so that in the process of data analysis and optimization, the effectiveness of the data is not lost, and the sensitive information of individuals and enterprises is protected to the greatest extent. In addition, the preset prediction network can be an LSTM prediction network, which includes an input layer, a hidden layer and an output layer. The input layer is used to receive the region data set, the hidden layer is used to capture the time sequence dependence relationship in the region data set, and the output layer is used to predict future freight demand. Specifically, the input layer: data set; the hidden layer: 3 LSTM units (128 neurons per layer) to capture time sequence dependence; the output layer: generate prediction values of feature data. It should be noted that the LSTM prediction network can be obtained by training a pre-constructed LSTM initial network based on a preset sample set. Each sample data in the preset sample set includes feature data and a prediction value corresponding to the feature data. When training the pre-constructed LSTM initial network, the feature data in each sample data is used as the input of the LSTM initial network, and the prediction value corresponding to the feature data in each sample data is used as the output of the LSTM initial network. Through training, the LSTM initial network can learn the corresponding relationship between all possible feature data and prediction values. The trained LSTM initial network is used as the LSTM prediction network.

[0072] The de-sensitized feature data in the de-sensitized region data set is input into the preset prediction network to obtain a second prediction value corresponding to each de-sensitized feature data, as described in step S2. The feature data in the region data set after preliminary de-sensitization is input into the preset prediction network to obtain a second prediction value corresponding to each de-sensitized feature data. The purpose is to capture patterns and features that may be ignored by the first prediction value by further processing the already de-sensitized data, because after data de-sensitization, some information may be converted from detailed data to more ambiguous forms, and the second prediction can provide additional information and perspectives for improving the model.

[0073] As described in step S3, the prediction error of each feature data is calculated according to the first prediction value and the second prediction value corresponding to each feature data. The prediction error is calculated according to the first prediction value and the second prediction value corresponding to each feature data. The prediction error is an important indicator for evaluating the performance of the model, which can quantify the gap between the model prediction and the reality, and its calculation formula is the absolute difference or relative difference between the prediction value and the actual value. In a specific embodiment, the mean absolute error can be calculated. The first prediction value and the second prediction value of each feature data are compared to calculate the error of each feature data.

[0074] As described in step S4, the de-sensitization level of each feature data is obtained according to the prediction error. The de-sensitization level of each feature data is obtained based on the calculated prediction error. The de-sensitization level is an important indicator for measuring data sensitivity and protection needs. Generally speaking, the greater the prediction error, the higher the sensitivity of the data, because these data have a greater impact on actual application and decision-making. A clear standardized rating system can be prepared in advance, and a grading mechanism can be developed according to the size of the prediction error, such as dividing the prediction error into low, medium and high levels, and the de-sensitization level corresponding to each level is predefined, such as high-level error requiring lower degree of de-sensitization to ensure that the data can be better used for scheme optimization, and low-level error requiring higher degree of de-sensitization to achieve better data de-sensitization and ensure that the data will not be leaked. Through such hierarchical management, cross-border logistics data will be more targeted and flexible in application, ensuring that sensitive information will not be leaked due to data use. Specifically, when the prediction error is > 10%, the de-sensitization level is 1 (the highest), 5%-10% is 2, and < 5% is 3.

[0075] According to the preset desensitization level and desensitization strategy correspondence table, the feature data in the region data set is desensitized to obtain the desensitized target region data set. According to the preset desensitization level and desensitization strategy correspondence table, the feature data in the region data set is desensitized to obtain the final target region data set. Generally speaking, the higher the prediction error is, the lower the desensitization level is. For data with high desensitization level, more stringent desensitization measures such as encryption, splitting or randomization processing can be used to prevent potential data leakage, for example, the data "Hefei-Hamburg Central Europe Train" is converted to "East China-Europe Railway Trunk". For data with low desensitization level, only simple de-identification processing is required, for example: the data "Yantian Port 2024-07-15 14:00 to port" is converted to "Southern Core Port Q3 Working Period". By constructing the desensitization strategy correspondence table, a clear desensitization method is formulated for each type of data feature, ensuring the systematicness and efficiency of the processing flow. For example, set desensitization level 1 to correspond to data encryption, desensitization level 2 to correspond to generalization processing, and desensitization level 3 to correspond to adding noise. The final target region data set will be more usable and reliable in actual business applications while ensuring data privacy.

[0076] As described in step S6 above, the target desensitization feature data in the target region data set is input into the preset global model to obtain the optimization scheme of the region to be optimized. The target desensitization feature data in the target region data set is input into the preset global model to obtain the optimization scheme of the region to be optimized. Using desensitized data, relevant global models are used to generate specific optimization suggestions, which can be targeted at logistics scheduling, warehouse management, transportation cost control and other aspects, helping to optimize the overall efficiency and economy of cross-border logistics. The construction of the global model is usually based on a large amount of historical data and business logic, so it has high accuracy and reliability. After processing the desensitized data, potential legal risks and privacy issues can be reduced, making the implementation of the optimization scheme smoother.

[0077] In one embodiment, before the step S6 of inputting the target desensitization feature data in the target region data set into the preset global model to obtain the optimization scheme of the region to be optimized, it further includes:

[0078] S501: Obtain model gradient parameters and feature contribution degree uploaded by a plurality of preset local models;

[0079] S502: Adjust the aggregation weight of the corresponding model gradient parameter based on the feature contribution degree of each preset local model;

[0080] S503: Obtain the preset global model by aggregating the model gradient parameters of each preset local model and the corresponding aggregation weights.

[0081] As described in step S501, the model gradient parameters uploaded by the plurality of preset local models and the feature contribution degrees are obtained. Each local model generates corresponding model parameter updates, i.e., model gradient parameters, during the training process. These gradient parameters reflect the response of the model to the loss function and further reflect the importance of each feature to the model prediction. The feature contribution degree is an evaluation of the influence or contribution degree of each feature to the model prediction result. By collecting and integrating each local model, the performance of different features in each model can be better understood. The feature contribution degree of the local model can be evaluated by various methods, such as ranking based on importance score or calculating the loss change caused by the feature during the training process. In an embodiment, the feature contribution degree can be quantified by SHAP (Shapley Additive Explanations) value. In some embodiments, the model gradient parameters may contain sensitive information, and therefore the model gradient parameters can be differentially private when uploaded by each preset local model.

[0082] As described in step S502, the aggregation weights of the model gradient parameters of each preset local model are adjusted based on the feature contribution degrees. The aggregation weights of the model gradient parameters of each preset local model are adjusted based on the feature contribution degrees. This process reflects the importance of features in federated learning and the idea of ensemble learning. When aggregating multiple models, simply averaging may ignore the performance of some models on features. The importance and stability of each local model are analyzed by the feature contribution degree. The parameter updates of the model with high feature contribution degree should be given higher aggregation weights, because these models perform better in feature expression and prediction and can provide more information and guidance. On the contrary, for the model with low feature contribution degree, its weight can be reduced, which can reduce the bias introduced by low-quality models. Through this weighted aggregation method, the system can better utilize the advantages of different models while reducing the noise and instability that may occur.

[0083] As described in step S503, the preset global model is obtained by aggregating the model gradient parameters of each preset local model and the corresponding aggregation weights. The gradient parameters of each model are weighted and summed using the preset aggregation weights. This can well focus on models with high feature contribution, so that these models occupy a larger proportion in the global model. In this way, the final global model can fully utilize the learning achievements of all local models and optimize on this basis. It is worth noting that during the aggregation process, the consistency of gradient information and model parameters needs to be ensured to generalize to new data distribution, while avoiding potential overfitting problems.

[0084] In one embodiment, before the step S503 of aggregating the model gradient parameters of each preset local model and the corresponding aggregation weights to obtain the preset global model, the method further comprises:

[0085] S5021: calculating an offset value corresponding to each model gradient parameter based on all model gradient parameters;

[0086] S5022: setting a model gradient parameter noise corresponding to each model gradient parameter according to the offset value corresponding to each model gradient parameter to obtain a target model gradient parameter of each preset local model.

[0087] As described in step S5021, the offset value corresponding to each model gradient parameter is calculated based on all model gradient parameters. The offset value is a measure of the difference between a certain model gradient and the overall average trend. Specifically, first, the gradient parameters of all local models need to be summarized and counted. For example, a baseline value (global average) can be obtained by calculating the arithmetic mean of all gradient parameters, and then each local model's gradient parameter is subtracted from this average to obtain the offset value.

[0088] As described in step S5022 above, the model gradient parameter noise corresponding to each model gradient parameter is set according to the offset value corresponding to the model gradient parameter, to obtain the target model gradient parameter of each preset local model. The model gradient parameter noise corresponding to each model gradient parameter is set according to the offset value corresponding to the model gradient parameter, to obtain the target model gradient parameter of each preset local model. The addition of model gradient parameter noise is generally based on the size of the offset value. The model gradient with a larger offset value will appropriately increase the noise to prevent over-reliance on the bias introduced by these models; on the contrary, models with smaller offset values can not need much noise or use smaller noise. In this way, the "smoothing" of the model can be achieved, making the learning process of the model more robust. Adding noise not only improves the generalization ability of the model, but also helps to prevent information leakage, which is particularly important in the federated learning scenario, and can ensure that even if the gradient of a certain local model is relatively high (or the offset is significant), it will not have a disproportionate impact on the global model.

[0089] In one embodiment, before the step S1 of inputting the feature data in the region data set specifying the region to be optimized into a preset prediction network to obtain a first prediction value corresponding to each feature data, and preliminarily desensitizing the feature data in the region data set to obtain a desensitized region data set, the step further comprises:

[0090] S001: obtaining a plurality of order text data of the specified region to be optimized;

[0091] S002: extracting feature data in each of the order text data through a preset graph attention network to form the region data set.

[0092] As described in step S001 above, a plurality of order text data of the specified region to be optimized are obtained. The order text data contains a large amount of actual business information, which usually includes multiple dimensions of data content, such as customer information, product categories, delivery and destination, transportation methods, order amounts, order dates, and estimated arrival times.

[0093] As described in step S002 above, feature data is extracted from each order text data by a preset graph attention network to form the regional data set. The graph attention network (GAT) is a graph structure-based data learning method that can handle highly correlated data, such as the relationships between elements in the order text. The order text is represented in graph form, with each node possibly representing a key feature (e.g., product name, order status, etc.), and the edges between nodes representing their relationships. The importance of different features is dynamically adjusted through an attention mechanism. First, the order text data is encoded to convert it into a form that can be input into the graph attention network. In this process, the preset graph attention network will traverse and aggregate information through multiple layers, extract key features, and generate sensitivity feature metrics. These extracted feature data typically include order time features, geographic features, and behavior features. The extracted feature data will constitute the regional data set and serve as the basis for subsequent desensitization processing and model prediction. This process not only improves the analyzability of the data but also ensures that subsequent steps can be based on high-quality features for more accurate model training and decision support. Through the effective application of the graph attention network, feature extraction can better capture the connections and patterns in complex data, providing a solid foundation for subsequent optimization scheme development. In some embodiments, considering the strong time sequence of cross-border logistics orders, a spatio-temporal graph convolution network can be used to collect feature data, which has a better representation form and can better extract data information, making the subsequent generated optimization scheme more optimal.

[0094] In one embodiment, the step S001 of obtaining the order text data of the specified region to be optimized comprises:

[0095] S0011: Obtain the order documents of the specified region to be optimized.

[0096] S0012: Identify each order document using a preset XLM-RoBERTa model to obtain the corresponding temporary text data for each order document.

[0097] S0013: Perform entity recognition on each temporary text data using a preset BiLSTM-CRF model to obtain the corresponding order text data for each temporary text data.

[0098] As described in step S0011 above, a plurality of order documents of the specified region to be optimized are obtained. Order documents generally refer to various records generated after customers place orders, including scanned data of paper documents or order records in electronic format. These order documents contain legal and commercial agreements between consumers and businesses and are an important data source indispensable to logistics operations. These order documents can be obtained by connecting to the corresponding database, file storage system or ERP (Enterprise Resource Planning) system, and performing specific query operations to filter out order records in the specified region.

[0099] As described in step S0012 above, each of the order documents is identified by a preset XLM-RoBERTa model to obtain corresponding temporary text data of each of the order documents. XLM-RoBERTa model is a large-scale pre-training language model based on Transformer architecture, which is widely popular due to its context understanding ability and cross-language support. Order documents may be presented in image or PDF format, and the system will use image processing techniques such as OCR (Optical Character Recognition) to convert these documents into machine-readable text. Then, the XLM-RoBERTa model will process and analyze the text. Specific operations include tokenization, embedding and context modeling of the text, so that the model can effectively identify key information and structure in the text. For example, the model can understand the product name, quantity, price and other related information in the order, and generate an initial text representation.

[0100] As described in step S0013 above, each of the temporary text data is identified by a preset BiLSTM-CRF model to obtain corresponding order text data of each of the temporary text data. BiLSTM-CRF model is a model combining BiLSTM (Bidirectional Long Short-Term Memory Network) and CRF (Conditional Random Field), which is suitable for processing Named Entity Recognition (NER) tasks in text. First, the BiLSTM model learns the word vectors in each temporary text data and considers the context information to capture the dependency between words. The bidirectional nature means that the network not only focuses on the content before the current word, but also on the words after it, so that it can more accurately understand the specific meaning of the words. Then, the CRF layer considers the feature probability distribution obtained by BiLSTM to optimize the selection of labels and ensure that the recognized entities follow a reasonable structure.

[0101] In one embodiment, before step S5 of desensitizing the feature data in the region data set according to the preset desensitization level and desensitization strategy correspondence table to obtain the desensitized target region data set, the method further comprises:

[0102] S401: Obtain the scene information corresponding to the region data set of the specified region to be optimized;

[0103] S402: Calculate the correlation value of each feature data in the region data set and the scene information;

[0104] S403: Record the feature data with a correlation value greater than the correlation threshold value as relevant feature data;

[0105] S404: Set the desensitization level for the relevant feature data.

[0106] As described in step S401 above, obtain the scene information corresponding to the region data set of the specified region to be optimized. The process of obtaining scene information involves data collection and research, including business-related documents, standard operating procedures, historical case analysis, etc. It can also extract scene tag data from logistics management systems. That is, through communication with business-related personnel or extracting data from existing management information systems, scene information closely related to specific business logic can be obtained. In addition, these scene information helps to enhance the effectiveness of subsequent data analysis, because different business scenarios have a significant impact on the sensitivity, value and application of some data. For example, in the "Middle East Ramadan preparation" scenario, the "Jeddah port daily throughput" data sensitivity rating is A level; in the "European Christmas logistics" scenario, the "Frankfurt airport cargo volume" sensitivity is reduced to B level.

[0107] As described in step S402 above, calculate the correlation value of each feature data in the region data set and the scene information. Correlation value usually uses statistical indicators, such as Pearson correlation coefficient, mutual information or other algorithms suitable for evaluating the relationship between features and scene information. By constructing mathematical models, the system can analyze the relationship between each feature data and the scene information, identify which features can provide more important information in the current scenario, or play a greater role in making decisions. For example, a certain feature such as transportation time may be a key indicator in a particular scenario, while in other scenarios it has a lower marginal impact.

[0108] As described in step S403 above, record the feature data with a correlation value greater than the correlation threshold value as relevant feature data. By setting the correlation threshold value, those feature data that do not have significant meaning in the current scenario are removed, focusing on data that is more representative and contextually meaningful. The screening of relevant feature data not only improves the efficiency of subsequent data processing, but also ensures that important data is not over-processed during desensitization, thereby affecting the accuracy and effectiveness of business decisions.

[0109] As described in step S404, the desensitization level is set for the relevant feature data. The system sets the desensitization level for the feature data that has been marked as relevant. The setting of the desensitization level is usually determined according to the evaluation of the criticality of the feature data to the business, the requirements of laws and regulations, and industry standards. For example, for features containing personal customer information, payment information, or other highly sensitive data, a high desensitization requirement can be set; while some data such as commodity categories or transportation methods can be set to a lower desensitization level.

[0110] In one embodiment, the step S5 of desensitizing the feature data in the region data set according to the preset desensitization level and desensitization strategy correspondence table to obtain the target region data set comprises:

[0111] S511: Desensitize the feature data in the region data set according to the preset desensitization level and desensitization strategy correspondence table to obtain a temporary region data set.

[0112] S512: Add Gaussian noise to the feature data in the temporary region data set with a desensitization level lower than the preset level to obtain the target region data set.

[0113] As described in step S511, the feature data in the region data set is desensitized according to the preset desensitization level and desensitization strategy correspondence table to obtain a temporary region data set. First, the preset desensitization level and desensitization strategy correspondence table defines the sensitivity level of each data feature and the corresponding desensitization processing method. This mapping relationship ensures that different types of data are properly processed. According to the desensitization level of the feature data, the system will use different desensitization techniques. For example, for high sensitivity data, complete de-identification processing, data replacement, or encryption methods can be used, while for low sensitivity data, only simple de-identification or blurring processing may be required. During processing, the system calls the corresponding algorithm to desensitize each feature. Through this desensitization process, the sensitivity of the feature data is reduced, so that the availability of the data can be preserved to some extent while protecting personal privacy. For example, in customer information, the real name may be replaced by characters or a randomly generated alternative name, while the specific amount in the order information may be ranged.

[0114] As described in step S512, the feature data with a preset desensitization level lower than the preset level in the temporary region data set is added with Gaussian noise to obtain the target region data set. Gaussian noise is a random signal or interference conforming to Gaussian (normal) distribution. A function for generating Gaussian noise is selected, and the mean and variance are set to obtain the completed Gaussian noise. For features with a lower desensitization level, the importance and sensitivity are relatively small, and therefore, the individual information can be effectively hidden by adding Gaussian noise to these features, the accuracy of the information is reduced, and the risk of data leakage in the analysis process is reduced. For example, assuming that a feature is the age information of a user, if the desensitization level is low, a normally distributed noise can be added to the actual age to make the final age information more blurred, but the statistical distribution can still be reflected.

[0115] The application further provides a cross-border logistics data collaborative optimization device, which comprises:

[0116] The optimization region data input module 902 is configured to input the feature data in the region data set of a specified region to be optimized into a preset prediction network to obtain a first prediction value corresponding to each feature data, and perform preliminary desensitization on the feature data in the region data set to obtain a desensitized region data set after desensitization; and the region data set comprises a plurality of feature data.

[0117] The desensitized feature data input module 904 is configured to input the desensitized feature data in the desensitized region data set into the preset prediction network to obtain a second prediction value corresponding to each desensitized feature data.

[0118] The prediction error calculation module 906 is configured to calculate a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to each feature data.

[0119] The desensitization level acquisition module 908 is configured to acquire a desensitization level of each feature data according to the prediction error.

[0120] The desensitization processing module 910 is configured to perform desensitization processing on the feature data in the region data set according to a preset desensitization level and desensitization strategy correspondence table to obtain a target region data set after desensitization.

[0121] The optimization scheme acquisition module 912 is configured to input target desensitized feature data in the target region data set into a preset global model to obtain an optimization scheme of the region to be optimized.

[0122] In one embodiment, the cross-border logistics data collaborative optimization device comprises:

[0123] The model gradient parameter acquisition module is configured to acquire model gradient parameters and feature contribution degrees uploaded by a plurality of preset local models.

[0124] an aggregation weight adjustment module configured to adjust an aggregation weight of a model gradient parameter of each of the preset local models based on a feature contribution degree of the preset local model;

[0125] a preset global model acquisition module configured to aggregate the model gradient parameters of each of the preset local models based on the corresponding aggregation weights to obtain the preset global model.

[0126] In one embodiment, the cross-border logistics data collaborative optimization device comprises:

[0127] an offset value calculation module configured to calculate an offset value corresponding to each of the model gradient parameters based on all the model gradient parameters;

[0128] a target model gradient parameter acquisition module configured to set a model gradient parameter noise corresponding to each of the model gradient parameters according to the offset value corresponding to the model gradient parameter to obtain a target model gradient parameter of each of the preset local models.

[0129] In one embodiment, the cross-border logistics data collaborative optimization device comprises:

[0130] an order text data acquisition module configured to acquire a plurality of order text data of the specified region to be optimized;

[0131] a region data set formation module configured to extract feature data in each of the order text data through a preset graph attention network to form the region data set.

[0132] In one embodiment, the cross-border logistics data collaborative optimization device comprises:

[0133] an order document acquisition module configured to acquire a plurality of order documents of the specified region to be optimized;

[0134] a first text data acquisition module configured to recognize each of the order documents through a preset XLM-RoBERTa model to obtain temporary text data corresponding to each of the order documents;

[0135] a second text data acquisition module configured to perform entity recognition on each of the temporary text data through a preset BiLSTM-CRF model to obtain the order text data corresponding to each of the temporary text data.

[0136] In one embodiment, the cross-border logistics data collaborative optimization device comprises:

[0137] a scene information acquisition module configured to acquire scene information corresponding to the region data set of the specified region to be optimized;

[0138] The correlation value calculation module is used to calculate the correlation value between each feature data in the regional dataset and the scene information;

[0139] The relevant feature data recording module is used to record feature data with a relevance value greater than the relevance threshold as relevant feature data;

[0140] The desensitization level setting module is used to set the desensitization level for the relevant feature data.

[0141] In one embodiment, the desensitization processing module 910 includes:

[0142] The feature data desensitization submodule is used to desensitize the feature data in the regional dataset according to a preset desensitization level and desensitization strategy correspondence table to obtain a temporary regional dataset.

[0143] The target region dataset acquisition submodule is used to add Gaussian noise to feature data in the temporary region dataset that has a preset desensitization level lower than a preset level, so as to obtain the target region dataset.

[0144] Figure 4 An internal structural diagram of an electronic device in one embodiment is shown. This electronic device can specifically be a terminal or a server, and more specifically, a computer device. Figure 4 As shown, the electronic device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a cross-border logistics data collaborative optimization method. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to implement the cross-border logistics data collaborative optimization method. Those skilled in the art will understand that… Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0145] In one embodiment, an electronic device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps:

[0146] The feature data of the specified region dataset to be optimized is input into a preset prediction network to obtain the first predicted value corresponding to each feature data, and the feature data in the region dataset is initially desensitized to obtain the desensitized region dataset; wherein, the region dataset includes multiple feature data;

[0147] The desensitized feature data in the desensitized region dataset is input into the preset prediction network to obtain a second predicted value corresponding to each desensitized feature data.

[0148] Calculate the prediction error for each feature data based on the first and second predicted values ​​corresponding to each feature data.

[0149] The desensitization level of each feature data is obtained based on the prediction error;

[0150] According to the preset desensitization level and desensitization strategy correspondence table, the feature data in the regional dataset is desensitized to obtain the desensitized target regional dataset.

[0151] The target desensitization feature data in the target region dataset is input into a preset global model to obtain the optimization scheme for the region to be optimized.

[0152] The system generates a first predicted value and a second predicted value after preliminary anonymization based on the feature data in the regional dataset. The anonymization level of each feature data is obtained by calculating the prediction error. According to the preset anonymization level and anonymization strategy correspondence table, corresponding anonymization strategies are formulated for different feature data to ensure that data validity is preserved to the greatest extent while protecting data privacy. This achieves the ability to dynamically adjust the anonymization granularity and effectively solves the shortcomings of traditional static anonymization methods in adjusting data value.

[0153] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following steps:

[0154] The feature data of the specified region dataset to be optimized is input into a preset prediction network to obtain the first predicted value corresponding to each feature data, and the feature data in the region dataset is initially desensitized to obtain the desensitized region dataset; wherein, the region dataset includes multiple feature data;

[0155] The desensitized feature data in the desensitized region dataset is input into the preset prediction network to obtain a second predicted value corresponding to each desensitized feature data.

[0156] Calculate the prediction error for each feature data based on the first and second predicted values ​​corresponding to each feature data.

[0157] obtain a desensitization level of each feature data according to the prediction error;

[0158] perform desensitization processing on the feature data in the region data set according to a preset desensitization level and desensitization strategy correspondence table, to obtain a target region data set after desensitization;

[0159] input target desensitization feature data in the target region data set into a preset global model, to obtain an optimization scheme for the region to be optimized.

[0160] generate a first prediction value and a second prediction value after preliminary desensitization based on the feature data in the region data set, obtain a desensitization level of each feature data by calculating a prediction error, and formulate a corresponding desensitization strategy for different feature data according to a preset desensitization level and desensitization strategy correspondence table, so as to ensure that data validity is maximally retained on the premise of guaranteeing data privacy, and realize the ability of dynamically adjusting desensitization granularity, effectively solving the deficiency of traditional static desensitization methods in data value adjustment.

[0161] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The program can be stored in a non-volatile computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.

[0162] The technical features of the above embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present disclosure.

[0163] The above-described embodiments are merely illustrative of several embodiments of the present application, which are described in more detail and in a specific manner, but should not be construed as limiting the scope of the patent of the present application. It should be noted that, for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for cross-border logistics data collaborative optimization, characterized in that, The method comprises: inputting feature data in a region data set specifying a region to be optimized into a preset prediction network to obtain a first prediction value corresponding to each feature data, and preliminarily desensitizing the feature data in the region data set to obtain a desensitized region data set after desensitization; wherein the region data set comprises a plurality of feature data, and the region data set is formed in the following manner: a plurality of order text data of the specified region to be optimized are acquired; feature data in each order text data is extracted through a preset graph attention network to form the region data set, the preset prediction network is an LSTM prediction network, and the preset prediction network comprises an input layer, a hidden layer and an output layer, wherein the input layer is used to receive the region data set, the hidden layer is used to capture the time sequence dependency in the region data set, and the output layer is used to predict future freight demand; inputting desensitized feature data in the desensitized region data set into the preset prediction network to obtain a second prediction value corresponding to each desensitized feature data; calculating a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to each feature data; obtaining a desensitization level of each feature data according to the prediction error; desensitizing the feature data in the region data set according to a preset desensitization level and desensitization strategy correspondence table to obtain a temporary region data set; adding Gaussian noise to feature data with a preset desensitization level lower than a preset level in the temporary region data set to obtain a target region data set; inputting target desensitized feature data in the target region data set into a preset global model to obtain an optimization scheme for the region to be optimized; Before the step of inputting the target desensitized feature data in the target region data set into the preset global model to obtain the optimization scheme for the region to be optimized, the method further comprises: acquiring model gradient parameters and feature contribution degrees uploaded by a plurality of preset local models; adjusting the aggregation weight of the model gradient parameters of each preset local model based on the feature contribution degree of the preset local model; calculating an offset value corresponding to each model gradient parameter based on all model gradient parameters; setting model gradient parameter noise corresponding to each model gradient parameter based on the offset value corresponding to each model gradient parameter to obtain target model gradient parameters of each preset local model; aggregating the model gradient parameters of each preset local model and the corresponding aggregation weight to obtain the preset global model.

2. The cross-border logistics data collaborative optimization method according to claim 1, characterized in that, The step of acquiring a plurality of order text data of the specified region to be optimized comprises: acquiring a plurality of order documents of the specified region to be optimized; identifying each order document through a preset XLM-RoBERTa model to obtain temporary text data corresponding to each order document; performing entity recognition on each temporary text data through a preset BiLSTM-CRF model to obtain order text data corresponding to each temporary text data, respectively.

3. A cross-border logistics data collaborative optimization device, characterized in that, The device comprises: The optimization area data input module is configured to input feature data in a region data set of a specified optimization area to a preset prediction network to obtain a first prediction value corresponding to each feature data, and to perform preliminary desensitization on the feature data in the region data set to obtain a desensitized region data set after desensitization; the region data set includes a plurality of feature data, and the region data set is formed in the following manner: a plurality of order text data of the specified optimization area are obtained; feature data in each order text data is extracted by a preset graph attention network to form the region data set, the preset prediction network is an LSTM prediction network, and the preset prediction network includes an input layer, a hidden layer, and an output layer, wherein the input layer is configured to receive the region data set, the hidden layer is configured to capture a time sequence dependency in the region data set, and the output layer is configured to predict future freight demand; The desensitization feature data input module is configured to input desensitization feature data in the desensitized region data set to the preset prediction network to obtain a second prediction value corresponding to each desensitization feature data; The prediction error calculation module is configured to calculate a prediction error of each feature data according to the first prediction value and the second prediction value corresponding to each feature data; The desensitization level acquisition module is configured to acquire a desensitization level of each feature data according to the prediction error; The desensitization processing module includes: a feature data desensitization submodule configured to desensitize the feature data in the region data set according to a preset desensitization level and a desensitization strategy correspondence table to obtain a temporary region data set; and a target region data set acquisition submodule configured to add Gaussian noise to feature data in the temporary region data set with a preset desensitization level lower than a preset level to obtain the target region data set; The optimization scheme acquisition module is configured to input target desensitization feature data in the target region data set to a preset global model to obtain an optimization scheme for the optimization area; The model gradient parameter acquisition module is configured to acquire model gradient parameters and feature contribution degrees uploaded by a plurality of preset local models; The aggregation weight adjustment module is configured to adjust an aggregation weight of the corresponding model gradient parameter based on the feature contribution degree of each preset local model; The offset value calculation module is configured to calculate an offset value corresponding to each model gradient parameter based on all model gradient parameters; The target model gradient parameter acquisition module is configured to set a model gradient parameter noise corresponding to each model gradient parameter according to the offset value corresponding to each model gradient parameter to obtain a target model gradient parameter of each preset local model; The preset global model acquisition module is configured to aggregate the model gradient parameters of each preset local model and the corresponding aggregation weight to obtain the preset global model.

4. A computer-readable storage medium, characterized in that, A computer program is stored, and the computer program is executed by a processor to enable the processor to perform the steps of the cross-border logistics data collaborative optimization method according to any one of claims 1 to 2.

5. An electronic device, comprising: The device comprises a memory and a processor, the memory stores a computer program, the computer program is executed by the processor, so that the processor executes the steps of the cross-border logistics data collaborative optimization method as claimed in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Log desensitization method and device, computer equipment and storage medium

    CN119203211A

  • Digital person business card interaction method, electronic equipment and storage medium

    CN119398728A