Logistics mode recognition method and device, computer equipment and storage medium

By acquiring logistics characteristic data and using clustering algorithms and Transformer models to automatically identify logistics development patterns, the data processing challenges of traditional logistics pattern analysis have been solved, enabling rapid and accurate identification of logistics development patterns and resource optimization.

CN121434840APending Publication Date: 2026-01-30GUANGDONG TRANSPORTATION PLANNING RES CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511307224.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-13
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Traditional logistics model analysis relies on manual surveys and experience-based judgments, making it difficult to quickly and comprehensively process massive amounts of complex data. Existing technologies have limitations in data integration and pattern recognition, and cannot dynamically adapt to the logistics development needs of different regions.

Method used

By acquiring logistics characteristic data, clustering algorithms are used to perform clustering, and the cluster centers are determined by combining the Calinski-Harabasz index and density peak detection algorithm. The Transformer classification model is then trained to achieve automatic identification of logistics development patterns.

Benefits of technology

It enables the classification and modeling of regional logistics development models within seconds, improving decision-making efficiency and accuracy, and providing strong support for logistics resource optimization and network design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434840A_ABST
    Figure CN121434840A_ABST
Patent Text Reader

Abstract

The invention relates to a logistics mode recognition method and device, computer equipment and a storage medium. The method comprises the following steps: obtaining logistics feature data of a target area; wherein the logistics feature data comprises at least one of distribution node data, policy related class data, infrastructure class data, operation efficiency class data and industry related class data; clustering the logistics feature data based on a clustering algorithm to obtain target clustering data; wherein each target cluster in the target clustering data is in one-to-one correspondence with the logistics development mode; training a classification recognition model by using the target clustering data to obtain a target recognition model; wherein the target identification model is used for identifying the logistics development mode of the to-be-identified area. Through the method, a three-dimensional logistics development evaluation system is constructed, automatic identification and classification modeling of the regional logistics development mode are realized, and the decision-making efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of logistics management technology, and in particular to a logistics pattern recognition method, apparatus, computer equipment, and storage medium. Background Technology

[0002] Against the backdrop of global economic integration and the rapid development of digital technology, the logistics industry, as a key link connecting production and consumption, plays a crucial role in enhancing regional economic competitiveness through the accurate identification and optimization of its development model. For example, rural logistics development is an important part of rural revitalization, and the rational division of rural logistics development models has a positive effect on promoting the refined operation and management of rural logistics.

[0003] Traditional logistics model analysis often relies on manual surveys and experience-based judgments or factor analysis, which struggles to quickly and comprehensively process massive amounts of complex data. With the increasing complexity of regional logistics networks, logistics characteristic data exhibits multi-source and heterogeneous characteristics. Existing technologies have limitations in data integration and pattern recognition, making it difficult to dynamically adapt to the logistics development needs of different regions. Therefore, a data-driven intelligent solution is urgently needed to achieve accurate analysis and prediction of logistics development models.

[0004] Therefore, how to extract potential patterns from complex logistics data and achieve scientific classification and intelligent identification of logistics development models has become an urgent problem to be solved in order to reduce costs and increase efficiency in the logistics industry and support high-quality regional economic development. Summary of the Invention

[0005] Therefore, it is necessary to provide a logistics pattern recognition method, apparatus, computer equipment, and storage medium to address the aforementioned technical problems.

[0006] In a first aspect, this application provides a logistics pattern recognition method, the method comprising: Obtain logistics characteristic data for the target area; wherein, the logistics characteristic data includes at least one of the following: delivery node data, policy-related data, infrastructure data, operational efficiency data, and industry-related data; The logistics feature data is clustered based on a clustering algorithm to obtain target cluster data; wherein, each target cluster in the target cluster data corresponds one-to-one with a logistics development model; A classification and recognition model is trained using the target clustered data to obtain a target recognition model; wherein, the target recognition model is used to identify the logistics development pattern of the area to be identified.

[0007] In one embodiment, the clustering of the logistics feature data based on a clustering algorithm to obtain target cluster data includes: An initial clustering range is preset. For each number of clusters included in the initial clustering range, the logistics feature data is clustered based on a clustering algorithm, and the Calinski-Harabasz index is calculated. The number of target clusters is determined based on the Calinski-Harabasz index. Based on the target number of clusters, the target cluster data is obtained using the clustering algorithm.

[0008] In one embodiment, determining the target number of clusters based on the Calinski-Harabasz index includes: After each iteration, the contour coefficients and the Calinski-Harabasz index are calculated synchronously. If the rate of change of both the silhouette coefficient and the Calinski-Harabasz exponent is less than a predetermined value, the number of clusters corresponding to the current iteration is determined as the target number of clusters.

[0009] In one embodiment, the method further includes: During the clustering process of the logistics feature data, candidate centroids are obtained based on the K-means algorithm; For each candidate center point, the local density and minimum distance are calculated based on the density peak detection algorithm; wherein, the local density characterizes the density of data points around the candidate center point; and the minimum distance indicates the shortest distance between the candidate center point and any data point with a higher density than the candidate center point. The weight of each candidate center point is determined based on the local density and the minimum distance; Target cluster centers are selected from the candidate center points according to the weights. The step of obtaining the target clustering data using the clustering algorithm based on the target cluster number includes: The target clustering data is obtained using the clustering algorithm based on the target cluster centers and the target number of clusters.

[0010] In one embodiment, acquiring the logistics characteristic data of the target area includes: Obtain the original logistics characteristic data of the target area; The original logistics feature data is normalized and dimensionality reduced to obtain the preprocessed logistics feature data.

[0011] In one embodiment, the method further includes: The target cluster data is divided into a training set, a test set, and a validation set; The validation set is input into the trained classification and recognition model to obtain the validation and recognition results; The classification and recognition model trained after the verification and recognition results are consistent with the expected results is determined as the target recognition model.

[0012] In one embodiment, determining the trained classification and recognition model, in which the verification and recognition results are consistent with the expected results, as the target recognition model includes: Based on the verification and identification results, the evaluation index results are determined; wherein, the evaluation index results include at least one of classification accuracy, recall, confusion matrix, index score and receiver operating characteristic curve; The training model whose evaluation index results are consistent with the expected results is the target recognition model.

[0013] Secondly, this application also provides a logistics pattern recognition device, the device comprising: The acquisition module is used to acquire logistics characteristic data of the target area; wherein, the logistics characteristic data includes at least one of delivery node data, policy-related data, infrastructure data, operational efficiency data, and industry-related data; the clustering module is used to cluster the logistics characteristic data based on a clustering algorithm to obtain target cluster data; wherein, each target cluster in the target cluster data corresponds one-to-one with a logistics development model; The modeling module is used to train a classification and recognition model using the target cluster data to obtain a target recognition model; wherein, the target recognition model is used to identify the logistics development pattern of the area to be identified.

[0014] Thirdly, this application also provides a computer device, including a processor and a memory for storing a computer program of the processor; wherein the processor is configured to, when executing the computer program, implement the steps of the method described in any embodiment of this application.

[0015] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the methods described in any embodiment of this application.

[0016] The aforementioned logistics pattern recognition method, on the one hand, integrates multi-source data such as delivery nodes, policies and regulations, infrastructure, operational efficiency, and industrial linkages to construct a three-dimensional logistics development evaluation system, avoiding the one-sidedness of single-indicator analysis and comprehensively depicting the characteristics of logistics development; on the other hand, it uses clustering algorithms to reduce the dimensionality and group multi-dimensional logistics data, automatically identifying hidden relationships between data. The trained target recognition model can quickly classify the logistics characteristic data of a new region and output its corresponding logistics development mode in seconds. In this way, it realizes the automatic identification and classification modeling of regional logistics development modes, which not only improves decision-making efficiency and accuracy, but also provides strong support for logistics resource optimization, policy formulation, and logistics network design. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a logistics pattern recognition method according to an exemplary embodiment; Figure 2 This is a schematic diagram illustrating the evaluation of K-means clustering results according to an exemplary embodiment; Figure 3 This is a schematic diagram of a training set confusion matrix according to an exemplary embodiment; Figure 4 This is a schematic diagram of a test set confusion matrix according to an exemplary embodiment; Figure 5 This is a flowchart illustrating a logistics pattern recognition method according to another exemplary embodiment; Figure 6 This is a structural block diagram of a logistics pattern recognition device according to an exemplary embodiment; Figure 7 This is an internal structural diagram of a computer device according to an exemplary embodiment. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] The terms "first," "second," and "third" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0020] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0021] In some embodiments, the logistics pattern recognition method provided in this application can be applied to a computer device. The computer device can be any mobile terminal or fixed terminal. The terminal can be a device that provides voice and / or data connectivity to a user. For example, the terminal can be an Internet of Things (IoT) terminal, such as a sensor device, a mobile phone or so-called "cellular" phone, and a computer with an IoT terminal, for example, it can be a fixed, portable, pocket-sized, handheld, or computer-embedded device.

[0022] In some embodiments, such as Figure 1 As shown, a logistics pattern recognition method is provided, the method comprising the following steps: S101, Obtain logistics characteristic data of the target area; wherein, the logistics characteristic data includes at least one of delivery node data, policy-related data, infrastructure data, operational efficiency data, and industry-related data.

[0023] In this embodiment of the application, the delivery node data may include, but is not limited to, at least one of the following: node name / identification data (used to identify different logistics centers, warehouses, sorting centers or delivery stations, etc.), geospatial data (address name, latitude and longitude), time dimension data (current time, service time window, peak period), contact data (contact phone number, email address), operational capacity data (indicating the node's ability to handle goods, such as daily processing volume, storage capacity), inventory status data (type and quantity of each type of goods), and abnormal record data (delay, lost goods).

[0024] In this embodiment of the application, policy-related data may include, but is not limited to, at least one of the following: regional logistics subsidy data, traffic control data, tax incentive data, market entity cultivation data, industry standard and specification data, and logistics technology innovation data.

[0025] In this application embodiment, infrastructure data may include, but is not limited to, at least one of the following: transportation network infrastructure data (railway / highway / waterway / air transport mileage, freight station scale, logistics lines), port throughput / quantity data, warehousing facility data (warehouse area / structure, automated warehousing coverage rate), and the popularity rate of logistics information platforms.

[0026] In this embodiment of the application, operational efficiency data may include, but is not limited to, trunk transportation efficiency data (vehicle load factor, empty run rate, transportation cost, on-time rate), multimodal transport efficiency data (transfer time, proportion of intermodal orders), warehousing operation efficiency data (warehouse utilization rate, order processing efficiency, inbound and outbound timeliness), inventory turnover efficiency data (inventory turnover rate, proportion of stagnant inventory, stockout rate), urban delivery efficiency data (delivery cost, delivery response time), rural delivery efficiency data (township delivery frequency, delivery cost), and automated equipment efficiency data (equipment utilization rate, sorting equipment failure rate).

[0027] In this embodiment of the application, industry-related data may include, but is not limited to, at least one of the following: agricultural production data (agricultural output value scale, production area distribution), agricultural product circulation efficiency (agricultural product loss rate, fresh food e-commerce penetration rate, supply chain length), and logistics financial service data (supply chain finance scale, insurance penetration rate).

[0028] S102, cluster the logistics feature data based on the clustering algorithm to obtain target cluster data; wherein, each target cluster in the target cluster data corresponds one-to-one with the logistics development model.

[0029] In this embodiment, clustering algorithms are unsupervised learning techniques in data mining and machine learning, revealing the inherent structure and distribution patterns of data. Clustering algorithms divide samples in a dataset into multiple clusters based on feature similarity, resulting in high similarity among data objects within the same cluster, while data objects in different clusters exhibit significant differences.

[0030] In some embodiments, the clustering algorithm may include, but is not limited to, at least one of the following: K-means algorithm, K-medoids algorithm, hierarchical clustering algorithm (Agglomerative Nesting, AGNES), and density-based clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, DBSCAN).

[0031] In some embodiments, clustering the logistics feature data based on a clustering algorithm to obtain target cluster data includes: Determine the number of clusters; Based on the number of clusters, the logistics feature data is clustered using the clustering algorithm to obtain the target clustered data.

[0032] In some embodiments, the computer device can iterate through the K values ​​(number of clusters) according to the elbow rule, calculate the sum of squared errors (SSE) within each cluster corresponding to each K value, plot the K-SSE curve, and select the K value at the inflection point as the number of clusters, i.e. the position where the SSE decreases significantly; or, it can be combined with a logistics cost model to replace SSE with a joint indicator of "transportation cost-loading rate".

[0033] In some embodiments, the computer device may use the silhouette coefficient method to determine the silhouette coefficient based on the average distance from a sample within each cluster to other samples within the same cluster and the average distance from a sample to the nearest cluster; calculate the average of each silhouette coefficient to obtain the average silhouette coefficient; and determine the number of clusters when the average silhouette coefficient is maximized by the value of K.

[0034] S103, a classification and recognition model is trained using the target cluster data to obtain a target recognition model; wherein, the target recognition model is used to identify the logistics development pattern of the area to be identified.

[0035] In some embodiments, the classification and recognition model may include, but is not limited to, at least one of the following: logistic regression model, decision tree model, random forest model, multilayer perceptron (MLP) and Transformer classification model.

[0036] In some embodiments, the classification and recognition model is a Transformer classification model. The computer device can hierarchically divide the target clustered data into a training set (80%) and a test set (20%) according to cluster labels; the classification and recognition model is trained using the training set, treating the normalized feature vector of each sample in the training set as a sequence of length T (T being the number of features), with each feature serving as a time step in the sequence; each feature value is mapped to a high-dimensional space through a linear layer to form an input embedding vector; a multi-layer Transformer encoder is used, with each layer containing a multi-head self-attention and feedforward network; a global average pooling layer is added after the encoder output, followed by a fully connected layer to output the class probability; based on the class probability, the logistics development model is determined.

[0037] For example, the calculation formula in the Transformer classification model is as follows: FNN(X) = Rule(W1X+b1)W2+b2.

[0038] Where Q indicates the query vector, representing the logistics feature that needs to be focused on; K indicates the key vector, representing the key information in the logistics feature; and V indicates the value vector, representing the specific value of the logistics feature. Indicates the scaling factor used to adjust the scale of the dot product result; d k The query vector dimension is indicated by X; multi-dimensional logistics features (such as transportation cost, timeliness, capacity, etc.) are indicated by X; W1 indicates the first layer weights, which is the feature mapping matrix; W2 indicates the second layer weights, which is the classification mapping matrix; Rule indicates the activation function of the feedforward network; b1 and b2 indicate the offsets.

[0039] The aforementioned logistics pattern recognition method, on the one hand, integrates multi-source data such as delivery nodes, policies and regulations, infrastructure, operational efficiency, and industrial linkages to construct a three-dimensional logistics development evaluation system, avoiding the one-sidedness of single-indicator analysis and comprehensively depicting the characteristics of logistics development; on the other hand, it uses clustering algorithms to reduce the dimensionality and group multi-dimensional logistics data, automatically identifying hidden relationships between data. The trained target recognition model can quickly classify the logistics characteristic data of a new region and output its corresponding logistics development mode in seconds. In this way, it realizes the automatic identification and classification modeling of regional logistics development modes, which not only improves decision-making efficiency and accuracy, but also provides strong support for logistics resource optimization, policy formulation, and logistics network design.

[0040] In some embodiments, clustering the logistics feature data based on a clustering algorithm to obtain target cluster data includes: An initial clustering range is preset. For each number of clusters included in the initial clustering range, the logistics feature data is clustered based on a clustering algorithm, and the Calinski-Harabasz index is calculated. The number of target clusters is determined based on the Calinski-Harabasz index. Based on the target number of clusters, the target cluster data is obtained using the clustering algorithm.

[0041] In some embodiments, the computer device can set upper and lower limits for the initial clustering range based on the sample size; or, it can determine the initial clustering range based on known logistics development model types; for example, if rural logistics is divided into three levels of development, then k can be set. min =3,k max =6. For each cluster number k∈[k min k max The K-means++ algorithm is used to generate initial centroids, and K-means clustering is performed until convergence to obtain the target cluster label c. k ={c0, c2, ..., c k} and cluster center coordinates. For example, the clustering method is as follows:

[0042] Where k indicates the number of cluster centers; Indicates the number of cluster center points; z j Indicates the j-th observation point; indicates the centroid of cluster i.

[0043] For example, such as Figure 2 As shown, Figure 2 This is a schematic diagram illustrating the evaluation of K-means clustering results; the horizontal axis represents the number of clusters, and the vertical axis represents the Calinski-Harabasz (CH) index. It can be seen that the CH index is largest when the number of clusters is 3, thus the target number of clusters can be determined to be 3.

[0044] In some embodiments, determining the target number of clusters based on the Calinski-Harabasz index includes: synchronously calculating the silhouette coefficient and the Calinski-Harabasz index after each iteration; If the rate of change of both the silhouette coefficient and the Calinski-Harabasz exponent is less than a predetermined value, the number of clusters corresponding to the current iteration is determined as the target number of clusters.

[0045] In one embodiment, the computer device determines the silhouette coefficient based on the average distance from a sample within each cluster to other samples within the same cluster and the average distance from a sample to the nearest cluster; determines the Calinski-Harabasz (CH) index based on the intra-cluster dispersion and inter-cluster dispersion; and stops the iteration and determines the current k value as the target number of clusters when the rate of change of both the silhouette coefficient and the Calinski-Harabasz index is less than a predetermined value (e.g., 5%, 6%, 8%).

[0046] In this embodiment, the combination of silhouette coefficient and CH index can reduce the error of a single index (such as the CH index may increase monotonically with the number of clusters, while the silhouette coefficient may decrease after a certain point). The dual verification ensures that the clustering results are both "compact and separate" and "cohesive and exclusive", thus improving the robustness and reliability of the clustering results.

[0047] In some embodiments, the method further includes: During the clustering process of the logistics feature data, candidate centroids are obtained based on the K-means algorithm; For each candidate center point, the local density and minimum distance are calculated based on the density peak detection algorithm; wherein, the local density characterizes the density of data points around the candidate center point; and the minimum distance indicates the shortest distance between the candidate center point and any data point with a higher density than the candidate center point. The weight of each candidate center point is determined based on the local density and the minimum distance; Target cluster centers are selected from the candidate center points according to the weights. The step of obtaining the target clustering data using the clustering algorithm based on the target cluster number includes: The target clustering data is obtained using the clustering algorithm based on the target cluster centers and the target number of clusters.

[0048] In some embodiments, the computer device uses the K-means++ strategy to initialize cluster centers. Compared with random selection, the next center point is selected by probability distribution (inversely proportional to the distance of the selected center point), which effectively avoids the initial center points being too dense and improves the convergence speed. The K-means algorithm is executed, and the sample allocation and center point update are performed alternately until the center point position is stable or the preset number of iterations is reached. The final k center points are used as the candidate center point set.

[0049] In some embodiments, the computer device uses a truncated kernel function to calculate the local density of each candidate center point in the candidate center point set, and calculates the minimum distance of all centers with densities higher than their own; it uses a product weighting method to calculate the weight of the candidate center points based on the local density and the minimum distance; it arranges the weights of each candidate center point in descending order, selects the weight at the inflection point of the curve as the weight threshold, and filters out the candidate center points with weights greater than the weight threshold as the target cluster centers.

[0050] In this embodiment, the density peak detection algorithm is used to perform secondary screening of K-means candidate centroids, which effectively corrects the local optimum problem caused by improper selection of initial centroids. When processing large-scale, multi-dimensional logistics feature data, a balance between processing efficiency and stability is achieved, significantly improving the accuracy and practicality of logistics feature data clustering.

[0051] In some embodiments, acquiring logistics characteristic data of the target area includes: Obtain the original logistics characteristic data of the target area; The original logistics feature data is normalized and dimensionality reduced to obtain the preprocessed logistics feature data.

[0052] In this embodiment, normalization is used to unify features of different dimensions or orders of magnitude to a specific range, so as to avoid large interference to the model due to the large range of feature values.

[0053] In this embodiment, data dimensionality reduction refers to mapping high-dimensional data to a low-dimensional space using a mathematical method, while preserving as much information as possible from the original data. Data dimensionality reduction can reduce redundant information, improve model efficiency, remove noise, and facilitate visualization, among other benefits.

[0054] In some embodiments, normalization may include, but is not limited to, at least one of min-max scaling, Z-score standardization, and vector normalization.

[0055] For example, one way to normalize raw logistics characteristic data is as follows: Where x' indicates the original logistics characteristic data; x min Indicates the minimum value in the original logistics characteristic data; x max x indicates the maximum value in the original logistics feature data; x indicates the normalized logistics feature data.

[0056] In some embodiments, data dimensionality reduction may include, but is not limited to, at least one of principal component analysis (PCA) and linear discriminant analysis (LDA).

[0057] In this embodiment, data normalization can solve the problem of imbalance between indicator units and magnitudes, unify data scale, and improve the convergence stability of the algorithm. Furthermore, logistics feature data often contains highly correlated indicators (e.g., warehouse turnover rate and inventory turnover rate are highly positively correlated), which can introduce noise and interfere with the accurate identification of cluster boundaries. Data dimensionality reduction can compress the original features into a comprehensive indicator with low correlation, reduce redundant feature interference, reduce computational complexity, and improve the efficiency of cluster analysis.

[0058] In some embodiments, the method further includes: The target cluster data is divided into a training set, a test set, and a validation set; The validation set is input into the trained classification and recognition model to obtain the validation and recognition results; The classification and recognition model trained after the verification and recognition results are consistent with the expected results is determined as the target recognition model.

[0059] In this embodiment, the validation set can be used to adjust the model's hyperparameters and to perform a preliminary evaluation of the model's capabilities. The validation set can reduce the model's overfitting to the training data.

[0060] In some embodiments, the computer device may divide the target cluster data into a training set, a test set, and a validation set according to a preset ratio; wherein, the preset ratio may be any suitable ratio, for example, the preset ratio may be 3:1:1, 4:1:1, or 8:1:1, etc.

[0061] In some embodiments, determining the trained classification and recognition model that verifies the recognition result is consistent with the expected result as the target recognition model includes: Based on the verification and identification results, the evaluation index results are determined; wherein, the evaluation index results include at least one of classification accuracy, recall, confusion matrix, index score and receiver operating characteristic curve; The training model whose evaluation index results are consistent with the expected results is the target recognition model.

[0062] In this embodiment of the application, the elements of the confusion matrix include true positives, true negatives, false positives, and false negatives. Specifically, true positives (TP) indicate samples correctly classified in the positive class; true negatives (TN) indicate samples correctly classified in the negative class; false positives (FP) indicate samples predicted as positive in the negative class; and false negatives (FN) indicate samples predicted as negative in the positive class.

[0063] In one embodiment, the computer device may determine a first value based on the sum of true positives and true negatives; determine a second value based on the sum of true positives, true negatives, false positives, and false negatives; and determine the classification accuracy based on the ratio of the first value to the second value.

[0064] In one embodiment, the computer device may determine a third value based on the sum of true positives and false negatives; and determine the recall rate based on the ratio of true positives to the third value.

[0065] In some embodiments, the Receiver Operating Characteristic Curve (ROC) can be determined based on the False Positive Rate (FPR) and the True Positive Rate (TPR). The area under the ROC curve (AUC) can also be used to determine whether the indicator results are consistent with the expected results.

[0066] In one embodiment, the computer device can also obtain a training set confusion matrix based on the training set and a test set confusion matrix based on the test set. For example... Figure 3 and Figure 4 As shown, Figure 3 This is a schematic diagram of the confusion matrix for the training set. Figure 4 This is a schematic diagram of the confusion matrix for the test set.

[0067] In this embodiment, the trained classification and recognition model is validated and evaluated using a validation set, which can ensure the model's generalization ability and reduce overfitting. Furthermore, by evaluating the model's performance using multi-dimensional indicators, the model's performance can be analyzed from different perspectives, thereby improving classification reliability.

[0068] In this application embodiment, specific examples are provided below in conjunction with any of the above embodiments: Specific example 1: Figure 5 This is an exemplary flowchart illustrating the implementation of the logistics pattern recognition method provided in any embodiment of this application, such as... Figure 5As shown, the execution steps of this logistics pattern recognition method in a computer device are as follows: S501, Obtain the original logistics characteristic data of the target area.

[0069] In one alternative embodiment, the target area can be a rural area, and the original logistics characteristic data can be original rural logistics characteristic data.

[0070] S502, normalize the original logistics characteristic data to obtain logistics characteristic data.

[0071] S503, set the initial clustering range, and cluster the logistics feature data based on the K-means algorithm.

[0072] S504 determines the target number of clusters based on the Calinski-Harabasz criterion and obtains the target clustering data based on the target number of clusters.

[0073] In an optional embodiment, the Calinski-Harabasz index is calculated for each number of clusters (k value) within the initial clustering range; based on the Calinski-Harabasz index, the target number of clusters is determined; based on the target number of clusters, the clustering results (target clustering data) of the rural logistics development model for each county in the target region can be determined.

[0074] S505, based on the target clustered data, divides the training set and test set, and establishes a Transformer classification and recognition model.

[0075] S506. Select evaluation indicators to evaluate the model recognition results and verify the effectiveness of the model.

[0076] In one alternative embodiment, the evaluation metrics may include, but are not limited to, at least one of classification accuracy, recall, and confusion matrix.

[0077] In this embodiment, on the one hand, by integrating multi-source data such as delivery nodes, policies and regulations, infrastructure, operational efficiency, and industrial linkages, a three-dimensional logistics development evaluation system is constructed to avoid the one-sidedness of single-indicator analysis and comprehensively depict the characteristics of logistics development. On the other hand, clustering algorithms are used to reduce the dimensionality and group multi-dimensional logistics data, automatically identifying hidden relationships between data. The trained target recognition model can quickly classify the logistics characteristic data of new regions and output the logistics development mode to which it belongs in seconds. In this way, automatic identification and classification modeling of regional logistics development modes are realized, which not only improves decision-making efficiency and accuracy, but also provides strong support for logistics resource optimization, policy formulation, and logistics network design.

[0078] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0079] Based on the same inventive concept, this application also provides a logistics pattern recognition device for implementing the logistics pattern recognition method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more logistics pattern recognition device embodiments provided below can be found in the limitations of the logistics pattern recognition method described above, and will not be repeated here.

[0080] In one embodiment, such as Figure 6 As shown, a logistics pattern recognition device is provided, the device comprising: The acquisition module 10 is used to acquire logistics characteristic data of the target area; wherein, the logistics characteristic data includes at least one of delivery node data, policy-related data, infrastructure data, operational efficiency data, and industry-related data; the clustering module 20 is used to cluster the logistics characteristic data based on a clustering algorithm to obtain target cluster data; wherein, each target cluster in the target cluster data corresponds one-to-one with a logistics development model; The modeling module 30 is used to train a classification and recognition model using the target cluster data to obtain a target recognition model; wherein the target recognition model is used to identify the logistics development pattern of the area to be identified.

[0081] In one embodiment, the clustering module 20 includes: The processing unit is used to preset an initial clustering range, and for each number of clusters contained in the initial clustering range, to cluster the logistics feature data based on a clustering algorithm and calculate the Calinski-Harabasz index. The determining unit is used to determine the number of target clusters based on the Calinski-Harabasz index; A clustering unit is used to obtain the target cluster data using the clustering algorithm based on the target number of clusters.

[0082] In one embodiment, the determining unit is configured to perform the following steps: After each iteration, the contour coefficients and the Calinski-Harabasz index are calculated synchronously. If the rate of change of both the silhouette coefficient and the Calinski-Harabasz exponent is less than a predetermined value, the number of clusters corresponding to the current iteration is determined as the target number of clusters.

[0083] In one embodiment, the apparatus further includes: An initialization module is used to obtain candidate centroids based on the K-means algorithm during the clustering process of the logistics feature data. The first calculation module is used to calculate the local density and minimum distance for each candidate center point based on the density peak detection algorithm; wherein the local density characterizes the density of data points around the candidate center point; and the minimum distance indicates the shortest distance between the candidate center point and any data point with a higher density than the candidate center point. The second calculation module is used to determine the weight of each candidate center point based on the local density and the minimum distance; the filtering module is used to filter out the target cluster center from the candidate center points according to the weight. The clustering unit is used to obtain the target clustering data using the clustering algorithm based on the target cluster centers and the target number of clusters.

[0084] In one embodiment, the acquisition module 10 is configured to perform the following steps: Obtain the original logistics characteristic data of the target area; The original logistics feature data is normalized and dimensionality reduced to obtain the preprocessed logistics feature data.

[0085] In one embodiment, the apparatus further includes: The partitioning module is used to divide the target cluster data into a training set, a test set, and a validation set. The verification module is used to input the verification set into the trained classification and recognition model to obtain the verification and recognition results; An evaluation module is used to determine the trained classification and recognition model that matches the verification and recognition results with the expected results as the target recognition model.

[0086] In one embodiment, the evaluation module is configured to perform the following steps: Based on the verification and identification results, the evaluation index results are determined; wherein, the evaluation index results include at least one of classification accuracy, recall, confusion matrix, index score and receiver operating characteristic curve; The training model whose evaluation index results are consistent with the expected results is the target recognition model.

[0087] Each module in the aforementioned logistics pattern recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in the processor of a computer device in hardware form or independent of the processor, or it can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0088] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, communication interface, display unit, and input device connected via a method bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores operating methods and computer programs. The internal memory provides an environment for the operation of the operating methods and computer programs stored in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an image processing method. The display screen can be an LCD screen or an e-ink display screen. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0089] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0090] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0091] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps performed by the processor of the computer device of any of the above.

[0092] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0093] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, compilable logic units, quantum computing-based data processing logic units, etc., and are not limited to these.

[0094] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0095] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A logistic pattern recognition method, characterized by, The method comprises: acquiring logistics feature data of a target region; wherein the logistics feature data comprises at least one of delivery node data, policy-related data, infrastructure data, operation efficiency data, and industry correlation data; performing clustering on the logistics feature data based on a clustering algorithm to obtain target cluster data; wherein each target cluster in the target cluster data corresponds to a logistics development mode; training a classification recognition model using the target cluster data to obtain a target recognition model; wherein the target recognition model is used to identify a logistics development mode of a region to be identified.

2. The method of claim 1, wherein, The method further comprises: presetting an initial clustering range, and for each clustering number contained in the initial clustering range, performing clustering on the logistics feature data based on a clustering algorithm and calculating a Calinski-Harabasz index; determining a target clustering number based on the Calinski-Harabasz index; obtaining the target cluster data using the clustering algorithm according to the target clustering number.

3. The method of claim 2, wherein, The method further comprises: synchronously calculating a silhouette coefficient and the Calinski-Harabasz index after each iteration; in the case where the change rates of the silhouette coefficient and the Calinski-Harabasz index are both less than a predetermined value, determining that the clustering number corresponding to the current iteration is the target clustering number.

4. The method of claim 2, wherein, The method further comprises: acquiring candidate center points based on a K-means algorithm during the clustering of the logistics feature data; calculating local density and minimum distance for each candidate center point based on a density peak value detection algorithm; wherein the local density represents the density of data points around the candidate center point, and the minimum distance indicates the shortest distance between the candidate center point and any data point with higher density than the candidate center point; determining the weight of each candidate center point based on the local density and the minimum distance; selecting a target clustering center from the candidate center points according to the weight. The method further comprises: obtaining the target cluster data using the clustering algorithm according to the target clustering center and the target clustering number.

5. The method of claim 1, wherein, The method further comprises: acquiring original logistics feature data of the target region; performing normalization and data dimensionality reduction on the original logistics feature data to obtain preprocessed logistics feature data.

6. The method of claim 1, wherein, The method further comprises: dividing the target cluster data into a training set, a test set, and a validation set; inputting the validation set into the trained classification recognition model to obtain a validation recognition result; determining the trained classification recognition model whose validation recognition result is consistent with an expected result as the target recognition model.

7. The method of claim 6, wherein, The trained classification recognition model consistent with the verification recognition result and the expected result is determined as the target recognition model, including: According to the verification recognition result, an evaluation index result is determined; wherein the evaluation index result includes at least one of classification accuracy, recall rate, confusion matrix, index score and receiver operating characteristic curve; The trained classification recognition model consistent with the evaluation index result and the expected result is determined as the target recognition model.

8. A logistics pattern recognition apparatus characterized by comprising: The device includes: An acquisition module is configured to acquire logistics feature data of a target region; wherein the logistics feature data includes at least one of delivery node data, policy-related data, infrastructure data, operation efficiency data and industry-related data; A clustering module is configured to cluster the logistics feature data based on a clustering algorithm to obtain target clustering data; wherein each target cluster in the target clustering data corresponds to a logistics development mode one by one; A modeling module is configured to train a classification recognition model using the target clustering data to obtain a target recognition model; wherein the target recognition model is used to identify the logistics development mode of a to-be-identified region.

9. A computer device, comprising: The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.