A regional business environment evaluation and analysis system and method based on big data analysis

Through the big data analysis system, abnormal changes in logistics multi-source customer service data are identified, local outlier factors and abnormal evaluation scores are calculated, and decision support reports are generated. This solves the problems of traditional methods that cannot identify detailed anomalies and ignore after-sales data, thereby improving logistics efficiency and corporate satisfaction.

CN119784240BActive Publication Date: 2025-09-09BEIJING MINSHENG THINK TANK TECHNOLOGY INFORMATION CONSULTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411853221.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-09-09
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

Traditional anomaly detection methods cannot accurately identify abnormal changes in details in logistics data analysis, especially when processing high-dimensional data, they ignore local anomalies and lack analysis of after-sales logistics data, making it difficult to optimize logistics efficiency in a targeted manner.

Method used

A regional business environment evaluation and analysis system based on big data analysis is adopted. Multi-source customer service data is obtained through the information collection platform. The identification module is used to calculate the approximation and local outlier factor (LOF) of the data. The comprehensive abnormality evaluation score is calculated in combination with the evaluation calculation module. The output processing module generates a decision support report to optimize logistics processes and customer service.

Benefits of technology

It improves the flexibility and accuracy of anomaly detection, can quickly identify potential problems and risks, improve logistics efficiency and enterprise terminal satisfaction, and optimize regional operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784240B_ABST
    Figure CN119784240B_ABST
Patent Text Reader

Abstract

The present invention discloses a regional business environment evaluation and analysis system and method based on big data analysis, wherein the method comprises: collecting multi-source customer service data; identifying abnormal growth points in the data by calculating the approximation between adjacent segmented time nodes; evaluating the degree of abnormality of the data points by calculating the local outlier factor (LOF) of the sampled data; at the same time, in the above scheme, based on the LOF and the abnormal growth evaluation value, calculating the comprehensive abnormality evaluation score of each data, more intuitively comparing the abnormality degrees of different data points, and enhancing the operability of the analysis; comparing the comprehensive abnormality evaluation score with a preset threshold, identifying the results and generating a report to provide decision support, thereby helping to optimize the overall operational efficiency in the region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information data analysis, and in particular to a system and method for evaluating and analyzing a regional business environment based on big data analysis. Background Art

[0002] For large-scale manufacturing enterprises, logistics efficiency within their region and even across the broader region is a crucial factor influencing their profitability and business environment. Only through an efficient and smooth logistics system can the goods produced by these enterprises be quickly transferred from the production line to the consumer's hands and realize their value. Logistics efficiency directly determines the time cost from production to sales. An efficient logistics system can shorten product turnover cycles, accelerate inventory turnover, and reduce the risk of overstocking, thereby saving companies significant operating costs. Furthermore, rapid logistics response can enhance customer satisfaction and strengthen the market competitiveness of large-scale manufacturing enterprises. Clearly, logistics efficiency has become a crucial component of a company's supply chain.

[0003] In regional economic development, the improvement of logistics efficiency and regional logistics management evaluation can drive the development of related industries, such as warehousing, transportation, information processing, etc., forming an industrial agglomeration effect. It is the key to improving the regional economy and one of the important means to optimize the business environment and promote regional economic development.

[0004] Research has found that regional logistics companies (logistics terminals) generate a large amount of multi-source customer service data during their operations. This data includes order data, logistics terminal coverage data, logistics terminal throughput data, logistics terminal transportation route data, customer feedback order response data, logistics terminal cost and efficiency indicator data, etc. Anomaly detection in logistics based on this multi-source customer service data provides important information for operational optimization, customer management and efficiency improvement in the logistics industry.

[0005] However, traditional anomaly detection methods rely primarily on overall statistical analysis or rule-based analysis, lacking sensitivity to dynamic, abnormal changes at specific time points. In practice, due to the high volatility of logistics terminal data, traditional methods often fail to accurately identify abnormal changes in details, resulting in a delay in providing timely warnings of potential operational issues.

[0006] Furthermore, traditional anomaly detection methods often rely solely on global analysis when processing high-dimensional data, ignoring the importance of local anomalies within the data. While the Local Outlier Factor (LOF) method can detect anomalies in localized data in some cases, traditional methods have not effectively integrated it into multi-source data analysis frameworks, nor have they combined it with other factors, such as anomaly growth assessment, to comprehensively analyze the occurrence of anomalies.

[0007] At the same time, traditional anomaly detection often only targets pre-sales data such as multi-source customer service data, and does not provide sufficient analysis of post-sales logistics data, making it difficult to make targeted adjustments to logistics efficiency. Summary of the Invention

[0008] The purpose of the present invention is to provide a regional business environment evaluation and analysis system and method based on big data analysis, which solves the above-mentioned technical problems pointed out in the prior art.

[0009] The present invention provides a regional business environment evaluation and analysis system based on big data analysis, comprising: an information collection platform, an identification module, a sampling calculation module, an evaluation calculation module, and an output processing module;

[0010] The information collection platform is used to collect logistics multi-source customer service data at subdivided time nodes within a preset time period; the logistics multi-source customer service data includes order data, logistics terminal coverage data, logistics terminal throughput data, logistics terminal transportation route data, customer feedback order response data, logistics terminal fees, and logistics terminal efficiency indicator data;

[0011] The identification module is used to obtain the approximation of the logistics multi-source customer service data of each segmented time node based on the logistics multi-source customer service data of two adjacent segmented time nodes, and to obtain the abnormal growth evaluation value of each data p' in the logistics multi-source customer service data of each segmented time node based on the approximation of the logistics multi-source customer service data of each segmented time node. ;

[0012] The sampling calculation module is used to obtain the sampling data p of each data p' in the logistics multi-source customer service data of each segmented time node, and then calculate the local outlier factor based on the neighborhood k obtained by cross-validation of the sampling data p, and output the local outlier factor evaluation value LOF of the sampling data p ;

[0013] The evaluation calculation module is used to evaluate the local outlier factor LOF based on the sampled data p. and the abnormal growth evaluation value of each data , calculate the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node ;

[0014] The output processing module is used to perform logistics benefit evaluation processing operations based on the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node.

[0015] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages:

[0016] From the analysis of the regional business environment evaluation and analysis system and method based on big data analysis provided by the present invention, it can be seen that in specific applications, multi-source customer service data, including orders, logistics terminal information, customer feedback, etc., are first collected to provide comprehensive basic data support, and data preprocessing is used to ensure high quality and consistency of data, laying a solid foundation for subsequent analysis; the above-mentioned high-quality basic data can improve the accuracy of subsequent analysis and provide a reliable basis for anomaly detection; further, by calculating the approximation between adjacent segmented time nodes, abnormal growth points in the data can be identified, which can quickly discover potential problems and risks, and provide technical data support and clues for further analysis and decision-making; by calculating the local outlier factor (LOF) of the sampled data ), evaluates the degree of abnormality of data points. The unsupervised learning method effectively captures local anomalies in the data set, improving the flexibility and accuracy of anomaly detection. At the same time, in the above scheme, based on LOF and abnormal growth evaluation value, the comprehensive anomaly evaluation score of each data is calculated. The comprehensive evaluation score provides a quantitative indicator, which can more intuitively compare the degree of abnormality of different data points and enhance the operability of analysis. The comprehensive anomaly evaluation score is compared with the preset threshold to identify customers or logistics links with abnormal performance, and generate reports to provide decision support and help optimize logistics processes and customer service. By identifying anomalies and providing targeted suggestions, it can significantly improve logistics efficiency and enterprise terminal satisfaction, and help optimize the overall operational efficiency of the region. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a schematic diagram of the main process of a regional business environment evaluation and analysis method based on big data analysis;

[0018] Figure 2 Calculate the local density of each sample data p based on the optimized neighborhood k in a regional business environment evaluation analysis method based on big data analysis Schematic diagram of the simulation effect;

[0019] Figure 3 A schematic diagram of the operation steps for a logistics benefit evaluation process based on big data analysis of a regional business environment evaluation method by introducing logistics after-sales variable information and combining it with a comprehensive abnormality evaluation score;

[0020] Figure 4 The Spearman rank correlation coefficient is calculated in a regional business environment evaluation analysis method based on big data analysis. Schematic diagram of the operation steps;

[0021] Figure 5 This is a schematic diagram of the simulation of variables in a regional business environment evaluation analysis method based on big data analysis;

[0022] Figure 6A schematic diagram of the operational steps for obtaining the transfer efficiency evaluation coefficient of a logistics transfer station in a regional business environment evaluation and analysis method based on big data analysis;

[0023] Figure 7 Calculating the Pearson correlation coefficient for the sub-factor combination in a regional business environment evaluation method based on big data analysis Schematic diagram of the effect simulation. DETAILED DESCRIPTION

[0024] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0025] The present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings.

[0026] The present invention provides a regional business environment evaluation and analysis system based on big data analysis, comprising: an information collection platform, an identification module, a sampling calculation module, an evaluation calculation module, and an output processing module;

[0027] The information collection platform is used to collect logistics multi-source customer service data at subdivided time nodes within a preset time period; the logistics multi-source customer service data includes order data, logistics terminal coverage data, logistics terminal throughput data, logistics terminal transportation route data, customer feedback order response data, logistics terminal fees, and logistics terminal efficiency indicator data;

[0028] The identification module is used to obtain the approximation of the logistics multi-source customer service data of each segmented time node based on the logistics multi-source customer service data of two adjacent segmented time nodes, and to obtain the abnormal growth evaluation value of each data p' in the logistics multi-source customer service data of each segmented time node based on the approximation of the logistics multi-source customer service data of each segmented time node. ;

[0029] The sampling calculation module is used to obtain the sampling data p of each data p' in the logistics multi-source customer service data of each segmented time node, and then calculate the local outlier factor based on the neighborhood k obtained by cross-validation of the sampling data p, and output the local outlier factor evaluation value LOF of the sampling data p ;

[0030] The evaluation calculation module is used to evaluate the local outlier factor LOF based on the sampled data p. and the abnormal growth evaluation value of each data , calculate the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node ;

[0031] The output processing module is used to perform logistics benefit evaluation processing operations based on the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node.

[0032] Preferably, the evaluation calculation module includes an initial processing module, a selection processing module, a domain calculation module, a correlation calculation module and an evaluation value calculation module:

[0033] The initial processing module is used to select a neighborhood k of an initial size; and evaluate the accuracy of the local outlier factor evaluation value based on the neighborhood k through cross-validation based on the neighborhood k;

[0034] The selection processing module is used to select the optimized neighborhood k based on the accuracy rate;

[0035] The domain calculation module is used to calculate the domain local density of each sample data p based on the optimized neighborhood k ;

[0036] The association calculation module is used to calculate the local density of the domain Calculate the correlation sensitivity factor of each of the sampled data ;

[0037] The evaluation value calculation module is used to calculate the value of the neighborhood k based on the optimized neighborhood k and the associated sensitivity factor. Calculate and obtain the local outlier factor evaluation value.

[0038] Preferably, the local density of the field The calculation method is:

[0039] ;

[0040] Where, is the neighborhood of the i-th sampling data; is the distance function, which represents the distance between sample data i and sample data j;

[0041] The associated sensitive factor The calculation method is:

[0042] ;

[0043] The local outlier factor evaluation value is calculated as follows:

[0044] ;

[0045] Where, is the correlation sensitivity factor of the j-th sampling data; is the correlation sensitivity factor of the i-th sampling data.

[0046] Preferably, the output processing module includes an acquisition module, a Spearman rank calculation module and an evaluation output processing module:

[0047] The acquisition module is used to collect and acquire logistics after-sales variable information; the logistics after-sales variable information includes logistics satisfaction, return rate and logistics complaint rate, logistics transfer station transfer efficiency evaluation coefficient, and delivery delay time;

[0048] The Spearman rank calculation module is used to calculate the Spearman rank correlation coefficient based on the logistics after-sales variable information and the comprehensive abnormality assessment score ;

[0049] The evaluation output processing module is used to determine the Spearman rank correlation coefficient Is it greater than 0 or less than 0; if so, based on the Spearman rank correlation coefficient The corresponding logistics after-sales variable information and the comprehensive abnormality evaluation score are subjected to a logistics benefit evaluation processing operation; if not, a logistics benefit evaluation processing operation is performed based on the comprehensive abnormality evaluation score.

[0050] Preferably, the Spearman rank calculation module specifically includes: a traversal processing module, a first calculation module, and a second calculation module;

[0051] The traversal processing module is used to traverse each of the logistics after-sales variable information, and obtain a variable pair based on the combination of the logistics after-sales variable information and the comprehensive abnormality evaluation score corresponding to the logistics after-sales variable information;

[0052] The first calculation module is configured to obtain a first variable sequence set by sorting the values ​​of all the logistics after-sales variable information from high to low; and obtain a second variable sequence set by sorting the values ​​of the comprehensive abnormality assessment scores from high to low;

[0053] The second calculation module is used to calculate the ranking difference of the variable pair based on the first variable sequence set and the second variable sequence set; and calculate the Spearman rank correlation coefficient based on the ranking difference. ;

[0054] The Spearman rank correlation coefficient The calculation method is:

[0055] ;

[0056] Where, is the ranking difference; n is the total value of the logistics after-sales variable information and the comprehensive abnormality evaluation score.

[0057] Preferably, the acquisition module includes a transfer data acquisition module, a matrix construction module and an evaluation coefficient calculation module:

[0058] The transit data acquisition module is used to collect and obtain logistics transit data; the logistics transit data includes transit station data , the throughput of each transfer station in the transfer station data , regional total throughput and the transfer station coverage area ;

[0059] The matrix construction module is used to construct a logistics transfer matrix based on the logistics transfer data; the logistics transfer matrix includes a transfer station throughput matrix P, a regional total throughput matrix vector R, and a coverage area matrix A;

[0060] Each element in the transfer station throughput matrix P Representative The transfer station The unit time throughput of the transfer station; each element in the total throughput matrix vector R of the region Representative The total item throughput per unit time of the area; each element of the coverage area matrix A represents the coverage area of ​​the jth transfer station in the i-th area;

[0061] The evaluation coefficient calculation module is used to calculate the transfer efficiency evaluation coefficient of the logistics transfer station based on the transfer station throughput matrix P, the regional total throughput matrix vector R and the coverage area matrix A. .

[0062] Preferably, the transfer efficiency evaluation coefficient of the logistics transfer station is The calculation method is:

[0063] ;

[0064] Where, Indicates in The sum of the throughput of items per unit time of all transfer stations in a transfer station; For the Total item throughput in each area; For the The transfer station is The coverage area of ​​a region.

[0065] Preferably, it also includes an optimization processing module; the optimization processing module is used to iteratively optimize the logistics multi-source customer service data based on the transfer efficiency evaluation coefficient of the logistics transfer station to obtain updated logistics multi-source customer service data; and re-perform logistics benefit evaluation processing operations based on the updated logistics multi-source customer service data.

[0066] Preferably, the optimization processing module specifically includes a sub-factor module, a Pearson correlation coefficient module, a target sub-factor combination processing module and an update processing module;

[0067] The sub-factor module is used to obtain multiple sub-factor combinations by combining sub-factors in pairs based on the transfer efficiency evaluation coefficient of the logistics transfer station and the logistics multi-source customer service data;

[0068] The Pearson correlation coefficient module is used to calculate the Pearson correlation coefficient based on the value of the logistics transfer station transfer efficiency evaluation coefficient sub-factor in the sub-factor combination and the value of the logistics multi-source customer service data sub-factor. ;

[0069] The Pearson correlation coefficient The calculation method is: ;

[0070] Where, is the value of the sub-factor of the transfer efficiency evaluation coefficient of the logistics transfer station; The values ​​of sub-factors in the logistics multi-source customer service data; is the average value of the sub-factors of the transfer efficiency evaluation coefficient of the logistics transfer station and is the average of the values ​​of the sub-factors in the logistics multi-source customer service data;

[0071] The target sub-factor combination processing module is used to screen and obtain the Pearson correlation coefficient The sub-factor combination corresponding to a value greater than or equal to the preset Pearson correlation coefficient threshold is the target sub-factor combination;

[0072] The update processing module is used to add the sub-factor of the transfer efficiency evaluation coefficient of the logistics transfer station in the target sub-factor combination to the logistics multi-source customer service data to obtain updated logistics multi-source customer service data.

[0073] Example 2

[0074] like Figure 1 As shown, the second embodiment of the present invention provides a method for evaluating and analyzing the regional business environment based on big data analysis, including the following steps:

[0075] Step S10: Collecting logistics multi-source customer service data at subdivided time nodes within a preset time period;

[0076] The logistics multi-source customer service data includes order data, logistics terminal coverage data, logistics terminal throughput data, logistics terminal (i.e., logistics company) transportation route data, customer feedback order response data (the customer here refers to the production enterprise in the current area, and its feedback on logistics orders in the area is the order response data), logistics terminal (i.e., logistics company) expenses and logistics terminal efficiency index data; logistics terminals specifically refer to logistics companies in the area;

[0077] It should be noted that in the above-mentioned embodiment of the present application, within a preset time period, the system collects logistics multi-source customer service data at segmented time nodes through multiple channels (such as logistics management systems, customer relationship management systems, social media, etc.), and it is also necessary to pre-process the logistics multi-source customer service data to make the data high-quality and consistent, thereby ensuring the accuracy of subsequent analysis.

[0078] Step S20: Obtain the approximation of the logistics multi-source customer service data at each segmented time node based on the logistics multi-source customer service data at two adjacent segmented time nodes, and obtain the abnormal growth evaluation value of each data p' in the logistics multi-source customer service data at each segmented time node based on the approximation of the logistics multi-source customer service data at each segmented time node. ;

[0079] It should be noted that the above-mentioned embodiment of the present application can use distance metrics such as Euclidean distance or cosine similarity to evaluate the similarity between data points (i.e., the aforementioned approximation), and combine the approximation to identify the abnormal growth evaluation value of each node (i.e., when the approximation is small, it proves that the growth of logistics multi-source customer service data at two adjacent segmented time nodes is abnormal, thereby analyzing and identifying to obtain the abnormal growth evaluation value), thereby identifying data points with significant fluctuations in the time series based on the abnormal growth evaluation value; specifically, the approximation is calculated as follows:

[0080] Where, and It is the logistics multi-source customer service data of two adjacent segmented time nodes. is a distance function, representing the logistics multi-source customer service data of segmented time nodes Logistics multi-source customer service data with segmented time nodes The distance between them.

[0081] Step S30: Obtain the sampled data p of each data p' in the logistics multi-source customer service data of each segmented time node, and then calculate the local outlier factor based on the neighborhood k obtained by cross-validation of the sampled data p, and output the local outlier factor evaluation value LOF of the sampled data p ;

[0082] It should be noted that the above-mentioned sampled data p and each data p' in the logistics multi-source customer service data at each subdivided time node are essentially the same data; the difference is that the sampled data p is sub-data obtained by sampling from each data p' in the logistics multi-source customer service data at each subdivided time node;

[0083] The above-mentioned local outlier factor (LOF) is an unsupervised learning algorithm for detecting local outliers in a data set. It evaluates the degree of abnormality of each point by comparing the density of a data point with its neighboring points. When calculating the local outlier factor evaluation value LOF of the sampled data, it is calculated by comparing the density of a point with the density of other points in its neighborhood. If the LOF value of a point is significantly greater than 1, it is usually regarded as an abnormal point. Among them, the neighborhood means that when calculating the local outlier factor evaluation value LOF of the sampled data, a neighborhood k is defined for each sampled data by setting a fixed radius or selecting a fixed number of nearest neighbors, and the local density of each point is calculated, that is, the number of data points in its neighborhood. Generally, points with high density are considered normal, and points with low density may be abnormal.

[0084] Step S40: local outlier factor evaluation value LOF based on the sampled data p and the abnormal growth evaluation value of each data , calculate the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node ;

[0085] The comprehensive anomaly assessment score The calculation method is: ;

[0086] Where, 、 is the weight coefficient, reflecting the importance of each indicator;

[0087] Step S50: performing a logistics benefit evaluation processing operation based on the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node;

[0088] Specifically, the above-mentioned embodiment of the present application performs logistics benefit evaluation processing operations based on the comprehensive abnormality evaluation score, which specifically means: comparing the comprehensive evaluation score with the preset threshold, identifying customers or logistics links with abnormal performance, and then generating reports based on the abnormal customers or logistics links to provide decision support and help optimize logistics processes and customer service.

[0089] It should be noted that the above embodiment of the present application first collects multi-source customer service data, including orders, logistics terminal information, customer feedback (or feedback from enterprises in the region on logistics orders), etc., to provide comprehensive basic data support; and ensures high data quality and consistency through data preprocessing, laying a solid foundation for subsequent analysis; the above high-quality basic data can improve the accuracy of subsequent analysis and provide a reliable basis for anomaly detection; further, by calculating the approximation between adjacent segmented time nodes, identifying abnormal growth points in the data, it is possible to quickly discover potential problems and risks, and provide technical data support and clues for further analysis and decision-making; by calculating the local outlier factor (LOF) of the sampled data, the data is evaluated The unsupervised learning method effectively captures the degree of abnormality of the stronghold, and improves the flexibility and accuracy of anomaly detection. At the same time, in the above scheme, based on LOF and abnormal growth evaluation value, the comprehensive abnormality evaluation score of each data is calculated. The comprehensive evaluation score provides a quantitative indicator, which can more intuitively compare the abnormality degree of different data points and enhance the operability of the analysis. The comprehensive abnormality evaluation score is compared with the preset threshold to identify customers or logistics links with abnormal performance, and generate reports to provide decision support and help optimize logistics processes and customer service. By identifying anomalies and providing targeted suggestions, it can significantly improve logistics efficiency and enterprise terminal satisfaction, and help optimize the overall operational efficiency of the region.

[0090] Specifically, if Figure 2 As shown, in step S30, the local outlier factor is calculated based on the neighborhood k obtained by cross-validation of the sample data p, and the local outlier factor evaluation value LOF of the sample data p is output. , including the following steps:

[0091] Step S31: selecting a neighborhood k of an initial size; and evaluating the accuracy of the local outlier factor evaluation value based on the neighborhood k through cross-validation based on the neighborhood k;

[0092] It should be noted that the above-mentioned embodiment of the present application is to collect sample sampling data with a preset local outlier factor evaluation value, and then divide the sample sampling data into a training set and a validation set with a fixed ratio (when the sample sampling data division is completed, the preset local outlier factor evaluation value is also divided into a training local outlier factor evaluation value and a validation local outlier factor evaluation value); further, based on the training set, a k-fold cross-validation method is used to evaluate the influence of the neighborhood k on the preset local outlier factor evaluation value (that is, calculate the training local outlier factor evaluation value on each fold of the k folds), and then use the validation local outlier factor evaluation value of the validation set to calculate the accuracy of the preset local outlier factor evaluation value of the training set;

[0093] Step S32: selecting an optimized neighborhood k based on the accuracy;

[0094] It should be noted that the above embodiment of the present application selects the corresponding neighborhood k with the highest accuracy as the optimized neighborhood k.

[0095] Step S33: Calculate the local density of each sample data p based on the optimized neighborhood k ;

[0096] The local density of the field The calculation method is:

[0097] ;

[0098] Where, is the neighborhood of the i-th sampling data; is the distance function, which represents the distance between sample data i and sample data j;

[0099] Step S34: Based on the local density of the field Calculate the correlation sensitivity factor of each of the sampled data ;

[0100] The associated sensitive factor The calculation method is:

[0101] ;

[0102] It should be noted that the above embodiment of the present application introduces the associated sensitive factor To express the relative density of sampled data i in its neighborhood, quantify the influence of sampled data i on its neighborhood, and thus reflect the importance of sampled data i in the neighborhood.

[0103] Step S35: Based on the optimized neighborhood k and the associated sensitivity factor Calculate and obtain the local outlier factor evaluation value;

[0104] The local outlier factor evaluation value is calculated as follows:

[0105] ;

[0106] Where, is the correlation sensitivity factor of the j-th sampling data; is the correlation sensitivity factor of the i-th sampling data;

[0107] It should be noted that the above-mentioned embodiment of the present application optimizes the neighborhood size k by introducing cross-validation and enhances the accuracy of the local outlier factor by calculating the associated sensitive factor, which can effectively improve the stability and accuracy of anomaly detection; it not only improves the performance of the model, but also lays the foundation for subsequent comprehensive anomaly evaluation and logistics benefit analysis.

[0108] During the specific implementation of the above embodiment of the present application, technical personnel discovered that the above comprehensive abnormality assessment score is obtained by analyzing multi-source customer service data. Therefore, when the technical solution adopted in the above embodiment of the present application is directly used to perform logistics benefit evaluation processing operations, only pre-sales data is evaluated and processed. However, this is basically pre-sales analysis data and does not participate in the use of post-sales data.

[0109] After a period of use, a variable that affects the comprehensive abnormality evaluation score will occur (this variable is the logistics after-sales variable data information), and it is impossible to perform targeted processing operations on this variable; therefore, when performing logistics benefit evaluation processing operations based on the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node, it is necessary not only to consider the comprehensive abnormality evaluation score, but also to consider the impact of the logistics after-sales variable data information on the comprehensive abnormality evaluation score, so as to perform targeted logistics benefit evaluation processing operations based on the comprehensive abnormality evaluation score.

[0110] Specifically, if Figure 3 As shown, in step S50, a logistics benefit evaluation processing operation is performed according to the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node, including the following operation steps:

[0111] Step S51: Collect and obtain logistics after-sales variable information; the logistics after-sales variable information includes logistics satisfaction, return rate and logistics complaint rate, logistics transfer station transfer efficiency evaluation coefficient, and delivery delay time;

[0112] It should be noted that, in the above-mentioned embodiment of the present application, enterprise terminal satisfaction refers to the satisfaction rating (e.g., 1 to 5 points) of customers (i.e., production enterprises in the region) on the service collected through questionnaires or online rating systems; return rate refers to the ratio of return logistics orders to total logistics orders calculated within a certain time period; logistics complaint rate refers to the ratio of the number of customer complaints received to the total number of orders processed within a certain period of time; the transshipment evaluation coefficient of the logistics transfer station refers to recording the number of transshipments at each transfer station and evaluating its impact on the overall logistics efficiency; delivery delay time refers to calculating the difference between the actual delivery time and the promised delivery time, and recording the number of delay days;

[0113] It should also be noted that after obtaining the logistics and after-sales variable information, the logistics and after-sales variable information needs to be preprocessed to make the data high-quality and consistent, thereby ensuring the accuracy of subsequent analysis.

[0114] Step S52: Calculate the Spearman rank correlation coefficient based on the logistics after-sales variable information and the comprehensive abnormality assessment score ;

[0115] It should be noted that the above Spearman rank correlation coefficient It refers to the relevant influence coefficient of each element in the logistics after-sales variable information on the comprehensive abnormality evaluation score, such as the relevant influence coefficient of logistics satisfaction on the comprehensive abnormality evaluation score.

[0116] Step S53: Determine the Spearman rank correlation coefficient Is it greater than 0 or less than 0; if so, based on the Spearman rank correlation coefficient The corresponding logistics after-sales variable information and the comprehensive abnormality evaluation score are subjected to a logistics benefit evaluation processing operation; if not, a logistics benefit evaluation processing operation is performed based on the comprehensive abnormality evaluation score.

[0117] It should be noted that, in the above embodiment of the present application, the Spearman rank correlation coefficient If the score is less than 0, it indicates that the comprehensive anomaly evaluation score is negatively correlated with a logistics after-sales variable (such as delivery delay time). Delayed delivery may be the main cause of the anomaly. In this case, when performing logistics benefit evaluation, the comprehensive anomaly evaluation score and the logistics after-sales variable information are comprehensively considered for logistics benefit evaluation. For example, if the comprehensive anomaly evaluation score is negatively correlated with delivery delay time, it is recommended to optimize the logistics process to reduce delivery delays.

[0118] The Spearman rank correlation coefficient When the value is greater than 0, it indicates that the comprehensive anomaly evaluation score is positively correlated with a logistics after-sales variable (such as enterprise terminal satisfaction). Low satisfaction may lead to anomalies. Therefore, when performing logistics benefit evaluation, the comprehensive anomaly evaluation score and the logistics after-sales variable information are comprehensively considered for logistics benefit evaluation. For example, if the comprehensive anomaly evaluation score is positively correlated with enterprise terminal satisfaction, it is recommended to improve customer service quality and feedback mechanisms.

[0119] Furthermore, in the Spearman rank correlation coefficient When it is equal to 0, it proves that the logistics post-sales variable information has no effect on the comprehensive abnormality evaluation score, and proves that the comprehensive abnormality evaluation score is obtained only because of the pre-sales data in the aforementioned process. Therefore, when performing the logistics benefit evaluation processing operation, the logistics benefit evaluation processing operation is only performed on the comprehensive abnormality evaluation score (that is, there is no need to refer to this logistics post-sales variable information).

[0120] Specifically, if Figure 4 As shown, in step S52, the Spearman rank correlation coefficient is calculated based on the logistics after-sales variable information and the comprehensive abnormality evaluation score. , including the following steps:

[0121] Step S521: traverse each of the logistics after-sales variable information, and obtain a variable pair based on the combination of the logistics after-sales variable information and the comprehensive abnormality evaluation score corresponding to the logistics after-sales variable information;

[0122] It should be noted that the above embodiment of the present application is to obtain variable pairs by combining the elements in the plurality of logistics after-sales variable information and the comprehensive abnormality assessment scores corresponding to the respective logistics after-sales variable information;

[0123] For example, there are currently three logistics after-sales variable information, which are represented as: after-sales 1 (logistics satisfaction 1, return rate and logistics complaint rate 1, logistics transfer station transfer efficiency evaluation coefficient 1, delivery delay time 1), after-sales 2 (logistics satisfaction 2, return rate and logistics complaint rate 2, logistics transfer station transfer efficiency evaluation coefficient 2, delivery delay time 2), and after-sales 3 (logistics satisfaction 3, return rate and logistics complaint rate 3, logistics transfer station transfer efficiency evaluation coefficient 3, delivery delay time 3). The comprehensive abnormality evaluation scores corresponding to these three logistics after-sales variable information are comprehensive score 1, comprehensive score 2, and comprehensive score 3, respectively.

[0124] Furthermore, each element in each logistics after-sales variable information is combined with the comprehensive abnormality evaluation score to obtain a variable pair, such as Figure 5 As shown, it is expressed as: logistics satisfaction 1-comprehensive score 1, return rate and logistics complaint rate 1-comprehensive score 1, logistics transfer station transfer efficiency evaluation coefficient 1-comprehensive score 1, delivery delay time 1-comprehensive score 1, logistics satisfaction 2-comprehensive score 2, return rate and logistics complaint rate 2-comprehensive score 2, logistics transfer station transfer efficiency evaluation coefficient 2-comprehensive score 2, delivery delay time 2-comprehensive score 2, logistics satisfaction 3-comprehensive score 3...;

[0125] Step S522: sorting the values ​​of all the logistics after-sales variable information (the values ​​of the logistics after-sales variable information are score values ​​or other forms of descriptive quantitative values) from high to low to obtain a first variable sequence set; and sorting the values ​​of the comprehensive abnormality assessment scores from high to low to obtain a second variable sequence set;

[0126] It should be noted that the above embodiment of the present application is to sort the values ​​of each element in each logistics after-sales variable information to obtain a first variable sequence set;

[0127] Continuing with the above example: the values ​​of the elements in after-sales service 1 are logistics satisfaction 1=2, return rate and logistics complaint rate 1=3, logistics transfer station transfer efficiency evaluation coefficient 1=5, and delivery delay time 1=1; the values ​​of the elements in after-sales service 2 are logistics satisfaction 2=3, return rate and logistics complaint rate 2=4, logistics transfer station transfer efficiency evaluation coefficient 2=6, and delivery delay time 2=2; the values ​​of the elements in after-sales service 3 are logistics satisfaction 3=7, return rate and logistics complaint rate 3=6, logistics transfer station transfer efficiency evaluation coefficient 3=2, and delivery delay time 3=7; the values ​​of the three comprehensive abnormality evaluation scores are comprehensive score 1=8, comprehensive score 2=6, and comprehensive score 3=4;

[0128] According to the technical solution adopted in the embodiment of the present application, the values ​​of the elements are sorted accordingly to obtain a first variable sequence set and a second variable sequence set;

[0129] Sorting by the value of logistics satisfaction, we get the first variable sequence set, expressed as: first variable sequence set 1 = {logistics satisfaction 3 is 7, logistics satisfaction 2 is 3, logistics satisfaction 1 is 2}, and we get the first first variable sequence set 1 corresponding to logistics satisfaction;

[0130] Sorting by the values ​​of return rate and logistics complaint rate, we get the first variable sequence set, expressed as: first variable sequence set 2 = {return rate and logistics complaint rate 3 is 6, return rate and logistics complaint rate 2 is 4, return rate and logistics complaint rate 1 is 3}, and we get the second first variable sequence set 2 corresponding to the return rate and logistics complaint rate;

[0131] The first variable sequence set is obtained by sorting the values ​​of the transfer efficiency evaluation coefficients of the logistics transfer stations, which is expressed as: first variable sequence set 3 = {transfer efficiency evaluation coefficient 2 of the logistics transfer station is 6, transfer efficiency evaluation coefficient 1 of the logistics transfer station is 5, and transfer efficiency evaluation coefficient 3 of the logistics transfer station is 2}, and the third first variable sequence set 3 corresponding to the transfer efficiency evaluation coefficient of the logistics transfer station is obtained;

[0132] Sorting by the delivery delay time value, the first variable sequence set is obtained, which is expressed as: first variable sequence set 4 = {delivery delay time 3 is 7, delivery delay time 2 is 2, delivery delay time 1 is 1}, and the fourth first variable sequence set 4 corresponding to the delivery delay time is obtained;

[0133] Sorting by the value of the comprehensive abnormality assessment score, the second variable sequence set is obtained, which is expressed as: the second variable sequence set = {comprehensive score 1 is 8, comprehensive score 2 is 6, comprehensive score 3 is 4};

[0134] Step S523: Calculating the ranking difference of the variable pairs based on the first variable sequence set and the second variable sequence set;

[0135] It should be noted that, in the above embodiment of the present application, the ranking difference of the variable pair is obtained by subtracting the ranking of the comprehensive abnormality assessment score corresponding to the value of the logistics after-sales variable information in the first variable sequence set from the ranking of the value of the logistics after-sales variable information in the second variable sequence set;

[0136] Continuing with the above example, according to the technical solution adopted in the above-mentioned embodiment of the present application, the ranking differences corresponding to each first variable sequence set and the second variable sequence set are calculated, that is, the ranking differences of the variable pairs are calculated using the first first variable sequence set 1 and the second variable sequence set corresponding to logistics satisfaction, where the variable pairs include logistics satisfaction 1-comprehensive score 1, logistics satisfaction 2-comprehensive score 2, and logistics satisfaction 3-comprehensive score 3; in the first variable sequence set 1, logistics satisfaction 1 ranks third, logistics satisfaction 2 ranks second, and logistics satisfaction 3 ranks third; in the second variable sequence set, comprehensive score 1 ranks first, comprehensive score 2 ranks second, and comprehensive score 3 ranks third. The calculated ranking differences are: the ranking difference of the first variable for logistics satisfaction 1-comprehensive score 1 is |3-1|=2; the ranking difference of the second variable for logistics satisfaction 2-comprehensive score 2 is |2-2|=0; the ranking difference of the third variable for logistics satisfaction 3-comprehensive score 3 is |3-1|=2...;

[0137] Step S524: Calculate the Spearman rank correlation coefficient based on the ranking difference ;

[0138] The Spearman rank correlation coefficient The calculation method is:

[0139] ;

[0140] Where, is the ranking difference; n is the total value of the logistics after-sales variable information and the comprehensive abnormality evaluation score.

[0141] It should be noted that the above-mentioned embodiment of the present application first combines each logistics after-sales variable information (such as logistics satisfaction, return rate, etc.) with its corresponding comprehensive abnormality evaluation score to form a "variable pair", thereby laying the foundation for subsequent sorting and correlation analysis; further, the variables are sorted according to their values ​​and comprehensive abnormality evaluation scores, and then the relative deviation of each variable is quantified by comparing the ranking difference of each logistics after-sales variable in the two sequences; further, the Spearman rank correlation coefficient is used to quantify the correlation between the variable pairs, and the Spearman rank correlation coefficient is used to deeply understand the relationship between the comprehensive abnormality evaluation score and other key variables, so as to more effectively identify the causes of abnormal performance and formulate corresponding improvement measures.

[0142] Specifically, if Figure 6 As shown, in step S51, obtaining the transfer efficiency evaluation coefficient of the logistics transfer station includes the following steps:

[0143] Step S511: Collect and obtain logistics transfer data; the logistics transfer data includes transfer station data , the throughput of each transfer station in the transfer station data , regional total throughput and the transfer station coverage area ;

[0144] It should be noted that, in the above embodiment of the present application, the transfer station data Refers to the In the region transfer stations; transfer station throughput Refers to the The transfer station The throughput of items per unit time of the transfer station; the total throughput of the area Refers to the Total item throughput per unit time in each area; Transit station coverage area Refers to the The transfer station is The coverage area of ​​each region;

[0145] Step S512: constructing a logistics transit matrix based on the logistics transit data;

[0146] The logistics transfer matrix includes the transfer station throughput matrix P, the regional total throughput matrix vector R and the coverage area matrix A;

[0147] Each element in the transfer station throughput matrix P Representative The transfer station The unit time throughput of the transfer station; each element in the total throughput matrix vector R of the region Representative The total item throughput per unit time of the area; each element of the coverage area matrix A represents the coverage area of ​​the jth transfer station in the i-th area;

[0148] Step S513: Calculate the transfer efficiency evaluation coefficient of the logistics transfer station based on the transfer station throughput matrix P, the regional total throughput matrix vector R and the coverage area matrix A ;

[0149] Transfer efficiency evaluation coefficient of the logistics transfer station The calculation method is:

[0150] ;

[0151] Where, Indicates in The sum of the throughput of items per unit time of all transfer stations in a transfer station; For the Total item throughput in each area; For the The transfer station is The coverage area of ​​a region.

[0152] It should be noted that the above-mentioned embodiment of the present application first collects data in steps S411 and S412 and converts it into a matrix form, which can ensure the comprehensiveness and operability of the data. The constructed matrix clearly shows the performance, regional requirements and coverage capabilities of each transfer station, so that subsequent analysis can be more detailed and accurate; further, through the calculation of step S413, the transfer efficiency of each transfer station is comprehensively evaluated.

[0153] Furthermore, during the specific implementation of the above-mentioned embodiment of the present application, the technicians also found that the acquisition of logistics multi-source customer service data will be more or less subjective due to the source of the data, such as logistics management system, customer relationship management system, social media, etc., and the transfer efficiency evaluation coefficient of the above-mentioned logistics transfer station In essence, it can be used to feedback and adjust the logistics multi-source customer service data, thereby improving logistics efficiency. For example, when the throughput of the transfer station is high, the transportation distance is shorter, and the customer feedback order response time is faster, thereby improving logistics efficiency and enterprise terminal satisfaction, and helping to optimize the overall operational efficiency of the region; therefore, when specifically obtaining logistics multi-source customer service data, the transfer efficiency evaluation coefficient of the above-mentioned logistics transfer station can be considered. The impact of the local outlier factor evaluation value LOF More accurate, thus making subsequent logistics benefit evaluation and processing operations better.

[0154] Specifically, after step S50, the method further includes iteratively optimizing the logistics multi-source customer service data based on the transfer efficiency evaluation coefficient of the logistics transfer station to obtain updated logistics multi-source customer service data; and re-performing a logistics benefit evaluation processing operation based on the updated logistics multi-source customer service data;

[0155] Specifically, if Figure 7 As shown, the iterative optimization of the logistics multi-source customer service data based on the transfer efficiency evaluation coefficient of the logistics transfer station to obtain updated logistics multi-source customer service data includes the following steps:

[0156] Step S501: obtaining a plurality of sub-factor combinations by combining sub-factors in pairs based on the transfer efficiency evaluation coefficient of the logistics transfer station and the logistics multi-source customer service data;

[0157] It should be noted that the above embodiment of the present application is to combine the sub-factors of the transfer efficiency evaluation coefficient of the logistics transfer station (i.e., the transfer station data , the throughput of each transfer station in the transfer station data , regional total throughput and the transfer station coverage area ) are combined in pairs with the sub-factors of logistics multi-source customer service data (i.e., order data, logistics terminal coverage data, logistics terminal throughput data, logistics terminal transportation path data, customer feedback order response data, logistics terminal (i.e., logistics company) fees and logistics terminal efficiency index data) to obtain multiple sub-factor combinations, which can provide a data basis for subsequent analysis; this application will not go into details.

[0158] Step S502: Calculate the Pearson correlation coefficient based on the value of the logistics transfer station transfer efficiency evaluation coefficient sub-factor in the sub-factor combination and the value of the logistics multi-source customer service data sub-factor ;

[0159] The Pearson correlation coefficient The calculation method is: ;

[0160] Where, is the value of the sub-factor of the transfer efficiency evaluation coefficient of the logistics transfer station (for example, the value of the sub-factor is: a score value or other form of descriptive quantitative value); The values ​​of sub-factors in the logistics multi-source customer service data; is the average value of the sub-factors of the transfer efficiency evaluation coefficient of the logistics transfer station and is the average of the values ​​of the sub-factors in the logistics multi-source customer service data;

[0161] Step S503: Screening and obtaining the Pearson correlation coefficient The sub-factor combination corresponding to a value greater than or equal to the preset Pearson correlation coefficient threshold is the target sub-factor combination;

[0162] Step S504: adding the sub-factor of the transfer efficiency evaluation coefficient of the logistics transfer station in the target sub-factor combination to the logistics multi-source customer service data to obtain updated logistics multi-source customer service data;

[0163] Then, based on the updated logistics multi-source customer service data, the above step S10 is returned to be executed again to obtain a comprehensive abnormality evaluation score, and a logistics benefit evaluation processing operation is performed according to the comprehensive abnormality evaluation score.

[0164] It should be noted that the above embodiment of the present application adds the sub-factor of the transfer efficiency evaluation coefficient of the logistics transfer station in the target sub-factor combination to the logistics multi-source customer service data in the target sub-factor combination, while retaining the sub-factor of the logistics multi-source customer service data in the target sub-factor combination, thereby obtaining the updated logistics multi-source customer service data. For example, a large logistics company A is responsible for logistics distribution in multiple cities across the country, has multiple transfer stations and a large number of customer orders. To improve efficiency, Company A decides to implement the above solution.

[0165] First, the sub-factors of the transfer efficiency evaluation coefficient of the logistics transfer station were combined with the sub-factors of the logistics multi-source customer service data in pairs to form multiple sub-factor combinations as the basis for subsequent analysis; further, by calculating the Pearson correlation coefficient, the company found that there was a strong correlation between certain factor combinations (such as "transfer station throughput" and "logistics terminal transportation path"), for example, it was found that the higher the transfer station throughput of certain transfer stations, the shorter the transportation path, and the faster the response time of customer orders, with a correlation coefficient of 0.85 (indicating a strong positive correlation); Company A set the threshold at 0.7 and only screened sub-factor combinations with a correlation coefficient greater than or equal to 0.7, and finally screened out the "transfer station throughput" and "transport path" combination as the target sub-factor combination; through this analysis, Company A added the throughput data of efficient transfer stations to the logistics multi-source customer service data, and used the updated data to help optimize customers' transportation paths and distribution arrangements, reducing unnecessary transfers and improving overall efficiency.

[0166] In summary, the present invention proposes a system and method for evaluating and analyzing the regional business environment based on big data analysis. The system first collects multi-source customer service data. Furthermore, it identifies abnormal growth points in the data by calculating the approximation between adjacent segmented time nodes. The local outlier factor (LOF) is calculated for the sampled data to evaluate the degree of abnormality of the data points. During the calculation process, the neighborhood size k is optimized by introducing cross-validation, and the accuracy of the local outlier factor is enhanced by calculating the associated sensitive factor, which can effectively improve the stability and accuracy of anomaly detection. This not only improves the performance of the model but also lays the foundation for subsequent comprehensive anomaly evaluation and logistics benefit analysis. At the same time, in the above scheme, based on the LOF and abnormal growth evaluation value, a comprehensive anomaly evaluation score is calculated for each data point, which more intuitively compares the degree of anomaly of different data points and enhances the operability of the analysis. The comprehensive anomaly evaluation score is compared with a preset threshold to identify customers or logistics links with abnormal performance, and a report is generated to provide decision support and help optimize logistics processes and customer service. By identifying anomalies and providing targeted suggestions, logistics efficiency and enterprise terminal satisfaction can be significantly improved, which helps to optimize the overall operational efficiency of the region and improve the level of regional business operations.

[0167] When performing logistics benefit evaluation and analysis, the Spearman rank correlation coefficient is obtained by introducing logistics after-sales variable information and comprehensive abnormal evaluation scores. To carry out targeted logistics benefit evaluation processing operations; in addition, the transfer efficiency evaluation coefficient of the logistics transfer station is introduced to iteratively optimize the logistics multi-source customer service data, thereby optimizing the regional logistics operation efficiency.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. A person skilled in the art may modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A regional business environment evaluation and analysis system based on big data analysis, characterized in that: include: Information collection platform, identification module, sampling calculation module, evaluation calculation module and output processing module; The information collection platform is used to collect logistics multi-source customer service data at subdivided time nodes within a preset time period; the logistics multi-source customer service data includes order data, logistics terminal coverage data, logistics terminal throughput data, logistics terminal transportation route data, customer feedback order response data, logistics terminal fees, and logistics terminal efficiency indicator data; The identification module is used to obtain the approximation of the logistics multi-source customer service data of each segmented time node based on the logistics multi-source customer service data of two adjacent segmented time nodes, and to obtain the abnormal growth evaluation value of each data p' in the logistics multi-source customer service data of each segmented time node based on the approximation of the logistics multi-source customer service data of each segmented time node. ; The sampling calculation module is used to obtain the sampling data p of each data p' in the logistics multi-source customer service data of each segmented time node, and then calculate the local outlier factor based on the neighborhood k obtained by cross-validation of the sampling data p, and output the local outlier factor evaluation value LOF of the sampling data p ; The evaluation calculation module is used to evaluate the local outlier factor LOF based on the sampled data p. and the abnormal growth evaluation value of each data , calculate the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node ; The output processing module is used to perform logistics benefit evaluation processing operations based on the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node.

2. The regional business environment evaluation and analysis system based on big data analysis according to claim 1 is characterized in that: The evaluation calculation module includes an initial processing module, a selection processing module, a domain calculation module, a correlation calculation module and an evaluation value calculation module: The initial processing module is used to select a neighborhood k of an initial size; and evaluate the accuracy of the local outlier factor evaluation value based on the neighborhood k through cross-validation based on the neighborhood k; The selection processing module is used to select the optimized neighborhood k based on the accuracy rate; The domain calculation module is used to calculate the domain local density of each sample data p based on the optimized neighborhood k ; The association calculation module is used to calculate the local density of the domain Calculate the correlation sensitivity factor of each of the sampled data ; The evaluation value calculation module is used to calculate the value of the neighborhood k based on the optimized neighborhood k and the associated sensitivity factor. Calculate and obtain the local outlier factor evaluation value.

3. The regional business environment evaluation and analysis system based on big data analysis according to claim 2 is characterized in that: The local density of the field The calculation method is: ; Where, is the neighborhood of the i-th sampling data; is the distance function, which represents the distance between sample data i and sample data j; The associated sensitive factor The calculation method is: ; The local outlier factor evaluation value is calculated as follows: ; Where, is the correlation sensitivity factor of the j-th sampling data; is the correlation sensitivity factor of the i-th sampling data.

4. The regional business environment evaluation and analysis system based on big data analysis according to claim 3 is characterized in that: The output processing module includes an acquisition module, a Spearman rank calculation module and an evaluation output processing module: The acquisition module is used to collect and acquire logistics after-sales variable information; the logistics after-sales variable information includes logistics satisfaction, return rate and logistics complaint rate, logistics transfer station transfer efficiency evaluation coefficient, and delivery delay time; The Spearman rank calculation module is used to calculate the Spearman rank correlation coefficient based on the logistics after-sales variable information and the comprehensive abnormality assessment score ; The evaluation output processing module is used to determine the Spearman rank correlation coefficient Is it greater than 0 or less than 0; if so, based on the Spearman rank correlation coefficient The corresponding logistics after-sales variable information and the comprehensive abnormality evaluation score are subjected to a logistics benefit evaluation processing operation; if not, a logistics benefit evaluation processing operation is performed based on the comprehensive abnormality evaluation score.

5. The regional business environment evaluation and analysis system based on big data analysis according to claim 4 is characterized in that: The Spearman rank calculation module specifically includes: a traversal processing module and a first calculation module and a second calculation module; The traversal processing module is used to traverse each of the logistics after-sales variable information, and obtain a variable pair based on the combination of the logistics after-sales variable information and the comprehensive abnormality evaluation score corresponding to the logistics after-sales variable information; The first calculation module is configured to obtain a first variable sequence set by sorting the values ​​of all the logistics after-sales variable information from high to low; and obtain a second variable sequence set by sorting the values ​​of the comprehensive abnormality assessment scores from high to low; The second calculation module is used to calculate the ranking difference of the variable pair based on the first variable sequence set and the second variable sequence set; and calculate the Spearman rank correlation coefficient based on the ranking difference. ; The Spearman rank correlation coefficient The calculation method is: ; Where, is the ranking difference; n is the total value of the logistics after-sales variable information and the comprehensive abnormality evaluation score.

6. The regional business environment evaluation and analysis system based on big data analysis according to claim 5 is characterized in that: The acquisition module includes a transfer data acquisition module, a matrix construction module and an evaluation coefficient calculation module: The transit data acquisition module is used to collect and obtain logistics transit data; the logistics transit data includes transit station data , the throughput of each transfer station in the transfer station data , regional total throughput and the transfer station coverage area ; The matrix construction module is used to construct a logistics transfer matrix based on the logistics transfer data; the logistics transfer matrix includes a transfer station throughput matrix P, a regional total throughput matrix vector R, and a coverage area matrix A; Each element in the transfer station throughput matrix P Representative The transfer station The unit time throughput of the transfer station; each element in the total throughput matrix vector R of the region Representative The total item throughput per unit time of the area; each element of the coverage area matrix A represents the coverage area of ​​the jth transfer station in the i-th area; The evaluation coefficient calculation module is used to calculate the transfer efficiency evaluation coefficient of the logistics transfer station based on the transfer station throughput matrix P, the regional total throughput matrix vector R and the coverage area matrix A. .

7. The regional business environment evaluation and analysis system based on big data analysis according to claim 6 is characterized in that: Transfer efficiency evaluation coefficient of the logistics transfer station The calculation method is: ; Where, Indicates in The sum of the throughput of items per unit time of all transfer stations in a transfer station; For the Total item throughput in each area; For the The transfer station is The coverage area of ​​a region.

8. The regional business environment evaluation and analysis system based on big data analysis according to claim 7 is characterized in that: It also includes an optimization processing module; the optimization processing module is used to iteratively optimize the logistics multi-source customer service data based on the transfer efficiency evaluation coefficient of the logistics transfer station to obtain updated logistics multi-source customer service data; and re-perform logistics benefit evaluation processing operations based on the updated logistics multi-source customer service data.

9. The regional business environment evaluation and analysis system based on big data analysis according to claim 8 is characterized in that: The optimization processing module specifically includes a sub-factor module, a Pearson correlation coefficient module, a target sub-factor combination processing module and an update processing module; The sub-factor module is used to obtain multiple sub-factor combinations by combining sub-factors in pairs based on the transfer efficiency evaluation coefficient of the logistics transfer station and the logistics multi-source customer service data; The Pearson correlation coefficient module is used to calculate the Pearson correlation coefficient based on the value of the logistics transfer station transfer efficiency evaluation coefficient sub-factor in the sub-factor combination and the value of the logistics multi-source customer service data sub-factor. ; The Pearson correlation coefficient The calculation method is: ; Where, is the value of the sub-factor of the transfer efficiency evaluation coefficient of the logistics transfer station; The values ​​of sub-factors in the logistics multi-source customer service data; is the average value of the sub-factors of the transfer efficiency evaluation coefficient of the logistics transfer station and is the average of the values ​​of the sub-factors in the logistics multi-source customer service data; The target sub-factor combination processing module is used to screen and obtain the Pearson correlation coefficient The sub-factor combination corresponding to a value greater than or equal to the preset Pearson correlation coefficient threshold is the target sub-factor combination; The update processing module is used to add the sub-factor of the transfer efficiency evaluation coefficient of the logistics transfer station in the target sub-factor combination to the logistics multi-source customer service data to obtain updated logistics multi-source customer service data.

10. A method for evaluating and analyzing regional business environment based on big data analysis, characterized in that: The steps are as follows: Collect logistics multi-source customer service data at subdivided time nodes within a preset time period; The logistics multi-source customer service data includes order data, logistics terminal coverage data, logistics terminal throughput data, logistics terminal transportation route data, customer feedback order response data, logistics terminal fees and logistics terminal efficiency index data; The approximation of the logistics multi-source customer service data at each segmented time node is obtained based on the logistics multi-source customer service data at two adjacent segmented time nodes. The abnormal growth evaluation value of each data p' in the logistics multi-source customer service data at each segmented time node is obtained based on the approximation of the logistics multi-source customer service data at each segmented time node. ; Get the sampled data p of each data p' in the logistics multi-source customer service data of each segmented time node, then calculate the local outlier factor based on the neighborhood k obtained by cross-validation based on the sampled data p, and output the local outlier factor evaluation value LOF of the sampled data p ; The local outlier factor evaluation value LOF based on the sampled data p and the abnormal growth evaluation value of each data , calculate the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node ; Logistics benefit evaluation processing operations are performed based on the comprehensive abnormality evaluation score of each data in the logistics multi-source customer service data at each segmented time node.

Citation Information

Patent Citations

  • Multi-dimensional data comprehensive analysis business environment monitoring method and system

    CN118333665A

  • Determining order lead time for a supply chain using a probability distribution for expected order lead time

    US20040230474A1