Customer data processing method based on artificial intelligence, computer equipment and medium
By using AI-based customer data processing methods and leveraging parcel volume data and geographical features to predict customer acquisition models, the problem of low customer acquisition efficiency in the traditional logistics industry has been solved, achieving efficient and accurate identification of potential customers and cost reduction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
The traditional logistics industry relies on human experience for customer acquisition, which is inefficient and unable to quickly respond to market changes and customer needs. Furthermore, the rule-based approach to extracting high-potential customers lacks flexibility and dynamism, making it difficult to fully consider the diversity and complexity of customer behavior, thus limiting the effectiveness of customer acquisition.
By using AI-based customer data processing methods, we acquire parcel volume data for time-series correlation analysis, capture geographical regional characteristics, and use pre-trained customer mining models to make predictions and identify potential customers.
It improves the efficiency and accuracy of customer acquisition in the logistics industry, enables efficient and accurate identification of potential customers, reduces customer acquisition costs, and adapts to the needs of large-scale data processing.
Smart Images

Figure CN121745974A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a customer data processing method, computer equipment, and medium based on artificial intelligence. Background Technology
[0002] In the highly competitive logistics industry, customer acquisition and relationship maintenance are crucial. By identifying customers and providing high-quality services and solutions tailored to their needs, logistics companies can achieve long-term, stable business growth and increased profits. Currently, the traditional logistics industry relies heavily on the experience of sales personnel for customer acquisition. However, manual customer acquisition is time-consuming and labor-intensive, making it difficult to respond quickly to rapidly changing market and customer demands. In other words, relying on manual experience for customer acquisition is currently inefficient. Therefore, improving the efficiency of customer acquisition in the logistics industry has become a pressing technical problem to be solved. Summary of the Invention
[0003] The main objective of this application is to propose a customer data processing method, computer device, and medium based on artificial intelligence, which aims to process customer data in the logistics field based on artificial intelligence for customer mining, thereby improving the efficiency of customer mining in the logistics industry.
[0004] To achieve the above objectives, a first aspect of this application proposes a customer data processing method based on artificial intelligence, the method comprising:
[0005] Obtain the shipment volume data of the customers to be explored;
[0006] A time-series correlation analysis is performed on the shipment volume data to obtain the analysis results, which are used to characterize the correlation between the shipment volume of the target customer and time.
[0007] Based on the shipment volume data, the geographic region characteristics of the customer to be mined are captured, and the geographic region characteristics include the embedded vector representation of the network nodes of the logistics network to which the customer to be mined belongs;
[0008] The model prediction result is obtained by using a pre-trained customer mining model to predict the analysis results and the geographical area characteristics. The model prediction result is used to characterize whether the customer to be mined is a potential customer.
[0009] In some embodiments, the analysis results include: trend analysis results;
[0010] The correlation analysis of the mail volume data over time series yields the following results:
[0011] The time series of the shipment volume data is decomposed based on a preset autocorrelation decomposition transformer to obtain the trend analysis results; the autocorrelation decomposition transformer includes a time series prediction model based on a self-attention mechanism.
[0012] In some embodiments, the analysis results further include: seasonal analysis results;
[0013] The step of performing a time-series correlation analysis on the mail volume data to obtain the analysis results also includes:
[0014] Based on the autocorrelation-based decomposition transformer, the time series of the dispatch volume data is seasonally decomposed to obtain the seasonal analysis results.
[0015] In some embodiments, capturing the geographic region characteristics of the customer to be mined based on the shipment volume data includes:
[0016] Based on the shipment volume data, determine the network nodes of the logistics network to which the customer to be mined belongs;
[0017] A fixed number of target neighbor nodes are randomly sampled from the neighbor nodes of the network node;
[0018] Aggregation processing is performed on the target neighbor nodes;
[0019] Based on the characteristics of the network node itself and the aggregation result obtained by aggregating the target neighbor node, feature update processing is performed;
[0020] The process of stacking multiple layers for aggregation and feature updating results in the embedding vector representation of the network nodes, and the embedding vector representation is used as the geographic region feature of the customer to be mined.
[0021] In some embodiments, the method further includes:
[0022] Obtain historical shipment volume data and profile tag data of sample customers;
[0023] Model training samples are generated based on the historical parcel volume data.
[0024] Customer tags are generated based on the profile tag data, and the customer tags are used to characterize whether the sample customer is a potential customer;
[0025] The customer tags are defined as the model prediction results of the preset model to be trained, and the model to be trained is trained based on the model training samples to obtain the customer mining model.
[0026] In some embodiments, generating model training samples based on the historical mail volume data includes:
[0027] A time-series correlation analysis is performed on the historical parcel volume data to obtain sample data analysis results, which are used to characterize the correlation between the historical parcel volume of the sample customer and historical time.
[0028] Based on the historical shipment volume data, the historical geographic region characteristics of the sample customers are captured. The historical geographic region characteristics are used to characterize the embedded vector representation of the network nodes of the historical logistics network to which the sample customers belong.
[0029] The training samples are generated based on the analysis results of the sample data and the characteristics of the historical geographical region.
[0030] In some embodiments, the model training samples include positive samples and negative samples; after generating model training samples based on the sample data analysis results and the historical geographic region features, the method further includes:
[0031] In the case of an imbalance between the positive samples and the negative samples, the preference degree between each first target sample and each second target sample is calculated; wherein, the number of first target samples is less than the number of second target samples, the first target sample is one of the positive samples and the negative samples, and the second target sample is the other of the positive samples and the negative samples;
[0032] Weights are assigned to each of the first target samples based on their respective preferences.
[0033] A synthetic sample of each first target sample is generated based on the weight of each first target sample;
[0034] The filtered synthetic samples obtained by filtering each of the synthetic samples, and the second target sample, are used as model training samples.
[0035] To achieve the above objectives, a second aspect of this application provides a customer data processing apparatus based on artificial intelligence, the apparatus comprising:
[0036] The acquisition module is used to acquire the shipment volume data of the customers to be explored;
[0037] The analysis module is used to perform time series correlation analysis on the shipment volume data to obtain analysis results, which are used to characterize the correlation between the shipment volume of the target customer and time.
[0038] The feature capture module is used to capture the geographic region features of the customer to be mined based on the shipment volume data. The geographic region features include the embedded vector representation of the network nodes of the logistics network to which the customer to be mined belongs.
[0039] The model prediction module is used to predict the analysis results and the geographical area characteristics using a pre-trained customer mining model to obtain model prediction results, which are used to characterize whether the customer to be mined is a potential customer.
[0040] To achieve the above objectives, a third aspect of the present application provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first aspect.
[0041] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0042] To achieve the above objectives, a fifth aspect of the present application provides a computer program product storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0043] This application proposes an artificial intelligence-based customer data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product, which acquires the parcel volume data of customers to be mined; performs time-series correlation analysis on the parcel volume data to obtain analysis results, the analysis results being used to characterize the correlation between the parcel volume of the customer to be mined and time; captures the geographical region characteristics of the customer to be mined based on the parcel volume data, the geographical region characteristics including the embedded vector representation of the network nodes of the logistics network to which the customer to be mined belongs; and uses a pre-trained customer mining model to predict the analysis results and the geographical region characteristics to obtain model prediction results, the model prediction results being used to characterize whether the customer to be mined is a potential customer.
[0044] Compared to relying on sales personnel's experience for customer acquisition, this embodiment of the application pre-trains a customer acquisition model. It then acquires the parcel volume data of the customer to be acquired and analyzes the time series of this data to obtain analytical results. Based on this parcel volume data, it captures geographical region characteristics. Therefore, the trained customer acquisition model, based on the analytical results and geographical region characteristics, can predict whether the customer to be acquired is a potential customer. Thus, this embodiment of the application can use artificial intelligence to process customer parcel volume data in the logistics field for customer acquisition, thereby improving the efficiency of customer acquisition in the logistics industry. Attached Figure Description
[0045] Figure 1 A flowchart illustrating the steps of the AI-based customer data processing method provided in some embodiments of this application;
[0046] Figure 2 for Figure 1 A detailed flowchart of step S103;
[0047] Figure 3 A schematic diagram of the model application process involved in the artificial intelligence-based customer data processing method provided in the embodiments of this application;
[0048] Figure 4 A flowchart illustrating the steps of the AI-based customer data processing method provided in this application in other embodiments;
[0049] Figure 5 for Figure 4 A detailed flowchart of step S402;
[0050] Figure 6 for Figure 4 A schematic diagram of another detailed step in step S402;
[0051] Figure 7 A flowchart illustrating the steps of the AI-based customer data processing method provided in some embodiments of this application;
[0052] Figure 8 A schematic diagram of the model training process involved in the artificial intelligence-based customer data processing method provided in the embodiments of this application;
[0053] Figure 9 This is a schematic diagram of the structure of the artificial intelligence-based customer data processing device provided in the embodiments of this application;
[0054] Figure 10 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0058] First, let's analyze some of the terms used in this application:
[0059] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0060] The overall concept of the embodiments of this application will be described next.
[0061] In the highly competitive logistics industry, identifying and maintaining relationships with high-potential clients is crucial for a logistics company's business growth and profitability. By identifying high-potential clients and providing high-quality services and solutions tailored to their specific needs, logistics companies can achieve long-term, stable business growth and profit increases. Furthermore, establishing solid partnerships with high-potential clients helps logistics companies expand their business scope and deepen cooperation, enabling them to jointly explore strategic cooperation opportunities. This, in turn, allows logistics companies to build a differentiated competitive advantage in the market, enhancing brand value and market position.
[0062] In the traditional logistics industry, identifying high-potential customers largely relies on the experience of sales personnel. However, relying on manual customer acquisition is time-consuming and labor-intensive, resulting in low efficiency and an inability to quickly respond to market changes and customer needs. Furthermore, this method is difficult to replicate and scale (e.g., it's difficult to quickly train new sales personnel), thus hindering the expansion and development of logistics companies. Additionally, the significant differences in experience and capabilities among sales personnel lead to inconsistent customer acquisition standards and results, making it difficult to establish unified strategies and standards.
[0063] Furthermore, the logistics industry also employs a rule-based approach to extract high-potential customers. However, this method often relies on pre-defined rules or standards for screening, lacking flexibility and dynamism, and making it difficult to respond to rapid changes in customer behavior and new market trends. Moreover, rule-based extraction of high-potential customers may depend too heavily on fixed rules and indicators, making it susceptible to the influence of specific factors and failing to fully consider the diversity and complexity of customer behavior, thus limiting the scope of high-potential customers identified. Additionally, this method fails to fully utilize hidden information and potential patterns within customer behavior data, lacking in-depth data mining and analysis capabilities, further limiting the effectiveness of high-potential customer identification. Going a step further, rule-based extraction of high-potential customers suffers from oversimplification; that is, it may oversimplify the customer behavior analysis process, neglecting the multidimensional information and complex circumstances behind customers, resulting in an insufficient understanding of the true nature of high-potential customers.
[0064] Therefore, how to continuously explore and optimize the methods or means of identifying potential customers in the logistics industry, so as to improve the accuracy and effectiveness of potential customer identification, achieve efficient and accurate identification and prediction of potential customers, improve the accuracy of potential customer identification and the competitiveness of communication service products, thereby reducing customer acquisition costs and adapting to the needs of large-scale data processing, has become a technical problem that urgently needs to be solved in the industry.
[0065] Based on this, embodiments of this application provide a customer data processing method, apparatus, computer device, computer-readable storage medium, and computer program product based on artificial intelligence. By pre-training a customer mining model, and then obtaining the parcel volume data of the customer to be mined and analyzing and predicting the time series of the parcel volume data to obtain the analysis results, and capturing the geographical features based on the parcel volume data, the trained customer mining model can make predictions based on the analysis results and geographical features to obtain the model prediction results characterizing whether the customer to be mined is a potential customer.
[0066] Compared to relying on sales personnel's experience for customer acquisition, this application's embodiments can process customer shipment volume data in the logistics field based on artificial intelligence for customer acquisition. By capturing the diversity and complexity of massive customer behaviors in the logistics field based on customer shipment volume data, it can deeply mine hidden information and potential patterns in customer behavior, thereby greatly improving the accuracy and effectiveness of potential customer acquisition. It achieves high-efficiency and high-accuracy identification and prediction of potential customers, thereby improving the accuracy of potential customer identification and the competitiveness of the applied communication service products, reducing customer acquisition costs, and adapting to the needs of large-scale data processing.
[0067] Based on the overall concept of the embodiments of this application described above, specific embodiments of the customer data processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on artificial intelligence provided in the embodiments of this application are proposed. First, the various specific embodiments of the customer data processing method based on artificial intelligence in the embodiments of this application are described in detail.
[0068] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0069] Foundational artificial intelligence technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0070] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0071] Furthermore, the AI-based customer data processing method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and AI platforms; the software can be an application implementing the AI-based customer data processing method, but is not limited to the above forms.
[0072] Furthermore, embodiments of this application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0073] For ease of understanding and explanation, the following description will use the AI-based customer data processing method provided in the embodiments of this application as an example for detailed explanation. The implementation of the AI-based customer data processing method provided in the embodiments of this application for any of the above-described subject matters can refer to the process described below for applying the AI-based customer data processing method to a terminal device.
[0074] Please refer to Figure 1 , Figure 1 The flowchart illustrates the steps of the artificial intelligence-based customer data processing method provided in some embodiments of this application. It should be understood that, although... Figure 1 The figure shows the execution order of some method steps, but based on different design needs of actual applications, the AI-based customer data processing method provided in this application embodiment can of course adopt a different execution order of method steps than that shown in the figure. That is, Figure 1 The order of the method steps shown does not constitute a limitation on the execution logic order of the AI-based customer data processing method provided in the embodiments of this application. Any other method based on... Figure 1 Reasonable changes to the sequence of steps shown should be included within the protection scope of the AI-based customer data processing method provided in this application embodiment.
[0075] like Figure 1 As shown, in some embodiments, the AI-based customer data processing method provided in this application may include, but is not limited to, steps S101 to S104.
[0076] Step S101: Obtain the shipment volume data of the customer to be mined.
[0077] When identifying and predicting potential customers in the logistics field, the terminal equipment first marks one or more customers that need to be identified and predicted as potential customers as potential customers. Then, it obtains the shipment volume data of each of the one or more potential customers in the past period from the business system of the logistics company used to provide shipment services to customers.
[0078] It should be noted that the terminal device can pre-establish a communication connection with the logistics company's business system, thereby obtaining the shipment volume data of the customer to be analyzed by sending data retrieval requests to the business system. Furthermore, the business system can record information about the customer and their shipments after each shipment management service is provided to that customer.
[0079] In some embodiments, the terminal device can also directly obtain the parcel volume data of the customers to be explored from a local database or an external data storage device. For example, logistics company staff can pre-upload the collected parcel volume data of the customers to be explored to the terminal device for storage. In this way, the terminal device can directly read the parcel volume data of the customers to be explored uploaded by the staff from the local database. For example, logistics company staff can pre-store the collected parcel volume data of the customers to be explored on a storage medium such as a portable hard drive or USB flash drive. Then, by connecting the storage medium to the device interface provided by the terminal device, the terminal device can read the parcel volume data of the customers to be explored from the external storage medium through the device interface.
[0080] It should be understood that, based on different design needs in practical applications, terminal devices may, of course, use other methods not listed here to obtain the parcel volume data of the customer to be mined. That is, the AI-based customer data processing method provided in this application embodiment does not limit the specific method by which the terminal device obtains customer parcel volume data.
[0081] Step S102: Perform time series correlation analysis on the mail volume data to obtain analysis results. The analysis results are used to characterize the correlation between the mail volume of the target customer and time.
[0082] After acquiring the parcel volume data of the customer to be explored, the terminal device can further analyze the parcel volume data to obtain corresponding analysis results. For example, for each customer to be explored, the terminal device can perform correlation analysis on the time series of the parcel volume data to determine the relationship between the parcel volume of the customer to be explored in the past period and the past period, thereby obtaining analysis results that characterize the relationship between the parcel volume of the customer to be explored and time.
[0083] Step S103: Capture the geographical region features of the customer to be mined based on the shipment volume data. The geographical region features include the embedded vector representation of the network nodes of the logistics network to which the customer to be mined belongs.
[0084] It should be noted that, in some embodiments, geographical region features may also be referred to as spatial region features.
[0085] When performing time-series correlation analysis on parcel volume data, the terminal device can also capture the geographical region characteristics of each customer to be analyzed based on their parcel volume data. This geographical region characteristic can then be used to characterize the embedded vector representation of the network nodes of the logistics network to which the customer belongs.
[0086] Step S104: The analysis results and the geographical area features are predicted by a pre-trained customer mining model to obtain the model prediction results. The model prediction results are used to characterize whether the customer to be mined is a potential customer.
[0087] After performing feature engineering on the parcel volume data of the target customer to obtain the aforementioned analysis results and geographic region features, the terminal device uses these results and features as model input. A pre-trained customer mining model then predicts the results and features, outputting a model prediction. Based on this prediction, the terminal device determines whether the target customer is a potential customer. For example, assuming the model prediction output includes 0 and 1, where 0 represents a potential customer and 1 represents a non-potential customer, the terminal device only needs to identify whether the model prediction is 0 or 1 to determine if the target customer is a potential customer.
[0088] In this embodiment, when identifying and predicting potential customers in the logistics field using a terminal device, one or more customers whose potential is to be identified and predicted are first marked as potential customers. Then, the device acquires the shipment volume data of each potential customer over a past period. Next, the terminal device further analyzes this shipment volume data to obtain corresponding analysis results. For example, for each potential customer's shipment volume data, the terminal device performs correlation analysis on the time series of the shipment volume data to determine the relationship between the potential customer's shipment volume over a past period and that past period, thereby obtaining an analysis result characterizing the relationship between the potential customer's shipment volume and time. Furthermore, when performing time series correlation analysis on the shipment volume data, the terminal device can also capture the geographical region characteristics of each potential customer based on their shipment volume data. This geographical region characteristic is then used to characterize the embedded vector representation of the network nodes of the logistics network to which the potential customer belongs. Finally, the terminal device uses the analysis results and geographical features as model inputs to predict the analysis results and geographical features using a pre-trained customer mining model, and obtains the model prediction results output by the customer mining model. Based on the model prediction results, it then determines whether the customer to be mined is a potential customer.
[0089] Thus, compared to relying on sales personnel's experience for customer mining, this application embodiment can process customer shipment volume data in the logistics field based on artificial intelligence for customer mining. Furthermore, by capturing the diversity and complexity of massive customer behaviors in the logistics field based on customer shipment volume data, it can deeply mine hidden information and potential patterns in customer behavior, thereby greatly improving the accuracy and effectiveness of mining potential customers. It achieves high-efficiency and high-accuracy identification and prediction of potential customers, thereby improving the accuracy of potential customer identification and the competitiveness of the applied communication service products, reducing customer mining costs, and adapting to the needs of large-scale data processing.
[0090] In some embodiments, the analysis results obtained by the terminal device may include trend analysis results. That is, the terminal device can perform trend analysis on the time series of the mail volume data of the customer to be mined, thereby obtaining trend analysis results.
[0091] Based on this, step S102 above, which involves performing a time-series correlation analysis on the mail volume data to obtain the analysis results, may include the following steps:
[0092] The time series of the shipment volume data is decomposed based on a preset autocorrelation decomposition transformer to obtain the trend analysis results; the autocorrelation decomposition transformer includes a time series prediction model based on a self-attention mechanism.
[0093] When performing correlation analysis on the time series of mail volume data, terminal equipment can use an autocorrelation decomposition transformer to decompose the time series of mail volume data into trends, thereby obtaining trend analysis results.
[0094] It should be noted that autocorrelation-based decomposition transformers can include: time series forecasting models based on self-attention mechanisms. The core idea of this time series forecasting model is to decompose the time series of mail volume data into trend and seasonal components, and to model these trend and seasonal components through a self-attention mechanism.
[0095] In some embodiments, the analysis results obtained by the terminal device may also include seasonal analysis results. That is, in addition to obtaining trend analysis results by performing trend analysis on the time series of the mail volume data of the customer to be mined, the terminal device also obtains seasonal analysis results by performing seasonal analysis on the time series of the mail volume data.
[0096] Based on this, step S102 above may also include the following steps:
[0097] Based on the autocorrelation-based decomposition transformer, the time series of the dispatch volume data is seasonally decomposed to obtain the seasonal analysis results.
[0098] When terminal equipment uses an autocorrelation-based decomposition transformer to perform trend decomposition on the time series of mail volume data, it can also use the same decomposition transformer to perform seasonal decomposition on the time series in a synchronous or asynchronous manner, thereby obtaining seasonal analysis results.
[0099] In some embodiments, the terminal device can perform trend decomposition on the time series of mail volume data using the trend decomposition module in the time series forecasting model to obtain the trend component in the time series. The terminal device then uses this trend component as the trend analysis result. Furthermore, the terminal device can also synchronously or asynchronously perform seasonal decomposition on the time series of mail volume data using the seasonality analysis module in the time series forecasting model to obtain the seasonal component in the time series. The terminal device then uses this seasonal component as the seasonality analysis result.
[0100] For example, the detailed process of using the autocorrelation-based decomposition transformer described above to perform trend and seasonal decomposition on the time series of the customer's mail volume data to be mined by the terminal device is as follows:
[0101] First, the time series of mail volume data X = {x1, x2, ..., x} L The trend is decomposed using the Trend Decomposition Block (TDB). The formula for calculating the trend component T in the TDB is shown below:
[0102] T = Conv1d(X; W)
[0103] Where Conv1d represents a one-dimensional convolution operation, and W is the convolution kernel (weights), the specific operation is as follows:
[0104]
[0105] Among them, T i w represents the value of the trend component at time point i. j These are the weights of the convolution kernel, k is the radius of the convolution kernel (which determines the size of the convolution kernel), and x... i+j It is the value of the input time series at position i+j.
[0106] Then, the time series of mail volume data X = {x1, x2, ..., x} LThe decomposition is performed using the Seasonal Decomposition Block (SDB). The seasonal component S in the SDB is mainly obtained by capturing the periodic changes in the time series X through a self-attention mechanism. The relevant calculation formula is shown below:
[0107] Input time series X = {x1, x2, ..., x} L The query vector Q, key vector K, and value vector V are obtained through linear transformation:
[0108] Q = XW Q K = XW K V = XW V
[0109] Among them, W Q W K and W v It is the weight matrix of the linear transformation.
[0110] In addition, the attention score is calculated using the softmax activation function:
[0111]
[0112] Where Attention(Q,K,V) represents the attention score, Q, K, and V are the query, key, and value vectors, respectively, and d k It is the dimension of the key vector.
[0113] A multi-head attention mechanism is used to compute multiple attention heads in parallel, and the results are concatenated to obtain the seasonal component S:
[0114] S = MultiHeadAttention(Q,K,V)
[0115] MultiHeadAttention represents a multi-head attention mechanism, where each head computes its own self-attention, and then the results are concatenated and subjected to a linear transformation to obtain the final output.
[0116] MultiHead(Q,K,V)=Concat(head1,…,head h W O
[0117] The calculation for each attention head is as follows:
[0118]
[0119] in, W is the weight matrix of the i-th attention head. O It is the output weight matrix.
[0120] In this embodiment, although the traditional time series forecasting algorithm Prophet can decompose time series into trend and seasonal terms, the algorithm relies on a piecewise linear model to fit the trend, which is not good at capturing nonlinear and complex time series dependencies. The static decomposition method cannot adapt to the dynamic changes of time series data, and the high computational complexity will further weaken the computational efficiency of long-term series, and it is difficult to cope with the massive data scenarios in the logistics field.
[0121] Therefore, this embodiment uses an autocorrelation-based decomposition transformer to perform trend and seasonal decomposition on the time series of mail volume data, and employs a self-attention mechanism to dynamically adjust attention weights for dynamic modeling. This effectively captures the long-term and short-term dependencies between customer mail volume and time, adapting to the dynamic changes in time series data. Furthermore, the trend-seasonal decomposition block in the decomposition transformer automatically decomposes the trend and seasonal components in the time series without requiring manual parameter setting for the decomposition model. The combination of self-attention and convolution can handle non-linear trends and seasonal components. The sparse self-attention mechanism reduces computational complexity, enabling efficient processing of long-term series data and significantly improving prediction performance and computational efficiency in large-scale complex scenarios.
[0122] In some embodiments, the terminal device may capture geographic region features of the customer to be mined by employing graph sampling and aggregation techniques.
[0123] Please refer to Figure 2 , Figure 2 for Figure 1 A detailed flowchart of step S103.
[0124] like Figure 2 As shown, step S103 above, which captures the geographical region characteristics of the customer to be mined based on the shipment volume data, may include steps S201 to S205 as shown below.
[0125] Step S201: Determine the network node of the logistics network to which the customer to be mined belongs based on the shipment volume data.
[0126] When the terminal device captures the geographical features of the customer based on the shipment volume data of the customer to be mined, it first determines the distribution area, province and city in the logistics network to which the customer belongs based on the shipment volume data, and regards the distribution area, province and city as network nodes in the network structure of the logistics network to which the customer belongs, so as to facilitate the subsequent use of graph sampling and aggregation technology to convert these network nodes into continuous vector representations.
[0127] Step S202: Randomly sample a fixed number of target neighbor nodes from the neighbor nodes of the network node.
[0128] Step S203: Perform aggregation processing on the target neighbor nodes.
[0129] After identifying the network nodes of the logistics network to which the customer belongs, the terminal device further employs image sampling technology to randomly sample a fixed number (e.g., k, where k is a positive integer greater than 1) of target neighbor nodes from among the neighbor nodes of each network node. Then, the terminal device uses aggregation technology to perform feature aggregation processing on the randomly sampled target neighbor nodes to obtain the aggregation result.
[0130] Step S204: Based on the network node's own characteristics and the aggregation result obtained by aggregating the target neighbor node, perform feature update processing.
[0131] After the terminal device performs aggregation processing on some target neighbor nodes of the network node to obtain the aggregation result, it further combines the network node's own characteristics with the aggregation result and generates a new node embedding representation through nonlinear transformation, thereby performing feature update processing.
[0132] Step S205 involves stacking multiple layers for aggregation and feature update to obtain the embedding vector representation of the network node, and using the embedding vector representation as the geographical region feature of the customer to be mined.
[0133] The terminal device obtains the final embedding vector representation of the network node through a process of aggregation and feature update by stacking multiple layers. The terminal device then uses this embedding vector representation as the captured geographic region feature of the customer to be mined.
[0134] It should be noted that when the terminal device treats the distribution areas, provinces, and cities in the logistics network to which the customer to be mined belongs as network nodes in the network structure, and then uses graph sampling and aggregation technology to convert them into continuous vector representations, the input of the graph embedding layer is a matrix of the connection relationships between nodes and a node feature matrix. The node feature matrix is represented by the number of shipments between the province or city to which the customer to be mined is located. The output of the graph embedding layer is a node embedding vector, which is a continuous vector. This vector learns the embedding representation of the node by sampling and aggregating the neighbor information of the node, which can efficiently process large-scale graph data.
[0135] For example, the relevant calculation formula is as follows:
[0136] (1) Initialize node characteristics: The terminal device denotes the initial characteristics of each network node v as follows:
[0137] (2) Neighbor Sampling: For each network node v, the terminal device randomly samples a fixed number of neighbors {u1, u2, ..., u3} from its neighbor node set N(v). k};
[0138] (3) Feature aggregation: The terminal device uses the mean aggregation function to generate new feature representations of the nodes. The formula is as follows:
[0139]
[0140] Among them, W k is the weight matrix of the k-th layer, σ is the non-linear activation function, concat represents the feature chaining operation, and mean represents mean aggregation;
[0141] (4) Feature update: The terminal device combines the features of network node v with the aggregation result and generates a new node embedding representation through nonlinear transformation:
[0142]
[0143] Here, Agg represents the aggregation function, and concat represents the feature connection operation.
[0144] (5) Final node representation: The terminal device obtains the final embedded representation of network node v through a stacked multi-layer aggregation and feature update process. Where L is the number of layers in the network.
[0145] In this embodiment, considering that traditional graph embedding methods (such as DeepWalk, node2vec, etc.) require random walks and global computations on the entire graph, fixed static features cannot dynamically capture the graph's evolution, making it difficult to extend to dynamic graphs or newly added nodes and inefficiently handle large-scale graph data, this embodiment introduces graph sampling and aggregation techniques to capture the geographical features of the target customer. Through neighbor sampling and feature aggregation, computational complexity and storage requirements are effectively reduced, enabling the processing of large-scale graph data. Furthermore, aggregation processing effectively captures the local structural information and attribute features of nodes by aggregating node features. It can also generate more comprehensive embeddings for new nodes without recalculating the entire graph, thus handling heterogeneous graph data and improving model scalability. This offers significant advantages in scenarios such as dynamic graphs and real-time recommendations.
[0146] The following is a complete embodiment of the customer data processing method based on artificial intelligence provided in this application, which uses a customer mining model to predict whether a customer to be mined is a potential customer.
[0147] Please refer to Figure 3 , Figure 3A schematic diagram of the model application process involved in the artificial intelligence-based customer data processing method provided in the embodiments of this application.
[0148] like Figure 3 As shown, in some embodiments, the terminal device, targeting potential customers in the logistics field, determines whether a customer is a potential customer through two major processes: data acquisition feature engineering and model prediction application. Specifically, the terminal first acquires the shipment volume data of the potential customer, then uses an autocorrelation decomposition transformer to perform trend and seasonal decomposition on the time series of the shipment volume data, thereby obtaining corresponding trend analysis results and seasonal analysis results. Furthermore, the terminal device also uses graph sampling and aggregation techniques to capture the geographical region characteristics of the potential customer. In addition, based on the trend analysis results, seasonal analysis results, and geographical region characteristics of the potential customer, the terminal device can also process imbalanced samples using an ensemble method based on SMOTE-IPF resampling, thereby completing the data acquisition feature engineering for the potential customer. After this, the terminal device can use the trend analysis results, seasonal analysis results, and geographical region characteristics of the potential customer, which have undergone the above data acquisition feature engineering, as model data. It can then call a pre-trained customer mining model to perform predictions, obtaining the model prediction results output by the customer mining model. These model prediction results, representing whether the potential customer is a potential customer, support business applications.
[0149] Before using a customer mining model to identify and predict whether a customer to be mined is a potential customer, the terminal device can train the model based on the historical shipment volume data and profile tag data of some sample customers to obtain the customer mining model. Then, the customer mining model is applied to predict the shipment volume data of the customer to be mined.
[0150] Please refer to Figure 4 , Figure 4 The flowcharts of the AI-based customer data processing method provided in this application are shown in other embodiments.
[0151] like Figure 4 As shown, in some embodiments, the AI-based customer data processing method provided in this application further includes steps S401 to S404 as shown below.
[0152] Step S401: Obtain historical shipment volume data and profile tag data of sample customers.
[0153] When training a customer mining model, the terminal device first obtains the historical shipment volume data and profile tag data of one or more sample customers.
[0154] It should be noted that sample customers can be pre-designated by logistics company staff, and their historical shipment volume data can be obtained one by one by the terminal device from the logistics company's business system used to provide shipment services to customers; alternatively, the terminal device can directly obtain the sample customers' historical shipment volume data from a local database or external data storage devices. Furthermore, the sample customer profile tags can be pre-built by logistics company staff and stored together with the historical shipment volume data, so that the terminal device can obtain the profile tag data simultaneously with the historical shipment volume data.
[0155] Step S402: Generate model training samples based on the historical mail volume data.
[0156] Step S403: Generate customer tags based on the profile tag data. The customer tags are used to characterize whether the sample customer is a potential customer.
[0157] Step S404: Define the customer tag as the model prediction result of the preset model to be trained, and train the model to be trained based on the model training samples to obtain a customer mining model.
[0158] After acquiring historical parcel volume data and profile tag data of sample customers, the terminal device can perform feature engineering on the historical parcel volume data, similar to the time-series correlation analysis and geographic feature capture operations described above, to generate model training samples. Furthermore, while generating model training samples based on historical parcel volume data, the terminal device can also generate customer tags needed for model training based on the sample customer profile tag data. Then, by defining the customer tags as the model prediction results of the preset model to be trained, the terminal device can begin training the model to be trained based on the generated model training samples to obtain the customer mining model.
[0159] It should be noted that the customer tags are similar to the prediction results of the aforementioned model, and are primarily used to characterize whether a sample customer is a potential customer. Furthermore, when training the model on the terminal device, a histogram-based algorithm can be used for feature value binning, and a leaf-based tree growth strategy can be selected for splitting. Then, the model is trained by minimizing the binary cross-entropy loss function.
[0160] In this embodiment, a customer mining model is pre-trained and then saved. Afterward, by deploying the saved model to the production environment and inputting the shipment volume data of the customer to be mined, predictions can be made to identify whether the customer is a potential customer. Thus, compared to relying on sales personnel experience for customer mining, this embodiment can use artificial intelligence to process customer shipment volume data in the logistics field for customer mining, thereby improving the efficiency of customer mining in the logistics industry.
[0161] Please refer to Figure 5 , Figure 5 for Figure 4 A detailed flowchart of step S402.
[0162] like Figure 5 As shown, in some embodiments, step S402 may include steps S501 to S503 as shown below.
[0163] Step S501: Perform time series correlation analysis on the historical mail volume data to obtain sample data analysis results. The sample data analysis results are used to characterize the correlation between the historical mail volume of the sample customer and historical time.
[0164] The terminal device performs trend decomposition on the time series of historical parcel volume data based on a preset autocorrelation decomposition transformer to obtain the trend component in the sample data analysis results. Furthermore, the terminal device also uses this decomposition transformer to perform seasonality analysis on the time series of historical parcel volume data, thereby obtaining the seasonality component in the sample data analysis results.
[0165] It should be noted that the terminal device uses an autocorrelation decomposition transformer to perform time series correlation analysis on historical parcel volume data, which is the same as step S102 and its detailed steps in the above embodiment. The same content will not be described again here.
[0166] Step S502: Capture the historical geographic region features of the sample customer based on the historical shipment volume data. The historical geographic region features are used to characterize the embedded vector representation of the network nodes of the historical logistics network to which the sample customer belongs.
[0167] Similar to the terminal device described above, which captures the geographical features of the customer to be mined based on the shipment volume data of the customer to be mined, the terminal device can also capture the historical geographical features of the sample customer based on the shipment volume data of the sample customer when performing time series correlation analysis on the historical shipment volume data of the sample customer. In this way, the embedded vector representation of the network node of the historical logistics network to which the sample customer belongs can be characterized by the historical geographical features.
[0168] It should be noted that the terminal device captures the historical geographical characteristics of the sample customer based on the sample customer's mailing volume data, which is the same as step S103 and its detailed steps in the above embodiment. The same content will not be described again here.
[0169] Step S503: Generate model training samples based on the analysis results of the sample data and the characteristics of the historical geographical area.
[0170] After obtaining the sample data analysis results and the historical geographic region characteristics of the sample customers, the terminal device can combine the sample data analysis results and the historical geographic region characteristics to construct and generate model training samples for model training.
[0171] In some embodiments, the model training samples generated by the terminal device include positive and negative samples. When the terminal device builds model training samples to train the customer mining model, if it finds that there is an imbalance between positive and negative samples in the model training samples, the terminal device can handle the imbalance problem based on the SMOTE-IPF resampling ensemble method.
[0172] Please refer to Figure 6 , Figure 6 for Figure 4 A schematic diagram of another detailed step in step S402.
[0173] like Figure 6 As shown, in step S503 above, after generating training samples for the model based on the sample data analysis results and the historical geographic area features, the customer data processing method based on artificial intelligence provided in this application embodiment may further include steps S601 to S604 as shown below.
[0174] Step S601: In the case of an imbalance between the positive samples and the negative samples, calculate the preference degree between each first target sample and each second target sample; wherein the number of first target samples is less than the number of second target samples, the first target sample is one of the positive samples and the negative samples, and the second target sample is the other of the positive samples and the negative samples.
[0175] Step S602: Assign weights to each of the first target samples based on their respective preference degrees.
[0176] Step S603: Generate a synthetic sample of each first target sample based on the weight of each first target sample.
[0177] Step S604: The filtered synthetic sample obtained by filtering each of the synthetic samples, and the second target sample, are used as model training samples.
[0178] After constructing training samples for the generated model, the terminal device detects the number of positive and negative samples. When an imbalance is found in the ratio of positive to negative samples, the terminal device addresses this imbalance problem using an ensemble method based on SMOTE-IPF resampling. Specifically, the terminal device first calculates the preference between the fewer first target samples (positive and negative samples) and the more numerous second target samples. Then, based on the preference of each first target sample, the terminal device assigns weights to them. Next, according to the assigned weights, the terminal device generates synthetic samples for each first target sample, and filters these synthetic samples to obtain filtered synthetic samples. Finally, the terminal device combines the filtered synthetic samples with the second target samples as training samples for the model.
[0179] It's important to note that the SMOTE-IPF resampling-based ensemble method combines resampling techniques with ensemble learning to address imbalanced datasets. It's suitable for datasets where the majority class significantly outnumbers the minority class, improving the model's classification performance on the minority class. SMOTE-IPF adjusts the weights of generated synthetic samples by calculating the preference between instances, effectively improving the quality and diversity of synthetic samples to address the challenges of imbalanced datasets. Ensemble learning, by combining the predictions of multiple classifiers, can significantly enhance overall classification performance.
[0180] For example, the relevant calculation process of the terminal device in handling the imbalance problem of positive and negative samples in the model training samples based on the SMOTE-IPF resampling ensemble method is as follows:
[0181] (1) Calculate instance preference:
[0182] The terminal device uses Euclidean distance to calculate the preference degree between each sample and other samples, given a small number of samples. The calculation formula is as follows:
[0183] Given two samples x i and x j Preference(x) i ,x j It can be defined as:
[0184] Where, dist(x) i ,x j ) is sample x i and xj The Euclidean distance or other distance measure between them.
[0185] (2) Weight allocation:
[0186] The terminal device assigns weights to each of the few samples based on the calculated preference degree, thereby determining the probability of generating a synthetic sample. The calculation formula for the weight allocation method based on instance preference degree normalization is shown below:
[0187]
[0188] Where, Preference(x i ) is sample x i The total preference, and ∑ j≠i Preference(x i ,x j ) is the sample x i The sum of the preferences for all other samples.
[0189] (3) Synthetic sample generation:
[0190] The terminal device generates synthetic samples for each sample, based on weighted allocation. That is, for each sample x... i Choose with x i The nearest neighbor minority class sample x nn and in x i With x nn Generate new synthetic samples x on the line segments between them new This is to increase the number and diversity of minority class samples.
[0191] (4) Filter out invalid synthesized samples:
[0192] After generating synthetic samples, the terminal device uses filters to remove those synthetic samples that are not helpful for the classification task or may introduce noise, thereby improving the quality of the synthetic samples and the stability of the model.
[0193] In this embodiment, the SMOTE method is traditionally used to handle imbalanced positive and negative samples. However, the SMOTE method, by overly averaging neighbor selection when generating synthetic samples, easily overlooks some important minority class samples. This prevents the model from fully learning the features of these important samples, and the SMOTE method may also generate some noise samples that are useless or even harmful to classification. Therefore, this embodiment uses the SMOTE-IPF algorithm to handle the imbalanced positive and negative sample problem. By calculating the preference and weight allocation of instances, the algorithm can focus on minority class samples that are important to the classification task, more accurately select samples to generate synthetic samples, and filter out those synthetic samples that are not helpful for classification, thereby reducing the negative impact of noise samples on the model and improving the model's ability to identify minority class samples. The attention to boundary samples and key samples further promotes the generated synthetic samples to better reflect the real data analysis, thereby improving the model's performance on complex distributed data. In addition, the dynamic adjustment of sample weights and the balancing of the dataset during ensemble learning can significantly improve the classifier's performance and efficiently handle the imbalanced sample problem of large-scale datasets.
[0194] After the terminal device builds and generates model training samples, and uses these samples to train the model to obtain a customer mining model by defining customer tags, it can further verify the effectiveness of the customer mining model based on the recall rate metric of the model.
[0195] Please refer to Figure 7 , Figure 7 A flowchart illustrating the steps of the AI-based customer data processing method provided in some other embodiments of this application.
[0196] like Figure 7 As shown, in some embodiments, after defining the customer tag as the model prediction result of the preset model to be trained in step S404 above, and after training the model to be trained based on the model training samples to obtain the customer mining model, the customer data processing method based on artificial intelligence provided in this application embodiment may further include steps S701 to S703 as shown below.
[0197] Step S701: Test the customer mining model to obtain all model test prediction results output by the customer mining model.
[0198] After training the model to be trained based on the model training samples to obtain the customer mining model, the terminal device inputs the pre-built test samples into the customer mining model to test the customer mining model. In addition, the terminal device also obtains all model test prediction results output by the customer mining model for the test samples.
[0199] Step S702: Divide the number of results representing test customers as potential customers in all model test prediction results by the total number of model test prediction results to obtain the recall rate index of the customer mining model.
[0200] After the terminal device obtains all the model test prediction results output by the customer mining model for the test samples, it further calculates the recall rate metric of the customer mining model based on all the model test prediction results. That is, the terminal device divides the number of results in all the model test prediction results that indicate that the test customer is a potential customer by the total number of model test prediction results to obtain the recall rate metric of the customer mining model.
[0201] Step S703: If the recall rate indicator does not meet the preset recall rate indicator requirement, the customer mining model is updated to obtain a new customer mining model until the recall rate indicator of the new customer mining model meets the recall rate indicator requirement.
[0202] It should be noted that the preset recall rate requirement can be that the recall rate is greater than or equal to a preset threshold (such as 80%).
[0203] After calculating the recall rate of the customer mining model, the terminal device compares this recall rate with a preset threshold (80%). If the comparison shows that the recall rate is less than the preset threshold (e.g., 80%), the terminal device confirms that the customer mining model's recall rate does not meet the recall requirement, thus confirming that the customer mining model's validity validation has failed. The terminal device then updates the customer mining model (e.g., adjusting model training parameters) to obtain a new customer mining model. Subsequently, the terminal device calculates the recall rate for the new customer mining model to determine if its recall rate meets the recall requirement. This process is repeated until the recall rate of the new customer mining model meets the recall requirement. In other words, if the recall rate is greater than or equal to the preset threshold (e.g., 80%), the customer mining model's recall rate is confirmed to meet the recall requirement. At this point, the terminal device confirms that the customer mining model's validity validation has passed and saves the customer mining model for application in the production environment for identifying and predicting potential customers.
[0204] In some embodiments, the terminal device can also combine the area under the receiver operating characteristic curve (ROC) (AUC) and the recall rate of the customer mining model to verify the effectiveness of the model, screen and save the model.
[0205] Next, we present a complete embodiment of the customer data processing method based on artificial intelligence proposed in this application, which includes constructing model training samples, training and validating the model, saving the customer mining model, and applying it to the production environment.
[0206] Please refer to Figure 8 , Figure 8 A schematic diagram of the model training process involved in the artificial intelligence-based customer data processing method provided in the embodiments of this application.
[0207] like Figure 8 As shown, when training the customer mining model, the terminal device first acquires historical parcel volume data and profile tag data of sample customers. Then, the terminal device uses an autocorrelation-based decomposition transformer to perform trend and seasonal decomposition on the time series of historical parcel volume data of sample customers, and captures the historical geographic region features of customers through graph sampling and aggregation techniques, thereby constructing training samples for the model. Furthermore, the terminal device uses an ensemble method based on SMOTE-IPF resampling to handle the imbalance between positive and negative samples. Then, the terminal device trains the model based on these training samples. During model training, a histogram-based algorithm is used for feature value binning, and a leaf-based tree growth strategy is selected for splitting. The model is then trained by minimizing the binary cross-entropy loss function, resulting in a trained customer mining model. For this customer mining model, the terminal device also uses AUC and recall metrics to verify the model's effectiveness, filter, and save the model. Finally, the terminal device deploys the saved customer mining model to the production environment, so that subsequent predictions can be made by inputting the parcel volume data of the customer to be mined, thus identifying whether the customer is a potential customer.
[0208] This embodiment employs an autocorrelation-based decomposition transformer to perform trend and seasonal decomposition on the time series of parcel volume data; captures spatial regional features through graph sampling and aggregation techniques; addresses the imbalance between positive and negative samples using an ensemble method based on SMOTE-IPF resampling; performs feature value binning using a histogram-based algorithm; selects a leaf-based tree growth strategy for splitting; trains the model by minimizing the binary cross-entropy loss function; and validates and filters the model's effectiveness by combining AUC and recall metrics. This approach effectively captures the diversity and complexity of massive customer behavior in the logistics field, deeply mining hidden information and potential patterns, thereby greatly improving the accuracy and effectiveness of potential customer identification. It achieves efficient and accurate identification and prediction of potential customers, enhancing the accuracy of identifying high-potential customers and the competitiveness of communication service products, while also reducing customer acquisition costs and adapting to the needs of large-scale data processing.
[0209] Please see Figure 9This application also provides an artificial intelligence-based customer data processing apparatus, which can implement the above-described artificial intelligence-based customer data processing method. The apparatus includes:
[0210] The acquisition module is used to acquire the shipment volume data of the customers to be explored;
[0211] The analysis module is used to perform time series correlation analysis on the shipment volume data to obtain analysis results, which are used to characterize the correlation between the shipment volume of the target customer and time.
[0212] The feature capture module is used to capture the geographic region features of the customer to be mined based on the shipment volume data. The geographic region features include the embedded vector representation of the network nodes of the logistics network to which the customer to be mined belongs.
[0213] The model prediction module is used to predict the analysis results and the geographical area characteristics using a pre-trained customer mining model to obtain model prediction results, which are used to characterize whether the customer to be mined is a potential customer.
[0214] It should be noted that the specific implementation of the AI-based customer data processing device provided in this application is basically the same as the specific implementation of the AI-based customer data processing method described above, and will not be repeated here.
[0215] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned artificial intelligence-based customer data processing method. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0216] Please see Figure 10 , Figure 10 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes:
[0217] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0218] The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 using the artificial intelligence-based client data processing method of the embodiments of this application.
[0219] Input / output interface 1003 is used to implement information input and output;
[0220] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0221] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0222] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0223] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described artificial intelligence-based customer data processing method.
[0224] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0225] This application also provides a computer program product that stores a computer program, which, when executed by a processor, implements the above-described artificial intelligence-based customer data processing method.
[0226] The customer data processing method, device, computer equipment, computer-readable storage medium, and computer program product based on artificial intelligence provided in this application embodiment pre-train a customer mining model, then obtain the parcel volume data of the customer to be mined and analyze and predict the time series of the parcel volume data to obtain the analysis result, and capture the geographical area characteristics based on the parcel volume data. Thus, the trained customer mining model can predict whether the customer to be mined is a potential customer based on the analysis result and the geographical area characteristics.
[0227] Compared to relying on sales personnel's experience for customer acquisition, this application's embodiments can process customer shipment volume data in the logistics field based on artificial intelligence for customer acquisition. By capturing the diversity and complexity of massive customer behaviors in the logistics field based on customer shipment volume data, it can deeply mine hidden information and potential patterns in customer behavior, thereby greatly improving the accuracy and effectiveness of potential customer acquisition. It achieves high-efficiency and high-accuracy identification and prediction of potential customers, thereby improving the accuracy of potential customer identification and the competitiveness of the applied communication service products, reducing customer acquisition costs, and adapting to the needs of large-scale data processing.
[0228] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0229] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0230] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0231] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0232] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0233] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0234] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0235] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0236] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0237] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0238] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An artificial intelligence-based customer data processing method, characterized by, The method comprises: obtaining shipment volume data of a to-be-mined customer; performing time series correlation analysis on the shipment volume data to obtain an analysis result, the analysis result being used to represent a correlation between shipment volume and time of the to-be-mined customer; capturing geographical area features of the to-be-mined customer based on the shipment volume data, the geographical area features comprising an embedding vector representation of a network node of a logistics network to which the to-be-mined customer belongs; performing prediction on the analysis result and the geographical area features by using a pre-trained customer mining model to obtain a model prediction result, the model prediction result being used to represent whether the to-be-mined customer is a potential customer.
2. The method of claim 1, wherein, The analysis result comprises a trend analysis result. The time series correlation analysis on the shipment volume data to obtain an analysis result comprises: performing trend decomposition on a time series of the shipment volume data based on a preset decomposition transformer with autocorrelation to obtain the trend analysis result; the decomposition transformer with autocorrelation comprises a time series prediction model based on a self-attention mechanism.
3. The method of claim 2, wherein, The analysis result further comprises a seasonality analysis result. The time series correlation analysis on the shipment volume data to obtain an analysis result further comprises: performing seasonality decomposition on the time series of the shipment volume data based on the decomposition transformer with autocorrelation to obtain the seasonality analysis result.
4. The method of claim 1, wherein, The capturing of the geographical area features of the to-be-mined customer based on the shipment volume data comprises: determining a network node of a logistics network to which the to-be-mined customer belongs based on the shipment volume data; randomly sampling a fixed number of target neighbor nodes from neighbor nodes of the network node; performing aggregation processing on the target neighbor nodes; performing feature updating processing based on self-features of the network node and an aggregation result obtained by performing the aggregation processing on the target neighbor nodes; stacking a process of the aggregation processing and a process of the feature updating processing to obtain an embedding vector representation of the network node, and taking the embedding vector representation as the geographical area features of the to-be-mined customer.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: obtaining historical shipment volume data and portrait label data of a sample customer; generating model training samples based on the historical shipment volume data; generating a customer label based on the portrait label data, the customer label being used to represent whether the sample customer is a potential customer; defining the customer label as a model prediction result of a preset to-be-trained model, and training the to-be-trained model based on the model training samples to obtain a customer mining model.
6. The method of claim 5, wherein, The generating of the model training samples based on the historical shipment volume data comprises: performing time series correlation analysis on the historical shipment volume data to obtain sample data analysis results, the sample data analysis results being used to represent a correlation between historical shipment volume and historical time of the sample customer; capturing historical geographical area features of the sample customer based on the historical shipment volume data, the historical geographical area features being used to represent an embedding vector representation of a network node of a historical logistics network to which the sample customer belongs; generating the model training sample based on the sample data analysis result and the historical geographical area feature generation model.
7. The method of claim 6, wherein, The model training sample includes positive samples and negative samples; after the model training sample is generated based on the sample data analysis result and the historical geographical area feature generation model, the method further includes: In the case of imbalance between the positive samples and the negative samples, the preference degrees between each first target sample and each second target sample are calculated; wherein the number of the first target samples is less than the number of the second target samples, the first target samples are one of the positive samples and the negative samples, and the second target samples are the other of the positive samples and the negative samples; Based on the respective preference degrees of each first target sample, a weight is assigned to each first target sample; Based on the weight of each first target sample, a synthetic sample of each first target sample is generated; The filtered synthetic samples obtained by filtering each synthetic sample and the second target samples are used as model training samples.
8. The method of claim 5, wherein, After the customer mining model is obtained by training the to-be-trained model based on the model training sample, the method further includes: Testing the customer mining model to obtain all model test prediction results output by the customer mining model; Dividing the number of results in the all model test prediction results representing that a test customer is a potential customer by the number of the all model test prediction results to obtain a recall rate index of the customer mining model; In the case that the recall rate index does not meet a preset recall rate index requirement, updating the customer mining model to obtain a new customer mining model until the recall rate index of the new customer mining model meets the recall rate index requirement.
9. A computer device, comprising: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the customer data processing method based on artificial intelligence in any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the customer data processing method based on artificial intelligence in any one of claims 1 to 8. The computer program is executed by the processor to implement the customer data processing method based on artificial intelligence in any one of claims 1 to 8.