Distributed big data customer information management and analysis method

By adopting a multi-level distributed architecture and a variety of data processing technologies in customer information management and analysis, the efficiency and accuracy problems of traditional methods when processing massive multi-modal customer data are solved, efficient customer information management and analysis are achieved, and the company's market competitiveness and customer service quality are improved.

CN120013546AInactive Publication Date: 2025-05-16SHENZHEN TIANYIYI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510113565.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When traditional customer information management and analysis methods face massive and multi-modal customer data, it is difficult to achieve efficient storage, rapid access, in-depth mining and real-time analysis, making it difficult for enterprises to obtain and utilize customer information in a timely manner, affecting market competitiveness and customer service quality.

Method used

A multi-level distributed architecture is adopted, including a multi-modal data acquisition layer, a data preprocessing layer, a data storage layer, a computing layer and a result display layer. Through user portraits, decision tree algorithms, load balancing algorithms, neural network models and distributed search algorithms, efficient processing and analysis of multi-modal customer information is achieved.

Benefits of technology

It realizes efficient processing and storage of massive multimodal customer information, accurately captures customer characteristics and behavior changes, improves the efficiency and accuracy of customer service management, and enhances the company's market competitiveness and customer service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013546A_ABST
    Figure CN120013546A_ABST
Patent Text Reader

Abstract

The invention provides a distributed big data customer information management and analysis method. The method comprises the following steps: constructing a multi-layer distributed architecture; carrying out user portraying and granularity division on the multi-modal customer information, and respectively storing the multi-modal customer information on different nodes; dividing the customer service management task into a plurality of customer service management sub-tasks according to a decision tree algorithm, and correspondingly matching different nodes; dynamically scheduling a matching result by using a load balancing algorithm, constructing and training a neural network model of a corresponding subtask, and performing dynamic fusion; searching multi-modal data related to the customer service management subtask by using a distributed search algorithm, and performing iterative training on the model; and processing the real-time data stream through stream-oriented calculation, and carrying out customer service management decision. The invention provides a set of comprehensive, efficient and intelligent customer information management and analysis solution, effectively solves many problems of a traditional method in a big data environment, and plays an important role in promoting sustainable development of enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data analysis and management, and in particular to a distributed big data customer information management and analysis method. Background Art

[0002] In today's digital business environment, the customer information accumulated by enterprises is growing explosively and has multimodal characteristics, covering text, images, audio, video, structured data, etc. Traditional customer information management and analysis methods have exposed many shortcomings when faced with such massive and complex data.

[0003] First, traditional data storage is mainly a centralized data storage architecture, which is difficult to bear the storage and rapid access requirements of large-scale data. As the amount of customer data continues to increase, the storage capacity and processing power of centralized databases quickly reach bottlenecks, resulting in extremely low data storage and retrieval efficiency, which seriously affects the timely acquisition and use of customer information by enterprises.

[0004] Secondly, traditional analysis methods are mostly designed based on small-scale, structured data, and are unable to cope with multimodal data. They lack the ability to effectively integrate and deeply mine data of different modalities, and are unable to fully extract the rich information contained in multimodal data, making it difficult for companies to fully and accurately understand customer needs, behavior patterns, and value characteristics.

[0005] Thirdly, in terms of data processing timeliness, traditional methods usually adopt batch processing mode, which cannot meet the urgent needs of enterprises for real-time data analysis and decision-making. In a rapidly changing market environment, enterprises need to adjust their strategies in a timely manner according to the latest behavior and feedback of customers, but the lag of traditional methods makes enterprises miss many business opportunities and put them at a disadvantage in market competition.

[0006] The increasing personalization and diversification of customer needs require companies to accurately segment and locate customers and provide customized products and services. However, due to the limitations of data processing and analysis capabilities, existing methods are unable to achieve fine classification of customers and personalized portrait construction, resulting in the lack of pertinence in the marketing and service strategies of enterprises, and difficulty in improving customer satisfaction and loyalty. With the intensification of market competition, companies are increasingly sensitive to customer churn. However, traditional customer management methods are difficult to accurately predict the risk of customer churn in advance, and often measures are taken only after customers have already churned, which is too late. Therefore, companies are in urgent need of an innovative and efficient distributed big data customer information management and analysis method to break through the difficulties of traditional technologies, achieve in-depth mining and efficient use of customer information, enhance the competitiveness of companies in the market and the quality of customer service, and effectively respond to increasingly complex and changing business challenges. Summary of the invention

[0007] The technical problem to be solved by the present invention is to provide a distributed big data customer information management and analysis method, which provides a comprehensive, efficient and intelligent customer information management and analysis solution, effectively solves many problems of traditional methods in the big data environment, and plays an important role in promoting the sustainable development of enterprises.

[0008] To solve the above technical problems, an embodiment of the present invention provides the following technical solution: a distributed big data customer information management and analysis method, comprising the following steps:

[0009] Build a multi-layer distributed architecture, including multimodal data collection layer, data preprocessing layer, data storage layer, computing layer and result display layer;

[0010] Create user portraits based on multimodal customer information, statistical, behavioral, and value attributes, divide the portrait data into granularities based on time and attribute dimensions, and store them on different nodes.

[0011] Dividing the customer service management task into a plurality of customer service management subtasks according to a decision tree algorithm, and matching the subtasks with the different nodes according to attribute association;

[0012] Use load balancing algorithms to dynamically schedule matching results, use the divided data as input, and customer service management decision strategies as task outputs to build and train neural network models for corresponding subtasks;

[0013] Dynamically integrate the neural network models of the corresponding subtasks to obtain a multimodal customer service management decision model;

[0014] Using a distributed search algorithm to search for multimodal data related to the customer service management subtask, and iteratively training the multimodal customer service management decision model;

[0015] Real-time data streams are processed through streaming computing, and the processed data stream information is input into the iteratively trained multimodal customer service management decision model to make customer service management decisions.

[0016] Furthermore, the multimodal customer information is used to create user profiles according to the attributes of statistics, behavior, and value, specifically:

[0017] Multimodal customer information is classified according to statistics, behavior, and value; statistics include age, gender, region, consumption preference, and consumption frequency; behaviors include browsing behavior, search behavior, interaction behavior, and purchase behavior; and values ​​include customer loyalty, consumption level, and consumption frequency;

[0018] Use clustering algorithms to cluster users and build user portraits;

[0019] Establish a user portrait update mechanism based on a time window to update the portrait in a timely manner according to the user's latest behavior and information.

[0020] Furthermore, the portrait data is divided into granularities based on the time dimension and the attribute dimension, specifically:

[0021] Analyze business needs and data characteristics, determine the appropriate time dimension division period, which includes day, week, month, quarter, and year; at the same time, sort out the key attributes for profiling, including statistical attributes, behavioral attributes, and value attributes;

[0022] Group the profile data according to the determined time period, calculate summary statistics for each time period group, and store the summary information together with the original profile data;

[0023] Create different storage structures to store data of different attribute categories respectively, and store the divided portrait data in corresponding storage media, which include distributed file systems, relational databases or NoSQL databases; during the storage process, create indexes based on the characteristics of the data and query requirements.

[0024] Furthermore, the customer service management task is divided into a plurality of customer service management subtasks according to the decision tree algorithm, specifically:

[0025] Determine the characteristics and target variables of the decision tree, collect a large amount of historical data related to customer service management tasks, including basic customer information, consumption behavior, service interaction records, and business results data; analyze and select characteristic variables that are relevant to customer service management decisions;

[0026] Select a decision tree algorithm, use training data to train a decision tree model, and prune the constructed decision tree;

[0027] According to the final structure and branch results of the decision tree, different customer service management subtasks are determined.

[0028] Furthermore, matching the subtasks with the different nodes according to attribute associations is specifically as follows:

[0029] Analyze the subtask attribute requirements, match the key attributes that each customer service management subtask depends on, and analyze the degree of dependence and correlation of each subtask on different attributes;

[0030] Identify the attribute distribution of data storage nodes, obtain the attribute characteristics of the stored data through metadata management tools or data directories, and classify and mark the data content of each node;

[0031] Establish attribute association mapping rules. According to the attribute requirements of the subtask and the attribute distribution of the data storage node, create the mapping rules between the two by building an association matrix. The mapping rules should be based on the semantics and business logic of the data.

[0032] The attribute association mapping rules are stored in a management database, and the management database uses a query and retrieval algorithm to query storage node information that matches the subtask attribute requirements;

[0033] Regularly evaluate the matching effect of subtasks and nodes. If some subtasks have data acquisition delays or incomplete data during execution, re-analyze the attribute association and storage node distribution, and dynamically adjust the matching results.

[0034] Furthermore, the use of a load balancing algorithm to dynamically schedule matching results is specifically as follows:

[0035] Real-time monitoring of the resource status of computing nodes, including at least key indicators such as CPU usage, memory usage, and network bandwidth. By deploying a resource monitoring agent on each computing node, resource usage data is regularly collected and fed back to the load balancing controller. At the same time, the I / O performance and storage capacity of storage nodes are evaluated and monitored.

[0036] Prioritize customer service management subtasks based on their urgency, importance, and business impact;

[0037] Select and configure the load balancing algorithm. When there are new matching results that need to be scheduled, the load balancing controller makes decisions based on resource monitoring information, task priority, and the selected load balancing algorithm.

[0038] After the task is assigned to the computing node, the execution status of the task is continuously tracked, feedback information on the execution of the task is collected, the effect of the load balancing algorithm is evaluated, and the parameters or strategies of the load balancing algorithm are adjusted according to the evaluation results.

[0039] Furthermore, the divided data is used as input, the customer service management decision strategy is used as task output, and a neural network model corresponding to the subtask is constructed and trained, specifically:

[0040] Divide the divided data into training set, validation set and test set according to a certain ratio;

[0041] Select different neural network models according to the subtask type, and set the number of network layers and nodes;

[0042] Use random initialization methods to assign weights and biases to neural networks, define loss functions and optimizers

[0043] Use the training set data to train the neural network model, pass the input data through each layer of the neural network in turn, calculate the loss between the predicted value and the true value, and then calculate the gradient through the back propagation algorithm to update the model parameters;

[0044] After each round of training, the performance of the model is evaluated using the validation set data. The loss value, accuracy, or mean square error indicator on the validation set is calculated to observe whether the model is overfitting or underfitting.

[0045] After the model training is completed, the test set data is used to perform a final evaluation of the model, and the model is tuned based on the evaluation results.

[0046] Furthermore, the neural network models of the corresponding subtasks are dynamically integrated to obtain a multimodal customer service management decision model, specifically:

[0047] Select the model involved in the fusion based on the correlation between the subtasks;

[0048] Use a unified test dataset to evaluate each selected subtask neural network model, and record the stability and generalization ability of each model on different sample subsets;

[0049] Select different fusion strategies according to different task attributes, the fusion strategy being weighted average fusion, model stacking fusion or voting fusion;

[0050] Regularly re-evaluate model performance, and based on the new evaluation results, use an adaptive algorithm to automatically adjust the fusion weights according to the model's prediction error on new data, and dynamically adjust the weights or parameters in the fusion strategy.

[0051] Furthermore, the distributed search algorithm is used to search for multimodal data related to the customer service management subtask, and the multimodal customer service management decision model is iteratively trained, specifically:

[0052] Analyze the requirements of each customer service management subtask, define the scope of distributed search, and determine the type and characteristics of the required multimodal data; the scope of the distributed search includes databases, file systems, and external data interfaces;

[0053] Select the appropriate distributed search algorithm based on the data distribution characteristics and search requirements;

[0054] Initialize the distributed search task, distribute the search request to each data storage node, use the selected search algorithm to retrieve data, each node will filter and process the searched relevant data, and return the data fragments related to the subtask; summarize the data returned by each node to form a multimodal data set related to the customer service management subtask;

[0055] Integrate the newly searched multimodal data with the original training data set, convert the data format and extract features according to the characteristics of the data and the input requirements of the model, so that the new data matches the input dimension and data type of the model;

[0056] The multimodal customer service management decision model is iteratively trained using the integrated data set until the model's performance on the validation set reaches a stable state or meets the preset performance indicator requirements, completing the update of the multimodal customer service management decision model.

[0057] Furthermore, the real-time data stream is processed by stream computing, specifically:

[0058] Acquire multimodal data from various data sources in real time, and configure corresponding collection strategies and format parsing rules for different types of data sources;

[0059] According to the needs of the customer service management subtask, feature extraction is performed on the real-time data stream. For text-based customer evaluation data, the natural language processing library is used to perform lexical analysis and syntactic analysis to extract features that are not limited to keywords and sentiment tendencies. For image data, the pre-trained convolutional neural network model is used to extract the key features of the image and convert the extracted features into a format suitable for model input.

[0060] When new feature data flows in, it is integrated with the original model training data to monitor the performance indicators of the multimodal customer service management decision model in real time;

[0061] The real-time prediction results of the multimodal customer service management decision model are fed back to the customer service management system for real-time processing, and the customer service strategy is dynamically adjusted based on the results of the real-time processing.

[0062] The beneficial effects of the above technical solution of the present invention are as follows:

[0063] 1. The present invention realizes efficient processing and storage of massive multimodal customer information through a multi-level distributed architecture. The multimodal data acquisition layer can widely collect data from a variety of data sources, and use efficient acquisition tools and message queues to ensure that the data is transmitted to the preprocessing layer in real time, effectively avoiding data congestion and loss. The data preprocessing layer uses advanced algorithms to clean, unify the format and extract preliminary features of the data to improve data quality and availability. In terms of storage, the reasonable combination of distributed file systems, relational databases and NoSQL databases can accurately store data based on data types and characteristics, give full play to the advantages of each storage medium, greatly improve data storage and retrieval efficiency, and break through the capacity and performance bottlenecks of traditional centralized architectures.

[0064] 2. The present invention builds user portraits based on statistical, behavioral and value attributes, and can accurately capture customer characteristics and behavioral changes through clustering algorithms and dynamic update mechanisms. This enables companies to gain a deep understanding of customer needs, preferences and values, and provide customers with highly personalized product recommendations, service solutions and marketing activities.

[0065] 3. The present invention rationally divides the customer service management task into multiple subtasks through the decision tree algorithm, and accurately matches the storage node through attribute association, thereby realizing efficient docking between tasks and data. The load balancing algorithm dynamically schedules according to the resource status of the computing node and the storage node and the task priority, ensuring that the system resources are fully and reasonably utilized, and avoiding idle or overloaded resources. This not only improves the task processing efficiency, but also reduces the system operation cost, enabling enterprises to handle more customer service tasks under limited resource conditions and improve overall operational efficiency.

[0066] 4. The neural network model constructed and trained for subtasks in this invention, combined with dynamic fusion strategy, can make full use of the rich information in multimodal data to improve the accuracy and generalization ability of the model. The application of distributed search algorithm and streaming computing enables the model to obtain the latest data in time and perform iterative training, and update the decision-making strategy in real time to adapt to the dynamic changes of the market and customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is an overall flow chart of the distributed big data customer information management and analysis method of the present invention;

[0068] Figure 2 A flowchart of the distributed big data customer information management and analysis customer service management subtask division of the present invention;

[0069] Figure 3 A flow chart of the distributed big data customer information management and analysis method of the present invention using a load balancing algorithm to dynamically schedule matching results;

[0070] Figure 4 A flow chart showing the dynamic fusion of neural network models of corresponding subtasks for the distributed big data customer information management and analysis method of the present invention;

[0071] Figure 5 The present invention is a flow chart of the distributed big data customer information management and analysis method for processing real-time data streams through streaming computing. DETAILED DESCRIPTION

[0072] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0073] like Figure 1As shown, the present invention proposes a distributed big data customer information management and analysis method, comprising the following steps:

[0074] S1. Build a multi-level distributed architecture, including multimodal data collection layer, data preprocessing layer, data storage layer, computing layer and result display layer;

[0075] S2. Create user portraits based on the attributes of statistics, behavior, and value for multimodal customer information, divide the portrait data into granularities based on the time dimension and attribute dimension, and store them on different nodes;

[0076] S3, dividing the customer service management task into a plurality of customer service management subtasks according to a decision tree algorithm, and matching the subtasks with the different nodes according to attribute association;

[0077] S4. Use the load balancing algorithm to dynamically schedule the matching results, take the divided data as input, take the customer service management decision strategy as the task output, and build and train the neural network model corresponding to the subtask;

[0078] S5. Dynamically integrate the neural network models of the corresponding subtasks to obtain a multimodal customer service management decision model;

[0079] S6. Using a distributed search algorithm to search for multimodal data related to the customer service management subtask, and iteratively training the multimodal customer service management decision model;

[0080] S7. Process the real-time data stream through streaming computing, input the processed data stream information into the iteratively trained multimodal customer service management decision model, and make customer service management decisions.

[0081] The following is a detailed description of how each of the above steps works:

[0082] Step S1 builds a multi-level distributed architecture, specifically:

[0083] (1) Multimodal data acquisition layer,

[0084] Diversified data sources: covering internal enterprise systems, such as basic customer information and communication records in the CRM system, order data in the transaction system; external data sources, such as public comments and feedback from customers on social media platforms, and sensor data generated by IoT devices (if it involves customer service scenarios).

[0085] Collection tool selection: Use Logstash to collect log file data; use Kafka Connect to capture change data from the database; for social media data, use the API provided by the corresponding platform combined with customized scripts for collection.

[0086] Data transmission: The collected data is transmitted to the data preprocessing layer in real time. Efficient and reliable data flow can be achieved through message queues such as Apache Kafka to ensure orderly transmission and processing of data.

[0087] (2) Data preprocessing layer

[0088] Data cleaning: remove duplicate data, such as marking processed data records through hash algorithms; correct erroneous data by setting a reasonable range for numerical data and verifying or correcting it if it exceeds the range; fill in missing data by using the mean, median or predicted value based on machine learning algorithms.

[0089] Format unification: Convert data in different formats into a standard format, such as unifying different time formats into the ISO8601 standard timestamp; for text data, perform unified encoding processing to ensure data consistency.

[0090] Preliminary feature extraction: Perform lexical and syntactic analysis on text data to extract keywords; for image data, use a simple image feature extraction algorithm to obtain basic features to prepare for subsequent more in-depth feature engineering.

[0091] (3) Data storage layer

[0092] Distributed file system: Use distributed file systems such as Ceph to store large-scale unstructured data, such as original images and audio files, and use their high scalability and fault tolerance to ensure secure storage and efficient access to data.

[0093] Relational database: Use relational databases such as MySQL and PostgreSQL to store structured data, such as basic customer information and transaction details, to facilitate complex queries and transaction processing.

[0094] NoSQL database: MongoDB is introduced to store semi-structured data, such as customer behavior logs. Its flexible document structure can adapt well to the diversity of data while providing efficient read and write performance.

[0095] (4) Computation layer

[0096] Distributed computing framework: Apache Spark, Apache Flink and other distributed computing frameworks can be used to process large-scale data in parallel and improve computing efficiency. For example, when processing massive amounts of customer behavior data, Spark's RDD (Resilient Distributed Dataset) or Flink's DataStream API can be used for efficient data conversion and analysis.

[0097] Task scheduling: According to the priority, data volume and computing resource requirements of the task, the computing tasks are reasonably allocated to different computing nodes. Through resource schedulers such as YARN (Yet Another Resource Negotiator), the computing resources can be optimized to avoid waste and overload of resources.

[0098] Model training and reasoning: This layer trains and reasoned the multimodal customer service management decision model. Distributed machine learning libraries such as Horovod are used to accelerate the model training process, and the model is updated in real time for new incoming data to support decision making.

[0099] (5) Result display layer

[0100] Data visualization tools: Use data visualization tools such as Tableau and PowerBI to display the analysis results of the calculation layer in the form of intuitive charts and reports. For example, use a bar chart to display the complaint rate of customers in different regions, and use a line chart to show the changing trend of customer satisfaction.

[0101] Interactive interface design: Design a user-friendly interactive interface that allows customer service personnel, managers and other different roles to filter and view data according to their needs. Provide real-time data update function to ensure that the information obtained by users is always the latest, facilitating timely decision-making.

[0102] Mobile support: Taking into account the needs of mobile office, a mobile-adapted display interface is provided to facilitate users to view key data and analysis results anytime and anywhere, thereby improving the timeliness and flexibility of decision-making.

[0103] In this embodiment, the step S2 generates user profiles based on the multimodal customer information according to the attributes of statistics, behavior, and value, and specifically includes the following steps:

[0104] (1) Classification of statistical information:

[0105] Age, gender, and region: extracted from the CRM system or user registration information to ensure data accuracy and completeness.

[0106] Consumption preferences: Analyze the transaction records of e-commerce platforms, count the frequency and amount of customers' purchases of different categories of goods, and determine their consumption preferences. For example, if the amount of electronic products purchased by a customer accounts for a high proportion of the total consumption amount, it can be judged that the customer has a preference for electronic products.

[0107] Consumption frequency: It is determined by counting the number of purchases made by a customer within a certain time frame (such as per month or per quarter).

[0108] (2) Behavioral information classification:

[0109] Browsing behavior: extracted from user behavior logs, recording information such as the pages viewed by customers on the website or app, dwell time, and browsing order.

[0110] Search behavior: Analyze the keywords users enter in the search box, search frequency, clicks on search results, etc.

[0111] Interaction behavior: includes operations such as likes, comments, and shares by customers on the website or app, as well as communication records with customer service.

[0112] Purchasing behavior: In addition to consumption frequency, it also includes information such as the time, place, payment method, and combination of purchased items.

[0113] (3) Classification of valuable information:

[0114] Customer loyalty: Comprehensively consider factors such as the customer's repeat purchase rate, purchase duration, and recommendation behavior. Customers with high repeat purchase rates, long purchase duration, and recommendation behavior have higher loyalty.

[0115] Consumption level: According to the customer's consumption amount, it can be divided into three levels: high, medium and low.

[0116] Consumption frequency: It is consistent with the consumption frequency in the statistical information, but in the value assessment, customers with high consumption frequency are generally considered to have higher value.

[0117] (4) After the data is classified, clustering algorithms are used to cluster users based on the three types of data to build user portraits. Clustering algorithms include K-means algorithm and DBSCAN algorithm:

[0118] K-means algorithm: is a commonly used clustering algorithm, suitable for situations where the amount of data is large and the distribution is relatively uniform. The algorithm randomly selects K initial cluster centers, iteratively calculates the distance from each data point to the cluster center, assigns the data point to the cluster with the closest distance, and updates the cluster center until the cluster center no longer changes or the maximum number of iterations is reached.

[0119] DBSCAN algorithm: If the data has large density differences, the DBSCAN algorithm is more suitable. It divides the density-connected data points into a cluster based on the density of the data points. It can find clusters of any shape and identify noise points in the data.

[0120] (5) After selecting the clustering algorithm, perform cluster analysis, specifically:

[0121] For example, for the K-means algorithm, the number of clusters K needs to be determined in advance. The appropriate K value can be selected by methods such as the elbow rule and the silhouette coefficient. The elbow rule calculates the sum of squared clustering errors (SSE) under different K values ​​and selects the K value at the inflection point of the SSE curve; the silhouette coefficient calculates the silhouette coefficient of each data point and selects the K value with the largest average silhouette coefficient. The preprocessed data is input into the selected clustering algorithm to obtain different user clusters. Each cluster represents a group of users with similar characteristics.

[0122] (6) After cluster analysis, construct a user profile, specifically:

[0123] Analyze and summarize the user characteristics in each cluster to form a user profile. For example, the users in a cluster are mainly between 25 and 35 years old, the gender is mainly male, the region is concentrated in first-tier cities, the consumption preference is electronic products, the consumption level is high, the consumption frequency is high, and the loyalty is high. The user profile of this cluster can be defined as "young male high-consumption electronic product enthusiasts in first-tier cities."

[0124] As business activities are constantly changing, it is necessary to determine the appropriate time window based on business needs and the frequency of data changes. For example, for businesses with high real-time requirements, you can set the time window to an hour or day level; for businesses with relatively slow data changes, you can set the time window to a week or month level. In each time window, collect the user's latest behavior and information. For example, in a one-day time window, collect the user's browsing behavior, search behavior, purchase behavior, etc.

[0125] Integrate the newly collected data with the original user portrait data. For new statistical information, such as age, gender, region, etc., if there is any change, update it; for new behavioral information, such as browsing behavior, search behavior, etc., add it to the original behavioral record; for new value information, such as customer loyalty, consumption level, etc., re-evaluate and update it according to the new data. When a certain amount of new data has been accumulated or a certain number of time windows have passed, re-cluster all user data. This is because user behavior and characteristics may change over time, and re-clustering can more accurately reflect the user's latest situation.

[0126] When re-clustering, the original number of clusters and cluster centers can be retained as initial values ​​to reduce the amount of calculation and improve the clustering effect. At the same time, according to the new data characteristics, the parameters of the clustering algorithm are adjusted and optimized. According to the results of re-clustering, the user portrait is updated. The user characteristics in each cluster are re-analyzed and summarized to ensure that the user portrait can accurately reflect the latest situation of the user. The updated user portrait is applied to the business, such as personalized recommendation, precision marketing, etc., to improve the business effect and user satisfaction.

[0127] After the above processing is completed, the portrait data is divided into granularities based on the time dimension and attribute dimension, specifically:

[0128] First, determine the time period and key attributes, analyze business needs and data characteristics, and determine the appropriate time dimension to divide the period, such as day, week, month, quarter, year, etc. At the same time, sort out the key attributes for portraits, including statistical attributes (such as age, consumption amount, etc.), behavioral attributes (such as browsing behavior, purchasing behavior, etc.) and value attributes (such as customer lifetime value, loyalty, etc.).

[0129] Secondly, group the profile data according to the determined time period. For example, group the daily customer profile data into a group and store them in a folder or data table named after the date. For each time period group, you can calculate some summary statistics, such as the average consumption amount and total number of purchases during the time period, and store these summary information together with the original profile data to facilitate subsequent trend analysis. Further subdivide the profile data according to key attributes. Create different storage structures (such as database tables or file directories) to store data of different attribute categories. For example, establish a table or directory dedicated to storing customer statistical attributes, and store statistical information such as age, gender, and region in it; set up another area to store customer behavior attributes, store browsing records, search history, purchase behavior sequence and other data; and create a section for storing value attributes to save information such as customer lifetime value and loyalty score.

[0130] Secondly, store the divided portrait data in the corresponding storage medium, such as a distributed file system (HDFS), a relational database (such as MySQL, PostgreSQL), or a NoSQL database (such as MongoDB). During the storage process, create appropriate indexes based on the characteristics of the data and query requirements. For example, when storing customer behavior data, if you often need to query behavior records based on customer ID and time, you can create a composite index based on customer ID and time to improve query efficiency.

[0131] The above method can effectively divide the portrait data into granularities according to the time dimension and attribute dimension, and store it in the appropriate location, providing a good data foundation for subsequent data analysis and customer management decisions.

[0132] like Figure 2 As shown, step S3 divides the customer service management task into multiple customer service management subtasks according to the decision tree algorithm, specifically including:

[0133] S31. Determine the characteristics and target variables of the decision tree. Collect a large amount of historical data related to customer service management tasks, including basic customer information (age, gender, region, etc.), consumer behavior (purchase frequency, consumption amount, type of product purchased, etc.), service interaction records (number of complaints, types of consulting questions, resolution time, etc.), and business results data (customer retention rate, repurchase rate, satisfaction score, etc.). Analyze this data and select characteristic variables that have an important impact on customer service management decisions. For example, use customer consumption amount, purchase frequency, and the time of the last purchase as key features, and determine the target variables to be predicted or classified, such as whether the customer has churned (binary classification) or the customer's value level (multi-classification).

[0134] S32. Build a decision tree model and select a suitable decision tree algorithm, such as C4.5, CART, etc. Taking the CART algorithm as an example, it selects the best feature splitting point based on the Gini Index or Information Gain. Use training data to train the decision tree model. During the training process, the decision tree will continuously divide branches according to the selected features to form a tree structure. For example, if the customer's consumption amount and purchase frequency are used as features, when the consumption amount is greater than a certain threshold and the purchase frequency is higher than a certain number of times, the customer is divided into a high-value customer branch; otherwise, it is divided into a low-value customer branch.

[0135] To prevent the decision tree from overfitting, it is necessary to prune the constructed decision tree. Pre-pruning or post-pruning methods can be used. Pre-pruning is to stop the growth of the decision tree in advance during the construction process. For example, when the number of samples of a node is too small or the information gain is less than the set threshold, the split will not continue. Post-pruning is to prune the tree according to certain evaluation indicators (such as accuracy and complexity on the validation set) after the decision tree is built. For example, by calculating the sum of the error rate of the subtree and the error rate of the leaf node, if the error rate of the subtree is not significantly reduced and the model complexity is increased, the subtree can be replaced with a leaf node.

[0136] S33. Determine different customer service management subtasks based on the final structure and branch results of the decision tree. For example, the path from the root node to the leaf node can be defined as different customer categories or service scenarios. For example, for high-value customer branches, it can be defined as customer maintenance and value-added service subtasks, including providing exclusive discounts, personalized recommendations, etc.; for low-value customer branches, it can be divided into customer activation and promotion subtasks, such as conducting targeted marketing activities and providing basic service optimization suggestions.

[0137] Through the above steps, the decision tree algorithm can be used to reasonably divide the customer service management task into multiple targeted subtasks, thereby improving the efficiency and effectiveness of customer service management and better meeting the needs of different customer groups. In actual applications, each step can be adjusted and optimized according to specific business conditions and data characteristics.

[0138] The above steps are used to divide multiple customer service management subtasks, and then the subtasks are matched with the different nodes according to attribute associations, specifically:

[0139] S34. Analyze the subtask attribute requirements. Match the key attributes that each customer service management subtask depends on. For example, for the customer churn warning subtask, you may need attributes such as the customer's recent consumption behavior, complaint frequency, and product usage time; for the high-value customer maintenance subtask, you may need attributes such as the customer's consumption amount, purchase frequency, and preferred product category. Through in-depth understanding of business logic and historical data analysis, clarify the degree of dependence and correlation of each subtask on different attributes.

[0140] S35. Identify the attribute distribution of data storage nodes. Check each node in the distributed storage system to understand the attribute characteristics of the data stored in each node. This information can be obtained through metadata management tools or data directories. For example, determine that a certain node mainly stores customer transaction records and consumption amount data, and another node stores customer behavior logs and product usage information. Classify and mark the data content of each node for subsequent matching.

[0141] S36. Establish attribute association mapping rules. Create mapping rules between the attribute requirements of the subtask and the attribute distribution of the data storage node. This can be achieved by building an association matrix or rule table. For example, if a subtask requires a customer's consumption amount and purchase frequency data, and these data are stored in node A and node B, then the subtask is associated with node A and node B. The mapping rules should be based on the semantics and business logic of the data to ensure accuracy and rationality.

[0142] S37, storing the attribute association mapping rules established above in a management database. The management database should have efficient query and retrieval functions. When it is necessary to match subtasks with nodes, query the database for storage node information that matches the subtask attribute requirements. For example, search for node records that meet specific attribute conditions in the database through SQL query statements to obtain corresponding node identifiers or address information.

[0143] S38. Dynamically adjust and optimize matching results. As the business develops and data changes, regularly evaluate the matching effect between subtasks and nodes. Monitor the data access efficiency and accuracy during the execution of subtasks. If it is found that some subtasks have data acquisition delays or incomplete data during execution, re-analyze the attribute association and storage node distribution, and dynamically adjust the matching results. For example, when a new data storage node is added to the system and contains important attributes related to a subtask, update the matching rules in time and associate the subtask with the new node to improve the overall performance of the system.

[0144] After completing the matching of subtasks and nodes, select some sample data for testing. Simulate the execution process of the subtask to check whether the required data can be obtained from the associated nodes accurately and quickly. Through actual data access and processing, verify whether the matching results meet the business requirements of the subtask. If problems are found, promptly trace back to the previous steps, find the causes and make corrections to ensure the effectiveness and reliability of the matching.

[0145] Through the above steps, customer service management subtasks can be systematically matched with different nodes according to attribute associations, achieving efficient use of data and smooth execution of tasks.

[0146] like Figure 3 As shown, step S4 uses a load balancing algorithm to dynamically schedule the matching results, specifically:

[0147] S41. Resource evaluation and monitoring: Real-time monitoring of the resource status of computing nodes, including key indicators such as CPU usage, memory usage, and network bandwidth. By deploying a resource monitoring agent on each computing node, resource usage data is regularly collected and fed back to the load balancing controller. For example, data is collected every 5 seconds to ensure that the load dynamics of the node can be grasped in a timely manner. At the same time, the I / O performance and storage capacity of the storage node are evaluated and monitored to understand the speed of data reading and writing and the remaining storage space, so as to reasonably allocate tasks and data storage locations.

[0148] S42. Task priority determination. Prioritize customer service management subtasks according to their urgency, importance, and business impact. For example, a customer churn warning task may have a higher priority because timely measures can recover potential losses; while some routine customer behavior analysis tasks have a relatively low priority. A priority weight can be set for each subtask, for example, from 1 to 10, with a higher weight indicating a higher priority, so that the load balancing algorithm gives priority to high-priority tasks when scheduling.

[0149] S43, load balancing algorithm selection and configuration. Select the appropriate load balancing algorithm according to the characteristics and requirements of the system, such as Round Robin, Weighted Round Robin, Least Connections, Weighted Least Connections, etc. For the Round Robin algorithm, tasks are assigned to each computing node in order to ensure that each node can get a relatively balanced task distribution; weighted round robin assigns different weights to nodes according to the performance or task processing capability of the node. Nodes with strong performance are assigned higher weights and receive more tasks; the Least Connections algorithm assigns tasks to the node with the least current connections to make full use of node resources; the Weighted Least Connections algorithm assigns tasks based on this and node weights. Configure the parameters of the selected algorithm according to the actual situation, such as setting weight values, connection thresholds, etc.

[0150] S44, dynamic scheduling execution. When there is a new matching result (i.e., a combination of subtasks and data) that needs to be scheduled, the load balancing controller makes a decision based on resource monitoring information, task priority, and the selected load balancing algorithm. For example, if a weighted round-robin algorithm is used, high-priority tasks are first judged and assigned to computing nodes with higher weights and relatively lower loads; for low-priority tasks, they are assigned to appropriate nodes in the order of weighted round-robin. During the allocation process, the node load information and task allocation records are updated in real time. If a computing node fails or is overloaded (exceeding a preset threshold, such as CPU usage reaching 80%), the new task will be automatically scheduled to other available and lightly loaded nodes to ensure the stability of the system and the smooth execution of tasks.

[0151] After the task is assigned to the computing node, the execution of the task is continuously tracked, and feedback information such as the execution time and resource consumption of the task is collected. Based on this feedback information, the effect of the load balancing algorithm is evaluated. If it is found that some nodes are in a high-load state for a long time or the task execution efficiency is low, the reasons are analyzed and the parameters or strategies of the load balancing algorithm are adjusted. For example, if the task execution time of a node is too long, it may be that the resource allocation is unreasonable. At this time, the weight of the node can be appropriately reduced or the task allocation strategy can be adjusted to optimize the system performance.

[0152] Through the above steps, the load balancing algorithm can be effectively used to dynamically schedule the matching results, fully utilize system resources, and improve the processing efficiency of customer service management tasks and the overall performance of the system.

[0153] In this embodiment, step S4 takes the divided data as input and the customer service management decision strategy as task output to construct and train a neural network model corresponding to the subtask, specifically:

[0154] First, divide the preprocessed data into training set, validation set and test set according to a certain ratio. Usually, the training set accounts for 60%-80%, the validation set accounts for 10%-20%, and the test set accounts for 10%-20%. For example, 70% of the data is used for training, 15% of the data is used for validation, and 15% of the data is used for testing.

[0155] Then, select the neural network model according to the subtask type: if the subtask is customer churn prediction, you can choose logistic regression, multi-layer perceptron (MLP) or more complex recurrent neural network (RNN) and its variants such as long short-term memory network (LSTM), etc. If it is customer service quality score prediction, you can consider using a fully connected neural network or a convolutional neural network (CNN), especially when the data contains features such as images or text, CNN can better extract features.

[0156] For simple subtasks, such as predicting whether a customer is a new customer based on their basic information, a simple MLP consisting of an input layer, one or two hidden layers, and an output layer can be constructed. The number of input layer nodes is determined by the number of input data features. For example, if there are 10 customer features, the input layer has 10 nodes. The number of hidden layer nodes can be determined by empirical formulas or by trying different values, such as 32, 64, etc. The number of output layer nodes is determined according to the task output type. If it is a binary classification problem (such as whether the customer has churned), the output layer has 1 node. If it is a multi-classification problem, the number of output layer nodes is the number of categories.

[0157] Use random initialization to assign values ​​to the weights of the neural network and perform bias initialization: Generally, the bias is initialized to 0 or a small constant, such as 0.01. This helps the model converge quickly in the early stages of training. Then choose a suitable loss function based on the nature of the subtask. For regression tasks, such as customer service response time prediction, the mean square error (MSE) loss function is commonly used. For classification tasks, such as customer complaint type classification, the cross entropy loss function is a common choice, such as the binary cross entropy loss function. Finally, determine the optimizer that optimizes the model parameters, such as stochastic gradient descent (SGD), Adagrad, Adadelta, RMSProp, Adam, etc.

[0158] After the model is initialized, the training set data is used to train the neural network model. During the training process, the input data is passed through each layer of the neural network in turn, the loss between the predicted value and the true value is calculated, and then the gradient is calculated through the back propagation algorithm to update the model parameters. For example, in each training batch, 128 samples are input into the model, the loss is calculated and the parameters are updated.

[0159] After each round of training, the performance of the model is evaluated using the validation set data. By calculating indicators such as the loss value, accuracy (for classification tasks), and mean square error (for regression tasks) on the validation set, we can observe whether the model is overfitting or underfitting. If the loss value on the validation set gradually increases during the training process, while the loss value on the training set continues to decrease, overfitting may occur. At this time, we can consider taking regularization measures, such as L1 or L2 regularization, to prevent the model from overfitting.

[0160] After the model training is completed, the test set data is used to perform a final evaluation of the model. Various performance indicators on the test set are calculated, such as accuracy, recall, F1 value (for classification tasks), root mean square error (RMSE) (for regression tasks), etc., to comprehensively measure the performance of the model.

[0161] like Figure 4 As shown, step S5 dynamically fuses the neural network models of the corresponding subtasks to obtain a multimodal customer service management decision model, which is specifically:

[0162] S51. Determine the models involved in the fusion. Select the most suitable models for the fusion based on the relevance and importance of the tasks and their performance in different scenarios. For example, for the customer churn prediction task, there may be an LSTM model trained based on customer behavior data and a multi-layer perceptron model trained based on customer value data. By comparing their accuracy, recall and other indicators on the validation set, determine whether to include both in the fusion system. Ensure that these models can provide valuable information for customer service management from different angles or dimensions, and avoid selecting models with highly overlapping functions.

[0163] S52. Use a unified test data set to evaluate each selected subtask neural network model and calculate a series of performance indicators. For classification tasks, such as customer complaint type judgment, calculate accuracy, recall, F1 value, etc.; for regression tasks, such as customer service response time prediction, calculate mean square error (MSE), mean absolute error (MAE), etc. Record the performance of each model on different sample subsets and analyze the stability and generalization ability of the model. For example, observe the performance fluctuations of the model in different time periods and on different customer groups to understand its ability to adapt to different scenarios.

[0164] S53. Assign a weight to each model according to the performance evaluation result of the model. The fusion methods include weighted fusion, model stacking fusion, and voting fusion, where:

[0165] Weighted fusion gives higher weights to models with better performance. For example, in the customer satisfaction prediction task, if the accuracy of model A is 80% and the accuracy of model B is 70%, then a weight of 0.6 can be assigned to model A and a weight of 0.4 to model B. When predicting, the output results of the two models are weighted and summed according to their respective weights to obtain the final prediction value.

[0166] Model stacking and fusion: The output of one or more models is used as the input of another model. For example, a neural network model based on the basic information of the customer is used to output the preliminary characteristics of the customer, and then these characteristics are input into a neural network model based on the customer's transaction data. Finally, the second model outputs the fused decision result.

[0167] Voting fusion: For classification tasks, each model predicts the classification of samples, and then the final classification result is determined by majority voting. For example, there are three models predicting whether a customer will buy a new product. Model A predicts "will buy", model B predicts "will not buy", and model C predicts "will buy", then the final decision result is "will buy".

[0168] S54. Dynamically adjust fusion weights or parameters. As new data is continuously generated and business scenarios change, model performance should be re-evaluated regularly. For example, a performance evaluation is performed on all models involved in the fusion every month. Based on the new evaluation results, the weights or parameters in the fusion strategy are dynamically adjusted. If it is found that the performance of a certain model on recent data has improved significantly, its weight in the weighted average fusion is increased accordingly; or in model stacking fusion, the input feature combination of the second layer model is adjusted. An adaptive algorithm, such as a gradient descent-based method, can be used to automatically adjust the fusion weights according to the prediction error of the model on new data to ensure that the fusion model always maintains good performance.

[0169] In this embodiment, the step S6 uses a distributed search algorithm to search for multimodal data related to the customer service management subtask, and iteratively trains the multimodal customer service management decision model, specifically:

[0170] First, analyze the requirements of each customer service management subtask and determine the type and characteristics of the required multimodal data. For example, for the customer churn prediction subtask, it is necessary to search for customer behavior data (such as purchase frequency, browsing history), consumption data (amount, consumption cycle), and feedback data (complaints, evaluations), etc. Then define the scope of distributed search to cover various data storage nodes within the enterprise, including databases, file systems, and external data interfaces. Clarify which nodes store data related to customer service management, such as the customer relationship management system (CRM) node storing basic customer information and communication records, and the transaction system node storing consumption data.

[0171] According to the distribution characteristics of data and search requirements, match the appropriate distributed search algorithm. For example, if the data is stored in a distributed hash table (DHT) in the form of key-value pairs, a DHT-based search algorithm, such as Chord, CAN, etc., can be used to take advantage of its efficient node positioning and data search mechanism. For searching large-scale text data, a distributed search algorithm based on inverted index can be used, such as the algorithm used by Elasticsearch. It indexes data by keywords and quickly locates data containing target keywords through distributed computing.

[0172] Next, the distributed search task is initialized and the search request is distributed to each data storage node. Each node uses the selected search algorithm to retrieve data based on the local data situation. For example, in a DHT-based search, the node quickly locates and searches based on the hash value of the key-value pair. Each node will preliminarily screen and process the searched relevant data and only return data fragments that are closely related to the subtask. For example, when searching for customer behavior data, only key behavior records in the recent period are returned to reduce the amount of data transmission. The data returned by each node is aggregated to form a multimodal data set related to the customer service management subtask.

[0173] Finally, the newly searched multimodal data is integrated with the original training data set. According to the characteristics of the data and the input requirements of the model, the data is converted and features are extracted. For example, the sentiment features of text-based customer evaluation data are extracted through natural language processing technology, and the image features of image data are extracted through convolutional neural networks. Ensure that the new data matches the input dimension and data type of the model. If the model input requires a fixed-length vector, the new data is encoded and padded accordingly.

[0174] The multimodal customer service management decision model is iteratively trained using the integrated data set until the model's performance on the validation set reaches a stable state or meets the preset performance indicator requirements, completing the update of the multimodal customer service management decision model.

[0175] like Figure 5 As shown, step S7 processes the real-time data stream through streaming computing, specifically:

[0176] S71. Deploy data collection tools to obtain multimodal data from various data sources in real time. For example, use Fluentd to collect customer operation logs from the company's log files, including login time, page browsing and other behavioral data; use Kafka Connect to obtain customer consumption data updates from the database change log. Configure corresponding collection strategies and format parsing rules for different types of data sources. For sensor data from IoT devices (if it involves customer service scenarios, such as the response time of intelligent customer service equipment, etc.), real-time parsing is required according to the data protocol of the device. Send the collected data to a streaming computing platform, such as an Apache Kafka message queue, for subsequent processing.

[0177] S72. Perform feature extraction on the real-time data stream according to the requirements of the customer service management subtask. For text-based customer evaluation data, use a natural language processing library (such as NLTK or SpaCy) to perform lexical analysis and syntactic analysis to extract features such as keywords and sentiment tendencies. For image data (if there is a real-time image stream, such as product problem pictures uploaded by customers), extract key features of the image, such as object recognition results, abnormal features in the image, etc., through a pre-trained convolutional neural network model. Convert the extracted features into a format suitable for model input, such as converting text features into word vectors and converting image features into fixed-length feature vectors.

[0178] S73. When new feature data flows in, integrate it with the original model training data. In a streaming computing environment, update the multimodal customer service management decision model in real time. Select an appropriate online learning algorithm based on the type of model. For example, for a neural network model based on gradient descent, use the online version of stochastic gradient descent (SGD) to calculate the gradient and update the model weights each time a new batch of data is received. Monitor the performance indicators of the model in real time, such as accuracy, recall (for classification tasks) or mean square error (for regression tasks). By setting a sliding window, evaluate the model's predictive effect on new data within a certain time range. If the performance indicators show a downward trend, adjust the model's training parameters in time, such as reducing the learning rate or increasing the regularization strength.

[0179] S74. Feedback the results of the real-time prediction of the model to the customer service management system. For example, for the customer churn prediction model, once a customer is predicted to have a high risk of churn, the early warning mechanism is immediately triggered to notify the customer service staff to intervene. According to the results of real-time processing, the customer service strategy is dynamically adjusted. If the complaint rate of a certain customer group is found to increase, the service process for this group is adjusted in a timely manner or additional service support is provided. At the same time, the model update status and decision results are recorded for subsequent data analysis and service optimization.

[0180] In summary, the present invention realizes efficient processing and storage of massive multimodal customer information through a multi-level distributed architecture. The multimodal data acquisition layer can widely collect data from a variety of data sources, and with the help of efficient acquisition tools and message queues, ensure that the data is transmitted to the preprocessing layer in real time, effectively avoiding data congestion and loss. The data preprocessing layer uses advanced algorithms to clean, unify the format and extract preliminary features of the data to improve data quality and availability. In terms of storage, the reasonable combination of distributed file systems, relational databases and NoSQL databases can accurately store data according to data types and characteristics, give full play to the advantages of each storage medium, greatly improve data storage and retrieval efficiency, and break through the capacity and performance bottlenecks of traditional centralized architectures.

[0181] Secondly, user portraits are built based on statistical, behavioral and value attributes, and through clustering algorithms and dynamic update mechanisms, customer characteristics and behavioral changes can be accurately captured. This enables companies to gain a deep understanding of customer needs, preferences and values, and provide customers with highly personalized product recommendations, service solutions and marketing activities. For example, companies can accurately push product information that meets the needs of different customer groups based on their consumption habits and interests, thereby increasing customers' willingness to buy and satisfaction and enhancing customer stickiness with the company.

[0182] Thirdly, the customer service management task is reasonably divided into multiple subtasks through the decision tree algorithm, and the task is accurately matched with the storage node through attribute association, so as to achieve efficient docking between the task and the data. The load balancing algorithm dynamically schedules the resource status of the computing node and the storage node and the task priority to ensure that the system resources are fully and reasonably utilized and avoid idle or overloaded resources. This not only improves the task processing efficiency, but also reduces the system operation cost, enabling enterprises to handle more customer service tasks under limited resource conditions and improve overall operational efficiency.

[0183] The neural network model built and trained for subtasks, combined with dynamic fusion strategies, can make full use of the rich information in multimodal data to improve the accuracy and generalization ability of the model. The application of distributed search algorithms and streaming computing enables the model to obtain the latest data in a timely manner and perform iterative training, and update the decision-making strategy in real time to adapt to the dynamic changes of the market and customers. Based on more accurate and timely decision-making models, enterprises can predict customer behavior and market trends in advance, such as accurately predicting customer churn risks and grasping changes in market demand, so as to adjust business strategies in a timely manner, seize the initiative in the fierce market competition, and enhance the competitiveness and profitability of enterprises.

[0184] In summary, the present invention provides enterprises with a comprehensive, efficient and intelligent customer information management and analysis solution, which effectively solves many problems of traditional methods in the big data environment and plays an important role in promoting the sustainable development of enterprises.

[0185] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A distributed big data customer information management and analysis method, characterized in that: The following steps are involved: Build a multi-layer distributed architecture, including multimodal data collection layer, data preprocessing layer, data storage layer, computing layer and result display layer; Create user portraits based on multimodal customer information, statistical, behavioral, and value attributes, divide the portrait data into granularities based on time and attribute dimensions, and store them on different nodes. Dividing the customer service management task into a plurality of customer service management subtasks according to a decision tree algorithm, and matching the subtasks with the different nodes according to attribute association; Use load balancing algorithms to dynamically schedule matching results, use the divided data as input, and customer service management decision strategies as task outputs to build and train neural network models for corresponding subtasks; Dynamically integrate the neural network models of the corresponding subtasks to obtain a multimodal customer service management decision model; Using a distributed search algorithm to search for multimodal data related to the customer service management subtask, and iteratively training the multimodal customer service management decision model; Real-time data streams are processed through streaming computing, and the processed data stream information is input into the iteratively trained multimodal customer service management decision model to make customer service management decisions.

2. The distributed big data customer information management and analysis method according to claim 1 is characterized in that: The multimodal customer information is used to create user profiles according to the attributes of statistics, behavior, and value, specifically: Multimodal customer information is classified according to statistics, behavior, and value; statistics include age, gender, region, consumption preference, and consumption frequency; behaviors include browsing behavior, search behavior, interaction behavior, and purchase behavior; and values ​​include customer loyalty, consumption level, and consumption frequency; Use clustering algorithms to cluster users and build user portraits; Establish a user portrait update mechanism based on a time window to update the portrait in a timely manner according to the user's latest behavior and information.

3. The distributed big data customer information management and analysis method according to claim 1 is characterized in that: The portrait data is divided into granularities based on the time dimension and the attribute dimension, specifically: Analyze business needs and data characteristics, determine the appropriate time dimension division period, which includes day, week, month, quarter, and year; at the same time, sort out the key attributes for profiling, including statistical attributes, behavioral attributes, and value attributes; Group the profile data according to the determined time period, calculate summary statistics for each time period group, and store the summary information together with the original profile data; Creating different storage structures to store data of different attribute categories respectively, and storing the divided portrait data in corresponding storage media, wherein the storage media includes a distributed file system, a relational database or a NoSQL database; During the storage process, indexes are created based on the characteristics of the data and query requirements.

4. The distributed big data customer information management and analysis method according to claim 1 is characterized in that: The customer service management task is divided into a plurality of customer service management subtasks according to the decision tree algorithm, specifically: Determine the characteristics and target variables of the decision tree, collect a large amount of historical data related to customer service management tasks, and analyze and select characteristic variables that are relevant to customer service management decisions; Select a decision tree algorithm, use training data to train a decision tree model, and prune the constructed decision tree; According to the final structure and branch results of the decision tree, different customer service management subtasks are determined.

5. The distributed big data customer information management and analysis method according to claim 1 is characterized in that: The matching of the subtasks with the different nodes according to the attribute association is specifically as follows: Analyze the subtask attribute requirements, match the key attributes that each customer service management subtask depends on, and analyze the degree of dependence and correlation of each subtask on different attributes; Identify the attribute distribution of data storage nodes, obtain the attribute characteristics of the stored data through metadata management tools or data directories, and classify and mark the data content of each node; Establish attribute association mapping rules. According to the attribute requirements of the subtask and the attribute distribution of the data storage node, create the mapping rules between the two by building an association matrix. The mapping rules should be based on the semantics and business logic of the data. The attribute association mapping rules are stored in a management database, and the management database uses a query and retrieval algorithm to query storage node information that matches the subtask attribute requirements; Regularly evaluate the matching effect of subtasks and nodes. If some subtasks have data acquisition delays or incomplete data during execution, re-analyze the attribute association and storage node distribution, and dynamically adjust the matching results.

6. The distributed big data customer information management and analysis method according to claim 1 is characterized in that: The dynamic scheduling matching result using the load balancing algorithm is specifically: Real-time monitoring of the resource status of computing nodes, including at least key indicators such as CPU usage, memory usage, and network bandwidth. By deploying a resource monitoring agent on each computing node, resource usage data is regularly collected and fed back to the load balancing controller. At the same time, the I / O performance and storage capacity of storage nodes are evaluated and monitored. Prioritize customer service management subtasks based on their urgency, importance, and business impact; Select and configure the load balancing algorithm. When there are new matching results that need to be scheduled, the load balancing controller makes decisions based on resource monitoring information, task priority, and the selected load balancing algorithm. After the task is assigned to the computing node, the execution status of the task is continuously tracked, feedback information on the execution of the task is collected, the effect of the load balancing algorithm is evaluated, and the parameters or strategies of the load balancing algorithm are adjusted according to the evaluation results.

7. The distributed big data customer information management and analysis method according to claim 1 is characterized in that: The divided data is used as input, the customer service management decision strategy is used as task output, and the neural network model corresponding to the subtask is constructed and trained, specifically: Divide the divided data into training set, validation set and test set according to a certain ratio; Select different neural network models according to the subtask type, and set the number of network layers and nodes; Use random initialization methods to assign weights and biases to neural networks, define loss functions and optimizers Use the training set data to train the neural network model, pass the input data through each layer of the neural network in turn, calculate the loss between the predicted value and the true value, and then calculate the gradient through the back propagation algorithm to update the model parameters; After each round of training, the performance of the model is evaluated using the validation set data. The loss value, accuracy, or mean square error indicator on the validation set is calculated to observe whether the model is overfitting or underfitting. After the model training is completed, the test set data is used to perform a final evaluation of the model, and the model is tuned based on the evaluation results.

8. The distributed big data customer information management and analysis method according to claim 1, characterized in that: The neural network models of the corresponding subtasks are dynamically integrated to obtain a multimodal customer service management decision model, which is specifically: Select the model involved in the fusion based on the correlation between the subtasks; Use a unified test dataset to evaluate each selected subtask neural network model, and record the stability and generalization ability of each model on different sample subsets; Select different fusion strategies according to different task attributes, the fusion strategy being weighted average fusion, model stacking fusion or voting fusion; Regularly re-evaluate model performance, and based on the new evaluation results, use an adaptive algorithm to automatically adjust the fusion weights according to the model's prediction error on new data, and dynamically adjust the weights or parameters in the fusion strategy.

9. The distributed big data customer information management and analysis method according to claim 1, characterized in that: The method of using a distributed search algorithm to search for multimodal data related to the customer service management subtask and iteratively training the multimodal customer service management decision model is specifically as follows: Analyze the requirements of each customer service management subtask, define the scope of distributed search, and determine the type and characteristics of the required multimodal data; the scope of the distributed search includes databases, file systems, and external data interfaces; Select the appropriate distributed search algorithm based on the data distribution characteristics and search requirements; Initialize the distributed search task, distribute the search request to each data storage node, use the selected search algorithm to retrieve data, each node will filter and process the searched relevant data, and return the data fragments related to the subtask; summarize the data returned by each node to form a multimodal data set related to the customer service management subtask; Integrate the newly searched multimodal data with the original training data set, convert the data format and extract features according to the characteristics of the data and the input requirements of the model, so that the new data matches the input dimension and data type of the model; The multimodal customer service management decision model is iteratively trained using the integrated data set until the model's performance on the validation set reaches a stable state or meets the preset performance indicator requirements, completing the update of the multimodal customer service management decision model.

10. The distributed big data customer information management and analysis method according to claim 1, characterized in that: The real-time data stream is processed by stream computing, specifically: Acquire multimodal data from various data sources in real time, and configure corresponding collection strategies and format parsing rules for different types of data sources; According to the needs of the customer service management subtask, feature extraction is performed on the real-time data stream. For text-based customer evaluation data, the natural language processing library is used to perform lexical analysis and syntactic analysis to extract features that are not limited to keywords and sentiment tendencies. For image data, the pre-trained convolutional neural network model is used to extract the key features of the image and convert the extracted features into a format suitable for model input. When new feature data flows in, it is integrated with the original model training data to monitor the performance indicators of the multimodal customer service management decision model in real time; The real-time prediction results of the multimodal customer service management decision model are fed back to the customer service management system for real-time processing, and the customer service strategy is dynamically adjusted based on the results of the real-time processing.

Citation Information

Cited By

  • Trusted industrial agent task distribution method and system

    CN120893794A