Big data mining analysis method

Through big data mining and analysis methods, in-depth analysis of customers' multi-source heterogeneous data is generated to generate potential needs of customers, solving the problem that existing customer management systems are difficult to deeply analyze customer data, and achieving the improvement of personalized shopping experience and customer satisfaction.

CN120179710APending Publication Date: 2025-06-20SUZHOU HUAYUAN CENTURY TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510248611.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

It is difficult for existing customer management systems to conduct in-depth analysis of customer data, making it difficult for enterprises to fully understand customer preferences, behavior patterns and unmet needs.

Method used

The big data mining and analysis method is adopted to collect multi-source heterogeneous data (such as transaction records, social media activities, market research feedback), perform data cleaning and preprocessing, use natural language processing technology to extract key elements, and perform cluster analysis through synonyms matching lexicon database and consumption scenario lexicon database to generate potential needs of customers.

Benefits of technology

It has achieved a deep understanding of customer behavior and preferences, helping enterprises accurately position target markets, optimize products and services, and create personalized shopping experiences, thereby improving customer satisfaction and loyalty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179710A_ABST
    Figure CN120179710A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital data processing, in particular to a big data mining analysis method, which comprises the following steps: collecting multi-source heterogeneous data of customers, the multi-source heterogeneous data comprising transaction records, social media activities and market investigation feedback; cleaning and preprocessing the multi-source heterogeneous data to obtain initial text data; extracting key elements in the initial text data by using a natural language processing technology, wherein the key elements comprise a commodity name, a commodity price, a commodity quantity and purchase time; performing clustering analysis on the key elements by adopting a synonym matching lexicon to obtain a client clustering result; performing scene analysis on key elements in the clustering result of each client by adopting a consumption scene word library to obtain a scene clustering result; and generating a potential demand of the customer based on a clustering analysis result and a scene analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital data processing, and in particular to a big data mining and analysis method. Background Art

[0002] Customer management in business operations refers to understanding and meeting customer needs and expectations through a systematic approach, thereby improving customer satisfaction and loyalty, and ultimately achieving the business goals of the company. It covers the entire process from market research, customer service, sales to after-sales support. Effective customer management not only focuses on acquiring new customers, but also pays attention to maintaining and developing existing customer relationships, using tools such as CRM (customer relationship management) systems to analyze customer data, identify potential business opportunities, optimize product and service offerings, and personalize marketing strategies to ensure that companies maintain their advantages in a highly competitive market environment.

[0003] In the current practice of customer management in enterprise operations, a significant challenge is to effectively analyze customer data to tap into customers' potential needs. Traditional or existing customer management systems often focus on recording basic customer information and transaction history, but lack the ability to conduct in-depth data analysis, which makes it difficult for enterprises to fully understand customer preferences, behavior patterns, and unmet needs. Summary of the invention

[0004] The purpose of the present invention is to provide a big data mining and analysis method, which aims to create a more personalized shopping experience for customers, thereby improving overall satisfaction and loyalty.

[0005] To achieve the above-mentioned object, the present invention provides a big data mining and analysis method, including collecting multi-source heterogeneous data of customers, the multi-source heterogeneous data including transaction records, social media activities, and market research feedback;

[0006] Cleaning and preprocessing the multi-source heterogeneous data to obtain initial text data;

[0007] Extract key elements from the initial text data using natural language processing technology, wherein the key elements include product name, product price, product quantity, and purchase time;

[0008] Use synonym matching word library to cluster key elements and obtain customer clustering results;

[0009] Use the consumption scenario word library to perform scenario analysis on the key elements in each customer clustering result to obtain the scenario clustering result;

[0010] Generate customer potential needs based on cluster analysis results and scenario analysis results.

[0011] The specific steps of collecting multi-source heterogeneous data of customers include:

[0012] Extract the sales information of customers through the enterprise's ERP and CRM systems;

[0013] Extract the text communication content of corresponding customers from chat tools;

[0014] Design and distribute electronic questionnaires to corresponding customers to collect their opinions on products;

[0015] Structure the sales information, text communication content, and opinions on products based on customer information to obtain multi-source heterogeneous data.

[0016] Among them, the specific steps of cleaning and preprocessing the multi-source heterogeneous data to obtain the initial text data include:

[0017] Check and delete the records that appear repeatedly in the multi-source heterogeneous data;

[0018] Delete the missing data points in the multi-source heterogeneous data;

[0019] Apply the stemming algorithm to normalize different forms of words into basic forms, reducing the number of lexical variants to obtain the initial text data.

[0020] Among them, the specific steps of applying the stemming algorithm to normalize different forms of words into basic forms, reducing the number of lexical variants to obtain the initial text data include:

[0021] Split the multi-source heterogeneous data into separate word data groups;

[0022] Select the Porter Stemmer algorithm to extract the stems in the word data group to obtain the stem data group;

[0023] Remove stop words after stemming to obtain the initial text data.

[0024] Among them, the specific steps of using natural language processing technology to extract key elements from the initial text data include:

[0025] Identify product names based on product catalog matching;

[0026] Match product prices and quantities based on regular expressions;

[0027] Identify the purchase time based on the date parsing library.

[0028] Among them, the specific steps of using a synonym matching thesaurus to perform clustering analysis on key elements to obtain customer clustering results include:

[0029] Create a dictionary containing key element words and their synonyms based on business requirements;

[0030] Perform a synonym expansion on the extracted key elements based on a dictionary;

[0031] Convert the text after the synonym expansion into a first numerical feature vector;

[0032] Use the hierarchical clustering algorithm to perform clustering analysis on the first numerical feature vector to obtain the customer clustering result.

[0033] Among them, the specific steps of using the hierarchical clustering algorithm to perform clustering analysis on the numerical feature vector to obtain the customer clustering result include:

[0034] According to the selected distance metric, calculate the distance between every two key elements and construct a distance matrix;

[0035] Define the inter-cluster distance as the distance between the two closest points in two clusters, and gradually construct a dendrogram by recursively merging the most similar objects;

[0036] Based on the dendrogram, cut the tree structure to determine the final number of clusters.

[0037] Among them, the specific steps of using the consumption scenario thesaurus to perform scenario analysis on the key elements in each customer clustering result to obtain the scenario clustering result include:

[0038] Based on industry knowledge, market research, and historical data, construct a consumption scenario thesaurus containing various consumption scenario keywords;

[0039] Match the key elements in the customer clustering result with the keywords in the consumption scenario thesaurus to find the consumption scenarios;

[0040] Convert the matched consumption scenarios into a second numerical feature vector to obtain a second feature dataset;

[0041] Select the K-Means clustering algorithm to perform clustering analysis on the second feature dataset to obtain the scenario clustering result.

[0042] Among them, the specific steps of selecting the K-Means clustering algorithm to perform clustering analysis on the numerical feature vector include:

[0043] Determine the optimal number of clusters through the elbow method;

[0044] Randomly select K sample points in the second feature dataset as the initial centroids;

[0045] Calculate the distance from each sample to each centroid, and assign the sample to the cluster to which the nearest centroid belongs;

[0046] Recalculate the centroid position of each cluster, calculate the distance from each sample to each centroid, and assign the sample to the cluster to which the nearest centroid belongs until the maximum number of iterations is reached.

[0047] A big data mining and analysis method of the present invention collects multi-source heterogeneous data of customers from a wide range of sources, including but not limited to transaction records, social media activities, and market research feedback. These data sources are not only diverse in form but also rich in content, and can comprehensively reflect the preferences and behavior patterns of customers. Then, the data is cleaned and preprocessed. In the data cleaning stage, noise data is removed, incorrect data is corrected, and missing values are filled, etc.; while preprocessing may involve operations such as data conversion and standardization, aiming to convert the original data into an initial text data format that can be used for analysis. Subsequently, the initial text data is deeply parsed to extract key elements. These key elements cover information such as product names, product prices, product quantities, and purchase times. To further deepen the analysis, a synonym matching thesaurus is used to perform clustering analysis on the above-extracted key elements. This method can effectively identify customer groups with similar purchase behaviors or preferences, forming customer clustering results. Then, a more detailed scenario analysis is performed on the key elements in each customer cluster using consumption scenario vocabulary, thereby obtaining clustering results under different consumption scenarios. Finally, based on the results of clustering analysis and scenario analysis, a potential demand report on customers is generated. This report can not only help merchants accurately target the market, optimize products and services, but also create a more personalized shopping experience for customers, thereby enhancing overall satisfaction and loyalty. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 is a flowchart of a big data mining and analysis method of the present invention.

[0050] Figure 2 is a flowchart of collecting multi-source heterogeneous data of customers of the present invention.

[0051] Figure 3 is a flowchart of cleaning and preprocessing the multi-source heterogeneous data to obtain initial text data of the present invention.

[0052] Figure 4It is a flowchart of the initial text data obtained by normalizing different forms of words into basic forms using the stem extraction algorithm of the present invention, reducing the number of lexical variants.

[0053] Figure 5 It is a flowchart of the present invention for extracting key elements from the initial text data using natural language processing technology.

[0054] Figure 6 It is a flowchart of the present invention for performing cluster analysis on key elements using a thesaurus of near-synonyms to obtain customer clustering results.

[0055] Figure 7 It is a flowchart of the present invention for performing cluster analysis on numerical feature vectors using a hierarchical clustering algorithm to obtain customer clustering results.

[0056] Figure 8 It is a flowchart of the present invention for performing scenario analysis on key elements in each customer clustering result using a consumption scenario thesaurus to obtain scenario clustering results.

[0057] Figure 9 It is a flowchart of the present invention for selecting the K-Means clustering algorithm to perform cluster analysis on numerical feature vectors. Detailed implementation manners

[0058] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.

[0059] The first embodiment

[0060] Please refer to Figures 1 to 9 , the present invention provides a big data mining and analysis method, including:

[0061] S101 Collect multi-source heterogeneous data of customers, and the multi-source heterogeneous data includes transaction records, social media activities, and market research feedback;

[0062] The specific steps include:

[0063] S201 Extract the sales information of customers through the enterprise's ERP and CRM systems;

[0064] To obtain detailed and accurate sales information, relevant data will be extracted from the enterprise's ERP (Enterprise Resource Planning) system and CRM (Customer Relationship Management) system. The ERP system contains detailed information about various aspects of enterprise operations, such as inventory levels, supply chain status, etc., while the CRM system focuses on customer interaction history, service requests, purchase behavior, etc.

[0065] S202 extracts the text communication content of the corresponding customer from the chat tool;

[0066] Modern customer service increasingly relies on instant messaging tools for communication. Therefore, extracting the text communication content related to specific customers from these chat tools becomes one of the key steps in understanding customer needs and opinions. This not only covers direct customer service conversations but also includes any brand interactions through social platforms. S203 designs and distributes electronic questionnaires to the corresponding customers to collect their opinions on the product;

[0067] To directly obtain customer feedback, targeted electronic questionnaires will also be designed and distributed. These questions will focus on the customers' views on the product or service, aiming to gain in-depth understanding of their satisfaction level, functional requirements, and desired improvement directions.

[0068] S204 structures the sales information, text communication content, and opinions on the product based on customer information to obtain multi-source heterogeneous data.

[0069] Integrate all the above-mentioned data sources and structure them based on the customer's identity information. This helps improve the efficiency of data analysis and also makes the collaboration between different departments smoother. In this way, a customer profile with rich dimensions can be created, providing strong support for enterprises to formulate precise marketing strategies and improve the user experience.

[0070] S102 cleans and preprocesses the multi-source heterogeneous data to obtain initial text data;

[0071] The specific steps include:

[0072] S301 checks and deletes the records that appear repeatedly in the multi-source heterogeneous data;

[0073] Check for duplicates in the dataset. Since the data comes from multiple channels (such as ERP, CRM systems, chat tools, and electronic questionnaires), there may be data duplication due to multiple interactions or synchronization issues. By identifying and deleting these duplicate records, we can ensure that each piece of information appears only once in the dataset, thus avoiding any potential biases caused by duplicate data.

[0074] S302 deletes the missing data points in the multi-source heterogeneous data;

[0075] Next, we need to handle the missing values in the dataset. Faced with missing data, we choose to directly delete the records that contain missing values. Although this method may reduce the amount of available data, it helps maintain the consistency and integrity of the dataset and prevents potential errors introduced by filling in missing values.

[0076] S303 applies a stemming algorithm to normalize different forms of words into their basic forms, reducing the number of lexical variants to obtain initial text data.

[0077] The specific steps include:

[0078] S401 splits the multi-source heterogeneous data into separate word data groups;

[0079] First, the original text data is split into individual words. This process is usually called tokenization, which converts a continuous text string into a series of discrete word units for subsequent processing.

[0080] S402 selects the Porter Stemmer algorithm to extract the stems from the word data groups, obtaining stem data groups;

[0081] Then, the Porter Stemmer algorithm is applied to perform stemming on each word. The Porter Stemmer is a rule-based algorithm designed to remove the suffixes of words and reduce them to their basic forms or stems. For example, "running", "runs", and "ran" will all be simplified to "run". This can effectively reduce the size of the vocabulary while retaining the core meaning of the words.

[0082] S403 removes stop words after stemming to obtain the initial text data.

[0083] After completing stemming, it is also necessary to remove the so-called "stop words". Stop words refer to words that frequently appear in the text but contribute little to expressing semantics, such as "the", "is", "at", "which", etc. By filtering out these words, our text data can be made more refined, highlighting the truly valuable content for analysis and laying a good foundation for subsequent text analysis.

[0084] S103 uses natural language processing techniques to extract key elements from the initial text data, and the key elements include product name, product price, product quantity, and purchase time;

[0085] The specific steps include:

[0086] S501 identifies the product name based on product catalog matching;

[0087] To accurately identify product names in text data, we will utilize the enterprise's existing product catalog as a reference benchmark. By matching the words in the text with the product names in the catalog, we can effectively identify the mentioned products. The specific approach is as follows: Organize and maintain a database containing all possible product names and their variants. Then apply fuzzy matching algorithms (such as Levenshtein distance or Jaccard similarity) to handle cases where product names do not exactly match due to spelling mistakes, abbreviations, or other reasons. Combine natural language processing techniques, especially named entity recognition (NER), to more precisely locate product names based on sentence structure and context information.

[0088] S502 matches product prices and quantities based on regular expressions;

[0089] Next, for the extraction of product prices and quantities, we choose to use regular expressions (Regex). This method is particularly suitable for extracting numerical information with relatively fixed patterns, such as currency amounts and quantity units. The following are the specific steps: Create a set of regular expression rules for matching different price formats (such as "$19.99", "19,99€") and quantity representations (such as "1 piece", "2 bottles"). Considering that price and quantity representations may vary in different regions, appropriate localization adjustments need to be made to the regular expressions. Before actual application, verify the matching accuracy and recall rate of the regular expressions through a test set to ensure that the target information can be effectively captured.

[0090] S503 identifies the purchase time based on a date parsing library.

[0091] When identifying the purchase time, a date parsing library will be used. These libraries can intelligently parse various date formats and convert them into a unified standard format. The operation details are as follows:

[0092] According to the different programming languages, select the corresponding date parsing library, such as dateutil.parser in Python or SimpleDateFormat in Java. Considering the possible time zone differences and cultural habits in global business, ensure that the date parsing process can flexibly handle various special situations.

[0093] S104 performs clustering analysis on key elements using a thesaurus of near synonyms to obtain the customer clustering results;

[0094] The specific steps include:

[0095] S601 creates a dictionary containing key element words and their synonyms based on business requirements;

[0096] First, according to the specific business scenarios and requirements, a detailed vocabulary or dictionary needs to be constructed. This vocabulary should not only include the key element words directly extracted from the data but also all their possible synonyms or related terms. For example, "purchase" can be extended to "procure", "shop", etc.; "mobile phone" can be extended to "smartphone", "mobile telephone", etc. The purpose of doing this is to ensure that all relevant text information can be comprehensively captured during the subsequent analysis process.

[0097] S602 performs a near-synonym expansion on the extracted key elements based on the dictionary.

[0098] Use the dictionary created in the previous step to perform a near-synonym expansion on the key elements extracted before. This means that for each key element, we not only need to consider its original expression form but also include all its synonyms in the analysis scope. This process helps to improve the accuracy and coverage of the analysis and avoid information loss or misjudgment caused by vocabulary differences.

[0099] S603 converts the text after near-synonym expansion into a first numerical feature vector.

[0100] After completing the near-synonym expansion, the text information needs to be converted into a numerical feature vector. This is achieved through some text representation methods, such as the Bag of Words model, TF-IDF (Term Frequency-Inverse Document Frequency), or Word2Vec, etc. Each method has its characteristics: the Bag of Words model is simple and direct but ignores the word order; TF-IDF can reflect the importance of words in the document; while Word2Vec can capture the semantic relationships between words. S604 uses a hierarchical clustering algorithm to perform a clustering analysis on the first numerical feature vector to obtain the customer clustering result.

[0101] The specific steps include:

[0102] S701 calculates the distance between every two key elements according to the selected distance metric and constructs a distance matrix.

[0103] Select a suitable method to measure the similarity or difference between different key elements, such as Euclidean distance, Manhattan distance, or cosine similarity, etc. Then, based on this distance metric, calculate the distance between every two key elements and construct a distance matrix reflecting these distance relationships.

[0104] S702 defines the inter-cluster distance as the distance between the two closest points in two clusters and gradually constructs a dendrogram by recursively merging the most similar objects.

[0105] Next, define the standard for the distance between clusters. Here, the Single Linkage method is adopted, that is, the distance between clusters is defined as the distance between the two closest points in the two clusters. Then, in the order of similarity from high to low, gradually merge the most similar objects to form a dendrogram. This dendrogram shows how the data points are merged into larger clusters step by step.

[0106] S703 cuts the tree structure based on the dendrogram to determine the final number of clusters.

[0107] Select a suitable height on the dendrogram according to actual needs or through certain criteria (such as the elbow method, silhouette coefficient, etc.) for cutting, so as to determine the final number of clusters. This cutting determines which data points belong to the same cluster, and thus forms the final customer clustering result.

[0108] S105 uses the consumption scenario thesaurus to conduct scenario analysis on the key elements in each customer clustering result to obtain the scenario clustering result;

[0109] The specific steps include:

[0110] S801 constructs a consumption scenario thesaurus containing various consumption scenario keywords based on industry knowledge, market research, and historical data;

[0111] According to industry expertise, market research, and historical transaction data, create a detailed consumption scenario keyword thesaurus. This thesaurus should cover various possible consumption scenarios, such as "holiday shopping", "purchase of household daily necessities", "upgrade of personal electronic products", etc. Each consumption scenario should be associated with a set of descriptive keywords that can help the system accurately identify the consumption scenarios mentioned in the text.

[0112] S802 matches the key elements in the customer clustering result with the keywords in the consumption scenario thesaurus to find out the consumption scenarios;

[0113] Next, use the consumption scenario thesaurus established in the previous step to match the key elements (such as product name, price, quantity, purchase time, etc.) in the customer clustering result with the keywords in the thesaurus. This process can be completed through text matching technology or more complex natural language processing methods, aiming to determine the most frequently occurring consumption scenarios in each cluster. For example, if most records in a certain cluster contain words such as "Christmas gifts", "discount season", etc., it can be classified as "holiday shopping".

[0114] S803 converts the matched consumption scenarios into the second numerical feature vector to obtain the second feature dataset;

[0115] Once the consumption scenarios corresponding to each cluster are determined, the next step is to convert this information into numerical form, namely the second numerical feature vector. One-Hot Encoding, TF-IDF or other suitable methods can be used to represent the presence and intensity of different consumption scenarios. This not only facilitates the processing of machine learning algorithms, but also effectively captures the relationship between different consumption scenarios.

[0116] S804 selects a K-Means clustering algorithm to perform cluster analysis on the second feature data set to obtain a scene clustering result.

[0117] The specific steps include:

[0118] S901 determines the optimal number of clusters by the elbow method;

[0119] The Elbow Method is used to determine the optimal number of clusters K. The Elbow Method plots the change curve of the intra-cluster sum of squared error (SSE) as the number of clusters increases, and finds the point where the curve begins to flatten as the optimal number of clusters. This is because when the number of clusters reaches a certain value, the error reduction caused by increasing the number of clusters will be significantly reduced.

[0120] S902 randomly selects K sample points in the second feature data set as initial centroids;

[0121] After selecting the optimal number of clusters K, K sample points are randomly selected from the second feature data set as initial centroids. These centroids represent the center position of each cluster.

[0122] S903 calculates the distance between each sample and each centroid, and assigns the sample to the cluster to which the nearest centroid belongs;

[0123] Calculate the distance from each sample point to all centroids and assign it to the cluster represented by the nearest centroid. This step is repeated until all samples are assigned to a cluster.

[0124] S904 recalculates the centroid position of each cluster, calculates the distance from each sample to each centroid, and assigns the sample to the cluster to which the nearest centroid belongs, until the maximum number of iterations is reached.

[0125] After each sample is assigned, the new centroid position of each cluster needs to be recalculated (i.e., the average value of the coordinates of all sample points in the cluster). Then the distances from all sample points to the new centroid are calculated again, and the clusters to which they belong are updated. This process is iterated continuously until the centroid no longer changes significantly or the preset maximum number of iterations is reached.

[0126] S106 generates potential needs of customers based on the cluster analysis results and scenario analysis results.

[0127] First, analyze the customer clustering results obtained through the K-Means algorithm and the secondary clustering results based on consumption scenarios. Each cluster represents a group of customers with similar behavior patterns or consumption preferences. It is necessary to clarify the main characteristics of each cluster, such as purchase frequency, preferred product categories, average consumption amount, main shopping time, etc. In addition, it is also necessary to analyze the differences between different clusters to understand which factors can best distinguish different customer groups.

[0128] Then, combine the results of the scenario analysis to refine the understanding of the potential needs of each customer group. For example, if a certain cluster is identified as "holiday shopping" enthusiasts, then this group may have a high interest in seasonal promotion activities, special holiday packages, or limited-edition products. Another example is that the "purchase of household daily necessities" cluster may indicate that this group values the cost-effectiveness, convenience, and brand loyalty of products. In this way, transform the abstract data analysis results into specific customer need insights.

[0129] Based on the above analysis, identify the potential needs of each customer group. This is not limited to the needs they have already shown (such as purchase history), but also includes those unmet or not yet realized needs. For example, provide customized recommendations according to the customer's purchase habits and preferences. For family users who are concerned about cost-effectiveness, a membership plan or a point reward system can be launched.

[0130] Finally, formulate corresponding strategies based on the identified potential needs. This may include, but is not limited to: developing new products or improving existing products based on the customer's demand for specific types of products. Design promotion activities that meet the preferences of the target customer group, such as special offers during holidays. Improve service quality, especially for those customer groups that value after-sales service and support.

[0131] The above disclosure is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A big data mining and analysis method, It is characterized in that Including: collecting multi-source heterogeneous data from customers, including transaction records, social media activities, and market research feedback; Cleaning and preprocessing the multi-source heterogeneous data to obtain initial text data; Extract key elements from the initial text data using natural language processing technology, wherein the key elements include product name, product price, product quantity, and purchase time; Use synonym matching word library to cluster key elements and obtain customer clustering results; Use the consumption scenario word library to perform scenario analysis on the key elements in each customer clustering result to obtain the scenario clustering result; Generate customer potential needs based on cluster analysis results and scenario analysis results.

2. A big data mining and analysis method as claimed in claim 1, characterized in that: The specific steps of collecting multi-source heterogeneous data of customers include: Extract customer sales information through the company's ERP and CRM systems; Extract the text communication content of the corresponding customer from the chat tool; Design and distribute electronic questionnaires to corresponding customers to collect their opinions on the product; Sales information, text communication content, and product opinion content are structured based on customer information to obtain multi-source heterogeneous data.

3. A big data mining and analysis method as claimed in claim 2, characterized in that: The specific steps of cleaning and preprocessing the multi-source heterogeneous data to obtain initial text data include: Check and delete duplicate records in multi-source heterogeneous data; Delete missing data points in multi-source heterogeneous data; The stemming algorithm is applied to normalize different forms of words into basic forms, reducing the number of vocabulary variants to obtain the initial text data.

4. A big data mining and analysis method as claimed in claim 3, characterized in that: The specific steps of applying the stem extraction algorithm to normalize different forms of words into basic forms and reduce the number of vocabulary variants to obtain initial text data include: Segment multi-source heterogeneous data into separate word data groups; Select the Porter Stemmer algorithm to extract the stems in the word data set to obtain the stem data set; After stemming, stop words are removed to obtain the initial text data.

5. A big data mining and analysis method as claimed in claim 4, characterized in that: The specific steps of extracting key elements from the initial text data using natural language processing technology include: Identify product names based on product catalog matching; Match product prices and product quantities based on regular expressions; Identify purchase time based on date parsing library.

6. A big data mining and analysis method as claimed in claim 5, characterized in that: The specific steps of using the synonym matching word library to perform cluster analysis on key elements and obtain customer clustering results include: Create a dictionary containing key element words and their synonyms based on business requirements; Expand the extracted key elements into synonyms based on the dictionary; Convert the text after synonym expansion into a first numerical feature vector; A hierarchical clustering algorithm is used to perform cluster analysis on the first numerical feature vector to obtain customer clustering results.

7. A big data mining and analysis method as claimed in claim 6, characterized in that: The specific steps of using the hierarchical clustering algorithm to perform cluster analysis on the numerical feature vector to obtain the customer clustering result include: According to the selected distance metric, the distance between each two key elements is calculated and a distance matrix is ​​constructed; Define the inter-cluster distance as the distance between the closest two points in two clusters, and gradually build a dendrogram by recursively merging the most similar objects; The tree structure was cut based on the dendrogram to determine the final number of clusters.

8. A big data mining and analysis method as claimed in claim 7, characterized in that: The specific steps of using the consumption scenario word library to perform scenario analysis on the key elements in each customer clustering result to obtain the scenario clustering result include: Build a consumer scenario vocabulary containing keywords for various consumer scenarios based on industry knowledge, market research and historical data; Match the key elements in the customer clustering results with the keywords in the consumption scenario vocabulary to find out the consumption scenarios; Convert the matched consumption scenario into a second numerical feature vector to obtain a second feature data set; The K-Means clustering algorithm is selected to perform cluster analysis on the second feature data set to obtain the scene clustering results.

9. A big data mining and analysis method as claimed in claim 8, characterized in that: The specific steps of selecting the K-Means clustering algorithm to perform cluster analysis on the numerical feature vector include: The optimal number of clusters was determined by the elbow method; Randomly select K sample points in the second feature data set as the initial centroid; Calculate the distance between each sample and each centroid, and assign the sample to the cluster to which the nearest centroid belongs; Recalculate the centroid position of each cluster and calculate the distance of each sample to each centroid, and assign the sample to the cluster to which the nearest centroid belongs until the maximum number of iterations is reached.