A method for constructing a customized production enterprise customer portrait
Patent Information
- Application Number
- CN202211422318.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-14
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2042-11-14
AI Technical Summary
现有技术能够让业务人员从客户标签体系中筛选出客户标签作为聚类特征,并进行简单配置后,即可自动完成客户聚类分群,将聚类分群结果展示到前台,呈现给业务人员,然而,该方案中,业务人员将需要进行分群的客户名单上传以获取待分群的客户数据集,操作冗杂耗时长,且客户信息挖掘不足,不能提供足够丰富的客户需求信息,难以满足商业应用的实际需要
[0051] By employing web crawling techniques to extract customer-related attribute characteristics from the internet, and integrating enterprise ERP system data with web crawling data and performing preprocessing, we can deeply mine and obtain sufficiently rich customer demand information, enriching the customer profile of manufacturing enterprises to better explore the potential value of customers. K-means clustering is used for cluster analysis to divide multiple customers with similar characteristics into several customer groups. Finally, the clustering results are output and displayed to assist business personnel in decision-making and meet the actual needs of commercial applications.
Smart Images

Figure CN115757917B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of customized manufacturing, and more specifically, to a method for constructing customer profiles for customized manufacturing enterprises. Background Technology
[0002] Customer profiling is built on a series of real target customer data. It involves selectively extracting and refining customer attribute characteristics to uncover and identify important customer features and form a representative, tagged model.
[0003] With the continuous application of big data technology, methods for refined analysis and intelligent decision-making using big data have gradually become a research hotspot. Customer profiling, as an effective tool for identifying target customers and improving decision-making efficiency, has been widely used in various fields. Scholars primarily build models from real datasets in areas such as social networks, e-commerce consumption, mobile communications, library resources, and banking wealth management to extract effective customer characteristics and abstract behavioral profiles of different customer groups. This allows for targeted services to specific customer groups, improving service efficiency.
[0004] When designing and developing products or services, companies often focus not on the preferences of a single customer, but on the preferences of one or several customer groups. Therefore, creating a single customer profile not only fails to provide sufficiently rich information on customer needs, but also requires a huge amount of computation, making it difficult to meet the actual needs of commercial applications.
[0005] Current research on customer profiling for customized manufacturing companies is limited. With the booming development of the Industrial Internet, artificial intelligence, smart agriculture, and agricultural mechanization, customized manufacturing companies specializing in agricultural equipment parts and electronic components are rapidly emerging, leading to increasingly fierce competition. Customers are the source of revenue for these companies, and maintaining good customer relationship management and improving service levels are crucial for maintaining a competitive edge. Data-driven customer profiles can recreate a complete picture of customers, providing the foundation for companies to uncover customer needs and value, segment customers, and implement precise marketing. It is foreseeable that applying profiling technology to customized manufacturing will greatly help business personnel gain a deeper understanding of customers, manage different customers differently, and better tap into their potential value, ultimately bringing profits to the company.
[0006] Existing technology discloses a customer segmentation method based on clustering analysis, including the following steps: establishing a tag profile system; obtaining a customer dataset to be segmented; selecting customer tags to generate an initial customer tag library; configuring the number of clusters K, and selecting whether to reduce the dimensionality of the tags in the initial customer tag library; using principal component analysis to reduce the dimensionality of continuous tags in the customer tag library to be analyzed, and performing One-Hot encoding on categorical tags to generate a final customer tag library; performing clustering analysis on the final customer tag library using the k-means algorithm, generating clustering results and displaying them. Existing technology allows business personnel to select customer tags from the customer tag system as clustering features, and after simple configuration, automatically complete customer clustering and display the clustering results on the front end for business personnel. However, in this solution, business personnel upload the list of customers to be segmented to obtain the customer dataset to be segmented, which is cumbersome and time-consuming, and the customer information mining is insufficient, failing to provide sufficiently rich customer demand information, making it difficult to meet the actual needs of commercial applications. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for constructing customized customer profiles for manufacturing enterprises. This method enriches the customer profiles of manufacturing enterprises, deeply mines and obtains sufficiently rich customer demand information, so as to better explore the potential value of customers and meet the actual needs of commercial applications.
[0008] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0009] A method for building customized customer profiles for manufacturing enterprises is provided, including the following steps:
[0010] S1: Establish a profile tagging system;
[0011] S2: Use web crawlers to extract customer attribute feature data from the Internet;
[0012] S3: Merge the order data from the enterprise ERP system and the feature data of customer attributes crawled in step S2 according to the customer ID name, and preprocess the merged data;
[0013] S4: Establish a clustering model, use the K-means algorithm to perform clustering analysis on the data preprocessed in step S3, output the clustering results and display them.
[0014] The customized customer profile construction method for manufacturing enterprises of this invention first establishes a profile tagging system, uses web crawling to crawl customer-related attribute features from the Internet, integrates enterprise ERP system data and web crawling data and performs preprocessing, deeply mines and obtains sufficiently rich customer demand information, enriching the customer profile of manufacturing enterprises in order to better explore the potential value of customers, and uses K-means clustering for cluster analysis, which can divide multiple customers with similar characteristics into several customer groups, and finally outputs and displays the clustering results to assist business personnel in decision-making and meet the actual needs of commercial applications.
[0015] Preferably, in step S1, the process of establishing a profile tagging system is as follows: sorting out and analyzing the domain characteristics of the customer profile scenario for customized production enterprises.
[0016] Preferably, in step S2, the workflow of the web crawler includes:
[0017] S21: Create a list of URLs;
[0018] S22: Determine if the URL list is empty. If yes, end data crawling; otherwise, proceed to step S23.
[0019] S23: Retrieve the URLs from the list in sequence, and send a network request to the server according to the URL address. The server responds according to the request method.
[0020] S24: Parse the data returned by the server, extract the relevant field information, store it in the database, and return to step S22.
[0021] Preferably, the specific process of step S3 is as follows:
[0022] S31: Integrate order data from the enterprise ERP system with web crawler data;
[0023] S32: Perform data cleaning on the merged data;
[0024] S33: Convert the data to timestamps;
[0025] S34: Construct image tags;
[0026] S35: Use one-hot encoding to convert and encode text tags in the image tags;
[0027] S36: Perform correlation analysis on the tag data;
[0028] S37: Normalize the label data.
[0029] Preferably, in step S4, the process of performing cluster analysis on the data is as follows:
[0030] S41: Select k sample points as the initial cluster centers {c1, c2, ..., c3} of K-means. k};
[0031] S42: Calculate the Euclidean distance from each sample point to the cluster center, and assign the sample points to the nearest cluster centers to obtain k clusters {s1, s2, ..., s3}. k};
[0032] S43: Recalculate the mean distance from each cluster sample point to the cluster center, and use it as the new cluster center c. j :
[0033]
[0034] In the formula, m′ represents the number of sample points in the j-th cluster; θ represents the number of sample points in the j-th cluster; s j Indicates the j-th cluster;
[0035] S44: Determine whether the cluster center has changed. If yes, repeat steps S42 and S43; otherwise, proceed to step S45.
[0036] S45: Calculate clustering evaluation indicators;
[0037] S46: Determine whether the optimal number of clusters k has been selected. If yes, output the clusters and the cluster centers of each cluster; otherwise, change the number of clusters k and return to step S41.
[0038] Preferably, in step S42, the Euclidean distance from the sample point to the cluster center is calculated as follows:
[0039]
[0040] In the formula, θ i c represents the i-th sample point, 1≤i≤m, where m is the number of samples; j The j-th cluster center is represented by θ; n represents the dimension of the sample points; θ it c represents the t-th dimension feature of the i-th sample point; jt This represents the t-th dimension feature of the j-th cluster center.
[0041] Preferably, in step S44, the stability of the sum of squared clustering errors is used to determine whether the cluster centers have changed. The calculation process for the sum of squared clustering errors is as follows:
[0042]
[0043] In the formula, SSE is the sum of squared clustering errors.
[0044] Preferably, in step S45, the clustering evaluation indicators include the silhouette coefficient (SC) and the CHI index, and the specific calculation method is as follows:
[0045]
[0046]
[0047] In the formula, m is the number of samples; a i b represents the average Euclidean distance between a sample point and other sample points in the same cluster; i B represents the average distance between a sample point and sample points within other clusters. k W represents the covariance matrix of inter-cluster sample data; k SC represents the covariance matrix of the sample data within the cluster; Tr represents the trace of the matrix, that is, the sum of the elements on the main diagonal of the matrix; the closer SC is to 1, the larger CHI is, and the better the clustering effect of the model.
[0048] Preferably, in step S4, a visual query interface is provided using a visual operating system to display the clustering results.
[0049] Preferably, the operation process of the visual operating system is as follows: log in to the visual operating system; visual operation query interface; visual correlation analysis; exit the visual operating system.
[0050] Compared with the prior art, the customized manufacturing enterprise customer profile construction method of the present invention has the following beneficial effects:
[0051] By employing web crawling techniques to extract customer-related attribute characteristics from the internet, and integrating enterprise ERP system data with web crawling data and performing preprocessing, we can deeply mine and obtain sufficiently rich customer demand information, enriching the customer profile of manufacturing enterprises to better explore the potential value of customers. K-means clustering is used for cluster analysis to divide multiple customers with similar characteristics into several customer groups. Finally, the clustering results are output and displayed to assist business personnel in decision-making and meet the actual needs of commercial applications. Attached Figure Description
[0052] Figure 1 This is a flowchart of the customized manufacturing enterprise customer profile construction method in Embodiment 1 of the present invention;
[0053] Figure 2 This is a flowchart of the web crawler's workflow in Embodiment 1 of the present invention;
[0054] Figure 3 This is a flowchart illustrating the data fusion and preprocessing process in Embodiment 2 of the present invention;
[0055] Figure 4 This is a flowchart of the clustering analysis of data in Embodiment 3 of the present invention;
[0056] Figure 5 This is a flowchart of the operation of the visual operating system in Embodiment 3 of the present invention. Detailed Implementation
[0057] The present invention will be further described below with reference to specific embodiments.
[0058] Example 1
[0059] A method for building customer profiles for customized manufacturing enterprises, such as Figure 1 As shown, it includes the following steps:
[0060] S1: Establish a profile tagging system;
[0061] S2: Use web crawlers to extract customer attribute feature data from the Internet;
[0062] S3: Merge the order data from the enterprise ERP system and the feature data of customer attributes crawled in step S2 according to the customer ID name, and preprocess the merged data;
[0063] S4: Establish a clustering model, use the K-means algorithm to perform clustering analysis on the data preprocessed in step S3, output the clustering results and display them.
[0064] The aforementioned method for constructing customized customer profiles for manufacturing enterprises first establishes a profile tagging system. It then uses web crawlers to extract relevant customer attributes from the internet, integrates enterprise ERP system data with the crawled data, and preprocesses it to deeply mine and obtain sufficiently rich customer demand information. This enriches the customer profiles of manufacturing enterprises, enabling better exploration of potential customer value. K-means clustering is then used for cluster analysis, dividing multiple customers with similar characteristics into several customer groups. Finally, the clustering results are output and displayed to assist business personnel in decision-making, meeting the practical needs of commercial applications.
[0065] In step S1, the domain characteristics of the customer profile scenario for customized production enterprises are analyzed, business is sorted out, and a profile tagging system is established.
[0066] Manufacturing companies' customers are often businesses themselves, not individual customers, representing a typical business-to-business transaction model. Their customer base is relatively small, and the transaction partners are relatively fixed. In practice, customer size, economic strength, industry, and regional influence are important factors for differentiated management by sales personnel, making the formation of customer profile tags for manufacturing companies more complex. In step S2, a web crawler is a program that automatically retrieves web page data by simulating manual login according to certain rules. This facilitates the extraction of customer attribute feature data. The workflow of a web crawler is as follows... Figure 2As shown, it includes:
[0067] S21: Create a list of URLs;
[0068] S22: Determine if the URL list is empty. If yes, end data crawling; otherwise, proceed to step S23.
[0069] S23: Retrieve the URLs from the list in sequence, and send a network request to the server according to the URL address. The server responds according to the request method.
[0070] S24: Parse the data returned by the server, extract the relevant field information, store it in the database, and return to step S22.
[0071] Example 2
[0072] This embodiment is similar to Embodiment 1, except that the specific process of step S3 is as follows: Figure 3 As shown, it includes the following steps:
[0073] S31: Integrate order data from the enterprise ERP system with web crawler data;
[0074] S32: Perform data cleaning on the merged data;
[0075] S33: Convert the data to timestamps;
[0076] S34: Construct image tags;
[0077] S35: Use one-hot encoding to convert and encode text tags in the image tags;
[0078] S36: Perform correlation analysis on the tag data;
[0079] S37: Normalize the label data.
[0080] In step S35, some text-based labels in the image tags cannot be directly input into the clustering model for calculation and require encoding processing to convert them into numerical labels. One-hot encoding is used for this conversion and encoding. This method uses an N-bit state register to encode N states, each with its own independent register bit. At any given time, only one bit is valid (i.e., only one bit is 1, and the rest are 0). One-hot encoding uses 0 and 1 to represent parameters and uses an N-bit state register to encode N states. This extends the values of discrete features to Euclidean space, where a value of a discrete feature corresponds to a point in Euclidean space. Using one-hot encoding for discrete features makes the distance calculation between features more reasonable.
[0081] In step S36, let x and y be two variables in the label data. The formulas involved in the correlation analysis are as follows:
[0082]
[0083] In the formula, r represents the Pearson correlation coefficient. Represents the covariance of variables x and y; The standard deviation of variable x; The standard deviation of variable y is represented by r. The larger the absolute value of r, the stronger the correlation, that is, the closer the correlation coefficient is to 1 or -1, the stronger the correlation. The closer the correlation coefficient is to 0, the weaker the correlation. The correlation strength of variables can be judged by the following value ranges: 0.8-1.0 is extremely strong correlation; 0.6-0.8 is strong correlation; 0.4-0.6 is moderate correlation; 0.2-0.4 is weak correlation; and 0.0-0.2 is extremely weak correlation or no correlation.
[0084] Since the various profile labels typically have different dimensions and orders of magnitude, when the levels of different indicators differ significantly, directly using the raw indicator values for analysis will highlight the role of indicators with higher numerical values in the comprehensive analysis, while relatively weakening the role of indicators with lower numerical levels. Therefore, to ensure the reliability of the results, it is necessary to normalize the label data. Let the label data be x. i and x j After data normalization, it becomes x ij :
[0085]
[0086] Data normalization does not change the distribution characteristics of the original sample data, but linearly maps all sample data to the interval [0,1].
[0087] Example 3
[0088] This embodiment is similar to Embodiment 2, except that in step S4, the K-means algorithm has a simple implementation principle, fast convergence speed, and superior clustering effect. Figure 4 As shown, the process of performing cluster analysis on the data is as follows:
[0089] S41: Select k sample points as the initial cluster centers {c1, c2, ..., c3} of K-means. k};
[0090] S42: Calculate the Euclidean distance from each sample point to the cluster center, and assign the sample points to the nearest cluster centers to obtain k clusters {s1, s2, ..., s3}. k The Euclidean distance from a sample point to the cluster center is calculated as follows:
[0091]
[0092] In the formula, θ i c represents the i-th sample point, 1≤i≤m, where m is the number of samples; j The j-th cluster center is represented by θ; n represents the dimension of the sample points; θ it c represents the t-th dimension feature of the i-th sample point; jt This represents the t-th dimension feature of the j-th cluster center;
[0093] S43: Recalculate the mean distance from each cluster sample point to the cluster center, and use it as the new cluster center c. j :
[0094]
[0095] In the formula, m′ represents the number of sample points in the j-th cluster; θ represents the number of sample points in the j-th cluster; s j Indicates the j-th cluster;
[0096] S44: The stability of the sum of squared clustering errors is used to determine whether the cluster centers have changed. The calculation process for the sum of squared clustering errors is as follows:
[0097]
[0098] In the formula, SSE is the sum of squared clustering errors;
[0099] If SSE is unstable, repeat steps S42 and S43; if SSE is stable, proceed to step S45.
[0100] S45: Calculate clustering evaluation indices; clustering evaluation indices include the silhouette coefficient (SC) and the CHI index, and the specific calculation methods are as follows:
[0101]
[0102]
[0103] In the formula, m is the number of samples; a i b represents the average Euclidean distance between a sample point and other sample points in the same cluster; i B represents the average distance between a sample point and sample points within other clusters. k W represents the covariance matrix of inter-cluster sample data; k SC represents the covariance matrix of the sample data within the cluster; Tr represents the trace of the matrix, that is, the sum of the elements on the main diagonal of the matrix; the closer SC is to 1, the larger CHI is, and the better the clustering effect of the model.
[0104] S46: Determine whether the optimal number of clusters k has been selected. If yes, output the clusters and the cluster centers of each cluster; otherwise, change the number of clusters k and return to step S41.
[0105] In step S4, a visual query interface is provided using a visual operating system to display the clustering results. For example... Figure 5 As shown, the operation flow of the visual operating system is as follows: Log in to the visual operating system; Visual operation query interface; Visual correlation analysis; Exit the visual operating system. The visual operating system is an extension of customer profile representations (radar charts, bar charts, pie charts), providing intuitive displays to assist business personnel in making relevant decisions.
[0106] In the specific implementation of the above embodiments, the technical features can be combined in any non-contradictory way. For the sake of brevity, not all possible combinations of the above technical features are described. However, as long as the combination of these technical features is not contradictory, it should be considered to be within the scope of this specification.
[0107] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for constructing customer profiles for customized manufacturing enterprises, characterized in that, Includes the following steps: S1: Establish a profile tagging system; S2: Use web crawlers to extract customer attribute feature data from the Internet; S3: Merge the order data from the enterprise ERP system and the feature data of customer attributes crawled in step S2 according to the customer ID name, and preprocess the merged data; S4: Establish a clustering model, use the K-means algorithm to perform clustering analysis on the data preprocessed in step S3, output the clustering results and display them; In step S4, the process of performing cluster analysis on the data is as follows: S41: Select 100 sample points are used as the initial cluster centers of K-means. , ; S42: Calculate the Euclidean distance from each sample point to the cluster center, and assign the sample points to the nearest cluster center. Cluster ; S43: Recalculate the mean distance from each cluster sample point to the cluster center, and use it as the new cluster center. : In the formula, Indicates the first Number of sample points in each cluster; Indicates the first Sample points in a cluster; Indicates the first A cluster; S44: Determine whether the cluster center has changed. If yes, repeat steps S42 and S43; otherwise, proceed to step S45. S45: Calculate clustering evaluation indicators; S46: Determine whether the optimal cluster number has been selected. If yes, output the clusters and their cluster centers; otherwise, change the number of clusters. Return to step S41.
2. The method for constructing customized customer profiles for manufacturing enterprises according to claim 1, characterized in that, In step S1, the process of establishing a profile tagging system involves sorting out and analyzing the domain characteristics of customized production enterprise customer profile scenarios.
3. The method for constructing customized production enterprise customer profiles according to claim 1, characterized in that, In step S2, the workflow of the web crawler includes: S21: Create a list of URLs; S22: Determine if the URL list is empty. If yes, end data crawling. If no, proceed to step S23. S23: Retrieve the URLs from the list in sequence, and send a network request to the server according to the URL address. The server responds according to the request method. S24: Parse the data returned by the server, extract the relevant field information, and store it in the database. Return to step S22.
4. The method for constructing customized production enterprise customer profiles according to claim 2, characterized in that, The specific process of step S3 is as follows: S31: Integrate order data from the enterprise ERP system with web crawler data; S32: Perform data cleaning on the merged data; S33: Convert the data to timestamps; S34: Construct image tags; S35: Use one-hot encoding to convert and encode text tags in the image tags; S36: Perform correlation analysis on the tag data; S37: Normalize the label data.
5. The method for constructing customized production enterprise customer profiles according to claim 1, characterized in that, In step S42, the Euclidean distance from the sample point to the cluster center is calculated as follows: In the formula, Indicates the first One sample point, , The number of samples; Indicates the first Cluster center; Indicates the dimension of the sample points; Indicates the first The first sample point Features in one dimension; Indicates the first The first cluster center Features in 100 dimensions.
6. The method for constructing customized production enterprise customer profiles according to claim 1, characterized in that, In step S45, the clustering evaluation metrics include the silhouette coefficient. and The index is calculated as follows: In the formula, The number of samples; This represents the average Euclidean distance between a sample point and other sample points in the same cluster. This represents the average distance between a sample point and sample points within other clusters; Represents the covariance matrix of inter-cluster sample data; This represents the covariance matrix of the sample data within the cluster; The trace of a matrix is the sum of the elements on its main diagonal. The closer to 1, The larger the value, the better the clustering effect of the model.
7. The method for constructing customized customer profiles for manufacturing enterprises according to any one of claims 1 to 6, characterized in that, In step S4, a visual query interface is provided using a visual operating system to display the clustering results.
8. The method for constructing customized production enterprise customer profiles according to claim 7, characterized in that, The operation process of the visual operating system is as follows: Log in to the visual operating system; Visual operation query interface; Visualize the correlation analysis; exit the visual operating system.
Citation Information
Patent Citations
Shopping center operation data management decision system and method
CN110413680A
Customer grouping implementation method based on clustering analysis
CN111159258A