Building industry Internet platform built based on data elements and clustering algorithm
By integrating multi-source data and using clustering algorithms, a spatiotemporal feature matrix is generated. By combining the elbow rule and the silhouette coefficient method to select clustering parameters, the problem of differentiated and refined management in existing supplier management systems is solved. This enables accurate supplier classification and risk warning, and improves the adaptability and efficiency of supply chain management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-27
AI Technical Summary
Existing supplier management systems in the construction industry struggle to achieve differentiated and refined management, especially when faced with highly heterogeneous data. Fixed parameters are difficult to adjust adaptively, leading to distorted clustering results and delayed risk response.
A spatiotemporal feature matrix is generated by integrating multi-source data. The elbow rule and the silhouette coefficient method are combined to select a clustering algorithm and generate a clustering parameter configuration set. This enables multidimensional evaluation of suppliers and quantification of cluster characteristics, and the construction of a risk warning mechanism to achieve accurate supplier classification and risk monitoring.
It enables precise clustering and differentiated management of suppliers, improves the accuracy of supplier classification and business adaptability, timely identifies high-risk suppliers, breaks through the fixed threshold response delay, and achieves precise and proactive prevention and control of supply chain risks.
Smart Images

Figure CN121743916A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of building industry industry chain supplier management, and particularly relates to a building industry industry internet platform based on data elements and a clustering algorithm. BACKGROUND
[0002] In the building industry, enterprises usually cooperate with a large number of suppliers, including building material suppliers, mechanical equipment suppliers, labor service suppliers, etc. Different types of suppliers have significant differences in product quality, price, delivery period, service level, etc. The existing supplier management method is usually difficult to realize differentiation and refinement management.
[0003] The building industry supply chain management generally adopts a multi-source data integration framework. The conventional method constructs a supplier portrait through a structured index system and divides the supplier groups. Industry practice usually relies on static parameter configuration and a single evaluation index, supplemented by historical performance data to set a static risk threshold. Such technology relies on the data center integration capability to realize the standardized execution of supplier resource allocation strategies and provides basic decision support for the engineering construction industry.
[0004] The current building industry supplier procurement platform generally relies on manual experience to preset parameters. When the logistics time efficiency standard deviation is greater than the high heterogeneity data, it is difficult for fixed parameters to adaptively adjust, resulting in distorted results. In addition, risk early warning relies on static thresholds and does not combine real-time fluctuation ranges within the cluster, which cannot capture sudden abnormalities, causing risk response lag.
[0005] In addition, the current building industry procurement platform still has the problem of insufficient dynamic adaptability of the grading model in supplier grading and classification management, including but not limited to the following aspects: 1. Static model lag: the K-means clustering model is retrained every quarter and cannot respond to sudden changes in suppliers (such as a waterproof material supplier whose production capacity utilization rate jumps from 60% to 85% due to production line upgrading, but the grading is not adjusted synchronously).
[0006] 2. Fixed index weight: the index weight determined by the Analytic Hierarchy Process (AHP) (such as the production capacity stability weight of 0.35) is not dynamically adjusted with the project stage (such as the delivery capacity weight of the waterproof material supplier should be increased to 0.45 before the rainy season).
[0007] 3. New supplier cold start problem: the logistic regression model relies on historical data, and the grading accuracy of new suppliers without transaction records is relatively low, which is only 50-60% in some cases. SUMMARY
[0008] The application provides a building industry industrial internet platform based on data elements and a clustering algorithm to realize differentiated and refined management of suppliers.
[0009] The application provides a building industry industrial internet platform based on data elements and a clustering algorithm, comprising: A multi-source data integration module is configured to obtain multi-source data of suppliers and perform preprocessing to generate a spatiotemporal feature matrix. A multi-dimensional clustering module is configured to select a clustering algorithm based on multi-dimensional evaluation of the spatiotemporal feature matrix, generate a clustering parameter configuration set by combining elbow rule and silhouette coefficient method. A cluster group feature quantification module is configured to perform cluster analysis on the suppliers by using the selected clustering algorithm based on the spatiotemporal feature matrix and the clustering parameter configuration set, and generate a cluster group feature quantification report. A hierarchical strategy generation module is configured to generate a supplier hierarchical label set by generating a hierarchical rule based on the cluster group feature quantification report, and generate a supplier hierarchical strategy execution instruction set by combining a preset differentiated resource allocation rule. A risk early warning module is configured to construct a risk index threshold rule set based on the supplier hierarchical strategy execution instruction set and the cluster group feature quantification report, and generate a high-risk supplier early warning list by a preset threshold comparison mechanism.
[0010] Optionally, the multi-dimensional evaluation of the spatiotemporal feature matrix to select a clustering algorithm, combined with elbow rule and silhouette coefficient method, to generate a clustering parameter configuration set, comprises: The data size of the spatiotemporal feature matrix is evaluated by using a capacity estimation method. The data uniformity of the spatiotemporal feature matrix is evaluated by using a density distribution histogram analysis method. The clustering algorithm is selected according to the data size and the data uniformity. A plurality of different cluster numbers are selected from a preset to-be-evaluated cluster number interval. Based on each of the cluster numbers, a first candidate cluster number is determined by using elbow rule. Based on each of the cluster numbers, a second candidate cluster number is determined by using silhouette coefficient method. The target cluster number is determined based on the first candidate cluster number and the second candidate cluster number, and a clustering parameter configuration set is generated by combining the selected clustering algorithm.
[0011] Optionally, the clustering algorithm is selected according to the data size and the data uniformity, comprising: When the data size is in a first preset size interval and the data uniformity meets a first preset uniformity standard, a K-means clustering algorithm is selected. When the data scale is in a second preset scale interval and a first preset service requirement is met, a hierarchical clustering algorithm is selected, the first preset service requirement being that a clustering hierarchy needs to be generated for a business scenario in the construction industry; When the data uniformity meets a second preset uniformity standard, a density-based clustering algorithm is selected.
[0012] Optionally, the elbow rule is used to determine a first candidate clustering number based on the clustering numbers, including: The sum of squared errors of clustering corresponding to each clustering number is calculated. A relationship curve between the clustering numbers and the sum of squared errors of clustering corresponding thereto is drawn. The clustering number corresponding to a curve slope mutation point of each relationship curve is selected as the first candidate clustering number.
[0013] Optionally, the silhouette coefficient method is used to determine a second candidate clustering number based on the clustering numbers, including: A plurality of silhouette coefficients corresponding to each clustering number is calculated. An average silhouette coefficient corresponding to each clustering number is determined according to the silhouette coefficients. The clustering number corresponding to a maximum value of the average silhouette coefficients is selected as the second candidate clustering number.
[0014] Optionally, a target clustering number is determined based on the first candidate clustering number and the second candidate clustering number, and a clustering parameter configuration set is generated in combination with the selected clustering algorithm, including: When the first candidate clustering number and the second candidate clustering number are consistent, the first candidate clustering number or the second candidate clustering number is determined as the target clustering number. When the first candidate clustering number and the second candidate clustering number are inconsistent, a clustering operation is performed on the spatio-temporal feature matrix using the second candidate clustering number. If a second clustering operation result meets a preset business scenario requirement, the second candidate clustering number is determined as the target clustering number. If the second clustering operation result does not meet the preset business scenario requirement, a clustering operation is performed on the spatio-temporal feature matrix using the first candidate clustering number. If the first clustering operation result meets the preset business scenario requirement, the first candidate clustering number is determined as the target clustering number. If the first clustering operation result does not meet the preset business scenario requirement, the first candidate clustering number is fine-tuned. The fine-tuned number of the first candidate clusters is used as the new number of the first candidate clusters. The process then jumps to the step of performing clustering operations on the spatiotemporal feature matrix using the first number of candidate clusters until the target number of clusters is output. Using the selected clustering algorithm as the key, a preset clustering parameter key-value library is retrieved to match the corresponding target clustering parameters; Integrate the target cluster number and the target cluster parameters to generate a cluster parameter configuration set.
[0015] Optionally, the step of performing cluster analysis on the suppliers based on the spatiotemporal feature matrix and the clustering parameter configuration set, using the selected clustering algorithm, and generating a cluster feature quantification report includes: Based on the clustering parameter configuration set, the corresponding cluster centers are generated by the selected clustering algorithm; Using the selected clustering algorithm, the cluster assignment distance from the supplier data points in the spatiotemporal feature matrix to each cluster center is calculated; Based on the cluster allocation distance, each supplier is assigned to the cluster to which the cluster center corresponding to the smallest cluster allocation distance belongs, thereby generating supplier cluster partitioning data; The mean and standard deviation of preset dimension indicators in each cluster are statistically analyzed. The preset dimension indicators include at least one of supply capacity, product quality, delivery timeliness, product price, and service quality. By comparing the differences in the preset dimensional indicators among the various clusters, the mean and standard deviation differences of the dimensions of each cluster are calculated to generate cluster difference data. By integrating the supplier cluster division data, the statistical results of preset dimension indicators of each cluster, and the cluster difference data, a cluster feature quantification report is generated.
[0016] Optionally, based on the cluster feature quantification report, a supplier grading label set is generated through preset grading rules, and combined with preset differentiated resource allocation rules, a supplier grading strategy execution instruction set is generated, including: Based on the statistical results of the preset dimension indicators in the cluster feature quantification report, and combined with the supplier's historical performance data, the threshold of the grading rules corresponding to each preset dimension is determined. The statistical results of the preset dimension indicators of each cluster are compared with the corresponding hierarchical rule thresholds, and a corresponding hierarchical label is matched for each cluster. The hierarchical labels of all clusters are then summarized to generate a supplier hierarchical label set. Based on the supplier cluster division data in the cluster feature quantification report, the number of suppliers corresponding to each cluster is counted. According to the correspondence relationship between the supplier hierarchical label set and the preset differentiated resource allocation rule, resource proportion intervals are allocated to each hierarchical label, and in combination with the number of suppliers and the total resource amount of a preset business scenario, resource actual allocation values of each cluster group are calculated to generate a differentiated resource allocation proportion table; The supplier hierarchical strategy execution instruction set is integrated with the differentiated resource allocation proportion table to generate a supplier hierarchical strategy execution instruction set.
[0017] Optionally, the risk indicator threshold rule set is constructed based on the supplier hierarchical strategy execution instruction set and the cluster group feature quantization report, and a high-risk supplier early warning list is generated through a preset threshold comparison mechanism, including: Based on the supplier hierarchical strategy execution instruction set, the cluster group feature quantization report and supplier historical performance data, a static threshold of a key risk indicator corresponding to a preset dimension indicator is determined; Based on the standard deviation of each preset dimension indicator in the cluster group feature quantization report, a dynamic threshold interval of the key risk indicator is determined; The static threshold and the dynamic threshold interval of the key risk indicator are integrated, and the corresponding risk level is marked to generate a risk indicator threshold rule set; Real-time risk indicator data of the supplier are obtained, and according to the risk indicator threshold rule set, a supplier whose risk indicator exceeds the static threshold and / or the dynamic threshold interval is screened out through a preset threshold comparison mechanism, the corresponding risk type is marked, and a high-risk supplier candidate set is generated; The high-risk supplier candidate set is associated with the supplier hierarchical label set, and supplier ID, cluster group label and hierarchical label information are supplemented to generate a high-risk supplier early warning list.
[0018] Optionally, the supplier multi-source data are obtained and preprocessed to generate a space-time feature matrix, including: Supplier multi-source data are obtained; The supplier multi-source data are preprocessed; The supplier multi-source data after data preprocessing are subjected to feature extraction to obtain time dimension features and space dimension features; The time dimension features and the space dimension features are used to construct a space-time feature matrix.
[0019] As can be seen from the above technical solutions, the present application has the following advantages: The multi-dimensional clustering module selects a clustering algorithm based on the spatiotemporal feature matrix through multi-dimensional evaluation, generates a clustering parameter configuration set in combination with the elbow rule and the silhouette coefficient method, can select a clustering algorithm based on data size and distribution characteristics, dynamically optimizes clustering parameters by combining the elbow rule and the silhouette coefficient method, thereby eliminating empirical preset bias, and significantly improves the accuracy and business adaptability of supplier classification. The risk early warning module generates a risk indicator threshold rule set based on the supplier grading strategy execution instruction set and the cluster feature quantization report, generates a high-risk supplier early warning list through a threshold comparison mechanism according to the risk indicator threshold rule set, constructs a composite risk rule set by integrating static thresholds and dynamic thresholds, monitors each indicator in real time, automatically marks high-risk suppliers when the data exceeds the threshold, generates an early warning list, breaks through the fixed threshold response delay, and realizes precise and proactive prevention and control of supply chain risks. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0021] Figure 1 A structural block diagram of a building industry industrial internet platform based on data elements and clustering algorithms provided by the embodiments of the present application; Figure 2 A supplier management flowchart provided by the embodiments of the present application; Figure 3 A multi-source data acquisition and preprocessing flowchart provided by the embodiments of the present application; Figure 4 A supplier cluster division parameter optimization flowchart provided by the embodiments of the present application; Figure 5 A supply chain risk dynamic early warning mechanism flowchart provided by the embodiments of the present application. DETAILED DESCRIPTION
[0022] The existing supplier management usually adopts a multi-source data integration and clustering analysis framework. The conventional method constructs a supplier portrait through a structured index system, and divides the supplier groups based on a K-means or hierarchical clustering algorithm. The industry practice usually relies on static parameter configuration and a single evaluation index, supplemented by historical performance data to set a static risk threshold. Such technology relies on the data middle platform integration capability to realize the standardized execution of the supplier resource allocation strategy, and provides basic decision support for the engineering construction industry. The current supplier management method in the construction industry generally faces the problem that the clustering parameters need to rely on artificial experience presetting, and when the standard deviation of the logistics time limit is greater than the highly heterogeneous data, the fixed parameters are difficult to adaptively adjust, resulting in distorted clustering results. In addition, the risk early warning relies on a static threshold, and does not combine the real-time fluctuation range in the cluster group, so it cannot capture sudden abnormalities, causing risk response lag. Based on this, the embodiment of the present application provides a building industry industrial internet platform based on data elements and clustering algorithms, which is used to solve the technical problem that the existing technology cannot realize the differentiated and refined management of suppliers.
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0024] Please refer to Figures 1-2 The building industry industrial internet platform based on data elements and clustering algorithms provided by the present application comprises: a multi-source data integration module for obtaining and preprocessing multi-source data of suppliers to generate a space-time feature matrix; a multi-dimensional clustering module for selecting a clustering algorithm based on multi-dimensional evaluation of the space-time feature matrix, generating a clustering parameter configuration set in combination with an elbow rule and a silhouette coefficient method; a cluster group feature quantization module for clustering analysis of the suppliers by using the selected clustering algorithm based on the space-time feature matrix and the clustering parameter configuration set to generate a cluster group feature quantization report; a hierarchical strategy generation module for generating a supplier hierarchical label set by generating a hierarchical rule based on the cluster group feature quantization report, and generating a supplier hierarchical strategy execution instruction set in combination with a preset differentiated resource allocation rule; and a risk early warning module for constructing a risk index threshold rule set based on the supplier hierarchical strategy execution instruction set and the cluster group feature quantization report, and generating a high-risk supplier early warning list through a preset threshold comparison mechanism.
[0025] Multi-source data integration module: One of the functional modules of the construction industry internet platform, responsible for acquiring and preprocessing multi-source supplier data to ultimately generate a spatiotemporal feature matrix; Multi-source supplier data: Supplier-related data obtained from different channels (such as transaction systems, logistics platforms, qualification registration databases, etc.), including information on transactions, logistics, qualifications, performance, etc.; Preprocessing: Operations such as cleaning, deduplication, format standardization, and missing value imputation performed on the multi-source supplier data to improve data quality; Spatiotemporal feature matrix: A structured data matrix constructed with suppliers as rows and time and spatial features as columns, serving as the basic data carrier for subsequent clustering analysis; Multi-dimensional clustering module: One of the platform's functional modules, responsible for spatiotemporal clustering... The feature matrix undergoes multidimensional evaluation to select a clustering algorithm, combining the elbow rule and silhouette coefficient method to generate a clustering parameter configuration set. Multidimensional evaluation involves a comprehensive assessment of the spatiotemporal feature matrix from dimensions such as data size and data uniformity to match suitable clustering algorithms. Clustering algorithms are used to divide supplier data into different clusters based on feature similarity; this scheme involves K-means clustering, hierarchical clustering, and density-based clustering algorithms. The elbow rule is a method that determines the number of clusters corresponding to the abrupt change in the slope of a curve by plotting the relationship between the number of clusters and the sum of squared clustering errors, used to determine the number of first-choice clusters. The silhouette coefficient method calculates the intra-cluster compactness and inter-cluster separation of individual data points to obtain the silhouette coefficient. The system employs several methods to determine the number of clusters for the second candidate clusters. These include: a clustering parameter configuration set (a structured set integrating the selected clustering algorithm, target number of clusters, and corresponding core parameters to guide subsequent clustering operations); a cluster feature quantification module (responsible for performing clustering operations based on the spatiotemporal feature matrix and clustering parameter configuration set, generating a cluster feature quantification report); a cluster feature quantification report containing supplier cluster partitioning data, cluster dimension index statistics, and cluster difference data to support subsequent grading and risk warning); and a grading strategy generation module (responsible for generating supplier grading label sets and grading strategies based on the cluster feature quantification report). The platform includes: a strategy execution instruction set; a supplier tiering tag set (e.g., a set of tiering tags corresponding to each cluster, such as "A-level / Core Supplier" and "B-level / Ordinary Supplier"); preset differentiated resource allocation rules (resource allocation standards based on supplier tiering tags, set through cross-analysis of business and performance dimensions); a supplier tiering strategy execution instruction set (execution instructions integrating supplier tiering tags and resource allocation values to guide actual resource allocation); a risk warning module (one of the platform's functional modules, responsible for building a risk indicator threshold rule set and generating a high-risk supplier warning list through threshold comparison); and a risk indicator threshold rule set (a set of rules integrating static thresholds, dynamic threshold ranges, and corresponding risk levels of key risk indicators).Pre-set threshold comparison mechanism: a mechanism for comparing supplier real-time risk indicator data with a risk indicator threshold rule set for screening high-risk suppliers; high-risk supplier early warning list: a list marking high-risk supplier ID, cluster group label, classification label, and risk type for business side risk control.
[0026] In the embodiment of the present application, the multi-source data integration module, the obtained supplier multi-source data includes at least one of supplier basic information, supplier transaction and contract data, supplier cooperation and service data, supplier credit evaluation information, supplier financial data, supplier industry evaluation data, and supplier logistics data; the supplier multi-source data is subjected to data cleaning and standardization processing to complete data preprocessing; the preprocessed supplier multi-source data is subjected to feature extraction to obtain time dimension features and space dimension features, and the time dimension features and the space dimension features are used to construct a space-time feature matrix.
[0027] In the embodiment of the present application, the multi-dimensional clustering module, based on the space-time feature matrix, evaluates data size by a capacity estimation method, evaluates data uniformity by a density distribution histogram analysis method, and selects a clustering algorithm according to the evaluation results; the elbow rule is combined to calculate the clustering error sum of squares corresponding to different clustering numbers, a relationship curve of clustering number and clustering error sum of squares is drawn, and a target clustering number corresponding to a curve slope mutation point is selected; at the same time, the clustering number is verified by a silhouette coefficient method, and a target clustering number corresponding to a maximum average silhouette coefficient is output; the target clustering numbers obtained by the above two methods are integrated to generate a clustering parameter configuration set; the clustering parameter configuration set includes a target clustering number and a target clustering parameter, and the target clustering parameter includes but is not limited to K-means algorithm parameters and DBSCAN algorithm parameters.
[0028] In the embodiment of the present application, the cluster group feature quantification module, based on the space-time feature matrix and the clustering parameter configuration set, calculates the distance from each time dimension feature and space dimension feature in the space-time feature matrix to the cluster center by the Euclidean distance formula, and assigns each supplier to the nearest cluster group; the mean and standard deviation of the dimension indicators in each cluster group are counted, and the dimension indicators include at least one of supplier capacity, product quality, delivery timeliness, product price, and service quality; the differences of the dimension indicators between the cluster groups are compared horizontally, and a cluster group feature quantification report is generated by summarizing; the cluster group feature quantification report includes a cluster group label, a supplier number, and a dimension indicator statistical table.
[0029] In the embodiment of the present application, the grading strategy generation module, based on the cluster group feature quantification report, defines grading rules combined with historical performance data to generate a supplier classification label set; based on the supplier classification label set, a differentiated resource allocation proportion table is generated combined with differentiated resource allocation rules to clearly allocate the proportion of resources; the supplier classification label set and the differentiated resource allocation proportion table are integrated by a rule engine to generate a supplier classification strategy execution instruction set.
[0030] In the embodiment of the present application, the risk early warning module, based on the supplier grading strategy execution instruction set and the cluster feature quantification report, sets static threshold rules and dynamic threshold rules, generates a risk indicator threshold rule set; the risk indicators include at least one of product quality unqualified rate, product quality defect rate, and delivery delay rate; the risk indicators of each supplier are obtained, and according to the risk indicator threshold rule set, the suppliers whose risk indicators exceed the static threshold and / or the dynamic threshold are screened out through a threshold comparison mechanism, and a high-risk supplier candidate set marked with a risk type is output; the high-risk supplier candidate set and the supplier grading label set are integrated to generate a high-risk supplier early warning list.
[0031] It should be noted that the building industry Internet of Things platform of the present application obtains supplier multi-source data through a multi-source data integration module and generates a space-time feature matrix after preprocessing, relies on a multi-dimensional clustering module to perform multi-dimensional evaluation on the space-time feature matrix to select an adaptive clustering algorithm, simultaneously determines a reasonable clustering number and generates a clustering parameter configuration set by combining the elbow rule and the silhouette coefficient method, and then a cluster feature quantification module quantifies the dimension indicators of each cluster based on the space-time feature matrix and the clustering parameter configuration set, and utilizes the selected clustering algorithm to divide the suppliers into clusters and generate a cluster feature quantification report. Subsequently, a grading strategy generation module obtains a supplier grading label set based on the preset grading rules of the report, and generates a supplier grading strategy execution instruction set combined with the preset differentiated resource allocation rules. Finally, a risk early warning module constructs a risk indicator threshold rule set based on the above execution instruction set and the cluster feature quantification report, generates a high-risk supplier early warning list through a threshold comparison mechanism, and overall forms a supplier management technical solution based on data elements and clustering algorithms.
[0032] It is worth mentioning that, in order to solve the problem that supplier differentiation and fine management cannot be realized in the prior art, the present application firstly cleans, standardizes and processes the dispersed supplier multi-source data through a multi-source data integration module to build a space-time feature matrix, solves the pain points of non-uniform data basis and information fragmentation in traditional management, and provides standardized and comprehensive data support for fine management; then the multi-dimensional clustering module realizes accurate cluster division of the suppliers through scientific algorithm selection and cluster number determination method, avoids the problems of extensive supplier classification and fuzzy group characteristics in the existing management; and the cluster feature quantization module quantitatively analyzes and compares the mean and standard deviation of the supply capacity, product quality and other dimensional indexes of each cluster, so that the characteristics of different supplier groups are clear and distinguishable, and provides accurate feature basis for differentiated management; on this basis, the hierarchical strategy generation module generates hierarchical labels based on the quantitative features and matches the differentiated resource allocation proportion, solves the problems of homogeneous resource allocation and inability to adapt to different supplier values in the existing management; the risk early warning module combines the hierarchical and quantitative features to build risk rules, identifies high-risk suppliers in time, further guarantees the fine degree of management, and finally realizes the whole process of differentiated and fine management of suppliers from data integration, accurate clustering, feature quantization to hierarchical resource allocation and risk early warning through the synergistic effect of each module, effectively filling the gap in the prior art.
[0033] Please refer to Figure 3 The application provides a building industry industrial internet platform based on data elements and clustering algorithms, which obtains supplier multi-source data and pre-processes the data to generate a space-time feature matrix, including: obtaining supplier multi-source data; data preprocessing of the supplier multi-source data; feature extraction of the data preprocessed supplier multi-source data to obtain time dimension features and space dimension features; and constructing a space-time feature matrix using the time dimension features and the space dimension features.
[0034] Time dimension features: time-related features (such as cooperation period, historical bid frequency) extracted from the supplier multi-source data; space dimension features: space-related features (such as registered address longitude, logistics trajectory coordinates) extracted from the supplier multi-source data.
[0035] In the embodiments of the present application, the supplier multi-source data is acquired, and the supplier multi-source data includes at least one of supplier basic information, supplier transaction and contract data, supplier cooperation and service data, supplier credit evaluation information, supplier financial data, supplier industry evaluation data, and supplier logistics data; the application does not limit the acquisition method. In a possible implementation manner, the supplier basic information, the supplier transaction and contract data, the supplier financial data, and the supplier cooperation and service data can be acquired by connecting an enterprise internal system (such as a material centralized procurement management system, a supplier management database, a data center, a business center, etc.) through an API interface; the supplier credit evaluation information and the supplier industry evaluation data can be collected from an industry public data platform through a web crawler technology; and the supplier logistics data such as supplier logistics operation data and cooperation and service data can be acquired by integrating a third-party logistics system through a data interface.
[0036] The supplier multi-source data is subjected to data preprocessing, including data cleaning and standardization processing of the supplier multi-source data. Optionally, the data cleaning can include deleting the supplier basic information, the supplier transaction and contract data, and the supplier cooperation and service data of the suppliers that have entered a blacklist in the past year, deleting the supplier basic information, the supplier transaction and contract data, and the supplier cooperation and service data of the suppliers that have a bid-winning record but no actual bid-winning amount, removing duplicate data in the supplier multi-source data, correcting error data, and processing missing values (for a small amount of missing values, a mean value or a median value method is used for filling; for a large amount of missing value samples, it is determined whether to delete according to the actual situation), so as to ensure the accuracy and integrity of the data. The supplier basic information can include but is not limited to a supplier name, a registered address, a time of establishment, an enterprise size, a business scope, and supplier business information, etc. The supplier transaction and contract data can include but is not limited to supplier historical bid-winning data, contract years, contract types, and the supplier historical bid-winning data can include bid-winning times and bid-winning total amount, etc. The supplier cooperation and service data can include but is not limited to cooperation years, cooperation project quantity and size, cooperation satisfaction evaluation, after-sales service response time, problem solving efficiency, technical support capability, etc. Optionally, the standardization processing includes but is not limited to Z-score standardization, Max-Abs standardization, Robust standardization, and Min-Max standardization, etc. Through the standardization processing, data of different magnitudes and dimensions can be converted into data of a unified standard.
[0037] The time dimension features and the space dimension features are extracted from the preprocessed supplier multi-source data: based on the preprocessed supplier multi-source data, the space dimension features can be extracted from the registration address, longitude and latitude included in the supplier basic information and the distribution track coordinates included in the supplier logistics data, and the supplier regional concentration (for example, 780 suppliers in the Yangtze River Delta, a total of 1200 suppliers → 65%) and the logistics timeliness stability are calculated through the geographic distribution density algorithm and the standard deviation calculation method; based on the preprocessed supplier multi-source data, the time series basic indexes (including cooperation years and historical bid frequency) can be extracted from the cooperation years, historical bid frequency and service level trend data included in the supplier cooperation and service data, the sliding mean of the service level trend data is calculated through the moving average method, and the trend slope is extracted; meanwhile, based on the periodic fluctuation in the supplier cooperation and service data, the seasonal intensity coefficient is calculated through the exponential smoothing method, the trend slope, the seasonal intensity coefficient and the time series basic indexes are integrated, and the time dimension features are generated.
[0038] The space-time feature matrix is constructed by using the time dimension features and the space dimension features: the space-time feature matrix can be constructed by integrating the time dimension features and the space dimension features of the preprocessed supplier multi-source data, taking the suppliers (for example, supplier 1-ID-S1001, supplier 2-ID-S1002) as rows, and taking a plurality of time dimension feature indexes and / or space dimension feature indexes (for example, supplier registration address longitude, latitude, regional concentration, cooperation years, historical bid frequency, logistics timeliness variation coefficient) as columns, to construct the space-time feature matrix.
[0039] It should be noted that in the present embodiment, the comprehensive collection of multi-source data covers the full-dimensional data of the suppliers from basic information to operation, credit and logistics, solves the problem of one-sided information and insufficient dimension in traditional management, and provides a complete data basis for subsequent precise management; secondly, the data cleaning operation eliminates invalid and redundant data and repairs data defects, the standardized processing unifies the data format and standard, effectively improves the accuracy and consistency of the data, and avoids the deviation of management decision caused by poor data quality; thirdly, the extraction of space-time double-dimensional features breaks through the limitation of traditional single dimension, not only describes the spatial distribution characteristics of the suppliers, but also captures the time trend and periodic change of the cooperation and service, so that the group and individual characteristics of the suppliers are more stereoscopic and accurate; and the structured integration of the space-time feature matrix converts the scattered features into standardized and high-value algorithm input data, provides reliable support for the precise application of clustering algorithm, solves the core problem of insufficient data support in the prior art from the data bottom, and lays a key data foundation for realizing the differentiation and fine management of the suppliers.
[0040] Please refer to Figure 4The application provides a building industry industrial internet platform based on a data element and a clustering algorithm, a multi-dimensional evaluation selects a clustering algorithm based on a space-time feature matrix, elbow rule and contour coefficient method are combined, and a clustering parameter configuration set is generated, including: a capacity estimation algorithm is used to evaluate the data size of the space-time feature matrix; a density distribution histogram analysis method is used to evaluate the data uniformity of the space-time feature matrix; the clustering algorithm is selected according to the data size and the data uniformity; a plurality of different clustering numbers are selected from a preset to-be-evaluated clustering number interval; the elbow rule is used to determine a first candidate clustering number based on the clustering numbers; the contour coefficient method is used to determine a second candidate clustering number based on the clustering numbers; the target clustering number is determined based on the first candidate clustering number and the second candidate clustering number, and the clustering parameter configuration set is generated in combination with the selected clustering algorithm.
[0041] The capacity estimation algorithm is a method for calculating the data size of the space-time feature matrix, that is, "data size = supplier number * feature dimension (time + space dimension feature number)"; the data size is the volume of the space-time feature matrix, which is determined by the product of the supplier number and the feature dimension; the density distribution histogram analysis method is a method for evaluating the data uniformity by drawing a histogram of the feature data to statistically analyze the frequency distribution and combining the frequency standard deviation; the data uniformity is the degree of concentration of the feature data in the value range, which is usually measured by the frequency standard deviation; the preset to-be-evaluated clustering number interval is a preset clustering number range (such as 2-8), which is used to limit the evaluation range of the elbow rule and the contour coefficient method; the clustering number is the total number of supplier clusters to be divided; the first candidate clustering number is the candidate clustering number obtained by the elbow rule; the second candidate clustering number is the candidate clustering number obtained by the contour coefficient method; the target clustering number is the total number of supplier clusters finally determined by verifying the first and second candidate clustering numbers.
[0042] In the embodiment of the application, the capacity estimation algorithm is used to evaluate the data size. The "row" of the space-time feature matrix corresponds to the supplier number, and the "column" corresponds to the extracted space-time feature dimension, and the data size is "supplier number * space-time feature dimension"; for example, the space-time feature matrix contains 1000 suppliers (rows) and 20 space-time features (columns, including registered address longitude, cooperation time, etc.), and the data size is 1000*20=20000 points.
[0043] The density distribution histogram analysis method is used to evaluate the data uniformity. Select a core feature (such as "delivery timeliness") in the space-time feature matrix and operate according to the following steps: Divide the value range (such as 0-100 points) of the feature into 10 intervals (each interval is 10 points); Count the number of supplier data points in each interval, and calculate the "frequency" of each interval (interval points / total supplier number), for example, the frequencies of the 10 intervals are 0.08, 0.10, 0.12, 0.15, 0.18, 0.15, 0.12, 0.10, 0.06, 0.04, respectively; Calculate the standard deviation of the frequency (reflecting the degree of concentration / discreteness of the distribution): ; Where, is the standard deviation of the frequency, is the frequency of the th interval, is the average frequency, is the number of intervals.
[0044] In a specific implementation, the average frequency is 0.1 (the average frequency of the 10 intervals), the number of intervals is 10, and the final frequency standard deviation is 0.15, which is used to evaluate the uniformity of the distribution of the feature.
[0045] Repeat the above steps to count the frequency standard deviation of at least 85% of the spatio-temporal features as a basis for the overall data uniformity.
[0046] Select a clustering algorithm according to the data size and uniformity: If the data size is >10000 points, and the frequency standard deviation of more than 85% of the features is <0.2: select K-means algorithm; for example, the data size is 12000 points, and the frequency standard deviation of 90% of the features is 0.18, which meets the condition; If the data size is 5000-10000 points, and the business needs to generate a hierarchical structure according to "core supplier-primary supplier-secondary supplier": select hierarchical clustering algorithm; for example, the data size is 8000 points, and the supplier hierarchy needs to be distinguished, which meets the condition; If the data distribution is not uniform (the frequency standard deviation of any core feature is >0.3): select DBSCAN algorithm (density-based clustering algorithm); for example, the frequency standard deviation of the "logistics time" feature is 0.35, which meets the condition.
[0047] Select the number of clusters to be evaluated: Set the preset interval of the number of clusters to be evaluated as 2-8, and select 2, 3, 4, 5, 6, 7, and 8 as the number of clusters (i.e. K value).
[0048] Determine the first candidate number of clusters based on the elbow rule: Calculate the sum of squared errors (SSE) for each K value: ; Where, is the sum of squared errors, For the th cluster, the centroid (feature mean of all points within the cluster), the supplier data points within the cluster, and the target number of clusters.
[0049] For example: , =2500; , =1800; , =1200; , =880; , =820; , =790; , =770; Plot the "K value-SSE" relationship curve: the curve drops significantly from to , with a significant drop in SSE (from 2500→1200), and then slows down (from 1200→770), so the slope of the curve at the point of inflection ("elbow") corresponds to , which is the first candidate number of clusters.
[0050] Determine the second candidate number of clusters based on the silhouette coefficient method: Calculate the average silhouette coefficient for each K value: The silhouette coefficient formula for a single data point is ; where is the silhouette coefficient, is the average distance of the point to other points in the same cluster, is the average distance of the point to all points in the "nearest different cluster".
[0051] For example , select a supplier data point: the average distance of other points in the same cluster is , and the average distance of the nearest different cluster is , so the silhouette coefficient of the point is .
[0052] Calculate the silhouette coefficients of all supplier data points and take the average: The average silhouette coefficient is 0.45 when , and 0.52 when 0.62; 0.58; then the average profile coefficient maximum value corresponds to The second candidate cluster number is 4.
[0053] Determine the target cluster number and generate a cluster parameter configuration set: Verify consistency: the first candidate is consistent with the second candidate , so the target cluster number is 4; Combine the selected cluster algorithm supplementary parameters: If K-means algorithm is selected: parameters include "cluster number 4" and "iteration number 50 times (to avoid insufficient convergence)"; If DBSCAN algorithm is selected: determine the neighborhood radius eps (take ) through k-distance graph, calculate the distance from each point to the nearest 4 points, sort by distance and find the inflection point of the curve, for example, the distance at the inflection point is 0.5, which is eps; The minimum number of neighborhood points is 3 (corresponding to 3-dimensional space-time features); The final generated cluster parameter configuration set includes: cluster number 4, K-means iteration number 50 (or DBSCAN's eps=0.5, minimum neighborhood point number=3).
[0054] It should be noted that the technical scheme of "multi-dimensional evaluation selection clustering algorithm based on space-time feature matrix, combined with elbow rule and silhouette coefficient method to generate clustering parameter configuration set" in this embodiment, the core is to evaluate the data size of the space-time feature matrix through the capacity estimation method, and evaluate the data uniformity through the density distribution histogram analysis method, according to the evaluation results of the two and combined with the business demand, select the adaptive clustering algorithm--when the data size is in the first preset size interval and the data uniformity meets the first preset uniformity standard, select K-means clustering algorithm, when the data size is in the second preset size interval and the first preset business demand of building industry needs to generate clustering hierarchy, select hierarchical clustering algorithm, when the data uniformity meets the second preset uniformity standard, select the clustering algorithm based on density; Then select a plurality of clustering numbers from the preset to be evaluated clustering number interval, respectively determine the first candidate clustering number through the elbow rule, determine the second candidate clustering number through the silhouette coefficient method, finally determine the target clustering number combined with the two types of candidate clustering number, and generate the clustering parameter configuration set with the selected clustering algorithm.The scheme first realizes accurate adaptation of the clustering algorithm through multi-dimensional evaluation of data size, data uniformity and business demand, avoiding the problem of "algorithm and data characteristics, business scene mismatch" caused by the selection of clustering algorithm only by experience in traditional management: when the data size is large and uniformly distributed, K-means algorithm is selected to ensure clustering efficiency, when the data size is moderate and the construction industry needs to distinguish core suppliers, first-level suppliers and second-level suppliers, hierarchical clustering algorithm is selected to accurately adapt to the hierarchical management needs of the construction industry, solving the problem that traditional algorithms cannot generate clustering hierarchical structure that fits the business scene, when the data is not uniformly distributed, DBSCAN algorithm is selected to ensure clustering accuracy, ensuring that the clustering process can accurately capture the spatial and temporal feature differences of suppliers under different characteristics and different business needs; secondly, the elbow rule (based on the change trend of clustering error square sum) and the silhouette coefficient method (based on the comprehensive evaluation of the compactness within the cluster and the separation between clusters) are used to determine the target clustering number, which avoids the problem of unclear slope mutation point of single elbow rule, and makes up for the defect of single silhouette coefficient method that may ignore the adaptability of business scene, so that the determined target clustering number can not only ensure the high similarity of the characteristics of the suppliers within the cluster, but also ensure the significant difference of the characteristics of the suppliers between the clusters, providing a scientific basis for the subsequent accurate cluster division of suppliers; the finally generated clustering parameter configuration set integrates the adapted algorithm type and reasonable clustering number, as well as the core parameters of the corresponding algorithm, laying a key foundation for the subsequent cluster feature quantization module to realize accurate clustering analysis of suppliers, solving the problem of extensive supplier classification and fuzzy group characteristics in the prior art from the source of algorithm selection and parameter configuration, not only meeting the business needs of the construction industry for hierarchical management of suppliers, but also adapting to the clustering accuracy requirements under different data characteristics, providing a reliable clustering basis for subsequent differentiated resource allocation and accurate risk warning of suppliers, effectively supporting the differentiated and refined management of suppliers, and filling the gap in the prior art in terms of algorithm adaptability and parameter scientificity in supplier clustering management.
[0055] Please refer to Figure 4 The application provides an industrial internet platform for the construction industry based on data elements and clustering algorithms, which selects a clustering algorithm according to data size and data uniformity, comprising: when the data size is in a first preset size interval and the data uniformity meets a first preset uniformity standard, a K-means clustering algorithm is selected; when the data size is in a second preset size interval and a first preset business demand is met, a hierarchical clustering algorithm is selected, and the first preset business demand is that the construction industry business scene needs to generate a clustering hierarchical structure; and when the data uniformity meets a second preset uniformity standard, a density-based clustering algorithm is selected.
[0056] Hierarchical clustering algorithm: a clustering algorithm that generates hierarchical cluster relationships, outputs a tree structure, and adapts to hierarchical management needs; clustering hierarchy: a cluster structure generated by a hierarchical clustering algorithm with hierarchical relationships (such as the hierarchy of core supplier clusters and first-tier supplier clusters); second preset uniformity standard: a pre-set low uniformity judgment standard (such as a frequency standard deviation of any core feature > 0.3); density-based clustering algorithm: an algorithm that divides clusters based on data density (such as DBSCAN), which can identify non-uniformly distributed data clusters, and is suitable for discrete data.
[0057] In an embodiment of the present application, the first preset scale interval: data scale > 10000 points (data scale = number of suppliers x spatiotemporal feature dimension, such as 1000 suppliers x 20 spatiotemporal features = 20000 points, which belongs to this interval); The second preset scale interval: the data scale is between 5000 and 10000 points (such as 800 suppliers x 10 spatiotemporal features = 8000 points, which belongs to this interval).
[0058] The first preset uniformity standard: more than 85% of the features in the spatiotemporal feature matrix (such as supply capacity, delivery timeliness, etc.) have a frequency standard deviation < 0.2 (frequency standard deviation calculation method: divide a single feature into 10 intervals, calculate the frequency of each interval, and then calculate the standard deviation of the frequency by the above formula; The second preset uniformity standard: the frequency standard deviation of any 1 core feature (such as logistics timeliness, product quality pass rate) in the spatiotemporal feature matrix is > 0.3.
[0059] The first preset business demand landing definition: the construction industry business scenario needs to generate a hierarchical clustering result of "core supplier - first-tier supplier - second-tier supplier", which is used to support business such as hierarchical resource allocation and hierarchical cooperation management (such as core suppliers being preferentially allocated large procurement orders, and second-tier suppliers being included in a cultivation system).
[0060] Scenario 1: Select K-means clustering algorithm
[0061] Implementation steps: First step: Calculate the data scale of the spatiotemporal feature matrix (number of suppliers x spatiotemporal feature dimension) by the capacity estimation method; Second step: Calculate the frequency standard deviation of at least 85% of the spatiotemporal features by the density distribution histogram analysis method; Third step: Determine whether the "data scale is in the first preset scale interval" and "the frequency standard deviation of more than 85% of the features is < 0.2" are met simultaneously, and if so, select the K-means clustering algorithm.
[0062] Example: The spatio-temporal feature matrix of a construction enterprise's suppliers contains 1200 suppliers (rows), 15 spatio-temporal features (columns, including latitude of registered address, cooperation time, etc.), data size = 1200 x 15 = 18000 points (in the first preset size interval > 10000 points); 13 core features (accounting for 86.7%) are selected to calculate the frequency standard deviation, the results are all 0.15-0.19 (all < 0.2, meeting the first preset uniformity standard); therefore, the K-means clustering algorithm is selected, which is suitable for efficient clustering of large-scale uniform data, and can quickly complete the basic cluster division of suppliers.
[0063] Scenario 2: Select hierarchical clustering algorithm
[0064] Implementation steps: First step: Calculate the data size of the spatio-temporal feature matrix through capacity estimation method; Second step: Determine whether the data size is in the second preset size interval (5000-10000 points); Third step: Confirm whether the construction industry business side has the demand of "generating clustering hierarchy" (such as procurement department needs to distinguish core / primary / secondary supplier levels), if both data size condition and business demand are met, select hierarchical clustering algorithm.
[0065] Example: The spatio-temporal feature matrix of a construction enterprise's suppliers contains 800 suppliers, 10 spatio-temporal features, data size = 800 x 10 = 8000 points (in the second preset size interval 5000-10000 points); the business side proposes "need to generate hierarchy according to supplier cooperation priority to develop differentiated procurement strategy" (meeting the first preset business demand); therefore, the hierarchical clustering algorithm is selected, which can directly generate the hierarchical relationship of "core suppliers (cluster 1)-primary suppliers (cluster 2)-secondary suppliers (cluster 3)" through tree clustering results, without additional processing to adapt to business needs.
[0066] Scenario 3: Select density-based clustering algorithm (such as DBSCAN algorithm)
[0067] Implementation steps: First step: Through density distribution histogram analysis method, focus on core features such as logistics timeliness and product quality pass rate, calculate the frequency standard deviation; Second step: Determine whether it meets the "frequency standard deviation of any core feature > 0.3" (second preset uniformity standard), if it meets, select density-based clustering algorithm; Supplementary note: This scenario does not need to strictly limit the data size, whether it is in the first or second preset size interval, as long as the data distribution is uneven, it is suitable.
[0068] Example: The space-time feature matrix of a construction enterprise's suppliers contains 1100 suppliers, 12 space-time features, and the data size = 1100 x 12 = 13200 points (in the first preset scale interval); calculate the frequency standard deviation of the core feature "logistics time limit": divide the logistics time limit (0-20 days) into 10 intervals, and calculate the frequency standard deviation = 0.35 (> 0.3, meet the second preset uniformity standard) after counting the frequency of each interval, which indicates that the logistics time limit data distribution is extremely uneven (some suppliers have extremely fast logistics time limit, and some suppliers have large fluctuations in logistics time limit); Therefore, the density-based clustering algorithm is selected, which can ignore the influence of uneven data distribution, accurately identify the supplier cluster group with stable logistics time limit, and exclude the discrete supplier points with abnormal logistics time limit.
[0069] It should be noted that scenario 4: select k-center point algorithm (such as K-medoids algorithm)
[0070] When the data size, data uniformity and business requirements do not meet the selection conditions of K-means clustering algorithm, hierarchical clustering algorithm and density-based clustering algorithm, K-medoids algorithm is automatically selected. This algorithm has strong compatibility for small amount of data (<5000 points) and moderately uniform data (85% or more feature frequency standard deviation between 0.2-0.3), and does not depend on the business requirement of "generating clustering hierarchy", which can cover all blank scenarios; At the same time, it uses actual supplier data points as cluster centers, which has better anti-outlier ability than K-means algorithm, and can adapt to the scenario of a small number of extreme values in the construction industry supplier data.
[0071] Please refer to Figure 4 The building industry industrial internet platform based on data elements and clustering algorithm provided by the application, based on the number of clusters, adopts elbow rule to determine the first candidate cluster number, including: calculating the cluster error sum of squares corresponding to each cluster number; draw the relationship curve between each cluster number and the corresponding cluster error sum of squares; select the cluster number corresponding to the curve slope mutation point of each relationship curve as the first candidate cluster number.
[0072] Cluster error sum of squares (SSE): the sum of squares of the distance from all supplier data points in the cluster to the cluster center, used to measure the compactness of the cluster; Relationship curve: a two-dimensional curve drawn with "cluster number" as the horizontal coordinate and "cluster error sum of squares" as the vertical coordinate, used to reflect the change relationship between the two; Curve slope mutation point: the turning point (commonly known as "elbow") where the relationship curve changes from "large downward" to "small flat".
[0073] In the embodiment of the present application, in combination with the conventional demand of the building industry supplier cluster division, a reasonable to-be-evaluated clustering number interval is preset, a plurality of continuous clustering numbers (referred to as K values) are selected from the interval, and a basis is provided for subsequent comparative analysis. For each selected K value, the supplier data is first preliminarily clustered in the number; then the difference degree of all supplier data in each cluster and the cluster core (the mean value of all supplier characteristics in the cluster) is calculated; finally, the difference degrees of all clusters are summarized to obtain the SSE corresponding to the K value. The core significance of SSE is to measure the compactness in the cluster, and the smaller the value is, the higher the similarity of the supplier characteristics in the same cluster is. Taking the "clustering number K" as the horizontal coordinate and the "SSE corresponding to the K value" as the vertical coordinate, a two-dimensional relationship curve is drawn. The core change trend of the curve is "first rapid decline, then gradually flat", and the principle behind it is: when the initial K value is small, increasing the K value can significantly reduce the difference in the cluster, so the SSE decreases rapidly; when the K value reaches a reasonable range, increasing the K value has little effect on reducing the difference in the cluster, and the SSE decreases slowly. By observing the change rule of the curve, the SSE decrease amplitudes corresponding to adjacent two clustering numbers are calculated. When the decrease amplitude suddenly changes from "large decrease" to "small decrease", the turning point is the inflection point (commonly known as "elbow point") of the curve. The clustering number corresponding to the inflection point is determined as the first candidate clustering number. If the inflection point of the curve is not obvious, a SSE decrease amplitude threshold can be preset, and the first clustering number with a decrease amplitude less than the threshold is selected as the first candidate clustering number, to ensure that there is no omission in the selection process and to ensure the reliability of the result.
[0074] Please refer to Figure 4 The present application provides a building industry industrial internet platform based on data elements and clustering algorithm, which determines a second candidate clustering number based on the clustering number using the silhouette coefficient method, including: calculating a plurality of silhouette coefficients corresponding to each clustering number; determining the average silhouette coefficient corresponding to each clustering number according to the silhouette coefficient; selecting the clustering number corresponding to the maximum average silhouette coefficient as the second candidate clustering number.
[0075] Silhouette coefficient: an index for measuring the rationality of clustering of a single supplier data point, with a value range of [-1, 1], and the closer to 1, the better the clustering effect; average silhouette coefficient: the arithmetic mean of the silhouette coefficients of all supplier data points under a certain clustering number, used to comprehensively evaluate the effect of the clustering number.
[0076] In the embodiment of the present application, consistent with the elbow rule for determining the first candidate cluster number, a plurality of consecutive cluster numbers (referred to as K values) are selected from a preset cluster number interval to be evaluated, so as to ensure the uniformity of the comparison basis of the two types of candidate numbers and avoid the influence of the K value range difference on the subsequent determination of the target cluster number. For each selected K value, the supplier data is first divided into clusters according to the number; then for each supplier data point, two key indicators are calculated, i.e., the average distance of the point to all other data points in the same cluster (reflecting the compactness of the cluster) and the average distance of the point to all data points in the nearest different cluster (reflecting the separation degree between clusters); then the silhouette coefficient of the data point is calculated according to the two indicators, and the coefficient value range is [-1, 1]. The closer to 1, the better the clustering effect of the data point, and the closer to -1, the worse the clustering effect. For each K value, a plurality of silhouette coefficients corresponding to the number of supplier data points are obtained, and for each K value, the silhouette coefficients of all individual supplier data points corresponding thereto are arithmetically averaged to obtain the average silhouette coefficient under the K value. The average silhouette coefficient can comprehensively reflect the overall clustering quality under the corresponding K value, effectively avoiding the interference of single data point outliers on the evaluation result, and by comparing the average silhouette coefficients corresponding to all K values to be evaluated, the K value corresponding to the average silhouette coefficient with the largest value is selected as the second candidate cluster number. The maximum value of the average silhouette coefficient means that the cluster division under the K value achieves the optimal balance between the compactness of the cluster and the separation degree between clusters, and the overall clustering effect is the best. If there are multiple K values with the same maximum average silhouette coefficient, the smallest cluster number is selected as the second candidate cluster number to simplify the subsequent supplier cluster management process and reduce the operation cost of supplier classification management in the construction industry.
[0077] Please refer to Figure 4The application provides a building industry industrial internet platform based on data elements and a clustering algorithm, determines a target clustering quantity based on a first candidate clustering quantity and a second candidate clustering quantity, and generates a clustering parameter configuration set in combination with a selected clustering algorithm, and the method comprises the following steps: when the first candidate clustering quantity is consistent with the second candidate clustering quantity, the first candidate clustering quantity or the second candidate clustering quantity is determined as the target clustering quantity; when the first candidate clustering quantity is inconsistent with the second candidate clustering quantity, a second clustering operation is performed on the space-time feature matrix by using the second candidate clustering quantity; if the second clustering operation result meets a preset business scenario requirement, the second candidate clustering quantity is determined as the target clustering quantity; if the second clustering operation result does not meet the preset business scenario requirement, a first clustering operation is performed on the space-time feature matrix by using the first candidate clustering quantity; if the first clustering operation result meets the preset business scenario requirement, the first candidate clustering quantity is determined as the target clustering quantity; if the first clustering operation result does not meet the preset business scenario requirement, the first candidate clustering quantity is fine-tuned; the fine-tuned first candidate clustering quantity is used as a new first candidate clustering quantity, and the step of performing the first clustering operation on the space-time feature matrix by using the first candidate clustering quantity is executed again until the target clustering quantity is output; the selected clustering algorithm is used as a key to search a preset clustering parameter key-value library, and corresponding target clustering parameters are matched; and the target clustering quantity and the target clustering parameters are integrated to generate the clustering parameter configuration set.
[0078] The preset business scenario requirement is an actual business requirement of the building industry for cluster division (for example, the number of clusters is adapted to a hierarchical management architecture, and the similarity within a cluster is greater than or equal to 80%); the preset clustering parameter key-value library is a database for storing different clustering algorithms and core parameters (including default values and adjustable ranges) thereof, and is used for quickly matching algorithm parameters; and the target clustering parameter is a core parameter (for example, the number of iterations of K-means) matched from the preset clustering parameter key-value library and adapted to the selected clustering algorithm.
[0079] In the embodiment of the application, the consistency of the source of the first candidate clustering quantity (from the elbow rule) and the second candidate clustering quantity (from the silhouette coefficient method) is determined, both of which are based on the same set of to-be-evaluated clustering quantity intervals and the same space-time feature matrix data, so as to avoid judgment deviation caused by reference difference; in combination with the core demand of the building industry supplier management, the preset requirement standard is that the cluster division result can directly support hierarchical management (for example, the division of core / ordinary / cultivated suppliers), the similarity of the features of the suppliers within a cluster is greater than or equal to 80% (measured by the proportion of feature variance), and there is no obvious feature overlap between clusters (the difference between the average features of the clusters is greater than or equal to 30%); the type of the clustering algorithm is used as a key to associate the storage of the core parameters (including default values and adjustable ranges) adapted to the data features of the building industry, for example: Key: K-means clustering algorithm -> value: number of iterations (default 50 times, adjustable range 30-100 times), convergence threshold (default 0.001); Key: hierarchical clustering algorithm → value: distance measurement method (default Euclidean distance), linkage type (default ward method, adaptive hierarchical division); Key: DBSCAN algorithm → value: neighborhood radius eps (default 0.5, adjustable range 0.3-0.8), minimum neighborhood point number (default 3, value according to feature dimension x 1.5).
[0080] Scenario 1: The first and second candidate cluster numbers are consistent
[0081] The consistent cluster number is directly determined as the target cluster number. For example, the elbow method obtains the first candidate K = 4, and the silhouette coefficient method obtains the second candidate K = 4. Both are consistent, so the target cluster number is 4, no additional verification is needed, and the parameter matching link is directly entered.
[0082] Scenario 2: The first and second candidate cluster numbers are inconsistent
[0083] Priority verification of the second candidate cluster number: use the second candidate cluster number, combine the selected clustering algorithm, and perform clustering operation on the spatiotemporal feature matrix to obtain the second clustering result; Business requirement verification: determine whether the second clustering result meets the preset business scenario requirements - specifically verify "whether the cluster number is suitable for hierarchical management architecture" "whether the intra-cluster supplier feature variance ratio is ≤20% (i.e., similarity ≥80%)" "whether the inter-cluster feature mean difference is ≥30%"; if all conditions are met, the second candidate cluster number is determined as the target cluster number; Verification of the first candidate cluster number: if the second clustering result does not meet the requirements, use the first candidate cluster number, combine the selected clustering algorithm, and perform clustering operation on the spatiotemporal feature matrix to obtain the first clustering result; Second business verification: similarly verify whether the first clustering result meets the preset business requirements, and if it does, determine the first candidate cluster number as the target cluster number; Fine-tuning iteration (bottom-up mechanism): if the first clustering result still does not meet the requirements, fine-tune the first candidate cluster number - the fine-tuning rule is "±1 each time (to ensure reasonable cluster number), and the fine-tuning range is limited to [first candidate - 2, first candidate + 2] (to avoid deviating from the reasonable range)"; Loop verification: use the fine-tuned cluster number as the new first candidate cluster number, jump to the step of "perform clustering operation with the first candidate cluster number", and repeat the business verification process until the clustering result meets the preset business requirements, and finally output the cluster number as the target cluster number.
[0084] Parameter matching: use the selected clustering algorithm (such as K-means, hierarchical clustering, etc.) as the search key to query the pre-set clustering parameter key-value library, and match the corresponding target clustering parameters (the default values can be directly reused, and if there are special business requirements, they can be fine-tuned); Integration generation: integrate the determined target clustering number with the matched target clustering parameters to form a structured clustering parameter configuration set. The configuration set needs to be clearly marked with "algorithm type, target clustering number, core parameter name and value" to facilitate direct calling by the subsequent cluster feature quantization module.
[0085] If the fine-tuning iteration exceeds 5 times and still does not meet the requirements, trigger the pre-set reasonable clustering number interval (such as 2-8) to output the intermediate value as the temporary target clustering number, and manually fine-tune the parameters after combining with the business requirements; the pre-set clustering parameter key-value library supports dynamic updating, and can supplement the adapted parameter combinations according to the supplier data characteristics of different sub-scenes in the construction industry (such as house building and municipal engineering), to improve the universality of the configuration set.
[0086] Please refer to Figure 5 The application provides a kind of based on data element and clustering algorithm to build the construction industry industrial internet platform, based on space-time feature matrix and clustering parameter configuration set, using selected clustering algorithm to carry out clustering analysis to supplier, generate cluster feature quantization report, including: based on clustering parameter configuration set, corresponding cluster center is generated by selected clustering algorithm;Using selected clustering algorithm, the cluster distribution distance of supplier data point in space-time feature matrix to each cluster center is calculated;According to each cluster distribution distance, each supplier is distributed to the cluster group to which the cluster center corresponding to the minimum cluster distribution distance belongs, to generate supplier cluster group division data;The mean and standard deviation of pre-set dimension index in each cluster group are counted, and the pre-set dimension index includes at least one of supply capacity, product quality, delivery timeliness, product price and service quality;Horizontally compare the differences of pre-set dimension index between each cluster group, calculate the dimension mean difference and standard deviation difference of each cluster group, and generate cluster group difference data;Integrate supplier cluster group division data, pre-set dimension index statistical results of each cluster group, and cluster group difference data to generate cluster feature quantization report.
[0087] Supplier data point: the feature vector corresponding to a single supplier in the space-time feature matrix, containing the time and space dimension characteristics of the supplier;Cluster center: the feature mean point of the cluster group, representing the core feature of the cluster group;Cluster distribution distance: the distance (such as Euclidean distance) from the supplier data point to each cluster center, used to determine the cluster group to which the supplier belongs;Supplier cluster group division data: records the result data of each supplier belonging to the cluster group;Pre-set dimension index: evaluates the core dimension of the supplier, including supply capacity, product quality, delivery timeliness, product price, service quality, etc.;Cluster group difference data: the mean difference and standard deviation difference of pre-set dimension index between each cluster group, used to reflect the feature differentiation degree between cluster groups.
[0088] In the embodiments of the present application, core information is extracted from the cluster parameter configuration set: including the selected clustering algorithm, the optimal cluster number K value, the distance calculation method (the Euclidean distance is used by default, and the Manhattan distance, Mahalanobis distance, Hamming distance, etc. are also supported, and the distance selection is not limited) : The quantification rules of the preset dimension indicators are clearly defined: Supply capacity: quantitatively generated based on supplier logistics data (including delivery cycle, order response speed, and order fulfillment accuracy) ; Product quality: calculated based on supplier industry evaluation data (including product pass rate, defect rate, and return rate) ; Delivery timeliness: statistically generated according to supplier transaction and contract data (including on-time delivery rate = on-time delivery times / total delivery times, and logistics time efficiency stability) ; Meanwhile, product price and service quality are also included in the dimension indicators.
[0089] Based on the cluster parameter configuration set, the optimal cluster number K value is extracted, and K cluster centers are initialized (as the initial feature representatives of each cluster group).
[0090] Using the selected clustering algorithm, the distances from each time dimension feature and space dimension feature in the time-space feature matrix to each cluster center are calculated through the Euclidean distance formula (only as an example, other distance methods can also be selected). All combinations of supplier data points and cluster centers are traversed to complete the distance calculation of each supplier data point to all cluster centers.
[0091] According to the cluster assignment distance, each supplier is assigned to the cluster group to which the cluster center with the "closest distance" belongs. For example, the distances of a supplier S1001 to the cluster centers of the four cluster groups are 1.2, 2.8, 3.1, and 0.9, respectively, so it is assigned to the fourth cluster group. In this way, the assignment of all suppliers is completed, and the supplier cluster group division data is generated - the cluster group is formed by assigning the supplier data points in the time-space feature matrix to the nearest cluster center through the clustering algorithm, and is composed of a collection of suppliers with similar spatial location concentration, cooperation time, and logistics time efficiency stability.
[0092] Through statistical calculation methods, the mean (arithmetic mean of the corresponding indicators of all suppliers in the cluster group) and standard deviation of the preset dimension indicators such as supply capacity, product quality, and delivery timeliness are calculated for each cluster group. In the example, the mean of the product quality dimension "batch pass rate" of cluster group C2 is 95.5%, and the standard deviation is 3.2%.
[0093] Transverse comparison of the differences between the preset dimension indicators of each cluster group: calculate the cross-cluster group difference of the mean of each dimension, and compare the standard deviation of each cluster group, to generate a cluster difference value; summarize the full-dimension statistical results of all cluster groups and the cluster difference value, and then generate the characteristic description of each cluster group based on spatial distribution, time stability, and performance (qualitative induction of the overall characteristics of the cluster group).
[0094] Integrate the supplier cluster division data, the preset dimension indicator statistical results (mean + standard deviation) of each cluster group, and the cluster difference data to generate a cluster characteristic quantization report; the report includes: cluster label (identifies the classification of the supplier grouping, obtained by matching the cluster characteristics); supplier quantity; dimension indicator statistical table (records the number of suppliers in each cluster group and the mean and standard deviation of each dimension, generated by dimension indicator calculation and table structuring); cluster difference value, cluster characteristic description.
[0095] Please refer to Figure 5 The application provides a building industry industrial internet platform based on data elements and clustering algorithms, generates a supplier classification tag set through preset definition classification rules based on the cluster characteristic quantization report, and generates a supplier classification strategy execution instruction set in combination with preset differentiated resource allocation rules, including: determining the classification rule threshold corresponding to each preset dimension based on the preset dimension indicator statistical results in the cluster characteristic quantization report and in combination with the historical performance data of the suppliers; comparing the preset dimension indicator statistical results of each cluster group with the corresponding classification rule threshold, matching the corresponding classification label for each cluster group, and summarizing the classification labels of all cluster groups to generate a supplier classification tag set; based on the supplier cluster division data in the cluster characteristic quantization report, counting the number of suppliers corresponding to each cluster group; according to the corresponding relationship between the supplier classification tag set and the preset differentiated resource allocation rules, allocating a resource proportion interval to each classification label, combining the number of suppliers and the total resource amount of the preset business scenario to calculate the actual allocation value of resources of each cluster group, and generating a differentiated resource allocation proportion table; integrating the supplier classification tag set and the differentiated resource allocation proportion table to generate a supplier classification strategy execution instruction set.
[0096] Classification rule threshold: the classification threshold corresponding to the preset dimension indicator (such as delivery timeliness ≥ 90%); resource proportion interval: the resource allocation proportion range corresponding to the classification label (such as 70%-80% of the procurement amount for A-level suppliers); total resource amount of the preset business scenario: the total resource (such as total procurement amount, total cooperation quota) under the building industry business scenario; actual allocation value of resources: the actual amount of resources obtained by each cluster group / single supplier, calculated from the total resource amount, the resource proportion interval, and the number of suppliers; differentiated resource allocation proportion table: a table recording the cluster label, the number of suppliers, the resource proportion interval, and the actual allocation value of resources.
[0097] In the embodiments of the present application, the historical performance data defines: the time series records of the delivery on time rate, product pass rate, order fulfillment accuracy rate and other indicators accumulated by the supplier in past cooperation, which are generated through quantitative calculation and data cleaning.
[0098] The preset differentiated resource allocation rule defines: through cross analysis of business dimensions (purchase amount proportion, supply risk level) and performance dimensions (quality score, delivery capacity), the mapping standard of resource allocation proportion interval is set.
[0099] The rule engine defines: an automated decision system separating business rules from program code, based on logical expressions and data structures, defining the binding rules of hierarchical labels and resource allocation fields.
[0100] Based on the full-dimension statistical results in the cluster feature quantification report, combined with the historical performance data of the supplier, the hierarchical rule threshold value corresponding to each preset dimension (supply capacity, product quality, delivery timeliness, etc.) is determined: The dimension rule is determined by setting the index threshold condition for the index value of the full dimension; the index value is defined based on the historical performance data (the value range is [0, 1]); The hierarchical rule threshold value is determined by analyzing the mean and standard deviation distribution range in the cluster feature quantification report, and the value is based on the report; in the example, the delivery timeliness dimension sets "mean ≥ 90% and standard deviation ≤ 5%" as the hierarchical rule threshold value corresponding to "A level".
[0101] Compare the full-dimension statistical results of each cluster with the corresponding hierarchical rule threshold value, and execute hierarchical rule matching for each cluster: when the full-dimension statistical results meet the preset dimension rule index threshold value, generate the corresponding combined hierarchical label; aggregate the hierarchical labels of all clusters to form a set of supplier hierarchical labels.
[0102] Based on the supplier cluster division data in the cluster feature quantification report, the number of suppliers corresponding to each cluster is counted.
[0103] According to the corresponding relationship between the set of supplier hierarchical labels and the preset differentiated resource allocation rule, allocate a resource proportion interval for each cluster label; based on the number of suppliers in the cluster feature quantification report and the total resource amount of the preset business scenario, calculate the actual resource allocation value of each cluster according to "resource actual allocation value = total resource amount × resource proportion interval corresponding proportion / number of suppliers in the cluster" (example: single supplier purchase amount = 10 million × 75% / 300 = 2.5 million); integrate the cluster label, the number of suppliers, the resource proportion interval and the actual allocation value to generate a differentiated resource allocation proportion table.
[0104] Integrate the supplier hierarchical label set and the differentiated resource allocation proportion table through the rule engine: establish the binding relationship between the hierarchical label and the resource allocation field, and check the rule priority in the multi-label scenario; according to the allocation value of the differentiated resource allocation proportion table, combine the number of suppliers and the total resource quantity of the cluster characteristic quantization report to calculate the resource allocation value of a single supplier; and aggregate the supplier ID, cluster label, resource type and allocation value to generate a supplier hierarchical strategy execution instruction set.
[0105] Referring to Figure 5 The application provides a building industry industrial internet platform based on a data element and a clustering algorithm, a risk indicator threshold rule set is constructed based on a supplier hierarchical strategy execution instruction set and a cluster characteristic quantization report, and a high-risk supplier early warning list is generated through a preset threshold comparison mechanism, including: based on the supplier hierarchical strategy execution instruction set, the cluster characteristic quantization report and supplier historical performance data, a static threshold of a key risk indicator is determined, the key risk indicator being a risk category indicator corresponding to a preset dimension indicator; based on the standard deviation of each preset dimension indicator in the cluster characteristic quantization report, a dynamic threshold interval of the key risk indicator is determined; the static threshold and the dynamic threshold interval of the key risk indicator are integrated, and the corresponding risk level is marked to generate the risk indicator threshold rule set; real-time risk indicator data of the supplier is obtained, and according to the risk indicator threshold rule set, a supplier whose risk indicator exceeds the static threshold and / or the dynamic threshold interval is screened out through the preset threshold comparison mechanism, the corresponding risk type is marked, and a high-risk supplier candidate set is generated; the high-risk supplier candidate set is associated with the supplier hierarchical label set, the supplier ID, the cluster label and the hierarchical label information are supplemented, and the high-risk supplier early warning list is generated.
[0106] Key risk indicator: a risk category indicator corresponding to a preset dimension indicator, including product quality unqualified rate, product quality defect rate, delivery delay rate, etc.; static threshold: a fixed critical value of the risk indicator determined based on the supplier historical performance data through quantile statistical method; dynamic threshold interval: a floating critical interval (such as ±2 times the standard deviation) of the risk indicator determined based on the standard deviation of the dimension indicator in the cluster characteristic quantization report; real-time risk indicator data: real-time collected supplier risk category indicator data (such as real-time delivery delay rate); high-risk supplier candidate set: a preliminary set of suppliers whose risk indicators exceed the threshold selected through threshold comparison.
[0107] In the embodiment of the application, the key risk indicator definition: refers to the risk category indicator corresponding to the preset dimension indicator, at least including product quality unqualified rate, product quality defect rate, delivery delay rate.
[0108] Quantile statistical method definition: a statistical method for analyzing the distribution characteristics of risk indicators, which can calculate quantiles based on historical data to determine thresholds.
[0109] Composite rule definition: A rule expression that combines static and dynamic threshold conditions using the logical operator "OR" to mark risk levels.
[0110] Based on the supplier tiered strategy execution instruction set and cluster characteristic quantitative report, combined with the supplier's historical performance data, the distribution characteristics of key risk indicators (product quality non-conformance rate, product quality defect rate, and delivery delay rate) are analyzed using quantile statistics. Based on the quantiles calculated by quantile statistics, the static thresholds of each key risk indicator are determined, and the value range is set based on the quantiles of historical performance data.
[0111] Based on the standard deviation of each preset dimension indicator in the cluster feature quantification report, determine the dynamic threshold range of key risk indicators (e.g., "logistics timeliness exceeds ±2 times the standard deviation range").
[0112] Integrate the static and dynamic threshold ranges of key risk indicators, and use the logical operator "OR" to form a compound rule condition expression; bind the compound rule with the key risk indicator name, threshold type, threshold range, and risk level to generate a structured field combination (i.e., compound rule) - for example, "mark as 'high risk' when delivery delay rate > 5% or logistics timeliness exceeds ±2 standard deviations"; finally, generate a risk indicator threshold rule set containing indicator name, threshold type, threshold range, and risk level fields.
[0113] Obtain real-time risk indicator data from suppliers, and based on a set of risk indicator threshold rules, filter suppliers whose risk indicators exceed the static and / or dynamic threshold ranges using a preset threshold comparison mechanism: If a supplier's delivery delay rate data point exceeds the static threshold, it is marked as "high risk of delivery". If the real-time logistics timeliness stability exceeds the dynamic threshold boundary, it is marked as "high risk of logistics stability"; For suppliers that trigger both static and dynamic thresholds, a risk type label (such as "high risk in both delivery and logistics") is overlaid; at the same time, the supplier's measured value and the degree of exceeding the standard (such as "delay rate 7% vs threshold 5%) are recorded to generate a high-risk supplier candidate set labeled with risk type.
[0114] The high-risk supplier candidate set is associated with the supplier classification label set. Using the supplier ID as the key, the risk records of the high-risk supplier candidate set are associated with the cluster labels and classification labels of the classification label set to form a complete risk profile. The supplier name, the cluster to which it belongs, the classification label and the specific risk description are supplemented. The supplier is sorted by risk level (e.g., "double high risk" > "single high risk") and degree of exceeding the standard, and a high-risk supplier warning list is output.
[0115] To sum up, the application can select a clustering algorithm based on data size and distribution characteristics through the multi-dimensional clustering module, dynamically optimize clustering parameters by combining the elbow rule and the silhouette coefficient method, and eliminate the preset deviation, thereby significantly improving the accuracy and business adaptability of supplier classification. The risk early warning module integrates static threshold and dynamic threshold, constructs a composite risk rule set, monitors each index in real time, automatically marks high-risk suppliers when the data exceeds the threshold in the composite risk rule set, generates an early warning list, breaks through the fixed threshold response delay, and realizes precise and proactive prevention and control of supply chain risks.
[0116] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0117] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0118] The above description and the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A building industry industrial internet platform based on data elements and clustering algorithms, characterized in that, The method comprises the following steps: A multi-source data integration module is used to obtain and preprocess multi-source data of suppliers, and generate a space-time feature matrix; A multi-dimensional clustering module is used to select a clustering algorithm based on multi-dimensional evaluation of the space-time feature matrix, generate a cluster parameter configuration set by combining elbow rule and silhouette coefficient method; A cluster group feature quantization module is used to perform cluster analysis on the suppliers by using the selected clustering algorithm based on the space-time feature matrix and the cluster parameter configuration set, and generate a cluster group feature quantization report; A hierarchical strategy generation module is used to generate a supplier hierarchical label set by a pre-defined hierarchical rule based on the cluster group feature quantization report, and generate a supplier hierarchical strategy execution instruction set by combining a pre-defined differentiated resource allocation rule; A risk early warning module is used to construct a risk index threshold rule set based on the supplier hierarchical strategy execution instruction set and the cluster group feature quantization report, and generate a high-risk supplier early warning list through a pre-set threshold comparison mechanism. 2.The building industry industrial Internet platform based on the data element and the clustering algorithm according to claim 1, characterized in that, The multi-dimensional evaluation of the space-time feature matrix based on the clustering algorithm, combined with elbow rule and silhouette coefficient method, generates a cluster parameter configuration set, which comprises: The data size of the space-time feature matrix is evaluated by using a volume estimation method; The data uniformity of the space-time feature matrix is evaluated by using a density distribution histogram analysis method; The clustering algorithm is selected according to the data size and the data uniformity; A plurality of different cluster numbers are selected from a pre-set cluster number interval to be evaluated; The first candidate cluster number is determined by using elbow rule based on each cluster number; The second candidate cluster number is determined by using silhouette coefficient method based on each cluster number; The target cluster number is determined based on the first candidate cluster number and the second candidate cluster number, and a cluster parameter configuration set is generated by combining the selected clustering algorithm. 3.The building industry industrial internet platform based on data elements and clustering algorithm according to claim 2, characterized in that, The selection of the clustering algorithm according to the data size and the data uniformity comprises: When the data size is in a first pre-set size interval and the data uniformity meets a first pre-set uniformity standard, the K-means clustering algorithm is selected; When the data size is in a second pre-set size interval and a first pre-set business requirement is met, the hierarchical clustering algorithm is selected, and the first pre-set business requirement is that the building industry business scenario needs to generate a clustering hierarchy; When the data uniformity meets a second pre-set uniformity standard, the density-based clustering algorithm is selected. 4.The building industry industrial internet platform based on data elements and clustering algorithm according to claim 2, characterized in that, The determination of the first candidate cluster number by using elbow rule based on each cluster number comprises: The cluster error sum of squares corresponding to each cluster number is calculated; The relationship curve between each cluster number and the corresponding cluster error sum of squares is drawn; The cluster number corresponding to the curve slope mutation point of each relationship curve is selected as the first candidate cluster number. 5.The building industry industrial internet platform based on data elements and clustering algorithm according to claim 2, characterized in that, The determination of the second candidate cluster number by using silhouette coefficient method based on each cluster number comprises: A plurality of silhouette coefficients corresponding to each cluster number are calculated; The average silhouette coefficient corresponding to each cluster number is determined according to the silhouette coefficients; The cluster number corresponding to the maximum value of the average silhouette coefficient is selected as the second candidate cluster number. 6.The building industry industrial Internet platform based on data elements and clustering algorithm according to claim 2, characterized in that, The target cluster quantity is determined based on the first candidate cluster quantity and the second candidate cluster quantity, and a cluster parameter configuration set is generated in combination with the selected cluster algorithm, including: When the first candidate cluster quantity is consistent with the second candidate cluster quantity, the first candidate cluster quantity or the second candidate cluster quantity is determined as the target cluster quantity; When the first candidate cluster quantity is inconsistent with the second candidate cluster quantity, the second candidate cluster quantity is used to perform a clustering operation on the spatio-temporal feature matrix; If the second clustering operation result meets the preset business scenario requirement, the second candidate cluster quantity is determined as the target cluster quantity; If the second clustering operation result does not meet the preset business scenario requirement, the first candidate cluster quantity is used to perform a clustering operation on the spatio-temporal feature matrix; If the first clustering operation result meets the preset business scenario requirement, the first candidate cluster quantity is determined as the target cluster quantity; If the first clustering operation result does not meet the preset business scenario requirement, the first candidate cluster quantity is fine-tuned; The fine-tuned first candidate cluster quantity is used as a new first candidate cluster quantity, and the step of performing a clustering operation on the spatio-temporal feature matrix using the first candidate cluster quantity is executed until a target cluster quantity is output; The selected cluster algorithm is used as a key to search a preset cluster parameter key-value library, and a corresponding target cluster parameter is matched; The target cluster quantity and the target cluster parameter are integrated to generate a cluster parameter configuration set. 7.The building industry industrial Internet platform based on data elements and clustering algorithm according to claim 1, characterized in that, Based on the spatio-temporal feature matrix and the cluster parameter configuration set, the selected cluster algorithm is used to perform a clustering analysis on the suppliers to generate a cluster group feature quantization report, including: Based on the cluster parameter configuration set, a corresponding cluster center is generated by the selected cluster algorithm; Using the selected cluster algorithm, the cluster assignment distance of the supplier data points in the spatio-temporal feature matrix to each cluster center is calculated; According to each cluster assignment distance, each supplier is assigned to a cluster group to which the cluster center corresponding to the minimum cluster assignment distance belongs, and supplier cluster group division data is generated; The mean and standard deviation of a preset dimension index in each cluster group are calculated, and the preset dimension index includes at least one of supplier capacity, product quality, delivery timeliness, product price, and service quality; The differences in the preset dimension index between each cluster group are compared horizontally, and the dimension mean difference and standard deviation difference of each cluster group are calculated to generate cluster group difference data; The supplier cluster group division data, the preset dimension index statistical result of each cluster group, and the cluster group difference data are integrated to generate a cluster group feature quantization report. 8.The building industry industrial Internet platform based on the data element and the clustering algorithm according to claim 1, characterized in that, Based on the cluster group feature quantization report, a supplier classification label set is generated by a preset definition classification rule, and a supplier classification strategy execution instruction set is generated in combination with a preset differentiated resource allocation rule, including: Based on the preset dimension index statistical result in the cluster group feature quantization report, in combination with the supplier historical performance data, the classification rule threshold corresponding to each preset dimension is determined; The preset dimension index statistical result of each cluster group is compared with the corresponding hierarchical rule threshold value, a corresponding hierarchical label is matched for each cluster group, and the hierarchical labels of all the cluster groups are summarized to generate a supplier hierarchical label set; Based on the supplier cluster division data in the cluster group feature quantization report, the number of suppliers corresponding to each cluster group is counted; According to the corresponding relationship between the supplier hierarchical label set and the preset differentiated resource allocation rule, a resource proportion interval is allocated to each hierarchical label, and the resource actual allocation value of each cluster group is calculated by combining the number of suppliers and the total resource amount of the preset business scenario, to generate a differentiated resource allocation proportion table; The supplier hierarchical label set and the differentiated resource allocation proportion table are integrated to generate a supplier hierarchical strategy execution instruction set. 9.The building industry industrial Internet platform based on the data element and the clustering algorithm according to claim 1, characterized in that, The risk indicator threshold value rule set is constructed based on the supplier hierarchical strategy execution instruction set and the cluster group feature quantization report, and a high-risk supplier early warning list is generated through a preset threshold value comparison mechanism, including: Based on the supplier hierarchical strategy execution instruction set, the cluster group feature quantization report, and supplier historical performance data, the static threshold value of a key risk indicator is determined, and the key risk indicator is a risk class indicator corresponding to a preset dimension indicator; Based on the standard deviation of each preset dimension indicator in the cluster group feature quantization report, a dynamic threshold value interval of the key risk indicator is determined; The static threshold value and the dynamic threshold value interval of the key risk indicator are integrated, and the corresponding risk level is marked to generate a risk indicator threshold value rule set; Real-time risk indicator data of the supplier is obtained, and a supplier whose risk indicator exceeds the static threshold value and / or the dynamic threshold value interval is screened out through a preset threshold value comparison mechanism according to the risk indicator threshold value rule set, the corresponding risk type is marked, and a high-risk supplier candidate set is generated; The high-risk supplier candidate set is associated with the supplier hierarchical label set, and the supplier ID, cluster label, and hierarchical label information are supplemented to generate a high-risk supplier early warning list. 10.The building industry industrial internet platform based on the data element and clustering algorithm according to any one of claims 1-9, characterized in that, The supplier multi-source data is obtained and preprocessed to generate a space-time feature matrix, including: Supplier multi-source data is obtained; The supplier multi-source data is preprocessed; Feature extraction is performed on the preprocessed supplier multi-source data to obtain time dimension features and space dimension features; The time dimension features and the space dimension features are used to construct a space-time feature matrix.