Enterprise digital management method and system
By constructing the initial data set in enterprise digital management and performing clustering and screening, calculating the importance of comprehensive exception indicators and targets, and performing incremental clustering and interpolation processing, the cold start problem in collaborative filtering recommendations is solved, and the rationality of the recommendation results is improved.
Patent Information
- Application Number
- CN202510208229.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-25
AI Technical Summary
In enterprise digital management, cold start problems often exist when recommending information based on collaborative filtering recommendation algorithms, resulting in poor rationality of recommendation results.
By obtaining the dimension data of each customer to be recommended in each preset dimension, building an initial data set, and clustering to determine the distribution change weight, filtering the target dimension. Then, the comprehensive anomaly indicators and target importance of each customer to be recommended are calculated, the target customer group is filtered, and the interpolation weight indicators are performed to determine the interpolation weight indicators. Finally, the data set is expanded through interpolation processing and the recommendations are collaboratively filtered.
The data set of information recommendation was expanded, the cold start problem in the collaborative filtering recommendation process was solved, and the rationality of the recommendation results were improved.
Smart Images

Figure CN120123604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital management, and particularly relates to an enterprise digital management method and system. Background Art
[0002] With the rapid development of information technology and the promotion of enterprise digital transformation, more and more enterprises begin to adopt digital management methods or systems to improve operational efficiency, optimize decision-making, and enhance customer satisfaction. In enterprise digital management, the CRM (Customer Relationship Management) system is widely used in aspects such as customer information management, sales management, and marketing. However, in large-scale customer data, how to accurately and personalized provide products or services to customers, that is, how to accurately recommend information to customers, is an important challenge. Currently, when making information recommendations, the commonly used method is: collecting data of multiple customers, and based on the data of all collected customers, through a collaborative filtering recommendation algorithm, making information recommendations for each customer. Among them, information recommendations can include but are not limited to: business recommendations and item recommendations.
[0003] However, when making information recommendations for each customer through a collaborative filtering recommendation algorithm based on the collected customer data, the following technical problems often exist:
[0004] The relevant information regarding newly added services or newly added projects in the collected customer data may be relatively scarce, which may lead to the cold start problem in the collaborative filtering recommendation process, and further lead to incorrect recommendation results, resulting in poor rationality of the recommendation results. Summary of the Invention
[0005] In order to solve the technical problem of poor rationality of the recommendation results, the present invention proposes an enterprise digital management method and system.
[0006] In a first aspect, the present invention provides an enterprise digital management method, which includes:
[0007] Obtaining the dimension data of each customer to be recommended under each preset dimension, and constructing an initial data set;
[0008] Clustering all customers to be recommended according to the dimension data of all customers to be recommended under each preset dimension, and obtaining an initial clustering cluster set corresponding to each preset dimension;
[0009] Determining the distribution change weight corresponding to each preset dimension according to the initial clustering cluster set corresponding to each preset dimension, and screening out the preset dimensions with distribution change weights greater than a preset distribution change threshold as target dimensions;
[0010] Determine the anomaly score corresponding to the data in each dimension, and determine the comprehensive anomaly index corresponding to each customer to be recommended according to the anomaly scores corresponding to the dimension data of each customer to be recommended under all preset dimensions;
[0011] Determine the target importance level corresponding to each customer to be recommended according to the comprehensive anomaly index corresponding to each customer to be recommended and the set of initial clustering clusters corresponding to all target dimensions;
[0012] Screen out the set of target customer groups from all customers to be recommended according to the target importance level;
[0013] Perform incremental clustering on the set of target customer groups according to the dimension data of all target customers in the set of target customer groups under all preset dimensions, and determine the interpolation weight index corresponding to each target customer according to the incremental clustering result;
[0014] Screen out the customers to be interpolated from the set of target customer groups according to the interpolation weight index, and perform interpolation processing according to the dimension data of all customers to be interpolated under all preset dimensions to obtain an interpolation data set;
[0015] Perform information recommendation for each customer to be recommended through a collaborative filtering recommendation algorithm based on the union of the initial data set and the interpolation data set.
[0016] Optionally, the formula for the distribution change weight corresponding to the preset dimension is:
[0017] where α i is the distribution change weight corresponding to the i-th preset dimension; i is the serial number of the preset dimension; norm() is a normalization function; d i is the mean value of the absolute values of the differences between the dimension data of the cluster centers of all initial clusters in the set of initial clusters corresponding to the i-th preset dimension under the i-th preset dimension; d i,max is the maximum value of the absolute values of the differences between the dimension data of the cluster centers of all initial clusters in the set of initial clusters corresponding to the i-th preset dimension under the i-th preset dimension; N i is the number of initial clusters in the set of initial clusters corresponding to the i-th preset dimension; j is the serial number of the initial cluster in the set of initial clusters corresponding to the i-th preset dimension; N ij is the number of customers to be recommended in the j-th initial cluster in the set of initial clusters corresponding to the i-th preset dimension; N is the total number of customers to be recommended; exp() is an exponential function with the natural constant as the base; is the variance of the target differences corresponding to all initial clusters in the set of initial clusters corresponding to the i-th preset dimension; d ijis the target difference corresponding to the j-th initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension; d ij1 is the reference difference of the j-th initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension under the i-th preset dimension; the reference difference of the j-th initial clustering cluster under the i-th preset dimension is the mean value of the absolute values of the differences between the dimension data of all recommended customers in the j-th initial clustering cluster under the i-th preset dimension; d ij2 is the mean value of the reference differences of the j-th initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension under all preset dimensions except the i-th preset dimension.
[0018] Optionally, determining the anomaly score corresponding to each dimension data, and determining the comprehensive anomaly index corresponding to each recommended customer according to the anomaly scores corresponding to the dimension data of each recommended customer under all preset dimensions, includes:
[0019] Determining the anomaly score corresponding to each dimension data through the isolation forest algorithm according to the initial data set;
[0020] Determining the mean value of the anomaly scores corresponding to the dimension data of the recommended customer under all preset dimensions as the comprehensive anomaly index corresponding to the recommended customer.
[0021] Optionally, the formula for the target importance corresponding to the recommended customer is:
[0022] where τ k is the target importance corresponding to the k-th recommended customer; k is the serial number of the recommended customer; exp() is the exponential function with the natural constant as the base; F k is the comprehensive anomaly index corresponding to the k-th recommended customer; n is the number of target dimensions; norm() is the normalization function; m is the serial number of the target dimension; v km is the mean value of the absolute values of the differences between the dimension data of the clustering center of the initial clustering cluster to which the k-th recommended customer belongs in the initial clustering cluster set corresponding to the n-th target dimension under the n-th target dimension and the dimension data of the clustering centers of all initial clustering clusters other than the initial clustering cluster to which the k-th recommended customer belongs in the initial clustering cluster set corresponding to the n-th target dimension under the n-th target dimension; || is the absolute value function; A km1 is the dimension data of the k-th recommended customer under the n-th target dimension; A km2 is the dimension data of the clustering center of the initial clustering cluster to which the k-th recommended customer belongs in the initial clustering cluster set corresponding to the n-th target dimension under the n-th target dimension.
[0023] Optionally, screening out a set of target customer groups from all customers to be recommended according to the target importance level, including:
[0024] Randomly forming a group of customers to be recommended with every first preset number of customers to be recommended;
[0025] Screening out the second preset number of customers to be recommended with the greatest target importance level from each group of customers to be recommended as target customers, and forming target customer groups, where the second preset number is less than the first preset number;
[0026] Combining all target customer groups to form a set of target customer groups.
[0027] Optionally, performing incremental clustering on the set of target customer groups according to the dimensional data of all target customers in the set of target customer groups under all preset dimensions, and determining the interpolation weight index corresponding to each target customer according to the incremental clustering result, including:
[0028] Randomly screening out a target customer group from the set of target customer groups as the initial clustering data set for incremental clustering;
[0029] Performing clustering on the initial clustering data set according to the dimensional data of all target customers in the initial clustering data set under all preset dimensions, and determining the result of clustering the initial clustering data set as the first clustering result in the incremental clustering process;
[0030] Determining each target customer group in the set of target customer groups except the initial clustering data set as the data set for each increment in the incremental clustering process;
[0031] According to the dimensional data of all target customers in the data set for each increment under all preset dimensions, adding the data set for each increment to the previous clustering result in the incremental clustering process, and determining the clustering result after addition as another clustering result in the incremental clustering process, where the number of clustering clusters in each clustering result in the incremental clustering process is the same;
[0032] Matching the clustering clusters in every two adjacent clustering results in the incremental clustering process to obtain a set of matching clustering cluster groups between every two adjacent clustering results in the incremental clustering process;
[0033] Determining the interpolation weight index corresponding to each target customer according to all the matching clustering clusters in the set of matching clustering cluster groups between all adjacent two clustering results in the incremental clustering process.
[0034] Optionally, the matching of the clustering clusters in every two adjacent clustering results in the incremental clustering process to obtain a set of matching clustering cluster groups between every two adjacent clustering results in the incremental clustering process includes:
[0035] Determine any two adjacent clustering results in the incremental clustering process as the first candidate clustering result and the second candidate clustering result respectively;
[0036] Determine each clustering cluster in the first candidate clustering result as the first candidate clustering cluster, and determine each clustering cluster in the second candidate clustering result as the second candidate clustering cluster;
[0037] Determine the target edge weight value between each first candidate clustering cluster and each second candidate clustering cluster according to the intersection-over-union ratio and cosine similarity between each first candidate clustering cluster and each second candidate clustering cluster, where both the intersection-over-union ratio and the cosine similarity are positively correlated with the target edge weight value;
[0038] According to the target edge weight values between all first candidate clustering clusters and all second candidate clustering clusters, use the KM algorithm to match all first candidate clustering clusters and all second candidate clustering clusters, and form a matching clustering cluster group from the mutually matched first candidate clustering clusters and second candidate clustering clusters, to obtain the set of matching clustering cluster groups between the first candidate clustering result and the second candidate clustering result.
[0039] Optionally, the formula for the interpolation weight index corresponding to the target customer is:
[0040]
[0041] where, δ q is the interpolation weight index corresponding to the q-th target customer; q is the serial number of the target customer; R is the number of clustering results in the incremental clustering process; n q is the minimum value among the serial numbers of all clustering results in which the q-th target customer exists in the incremental clustering process; b is the serial number of the clustering result in the incremental clustering process; δ q,b,b+1 is the sub-weight index of the q-th target customer between the b-th clustering result and the (b + 1)-th clustering result in the incremental clustering process; exp() is the exponential function with the natural constant as the base; || is the absolute value function; g q,b is the cosine similarity between the dimensional data of the clustering center of the cluster to which the q-th target customer belongs in the b-th clustering result in the incremental clustering process in all preset dimensions and the dimensional data of the q-th target customer in all preset dimensions; g q,b+1 is the cosine similarity between the dimensional data of the clustering center of the cluster to which the q-th target customer belongs in the (b + 1)-th clustering result in the incremental clustering process in all preset dimensions and the dimensional data of the q-th target customer in all preset dimensions; NMI() is the normalized mutual information function; Q q,bis the cluster to which the q-th target customer belongs in the b-th clustering result during the incremental clustering process; Q q,b+1 is the cluster to which the q-th target customer belongs in the (b + 1)-th clustering result during the incremental clustering process; if the clusters to which the q-th target customer belongs in the b-th and (b + 1)-th clustering results during the incremental clustering process belong to the same matching cluster group, then set f q,b,b+1 to be equal to the first preset value; if the clusters to which the q-th target customer belongs in the b-th and (b + 1)-th clustering results during the incremental clustering process do not belong to the same matching cluster group, then set f q,b,b+1 to be equal to the second preset value; the first preset value is greater than the second preset value.
[0042] Optionally, the screening of the customers to be interpolated from the set of target customer groups according to the interpolation weight index includes:
[0043] Screen out the target customers with the largest preset proportion of the interpolation weight index from the set of target customer groups as the customers to be interpolated.
[0044] In a second aspect, the present invention provides an enterprise digital management system, including a processor and a memory, and the processor is used to process the instructions stored in the memory to implement the above-mentioned enterprise digital management method.
[0045] The present invention has the following beneficial effects:
[0046] An enterprise digital management method of the present invention expands the dataset for information recommendation, solves the cold start problem in the collaborative filtering recommendation process to a certain extent, and improves the rationality of the recommendation result. First, dimension data of each customer to be recommended under each preset dimension is obtained, and an initial dataset to be expanded can be constructed. Next, the obtained set of initial clustering clusters corresponding to each preset dimension can reflect the data distribution of the customers to be recommended under this preset dimension to a certain extent. Then, based on the set of initial clustering clusters corresponding to the preset dimension, the greater the quantified distribution change weight corresponding to the preset dimension, the greater the amount of effective information under this preset dimension, and the more important this preset dimension is; the target dimension is often a relatively important preset dimension. Continuing, based on the anomaly scores corresponding to the dimension data of the customers to be recommended under all preset dimensions, the smaller the quantified comprehensive anomaly index corresponding to the customer to be recommended, the greater the amount of effective information contained in the dimension data of this customer to be recommended under all preset dimensions, and the more this customer to be recommended can be used for subsequent interpolation. Secondly, based on the comprehensive anomaly index corresponding to the customer to be recommended and the set of initial clustering clusters corresponding to all target dimensions, the greater the quantified target importance corresponding to the customer to be recommended, the more important this customer to be recommended is for subsequent interpolation processing, and the more this customer to be recommended can be used for subsequent interpolation processing. Furthermore, based on the target importance, the target customers in the set of target customer groups selected can be relatively important customers. After that, based on the dimension data of all target customers in the set of target customer groups under all preset dimensions, incremental clustering is performed on the set of target customer groups, and according to the incremental clustering result, the greater the quantified interpolation weight index corresponding to the target customer, the more important this target customer is for subsequent interpolation processing, and the more this target customer can be used for subsequent interpolation processing. Then, based on the interpolation weight index, the customers to be interpolated are selected from the set of target customer groups, and interpolation processing is performed according to the dimension data of all customers to be interpolated under all preset dimensions. The obtained interpolated dataset can be an adaptively newly added dataset. Finally, based on the union of the initial dataset and the interpolated dataset, through the collaborative filtering recommendation algorithm, information recommendation for each customer to be recommended is realized. Compared with directly using the initial dataset for collaborative filtering recommendation, the present invention realizes the expansion of the dataset, enriches the data in the dataset, solves the cold start problem in the collaborative filtering recommendation process to a certain extent, and improves the rationality of the recommendation result. Description of the Drawings
[0047] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 It is a flowchart of a method for enterprise digital management of the present invention. Specific embodiments
[0049] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail the specific embodiments, structures, features, and effects of the technical solutions proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0051] The present invention provides a method for enterprise digital management, and the method includes the following steps:
[0052] Obtain the dimension data of each customer to be recommended under each preset dimension, and construct an initial data set;
[0053] Cluster all customers to be recommended according to the dimension data of all customers to be recommended under each preset dimension, and obtain an initial cluster set corresponding to each preset dimension;
[0054] According to the initial cluster set corresponding to each preset dimension, determine the distribution change weight corresponding to each preset dimension, and screen out the preset dimensions with distribution change weights greater than the preset distribution change threshold as target dimensions;
[0055] Determine the anomaly score corresponding to each dimension data, and determine the comprehensive anomaly index corresponding to each customer to be recommended according to the anomaly scores corresponding to the dimension data of each customer to be recommended under all preset dimensions;
[0056] Determine the target importance level corresponding to each customer to be recommended according to the comprehensive anomaly index corresponding to each customer to be recommended and the initial cluster set corresponding to all target dimensions;
[0057] Screen out the set of target customer groups from all customers to be recommended according to the target importance level;
[0058] Based on the dimensional data of all target customers in the set of target customer groups under all preset dimensions, perform incremental clustering on the set of target customer groups, and determine the interpolation weight index corresponding to each target customer according to the incremental clustering result;
[0059] According to the interpolation weight index, screen out the customers to be interpolated from the set of target customer groups, and perform interpolation processing based on the dimensional data of all customers to be interpolated under all preset dimensions to obtain an interpolation data set;
[0060] According to the union of the initial data set and the interpolation data set, perform information recommendation for each customer to be recommended through a collaborative filtering recommendation algorithm.
[0061] The following expands each of the above steps in detail:
[0062] Reference Figure 1 , which shows the flow of some embodiments of an enterprise digital management method of the present invention. The enterprise digital management method includes the following steps:
[0063] Step S1, obtain the dimensional data of each customer to be recommended under each preset dimension, and construct an initial data set.
[0064] In some embodiments, the dimensional data of each customer to be recommended under each preset dimension can be obtained to construct an initial data set.
[0065] Among them, the customer to be recommended can be a customer to whom information recommendation is to be performed, that is, a user to whom information recommendation is to be performed. Information recommendation can be, but is not limited to: business recommendation and product recommendation. The preset dimension can be a dimension related to information recommendation set in advance. The dimensional data can be the normalized preset dimension value. For example, if the customer to be recommended is a customer to whom product recommendation is to be performed, a preset dimension can be the dimension where the number of times the customer to be recommended purchases a certain product is located, and the dimensional data of the customer to be recommended under this preset dimension can be: the normalized value of the number of times the customer to be recommended purchases this product. A customer to be recommended can have one dimensional data under a preset dimension.
[0066] It should be noted that by obtaining the dimensional data of each customer to be recommended under each preset dimension, an initial data set to be expanded can be constructed.
[0067] As an example, through an enterprise management system, such as an ERP (Enterprise Resource Planning) system, a CRM (Customer Relationship Management) system, a financial system, etc., the dimension data of each customer to be recommended under each preset dimension can be obtained, and the dimension data of all customers to be recommended under all preset dimensions are combined into an initial data set.
[0068] It should be noted that the CRM system is a customer relationship management system, which can facilitate accurately understanding the needs and preferences of customers, and then recommending new services to customers. The data in the CRM system can include: basic information of customers, historical services of customers (such as interaction records and purchase records), and customer feedback, etc.
[0069] Step S2: Cluster all customers to be recommended according to the dimension data of all customers to be recommended under each preset dimension, and obtain an initial cluster set corresponding to each preset dimension.
[0070] In some embodiments, all customers to be recommended can be clustered according to the dimension data of all customers to be recommended under each preset dimension, and an initial cluster set corresponding to each preset dimension is obtained.
[0071] It should be noted that based on the dimension data of all customers to be recommended under each preset dimension, clustering all customers to be recommended, and the obtained initial cluster set corresponding to each preset dimension can reflect the data distribution of customers to be recommended under this preset dimension to a certain extent.
[0072] As an example, according to the dimension data of all customers to be recommended under a certain preset dimension, through the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm, all customers to be recommended can be clustered, and each cluster obtained at this time is used as an initial cluster, and an initial cluster set corresponding to this preset dimension is obtained.
[0073] For example, if a certain preset dimension represents an address, then through the DBSCAN algorithm, customers to be recommended with similar addresses can be clustered into the same cluster as an initial cluster, and all the initial clusters obtained at this time form an initial cluster set corresponding to this preset dimension.
[0074] Step S3: Based on the initial cluster set corresponding to each preset dimension, determine the distribution change weight corresponding to each preset dimension, and filter out the preset dimensions with distribution change weights greater than the preset distribution change threshold as the target dimensions.
[0075] In some embodiments, based on the initial cluster set corresponding to each preset dimension, determine the distribution change weight corresponding to each preset dimension, and filter out the preset dimensions with distribution change weights greater than the preset distribution change threshold as the target dimensions.
[0076] Among them, the preset distribution change threshold can be a pre-set threshold. For example, the preset distribution change threshold can be 0.65.
[0077] It should be noted that, based on the initial cluster set corresponding to the preset dimension, the larger the quantified distribution change weight corresponding to the preset dimension, the greater the effective information volume under the preset dimension, and the more important the preset dimension. The target dimension is often a relatively important preset dimension.
[0078] As an example, the formula for determining the distribution change weight corresponding to the preset dimension can be:
[0079] where α i is the distribution change weight corresponding to the i-th preset dimension. i is the serial number of the preset dimension. norm() is the normalization function. d i is the mean of the absolute values of the differences between the dimensional data of the cluster centers of all initial clusters in the initial cluster set corresponding to the i-th preset dimension under the i-th preset dimension. d i,max is the maximum value of the absolute values of the differences between the dimensional data of the cluster centers of all initial clusters in the initial cluster set corresponding to the i-th preset dimension under the i-th preset dimension. N i is the number of initial clusters in the initial cluster set corresponding to the i-th preset dimension. j is the serial number of the initial cluster in the initial cluster set corresponding to the i-th preset dimension. N ij is the number of customers to be recommended in the j-th initial cluster in the initial cluster set corresponding to the i-th preset dimension. N is the total number of customers to be recommended. exp() is the exponential function with the natural constant as the base. is the variance of the target differences corresponding to all initial clusters in the initial cluster set corresponding to the i-th preset dimension. d ij is the target difference corresponding to the j-th initial cluster in the initial cluster set corresponding to the i-th preset dimension. d ij1is the reference difference of the j-th initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension under the i-th preset dimension. The reference difference of the j-th initial clustering cluster under the i-th preset dimension is the average value of the absolute values of the differences between the dimension data of all customers to be recommended in the j-th initial clustering cluster under the i-th preset dimension. d ij2 is the average value of the reference differences of the j-th initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension under all preset dimensions except the i-th preset dimension.
[0080] It should be noted that when is larger, it often indicates that based on all the dimension data under the i-th preset dimension, the classification results of clustering all customers to be recommended are relatively more distinct; it often indicates that the classification results of clustering all the dimension data under the i-th preset dimension are more distinct; it often indicates that the effective information volume under the i-th preset dimension is relatively larger, it often indicates that the i-th preset dimension is relatively more important, and it often indicates that a larger weight should be set for the i-th preset dimension. can represent the weight value of the j-th initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension. The larger its value, the more information it often indicates in the j-th initial clustering cluster. can represent the data distribution difference between the distribution of the j-th initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension in the i-th preset dimension and other preset dimensions. The larger its value, the more obvious the data distribution difference between the distribution of this initial clustering cluster in the i-th preset dimension and other preset dimensions often is, and it often indicates that a smaller weight should be set for the i-th preset dimension. Therefore, when is larger, it often indicates that the effective information volume under the i-th preset dimension is relatively larger, and it often indicates that a larger weight should be set for the i-th preset dimension. Therefore, when α i is larger, it often indicates that the i-th preset dimension is relatively more important, and it often indicates that a larger weight should be set for the i-th preset dimension.
[0081] Step S4, determine the anomaly score corresponding to each dimension data, and determine the comprehensive anomaly index corresponding to each customer to be recommended according to the anomaly scores corresponding to the dimension data of each customer to be recommended under all preset dimensions.
[0082] In some embodiments, the anomaly score corresponding to each dimension data can be determined, and the comprehensive anomaly index corresponding to each customer to be recommended can be determined according to the anomaly scores corresponding to the dimension data of each customer to be recommended under all preset dimensions.
[0083] It should be noted that, based on the anomaly scores corresponding to the dimension data of the customers to be recommended under all preset dimensions, the smaller the comprehensive anomaly index corresponding to the quantified customers to be recommended, the greater the amount of effective information contained in the dimension data of the customers to be recommended under all preset dimensions, and the more suitable the customers to be recommended are for subsequent interpolation.
[0084] As an example, this step may include the following steps:
[0085] First step, according to the above initial data set, determine the anomaly score corresponding to each dimension data through the isolation forest algorithm.
[0086] Second step, determine the mean value of the anomaly scores corresponding to the dimension data of the above customers to be recommended under all preset dimensions as the comprehensive anomaly index corresponding to the above customers to be recommended.
[0087] Step S5, determine the target importance level corresponding to each customer to be recommended according to the comprehensive anomaly index corresponding to each customer to be recommended and the set of initial clustering clusters corresponding to all target dimensions.
[0088] In some embodiments, the target importance level corresponding to each customer to be recommended may be determined according to the comprehensive anomaly index corresponding to each customer to be recommended and the set of initial clustering clusters corresponding to all target dimensions.
[0089] It should be noted that, based on the comprehensive anomaly index corresponding to the customers to be recommended and the set of initial clustering clusters corresponding to all target dimensions, the greater the target importance level corresponding to the quantified customers to be recommended, the more important the customers to be recommended are for subsequent interpolation processing, and the more suitable the customers to be recommended are for subsequent interpolation processing.
[0090] As an example, the formula for determining the target importance level corresponding to the customers to be recommended may be:
[0091] where τ k is the target importance level corresponding to the kth customer to be recommended. k is the serial number of the customer to be recommended. exp() is the exponential function with the natural constant as the base. F k is the comprehensive anomaly index corresponding to the kth customer to be recommended. n is the number of target dimensions. norm() is the normalization function. m is the serial number of the target dimension. v kmis the mean of the absolute values of the differences between the dimensional data of the clustering center of the k-th to-be-recommended customer belonging to the k-th initial clustering cluster in the initial clustering cluster set corresponding to the n-th target dimension under the n-th target dimension and the dimensional data of the clustering centers of all initial clustering clusters other than the initial clustering cluster to which the k-th to-be-recommended customer belongs under the n-th target dimension; that is, in the initial clustering cluster set corresponding to the n-th target dimension, it is the mean of the absolute values of the differences between the dimensional data of the clustering center of the initial clustering cluster to which the k-th to-be-recommended customer belongs under the n-th target dimension and the dimensional data of the clustering centers of all other initial clustering clusters under the n-th target dimension. || is the absolute value function. A km1 is the dimensional data of the k-th to-be-recommended customer under the n-th target dimension. A km2 is the dimensional data of the clustering center of the initial clustering cluster to which the k-th to-be-recommended customer belongs in the initial clustering cluster set corresponding to the n-th target dimension under the n-th target dimension.
[0092] It should be noted that when F k is smaller, it often indicates that the information-containing ability of the k-th to-be-recommended customer is relatively greater, and it often indicates that the k-th to-be-recommended customer is more suitable for subsequent interpolation processing. When v km is smaller, it often indicates that the importance of the initial clustering cluster to which the k-th to-be-recommended customer belongs in the initial clustering cluster set corresponding to the n-th target dimension for the data distribution characteristics of other initial clustering clusters under this dimension is relatively greater, it often indicates that the k-th to-be-recommended customer is more important for subsequent interpolation processing, and it often indicates that this to-be-recommended customer is more suitable for subsequent interpolation processing. When |A km1 -A km2 | is smaller, it often indicates that the data of the k-th to-be-recommended customer in the initial clustering cluster set corresponding to the n-th target dimension under this dimension is closer to the data of its clustering center under this dimension, it often indicates that the k-th to-be-recommended customer can better represent the clustering center to which it belongs under this dimension, it often indicates that the k-th to-be-recommended customer is more important for subsequent interpolation processing, and it often indicates that the k-th to-be-recommended customer is more suitable for subsequent interpolation processing. can be used as the weight value of exp(-F k ). When τ k is larger, it often indicates that the k-th to-be-recommended customer is more important for subsequent interpolation processing, and it often indicates that the k-th to-be-recommended customer is more suitable for subsequent interpolation processing.
[0093] Step S6, screen out the set of target customer groups from all to-be-recommended customers according to the target importance.
[0094] In some embodiments, the set of target customer groups can be screened out from all to-be-recommended customers according to the target importance.
[0095] It should be noted that, based on the importance of the targets, the target customers in the set of target customer groups selected can be relatively important customers.
[0096] As an example, this step may include the following steps:
[0097] In the first step, randomly form a group of customers to be recommended with every first preset number of customers to be recommended.
[0098] Among them, the first preset number can be a preset number. For example, the first preset number can be 100.
[0099] For example, if the first preset number is 100, then every 100 customers to be recommended can be randomly grouped into a group of customers to be recommended.
[0100] Optionally, all customers to be recommended can be equally divided, and each divided group can be used as a group of customers to be recommended.
[0101] In the second step, screen out the second preset number of customers to be recommended with the greatest target importance from each group of customers to be recommended as target customers, and form a target customer group.
[0102] Among them, the second preset number can be less than the first preset number. For example, if the first preset number is 100, the second preset number can be 50. There is one target customer group in each group of customers to be recommended.
[0103] In the third step, form a set of target customer groups with all the target customer groups.
[0104] Step S7: Based on the dimensional data of all target customers in the set of target customer groups in all preset dimensions, perform incremental clustering on the set of target customer groups, and determine the interpolation weight index corresponding to each target customer according to the incremental clustering result.
[0105] In some embodiments, based on the dimensional data of all target customers in the set of target customer groups in all preset dimensions, perform incremental clustering on the set of target customer groups, and determine the interpolation weight index corresponding to each target customer according to the incremental clustering result.
[0106] It should be noted that, based on the dimensional data of all target customers in the set of target customer groups in all preset dimensions, perform incremental clustering on the set of target customer groups, and the larger the interpolation weight index corresponding to the quantified target customer according to the incremental clustering result, the more important the target customer is usually for subsequent interpolation processing, and the more the target customer can usually be used for subsequent interpolation processing.
[0107] As an example, this step may include the following steps:
[0108] First, randomly select a target customer group from the above-mentioned set of target customer groups as the initial clustering data set for incremental clustering.
[0109] Among them, the initial clustering data set can be the data set during the first clustering in the incremental clustering process.
[0110] Second, perform clustering on the above-mentioned initial clustering data set according to the dimensional data of all target customers in the initial clustering data set under all preset dimensions, and determine the result of clustering the above-mentioned initial clustering data set as the first clustering result in the incremental clustering process.
[0111] For example, first, the vector composed of the dimensional data of the target customer under all preset dimensions can be used as the feature vector of the target customer. Then, according to the feature vectors of all target customers in the initial clustering data set, the initial clustering data set can be clustered through the k-means (k-means clustering algorithm) algorithm. Target customers with similar feature vectors in the initial clustering data set can be clustered into the same clustering cluster to obtain the first clustering result in the incremental clustering process. Among them, the number of clustering clusters in the k-means algorithm can be set to 8.
[0112] Third, determine each target customer group in the above-mentioned set of target customer groups except the initial clustering data set as the data set for each increment in the incremental clustering process.
[0113] Fourth, according to the dimensional data of all target customers in the data set for each increment under all preset dimensions, add the data set for each increment to the previous clustering result in the incremental clustering process, and determine the clustering result after addition as another clustering result in the incremental clustering process.
[0114] Among them, the number of clustering clusters in each clustering result in the incremental clustering process is the same.
[0115] For example, according to the feature vectors of all target customers in the new increment data set, through the k-means algorithm, the new increment data set can be added to the previous clustering result in the incremental clustering process to achieve clustering of target customers with similar feature vectors into the same clustering cluster and obtain a clustering result.
[0116] Fifth, match the clustering clusters in every two adjacent clustering results in the incremental clustering process. Obtaining the set of matching clustering cluster groups between every two adjacent clustering results in the incremental clustering process can include the following sub-steps:
[0117] The first sub-step is to respectively determine any two adjacent clustering results in the incremental clustering process as the first candidate clustering result and the second candidate clustering result.
[0118] The second sub-step is to determine each clustering cluster in the first candidate clustering result as the first candidate clustering cluster, and determine each clustering cluster in the second candidate clustering result as the second candidate clustering cluster.
[0119] The third sub-step is to determine the target edge weight value between each first candidate clustering cluster and each second candidate clustering cluster according to the intersection-over-union ratio and cosine similarity between each first candidate clustering cluster and each second candidate clustering cluster.
[0120] Among them, both the intersection-over-union ratio and the cosine similarity can be positively correlated with the target edge weight value. The target customers in the first candidate clustering cluster can be characterized by the feature vectors of the target customers. The target customers in the second candidate clustering cluster can be characterized by the feature vectors of the target customers.
[0121] It should be noted that the larger the intersection-over-union ratio between two clustering clusters, the more it often indicates the more overlap between these two clustering clusters, and the more it often indicates the better match between these two clustering clusters. The larger the cosine similarity between two clustering clusters, the more it often indicates the more similar these two clustering clusters are, and the more it often indicates the better match between these two clustering clusters. Therefore, the larger the target edge weight value between two clustering clusters, the more it often indicates the better match between these two clustering clusters.
[0122] The fourth sub-step is to match all the first candidate clustering clusters and all the second candidate clustering clusters through the KM (Kuhn and Munkres, bipartite graph matching) algorithm according to the target edge weight values between all the first candidate clustering clusters and all the second candidate clustering clusters, and form a matching clustering cluster group for the mutually matching first candidate clustering cluster and second candidate clustering cluster, to obtain a set of matching clustering cluster groups between the first candidate clustering result and the second candidate clustering result, where a matching clustering cluster group includes a first candidate clustering cluster and a second candidate clustering cluster.
[0123] The sixth step is to determine the interpolation weight index corresponding to each target customer according to all the matching clustering clusters in the set of matching clustering cluster groups between all adjacent two clustering results in the incremental clustering process.
[0124] For example, the formula for determining the interpolation weight index corresponding to the target customer can be:
[0125]
[0126] where δ q is the interpolation weight index corresponding to the qth target customer. q is the serial number of the target customer. R is the number of clustering results in the incremental clustering process. nq is the minimum value among the sequence numbers of all clustering results where the q-th target customer exists in the incremental clustering process. b is the sequence number of the clustering result in the incremental clustering process, and the sequence number of the clustering result in the incremental clustering process can be the sequence number obtained according to the clustering order from early to late. δ q,b,b+1 is the q-th target customer, and it is the sub-weight index between the b-th clustering result and the (b + 1)-th clustering result in the incremental clustering process. exp() is the exponential function with the natural constant as the base. || is the absolute value function. g q,b is the cosine similarity between the dimensional data of the cluster center of the cluster to which the q-th target customer belongs in the b-th clustering result in the incremental clustering process in all preset dimensions and the dimensional data of the q-th target customer in all preset dimensions; that is, the cosine similarity between the feature vector of the cluster center of the cluster to which the q-th target customer belongs in the b-th clustering result in the incremental clustering process and the feature vector of the q-th target customer, which can characterize the central degree of the q-th target customer in the b-th clustering result in the incremental clustering process. g q,b+1 is the cosine similarity between the dimensional data of the cluster center of the cluster to which the q-th target customer belongs in the (b + 1)-th clustering result in the incremental clustering process in all preset dimensions and the dimensional data of the q-th target customer in all preset dimensions, that is, the cosine similarity between the feature vector of the cluster center of the cluster to which the q-th target customer belongs in the (b + 1)-th clustering result in the incremental clustering process and the feature vector of the q-th target customer, which can characterize the central degree of the q-th target customer in the (b + 1)-th clustering result in the incremental clustering process. NMI() is the normalized mutual information function. Q q,b is the cluster to which the q-th target customer belongs in the b-th clustering result in the incremental clustering process, and it can be characterized by the feature vectors of all target customers in the cluster to which the q-th target customer belongs in the b-th clustering result in the incremental clustering process. Q q,b+1 is the cluster to which the q-th target customer belongs in the (b + 1)-th clustering result in the incremental clustering process, and it can be characterized by the feature vectors of all target customers in the cluster to which the q-th target customer belongs in the (b + 1)-th clustering result in the incremental clustering process. If the clusters to which the q-th target customer belongs in the b-th clustering result and the (b + 1)-th clustering result in the incremental clustering process belong to the same matching cluster group, then set f q,b,b+1 to be equal to the first preset value. If the clusters to which the q-th target customer belongs in the b-th clustering result and the (b + 1)-th clustering result in the incremental clustering process do not belong to the same matching cluster group, then set f q,b,b+1 to be equal to the second preset value. The first preset value is greater than the second preset value. For example, the first preset value can be 1. The second preset value can be -1.
[0127] It should be noted that in the incremental clustering process, if the two clustering clusters to which the same customer belongs in two adjacent clustering results are in the same matching clustering cluster group and the difference between the two clustering clusters is relatively large, it indicates that the information contained in this customer is relatively large, that is, the corresponding information weight is relatively large; conversely, if the two clustering clusters to which the same customer belongs in two adjacent clustering results are not in the same matching clustering cluster group, it indicates that the distribution characteristics of this customer change differently. If the centrality of the clustering clusters before and after is similar, it indicates that the existence of this customer will disrupt the clustering result, and the information contained in this customer is often less, that is, the corresponding information weight often needs to be adjusted downwards. When f q,b,b+1 is larger, it often means that the probability that the clustering clusters to which the q-th target customer belongs in the b-th clustering result and the (b + 1)-th clustering result belong to the same matching clustering cluster group is relatively larger. When is smaller, it often means that the centrality of the clustering clusters to which the q-th target customer belongs in the b-th clustering result and the (b + 1)-th clustering result is more similar, often indicating that the existence of the q-th target customer is more likely to disrupt the clustering result, and the corresponding information weight often needs to be adjusted downwards. When NMI(Q q,b ,Q q,b+1 ) is smaller, it often means that the clustering clusters to which the q-th target customer belongs in the b-th clustering result and the (b + 1)-th clustering result are less similar, often indicating that the corresponding information weight needs to be adjusted downwards more. Therefore, when δ q is larger, it often means that the corresponding information weight needs to be adjusted upwards more, often indicating that the q-th target customer is more important for subsequent interpolation processing, and often indicating that the q-th target customer can be used for subsequent interpolation processing. Secondly, the target customers newly added in the last clustering in the incremental clustering process can not participate in the calculation of the interpolation weight index and can not participate in subsequent interpolation processing.
[0128] Step S8: According to the interpolation weight index, screen out the customers to be interpolated from the set of target customer groups, and perform interpolation processing based on the dimension data of all customers to be interpolated in all preset dimensions to obtain an interpolation data set.
[0129] In some embodiments, the customers to be interpolated can be screened out from the set of target customer groups according to the interpolation weight index, and interpolation processing is performed based on the dimension data of all customers to be interpolated in all preset dimensions to obtain an interpolation data set.
[0130] It should be noted that based on the interpolation weight index, the customers to be interpolated are screened out from the set of target customer groups, and the obtained interpolation data set can be an adaptively newly added data set through interpolation processing based on the dimension data of all customers to be interpolated in all preset dimensions.
[0131] As an example, this step may include the following steps:
[0132] First step, screen out the target customers with the largest interpolation weight index in a preset proportion from the set of target customer groups as the customers to be interpolated.
[0133] Among them, the preset proportion can be a ratio set in advance. For example, the preset proportion can be
[0134] For example, if the preset proportion is then the target customers in the set of target customer groups can be sorted in descending order according to the interpolation weight index to obtain a target customer sequence, and the first target customers in the target customer sequence are used as the customers to be interpolated.
[0135] Second step, perform interpolation processing on the dimension data of all customers to be interpolated under all preset dimensions to obtain an interpolation data set.
[0136] For example, first, linear interpolation can be performed on the dimension data of all customers to be interpolated under each preset dimension to obtain an interpolation data sequence under each preset dimension. Then, when the interpolation data in the interpolation data sequence under a preset dimension is greater than the maximum value in the preset range under this preset dimension, the interpolation data is updated to the maximum value in the preset range under this preset dimension. When the interpolation data in the interpolation data sequence under a preset dimension is less than the minimum value in the preset range under this preset dimension, the interpolation data is updated to the minimum value in the preset range under this preset dimension. When the interpolation data in the interpolation data sequence under a preset dimension belongs to the preset range under this preset dimension, the interpolation data does not need to be updated. Among them, the preset range under a preset dimension can be [H max +0.2×H, H min +0.2×H], where H max is the largest dimension data under the preset dimension. H min is the smallest dimension data under the preset dimension. H = H max -H min . Then, the vectors formed by the interpolation data at the same positions in the final interpolation data sequences under all preset dimensions can be used as target vectors. Among them, a target vector can represent a virtual customer obtained by interpolation, that is, the elements in the target vector can be the dimension data of the virtual customer under the preset dimensions.
[0137] Step S9, according to the union of the initial data set and the interpolation data set, perform information recommendation for each customer to be recommended through a collaborative filtering recommendation algorithm.
[0138] In some embodiments, information recommendation can be performed for each customer to be recommended through a collaborative filtering recommendation algorithm according to the union of the initial data set and the interpolation data set.
[0139] It should be noted that, based on the union of the initial data set and the interpolation data set, information recommendation for each customer to be recommended is realized through the collaborative filtering recommendation algorithm.
[0140] As an example, if the content to be recommended is business, the union of the initial data set and the interpolation data set can be used as the final data set. Based on the final data set, the scoring values of multiple businesses are calculated through the collaborative filtering recommendation algorithm. The businesses are sorted according to the scoring values, and business recommendation is made for the customers to be recommended according to the sorting result.
[0141] Based on the same inventive concept as the above method embodiment, the present invention provides an enterprise digital management system, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of an enterprise digital management method are realized.
[0142] In summary, compared with directly using the initial data set for collaborative filtering recommendation, the present invention realizes the expansion of the data set, enriches the data in the data set, solves the cold start problem in the collaborative filtering recommendation process to a certain extent, and improves the rationality of the recommendation result.
[0143] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A digital management method for an enterprise, characterized in that: The following steps are involved: Obtain the dimension data of each customer to be recommended under each preset dimension and construct an initial data set; According to the dimension data of all the customers to be recommended under each preset dimension, all the customers to be recommended are clustered to obtain an initial cluster set corresponding to each preset dimension; According to the initial cluster set corresponding to each preset dimension, determine the distribution change weight corresponding to each preset dimension, and select the preset dimension whose distribution change weight is greater than the preset distribution change threshold as the target dimension; Determine the abnormality score corresponding to each dimension data, and determine the comprehensive abnormality index corresponding to each customer to be recommended based on the abnormality score corresponding to the dimension data of each customer to be recommended under all preset dimensions; Determine the target importance of each customer to be recommended based on the comprehensive abnormal index corresponding to each customer to be recommended and the initial cluster set corresponding to all target dimensions; According to the importance of the target, select the target customer group from all the customers to be recommended; According to the dimension data of all target customers in the target customer group set under all preset dimensions, incremental clustering is performed on the target customer group set, and according to the incremental clustering result, an interpolation weight index corresponding to each target customer is determined; According to the interpolation weight index, the customers to be interpolated are screened out from the target customer group set, and interpolation processing is performed according to the dimensional data of all the customers to be interpolated under all preset dimensions to obtain an interpolation data set; According to the union of the initial data set and the interpolated data set, information is recommended to each customer to be recommended through the collaborative filtering recommendation algorithm.
2. The enterprise digital management method according to claim 1, characterized in that: The formula corresponding to the distribution change weight corresponding to the preset dimension is: Among them, α i is the distribution change weight corresponding to the i-th preset dimension; i is the serial number of the preset dimension; norm() is the normalization function; d i It is the mean of the absolute values of the differences between the dimensional data of the cluster centers of all initial clusters under the i-th preset dimension in the initial cluster cluster set corresponding to the i-th preset dimension; d i,max It is the maximum value of the absolute value of the difference between the dimensional data of the cluster centers of all initial clusters under the i-th preset dimension in the initial cluster cluster set corresponding to the i-th preset dimension; N i is the number of initial clustering clusters in the initial clustering cluster set corresponding to the i-th preset dimension; j is the sequence number of the initial clustering cluster in the initial clustering cluster set corresponding to the i-th preset dimension; N ij is the number of customers to be recommended in the jth initial cluster in the initial cluster set corresponding to the i-th preset dimension; N is the total number of customers to be recommended; exp() is an exponential function with a natural constant as the base; is the variance of the target differences corresponding to all initial clustering clusters in the initial clustering cluster set corresponding to the i-th preset dimension; d ij is the target difference corresponding to the jth initial cluster in the initial cluster set corresponding to the i-th preset dimension; d ij1 is the reference difference of the jth initial cluster in the initial cluster set corresponding to the i-th preset dimension under the i-th preset dimension; the reference difference of the jth initial cluster under the i-th preset dimension is the mean of the absolute values of the differences between the dimension data of all the recommended customers in the jth initial cluster under the i-th preset dimension; d ij2 It is the mean of the reference differences of the jth initial cluster in the initial cluster set corresponding to the i-th preset dimension under all preset dimensions except the i-th preset dimension.
3. The enterprise digital management method according to claim 1, characterized in that: The step of determining the abnormality score corresponding to each dimension data, and determining the comprehensive abnormality index corresponding to each customer to be recommended according to the abnormality score corresponding to the dimension data of each customer to be recommended under all preset dimensions, includes: According to the initial data set, an abnormal score corresponding to each dimension data is determined by using an isolation forest algorithm; The average of the abnormality scores corresponding to the dimension data of the customer to be recommended under all preset dimensions is determined as the comprehensive abnormality index corresponding to the customer to be recommended.
4. The enterprise digital management method according to claim 1, characterized in that: The formula for the target importance of the recommended customers is: Among them, τ k is the target importance corresponding to the kth customer to be recommended; k is the serial number of the customer to be recommended; exp() is an exponential function with a natural constant as the base; F k is the comprehensive abnormality index corresponding to the kth customer to be recommended; n is the number of target dimensions; norm() is the normalization function; m is the sequence number of the target dimension; v km is the average of the absolute values of the differences between the dimensional data of the cluster center of the initial cluster to which the kth customer to be recommended belongs in the initial cluster set corresponding to the nth target dimension under the nth target dimension and the dimensional data of the cluster centers of all initial clusters other than the initial cluster to which the kth customer to be recommended belongs in the initial cluster set corresponding to the nth target dimension under the nth target dimension; || is the absolute value function; A km1 is the dimension data of the kth customer to be recommended under the nth target dimension; A km2 It is the dimensional data under the nth target dimension of the cluster center of the initial cluster to which the kth customer to be recommended belongs in the initial cluster set corresponding to the nth target dimension.
5. The enterprise digital management method according to claim 1, characterized in that: The target customer group set is screened from all customers to be recommended according to the target importance, including: Randomly forming a to-be-recommended customer group with each first preset number of to-be-recommended customers; Screening out a second preset number of recommended customers with the greatest target importance from each group of recommended customers as target customers to form a target customer group, wherein the second preset number is less than the first preset number; All target customer groups constitute a target customer group set.
6. The enterprise digital management method according to claim 1, characterized in that: The step of performing incremental clustering on the target customer group set according to the dimensional data of all target customers in the target customer group set under all preset dimensions, and determining the interpolation weight index corresponding to each target customer according to the incremental clustering result, includes: Randomly selecting a target customer group from the target customer group set as an initial clustering data set for incremental clustering; Clustering the initial clustering data set according to the dimension data of all target customers in the initial clustering data set under all preset dimensions, and determining the result of clustering the initial clustering data set as the first clustering result in the incremental clustering process; Determine each target customer group in the target customer group set, except the initial clustering data set, as a data set to be incremented each time in the incremental clustering process; According to the dimension data of all target customers in each incremental data set under all preset dimensions, each incremental data set is added to the previous clustering result in the incremental clustering process, and the clustering result after the addition is determined as another clustering result in the incremental clustering process, wherein the number of clusters in each clustering result in the incremental clustering process is the same; Matching the clusters in every two adjacent clustering results in the incremental clustering process to obtain a set of matching cluster groups between every two adjacent clustering results in the incremental clustering process; According to all matching clustering clusters in the matching clustering cluster group set between all two adjacent clustering results in the incremental clustering process, the interpolation weight index corresponding to each target customer is determined.
7. The enterprise digital management method according to claim 6, characterized in that: The step of matching the clusters in each two adjacent clustering results in the incremental clustering process to obtain a set of matching cluster groups between each two adjacent clustering results in the incremental clustering process includes: Determine any two adjacent clustering results in the incremental clustering process as the first candidate clustering result and the second candidate clustering result respectively; Determine each cluster in the first candidate clustering result as a first candidate clustering cluster, and determine each cluster in the second candidate clustering result as a second candidate clustering cluster; Determine a target edge weight between each first candidate cluster and each second candidate cluster according to an intersection-over-union ratio and a cosine similarity between each first candidate cluster and each second candidate cluster, wherein both the intersection-over-union ratio and the cosine similarity are positively correlated with the target edge weight; According to the target edge weights between all the first candidate clustering clusters and all the second candidate clustering clusters, the KM algorithm is used to match all the first candidate clustering clusters and all the second candidate clustering clusters, and the mutually matching first candidate clustering clusters and second candidate clustering clusters form a matching clustering cluster group, thereby obtaining a matching clustering cluster group set between the first candidate clustering results and the second candidate clustering results.
8. The enterprise digital management method according to claim 6, characterized in that: The formula for the interpolation weight indicator corresponding to the target customer is: Among them, δ q is the interpolation weight index corresponding to the qth target customer; q is the serial number of the target customer; R is the number of clustering results in the incremental clustering process; n q is the minimum value among the serial numbers of all clustering results with the qth target customer in the incremental clustering process; b is the serial number of the clustering result in the incremental clustering process; δ q,b,b+1 is the sub-weight index between the b-th clustering result and the b+1-th clustering result of the q-th target customer in the incremental clustering process; exp() is an exponential function with a natural constant as the base; || is an absolute value function; g q,b is the cosine similarity between the dimensional data of the cluster center of the cluster to which the qth target customer belongs in the bth clustering result in the incremental clustering process under all preset dimensions and the dimensional data of the qth target customer under all preset dimensions; g q,b+1 is the cosine similarity between the dimensional data of the cluster center of the cluster to which the qth target customer belongs in the b+1th clustering result in the incremental clustering process under all preset dimensions and the dimensional data of the qth target customer under all preset dimensions; NMI() is the normalized mutual information function; Q q,b is the cluster to which the qth target customer belongs in the bth clustering result during the incremental clustering process; Q q,b+1 is the cluster to which the qth target customer belongs in the b+1th clustering result in the incremental clustering process; if the cluster to which the qth target customer belongs in the bth clustering result and the b+1th clustering result in the incremental clustering process belongs to the same matching cluster group, then set f q,b,b+1 is equal to the first preset value; if the clustering clusters to which the qth target customer belongs in the bth clustering result and the b+1th clustering result in the incremental clustering process do not belong to the same matching clustering cluster group, then set f q,b,b+1 is equal to the second preset value; the first preset value is greater than the second preset value.
9. The enterprise digital management method according to claim 1, characterized in that: The step of selecting the customers to be interpolated from the target customer group set according to the interpolation weight index includes: Target customers with the largest preset proportion of interpolation weight index are selected from the target customer group set as customers to be interpolated.
10. An enterprise digital management system, characterized in that: It comprises a processor and a memory, wherein the processor is used to process instructions stored in the memory to implement an enterprise digital management method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Industry user recommendation method and device based on operator data
CN108052639A
Single-layer enhanced negative sample generation algorithm for graph neural network training
CN115858924A
System and method for recommendation using multivariate data learning through collaborative filtering
FR3144362A1