A method and apparatus for classifying customers

CN115640397BActive Publication Date: 2026-08-28INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211226436.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2026-08-28
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

可以通过半监督聚类算法对客户进行分类,为相同类型的客户推荐相同类型的理财产品,然而在实际应用中,存在由于客户相似和不相似的成对约束信息较少,导致客户聚类的结果不够准确,影响理财产品的推荐效果

Benefits of technology

[0018]The customer classification method and apparatus provided in this invention can acquire the attribute information of each customer in the prior customer pair constraint set and the mandatory constraint subset. The attribute information of each customer in the mandatory constraint subset is preprocessed to obtain attribute feature data for each customer in the pair constraint set. Based on the mandatory customer pairs in the mandatory constraint subset, the non-mandatory customer pairs in the non-mandatory constraint subset, and the attribute feature data of each customer in the mandatory constraint subset, the mandatory constraint subset is expanded to obtain an expanded mandatory constraint subset, and the non-mandatory constraint subset is expanded to obtain an expanded non-mandatory constraint subset. Based on the expanded mandatory constraint subset, the expanded non-mandatory constraint subset, and the attribute feature data of the customers to be classified, customer clustering is performed using a pair constraint semi-supervised clustering algorithm to obtain the customer classification result. Because the pair constraint information required in the pair constraint semi-supervised clustering algorithm is added, the accuracy of customer classification is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115640397B_ABST
    Figure CN115640397B_ABST
Patent Text Reader

Abstract

The application provides a customer classification method and device, which can be used in the technical field of big data. The method comprises the following steps: acquiring a pair constraint set of prior customers and attribute information of each customer in a must-connect constraint subset; preprocessing the attribute information of each customer in the must-connect constraint subset to obtain attribute feature data of each customer; expanding the must-connect constraint subset to obtain an expanded must-connect constraint subset and expanding the must-not-connect constraint subset to obtain an expanded must-not-connect constraint subset according to each must-connect customer pair in the must-connect constraint subset, each must-not-connect customer pair in the must-not-connect constraint subset and the attribute feature data of each customer in the must-connect constraint subset; and clustering to obtain a classification result of the customers according to the expanded must-connect constraint subset, the expanded must-not-connect constraint subset and attribute feature data of a customer to be classified. The device is used to execute the above method. The customer classification method and device provided by the application embodiment improve the accuracy of customer classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data technology, specifically to a customer classification method and apparatus. Background Technology

[0002] Semi-supervised clustering guides the clustering process by using a small amount of prior information about the class structure, with the aim of ensuring that the clustering results meet expectations.

[0003] Currently, numerous semi-supervised clustering algorithms have been proposed, and some methods have been successfully applied in various fields. For example, banks can use clustering algorithms to classify customers, enabling them to recommend suitable wealth management products to different customer types. Banks offer many different wealth management products, each with varying investment targets, investment periods, investment amounts, expected returns, and risk levels. Customers also differ in their investment amounts, risk tolerance, investment preferences, and expected returns. While semi-supervised clustering algorithms can classify customers and recommend similar wealth management products to similar customers, in practical applications, the limited pairwise constraints between similar and dissimilar customers can lead to inaccurate clustering results, affecting the effectiveness of wealth management product recommendations. Summary of the Invention

[0004] To address the problems in the prior art, embodiments of the present invention provide a customer classification method and apparatus that can at least partially solve the problems existing in the prior art.

[0005] In a first aspect, the present invention proposes a customer classification method, comprising:

[0006] Obtain the set of paired constraints for prior customers and the attribute information of each customer in the mandatory connection constraint subset; wherein, the set of paired constraints includes the mandatory connection constraint subset and the non-connection constraint subset; the mandatory connection constraint subset includes a first preset number of mandatory connection customer pairs, and the non-connection constraint subset includes a second preset number of non-connection customer pairs;

[0007] Preprocess the attribute information of each customer in the mandatory constraint subset to obtain the attribute feature data of each customer in the pairwise constraint set;

[0008] Based on the mandatory connection customer pairs in the mandatory connection constraint subset, the non-connection customer pairs in the non-connection constraint subset, and the attribute feature data of each customer in the mandatory connection constraint subset, the mandatory connection constraint subset is expanded to obtain an expanded mandatory connection constraint subset, and the non-connection constraint subset is expanded to obtain an expanded non-connection constraint subset.

[0009] Based on the extended mandatory-connection constraint subset, the extended non-connection constraint subset, and the attribute feature data of the customers to be classified, customer clustering is performed using a pairwise constraint semi-supervised clustering algorithm to obtain the customer classification results.

[0010] Secondly, the present invention provides a customer sorting device, comprising:

[0011] The acquisition module is used to acquire the attribute information of each customer in the prior customer pair constraint set and the mandatory connection constraint subset; wherein, the pair constraint set includes the mandatory connection constraint subset and the non-connection constraint subset; the mandatory connection constraint subset includes a first preset number of mandatory connection customer pairs, and the non-connection constraint subset includes a second preset number of non-connection customer pairs;

[0012] The preprocessing module is used to preprocess the attribute information of each customer in the mandatory constraint subset to obtain the attribute feature data of each customer in the pairwise constraint set.

[0013] An expansion module is used to expand the mandatory connection constraint subset to obtain an expanded mandatory connection constraint subset and expand the non-connection constraint subset to obtain an expanded non-connection constraint subset based on each mandatory connection customer pair in the mandatory connection constraint subset, each non-connection customer pair in the non-connection constraint subset, and the attribute feature data of each customer in the mandatory connection constraint subset.

[0014] The clustering module is used to cluster customers based on the extended mandatory-connection constraint subset, the extended non-connection constraint subset, and the attribute feature data of the customers to be classified, using a pairwise constraint-based semi-supervised clustering algorithm to obtain the customer classification results.

[0015] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the customer classification method described in any of the above embodiments.

[0016] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the customer classification method described in any of the above embodiments.

[0017] Fifthly, the present invention provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the customer classification method described in any of the above embodiments.

[0018] The customer classification method and apparatus provided in this invention can acquire the attribute information of each customer in the prior customer pair constraint set and the mandatory constraint subset. The attribute information of each customer in the mandatory constraint subset is preprocessed to obtain attribute feature data for each customer in the pair constraint set. Based on the mandatory customer pairs in the mandatory constraint subset, the non-mandatory customer pairs in the non-mandatory constraint subset, and the attribute feature data of each customer in the mandatory constraint subset, the mandatory constraint subset is expanded to obtain an expanded mandatory constraint subset, and the non-mandatory constraint subset is expanded to obtain an expanded non-mandatory constraint subset. Based on the expanded mandatory constraint subset, the expanded non-mandatory constraint subset, and the attribute feature data of the customers to be classified, customer clustering is performed using a pair constraint semi-supervised clustering algorithm to obtain the customer classification result. Because the pair constraint information required in the pair constraint semi-supervised clustering algorithm is added, the accuracy of customer classification is improved. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0020] Figure 1 This is a flowchart illustrating the customer classification method provided in the first embodiment of the present invention.

[0021] Figure 2 This is a flowchart illustrating the customer classification method provided in the second embodiment of the present invention.

[0022] Figure 3 This is a schematic diagram of the customer's K-nearest neighbor graph provided in the third embodiment of the present invention.

[0023] Figure 4 This is a schematic diagram of the customer classification device provided in the fourth embodiment of the present invention.

[0024] Figure 5 This is a schematic diagram of the customer sorting device provided in the fifth embodiment of the present invention.

[0025] Figure 6 This is a schematic diagram of the customer classification device provided in the sixth embodiment of the present invention.

[0026] Figure 7 This is a schematic diagram of the physical structure of the electronic device provided in the seventh embodiment of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain the present invention, but are not intended to limit the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant means, and the acquisition, storage, use, and processing of user information have been authorized and agreed upon by the customer.

[0028] To facilitate understanding of the technical solution provided in this application, the relevant content of the technical solution in this application will be explained below.

[0029] This invention provides a customer classification method that, during customer classification, expands prior customer information by using existing prior customer information, namely, pairwise constraints on similar and dissimilar customers. This expanded prior customer information then guides the customer clustering process, thereby improving the accuracy of customer clustering. The pairwise constraints defined for customers refer to the determination, through manual comparison, of which customers belong to the same type and which do not. For example, two customers with similar asset sizes and who have purchased the same type of product are grouped into the same type, forming a must-link constraint pair; conversely, two customers, one who only purchased bank deposits and the other who only purchased bank wealth management products, are grouped into different types, forming a cannot-link constraint pair.

[0030] The following describes the specific implementation process of the customer classification method provided in this embodiment of the invention, using a server as the execution subject as an example.

[0031] Figure 1 This is a flowchart illustrating the customer classification method provided in the first embodiment of the present invention, as shown below. Figure 1 As shown, the customer classification method provided in this embodiment of the invention includes:

[0032] S101. Obtain the set of paired constraints for prior customers and the attribute information of each customer in the set of paired constraints; wherein, the set of paired constraints includes a subset of mandatory constraints and a subset of non-mandatory constraints; the subset of mandatory constraints includes a first preset number of mandatory customer pairs, and the subset of non-mandatory constraints includes a second preset number of non-mandatory customer pairs.

[0033] Specifically, a set of paired constraints for prior customers and the attribute information of each customer in the paired constraint set are collected in advance. The server can obtain the set of paired constraints for prior customers and the attribute information of each customer in the paired constraint set. The set of paired constraints includes a mandatory constraint subset and a non-mandatory constraint subset. The mandatory constraint subset includes a first preset number of mandatory customer pairs, which includes two customers belonging to the same category. The non-mandatory constraint subset includes a second preset number of non-mandatory customer pairs, which includes two customers belonging to different categories. The first preset number is set according to actual needs and is not limited in this embodiment of the invention. The second preset number is set according to actual needs and is not limited in this embodiment of the invention. Whether the two customers are of the same type or different types is set according to the actual situation and is not limited in this embodiment of the invention. The customer attribute information includes, but is not limited to, the customer's total assets, whether they have purchased a certain product, the amount purchased for bank wealth management products, fund purchases, deposit product purchases, the interest rate of the purchased product, the term of the purchased product, the customer's family situation (e.g., whether they have children, whether they need to support parents), the customer's risk assessment level, and the time period during which the customer purchased the product.

[0034] In this context, prior customers can be a subset of the customers to be categorized. A set of paired constraints for prior customers can be established manually or through other means. For example, 5% of the customers to be categorized can be selected as prior customers.

[0035] S102. Preprocess the attribute information of each customer in the pairwise constraint set to obtain the attribute feature data of each customer in the pairwise constraint set;

[0036] Specifically, the server preprocesses the attribute information of each customer in the pairwise constraint set to obtain attribute feature data for each customer in the pairwise constraint set. The preprocessing includes, but is not limited to, data normalization, one-hot encoding, and binarization, and is configured according to actual needs; this embodiment of the invention does not impose limitations on these configurations.

[0037] For example, numerical data such as a customer's total assets, the amount purchased with bank wealth management products, the amount purchased with funds, the amount purchased with deposit products, and the interest rates of the purchased products can be normalized, scaling the data to between 0 and 1 to unify the measurement standards for different data. For the terms corresponding to the purchased products, such as 3 months, 6 months, 1 year, 2 years, 3 years, and 5 years, each term can be uniquely encoded, assigning a unique code to each term. Whether a customer has purchased product A can be binary-coded: purchasing product A is represented as 1, and not purchasing product A is represented as 0.

[0038] S103. Based on each pair of mandatory connected customers in the mandatory connection constraint subset, each pair of non-connected customers in the non-connected constraint subset, and the attribute feature data of each customer in the mandatory connection constraint subset, the mandatory connection constraint subset is expanded to obtain an expanded mandatory connection constraint subset, and the non-connected constraint subset is expanded to obtain an expanded non-connected constraint subset.

[0039] Specifically, the server expands the mandatory connection constraint subset based on each mandatory connection customer pair in the mandatory connection constraint subset, each non-connection customer pair in the non-connection constraint subset, and the attribute feature data of each customer in the mandatory connection constraint subset, to obtain more mandatory connection customer pairs to form an expanded non-connection constraint subset, and at the same time obtain more non-connection customer pairs to form an expanded non-connection constraint subset.

[0040] S104. Based on the extended mandatory constraint subset, the extended non-mandatory constraint subset, and the attribute feature data of the customers to be classified, customer clustering is performed using a pairwise constraint semi-supervised clustering algorithm to obtain the customer classification results.

[0041] Specifically, the server uses each pair of mandatory customers in the expanded mandatory-connection constraint subset and each pair of non-connection customers in the expanded non-connection constraint subset as pairwise constraints in a pairwise constraint semi-supervised clustering algorithm to cluster the attribute feature data of the customers to be classified, thereby obtaining the customer classification results. The customer classification results may include multiple customer categories and customers under each category. The specific process of clustering data using the pairwise constraint semi-supervised clustering algorithm is existing technology and will not be elaborated here.

[0042] The customer classification method provided in this invention can obtain the attribute information of each customer in the prior customer pair constraint set and the mandatory constraint subset. It preprocesses the attribute information of each customer in the mandatory constraint subset to obtain attribute feature data for each customer in the pair constraint set. Based on the mandatory customer pairs in the mandatory constraint subset, the non-mandatory customer pairs in the non-mandatory constraint subset, and the attribute feature data of each customer in the mandatory constraint subset, the mandatory constraint subset is expanded to obtain an expanded mandatory constraint subset, and the non-mandatory constraint subset is expanded to obtain an expanded non-mandatory constraint subset. Based on the expanded mandatory constraint subset, the expanded non-mandatory constraint subset, and the attribute feature data of the customers to be classified, customer clustering is performed using a pair constraint semi-supervised clustering algorithm to obtain the customer classification result. Because the pair constraint information required in the pair constraint semi-supervised clustering algorithm is added, the accuracy of customer classification is improved.

[0043] Figure 2 This is a flowchart illustrating the customer classification method provided in the second embodiment of the present invention, as shown below. Figure 2As shown, based on the above embodiments, further, the step of expanding the mandatory connection constraint subset to obtain an expanded mandatory connection constraint subset and expanding the non-connection constraint subset to obtain an expanded non-connection constraint subset based on the attribute feature data of each mandatory connection customer pair in the mandatory connection constraint subset, each non-connection customer pair in the non-connection constraint subset, and each customer in the pairwise constraint set includes:

[0044] S201. Obtain two sets of mandatory objects from the set of mandatory targets; wherein, the initial data of the set of mandatory targets is the subset of mandatory constraints, and the set of mandatory objects is a pair of mandatory clients or a set of mandatory clients;

[0045] Specifically, the server obtains two sets of mandatory connection objects from the mandatory connection target set, which includes multiple sets of mandatory connection objects. Each set of mandatory connection objects is either a pair of mandatory connection clients or a set of mandatory connection clients, where the set of mandatory connection clients consists of clients from multiple pairs of mandatory connection clients that satisfy the merging rules. The initial data of the mandatory connection target set is the subset of mandatory connection constraints.

[0046] S202. If it is determined that the two sets of mandatory objects satisfy the merging rule based on the attribute feature data of each customer in the two sets of mandatory objects and the non-connection customer pairs in the non-connection constraint subset, then the customers in the two sets of mandatory objects are merged into one set of mandatory customers; if it is determined that the customers in the two sets of mandatory objects satisfy the non-connection confirmation rule, then a new pair of non-connection customers is generated based on the customers in the two sets of mandatory objects.

[0047] Specifically, the server determines whether the two sets of mandatory connections satisfy the merging rule based on the attribute feature data of each customer in the two sets of mandatory connections and the pairs of non-connecting customers in the non-connecting constraint subset. If the merging rule is satisfied, it means that any two customers in the two sets of mandatory connections can form a mandatory customer pair, and the customers in the two sets of mandatory connections will be merged into a single set of mandatory customers. If the two sets of mandatory connections do not satisfy the merging rule, the customers in the two sets of mandatory connections will not be merged into a single set of mandatory customers. The merging rule is preset and is used to determine under what circumstances the two sets of mandatory connections can be merged. It is understood that each customer in the set of mandatory customers is unique.

[0048] The server determines whether the customers in the two sets of mandatory connections meet the do-not-connect confirmation rule. If the customers in the two sets of mandatory connections meet the do-not-connect confirmation rule, it means that the customers from one set of mandatory connections and the other set of mandatory connections can form a do-not-connect customer pair, and a new do-not-connect customer pair can be generated based on the customers in the two sets of mandatory connections.

[0049] For example, the set of objects that must be connected, A, includes customers a1 and a2, and the set of objects that must be connected, B, includes customers b1 and b2. If it is determined that the set of objects that must be connected, A and B satisfy the merging rule, then the set of objects that must be connected, A and B are combined into a set of customers that must be connected, C. The set of customers that must be connected, C includes customers a1, a2, b1, and b2.

[0050] For example, the set of objects that must be connected, E, includes customers e1 and e2, and the set of objects that must be connected, F, includes customers f1 and f2. If it is determined that the set of objects that must be connected, E, and F satisfy the do-not-connect confirmation rule, then customers e1 and f1 form a new do-not-connect customer pair, customers e1 and f2 form a new do-not-connect customer pair, customers e2 and f1 form a new do-not-connect customer pair, and customers e2 and f2 form a new do-not-connect customer pair.

[0051] S203. Update the set of mandatory targets based on the remaining set of mandatory objects in the set of mandatory targets and the merged set of mandatory customers; wherein, the remaining set of mandatory objects in the set of mandatory targets refers to the set of mandatory objects remaining after removing the two sets of mandatory objects used to merge into the set of mandatory customers from the set of mandatory targets.

[0052] Specifically, the server retains the remaining set of mandatory connection objects in the mandatory connection target set and adds a newly merged set of mandatory connection clients as the mandatory connection object set, thereby updating the mandatory connection target set. Here, the remaining set of mandatory connection objects in the mandatory connection target set refers to the set of mandatory connection objects remaining after removing the two sets of mandatory connection objects used to merge into a mandatory connection client set. It is understood that if, after step S202, the two sets of mandatory connection objects are not merged into a single mandatory connection client set, then the mandatory connection target set remains unchanged.

[0053] For example, the set of required connected objects A and the set of required connected objects B are combined into a set of required connected clients C. The set of required connected objects A and B are removed from the set of required connected targets, and the set of required connected clients C is added, resulting in an updated set of required connected targets.

[0054] S204. If it is determined that the number of required object sets in the required target set is greater than the number threshold, then continue to obtain two required object sets from the required target set for merging judgment, until the number of required object sets in the required target set is equal to the number threshold or any two required object sets in the required target set do not meet the merging rule;

[0055] Specifically, the server counts the number of mandatory connection object sets in the mandatory connection target set, obtains the number of mandatory connection object sets, and compares the number of mandatory connection object sets with a quantity threshold. If the number of mandatory connection object sets is greater than the quantity threshold, then two mandatory connection object sets are obtained from the mandatory connection target set for merging judgment, that is, steps S201 and S202 are repeated. If the two mandatory connection object sets meet the merging rules, then they are merged into a mandatory connection customer set. If the two mandatory connection object sets meet the do-not-connection confirmation rules, then a new do-not-connection customer pair is generated.

[0056] If the number of required connected object sets equals the specified threshold, no further merging of required connected object sets will be performed, and this step ends. This is to avoid inaccurate or even unreliable clustering results for customers due to too few generated required connected customer pairs. Alternatively, if any two required connected object sets in the required connected target set do not satisfy the merging rule, it means that further merging of required connected object sets is not possible, and this step ends. The specified threshold can be pre-set based on experience or obtained through a modularity algorithm.

[0057] Understandably, for two sets of objects that have already undergone the judgment of whether they meet the merging rules in step S202, step S202 will not be repeated. That is, any two sets of objects that must be connected will only be judged once to see if they meet the merging rules.

[0058] S205. Construct extended mandatory connection customer pairs based on each customer in each mandatory connection object set in the mandatory connection target set, and obtain the extended mandatory connection constraint subset; obtain the extended non-connection constraint subset based on the newly generated non-connection customer pairs and each non-connection customer pair in the non-connection constraint subset.

[0059] Specifically, any two customers in each set of mandatory connection objects within the mandatory connection target set can constitute a mandatory connection customer pair. The server obtains all combinations of two customers for each customer in each mandatory connection object set, resulting in an expanded mandatory connection customer pair corresponding to each mandatory connection object set. The expanded mandatory connection customer pairs corresponding to each mandatory connection object set form the expanded mandatory connection constraint subset. The server combines the newly generated non-connection customer pairs with each non-connection customer pair in the non-connection constraint subset to form an expanded non-connection constraint subset.

[0060] For example, the final set of mandatory target connections includes a set of mandatory object connections, D, which includes customers d1, d2, d3, and d4. Any two customers in the set of mandatory object connections, D, can be combined to obtain the following extended mandatory customer pairs: {customer d1, customer d2}, {customer d1, customer d3}, {customer d1, customer d4}, {customer d2, customer d3}, {customer d3, customer d4}, {customer d3, customer d4}.

[0061] Based on the above embodiments, the merging rules further include: any two customers in the two mandatory connection object sets do not constitute a pair of non-connected customers in the non-connection constraint subset; there are duplicate customers in the two mandatory connection object sets; and the distance between any two customers obtained based on the attribute feature data of any two customers in the two mandatory connection object sets is less than a preset value.

[0062] Specifically, any two customers in the two sets of mandatory objects do not constitute a pair of non-mandatory customers in the subset of non-mandatory constraints. That is, if a customer is randomly selected from one of the sets of mandatory objects and another customer is randomly selected from the other set of mandatory objects, these two customers will not be the same as the two customers in any pair of non-mandatory customers in the subset of non-mandatory constraints.

[0063] The customers in two sets of mandatory connected objects are compared. If the two sets contain the same customer, attribute feature data of any customer from one set and any customer from the other set are obtained. The distance between these two customers is calculated based on their attribute feature data. The distance between any two customers within the same set of mandatory connected objects is less than a preset value. If the distance between any two customers in both sets of mandatory connected objects is less than the preset value, and no two customers in these sets constitute a pair of non-connected customers in the non-connection constraint subset, then these two sets of mandatory connected objects satisfy the merging rule and can be merged into a single set of mandatory connected customers. The preset value is set according to actual needs and is not limited in this embodiment. The distance between two customers can be measured using geodesic distance.

[0064] Suppose X = {x1, ..., x} n Let M = {x} represent a dataset with n data objects, where L = {1, ..., k} is the corresponding set of class labels, and k represents the number of clusters for the n data objects. i ,x j ):y i =y j The set of constraints for a group of data objects that must be connected is denoted as {x, 1 ≤ i ≤ n, 1 ≤ j ≤ n}. i ,x j ) is a mandatory connection pair, y i For x i The corresponding class tag, y j For x j The corresponding class tag. Let C = {(x e ,x f ):y e =y f,1≤e≤n,1≤f≤n} represents a set of constraints for a set of unconnected data objects, (x e ,x f ) is a do-not-join constraint pair, y e For x e The corresponding class tag, y f For x f The corresponding class tag. In this embodiment of the invention, the data object is the customer.

[0065] Because a mandatory connection constraint represents an equivalence relationship between two data objects, it possesses reflexivity, symmetry, and transitivity. Based on the transitivity of mandatory connections, for data object x... i x j and x h The following relationship exists:

[0066] (x i ,x j )∈M,(x j ,x h )∈M=>(x i ,x h )∈M

[0067] (x i ,x j (x) is a mandatory connection pair. j ,x h If x is another required pair of constraints, then x i and x h This forms a mandatory pair of constraints, namely x i x j and x h There is an equivalence relation between them, x i x j and x h Those belonging to the same class label belong to the same category. Therefore, by extending all mandatory connection pairs through transitivity, we can obtain transitive closures corresponding to different class structures.

[0068] Based on the definitions of mandatory and non-mandatory constraints, pairwise constraints also have the following properties:

[0069] (x i ,x e )∈M,(x e ,x f )∈C=>(x i ,x f )∈C

[0070] (x i ,x e (x) is a mandatory connection pair. e ,x f) is a do-not-join constraint pair, since (x i ,x e ) and (x e ,x f ) Store a common data object x e Therefore (x i ,x f A do-not-join constraint pair can be formed. Therefore, if there is one or more do-not-join constraints between data objects in two transitive closures, the do-not-join constraints can be expanded by adding do-not-join constraints between all objects in the two transitive closures.

[0071] Although pairwise constraints can be effectively expanded based on transitivity and their properties, when the number of pairwise constraints is too small, the expanded constraints may have a negligible impact on the clustering performance compared to the initial pairwise constraints. To effectively expand pairwise constraints and address the impact of pairwise constraint sparsity on the performance of semi-supervised clustering algorithms, and to apply the expanded pairwise constraints to different semi-supervised clustering algorithms, further research is needed.

[0072] This invention modifies the similarity between different transitive closures to merge them in a safe manner, thereby expanding them into pairwise constraints. To better understand the technical solution of this application, a definition of transitive closure similarity is first given.

[0073] Let R = {r1, r2, ... r} m Let} represent the set of m transitive closures obtained by the transitivity of pre-defined paired constraints through mandatory connections, and S represent the m×m transitive closure similarity matrix. Two transitive closures r α ,r β The similarity between ∈R (α≠β) is defined as follows:

[0074]

[0075] Where d(r) α ,r β ) represents the distance between the α-th transitive closure and the β-th transitive closure, exp represents the exponential function with the natural constant e as the base, α is a positive integer and α is less than or equal to m, and β is a positive integer and β is less than or equal to m.

[0076]

[0077] a(r α ) represents r α The a-th data object in the data, b(r) β ) represents r β The b-th data object in the... Represents a(r) α ) and b(rβ The distance between r, a∈(1,2,…θ), b∈(1,2,…ρ), where θ represents r α The number of data objects in r, ρ represents r β The number of data objects in the data.

[0078] Since transitive closures are derived through the transitivity of pairwise constraints, the data objects within each transitive closure have the same class structure; that is, each transitive closure can be viewed as a known class. However, pairwise constraints represent the relationship between two data objects, and the class label of a data object cannot be inferred from pairwise constraints. Therefore, different transitive closures may correspond to the same class label.

[0079] Based on the transitivity of mandatory connections, it can be observed that similar transitive closures can be merged by calculating their similarity, and the property of pairwise constraints can be used to expand the constraints. Simultaneously, the connectivity of transitive closures is effectively utilized, using the shortest distance between data objects within two transitive closures as their overall distance. To better reflect the differences between data objects in the sample space, geodesic distance is preferably used as the metric for the distance between data objects. In this embodiment of the invention, the geodesic distance between two customers is calculated based on their attribute feature data.

[0080] a(r α ) and b(r β Geodetic distance between ) By constructing r α and r β The k-nearest neighbor graph of the data object is obtained by calculating the shortest path distance between two points in the graph.

[0081] While closer proximity between two transitive closures indicates greater similarity, directly merging two transitive closures and expanding pairwise constraints carries the risk of introducing erroneous pairwise constraints, thus affecting the clustering results. To more reliably merge transitive closures and minimize risky merging attempts, a preset value is set to determine whether to merge transitive closures.

[0082] The maximum geodesic distance between the data objects in the two transitive closures is compared with a preset value. If it is higher than the preset value, it indicates that merging the two transitive closures may lead to inaccurate clustering. Therefore, the distance between the two transitive closures is set to infinity. Otherwise, the original geodesic distance is retained.

[0083] Based on the above embodiments, the distance between any two customers is further obtained by calculating the geodesic distance between the two customers.

[0084] Specifically, the distance between any two customers in two necessarily connected object sets can be calculated by constructing a K-nearest neighbor graph of each customer in the two sets and then using the shortest path between the two customers as the distance between them. Geodesic distance can better reflect the global connectivity between the two customers. The shortest path between the two customers can be obtained by calculating the Euclidean distance using the attribute feature data of the two customers.

[0085] For example, for two necessarily connected object sets G and H, where G includes customers g1 and g2, and H includes customers g1, h1, and h2, since both G and H include the same customer g1, a K-nearest neighbor graph is constructed for customers g1, g2, h1, and h2 as follows: Figure 3 As shown. The distance between customer g2 and customer h1 is equal to the sum of the distance between customer g2 and customer g1 and the distance between customer g1 and customer h1. The distance between customer g2 and customer h2 is equal to the sum of the distance between customer g2 and customer g1 and the distance between customer g1 and customer h2. The distance between two adjacent customers in the K-nearest neighbor graph can be calculated using Euclidean distance based on the attribute data of these two customers.

[0086] Based on the above embodiments, the do-not-connect confirmation rule further includes the existence of two customers in the two mandatory-connection object sets that constitute a do-not-connection customer pair in the do-not-connection constraint subset.

[0087] Specifically, a customer is obtained from one of the two mandatory connection object sets, and a customer is obtained from the other mandatory connection object set. If a certain non-connect customer pair in the non-connect constraint subset includes the above two customers, it means that there are two customers in the two mandatory connection object sets that constitute a non-connect customer pair in the non-connect constraint subset. Then, the two mandatory connection object sets satisfy the non-connect confirmation rule.

[0088] Based on the above embodiments, the quantity threshold is further obtained by dividing each pair of mandatory customer pairs according to the attribute feature data of customers in each mandatory customer pair in the mandatory constraint subset and the modularity algorithm.

[0089] Specifically, each pair of mandatory connected clients in the mandatory connection constraint subset is treated as an independent community, and each community is treated as a node, forming an initial network. The criterion of maximizing modularity increment determines which adjacent communities should be merged. After the first phase of scanning, the second phase begins, treating the merged communities and unmerged communities from the first phase as nodes again to construct a new network. The first phase is then repeated on this new network. These two phases are repeated until the modularity of the network's community divisions no longer increases; the number of communities in the resulting network is the aforementioned threshold. The modularity of the initial network is calculated using the attribute feature data of the two clients in each mandatory connected client pair. The modularity of the network after merging communities is calculated using the attribute feature data of each client at each node.

[0090] Based on the above embodiments, the customer classification method provided by the embodiments of the present invention further includes:

[0091] If it is determined that the two sets of objects that must be connected do not satisfy the merging rule and do not satisfy the do not connect confirmation rule, then the set of objects that must be connected remains unchanged.

[0092] Specifically, the server will determine whether the two sets of mandatory connections meet the merging rule based on the attribute feature data of each customer in the two sets of mandatory connections and the pairs of non-connecting customers in the non-connecting constraint subset. If the merging rule is not met, the server will continue to determine whether the customers in the two sets of mandatory connections meet the non-connecting confirmation rule. If the customers in the two sets of mandatory connections also do not meet the non-connecting confirmation rule, it means that there is no need to update the set of mandatory connections, and the set of mandatory connections remains unchanged.

[0093] Compared to existing methods that expand pairwise constraints based on transitivity, the customer classification method provided in this invention can effectively expand pairwise constraints and apply the expanded pairwise constraints to different semi-supervised clustering algorithms. Furthermore, the customer classification method provided in this invention can expand pairwise constraints more reliably, rather than simply pursuing a higher number of expanded pairwise constraints.

[0094] This invention expands upon prior information derived from a small number of known customers of the same type, utilizing the similarities between various customer attributes to further enrich this prior information, thereby achieving customer classification. Based on the preferences of different customer types for banking products, suitable banking products can be recommended to different customer groups. This invention has significant application value for analyzing existing customers and providing them with better financial product recommendations.

[0095] Figure 4 This is a schematic diagram of the customer sorting device provided in the fourth embodiment of the present invention, as shown below. Figure 4As shown, the customer classification device provided in this embodiment of the invention includes an acquisition module 401, a preprocessing module 402, an expansion module 403, and a clustering module 404, wherein:

[0096] The acquisition module 401 is used to acquire the attribute information of each customer in the prior customer pair constraint set and the mandatory connection constraint subset; wherein, the pair constraint set includes the mandatory connection constraint subset and the non-connection constraint subset; the mandatory connection constraint subset includes a first preset number of mandatory connection customer pairs, and the non-connection constraint subset includes a second preset number of non-connection customer pairs; the preprocessing module 402 is used to preprocess the attribute information of each customer in the mandatory connection constraint subset to obtain the attribute feature data of each customer in the pair constraint set; the expansion module 403 is used to... Based on the required customer pairs in the required connection constraint subset, the non-connected customer pairs in the non-connected constraint subset, and the attribute feature data of each customer in the required connection constraint subset, the required connection constraint subset is expanded to obtain an expanded required connection constraint subset, and the non-connected constraint subset is expanded to obtain an expanded non-connected constraint subset; the clustering module 404 is used to perform customer clustering based on the expanded required connection constraint subset, the expanded non-connected constraint subset, and the attribute feature data of the customers to be classified, using a pairwise constraint semi-supervised clustering algorithm to obtain the customer classification results.

[0097] Specifically, a set of paired constraints for prior customers and attribute information of each customer in the paired constraint set are collected in advance. The acquisition module 401 can acquire the set of paired constraints for prior customers and attribute information of each customer in the paired constraint set. The set of paired constraints includes a mandatory constraint subset and a non-mandatory constraint subset. The mandatory constraint subset includes a first preset number of mandatory customer pairs, each pair consisting of two customers belonging to the same category. The non-mandatory constraint subset includes a second preset number of non-mandatory customer pairs, each pair consisting of two customers belonging to different categories. The first preset number is set according to actual needs, and this embodiment of the invention does not limit it. The second preset number is also set according to actual needs, and this embodiment of the invention does not limit it.

[0098] The preprocessing module 402 preprocesses the attribute information of each customer in the pairwise constraint set to obtain the attribute feature data of each customer in the pairwise constraint set. The preprocessing includes, but is not limited to, normalization, one-hot encoding, and binarization of the data, and is set according to actual needs; this embodiment of the invention does not impose limitations on these settings.

[0099] The expansion module 403 expands the mandatory connection constraint subset based on each mandatory connection customer pair in the mandatory connection constraint subset, each non-connection customer pair in the non-connection constraint subset, and the attribute feature data of each customer in the mandatory connection constraint subset, to obtain more mandatory connection customer pairs to form an expanded non-connection constraint subset, and at the same time obtain more non-connection customer pairs to form an expanded non-connection constraint subset.

[0100] The clustering module 404 uses each required customer pair in the expanded required-connection constraint subset and each non-connected customer pair in the expanded non-connection constraint subset as pairwise constraints in a pairwise constraint semi-supervised clustering algorithm to cluster the attribute feature data of the customers to be classified, thereby obtaining the customer classification results. The customer classification results may include multiple customer categories and customers under each category. The specific process of clustering data using the pairwise constraint semi-supervised clustering algorithm is existing technology and will not be elaborated here.

[0101] The customer classification device provided in this embodiment of the invention can acquire the attribute information of each customer in the prior customer pair constraint set and the mandatory constraint subset. It preprocesses the attribute information of each customer in the mandatory constraint subset to obtain attribute feature data for each customer in the pair constraint set. Based on the mandatory customer pairs in the mandatory constraint subset, the non-mandatory customer pairs in the non-mandatory constraint subset, and the attribute feature data of each customer in the mandatory constraint subset, it expands the mandatory constraint subset to obtain an expanded mandatory constraint subset and expands the non-mandatory constraint subset to obtain an expanded non-mandatory constraint subset. Based on the expanded mandatory constraint subset, the expanded non-mandatory constraint subset, and the attribute feature data of the customers to be classified, it performs customer clustering using a pair constraint semi-supervised clustering algorithm to obtain the customer classification result. Because it adds the pair constraint information required in the pair constraint semi-supervised clustering algorithm, it improves the accuracy of customer classification.

[0102] Figure 5 This is a schematic diagram of the customer sorting device provided in the fifth embodiment of the present invention, as shown below. Figure 5 As shown, based on the above embodiments, the expansion module 403 further includes an acquisition unit 4031, a first judgment unit 4032, an update unit 4033, a second judgment unit 4034, and an expansion unit 4035, wherein:

[0103] The acquisition unit 4031 is used to acquire two sets of mandatory connection objects from the mandatory connection target set; wherein, the initial data of the mandatory connection target set is the mandatory connection constraint subset, and the mandatory connection object set is a mandatory connection customer pair or a mandatory connection customer set; the first judgment unit 4032 is used to, if it is determined based on the attribute feature data of each customer in the two mandatory connection object sets and each non-connection customer pair in the non-connection constraint subset that the two mandatory connection object sets meet the merging rule, then merge the customers in the two mandatory connection object sets into a mandatory connection customer set; if it is determined that the customers in the two mandatory connection object sets meet the non-connection confirmation rule, then generate a new non-connection customer pair based on the customers in the two mandatory connection object sets; the update unit 4033 is used to update the mandatory connection target set based on the remaining mandatory connection object sets in the mandatory connection target set and the merged mandatory connection customer set; wherein The remaining set of mandatory objects in the mandatory target set refers to the set of mandatory objects remaining after removing the two sets of mandatory objects used to merge into a set of mandatory customers from the mandatory target set. The second judgment unit 4034 is used to continue to obtain two sets of mandatory objects from the mandatory target set for merging judgment if it is determined that the number of mandatory object sets in the mandatory target set is greater than the number threshold, until the number of mandatory object sets in the mandatory target set is equal to the number threshold or any two sets of mandatory objects in the mandatory target set do not meet the merging rule. The expansion unit 4035 is used to construct expanded mandatory customer pairs according to each customer in each set of mandatory objects in the mandatory target set to obtain the expanded mandatory constraint subset; and to obtain the expanded non-connection constraint subset according to the newly generated non-connection customer pairs and each non-connection customer pair in the non-connection constraint subset.

[0104] Based on the above embodiments, the merging rules further include: any two customers in the two mandatory connection object sets do not constitute a pair of non-connected customers in the non-connection constraint subset; there are duplicate customers in the two mandatory connection object sets; and the distance between any two customers obtained based on the attribute feature data of any two customers in the two mandatory connection object sets is less than a preset value.

[0105] Based on the above embodiments, the distance between any two customers is further obtained by calculating the geodesic distance between the two customers.

[0106] Based on the above embodiments, the do-not-connect confirmation rule further includes the existence of two customers in the two mandatory-connection object sets that constitute a do-not-connection customer pair in the do-not-connection constraint subset.

[0107] Based on the above embodiments, the quantity threshold is further obtained by dividing each mandatory customer pair into components according to the attribute feature data of the customers in each mandatory customer pair in the mandatory constraint subset and the modularity algorithm.

[0108] Figure 6 This is a schematic diagram of the customer sorting device provided in the sixth embodiment of the present invention, as shown below. Figure 6 As shown, based on the above embodiments, the customer sorting device provided in this embodiment of the invention further includes:

[0109] The holding unit 4036 is used to keep the set of required targets unchanged if it is determined that the two sets of required objects do not meet the merging rule and do not meet the do not connect confirmation rule.

[0110] The embodiments of the device provided in this invention can be used to execute the processing flow of the above-described method embodiments. Its functions will not be repeated here, but can be referred to the detailed description of the above-described method embodiments.

[0111] It should be noted that the customer classification method and apparatus provided in the embodiments of the present invention can be used in the financial field, or in any technical field other than the financial field. The embodiments of the present invention do not limit the application field of the customer classification method and apparatus.

[0112] Figure 7 This is a schematic diagram of the physical structure of the electronic device provided in the seventh embodiment of the present invention, as shown below. Figure 7 As shown, the electronic device may include: a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704. The processor 701 can call logical instructions in the memory 703 to execute the following method: obtaining the set of paired constraints for prior customers and the attribute information of each customer in the mandatory-connection constraint subset; wherein the set of paired constraints includes the mandatory-connection constraint subset and the non-connection constraint subset; the mandatory-connection constraint subset includes a first preset number of mandatory-connection customer pairs, and the non-connection constraint subset includes a second preset number of non-connection customer pairs; preprocessing the attribute information of each customer in the mandatory-connection constraint subset to obtain attribute feature data for each customer in the set of paired constraints; expanding the mandatory-connection constraint subset to obtain an expanded mandatory-connection constraint subset and expanding the non-connection constraint subset to obtain an expanded non-connection constraint subset based on the mandatory-connection constraint subset, the expanded non-connection constraint subset, and the attribute feature data of each customer in the mandatory-connection constraint subset; and performing customer clustering using a pairwise constraint semi-supervised clustering algorithm based on the expanded mandatory-connection constraint subset, the expanded non-connection constraint subset, and the attribute feature data of the customers to be classified, to obtain the customer classification result.

[0113] Furthermore, the logical instructions in the aforementioned memory 703 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0114] This embodiment discloses a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by the computer, the computer can perform the methods provided in the above-described method embodiments, such as: obtaining the attribute information of each customer in the prior customer pair constraint set and the mandatory connection constraint subset; wherein, the pair constraint set includes the mandatory connection constraint subset and the non-connection constraint subset; the mandatory connection constraint subset includes a first preset number of mandatory connection customer pairs, and the non-connection constraint subset includes a second preset number of non-connection customer pairs; and the mandatory connection constraints... The attribute information of each customer in the subset is preprocessed to obtain the attribute feature data of each customer in the pairwise constraint set. Based on the attribute feature data of each mandatory customer pair in the mandatory constraint subset, each non-mandatory customer pair in the non-mandatory constraint subset, and each customer in the mandatory constraint subset, the mandatory constraint subset is expanded to obtain an expanded mandatory constraint subset, and the non-mandatory constraint subset is expanded to obtain an expanded non-mandatory constraint subset. Based on the expanded mandatory constraint subset, the expanded non-mandatory constraint subset, and the attribute feature data of the customers to be classified, customer clustering is performed using a pairwise constraint semi-supervised clustering algorithm to obtain the customer classification results.

[0115] This embodiment provides a computer-readable storage medium storing a computer program that causes a computer to execute the methods provided in the above-described method embodiments. For example, the methods include: obtaining a set of pairwise constraints for prior clients and attribute information of each client in a subset of mandatory connection constraints; wherein the set of pairwise constraints includes the subset of mandatory connection constraints and a subset of non-connection constraints; the subset of mandatory connection constraints includes a first preset number of mandatory client pairs, and the subset of non-connection constraints includes a second preset number of non-connection client pairs; and preprocessing the attribute information of each client in the subset of mandatory connection constraints. The algorithm first obtains the attribute feature data of each customer in the pairwise constraint set; then, based on the attribute feature data of each mandatory customer pair in the mandatory constraint subset, each non-mandatory customer pair in the non-mandatory constraint subset, and each customer in the mandatory constraint subset, it expands the mandatory constraint subset to obtain an expanded mandatory constraint subset and expands the non-mandatory constraint subset to obtain an expanded non-mandatory constraint subset; finally, it performs customer clustering using a pairwise constraint semi-supervised clustering algorithm based on the expanded mandatory constraint subset, the expanded non-mandatory constraint subset, and the attribute feature data of the customers to be classified, to obtain the customer classification results.

[0116] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0118] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0119] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0120] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0121] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A customer classification method, characterized in that, include: Obtain the set of paired constraints for prior customers and the attribute information of each customer in the mandatory connection constraint subset; wherein, the set of paired constraints includes the mandatory connection constraint subset and the non-connection constraint subset; the mandatory connection constraint subset includes a first preset number of mandatory connection customer pairs, and the non-connection constraint subset includes a second preset number of non-connection customer pairs; Preprocess the attribute information of each customer in the mandatory constraint subset to obtain the attribute feature data of each customer in the pairwise constraint set; Obtain two sets of mandatory objects from the set of mandatory targets; wherein the initial data of the set of mandatory targets is the subset of mandatory constraints, and the set of mandatory objects is a pair of mandatory clients or a set of mandatory clients; If, based on the attribute feature data of each customer in the two mandatory connection object sets and the non-connection customer pairs in the non-connection constraint subset, it is determined that the two mandatory connection object sets satisfy the merging rule, then the customers in the two mandatory connection object sets are merged into one mandatory connection customer set; if it is determined that the customers in the two mandatory connection object sets satisfy the non-connection confirmation rule, then a new non-connection customer pair is generated based on the customers in the two mandatory connection object sets; wherein, the merging rule includes that any two customers in the two mandatory connection object sets do not constitute a non-connection customer pair in the non-connection constraint subset; there are duplicate customers in the two mandatory connection object sets, and the distance between any two customers obtained based on the attribute feature data of any two customers in the two mandatory connection object sets is less than a preset value. The set of mandatory targets is updated based on the remaining set of mandatory objects in the set of mandatory targets and the merged set of mandatory customers; wherein, the remaining set of mandatory objects in the set of mandatory targets refers to the set of mandatory objects remaining after removing the two sets of mandatory objects used to merge into the set of mandatory customers from the set of mandatory targets; If it is determined that the number of required object sets in the required target set is greater than the number threshold, then two more required object sets are obtained from the required target set for merging and judgment, until the number of required object sets in the required target set is equal to the number threshold or any two required object sets in the required target set do not meet the merging rule; Based on each customer in the set of required connected objects in the set of required connected targets, construct an expanded set of required connected customers to obtain an expanded set of required connected constraints; based on the newly generated set of non-connected customers and each set of non-connected customers in the set of non-connected constraints, obtain an expanded set of non-connected constraints. Based on the extended mandatory-connection constraint subset, the extended non-connection constraint subset, and the attribute feature data of the customers to be classified, customer clustering is performed using a pairwise constraint semi-supervised clustering algorithm to obtain the customer classification results.

2. The method according to claim 1, characterized in that, The distance between any two customers is obtained by calculating the geodesic distance between the two customers.

3. The method according to claim 1, characterized in that, The do-not-connect confirmation rule includes the existence of two customers in the two sets of mandatory-connection objects that constitute a do-not-connection customer pair in the subset of the do-not-connection constraint.

4. The method according to claim 1, characterized in that, The quantity threshold is obtained by dividing each mandatory customer pair into components based on the attribute feature data of the customers in each mandatory customer pair in the mandatory constraint subset and the modularity algorithm.

5. The method according to any one of claims 1 to 4, characterized in that, Also includes: If it is determined that the two sets of objects that must be connected do not satisfy the merging rule and do not satisfy the do not connect confirmation rule, then the set of objects that must be connected remains unchanged.

6. A customer sorting device, characterized in that, include: The acquisition module is used to acquire the attribute information of each customer in the prior customer pair constraint set and the mandatory connection constraint subset; wherein, the pair constraint set includes the mandatory connection constraint subset and the non-connection constraint subset; the mandatory connection constraint subset includes a first preset number of mandatory connection customer pairs, and the non-connection constraint subset includes a second preset number of non-connection customer pairs; The preprocessing module is used to preprocess the attribute information of each customer in the mandatory constraint subset to obtain the attribute feature data of each customer in the pairwise constraint set. An expansion module is used to expand the mandatory connection constraint subset to obtain an expanded mandatory connection constraint subset and expand the non-connection constraint subset to obtain an expanded non-connection constraint subset based on each mandatory connection customer pair in the mandatory connection constraint subset, each non-connection customer pair in the non-connection constraint subset, and the attribute feature data of each customer in the mandatory connection constraint subset. The clustering module is used to cluster customers based on the extended mandatory connection constraint subset, the extended non-connection constraint subset, and the attribute feature data of the customers to be classified, using a pairwise constraint semi-supervised clustering algorithm to obtain the customer classification results. The expansion module includes: The acquisition unit is used to acquire two sets of mandatory objects from the set of mandatory targets; wherein the initial data of the set of mandatory targets is the subset of mandatory constraints, and the set of mandatory objects is a pair of mandatory clients or a set of mandatory clients; The first judgment unit is configured to, if it is determined based on the attribute feature data of each customer in the two mandatory connection object sets and the non-connection customer pairs in the non-connection constraint subset that the two mandatory connection object sets satisfy the merging rule, then merge the customers in the two mandatory connection object sets into a mandatory connection customer set; if it is determined that the customers in the two mandatory connection object sets satisfy the non-connection confirmation rule, then generate a new non-connection customer pair based on the customers in the two mandatory connection object sets; wherein, the merging rule includes that any two customers in the two mandatory connection object sets do not constitute a non-connection customer pair in the non-connection constraint subset; there are duplicate customers in the two mandatory connection object sets, and the distance between any two customers obtained based on the attribute feature data of any two customers in the two mandatory connection object sets is less than a preset value. The update unit is used to update the set of mandatory targets based on the remaining set of mandatory objects in the set of mandatory targets and the merged set of mandatory customers; wherein, the remaining set of mandatory objects in the set of mandatory targets refers to the set of mandatory objects remaining after removing the two sets of mandatory objects used to merge into the set of mandatory customers from the set of mandatory targets; The second judgment unit is used to determine if the number of ... An expansion unit is used to construct expanded mandatory connection customer pairs based on each customer in each mandatory connection object set in the mandatory connection target set, thereby obtaining the expanded mandatory connection constraint subset; and to obtain an expanded non-connection constraint subset based on newly generated non-connection customer pairs and each non-connection customer pair in the non-connection constraint subset.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 5.

9. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Weibo user gender deduction method and system based on combined theme

    CN106327341A

  • Semi-supervised clustering method combining pairwise constraint and scale constraint

    CN108446736A