Customer group segmentation method, device, equipment and storage medium
By clustering and feature identification of user groups, and combining user behavior to generate group classification model, the problem of being unable to accurately divide user groups in the existing technology is solved, and high-accurate customer group division is achieved.
Patent Information
- Application Number
- CN202210703473.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-06-21
AI Technical Summary
When dividing user groups, the prior art ignores the characteristics of users in the group, resulting in the inability to accurately realize group division.
By receiving customer group division requests, obtaining attribute information of users to be classified, performing clustering processing to generate initial groups, identifying group characteristics of each user to be classified, generating group classification models, collecting user behavior information, identifying group categories and merging users to obtain customer groups.
Accurate customer group division based on user characteristics and behaviors is realized, and the accuracy and effectiveness of group division are improved.
Smart Images

Figure CN115222443B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a customer group segmentation method, device, equipment and storage medium. Background Art
[0002] In order to improve the activity experience and conversion rate, it is usually necessary to analyze the potential customers of user groups. At present, when dividing user groups, users are usually grouped based on pre-configured indicators, while ignoring the characteristics of users in groups, resulting in the inability to accurately achieve group division. Summary of the invention
[0003] In view of the above, it is necessary to provide a customer group segmentation method, device, equipment and storage medium that can solve the technical problem of being unable to accurately implement group segmentation.
[0004] In one aspect, the present invention provides a method for dividing customer groups, the method comprising:
[0005] Receiving a customer group classification request, and acquiring attribute information of a plurality of users to be classified according to the customer group classification request;
[0006] Clustering the multiple users to be classified according to the attribute information to obtain multiple initial groups;
[0007] Identifying group characteristics corresponding to each user to be classified based on the multiple initial groups;
[0008] Generate a group classification model for each group feature according to the user behaviors of multiple annotated users on a preset platform and the annotated category of each annotated user;
[0009] Collecting behavior information of each user to be classified on the preset platform;
[0010] Identify the behavior information based on the group classification model to obtain the group category corresponding to each user to be classified;
[0011] The multiple users to be classified are merged based on the group category to obtain a customer group.
[0012] According to a preferred embodiment of the present invention, clustering the multiple users to be classified according to the attribute information to obtain multiple initial groups includes:
[0013] Encoding the attribute information to obtain a representation vector for each user to be classified;
[0014] extracting a plurality of first users from the plurality of users to be classified according to the number of users of the plurality of users to be classified, and determining the remaining users of the plurality of users to be classified except the plurality of first users as a plurality of second users;
[0015] For any second user, calculating the user similarity between the any second user and each first user according to the characterization vector of the any second user and the characterization vector of each first user;
[0016] Classify the arbitrary second user into the category of the first user with the highest user similarity, to obtain multiple basic groups;
[0017] The multiple basic groups are merged based on the characterization vectors of group users in each basic group to obtain the multiple initial groups.
[0018] According to a preferred embodiment of the present invention, merging the multiple basic groups based on the characterization vectors of group users in each basic group to obtain the multiple initial groups includes:
[0019] Calculate the average value of the representation vectors of group users in each basic group as the group feature of each basic group;
[0020] Calculating the group similarity between any two groups in the multiple basic groups based on the group characteristics;
[0021] The multiple basic groups are merged based on the group similarities to obtain the multiple initial groups.
[0022] According to a preferred embodiment of the present invention, the identifying the group characteristics corresponding to each user to be classified based on the multiple initial groups includes:
[0023] Calculate the average value of the representation vectors of group users in each initial group as the encoding information of each initial group;
[0024] Extracting target features from the encoded information based on a pre-trained convolutional model;
[0025] The attribute factor corresponding to the target feature is determined as the group feature.
[0026] According to a preferred embodiment of the present invention, the group classification model for generating each group feature according to the user behaviors of multiple annotated users on a preset platform and the annotated category of each annotated user includes:
[0027] For each group feature, collecting the buried point information of the marked users corresponding to the group feature from the preset platform as the user behavior;
[0028] Classifying the user behaviors into events to obtain the event category corresponding to each marked user;
[0029] For each event category, determine the annotated user of the event category as the event user;
[0030] Based on the marked category, a first number of the event users that are positive samples is counted, and based on the marked category, a second number of the event users that are negative samples is counted;
[0031] Based on the annotation category, counting a third number of the plurality of annotated users belonging to positive samples, and counting a fourth number of the plurality of annotated users belonging to negative samples;
[0032] Calculate the category importance of each event category based on the first number, the second number, the third number, and the fourth number;
[0033] Construct a decision layer for each event category based on the user behavior corresponding to each event category and the annotated category corresponding to the annotated user of the user behavior;
[0034] The group classification model is obtained by splicing the multiple decision layers in descending order of the importance of the categories.
[0035] According to a preferred embodiment of the present invention, the calculating the category importance of each event category based on the first number, the second number, the third number, and the fourth number includes:
[0036] The interference degree of each event category is calculated based on the first number and the second number, and the calculation formula of the interference degree is: Wherein, R represents the interference degree, n1 represents the first number, and n2 represents the second number;
[0037] The target importance is calculated based on the third number and the fourth number. If the third number is less than the fourth number, the calculation formula of the target importance is:
[0038] Wherein, K represents the target importance, n3 represents the third number, and n4 represents the fourth number;
[0039] The difference between the target importance and each interference degree is calculated to obtain the category importance of each event category.
[0040] According to a preferred embodiment of the present invention, the behavior information is identified based on the group classification model to obtain the group category corresponding to each user to be classified, including:
[0041] For each user to be classified, a group classification model corresponding to the group characteristics of the user to be classified is selected as a target classification model, and group classification models other than the characteristic classification model are determined as a plurality of characteristic classification models;
[0042] Identify the behavior information based on the first decision layer in the target classification model to obtain an identification result;
[0043] If the recognition result is not the preset result, a second decision layer is selected from the target classification model according to the category importance to identify the behavior information, until the termination decision layer with the smallest category importance in the target classification model completes the recognition of the behavior information, or the recognized result is the preset result, and the target prediction result of the target classification model for the user to be classified is obtained;
[0044] Identify the behavior information based on the multiple feature classification models to obtain feature prediction results of each feature classification model for the user to be classified;
[0045] The target prediction result and the multiple feature prediction results are weighted and calculated based on a first preset weight and multiple second preset weights to obtain the group category, and the first preset weight is greater than the multiple second preset weights.
[0046] On the other hand, the present invention further provides a customer group classification device, the customer group classification device comprising:
[0047] An acquisition unit, configured to receive a customer group classification request, and acquire attribute information of a plurality of users to be classified according to the customer group classification request;
[0048] A clustering unit, configured to perform clustering processing on the plurality of users to be classified according to the attribute information to obtain a plurality of initial groups;
[0049] an identification unit, configured to identify a group feature corresponding to each user to be classified based on the multiple initial groups;
[0050] A generating unit, used to generate a group classification model of each group feature according to user behaviors of multiple annotated users on a preset platform and annotated categories of each annotated user;
[0051] A collection unit, used to collect behavior information of each user to be classified on the preset platform;
[0052] The identification unit is further used to identify the behavior information based on the group classification model to obtain the group category corresponding to each user to be classified;
[0053] A merging unit is used to merge the multiple users to be classified based on the group category to obtain a customer group.
[0054] On the other hand, the present invention further provides an electronic device, comprising:
[0055] a memory storing computer readable instructions; and
[0056] A processor executes the computer-readable instructions stored in the memory to implement the customer group segmentation method.
[0057] On the other hand, the present invention further proposes a computer-readable storage medium, in which computer-readable instructions are stored. The computer-readable instructions are executed by a processor in an electronic device to implement the customer group segmentation method.
[0058] It can be seen from the above technical solutions that the present invention clusters the users to be classified based on the attribute information, and can quickly generate the multiple initial groups. Then, combined with the group characteristics corresponding to each user to be classified and the behavior information of the user to be classified on the preset platform, the group category can be accurately detected, thereby improving the accuracy of the division of the customer group. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flow chart of a preferred embodiment of the method for dividing customer groups of the present invention.
[0060] Figure 2 It is a functional module diagram of a preferred embodiment of the customer group segmentation device of the present invention.
[0061] Figure 3 It is a structural schematic diagram of an electronic device of a preferred embodiment of the method for implementing customer group segmentation of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0063] like Figure 1 FIG. 1 is a flow chart of a preferred embodiment of the method for dividing customer groups of the present invention. According to different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted.
[0064] The customer group segmentation method can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0065] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0066] The customer group segmentation method is applied to one or more electronic devices, which are devices that can automatically perform numerical calculations and / or information processing according to pre-set or stored computer-readable instructions, and whose hardware includes but is not limited to microprocessors, application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), digital signal processors (DSP), embedded devices, etc.
[0067] The electronic device may be any electronic product that can perform human-computer interaction with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.
[0068] The electronic device may include a network device and / or a user device. The network device includes, but is not limited to, a single network electronic device, an electronic device group consisting of multiple network electronic devices, or a cloud consisting of a large number of hosts or network electronic devices based on cloud computing.
[0069] The network where the electronic device is located includes, but is not limited to: the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0070] 101 , receiving a customer group classification request, and acquiring attribute information of a plurality of users to be classified according to the customer group classification request.
[0071] In at least one embodiment of the present invention, the customer group classification request may be triggered and generated by a business person of a preset platform. The customer group classification request carries user identification codes of the multiple users to be classified.
[0072] In at least one embodiment of the present invention, the attribute information refers to the information corresponding to the multiple users to be classified on preset attribute factors, wherein the preset attribute factors may include, but are not limited to: age, zodiac sign, date of birth, height, gender and device model, etc.
[0073] In at least one embodiment of the present invention, the electronic device may obtain attribute information of the multiple users to be classified from a user information database.
[0074] 102 , clustering the multiple users to be classified according to the attribute information to obtain multiple initial groups.
[0075] In at least one embodiment of the present invention, the multiple initial groups refer to groups obtained after initially dividing the multiple users to be classified according to the similarity of the attribute information.
[0076] In at least one embodiment of the present invention, the electronic device performs clustering processing on the multiple users to be classified according to the attribute information, and obtains multiple initial groups including:
[0077] Encoding the attribute information to obtain a representation vector for each user to be classified;
[0078] extracting a plurality of first users from the plurality of users to be classified according to the number of users of the plurality of users to be classified, and determining the remaining users of the plurality of users to be classified except the plurality of first users as a plurality of second users;
[0079] For any second user, calculating the user similarity between the any second user and each first user according to the characterization vector of the any second user and the characterization vector of each first user;
[0080] Classify the arbitrary second user into the category of the first user with the highest user similarity, to obtain multiple basic groups;
[0081] The multiple basic groups are merged based on the characterization vectors of group users in each basic group to obtain the multiple initial groups.
[0082] The number of users may be the total number of the plurality of users to be classified at a preset ratio, for example, if the total number of users is 100 and the preset ratio is 0.1, the number of users is 10. The number of users is usually an integer greater than 1.
[0083] The number of the plurality of first users is the number of users.
[0084] The user similarity may be calculated based on the cosine value of the characterization vector of any second user and the characterization vector of each first user.
[0085] By encoding the user to be classified through the attribute information, the comprehensiveness of the characterization of the user to be classified by the characterization vector can be improved. The user similarity between any second user and each first user can be accurately quantified through the characterization vector, so that based on the user similarity, any second user can be reasonably classified into the corresponding first user, thereby improving the accuracy of the multiple basic groups, thereby improving the clustering accuracy of the multiple initial groups.
[0086] Specifically, the electronic device merges the multiple basic groups based on the characterization vectors of group users in each basic group to obtain the multiple initial groups including:
[0087] Calculate the average value of the representation vectors of group users in each basic group as the group feature of each basic group;
[0088] Calculating the group similarity between any two groups in the multiple basic groups based on the group characteristics;
[0089] The multiple basic groups are merged based on the group similarities to obtain the multiple initial groups.
[0090] The similarity between any two groups among the multiple initial groups is less than a preset similarity threshold, and the preset similarity threshold can be set according to actual needs.
[0091] By merging the multiple basic groups according to the group similarity, it is possible to avoid the situation where there are groups with too high similarity among the multiple initial groups, which may cause the subsequent inability to accurately extract group features.
[0092] 103 : Identify a group feature corresponding to each user to be classified based on the multiple initial groups.
[0093] In at least one embodiment of the present invention, the group characteristics refer to key attribute factors corresponding to the initial group to which each user to be classified belongs.
[0094] In at least one embodiment of the present invention, the electronic device identifying the group characteristics corresponding to each user to be classified based on the multiple initial groups includes:
[0095] Calculate the average value of the representation vectors of group users in each initial group as the encoding information of each initial group;
[0096] Extracting target features from the encoded information based on a pre-trained convolutional model;
[0097] The attribute factor corresponding to the target feature is determined as the group feature.
[0098] Through the above implementation, the influence of irrelevant features on customer group segmentation can be eliminated.
[0099] Specifically, the electronic device extracts target features from the encoded information based on a pre-trained convolutional model, including:
[0100] Obtaining a model threshold in the convolutional model;
[0101] A coding value having a value greater than the model threshold is obtained from the coding information as the target feature.
[0102] 104 , generating a group classification model of each group feature according to user behaviors of multiple annotated users on a preset platform and the annotated category of each annotated user.
[0103] In at least one embodiment of the present invention, the preset platform may be any operating platform, and the preset platform may include multiple application programs.
[0104] The user behavior may include login operations of the multiple marked users on a game application in the preset platform, etc.
[0105] The marked category may include potential for purchase of a certain product, etc.
[0106] The group classification model includes a first decision layer and a plurality of second decision layers, wherein the category importance of the first decision layer is greater than the category importance of the plurality of second decision layers.
[0107] In at least one embodiment of the present invention, the electronic device generates a group classification model for each group feature according to the user behaviors of multiple annotated users on a preset platform and the annotated category of each annotated user, including:
[0108] For each group feature, collecting the buried point information of the marked users corresponding to the group feature from the preset platform as the user behavior;
[0109] Classifying the user behaviors into events to obtain the event category corresponding to each marked user;
[0110] For each event category, determine the annotated user of the event category as the event user;
[0111] Based on the marked category, a first number of the event users that are positive samples is counted, and based on the marked category, a second number of the event users that are negative samples is counted;
[0112] Based on the annotation category, counting a third number of the plurality of annotated users belonging to positive samples, and counting a fourth number of the plurality of annotated users belonging to negative samples;
[0113] Calculate the category importance of each event category based on the first number, the second number, the third number, and the fourth number;
[0114] Construct a decision layer for each event category based on the user behavior corresponding to each event category and the annotated category corresponding to the annotated user of the user behavior;
[0115] The group classification model is obtained by splicing the multiple decision layers in descending order of the importance of the categories.
[0116] The event categories include login events, purchase events, browsing events, etc.
[0117] The positive samples refer to users whose labeled category is that they have the potential to purchase a certain product, and the negative samples refer to users whose labeled category is that they do not have the potential to purchase a certain product.
[0118] By combining the first number, the second number, the third number and the fourth number, the category importance of each event category can be accurately quantified, and further by splicing multiple decision layers according to the category importance, the accuracy of the group classification model can be improved.
[0119] Specifically, the electronic device calculates the category importance of each event category based on the first number, the second number, the third number, and the fourth number, including:
[0120] The interference degree of each event category is calculated based on the first number and the second number, and the calculation formula of the interference degree is: Wherein, R represents the interference degree, n1 represents the first number, and n2 represents the second number;
[0121] The target importance is calculated based on the third number and the fourth number. If the third number is less than the fourth number, the calculation formula of the target importance is:
[0122] Wherein, K represents the target importance, n3 represents the third number, and n4 represents the fourth number;
[0123] The difference between the target importance and each interference degree is calculated to obtain the category importance of each event category.
[0124] The category importance of each event category can be accurately quantified through the relationship between the target importance and each interference degree.
[0125] Specifically, the electronic device constructs a behavior category mapping relationship according to the user behavior corresponding to each event category and the annotated category corresponding to the annotated user of the user behavior, and then generates a decision layer for each event category according to the category mapping relationship.
[0126] 105. Collecting behavior information of each user to be classified on the preset platform.
[0127] In at least one embodiment of the present invention, the behavior information can be obtained based on information embedding for each user to be classified on the preset platform.
[0128] 106 , identifying the behavior information based on the group classification model to obtain a group category corresponding to each user to be classified.
[0129] It should be emphasized that in order to further ensure the privacy and security of the above group categories, the above group categories can also be stored in a node of a blockchain.
[0130] In at least one embodiment of the present invention, the group category refers to the category corresponding to the to-be-classified users, and the group category includes whether there is a purchasing potential for a certain product.
[0131] In at least one embodiment of the present invention, the electronic device identifies the behavior information based on the group classification model, and obtains the group category corresponding to each user to be classified, including:
[0132] For each user to be classified, a group classification model corresponding to the group characteristics of the user to be classified is selected as a target classification model, and group classification models other than the characteristic classification model are determined as a plurality of characteristic classification models;
[0133] Identify the behavior information based on the first decision layer in the target classification model to obtain an identification result;
[0134] If the recognition result is not the preset result, a second decision layer is selected from the target classification model according to the category importance to identify the behavior information, until the termination decision layer with the smallest category importance in the target classification model completes the recognition of the behavior information, or the recognized result is the preset result, and the target prediction result of the target classification model for the user to be classified is obtained;
[0135] Identify the behavior information based on the multiple feature classification models to obtain feature prediction results of each feature classification model for the user to be classified;
[0136] The target prediction result and the multiple feature prediction results are weighted and calculated based on a first preset weight and multiple second preset weights to obtain the group category, and the first preset weight is greater than the multiple second preset weights.
[0137] By setting the weight of the target classification model corresponding to the group characteristics to be greater than the weight of the feature classification model, the target classification model can make a greater contribution to the group classification, thereby improving the accuracy of the group classification. At the same time, since the classification model is predicted based on multiple group classification models, the voting results of multiple group classification models can be combined to further improve the accuracy of the group classification.
[0138] 107 , merging the multiple users to be classified based on the group category to obtain a customer group.
[0139] In at least one embodiment of the present invention, the same customer group refers to users to be classified in the same category.
[0140] The customer group may include a user group that has the potential to purchase a certain product, and the customer group may also include another user group that does not have the potential to purchase a certain product.
[0141] It can be seen from the above technical solutions that the present invention clusters the users to be classified based on the attribute information, and can quickly generate the multiple initial groups. Then, combined with the group characteristics corresponding to each user to be classified and the behavior information of the user to be classified on the preset platform, the group category can be accurately detected, thereby improving the accuracy of the division of the customer group.
[0142] like Figure 2 , which is a functional module diagram of a preferred embodiment of the customer group segmentation device of the present invention. The customer group segmentation device 11 includes an acquisition unit 110, a clustering unit 111, an identification unit 112, a generation unit 113, a collection unit 114 and a merging unit 115. The module / unit referred to in the present invention refers to a series of computer-readable instruction segments that can be acquired by the processor 13 and can perform fixed functions, which are stored in the memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0143] The acquisition unit 110 receives a customer group classification request, and acquires attribute information of a plurality of users to be classified according to the customer group classification request.
[0144] In at least one embodiment of the present invention, the customer group classification request may be triggered and generated by a business person of a preset platform. The customer group classification request carries user identification codes of the multiple users to be classified.
[0145] In at least one embodiment of the present invention, the attribute information refers to the information corresponding to the multiple users to be classified on preset attribute factors, wherein the preset attribute factors may include, but are not limited to: age, zodiac sign, date of birth, height, gender and device model, etc.
[0146] In at least one embodiment of the present invention, the acquisition unit 110 may acquire the attribute information of the plurality of users to be classified from a user information database.
[0147] The clustering unit 111 performs clustering processing on the multiple users to be classified according to the attribute information to obtain multiple initial groups.
[0148] In at least one embodiment of the present invention, the multiple initial groups refer to groups obtained after initially dividing the multiple users to be classified according to the similarity of the attribute information.
[0149] In at least one embodiment of the present invention, the clustering unit 111 performs clustering processing on the multiple users to be classified according to the attribute information, and obtains multiple initial groups including:
[0150] Encoding the attribute information to obtain a representation vector for each user to be classified;
[0151] extracting a plurality of first users from the plurality of users to be classified according to the number of users of the plurality of users to be classified, and determining the remaining users of the plurality of users to be classified except the plurality of first users as a plurality of second users;
[0152] For any second user, calculating the user similarity between the any second user and each first user according to the characterization vector of the any second user and the characterization vector of each first user;
[0153] Classify the arbitrary second user into the category of the first user with the highest user similarity, to obtain multiple basic groups;
[0154] The multiple basic groups are merged based on the characterization vectors of group users in each basic group to obtain the multiple initial groups.
[0155] The number of users may be the total number of the plurality of users to be classified at a preset ratio, for example, if the total number of users is 100 and the preset ratio is 0.1, the number of users is 10. The number of users is usually an integer greater than 1.
[0156] The number of the plurality of first users is the number of users.
[0157] The user similarity may be calculated based on the cosine value of the characterization vector of any second user and the characterization vector of each first user.
[0158] By encoding the user to be classified through the attribute information, the comprehensiveness of the characterization of the user to be classified by the characterization vector can be improved. The user similarity between any second user and each first user can be accurately quantified through the characterization vector, so that based on the user similarity, any second user can be reasonably classified into the corresponding first user, thereby improving the accuracy of the multiple basic groups, thereby improving the clustering accuracy of the multiple initial groups.
[0159] Specifically, the clustering unit 111 merges the multiple basic groups based on the characterization vectors of group users in each basic group, and obtains the multiple initial groups including:
[0160] Calculate the average value of the representation vectors of group users in each basic group as the group feature of each basic group;
[0161] Calculating the group similarity between any two groups in the multiple basic groups based on the group characteristics;
[0162] The multiple basic groups are merged based on the group similarities to obtain the multiple initial groups.
[0163] The similarity between any two groups among the multiple initial groups is less than a preset similarity threshold, and the preset similarity threshold can be set according to actual needs.
[0164] By merging the multiple basic groups according to the group similarity, it is possible to avoid the situation where there are groups with too high similarity among the multiple initial groups, which may cause the subsequent inability to accurately extract group features.
[0165] The identification unit 112 identifies the group characteristics corresponding to each user to be classified based on the multiple initial groups.
[0166] In at least one embodiment of the present invention, the group characteristics refer to key attribute factors corresponding to the initial group to which each user to be classified belongs.
[0167] In at least one embodiment of the present invention, the identification unit 112 identifies the group characteristics corresponding to each user to be classified based on the multiple initial groups, including:
[0168] Calculate the average value of the representation vectors of group users in each initial group as the encoding information of each initial group;
[0169] Extracting target features from the encoded information based on a pre-trained convolutional model;
[0170] The attribute factor corresponding to the target feature is determined as the group feature.
[0171] Through the above implementation, the influence of irrelevant features on customer group segmentation can be eliminated.
[0172] Specifically, the recognition unit 112 extracts target features from the encoded information based on the pre-trained convolutional model, including:
[0173] Obtaining a model threshold in the convolutional model;
[0174] A coding value having a value greater than the model threshold is obtained from the coding information as the target feature.
[0175] The generating unit 113 generates a group classification model of each group feature according to the user behaviors of multiple annotated users on the preset platform and the annotated category of each annotated user.
[0176] In at least one embodiment of the present invention, the preset platform may be any operating platform, and the preset platform may include multiple application programs.
[0177] The user behavior may include login operations of the multiple marked users on a game application in the preset platform, etc.
[0178] The marked category may include potential for purchase of a certain product, etc.
[0179] The group classification model includes a first decision layer and a plurality of second decision layers, wherein the category importance of the first decision layer is greater than the category importance of the plurality of second decision layers.
[0180] In at least one embodiment of the present invention, the generating unit 113 generates a group classification model for each group feature according to the user behaviors of multiple annotated users on a preset platform and the annotated category of each annotated user, including:
[0181] For each group feature, collecting the buried point information of the marked users corresponding to the group feature from the preset platform as the user behavior;
[0182] Classifying the user behaviors into events to obtain the event category corresponding to each marked user;
[0183] For each event category, determine the annotated user of the event category as the event user;
[0184] Based on the marked category, a first number of the event users that are positive samples is counted, and based on the marked category, a second number of the event users that are negative samples is counted;
[0185] Based on the annotation category, counting a third number of the plurality of annotated users belonging to positive samples, and counting a fourth number of the plurality of annotated users belonging to negative samples;
[0186] Calculate the category importance of each event category based on the first number, the second number, the third number, and the fourth number;
[0187] Construct a decision layer for each event category based on the user behavior corresponding to each event category and the annotated category corresponding to the annotated user of the user behavior;
[0188] The group classification model is obtained by splicing the multiple decision layers in descending order of the importance of the categories.
[0189] The event categories include login events, purchase events, browsing events, etc.
[0190] The positive samples refer to users whose labeled category is that they have the potential to purchase a certain product, and the negative samples refer to users whose labeled category is that they do not have the potential to purchase a certain product.
[0191] By combining the first number, the second number, the third number and the fourth number, the category importance of each event category can be accurately quantified, and further by splicing multiple decision layers according to the category importance, the accuracy of the group classification model can be improved.
[0192] Specifically, the generating unit 113 calculates the category importance of each event category based on the first number, the second number, the third number, and the fourth number, including:
[0193] The interference degree of each event category is calculated based on the first number and the second number, and the calculation formula of the interference degree is: Wherein, R represents the interference degree, n1 represents the first number, and n2 represents the second number;
[0194] The target importance is calculated based on the third number and the fourth number. If the third number is less than the fourth number, the calculation formula of the target importance is:
[0195] Wherein, K represents the target importance, n3 represents the third number, and n4 represents the fourth number;
[0196] The difference between the target importance and each interference degree is calculated to obtain the category importance of each event category.
[0197] The category importance of each event category can be accurately quantified through the relationship between the target importance and each interference degree.
[0198] Specifically, the generating unit 113 constructs a behavior category mapping relationship according to the user behavior corresponding to each event category and the annotated category corresponding to the annotated user of the user behavior, and then generates a decision layer for each event category according to the category mapping relationship.
[0199] The collecting unit 114 collects the behavior information of each user to be classified on the preset platform.
[0200] In at least one embodiment of the present invention, the behavior information can be obtained based on information embedding for each user to be classified on the preset platform.
[0201] The identification unit 112 identifies the behavior information based on the group classification model to obtain the group category corresponding to each user to be classified.
[0202] It should be emphasized that in order to further ensure the privacy and security of the above group categories, the above group categories can also be stored in a node of a blockchain.
[0203] In at least one embodiment of the present invention, the group category refers to the category corresponding to the to-be-classified users, and the group category includes whether there is a purchasing potential for a certain product.
[0204] In at least one embodiment of the present invention, the identification unit 112 identifies the behavior information based on the group classification model, and obtains the group category corresponding to each user to be classified, including:
[0205] For each user to be classified, a group classification model corresponding to the group characteristics of the user to be classified is selected as a target classification model, and group classification models other than the characteristic classification model are determined as a plurality of characteristic classification models;
[0206] Identify the behavior information based on the first decision layer in the target classification model to obtain an identification result;
[0207] If the recognition result is not the preset result, a second decision layer is selected from the target classification model according to the category importance to identify the behavior information, until the termination decision layer with the smallest category importance in the target classification model completes the recognition of the behavior information, or the recognized result is the preset result, and the target prediction result of the target classification model for the user to be classified is obtained;
[0208] Identify the behavior information based on the multiple feature classification models to obtain feature prediction results of each feature classification model for the user to be classified;
[0209] The target prediction result and the multiple feature prediction results are weighted and calculated based on a first preset weight and multiple second preset weights to obtain the group category, and the first preset weight is greater than the multiple second preset weights.
[0210] By setting the weight of the target classification model corresponding to the group characteristics to be greater than the weight of the feature classification model, the target classification model can make a greater contribution to the group classification, thereby improving the accuracy of the group classification. At the same time, since the classification model is predicted based on multiple group classification models, the voting results of multiple group classification models can be combined to further improve the accuracy of the group classification.
[0211] The merging unit 115 merges the multiple users to be classified based on the group category to obtain a customer group.
[0212] In at least one embodiment of the present invention, the same customer group refers to users to be classified in the same category.
[0213] The customer group may include a user group that has the potential to purchase a certain product, and the customer group may also include another user group that does not have the potential to purchase a certain product.
[0214] It can be seen from the above technical solutions that the present invention clusters the users to be classified based on the attribute information, and can quickly generate the multiple initial groups. Then, combined with the group characteristics corresponding to each user to be classified and the behavior information of the user to be classified on the preset platform, the group category can be accurately detected, thereby improving the accuracy of the division of the customer group.
[0215] like Figure 3 FIG. 1 is a schematic diagram of the structure of an electronic device of a preferred embodiment of the method for implementing customer group classification of the present invention.
[0216] In one embodiment of the present invention, the electronic device 1 includes, but is not limited to, a memory 12, a processor 13, and computer-readable instructions stored in the memory 12 and executable on the processor 13, such as a customer group segmentation program.
[0217] Those skilled in the art will understand that the schematic diagram is merely an example of the electronic device 1 and does not constitute a limitation on the electronic device 1. The electronic device 1 may include more or fewer components than shown in the diagram, or a combination of certain components, or different components. For example, the electronic device 1 may also include input and output devices, network access devices, buses, etc.
[0218] The processor 13 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor 13 is the computing core and control center of the electronic device 1, and uses various interfaces and lines to connect various parts of the entire electronic device 1, and executes the operating system of the electronic device 1 and various installed applications, program codes, etc.
[0219] Exemplarily, the computer-readable instructions may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of completing specific functions, which are used to describe the execution process of the computer-readable instructions in the electronic device 1. For example, the computer-readable instructions may be divided into an acquisition unit 110, a clustering unit 111, an identification unit 112, a generation unit 113, a collection unit 114, and a merging unit 115.
[0220] The memory 12 can be used to store the computer-readable instructions and / or modules, and the processor 13 implements various functions of the electronic device 1 by running or executing the computer-readable instructions and / or modules stored in the memory 12, and calling the data stored in the memory 12. The memory 12 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device, etc. The memory 12 can include non-volatile and volatile memories, such as: a hard disk, a memory, a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other storage devices.
[0221] The memory 12 may be an external memory and / or an internal memory of the electronic device 1. Furthermore, the memory 12 may be a memory in a physical form, such as a memory stick, a TF card (Trans-flash Card), and the like.
[0222] If the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also instruct the relevant hardware to complete it through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed by the processor, the steps of the above-mentioned various method embodiments can be implemented.
[0223] The computer-readable instructions include computer-readable instruction codes, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer-readable instruction code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM).
[0224] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0225] Combination Figure 1 The memory 12 in the electronic device 1 stores computer-readable instructions to implement a customer group segmentation method, and the processor 13 can execute the computer-readable instructions to implement:
[0226] Receiving a customer group classification request, and acquiring attribute information of a plurality of users to be classified according to the customer group classification request;
[0227] Clustering the multiple users to be classified according to the attribute information to obtain multiple initial groups;
[0228] Identifying group characteristics corresponding to each user to be classified based on the multiple initial groups;
[0229] Generate a group classification model for each group feature according to the user behaviors of multiple annotated users on a preset platform and the annotated category of each annotated user;
[0230] Collecting behavior information of each user to be classified on the preset platform;
[0231] Identify the behavior information based on the group classification model to obtain the group category corresponding to each user to be classified;
[0232] The multiple users to be classified are merged based on the group category to obtain a customer group.
[0233] Specifically, the specific implementation method of the processor 13 for the above-mentioned computer readable instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0234] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0235] The computer-readable storage medium stores computer-readable instructions, wherein the computer-readable instructions are used to implement the following steps when executed by the processor 13:
[0236] Receiving a customer group classification request, and acquiring attribute information of a plurality of users to be classified according to the customer group classification request;
[0237] Clustering the multiple users to be classified according to the attribute information to obtain multiple initial groups;
[0238] Identifying group characteristics corresponding to each user to be classified based on the multiple initial groups;
[0239] Generate a group classification model for each group feature according to the user behaviors of multiple annotated users on a preset platform and the annotated category of each annotated user;
[0240] Collecting behavior information of each user to be classified on the preset platform;
[0241] Identify the behavior information based on the group classification model to obtain the group category corresponding to each user to be classified;
[0242] The multiple users to be classified are merged based on the group category to obtain a customer group.
[0243] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0244] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0245] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0246] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices mentioned can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0247] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for dividing customer groups, characterized in that: The customer group segmentation method includes: Receiving a customer group classification request, and acquiring attribute information of a plurality of users to be classified according to the customer group classification request; Clustering the multiple users to be classified according to the attribute information to obtain multiple initial groups; Identifying group characteristics corresponding to each user to be classified based on the multiple initial groups; Generate a group classification model for each group feature according to the user behaviors of multiple labeled users on a preset platform and the labeled categories of each labeled user, including: for each group feature, collect the buried point information of the labeled users corresponding to the group feature from the preset platform as the user behavior; classify the user behavior into events to obtain the event category corresponding to each labeled user; for each event category, determine the labeled users of the event category as event users; count the first number of positive samples of the event users based on the labeled category, and count the second number of negative samples of the event users based on the labeled category; count the third number of positive samples of the multiple labeled users based on the labeled category, and count the fourth number of negative samples of the multiple labeled users; calculate the category importance of each event category based on the first number, the second number, the third number and the fourth number; construct a decision layer for each event category according to the user behavior corresponding to each event category and the labeled category corresponding to the labeled user of the user behavior; splice the multiple decision layers in descending order of the category importance to obtain the group classification model; Collecting behavior information of each user to be classified on the preset platform; Based on the group classification model, the behavior information is identified to obtain the group category corresponding to each user to be classified, including: for each user to be classified, the group classification model corresponding to the group characteristics of the user to be classified is selected as the target classification model, and the group classification models other than the target classification model are determined as multiple feature classification models; based on the first decision layer in the target classification model, the behavior information is identified to obtain a recognition result; if the recognition result is not a preset result, a second decision layer is selected from the target classification model according to the category importance to identify the behavior information, until the termination decision layer with the smallest category importance in the target classification model completes the recognition of the behavior information, or the recognized result is the preset result, and the target prediction result of the target classification model for the user to be classified is obtained; based on the multiple feature classification models, the behavior information is identified to obtain the feature prediction result of each feature classification model for the user to be classified; based on the first preset weight and multiple second preset weights, the target prediction result and the multiple feature prediction results are weighted and calculated to obtain the group category, and the first preset weight is greater than the multiple second preset weights; The multiple users to be classified are merged based on the group category to obtain a customer group.
2. The method for dividing customer groups according to claim 1, characterized in that: The clustering of the multiple users to be classified according to the attribute information to obtain multiple initial groups includes: Encoding the attribute information to obtain a representation vector for each user to be classified; extracting a plurality of first users from the plurality of users to be classified according to the number of users of the plurality of users to be classified, and determining the remaining users of the plurality of users to be classified except the plurality of first users as a plurality of second users; For any second user, calculating the user similarity between the any second user and each first user according to the characterization vector of the any second user and the characterization vector of each first user; Classify the arbitrary second user into the category of the first user with the highest user similarity, to obtain multiple basic groups; The multiple basic groups are merged based on the characterization vectors of group users in each basic group to obtain the multiple initial groups.
3. The method for dividing customer groups according to claim 2, characterized in that: The step of merging the multiple basic groups based on the characterization vectors of group users in each basic group to obtain the multiple initial groups comprises: Calculate the average value of the representation vectors of group users in each basic group as the group feature of each basic group; Calculating the group similarity between any two groups in the multiple basic groups based on the group characteristics; The multiple basic groups are merged based on the group similarities to obtain the multiple initial groups.
4. The method for dividing customer groups according to claim 2, characterized in that: The step of identifying the group characteristics corresponding to each user to be classified based on the multiple initial groups includes: Calculate the average value of the representation vectors of group users in each initial group as the encoding information of each initial group; Extracting target features from the encoded information based on a pre-trained convolutional model; The attribute factor corresponding to the target feature is determined as the group feature.
5. The method for dividing customer groups according to claim 1, characterized in that: The calculating the category importance of each event category based on the first number, the second number, the third number, and the fourth number includes: The interference degree of each event category is calculated based on the first number and the second number, and the calculation formula of the interference degree is: ,in, represents the interference degree, represents the first quantity, represents said second quantity; The target importance is calculated based on the third number and the fourth number. If the third number is less than the fourth number, the calculation formula of the target importance is: ,in, Indicates the importance of the target, represents the third quantity, represents the fourth quantity; The difference between the target importance and each interference degree is calculated to obtain the category importance of each event category.
6. A customer group classification device, characterized in that: The customer group classification device comprises: An acquisition unit, configured to receive a customer group classification request, and acquire attribute information of a plurality of users to be classified according to the customer group classification request; A clustering unit, configured to perform clustering processing on the plurality of users to be classified according to the attribute information to obtain a plurality of initial groups; an identification unit, configured to identify a group feature corresponding to each user to be classified based on the multiple initial groups; A generating unit, for generating a group classification model for each group feature according to the user behaviors of multiple annotated users on a preset platform and the annotated categories of each annotated user, including: for each group feature, collecting the buried point information of the annotated users corresponding to the group feature from the preset platform as the user behavior; classifying the user behavior into events to obtain the event category corresponding to each annotated user; for each event category, determining the annotated users of the event category as event users; counting the first number of the event users belonging to positive samples based on the annotated category, and counting the second number of the event users belonging to negative samples based on the annotated category; counting the third number of the multiple annotated users belonging to positive samples based on the annotated category, and counting the fourth number of the multiple annotated users belonging to negative samples; calculating the category importance of each event category based on the first number, the second number, the third number and the fourth number; constructing a decision layer for each event category according to the user behavior corresponding to each event category and the annotated category corresponding to the annotated user of the user behavior; splicing the multiple decision layers in descending order of the category importance to obtain the group classification model; A collection unit, used to collect behavior information of each user to be classified on the preset platform; The identification unit is also used to identify the behavior information based on the group classification model to obtain the group category corresponding to each user to be classified, including: for each user to be classified, selecting the group classification model corresponding to the group characteristics of the user to be classified as the target classification model, and determining the group classification models other than the target classification model as multiple feature classification models; identifying the behavior information based on the first decision layer in the target classification model to obtain a recognition result; if the recognition result is not a preset result, selecting a second decision layer from the target classification model according to the category importance to identify the behavior information until the termination decision layer with the smallest category importance in the target classification model completes the recognition of the behavior information, or the recognized result is the preset result, and the target prediction result of the target classification model for the user to be classified is obtained; identifying the behavior information based on the multiple feature classification models to obtain the feature prediction result of each feature classification model for the user to be classified; performing a weighted sum operation on the target prediction result and the multiple feature prediction results based on a first preset weight and multiple second preset weights to obtain the group category, and the first preset weight is greater than the multiple second preset weights; A merging unit is used to merge the multiple users to be classified based on the group category to obtain a customer group.
7. An electronic device, characterized in that: The electronic device comprises: a memory storing computer-readable instructions; and A processor executes computer-readable instructions stored in the memory to implement the customer group segmentation method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions are executed by a processor in an electronic device to implement the customer group segmentation method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Microblog specific event attention group identification method
CN111026976A
Community division method and device based on artificial intelligence, equipment and storage medium
CN113570391A