Target object grouping and layering processing method and device and electronic equipment
Through the construction of pre-trained hierarchical model and loss function, combined with the gradient descent algorithm, the hierarchical boundary value of the target object is dynamically determined, which solves the problem of poor hierarchical results caused by manual division and improves the effectiveness of hierarchical results.
Patent Information
- Application Number
- CN202510134389.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-07
AI Technical Summary
In the prior art, the hierarchical boundary values of the target object need to be divided manually, resulting in poor use of the hierarchical results.
By obtaining the user characteristics of the user to be identified, input them into the pre-trained hierarchical model, the customer group and customer hierarchy to which the user belongs. Then, a functional relationship is constructed based on the statistical indicators of customer hierarchy, and a loss function is constructed based on the expected number of users, the number of users meeting the standard and the overall risk, and a gradient descent algorithm is used to determine the hierarchical lower bound parameters.
The dynamic determination of the hierarchical boundary value is achieved, which improves the use of hierarchical results, making the hierarchical results more reasonable and more suitable for practical application scenarios.
Smart Images

Figure CN120125259A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, and electronic device for clustering and hierarchical processing of target objects. Background Art
[0002] In the technical field of data processing, the division of the group to which a target object belongs and the division of the hierarchy to which it belongs are the keys to the refined management of the target object. Among them, the target object can be a user or a commodity. Taking the user as the target object as an example, how to accurately determine the customer group to which the user belongs and the priority of the user in the customer group, and then formulate corresponding marketing strategies and advertising strategies for the user according to the determined customer group and priority, becomes the key to providing targeted services to the user and is also an important factor in saving the marketing cost and publicity cost of the service provider.
[0003] However, for how many levels to divide and how to set the values of the hierarchical boundary between each level, the existing technical solutions all adopt manual hierarchical division and manually set the hierarchical boundary values. Summary of the Invention
[0004] In view of this, the embodiments of this application provide a method, apparatus, and electronic device for clustering and hierarchical processing of target objects to dynamically determine the hierarchical boundary value and further improve the use effect of the hierarchical result.
[0005] In a first aspect, the embodiments of this application provide a method for clustering and hierarchical processing of target objects, where the method includes:
[0006] Obtain the user characteristics of the user to be identified, input the user characteristics into a pre-trained hierarchical model, and based on the input user characteristics, the hierarchical model determines the customer group and customer hierarchy to which the user belongs;
[0007] According to the customer hierarchical result, determine the statistical indicators corresponding to each customer hierarchy according to a preset statistical algorithm, where the types of the statistical indicators are: the user quantity of the customer hierarchy, the number of qualified users in the customer hierarchy, and the overall risk of the customer hierarchy;
[0008] Among them, the customer hierarchy is obtained by dividing with different lower hierarchical boundary parameters, and the lower hierarchical boundary parameters are determined in advance according to the following steps:
[0009] Based on the statistical indicators corresponding to each of the customer stratifications, construct a first functional relationship, a second functional relationship, and a third functional relationship respectively, where the first functional relationship is the functional relationship between the user quantity level and the lower bound parameter of the stratification, the second functional relationship is the functional relationship between the number of users meeting the standard and the lower bound parameter of the stratification, and the third functional relationship is the functional relationship between the overall risk and the lower bound parameter of the stratification;
[0010] According to the expected user quantity of each of the customer stratifications, the expected number of users meeting the standard of each of the customer stratifications, and the expected overall risk of each of the customer stratifications, construct a user quantity level loss function, a loss function for users meeting the standard, and an overall risk loss function respectively; among them, the user quantity level loss function is the loss function between the first functional relationship and the expected user quantity, the loss function for users meeting the standard is the loss function between the second functional relationship and the expected number of users meeting the standard, and the overall risk loss function is the loss function between the third functional relationship and the expected overall risk;
[0011] Based on the user quantity level loss function, the loss function for users meeting the standard, and the overall risk loss function, construct a comprehensive loss function, and use a preset gradient descent algorithm to determine that the lower bound parameter of each customer stratification when the comprehensive loss function converges is the target lower bound parameter.
[0012] In some possible embodiments, the constructing the first functional relationship, the second functional relationship, and the third functional relationship respectively based on the statistical indicators corresponding to each of the customer stratifications includes:
[0013] Based on the statistical indicators corresponding to each of the customer stratifications, construct a target parameter matrix, where each column vector in the target parameter matrix is: the lower bound value of each layer corresponding to each customer stratification of the same customer group; each row vector in the target parameter matrix is: the lower bound value of each of the customer groups in the same customer stratification;
[0014] According to the target parameter matrix, using each of the row vectors as the independent variable and the dimension corresponding to each of the statistical indicators as the dependent variable, construct the first functional relationship, the second functional relationship, and the third functional relationship respectively.
[0015] In some possible embodiments, the constructing the user quantity level loss function, the loss function for users meeting the standard, and the overall risk loss function respectively according to the expected user quantity of each of the customer stratifications, the expected number of users meeting the standard of each of the customer stratifications, and the expected overall risk of each of the customer stratifications includes:
[0016] Obtain the expected user quantity, the expected number of users meeting the standard, and the expected overall risk input by the user;
[0017] The mean square error loss function loss between the expected number of users and the first functional relationship y is determined as the user magnitude loss function;
[0018] The mean square error loss function loss between the expected number of qualified users and the second functional relationship τ is determined as the qualified user loss function;
[0019] The mean square error loss function loss between the expected overall risk and the third functional relationship γ is determined as the overall risk loss function.
[0020] In some possible embodiments, constructing the comprehensive loss function based on the user magnitude loss function, the qualified user loss function, and the overall risk loss function includes:
[0021] The comprehensive loss function Loss is determined based on the following formula:
[0022] loss = loss y + loss τ + loss γ
[0023] Using the preset gradient descent algorithm to determine the lower bound parameters of each customer layer when the comprehensive loss function converges as the target lower bound parameters of the layer includes:
[0024] Taking the derivative of the comprehensive loss function in the three direction angles of y, τ, and γ to respectively determine the gradient of the comprehensive loss function in the y direction The gradient of the comprehensive loss function in the τ direction and the gradient of the comprehensive loss function in the γ direction
[0025] According to the preset iteration step α, iterate the following parameter iteration formula:
[0026]
[0027] Until the comprehensive loss function converges or until the number of iterations reaches the preset number of iterations;
[0028] Taking the lower bound parameter x of the layer when the comprehensive loss function converges or reaches the preset number of iterations l as the target lower bound parameter of the layer.
[0029] In some possible embodiments, in the process of constructing the first functional relationship, the second functional relationship, and the third functional relationship respectively based on the statistical indicators corresponding to each customer layer, the method further includes:
[0030] Execute a preset data aggregation algorithm from the database to obtain the first sample data;
[0031] Based on the first sample data, execute a preset prefix sum processing algorithm to calculate the statistical indicators corresponding to each historical customer stratification;
[0032] Based on the statistical indicators corresponding to each historical customer stratification, construct the first functional relationship, the second functional relationship, and the third functional relationship respectively.
[0033] In some possible embodiments, the method further includes:
[0034] Based on the expected number of users, the expected number of qualified users, and the expected overall risk of each customer stratification, and according to the progressive relationship of the customer stratifications, perform the following steps for each customer stratification layer by layer:
[0035] Determine the current stratification parameter corresponding to the current customer stratification, and based on the comprehensive loss function, calculate the gradients in the three direction angles of y, τ, and γ corresponding to the current customer stratification;
[0036] Based on the gradients in the three direction angles of y, τ, and γ corresponding to the current customer stratification, perform cyclic iteration using the parameter iteration formula;
[0037] Determine the customer stratification parameter corresponding to when the preset number of iterations is reached as the target stratification lower bound parameter.
[0038] In some possible embodiments, the user characteristics include: identity characteristics, behavior characteristics, consumption characteristics, and interest and hobby characteristics.
[0039] In a second aspect, an embodiment of the present application provides a clustering and stratification processing device for a target object, where the device includes:
[0040] A stratification processing module; obtain the user characteristics of the user to be identified, input the user characteristics into a pre-trained stratification model, and based on the input user characteristics, the stratification model determines the customer group and customer stratification to which the user belongs, where the customer stratification is obtained by dividing with different stratification lower bound parameters;
[0041] A statistics module, according to the customer stratification result, determine the statistical indicators corresponding to each customer stratification according to a preset statistical algorithm, where the types of the statistical indicators are: the user quantity of the customer stratification, the number of qualified users of the customer stratification, and the overall risk of the customer stratification;
[0042] A stratification lower bound parameter calculation module, used to determine the stratification lower bound parameter in advance according to the following steps:
[0043] Based on the statistical indicators corresponding to each customer stratification, construct a first functional relationship, a second functional relationship, and a third functional relationship respectively, where the first functional relationship is the functional relationship between the user quantity level and the lower bound parameter of the stratification, the second functional relationship is the functional relationship between the number of qualified users and the lower bound parameter of the stratification, and the third functional relationship is the functional relationship between the overall risk and the lower bound parameter of the stratification;
[0044] According to the expected user quantity of each customer stratification, the expected number of qualified users of each customer stratification, and the expected overall risk of each customer stratification, construct a user quantity loss function, a qualified user loss function, and an overall risk loss function respectively; where the user quantity loss function is the loss function between the first functional relationship and the expected user quantity, the qualified user loss function is the loss function between the second functional relationship and the expected number of qualified users, and the overall risk loss function is the loss function between the third functional relationship and the expected overall risk;
[0045] Based on the user quantity loss function, the qualified user loss function, and the overall risk loss function, construct a comprehensive loss function, and use a preset gradient descent algorithm to determine that the lower bound parameter of each customer stratification when the comprehensive loss function converges is the target lower bound parameter.
[0046] Combined with the second aspect, in some possible embodiments, the lower bound parameter calculation module is specifically used for:
[0047] Based on the statistical indicators corresponding to each customer stratification, construct a target parameter matrix, where each column vector in the target parameter matrix is: the lower bound value of each stratification corresponding to each customer stratification in the same customer group; each row vector in the target parameter matrix is: the lower bound value of each customer group in the same customer stratification;
[0048] According to the target parameter matrix, using each row vector as the independent variable and the dimension corresponding to each statistical indicator as the dependent variable, construct the first functional relationship, the second functional relationship, and the third functional relationship respectively.
[0049] Combined with the second aspect, in some possible embodiments, the user quantity loss function, the qualified user loss function, and the overall risk loss function are constructed respectively according to the expected user quantity of each customer stratification, the expected number of qualified users of each customer stratification, and the expected overall risk of each customer stratification;
[0050] Obtain the expected user quantity, the expected number of qualified users, and the expected overall risk input by the user;
[0051] The mean square error loss function loss between the expected number of users and the first functional relationship y is determined as the user magnitude loss function;
[0052] The mean square error loss function loss between the expected number of qualified users and the second functional relationship τ is determined as the qualified user loss function;
[0053] The mean square error loss function loss between the expected overall risk and the third functional relationship γ is determined as the overall risk loss function.
[0054] Combined with the second aspect, in some possible embodiments, the hierarchical lower bound parameter calculation module is specifically configured to:
[0055] The comprehensive loss function Loss is determined based on the following formula:
[0056] Loss = loss y + loss τ + loss γ
[0057] Derive the comprehensive loss function with respect to the three direction angles of y, τ, and γ, and respectively determine the gradient of the comprehensive loss function in the y direction the gradient of the comprehensive loss function in the τ direction and the gradient of the comprehensive loss function in the γ direction
[0058] According to the preset iteration step α, iterate the following parameter iteration formula:
[0059]
[0060] until the comprehensive loss function converges or until the number of iterations reaches the preset number of iterations;
[0061] The hierarchical lower bound parameter x when the comprehensive loss function converges or reaches the preset number of iterations l is the target hierarchical lower bound parameter.
[0062] Combined with the second aspect, in some possible embodiments, the apparatus further includes: a data preprocessing module, configured to:
[0063] Execute a preset data aggregation algorithm from the database to obtain first sample data;
[0064] Based on the first sample data, execute a preset prefix sum processing algorithm to calculate the statistical indicators corresponding to each historical customer stratification;
[0065] Construct the first functional relationship, the second functional relationship, and the third functional relationship respectively based on the statistical indicators corresponding to each of the historical customer stratifications.
[0066] In combination with the second aspect, in some possible embodiments, the stratification lower bound parameter calculation module is further configured to:
[0067] Based on the expected number of users, the expected number of qualified users, and the expected overall risk of each customer stratification, and according to the progressive relationship of the customer stratifications, perform the following steps layer by layer on each customer stratification:
[0068] Determine the current stratification parameter corresponding to the current customer stratification, and calculate the gradients in the three direction angles of y, τ, and γ corresponding to the current customer stratification based on the comprehensive loss function;
[0069] Perform cyclic iteration using the parameter iteration formula based on the gradients in the three direction angles of y, τ, and γ corresponding to the current customer stratification;
[0070] Determine the customer stratification parameter corresponding to when the preset number of iterations is reached as the target stratification lower bound parameter.
[0071] In combination with the second aspect, in some possible embodiments, the user characteristics include: identity characteristics, behavior characteristics, consumption characteristics, and hobby characteristics.
[0072] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes: a processor; and a memory storing a program; where the program includes instructions that, when executed by the processor, cause the processor to execute the method for clustering and stratifying the target object described in the first aspect.
[0073] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to cause a computer to execute the method for clustering and stratifying the target object described in the first aspect.
[0074] Advantages of the present application:
[0075] The present application provides a method, an apparatus, and an electronic device for clustering and stratifying a target object. Among them, the method obtains user characteristics of a user to be identified, inputs the user characteristics into a pre-trained stratification model, and the stratification model determines the customer group and customer stratification to which the user belongs based on the input user characteristics. Then, according to the customer stratification result, according to a preset statistical algorithm, the user quantity, the number of qualified users, and the overall risk corresponding to each customer stratification are calculated. Among them, in the embodiments of the present application, customer stratification is obtained by dividing with different lower stratification boundary parameters, and the lower stratification boundary parameters are pre-constructed with a function relationship between the user quantity, the qualified users, and the overall risk corresponding to each customer stratification and the lower stratification boundary parameters. Then, based on the constructed function relationship and the set expected number of users, expected number of qualified users, and expected overall risk, a user quantity loss function, a qualified user loss function, and an overall risk loss function are respectively constructed. Finally, a comprehensive loss function is constructed based on each loss function, and a preset gradient descent algorithm is used to determine that when the comprehensive loss function converges, the lower stratification boundary parameters of each customer stratification are the target lower stratification boundary parameters corresponding to each customer stratification.
[0076] Selecting the embodiments of the present application, abandoning the solution of the prior art that manually sets the lower stratification boundary parameters to divide different customer stratifications, continuously and dynamically obtaining the expected number of users, the expected number of qualified users, and the expected overall risk to be achieved corresponding to each customer stratification, and using a preset gradient descent algorithm to solve the corresponding lower stratification boundary parameters when the expected effect is achieved. The lower stratification boundary parameters set in this way are not fixed, but continuously adapt to the statistical indicators required by the application scenario, and the obtained lower stratification boundary parameters are more reasonable. When the stratification results of different customer stratifications under different customer groups are actually used, they can better help enterprises and institutions formulate more reasonable and cost-effective marketing strategies and advertising strategies, and the application effect is better. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In the following description of the exemplary embodiments in conjunction with the drawings, more details, features, and advantages of the present application are disclosed. In the drawings:
[0078] Figure 1 A flowchart showing a method for clustering and stratifying a target object provided by an embodiment of the present application is shown;
[0079] Figure 2 A flowchart showing a method for determining the lower stratification boundary parameters provided by an embodiment of the present application is shown;
[0080] Figure 3a Another flowchart showing a method for clustering and stratifying a target object provided by an embodiment of the present application is shown;
[0081] Figure 3bShows a schematic diagram of the prefix sum processing provided by an embodiment of the present application;
[0082] Figure 4 Shows another schematic flowchart of the method for clustering and stratifying the target object provided by an embodiment of the present application;
[0083] Figure 5 Shows another schematic flowchart of the method for clustering and stratifying the target object provided by an embodiment of the present application;
[0084] Figure 6 Shows another schematic flowchart of the method for clustering and stratifying the target object provided by an embodiment of the present application;
[0085] Figure 7 Shows a schematic logical structure diagram of the device for clustering and stratifying the target object provided by an embodiment of the present application;
[0086] Figure 8 Shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present application. Detailed implementation manners
[0087] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0088] It should be understood that the steps described in the method embodiments of the present application can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.
[0089] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or the interdependent relationship.
[0090] It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0091] In a first aspect, this application provides a method for clustering and stratifying target objects. This method is applied to any electronic device with the function of clustering and stratifying target objects, including but not limited to personal mobile terminals, computers, servers, etc. As Figure 1 shown, this method includes the following steps:
[0092] S11. Obtain the user characteristics of the user to be identified, input the user characteristics into a pre-trained stratification model, and based on the input user characteristics, the stratification model determines the customer group and customer stratification to which the user belongs;
[0093] S12. According to the customer stratification result, determine the statistical indicators corresponding to each customer stratification according to a preset statistical algorithm. Among them, the types of the statistical indicators are: the user quantity level of the customer stratification, the number of qualified users in the customer stratification, and the overall risk of the customer stratification.
[0094] Among them, the above customer stratification is obtained by dividing different stratification lower bound parameters. As Figure 2 shown, the stratification lower bound parameter is determined in advance according to the following steps:
[0095] S21. Based on the statistical indicators corresponding to each customer stratification, respectively construct a first function relationship, a second function relationship, and a third function relationship. Among them, the first function relationship is the function relationship between the user quantity level and the stratification lower bound parameter, the second function relationship is the function relationship between the number of qualified users and the stratification lower bound parameter, and the third function relationship is the function relationship between the overall risk and the stratification lower bound parameter;
[0096] S22. According to the expected user quantity of each customer stratification, the expected number of qualified users of each customer stratification, and the expected overall risk of each customer stratification, respectively construct a user quantity loss function, a qualified user loss function, and an overall risk loss function; among them, the user quantity loss function is the loss function between the first function relationship and the expected user quantity, the qualified user loss function is the loss function between the second function relationship and the expected number of qualified users, and the overall risk loss function is the loss function between the third function relationship and the expected overall risk;
[0097] S23. Construct a comprehensive loss function based on the user volume loss function, the compliant user loss function, and the overall risk loss function, and use a preset gradient descent algorithm to determine the lower bound parameters of each customer stratification when the comprehensive loss function converges as the target lower bound parameters of each customer stratification.
[0098] This method obtains the user characteristics of the user to be identified, inputs the user characteristics into a pre-trained stratification model, and the stratification model determines the customer group and customer stratification to which the user belongs based on the input user characteristics. Then, according to the customer stratification result, the user volume, the number of compliant users, and the overall risk corresponding to each customer stratification are calculated according to a preset statistical algorithm. Among them, in the embodiment of the present application, the customer stratification is obtained by dividing with different lower bound parameters of stratification, and the lower bound parameters of stratification are pre-constructed with the function relationships between the user volume, the compliant users, and the overall risk corresponding to each customer stratification and the lower bound parameters of stratification respectively. Then, based on the constructed function relationships and the set expected number of users, the expected number of compliant users, and the expected overall risk, a user volume loss function, a compliant user loss function, and an overall risk loss function are constructed respectively. Finally, a comprehensive loss function is constructed based on each loss function, and a preset gradient descent algorithm is used to determine the lower bound parameters of each customer stratification when the comprehensive loss function converges as the target lower bound parameters corresponding to each customer stratification.
[0099] By selecting the embodiment of the present application and abandoning the existing solution of manually setting the lower bound parameters of stratification to divide different customer stratifications, by continuously and dynamically obtaining the expected number of users, the expected number of compliant users to be achieved, and the overall risk to be controlled corresponding to each customer stratification, and using a preset gradient descent algorithm to solve the corresponding lower bound parameters of stratification when the expected effect is achieved. In this way, the set lower bound parameters of stratification are not fixed, but continuously adapt to the statistical indicators required by the application scenario, and the obtained lower bound parameters of stratification are more reasonable. When the stratification results of different customer stratifications under different customer groups are actually used, they can better help enterprises and institutions formulate more reasonable and cost-effective marketing strategies and advertising promotion strategies, and the application effect is better.
[0100] The following will elaborate on the above steps S11, S12, and S21 - S23 with specific examples:
[0101] Before elaborating on the above steps S11 to S12 and steps S21 to S23 in detail, the execution order between each step is explained here: In the embodiment of this application, steps S11 and S12 are to use a pre-trained stratification model to divide the input user to be recognized after steps S21 to S23 are executed to determine the corresponding lower bound parameters of each customer stratification, so as to determine the customer group to which the user belongs and which specific customer stratification in this customer group the user belongs to. If the application scenario of this application is analogized to an Excel table and each column corresponds to a customer group, then each row in this column corresponds to each customer stratification in this customer group. Executing steps S11 to S12 is to locate the attribution of the input user to be recognized and determine which column and which row the user specifically belongs to. The specific division criteria between specific rows need to be determined in advance by executing steps S21 to S23.
[0102] After understanding the execution order between each step, it can be seen that in the embodiment of this application, steps S21 to S23 are the core key to dynamically determining customer stratification. Therefore, steps S21 to S23 will be described first below, and then steps S11 to S12 will be described.
[0103] In the embodiment of this application, a customer group refers to a group of customers, specifically a group of customers with some common characteristics that the products and services of an enterprise or organization are targeted at. Customer stratification is a subset of customer groups that are further segmented according to different dimensions of a certain characteristic within the same customer group. Exemplarily, taking the lending business scenario as an example, the customer group can be divided into: high-net-worth customer group, medium-net-worth customer group, low-net-worth customer group. Then, for the customers within the high-net-worth customer group, they can be further divided into first-priority high-net-worth customers, second-priority high-net-worth customers, third-priority high-net-worth customers... according to the specific asset scale. In the embodiment of this application, the specific division methods of the customer group and customer stratification depend on the division characteristics selected in the actual application process. In other words, the customer group and customer stratification are flexibly set according to the actual application scenario requirements. For example, it can be stipulated that the number of customer groups to be divided and the number of customer stratifications within each customer group can be set by the enterprise or organization according to the operation requirements. If refined management and operation are required, a larger number of customer groups and a larger number of customer stratifications can be set. If the operation management cost needs to be saved, a smaller number of customer groups and a smaller number of customer stratifications can be set. The specific types of specific customer groups and customer stratifications will not be strictly specified in this article.
[0104] Both the customer group and customer stratification serve the actual application scenarios. In the embodiments of the present application, the application scenario of this method can be a lending scenario. By extracting features of credit users, and then combining the features of the users, it can be determined which credit risk customer group the user specifically belongs to, and which specific credit stratification within the credit risk customer group the credit user is in, which helps credit institutions or enterprises provide more accurate credit services, and is conducive to credit institutions or enterprises controlling credit risks and maximizing resource benefits, so that the credit business can find a balance between safety and growth space.
[0105] That is to say, the method provided in the present application can be applied to risk control customer group stratification. Customers are divided into multiple levels such as loan level 1 and loan level 2 based on the value and risk of the users. Among them, the higher the value of the users in the loan level 1 stratification, as the level progresses, the user value gradually decreases, and the risk increases in turn. Based on this, in the risk control process, different service resources can be allocated to customers in different customer stratifications. Exemplarily, high-value customers with a more forward position in the level can obtain higher-level service strategies and attentions, as well as more refined operation services. For credit institutions, this way can avoid investing too many resources in low-value or high-risk customers, and reduce the input costs of the institution or enterprise.
[0106] Based on this, when performing step S21, the customers can be pre-initialized and stratified. The customer group and customer stratification to which the customer belongs are pre-divided through a machine learning model, and then the initialized stratification result is obtained. In this initialized stratification result, each customer stratification also contains corresponding users. Then, in the same way as performing step S12, the statistical indicators under each initialized customer stratification are counted. In some possible embodiments, the first function relationship, the second function relationship, and the third function relationship can be constructed based on the statistical indicators corresponding to each customer stratification pre-stored in the database. As an implementation manner, when performing step S21, it can be realized through the following steps:
[0107] S21-1: Execute a preset data aggregation algorithm from the database to obtain the first sample data.
[0108] S21-2: Based on the first sample data, execute a preset prefix sum processing algorithm to calculate the statistical indicators corresponding to each historical customer stratification.
[0109] S21-3: Based on the statistical indicators corresponding to each historical customer stratification, respectively construct the first function relationship, the second function relationship, and the third function relationship.
[0110] In the embodiments of the present application, the type of the database may be a distributed database. The historical user stratification records, the stratification results, and the real-time generated user stratification results are all scattered and recorded in different data sources of the distributed database. When performing step S21-1, data is obtained from each data source, and then the obtained data is merged and refined to obtain a more complete and valuable sample data for determining the lower bound parameter of stratification. This sample data is the first sample data.
[0111] Among them, the lower bound parameter of stratification is the basis for dividing different customer stratifications, specifically the boundary value between two adjacent customer stratifications. Exemplarily, it can be as Figure 3a shown. If it is necessary to divide user data into three customer groups: customer group A, customer group B, and customer group C, and then each customer group is further subdivided into n customer stratifications x 1A ~x nA 、x 1B ~x nB 、x 1C ~x nC . The boundary value between each customer stratification is the lower bound parameter of stratification. Among them, the first lower bound parameter of stratification is the demarcation value between the first customer stratification and the second customer stratification.
[0112] In the embodiments of the present application, the preset data aggregation algorithm can be any type of data aggregation algorithm. Exemplarily, the preset data aggregation algorithm can be a tree-based aggregation algorithm, a data cluster-based aggregation algorithm, or a distributed data aggregation algorithm. The specific type can be flexibly set according to actual needs, and the present application does not make strict limitations. Among them, compared with the data in the database, the first sample data realizes sample compression of tens of millions of user data according to the model dimensions.
[0113] Among them, as an implementation manner, model score refers to the score corresponding to the data aggregated by the preset data aggregation algorithm, usually the result obtained after being evaluated and scored by a machine learning model or a statistical model on the compressed data. In the embodiments of the present application, the user data in the database is aggregated and compressed by adopting the preset data aggregation algorithm. Among them, the user data in the database includes various historical credit record data of users, such as the repayment history of users, the debt situation, the number of credit accounts, etc. A credit score is calculated by a preset credit evaluation model based on the data obtained according to the preset mathematical aggregation algorithm from the historical credit record data of users. This credit score is the model score, and the value range of the model score is 0 to 1000. The higher the score, the better the credit status of the user. The first sample data after aggregation still contains information such as the number of people in each customer group, the number of non-overdue borrowers, the outstanding balance, and the overdue amount.
[0114] Further, perform step S21-2. According to the aggregated user data, i.e., the first sample data, execute a preset prefix sum processing algorithm to calculate the statistical indicators corresponding to each historical customer stratification. Among them, the statistical indicators are statistical values set according to the actual application scenario. Exemplarily, taking the credit scenario as an example, the statistical indicators may include: the user quantity, the number of qualified users, and the overall risk. Among them, a qualified user may refer to a user who meets the set criteria. In the credit scenario, the qualified user may be a non-overdue user with an active loan. Non-overdue means that no overdue has occurred yet, and having an active loan means that the credit service is still being used. That is, by performing step S21-2, the user quantity, the number of non-overdue users with active loans, and the overall risk corresponding to each historical customer stratification can be obtained.
[0115] In the embodiment of the present application, although after being processed by the preset data aggregation algorithm, the amount of user data is greatly reduced compared to the original data amount in the database corresponding to the first sample data, the time consumed for calculating the statistical indicators corresponding to one customer stratification is still very large. For example, when calculating the user quantity, it is necessary to add up the number of people in each model score corresponding to each customer stratification of each customer group, and the computational complexity is O(N), where N is the number of model scores included in the customer stratification. Similarly, when calculating the overall risk, it is also necessary to add up the data where the model score is in a certain interval, and the corresponding computational complexity is also O(N), and the required time consumption is relatively long. To save time, in the embodiment of the present application, for this first sample data, a preset prefix sum processing algorithm is executed, and the first sample data is traversed according to the difference between the upper and lower bounds of the stratification, and the statistical indicators of each customer stratification can be quickly calculated.
[0116] Specifically, as shown in Figure 3b For a certain model score interval, assume it is [220, 890]. The vertical coordinate represents the number of users falling into this interval, and the horizontal coordinate represents the size of the model score of the first sample data. Then each point on the curve represents how many people have a model score greater than this model score. A, B, and C correspond to different customer groups. Further, if the abscissas corresponding to point F and point E in the figure are regarded as the upper bound parameter and the lower bound parameter of a certain customer stratification corresponding to the stratification, the stratification scores of each customer stratification can be directly calculated by the difference between the vertical coordinates of the two points, avoiding directly traversing and calculating the data of each person. The time complexity corresponding to this calculation method is O(1), effectively saving the calculation time. With the same calculation idea, the overall risk and the number of non-overdue users with active loans of each customer stratification can be quickly calculated.
[0117] As an example, as shown in Figure 4As shown, after data aggregation and data preprocessing of the current user data in the database by steps S21-1 and S21-2, the distribution results of the current user data on different customer hierarchies are obtained, that is, statistical indicators. Then, step S21-3 is executed to construct the correlation function between each statistical indicator and the lower bound parameter of the hierarchy, so as to perform gradient descent according to the calculated correlation function and dynamically determine the most appropriate lower bound parameter of the hierarchy.
[0118] As another implementation manner, the embodiment of the present application retrains the hierarchical model used in step S11 by executing steps S21 to S23, which can improve the hierarchical accuracy of the hierarchical model. Specifically, the hierarchical model can perform hierarchical processing on a large number of input sample data to determine the customer groups and customer hierarchies corresponding to each sample data, and then further calculate the statistical indicators corresponding to each customer hierarchy according to a preset statistical algorithm. Then, step S21 is performed with the help of the statistical indicators to construct the first function relationship, the second function relationship, and the third function relationship.
[0119] In some possible embodiments, the construction of the first function relationship, the second function relationship, and the third function relationship can be obtained through the following steps:
[0120] Based on the statistical indicators corresponding to each customer hierarchy, a target parameter matrix is constructed. Each column vector in the target parameter matrix is: the lower bound values of each hierarchy corresponding to each customer hierarchy in the same customer group; each row vector in the target parameter matrix is: the lower bound values of each customer group in the same customer hierarchy.
[0121] According to the target parameter matrix, with each row vector as the independent variable and the dimension corresponding to each statistical indicator as the dependent variable, the first function relationship, the second function relationship, and the third function relationship are respectively constructed.
[0122] Among them, the target parameter matrix can be x l =[x lA , x lB , x lc T . Among them, l represents the level of the customer hierarchy, and A, B, and C respectively represent different customer groups. Exemplarily, it can be as Figure 3a , and this target parameter matrix corresponds to each customer group and each customer hierarchy. Each column vector corresponds to the lower bound parameters of each customer hierarchy of a customer group, and each row vector corresponds to the lower bound parameters of each customer group in the same customer hierarchy. For example, the first column is the lower bound parameters of each layer from the first layer to the nth layer of customer group A, and the first row is the lower bound parameters of the first layer of customer group A, customer group B, and customer group C respectively.
[0123] Among them, the first function relationship is the user quantity yl The functional relationship with the hierarchical lower bound parameter x l is as follows: The second functional relationship is the number of compliant users z l and the hierarchical lower bound parameter x l The functional relationship between them is: z l = γ l (x l ). The third functional relationship is the overall risk r l and the hierarchical lower bound parameter x l The functional relationship between them is: r l = τ l (x l ). Among them, when the hierarchical lower bound parameters of the previous l - 1 hierarchies have been determined, is a functional relationship of the user scale in the customer group dimension, and specifically can be γ l is a functional relationship of the number of compliant users in the customer group dimension, and specifically can be γ lA , γ lB , γ lc , τ l is a functional relationship of the overall risk in the customer group dimension, and specifically can be τ lA , τ lB , τ lC . Among them, γ l , τ l can be similar to a black box, and the specific functional expressions can be obtained by fitting a large amount of data, and are not strictly limited here.
[0124] Furthermore, in some possible embodiments, the loss functions corresponding to each z l = γ l (x l ), r l = τ l (x l ) can be constructed through the following steps. Among them, the loss function corresponding to the first functional relationship is the user scale loss function loss y , the loss function corresponding to the second functional relationship is the compliant user loss function loss τ , and the loss function corresponding to the third functional relationship is the overall risk loss function loss γ . Each loss function can be constructed through the following steps:
[0125] Obtain the expected user quantity expected number of compliant users expected overall risk
[0126] The mean square error loss function loss between the expected number of users and the first functional relationship y is determined as the user magnitude loss function;
[0127] The mean square error loss function loss between the expected number of qualified users and the second functional relationship τ is determined as the qualified user loss function;
[0128] The mean square error loss function loss between the expected overall risk and the third functional relationship γ is determined as the overall risk loss function.
[0129] Specifically, for the mean square error loss function loss of the user magnitude y satisfies the following formula:
[0130]
[0131] For the mean square error loss function loss of the qualified users τ satisfies the following formula:
[0132]
[0133] For the mean square error loss function loss of the overall risk γ satisfies the following formula:
[0134]
[0135] In the embodiments of the present application, if the number of levels of customer stratification is N, the number of lower bound parameters of the corresponding constraints should be controlled within N - 1, and this condition should be satisfied for each customer group. Further, the goal of constructing the loss function is to make the number of users, the number of qualified users, and the overall risk distribution within the actually divided customer stratification conform to the expected effect of the enterprise or organization. Based on this, further, the sum function of each mean square error loss function can be calculated as the comprehensive loss function, that is, the comprehensive loss function Loss is determined based on the following formula:
[0136] Loss = loss y + loss τ + loss γ
[0137] Based on this, when performing step S22 and step S23, the target lower bound parameter can be determined through the following steps:
[0138] Derive the comprehensive loss function with respect to the three direction angles of y, τ, and γ, and respectively determine the gradient of the comprehensive loss function in the y direction The gradient of the comprehensive loss function in the τ direction and the gradient of the comprehensive loss function in the γ direction
[0139] According to the preset iteration step α, iterate on the following parameter iteration formula:
[0140]
[0141] Until the comprehensive loss function converges or until the number of iterations reaches the preset number of iterations;
[0142] Take the hierarchical lower bound parameter x when the comprehensive loss function converges or reaches the preset number of iterations l as the target hierarchical lower bound parameter.
[0143] Specifically, after taking the derivative of the comprehensive loss function in three directions, the gradient of the entire comprehensive loss function satisfies the following formula:
[0144]
[0145] where
[0146] where τ’ l and γ’ l are respectively τ l and γ l derivatives. Iterate according to the gradient descent algorithm. Let the iteration step be α each time. The value of the iteration step can be set flexibly. As a preferred implementation, it can be set to 0.1. The parameter iteration formula for layer l can be obtained:
[0147]
[0148] And perform gradient estimation. Since the statistical information of all users is known, so among them τ l (x l ), γ l (x l ) are known scalars, and the change rate of the number of users and risk near x l can be estimated as the gradient, and the vector x l is split into each customer group dimension. Taking the hierarchical parameter of customer group A in layer l as an example, the iteration formula for x lA can be obtained:
[0149]
[0150] Where Δ is a relatively small value used to estimate the gradient in the case of a continuous function. In the case where the model is divided into integers, it is further simplified. That is, the gradient is approximated as the difference.
[0151]
[0152] Among them, τ l (x l ) = τ lA (x lA ) + τ lB (x lB ) + τ lC (x lC ), γ l (x l ) = γ lA (x lA ) + γ lB (x lB ) + γ lC (x lC ). Thus, it has all been converted into computable values. The optimization of the boundary parameters for different customer groups and different stratifications is the same.
[0153] It can be combined with the flow chart as shown in Figure 5 to understand the above gradient descent calculation process:
[0154] Receive the expected values of each customer stratification passed in by the user, and then calculate the lower bound parameters of each customer stratification layer by layer.
[0155] First, calculate the lower bound parameter of the L-th layer, initialize the current lower bound parameter of the L-th layer, and solve the gradients in each direction of the L-th layer. Then update the lower bound parameter of the customer stratification of the L-th layer, and then calculate the error corresponding to the comprehensive loss function. Determine whether to return and execute the initialization of the current lower bound parameter of the L-th layer according to the error result, and iterate this process k times until the comprehensive loss function converges or the number of iterations reaches the set number of iterations. The set number of iterations can be flexibly set according to actual experience, and this application does not make strict restrictions.
[0156] There are two loops in the whole process:
[0157] Loop 1: Solve the lower bound parameter of each customer stratification layer by layer.
[0158] Loop 2: Calculate the lower bound parameter of each layer by gradient descent.
[0159] Among them, when initializing the current hierarchical lower bound parameter, the solution space divided by the model needs to be considered. For example, the parameter of the hierarchical lower bound parameter of layer 2 cannot exceed that of layer 1. That is, initializing the current hierarchical wire harness needs to be randomly initialized within a specific range, and the specific range size can be adjusted manually as a hyperparameter. The steps of solving the gradient and updating the hierarchical lower bound parameter are to solve the parameters of multiple customer groups in parallel according to the above gradient descent formula. When the comprehensive loss function converges, that is, the error value of the comprehensive loss function reaches the set threshold, or the number of iterations reaches the set number, the loop stops.
[0160] As described above, the number of customer hierarchies can be set according to the enterprise or organization. However, specifically, the way to determine the customer hierarchies that meet the number of customer hierarchies needs to be implemented by executing the above steps S21 to S23. Specifically,
[0161] After determining the hierarchical lower bound parameters of each customer hierarchy as described above, the hierarchical lower bound parameters can be applied to classify the subsequent input users to be recognized. In the subsequent application process, first execute step S11 to obtain the user characteristics of the user to be recognized. The user characteristics can be the user characteristics classified in real time by means of a user characteristic classification model, or the user characteristics that have been recognized and stored in advance obtained from a specified path. Among them, the user characteristics are a set of information used to describe various attributes and behavioral characteristics of the user. As an implementation manner, the user characteristics in the embodiments of the present application may be: the identity characteristics of the user, the behavioral characteristics of the user, the consumption characteristics of the user, the interest and hobby characteristics of the user, and so on. Exemplarily, as shown in 5, the user identity characteristics, behavioral characteristics, consumption characteristics, interests and hobbies, etc. can be obtained in advance to achieve accurate classification of the customer group and customer hierarchy to which the user to be recognized belongs.
[0162] Among them, the identity characteristics of the user are usually the characteristics at the demographic level of the user, and may include the gender, age, region, education level, occupation, etc. of the user. The behavioral characteristics of the user are the specific operation behaviors when the user uses the online service, such as clicking, browsing, downloading, and watching. The consumption characteristics of the user refer to the characteristics when the user generates consumption when using the online service, such as the purchase frequency, purchase price, loyalty to a certain brand, and so on. The interest and hobby characteristics of the user mainly represent the psychological preferences of the user. For example, the attitude towards a certain item or event is positive or negative. A positive attitude indicates that the user's psychological preference is liking.
[0163] Similarly, it can be as Figure 6As shown, the pre-trained hierarchical model in this application can be any type of machine learning model, specifically it can be a supervised learning model or an unsupervised learning model. The model structure and training process of the specific hierarchical model are not the focus of this application and will not be elaborated here. However, the pre-trained hierarchical model can learn the association relationship between user features and the customer groups and customer hierarchies to which the users belong based on the input user features, calculate the model scores of the users and the customer groups to which they belong, and then further divide the current users into specific hierarchies according to the hierarchical algorithm preset in the hierarchical model. This specific hierarchy is the customer hierarchy to which the user belongs. As described above, the customer hierarchy is flexibly set according to the actual application scenario requirements. As an example, it can be as Figure 6 shown, the customer hierarchy is specifically divided into: loan level 1, loan level 2, loan level 3, and so on.
[0164] Further, perform step S12. After all the user stratification results are completed, determine the statistical indicators corresponding to each customer hierarchy according to the preset statistical algorithm. Among them, the preset statistical algorithm varies according to different statistical indicators. Exemplarily, as Figure 6 shown, taking loan level 1 as an example, if the statistical indicator is the user quantity, the preset statistical algorithm for this user quantity can be summation. If the statistical indicator is the overall risk, the preset statistical algorithm for this overall risk can be to calculate the overdue probabilities of each user in the entire loan level 1 and then perform weighted summation. If the statistical indicator is the number of non-overdue borrowers with outstanding loans, the statistical algorithm for this number of non-overdue borrowers with outstanding loans is to sum up the number of people who have not overdue and still have outstanding loans in the entire loan level 1. Further, judge whether the current customer hierarchy meets the business requirements according to the statistical indicators corresponding to the customer hierarchy. If it meets, corresponding service strategies can be further provided for this customer hierarchy that meets the business requirements for service, so as to achieve targeted customer service, ensure the customer experience, optimize the service strategy, and save the service cost on the enterprise side.
[0165] In the second aspect, this application provides a device for clustering and stratifying target objects. Among them, as Figure 7 shown, the device includes:
[0166] A stratification processing module 701; obtain the user features of the user to be identified, input the user features into the pre-trained hierarchical model, and the hierarchical model determines the customer group and customer hierarchy to which the user belongs based on the input user features, where the customer hierarchy is obtained by dividing with different lower stratification boundary parameters;
[0167] A statistics module 702 determines statistical indicators corresponding to each customer stratification according to the customer stratification results and a preset statistical algorithm. The types of the statistical indicators are: the user quantity level of the customer stratification, the number of qualified users in the customer stratification, and the overall risk of the customer stratification.
[0168] A lower bound parameter calculation module 703 for the stratification is used to determine the lower bound parameter of the stratification in advance according to the following steps:
[0169] Based on the statistical indicators corresponding to each customer stratification, a first functional relationship, a second functional relationship, and a third functional relationship are respectively constructed. Among them, the first functional relationship is the functional relationship between the user quantity level and the lower bound parameter of the stratification, the second functional relationship is the functional relationship between the number of qualified users and the lower bound parameter of the stratification, and the third functional relationship is the functional relationship between the overall risk and the lower bound parameter of the stratification.
[0170] According to the expected user quantity of each customer stratification, the expected number of qualified users of each customer stratification, and the expected overall risk of each customer stratification, a user quantity loss function, a qualified user loss function, and an overall risk loss function are respectively constructed. Among them, the user quantity loss function is the loss function between the first functional relationship and the expected user quantity, the qualified user loss function is the loss function between the second functional relationship and the expected number of qualified users, and the overall risk loss function is the loss function between the third functional relationship and the expected overall risk.
[0171] Based on the user quantity loss function, the qualified user loss function, and the overall risk loss function, a comprehensive loss function is constructed, and a preset gradient descent algorithm is used to determine that the lower bound parameter of each customer stratification when the comprehensive loss function converges is the target lower bound parameter of the stratification.
[0172] Combined with the second aspect, in some possible embodiments, the lower bound parameter calculation module for the stratification is specifically used for:
[0173] Based on the statistical indicators corresponding to each customer stratification, a target parameter matrix is constructed. Each column vector in the target parameter matrix is: the lower bound value of each stratification corresponding to each customer stratification in the same customer group; each row vector in the target parameter matrix is: the lower bound value of each customer group in the same customer stratification.
[0174] According to the target parameter matrix, with each row vector as the independent variable and the dimension corresponding to each statistical indicator as the dependent variable, the first functional relationship, the second functional relationship, and the third functional relationship are respectively constructed.
[0175] In combination with the second aspect, in some possible embodiments, user magnitude loss functions, qualified user loss functions, and overall risk loss functions are respectively constructed according to the expected number of users in each customer stratification, the expected number of qualified users in each customer stratification, and the expected overall risk of each customer stratification.
[0176] Obtain the expected number of users, the expected number of qualified users, and the expected overall risk input by the user.
[0177] The mean square error loss function loss between the expected number of users and the first functional relationship y is determined as the user magnitude loss function.
[0178] The mean square error loss function loss between the expected number of qualified users and the second functional relationship τ is determined as the qualified user loss function.
[0179] The mean square error loss function loss between the expected overall risk and the third functional relationship γ is determined as the overall risk loss function.
[0180] In combination with the second aspect, in some possible embodiments, the lower bound parameter calculation module for stratification is specifically configured to:
[0181] The comprehensive loss function Loss is determined based on the following formula:
[0182] Loss = loss y + loss τ + loss γ
[0183] Derive the comprehensive loss function with respect to the three direction angles of y, τ, and γ to respectively determine the gradient of the comprehensive loss function in the y direction the gradient of the comprehensive loss function in the τ direction and the gradient of the comprehensive loss function in the γ direction
[0184] According to the preset iteration step α, iterate the following parameter iteration formula:
[0185]
[0186] Until the comprehensive loss function converges or until the number of iterations reaches the preset number of iterations;
[0187] Take the lower bound parameter x of the comprehensive loss function when it converges or reaches the preset number of iterations l as the target lower bound parameter for stratification.
[0188] In combination with the second aspect, in some possible embodiments, the apparatus further includes: a data preprocessing module, configured to:
[0189] Execute a preset data aggregation algorithm from a database to obtain first sample data;
[0190] Based on the first sample data, execute a preset prefix sum processing algorithm to calculate statistical indicators corresponding to each historical customer stratification;
[0191] Based on the statistical indicators corresponding to each historical customer stratification, respectively construct the first functional relationship, the second functional relationship, and the third functional relationship.
[0192] In combination with the second aspect, in some possible embodiments, the stratification lower bound parameter calculation module is further configured to:
[0193] Based on the expected number of users, the expected number of qualified users, and the expected overall risk of each customer stratification, according to the progressive relationship of the customer stratification, perform the following steps for each customer stratification layer by layer:
[0194] Determine the current stratification parameter corresponding to the current customer stratification, and based on the comprehensive loss function, calculate the gradients in the y, τ, and γ direction angles corresponding to the current customer stratification;
[0195] Based on the gradients in the y, τ, and γ direction angles corresponding to the current customer stratification, perform cyclic iteration using the parameter iteration formula;
[0196] Determine the customer stratification parameter corresponding to when the preset number of iterations is reached as the target stratification lower bound parameter.
[0197] In combination with the second aspect, in some possible embodiments, the user characteristics include: identity characteristics, behavior characteristics, consumption characteristics, and hobby characteristics.
[0198] Wherein, in this application, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0199] The names of the messages or information exchanged between multiple devices in the embodiments of this application are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0200] In a third aspect, an exemplary embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present application.
[0201] An exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present application.
[0202] An exemplary embodiment of the present application further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present application.
[0203] Referring to Figure 8 , a block diagram of an electronic device 800 that can be used as a server or a client of the present application will now be described. It is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0204] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM 802) or a computer program loaded from a storage unit 808 into a random access memory (RAM 803). In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output interface (I / O interface 805) is also connected to the bus 804.
[0205] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into the electronic device 800. The input unit 806 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 807 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, magnetic disks and optical discs. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0206] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above. For example, in some embodiments, the method for clustering and hierarchical processing of the aforementioned target objects can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to execute the method for clustering and hierarchical processing of the aforementioned target objects in any other suitable manner (e.g., by means of firmware).
[0207] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0208] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0209] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device (PLD)) that can be used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that can be used to provide machine instructions and / or data to a programmable processor.
[0210] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0211] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0212] A computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other.
Claims
1. A method for grouping and stratifying target objects, characterized in that: The method comprises: Acquire user features of the user to be identified, input the user features into a pre-trained hierarchical model, and determine the customer group and customer stratification to which the user belongs based on the input user features by the hierarchical model; According to the customer stratification results, according to a preset statistical algorithm, the statistical indicators corresponding to each customer stratification are determined, wherein the types of the statistical indicators include: the user level of the customer stratification, the number of qualified users of the customer stratification, and the overall risk of the customer stratification; The customer stratification is obtained by dividing the customer stratification by different stratification lower limit parameters, and the stratification lower limit parameters are pre-determined according to the following steps: Based on the statistical indicators corresponding to each of the customer strata, a first functional relationship, a second functional relationship, and a third functional relationship are respectively constructed, wherein the first functional relationship is a functional relationship between the user magnitude and the stratification lower limit parameter, the second functional relationship is a functional relationship between the number of qualified users and the stratification lower limit parameter, and the third functional relationship is a functional relationship between the overall risk and the stratification lower limit parameter; According to the expected number of users of each customer stratification, the expected number of qualified users of each customer stratification, and the expected overall risk of each customer stratification, a user level loss function, a qualified user loss function, and an overall risk loss function are respectively constructed; wherein the user level loss function is a loss function between the first functional relationship and the expected number of users, the qualified user loss function is a loss function between the second functional relationship and the expected number of qualified users, and the overall risk loss function is a loss function between the third functional relationship and the expected overall risk; A comprehensive loss function is constructed based on the user level loss function, the qualified user loss function, and the overall risk loss function. A preset gradient descent algorithm is used to determine the stratification lower bound parameters of each customer stratum when the comprehensive loss function reaches convergence as the target stratification lower bound parameters.
2. The method according to claim 1, characterized in that The first functional relationship, the second functional relationship, and the third functional relationship are respectively constructed based on the statistical indicators corresponding to each of the customer stratifications, including: Based on the statistical indicators corresponding to each of the customer strata, a target parameter matrix is constructed, wherein each column vector in the target parameter matrix is: the lower bound values of each stratum corresponding to each customer stratum of the same customer group; each row vector in the target parameter matrix is: the lower bound values of each stratum of each of the customer groups in the same customer stratum; According to the target parameter matrix, the first functional relationship, the second functional relationship, and the third functional relationship are constructed respectively with each row vector as an independent variable and the dimensions corresponding to each statistical indicator as a dependent variable.
3. The method according to claim 1, characterized in that According to the expected number of users of each customer stratification, the expected number of qualified users of each customer stratification, and the expected overall risk of each customer stratification, a user magnitude loss function, a qualified user loss function, and an overall risk loss function are respectively constructed; include: Obtain the expected number of users, the expected number of users who meet the requirements, and the expected overall risk input by the user; The mean square error loss function loss between the expected number of users and the first function relationship y , determined as the user-level loss function; The mean square error loss function loss between the expected number of users meeting the target and the second function τ , is determined as the qualified user loss function; The mean square error loss function loss between the expected overall risk and the third function relationship γ , determined as the overall risk loss function.
4. The method according to claim 3, characterized in that The constructing of a comprehensive loss function based on the user level loss function, the qualified user loss function, and the overall risk loss function includes: The comprehensive loss function Loss is determined based on the following formula: Loss=loss y +loss τ +loss γ The preset gradient descent algorithm is used to determine the stratification lower bound parameter of each customer stratification when the comprehensive loss function reaches convergence as the target stratification lower bound parameter, including: The comprehensive loss function is derived in the three directions of y, τ, and γ to determine the gradient of the comprehensive loss function in the y direction. The gradient of the comprehensive loss function in the τ direction And the gradient of the comprehensive loss function in the γ direction According to the preset iteration step α, the following parameter iteration formula is iterated: Until the comprehensive loss function converges, or until the number of iterations reaches a preset number of iterations; The layered lower bound parameter x is calculated when the comprehensive loss function converges or reaches the preset number of iterations. l The lower bound parameters for the target stratification.
5. The method according to claim 1, characterized in that In the process of respectively constructing the first functional relationship, the second functional relationship, and the third functional relationship based on the statistical indicators corresponding to each of the customer stratifications, the method further includes: Execute a preset data aggregation algorithm from the database to obtain first sample data; Based on the first sample data, a preset prefix and processing algorithm is executed to calculate the statistical indicators corresponding to each of the historical customer stratifications; Based on the historical statistical indicators corresponding to each of the customer stratifications, the first functional relationship, the second functional relationship, and the third functional relationship are respectively constructed.
6. The method according to claim 4, characterized in that The method further comprises: Based on the expected number of users, the expected number of users who meet the requirements, and the expected overall risk of each customer layer, and according to the progressive relationship of the customer layers, the following steps are performed for each customer layer layer by layer: Determine the current stratification parameters corresponding to the current customer stratification, and calculate the gradients of the three directions of y, τ, and γ corresponding to the current customer stratification based on the comprehensive loss function; Based on the gradients of the three directions of y, τ, and γ corresponding to the current customer stratification, the parameter iteration formula is used to perform loop iteration; The customer stratification parameter corresponding to when the preset number of iterations is reached is determined as the target stratification lower limit parameter.
7. The method according to claim 1, characterized in that The user characteristics include: identity characteristics, behavior characteristics, consumption characteristics, and interest and hobby characteristics.
8. A device for grouping and stratifying target objects, characterized in that: The device comprises: A stratification processing module; obtaining user features of a user to be identified, inputting the user features into a pre-trained stratification model, and having the stratification model determine the customer group and customer stratification to which the user belongs based on the input user features, wherein the customer stratification is obtained by dividing the customer stratification using different stratification lower bound parameters; A statistical module, based on the customer stratification results and in accordance with a preset statistical algorithm, determines statistical indicators corresponding to each customer stratification, wherein the types of statistical indicators include: user level of the customer stratification, number of qualified users of the customer stratification, and overall risk of the customer stratification; The layered lower bound parameter calculation module is used to determine the layered lower bound parameter in advance according to the following steps: Based on the statistical indicators corresponding to each of the customer strata, a first functional relationship, a second functional relationship, and a third functional relationship are respectively constructed, wherein the first functional relationship is a functional relationship between the user magnitude and the stratification lower limit parameter, the second functional relationship is a functional relationship between the number of qualified users and the stratification lower limit parameter, and the third functional relationship is a functional relationship between the overall risk and the stratification lower limit parameter; According to the expected number of users of each customer stratification, the expected number of qualified users of each customer stratification, and the expected overall risk of each customer stratification, a user level loss function, a qualified user loss function, and an overall risk loss function are respectively constructed; wherein the user level loss function is a loss function between the first functional relationship and the expected number of users, the qualified user loss function is a loss function between the second functional relationship and the expected number of qualified users, and the overall risk loss function is a loss function between the third functional relationship and the expected overall risk; A comprehensive loss function is constructed based on the user level loss function, the qualified user loss function, and the overall risk loss function. A preset gradient descent algorithm is used to determine the stratification lower bound parameters of each customer stratum when the comprehensive loss function reaches convergence as the target stratification lower bound parameters.
9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing a program; wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Customer layering method and device
CN113643118A
Marketing customer group screening method and device
CN115760180A
Client classification method and device based on artificial intelligence, equipment and storage medium
CN116680612A
Credit risk prediction method and device and electronic equipment
CN118037418A
Data processing method and continuous risk control data discretization method and device
CN118260533A