A method, device and electronic equipment for grouping and layering target objects
By dynamically determining the layer boundary values in the layered model and optimizing the lower bound parameters of the layer using the gradient descent algorithm, the shortcomings of manually setting the layer are solved, achieving more reasonable and efficient customer layer management and improving the effectiveness of marketing and advertising strategies.
Patent Information
- Application Number
- CN202510134389.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-02-07
AI Technical Summary
In existing technologies, the segmentation of target objects relies on manual settings, which lacks dynamism and precision, resulting in poor effectiveness of marketing and advertising strategies.
By obtaining user feature inputs to pre-train a hierarchical model, the hierarchical boundary values are dynamically determined, the gradient descent algorithm is used to optimize the lower bound parameters of the hierarchical model, and a loss function is constructed based on the user scale, the number of qualified users, and the overall risk, thereby achieving automated and precise hierarchical management.
Dynamically determined lower bound parameters are better suited to application scenarios, improving the rationality and effectiveness of marketing and advertising strategies while reducing costs.
Smart Images

Figure CN120125259B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a target object grouping and hierarchical processing method and device and electronic equipment. BACKGROUND
[0002] In the field of data processing, the division of the group and the division of the level to which a target object belongs become the key to the fine management of the target object. The target object can be a user or a commodity. Taking the user as the target object as an example, how to accurately determine the customer group to which the user belongs and the priority of the user in the customer group, and then develop corresponding marketing strategies and advertising strategies for the user according to the determined customer group and priority, become the key to providing targeted services to the user, and are also important factors for saving the marketing and advertising costs of the service provider.
[0003] However, the existing technical solutions all adopt manual division of levels and manual setting of hierarchical boundary values for how many levels are divided and how to set the values of the hierarchical boundaries between each level. SUMMARY
[0004] Therefore, the embodiments of the present application provide a target object grouping and hierarchical processing method, device and electronic equipment to dynamically determine the hierarchical boundary value and further improve the use effect of the hierarchical result.
[0005] In a first aspect, the embodiments of the present application provide a target object grouping and hierarchical processing method, and the method comprises the following steps.
[0006] Obtaining a user feature of a to-be-recognized user, inputting the user feature into a pre-trained hierarchical model, and determining a customer group and a customer hierarchy to which the user belongs based on the input user feature by the hierarchical model;
[0007] According to the customer hierarchy result, determining a statistical index corresponding to each customer hierarchy according to a preset statistical algorithm, wherein the type of the statistical index includes a user level of the customer hierarchy, a number of users meeting the standard of the customer hierarchy, and an overall risk of the customer hierarchy.
[0008] The customer hierarchy is divided by different hierarchical lower boundary parameters, and the hierarchical lower boundary parameters are determined in advance according to the following steps.
[0009] construct a first function relationship, a second function relationship, and a third function relationship based on the statistical indicators corresponding to each of the customer stratifications, wherein the first function relationship is a function relationship between the user magnitude and the stratification lower bound parameter, the second function relationship is a function relationship between the number of users meeting the standard and the stratification lower bound parameter, and the third function relationship is a function relationship between the overall risk and the stratification lower bound parameter;
[0010] construct a user magnitude loss function, a user meeting the standard loss function, and an overall risk loss function based on the expected number of users of each of the customer stratifications, the expected number of users meeting the standard of each of the customer stratifications, and the expected overall risk of each of the customer stratifications, wherein the user magnitude loss function is a loss function between the first function relationship and the expected number of users, the user meeting the standard loss function is a loss function between the second function relationship and the expected number of users meeting the standard, and the overall risk loss function is a loss function between the third function relationship and the expected overall risk;
[0011] construct a comprehensive loss function based on the user magnitude loss function, the user meeting the standard loss function, and the overall risk loss function, and determine the stratification lower bound parameter of each customer stratification as a target stratification lower bound parameter when the comprehensive loss function converges by using a preset gradient descent algorithm.
[0012] In some possible embodiments, the first function relationship, the second function relationship, and the third function relationship are respectively constructed based on the statistical indicators corresponding to each of the customer stratifications, including:
[0013] construct a target parameter matrix based on the statistical indicators corresponding to each of the customer stratifications, wherein each column vector in the target parameter matrix is a lower bound value of each stratification corresponding to a same customer stratification of a same customer group, and each row vector in the target parameter matrix is a stratification lower bound value of each of the customer groups in a same customer stratification.
[0014] construct the first function relationship, the second function relationship, and the third function relationship based on the target parameter matrix, with each of the row vectors as an independent variable and the dimensions corresponding to each of the statistical indicators as a dependent variable.
[0015] In some possible embodiments, the user magnitude loss function, the user meeting the standard loss function, and the overall risk loss function are respectively constructed based on the expected number of users of each of the customer stratifications, the expected number of users meeting the standard of each of the customer stratifications, and the expected overall risk of each of the customer stratifications, including:
[0016] obtain the expected number of users, the expected number of users meeting the standard, and the expected overall risk input by a user;
[0017] a mean squared error loss function between the expected number of users and the first function relationship determining the user magnitude loss function
[0018] a mean squared error loss function between the expected number of users meeting the target and the second function relationship determining the target meeting user loss function
[0019] a mean squared error loss function between the expected overall risk and the third function relationship determining the overall risk loss function
[0020] In some possible embodiments, the method further comprises:
[0021] determining the comprehensive loss function based on the user magnitude loss function, the target meeting user loss function, and the overall risk loss function Loss based on the following formula:
[0022]
[0023] determining the target stratification lower bound parameter as the stratification lower bound parameter when the comprehensive loss function converges, by using a preset gradient descent algorithm, comprises:
[0024] deriving the comprehensive loss function in three direction angles, respectively determining the gradient of the comprehensive loss function in the direction , the gradient of the comprehensive loss function in the direction , and the gradient of the comprehensive loss function in the direction ; ;
[0025] iterating the following parameter iteration formula according to a preset iteration step size α:
[0026]
[0027] until the comprehensive loss function converges, or until the iteration number reaches a preset iteration number
[0028] the stratification lower bound parameter when the comprehensive loss function converges or reaches the preset iteration number is the target stratification lower bound parameter.
[0029] In some possible embodiments, in the process of constructing the first function relationship, the second function relationship, and the third function relationship based on the statistical indicators corresponding to each of the customer stratifications, respectively, the method further comprises:
[0030] performing a preset data aggregation algorithm on the database to obtain first sample data;
[0031] performing a preset prefix and processing algorithm based on the first sample data to calculate statistical indicators corresponding to each of the historical customer stratifications;
[0032] Based on the statistical indicators corresponding to each of the historical customer stratifications, the first function relationship, the second function relationship, and the third function relationship are respectively constructed.
[0033] In some possible embodiments, the method further comprises:
[0034] Based on the expected number of users, the expected number of users meeting the target, and the expected overall risk of each customer stratification, and according to the progressive relationship of the customer stratifications, the following steps are performed on each customer stratification layer by layer:
[0035] determining a current stratification parameter corresponding to a current customer stratification, and calculating a gradient in three direction angles corresponding to the current customer stratification based on the comprehensive loss function;
[0036] Based on the gradient in three direction angles corresponding to the current customer stratification, the parameter iterative formula is used for cyclic iteration;
[0037] When the preset number of iterations is reached, the customer stratification parameter corresponding thereto is determined as the target stratification lower bound parameter.
[0038] In some possible embodiments, the user features include identity features, behavior features, consumption features, and interest and hobby features.
[0039] In a second aspect, an embodiment of the present application provides a target object grouping and stratification processing apparatus, and the apparatus comprises:
[0040] a stratification processing module that obtains user features of a user to be identified, inputs the user features into a pre-trained stratification model, and determines a customer group to which the user belongs and a customer stratification based on the input user features by using the stratification model, wherein the customer stratification is divided by different stratification lower bound parameters;
[0041] a statistics module that determines statistical indicators corresponding to each customer stratification according to a preset statistical algorithm based on the customer stratification result, wherein the types of the statistical indicators include a user magnitude of a customer stratification, a number of users meeting a target of a customer stratification, and an overall risk of a customer stratification;
[0042] The hierarchical lower bound parameter calculation module is configured to determine the hierarchical lower bound parameters according to the following steps in advance:
[0043] Based on the statistical indicators corresponding to each customer hierarchy, a first function relationship, a second function relationship, and a third function relationship are respectively constructed, wherein the first function relationship is a function relationship between the user magnitude and the hierarchical lower bound parameter, the second function relationship is a function relationship between the number of users meeting the standard and the hierarchical lower bound parameter, and the third function relationship is a function relationship between the overall risk and the hierarchical lower bound parameter.
[0044] According to the expected number of users of each customer hierarchy, the expected number of users meeting the standard of each customer hierarchy, and the expected overall risk of each customer hierarchy, a user magnitude loss function, a user meeting the standard loss function, and an overall risk loss function are respectively constructed, wherein the user magnitude loss function is a loss function between the first function relationship and the expected number of users, the user meeting the standard loss function is a loss function between the second function relationship and the expected number of users meeting the standard, and the overall risk loss function is a loss function between the third function relationship and the expected overall risk.
[0045] Based on the user magnitude loss function, the user meeting the standard loss function, and the overall risk loss function, a comprehensive loss function is constructed, and a preset gradient descent algorithm is used to determine that the hierarchical lower bound parameters of each customer hierarchy when the comprehensive loss function reaches convergence are target hierarchical lower bound parameters.
[0046] In combination with the second aspect, in some possible embodiments, the hierarchical lower bound parameter calculation module is specifically configured to:
[0047] Based on the statistical indicators corresponding to each customer hierarchy, a target parameter matrix is constructed, each column vector in the target parameter matrix is a lower bound value of each customer hierarchy of a same customer group, and each row vector in the target parameter matrix is a hierarchical lower bound value of each customer group in a same customer hierarchy.
[0048] According to the target parameter matrix, each row vector is taken as an independent variable, and a dimension corresponding to each statistical indicator is taken as a dependent variable, and the first function relationship, the second function relationship, and the third function relationship are respectively constructed.
[0049] In combination with the second aspect, in some possible embodiments, the user magnitude loss function, the user meeting the standard loss function, and the overall risk loss function are respectively constructed according to the expected number of users of each customer hierarchy, the expected number of users meeting the standard of each customer hierarchy, and the expected overall risk of each customer hierarchy.
[0050] obtaining the expected user quantity, the expected qualified user quantity and the expected overall risk input by the user;
[0051] determining the expected user quantity and the first function relationship as a mean square error loss function , and determining the user magnitude loss function;
[0052] determining the expected qualified user quantity and the second function relationship as a mean square error loss function , and determining the qualified user loss function;
[0053] determining the expected overall risk and the third function relationship as a mean square error loss function , and determining the overall risk loss function.
[0054] In combination with the second aspect, in some possible embodiments, the hierarchical lower bound parameter calculation module is specifically configured to:
[0055] determining the comprehensive loss function Loss based on the following formula:
[0056]
[0057] deriving the comprehensive loss function in three direction angles, respectively determining the gradient of the comprehensive loss function in the direction , the gradient of the comprehensive loss function in the direction and the gradient of the comprehensive loss function in the direction ; ;
[0058] iterating the following parameter iteration formula according to a preset iteration step size α:
[0059]
[0060] until the comprehensive loss function converges or until the iteration number reaches a preset iteration number;
[0061] the hierarchical lower bound parameter when the comprehensive loss function converges or reaches the preset iteration number is the target hierarchical lower bound parameter.
[0062] In combination with the second aspect, in some possible embodiments, the apparatus further includes a data preprocessing module configured to:
[0063] executing a preset data aggregation algorithm on a database to obtain first sample data;
[0064] Based on the first sample data, a preset prefix sum processing algorithm is executed to calculate the statistical indicators corresponding to each of the historical customer segments.
[0065] Based on the statistical indicators corresponding to each of the historical customer segments, the first functional relationship, the second functional relationship, and the third functional relationship are constructed respectively.
[0066] In conjunction with the second aspect, in some possible embodiments, the hierarchical lower bound parameter calculation module is further configured to:
[0067] Based on the expected number of users, the expected number of users meeting the target, and the expected overall risk for each customer segment, the following steps are performed layer by layer for each customer segment according to the progressive relationship of the customer segments:
[0068] Determine the current customer segment parameters corresponding to the current segment, and calculate the corresponding parameters for the current customer segment based on the comprehensive loss function. Gradient in three angular directions;
[0069] Based on the current customer segmentation The gradients at the three directional angles are iterated cyclically using the aforementioned parameter iteration formula;
[0070] The customer segmentation parameter corresponding to the preset number of iterations is determined as the target segmentation lower bound parameter.
[0071] In conjunction with the second aspect, in some possible embodiments, the user characteristics include: identity characteristics, behavioral characteristics, consumption characteristics, and interest and hobby characteristics.
[0072] Thirdly, embodiments of this application provide an electronic device, wherein the electronic device includes: a processor; and a memory storing a program; wherein the program includes instructions, which, when executed by the processor, cause the processor to perform the method for grouping and hierarchical processing of target objects as described in the first aspect.
[0073] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to cause a computer to execute the method for grouping and hierarchical processing of the target object as described in the first aspect.
[0074] The beneficial effects of this application are:
[0075] The application provides a target object grouping and hierarchical processing method, device and electronic equipment, wherein the method comprises the following steps: obtaining a user feature of a to-be-identified user, inputting the user feature into a pre-trained hierarchical model, determining a customer group to which the user belongs and a customer hierarchy based on the input user feature by the hierarchical model, and calculating a user level, a number of qualified users and an overall risk corresponding to each customer hierarchy according to a preset statistical algorithm according to the customer hierarchy result. In the embodiment of the application, the customer hierarchy is divided by different hierarchical lower bound parameters, and the hierarchical lower bound parameters are pre-constructed by the user level, the number of qualified users and the overall risk corresponding to each customer hierarchy, respectively, and then the user level loss function, the qualified user loss function and the overall risk loss function are constructed according to the constructed function relationship and the set expected user quantity, the expected number of qualified users and the expected overall risk, respectively. Finally, based on each loss function, a comprehensive loss function is constructed, and a preset gradient descent algorithm is used to determine that the hierarchical lower bound parameters of each customer hierarchy are the target hierarchical lower bound parameters corresponding to each customer hierarchy when the comprehensive loss function reaches convergence.
[0076] By selecting the embodiment of the application, the scheme of manually setting hierarchical lower bound parameters to divide different customer hierarchies in the prior art is abandoned, the expected number of users to be achieved, the expected number of qualified users to be achieved and the expected overall risk to be controlled corresponding to each customer hierarchy are dynamically obtained, the preset gradient descent algorithm is used to solve the corresponding hierarchical lower bound parameters when the expected effect is achieved, and the hierarchical lower bound parameters obtained in this way are not fixed, but adapt to the statistical indicators required by the application scenario, so that the hierarchical lower bound parameters obtained are more reasonable, and the hierarchical results of different customer hierarchies under different customer groups are more helpful for enterprises and institutions to develop more reasonable and cost-effective marketing strategies and advertising strategies, and the application effect is better. BRIEF DESCRIPTION OF DRAWINGS
[0077] In the following description of exemplary embodiments in conjunction with the accompanying drawings, more details, features and advantages of the application are disclosed, and in the drawings:
[0078] Figure 1 A flowchart of a method for grouping and hierarchical processing of target objects provided by an embodiment of the application is shown;
[0079] Figure 2 A flowchart of a method for determining hierarchical lower bound parameters provided by an embodiment of the application is shown;
[0080] Figure 3a Another flowchart of a method for grouping and hierarchical processing of target objects provided by an embodiment of the application is shown;
[0081] Figure 3bFig. 1 shows a processing schematic diagram of the prefix and processing provided by an embodiment of the present application;
[0082] Figure 4 Fig. 2 shows another flow schematic diagram of the method for the grouping and hierarchical processing of the target objects provided by an embodiment of the present application;
[0083] Figure 5 Fig. 3 shows another flow schematic diagram of the method for the grouping and hierarchical processing of the target objects provided by an embodiment of the present application;
[0084] Figure 6 Fig. 4 shows another flow schematic diagram of the method for the grouping and hierarchical processing of the target objects provided by an embodiment of the present application;
[0085] Figure 7 Fig. 5 shows a logic structure schematic diagram of the apparatus for the grouping and hierarchical processing of the target objects provided by an embodiment of the present application;
[0086] Figure 8 Fig. 6 shows a structure block diagram of an exemplary electronic device that can be used to implement embodiments of the present application. DETAILED DESCRIPTION
[0087] Embodiments of the present application will be described in more detail with reference to the drawings. While certain embodiments of the present application will be shown in the drawings and described below, it should be understood that the present application can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the application to those skilled in the art.
[0088] It should be understood that the various steps of the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present application is not limited in this respect.
[0089] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising but not limited to." The term "based on" is "based at least in part on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related terms are defined in the description that follows. It should be noted that reference to a "first," "second," etc. concept does not limit the scope of these concepts in the claims to the concepts with the corresponding ordinal numbering. Rather, such references are merely intended to distinguish different concepts from one another.
[0090] It should be noted that the modification of "one" and "multiple" mentioned in the present application is illustrative but not restrictive, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0091] In a first aspect, the present application provides a method for grouping and stratifying target objects, which is applied to any electronic device with the function of grouping and stratifying target objects, including but not limited to personal mobile terminals, computers or servers, etc. As shown in the figure, the method comprises the following steps: Figure 1
[0092] S11, obtaining the user characteristics of the user to be identified, inputting the user characteristics into the stratification model trained in advance, and determining the customer group and customer stratification to which the user belongs based on the input user characteristics by the stratification model;
[0093] S12, according to the customer stratification result, determining the statistical indicators corresponding to each customer stratification according to the preset statistical algorithm, wherein the types of statistical indicators include the user level of customer stratification, the number of users meeting the standard of customer stratification, and the overall risk of customer stratification.
[0094] Among them, the above-mentioned customer stratification is obtained by different stratification lower bound parameters, as shown in the figure, the stratification lower bound parameter is determined in advance according to the following steps: Figure 2
[0095] S21, based on the statistical indicators corresponding to each customer stratification, respectively constructing a first function relationship, a second function relationship and a third function relationship, wherein the first function relationship is the function relationship between the user level and the stratification lower bound parameter, the second function relationship is the function relationship between the number of users meeting the standard and the stratification lower bound parameter, and the third function relationship is the function relationship between the overall risk and the stratification lower bound parameter;
[0096] S22, according to the expected number of users of each customer stratification, the expected number of users meeting the standard of each customer stratification, and the expected overall risk of each customer stratification, respectively constructing a user level loss function, a user meeting the standard loss function, and an overall risk loss function; wherein the user level loss function is the loss function between the first function relationship and the expected number of users, the user meeting the standard loss function is the loss function between the second function relationship and the expected number of users meeting the standard, and the overall risk loss function is the loss function between the third function relationship and the expected overall risk;
[0097] S23, constructing a comprehensive loss function based on the user level loss function, the target user loss function, and the overall risk loss function, and determining the target stratification lower bound parameter of each customer stratification when the comprehensive loss function reaches convergence by using a preset gradient descent algorithm.
[0098] The method obtains user features of a user to be identified, inputs the user features into a pre-trained stratification model, determines a customer group to which the user belongs and a customer stratification based on the input user features by the stratification model, and calculates the user level, the number of target users, and the overall risk corresponding to each customer stratification according to a preset statistical algorithm according to the customer stratification result. In the embodiments of the present application, the stratification lower bound parameter is obtained by dividing different customer stratifications, and the stratification lower bound parameter is pre-constructed by the user level, the number of target users, and the overall risk corresponding to each customer stratification, respectively, and the function relationship between the stratification lower bound parameter, and then the function relationship is constructed according to the constructed function relationship and the expected number of users, the expected number of target users, and the expected overall risk, respectively. The user level loss function, the target user loss function, and the overall risk loss function are constructed, and finally the comprehensive loss function is constructed based on each loss function, and the stratification lower bound parameter of each customer stratification when the comprehensive loss function reaches convergence is determined as the target stratification lower bound parameter corresponding to each customer stratification by using a preset gradient descent algorithm.
[0099] The embodiments of the present application abandon the scheme of manually setting stratification lower bound parameters to divide different customer stratifications in the prior art, and continuously dynamically obtain the expected number of users, the expected number of target users, and the expected overall risk corresponding to each customer stratification, and solve the stratification lower bound parameter corresponding to the expected effect by using a preset gradient descent algorithm. The stratification lower bound parameter obtained in this way is not fixed, but continuously adapts to the statistical indicators required by the application scenario, and the stratification lower bound parameter obtained is more reasonable. The stratification results of different customer stratifications under different customer groups in actual use can help enterprises and institutions to develop more reasonable and cost-effective marketing strategies and advertising strategies, and the application effect is better.
[0100] The above steps S11, S12, and S21-S23 will be described in detail below with specific examples:
[0101] Before the steps S11 to S12 and the steps S21 to S23 are described in detail, the execution order between the steps is explained here: in the embodiment of the present application, the steps S11 and S12 are to divide the inputted to-be-recognized user by using the pre-trained hierarchical model after the execution of the steps S21 to S23 determines the corresponding hierarchical lower bound parameters between the customer hierarchies of each customer group, to determine which customer group the user belongs to and which customer hierarchy in the customer group the user belongs to. If the application scenario of the present application is compared to an Excel table, each column corresponds to a customer group, and each row in the column corresponds to each customer hierarchy in the customer group. The steps S11 to S12 are to attribute and locate the inputted to-be-recognized user, to determine which column and which row the user belongs to. The division standard between the specific rows needs to be determined in advance by executing the steps S21 to S23.
[0102] After understanding the execution order between the steps, it can be seen that in the embodiment of the present application, the steps S21 to S23 are the core and key to dynamically determine the customer hierarchy, so the steps S21 to S23 will be described first, and then the steps S11 to S12 will be described.
[0103] In the embodiment of the present application, the customer group refers to the customer group, specifically to the customer group that the products and services of an enterprise or organization face and has some common characteristics. The customer hierarchy is a subset of the customer group that is further subdivided according to different dimensions of a certain characteristic in the same customer group. For example, in the loan business scenario, the customer group can be divided into: high net worth customer group, medium net worth customer group, and low net worth customer group. Then, the customers in the high net worth customer group can be further divided into first priority high net worth customers, second priority high net worth customers, third priority high net worth customers, and so on according to the specific asset size. In the embodiment of the present application, the specific division method of the customer group and the customer hierarchy depends on the selected division characteristic in the actual application process, in other words, the customer group and the customer hierarchy are flexibly set according to the actual application scenario requirements. For example, the number of divided customer groups and the number of customer hierarchies in each customer group can be set by the enterprise or organization according to the operation requirements. If fine-grained management and operation are required, a larger number of customer groups and a larger number of customer hierarchies can be set. If the operation and management cost needs to be saved, a smaller number of customer groups and a smaller number of customer hierarchies can be set. The specific types of the customer group and the customer hierarchy will not be strictly specified in this paper.
[0104] The customer stratification or the customer layering is for the actual application scene. In the embodiment of the application, the application scene of the method can be a lending scene. The characteristics of the credit users can be extracted, and then the specific credit risk customer group to which the user belongs and the specific credit stratification of the credit user in the credit risk customer group can be determined based on the characteristics of the user. This is helpful for the credit institution or enterprise to provide more accurate credit services and control credit risks, maximize resource benefits, and balance the safety and growth space of the credit business.
[0105] That is, the method provided in the application can be applied to risk control customer stratification. The customers are divided into multiple levels such as loan one and loan two based on the value and risk of the user. The higher the value of the user in the loan one stratification, the lower the value of the user and the higher the risk as the level progresses. Based on this, different service resources can be allocated to customers in different customer layers during the risk control process. For example, high-value customers in the earlier level can obtain higher-level service strategies and attention, as well as more detailed operation services. For the credit institution, this approach can avoid investing too many resources in low-value or high-risk customers and reduce the investment cost of the institution or enterprise.
[0106] Based on this, when step S21 is performed, the customer can be initialized and stratified in advance. The customer group and customer stratification to which the customer belongs are divided in advance by the machine learning model, and then the initialization stratification result is obtained. In the initialization stratification result, each customer stratification also contains the corresponding user. Then, the statistical indicators of each customer stratification obtained in the initialization are counted in the same way as step S12. In some possible embodiments, the first function relationship, the second function relationship, and the third function relationship can be constructed based on the statistical indicators corresponding to each customer stratification stored in the database in advance. As an implementation, when step S21 is performed, the following steps can be implemented:
[0107] S21-1, execute a preset data aggregation algorithm from the database to obtain first sample data.
[0108] S21-2, based on the first sample data, execute a preset prefix and processing algorithm to calculate the statistical indicators corresponding to each of the historical customer stratifications.
[0109] S21-3, based on the statistical indicators corresponding to each of the historical customer stratifications, respectively construct the first function relationship, the second function relationship, and the third function relationship.
[0110] In the embodiment of the present application, the type of the database can be a distributed database, and the historical user stratification records and the stratification results are recorded in different data sources of the distributed database. In step S21-1, data is obtained from each data source, and then the obtained data is combined and refined to obtain a more complete and valuable sample data for determining the stratification lower bound parameter, which is the first sample data.
[0111] The stratification lower bound parameter is the basis for dividing different customer stratifications, and specifically is the boundary value between two adjacent customer stratifications. For example, if it is required to divide user data into three customer groups, customer group A, customer group B, and customer group C, and then each customer group is further divided into n customer stratifications x Figure 3a 1A nA 1B nB 1C nC The boundary value between each customer stratification is the stratification lower bound parameter, wherein the first stratification lower bound parameter is the boundary value between the first customer stratification and the second customer stratification.
[0112] In the embodiment of the present application, the preset data aggregation algorithm can be any type of data aggregation algorithm. For example, the preset data aggregation algorithm can be a tree structure-based aggregation algorithm, a data cluster-based aggregation algorithm, or a distributed data aggregation algorithm. The specific type can be flexibly set according to actual needs, and the present application is not strictly limited. Compared with the data in the database, the first sample data realizes sample compression of ten million user data according to the model dimension.
[0113] As an embodiment, the model score refers to the score corresponding to the data aggregated by the preset data aggregation algorithm. The result is usually obtained by evaluating and scoring the compressed data by a machine learning model or a statistical model. In the embodiment of the present application, the user data in the database is aggregated and compressed by using the preset data aggregation algorithm, wherein the user data in the database includes various historical credit record data of the user, such as the user's repayment history, debt situation, and credit account number. The data obtained by the preset credit evaluation model based on the historical credit record data of the user according to the preset mathematical aggregation algorithm is calculated, and a credit score is calculated, which is the model score. The model score has a value range of 0-1000, and the higher the score, the better the credit status of the user. The aggregated first sample data still contains the number of people in each customer group, the number of non-overdue loan, the loan balance, and the overdue amount.
[0114] Further, step S21-2 is performed to execute a preset prefix sum processing algorithm according to the converged user data, i.e., the first sample data, to calculate statistical indicators corresponding to each customer layer. The statistical indicators are statistical values set according to actual application scenarios. For example, in a credit scenario, the statistical indicators can include user magnitude, number of users meeting a standard, and overall risk. The users meeting the standard can refer to users meeting a set standard. In the credit scenario, the users meeting the standard can refer to non-overdue users, i.e., users who have not generated overdue credit. The users still using credit services are non-overdue users. That is, step S21-2 can obtain the user magnitude, number of non-overdue users, and overall risk corresponding to each historical customer layer.
[0115] In the embodiment of the present application, although the amount of user data is greatly reduced from the original amount of data in the database to the amount of data corresponding to the first sample data after the preset data convergence algorithm is processed, the time consumption for calculating the statistical indicators corresponding to each customer layer is still large. For example, when calculating the user magnitude, the number of users corresponding to each model segment of each customer layer needs to be added and summed, and the calculation complexity is O(N), where N is the number of model segments containing the customer layer. Similarly, when calculating the overall risk, the data in the model segment in a certain interval also needs to be added and summed, and the corresponding calculation complexity is also O(N), which requires a long time. In order to save time, in the embodiment of the present application, a preset prefix sum processing algorithm is executed for the first sample data, and the first sample data is traversed according to the difference between the upper and lower bounds of the layer to quickly calculate the statistical indicators of each customer layer.
[0116] Specifically, as shown in Figure 3b , for a certain model segment interval, assume that the interval is [220, 890], the vertical coordinate represents the number of users falling into the interval, and the horizontal coordinate represents the size of the model segment of the first sample data. Each point on the curve represents the number of users whose model segment is greater than the model segment. A, B, and C correspond to different customer groups. Further, if the horizontal coordinates of points F and E in the figure are regarded as the layer upper bound parameter and the layer lower bound parameter corresponding to each customer layer, the layer score of each customer layer can be directly calculated by the difference between the vertical coordinates of the two points, avoiding direct traversal calculation of each user's data. The time complexity of this calculation method is O(1), which effectively saves the calculation time. Using the same calculation method, the overall risk and the number of non-overdue users of each customer layer can be quickly calculated.
[0117] As an example, as shown in Figure 4As shown, after the data aggregation and data preprocessing of the current user data in the database by steps S21-1 and S21-2, the distribution of the current user data on different customer hierarchies, i.e., the statistical indicators, are obtained, and then step S21-3 is performed to construct the correlation function between each statistical indicator and the lower bound parameter of the hierarchy, so that the most appropriate lower bound parameter of the hierarchy is dynamically determined by gradient descent according to the calculated correlation function.
[0118] As another implementation, the embodiment of the present application can improve the hierarchy precision of the hierarchy model used in step S11 by performing steps S21 to S23 to perform secondary training on the hierarchy model. Specifically, the hierarchy model can be used to perform hierarchy processing on a large amount of sample data input, determine the customer group and customer hierarchy corresponding to each sample data, and then further calculate the statistical indicators corresponding to each customer hierarchy according to the preset statistical algorithm. Then, the statistical indicators are used to perform step S21 to construct the first function relationship, the second function relationship, and the third function relationship.
[0119] In some possible embodiments, the first function relationship, the second function relationship, and the third function relationship can be obtained by the following steps:
[0120] Based on the statistical indicators corresponding to each customer hierarchy, a target parameter matrix is constructed, each column vector in the target parameter matrix is the lower bound value of each hierarchy corresponding to each customer hierarchy of the same customer group, and each row vector in the target parameter matrix is the hierarchy lower bound value of each customer group in the same customer hierarchy.
[0121] According to the target parameter matrix, each row vector is taken as an independent variable, and the dimension corresponding to each statistical indicator is taken as a dependent variable, and the first function relationship, the second function relationship, and the third function relationship are constructed respectively.
[0122] The target parameter matrix can be wherein, l represents the level of the customer hierarchy, and A, B, and C represent different customer groups. For example, Figure 3a The target parameter matrix corresponds to each customer group and each customer hierarchy. Each column vector corresponds to the hierarchy lower bound parameter of each customer hierarchy of a customer group, and each row vector corresponds to the hierarchy lower bound parameter of each customer group in the same customer hierarchy. For example, the first column is the hierarchy lower bound parameter of the first layer to the nth layer of the customer group A, and the first row is the hierarchy lower bound parameter of the first layer of the customer group A, the customer group B, and the customer group C.
[0123] The first function relationship is the function relationship between the user level and the hierarchy lower bound parameter , the second function relationship is the number of users reaching the target and the hierarchical lower bound parameter between them: , the third function relationship is the overall risk and the hierarchical lower bound parameter between them: . Wherein, in the case that the hierarchical lower bound parameter of the previous l -1 hierarchy has been determined, is the function relationship of the user magnitude in the customer group dimension, which can be , is the function relationship of the number of users reaching the target in the customer group dimension, which can be , is the function relationship of the overall risk in the customer group dimension, which can be . Wherein, , , can be similar to a black box, and the specific function expression can be obtained by fitting a large amount of data, which is not strictly limited here.
[0124] Further, in some possible embodiments, each , , corresponding loss function can be constructed by the following steps. Wherein, the loss function corresponding to the first function relationship is the user magnitude loss function , the loss function corresponding to the second function relationship is the user reaching the target loss function , and the loss function corresponding to the third function relationship is the overall risk loss function . Each loss function can be constructed by the following steps:
[0125] obtaining the expected number of users , the expected number of users reaching the target , and the expected overall risk ;
[0126] determining the mean square error loss function between the expected number of users and the first function relationship as the user magnitude loss function;
[0127] determining the mean square error loss function between the expected number of users reaching the target and the second function relationship as the user reaching the target loss function;
[0128] determining the mean square error loss function between the expected overall risk and the third function relationship as the overall risk loss function.
[0129] Specifically, for the mean squared error loss function at the user scale. Satisfy the following formula:
[0130]
[0131] For qualified users, the mean squared error loss function Satisfy the following formula:
[0132]
[0133] For the mean squared loss function of overall risk Satisfy the following formula:
[0134]
[0135] In this embodiment, if the number of customer stratification levels is N, then the number of lower bound parameters for each stratification level corresponding to the constraint terms should be controlled at N-1, and this condition must be met for each customer group. Furthermore, the goal of constructing the loss function is to ensure that the number of users, qualified users, and overall risk distribution within the actual customer stratification matches the expected results of the enterprise or organization. Based on this, the comprehensive loss function can be calculated by summing the mean squared error loss functions, i.e., determined based on the following formula. Loss :
[0136]
[0137] Based on this, when performing steps S22 and S23, the target layer lower bound parameter can be determined through the following steps:
[0138] For the aforementioned comprehensive loss function, in Differentiate the functions in three directions to determine the comprehensive loss function at each angle. gradient in direction The comprehensive loss function is in gradient in direction And the comprehensive loss function in gradient in direction ;
[0139] Based on the preset iteration step size α, the following parameter iteration formula is iterated:
[0140]
[0141] Until the comprehensive loss function converges, or until the number of iterations reaches the preset number of iterations;
[0142] The comprehensive loss function converges, or the hierarchical lower bound parameter reaches the preset number of iterations. the target hierarchical lower bound parameter.
[0143] Specifically, after the comprehensive loss function is derived in three directions, the gradient of the entire comprehensive loss function is satisfies the following formula:
[0144]
[0145] wherein, ; ; .
[0146] wherein , and are the derivatives of , and respectively. According to the gradient descent algorithm iteration, let the iteration step size be , the value of the iteration step size can be flexibly set, and as a preferred embodiment, it can be set to 0.1, and the parameter iteration formula of the hierarchy can be obtained:
[0147]
[0148] and perform gradient estimation, because the statistical information of all users is known, wherein , , are known scalars, and the number of users near and the change rate of risk can be estimated as the gradient, and the vector is split into each customer group dimension. Taking the A customer group hierarchical parameter as an example, the iteration formula of can be obtained:
[0149]
[0150] wherein is a relatively small value, which is used to estimate the gradient in the case of continuous function. In the case of integer model division, it is further simplified. That is, the gradient is approximated as the difference.
[0151]
[0152] wherein, , , , and thus all have been converted into calculable values. The optimization of the boundary parameters of different customer groups and different hierarchies is the same.
[0153] can be combined with, for example Figure 5The flowchart shown, understand the gradient descent calculation process described above:
[0154] Receive the user input of each customer stratification expected value, and then calculate the stratification of each customer stratification lower bound parameters.
[0155] First calculate the stratification of the L layer lower bound parameters, initialize the current stratification of the L layer lower bound parameters, and solve the gradient of each direction of the L layer, then update the stratification of the customer stratification of the L layer, and then calculate the error corresponding to the comprehensive loss function, according to the error result to determine whether to return to execute the initialization of the current stratification of the L layer lower bound parameters, and this process is iterated k times, until the comprehensive loss function converges, or the number of iterations reaches the set iteration number. The set iteration number can be flexibly set according to actual experience, and the present application is not strictly limited.
[0156] There are two loops in the whole process:
[0157] Loop 1: Solve the stratification lower bound parameters layer by layer according to customer stratification.
[0158] Loop 2: Gradient descent calculation of stratification lower bound parameters of each layer.
[0159] Where, when initializing the current stratification lower bound parameters, the solution space of the model should be considered, for example, the stratification lower bound parameters of stratification 2 cannot exceed the stratification lower bound parameters of stratification 1. That is, the initialization of the current stratification line bundle needs to be randomly initialized within a certain range, and the specific range size can be adjusted by manual as a hyperparameter. The steps of solving the gradient and updating the stratification lower bound parameters are as described above. The gradient descent formula is used to solve the parameters of multiple customer groups in parallel. When the comprehensive loss function converges, that is, the error value of the comprehensive loss function reaches the set threshold, and the number of iterations reaches the set number, stop the loop.
[0160] As described above, the number of customer stratifications can be set by enterprises or organizations, but the specific way to determine the customer stratification that meets the number of customer stratifications needs to be implemented by executing the above steps S21 to S23. Specifically,
[0161] After determining the lower bound parameters for each customer segment, these parameters can be used to classify the subsequently input users to be identified. In the subsequent application process, step S11 is executed first to obtain the user characteristics of the users to be identified. These user characteristics can be user characteristics obtained in real-time using a user characteristic classification model, or they can be user characteristics that have been pre-identified and stored from a specified path. User characteristics are a set of information describing various attributes and behavioral characteristics of a user. As one implementation method, the user characteristics in this embodiment can be: user identity characteristics, user behavioral characteristics, user consumption characteristics, user interest characteristics, etc. For example, as shown in Figure 5, by pre-obtaining user identity characteristics, behavioral characteristics, consumption characteristics, interests, etc., accurate classification of the customer group and customer segment to which the user to be identified belongs can be achieved.
[0162] User identity characteristics typically refer to demographic features, including gender, age, region, education level, and occupation. User behavioral characteristics refer to specific actions taken while using online services, such as clicking, browsing, downloading, and watching. User consumption characteristics refer to the characteristics of a user's spending when using online services, such as purchase frequency, purchase price, and brand loyalty. User interest characteristics primarily represent psychological preferences, such as whether a user holds a positive or negative attitude towards a particular item or event; a positive attitude indicates a liking or preference.
[0163] Similarly, Figure 6 As shown, the pre-trained hierarchical model in this application can be any type of machine learning model, specifically a supervised learning model or an unsupervised learning model. The specific model structure and training process are not the focus of this application and will not be elaborated upon here. However, this pre-trained hierarchical model can learn the correlation between user features and the user's customer group and customer stratum based on the input user features, calculate the user's model score and customer group, and then further classify the current user into a specific stratum according to the pre-set hierarchical algorithm in the hierarchical model. This specific stratum is the user's customer stratum. As described above, customer stratification is flexibly set according to the needs of actual application scenarios. As an example, it can be as follows: Figure 6 As shown, customer tiers are specifically divided into: Loan Tier 1, Loan Tier 2, Loan Tier 3, etc.
[0164] Further, in step S12, after all user segmentation results have been obtained, statistical indicators corresponding to each customer segment are determined according to a preset statistical algorithm. The preset statistical algorithm varies depending on the different statistical indicators; for example, such as... Figure 6As shown, taking the loan one layer as an example, if the statistical index is the user level, the preset statistical algorithm for the user level can be summation, if the statistical index is the overall risk, the preset statistical algorithm for the overall risk can be to calculate the overdue probability of each user in the entire loan one layer, and then perform weighted summation. If the statistical index is the number of non-overdue borrowers, the statistical algorithm for the number of non-overdue borrowers is to sum the number of people who have no overdue and still have outstanding loans in the entire loan one layer. Further, according to the statistical index corresponding to the customer layer, it is judged whether the current customer layer meets the business requirements, if yes, the corresponding service strategy can be provided for the customer layer meeting the business requirements to serve, so as to realize targeted service for customers, guarantee customer experience, optimize service strategy and save service cost of the enterprise end.
[0165] In a second aspect, the application provides a target object grouping and layering processing device, wherein, as shown, Figure 7 The device comprises:
[0166] a layering processing module 701, which acquires user features of a user to be identified, inputs the user features into a pre-trained layering model, and determines a customer group and a customer layer to which the user belongs based on the input user features by using the layering model, wherein the customer layer is divided by different layering lower bound parameters;
[0167] a statistical module 702, which determines statistical indexes corresponding to each customer layer according to a preset statistical algorithm based on the customer layering result, wherein the types of the statistical indexes include a user level of a customer layer, a number of qualified users of a customer layer, and an overall risk of a customer layer;
[0168] a layering lower bound parameter calculation module 703, which is configured to determine the layering lower bound parameters according to the following steps in advance:
[0169] construct a first function relationship, a second function relationship, and a third function relationship based on the statistical indexes corresponding to each customer layer, wherein the first function relationship is a function relationship between the user level and the layering lower bound parameter, the second function relationship is a function relationship between the number of qualified users and the layering lower bound parameter, and the third function relationship is a function relationship between the overall risk and the layering lower bound parameter;
[0170] construct a user magnitude loss function, a target user loss function, and an overall risk loss function according to the expected number of users of each customer layer, the expected number of target users of each customer layer, and the expected overall risk of each customer layer, respectively; the user magnitude loss function is a loss function between a first function relationship and the expected number of users, the target user loss function is a loss function between a second function relationship and the expected number of target users, and the overall risk loss function is a loss function between a third function relationship and the expected overall risk;
[0171] construct a comprehensive loss function based on the user magnitude loss function, the target user loss function, and the overall risk loss function, and determine the layer lower bound parameters of each customer layer as target layer lower bound parameters when the comprehensive loss function converges by using a preset gradient descent algorithm.
[0172] In combination with the second aspect, in some possible embodiments, the layer lower bound parameter calculation module is specifically configured to:
[0173] construct a target parameter matrix based on the statistical indicators corresponding to each customer layer, each column vector in the target parameter matrix being a lower bound value of each customer layer corresponding to a same customer group, and each row vector in the target parameter matrix being a respective layer lower bound value of each customer group in a same customer layer;
[0174] construct the first function relationship, the second function relationship, and the third function relationship according to the target parameter matrix, each row vector being an independent variable, and the dimensions corresponding to each statistical indicator being a dependent variable.
[0175] In combination with the second aspect, in some possible embodiments, the user magnitude loss function, the target user loss function, and the overall risk loss function are respectively constructed according to the expected number of users of each customer layer, the expected number of target users of each customer layer, and the expected overall risk of each customer layer.
[0176] obtain the expected number of users, the expected number of target users, and the expected overall risk input by a user;
[0177] determine the user magnitude loss function as a mean square error loss function between the expected number of users and the first function relationship;
[0178] determine the target user loss function as a mean square error loss function between the expected number of target users and the second function relationship;
[0179] determine the overall risk loss function as a mean square error loss function between the expected overall risk and the third function relationship. , determine the overall risk loss function.
[0180] In combination with the second aspect, in some possible embodiments, the hierarchical lower bound parameter calculation module is specifically configured to:
[0181] The comprehensive loss function Loss is determined based on the following formula:
[0182]
[0183] For the comprehensive loss function, the gradients of the comprehensive loss function in the three direction angles are determined respectively ;
[0184] According to a preset iteration step α, the following parameter iteration formula is iterated:
[0185]
[0186] until the comprehensive loss function converges, or until the number of iterations reaches a preset number of iterations;
[0187] The hierarchical lower bound parameter when the comprehensive loss function converges or reaches the preset number of iterations is the target hierarchical lower bound parameter.
[0188] In combination with the second aspect, in some possible embodiments, the device further includes a data preprocessing module configured to:
[0189] execute a preset data aggregation algorithm from a database to obtain first sample data;
[0190] based on the first sample data, execute a preset prefix sum processing algorithm to calculate statistical indicators corresponding to each of the historical customer hierarchies;
[0191] based on the statistical indicators corresponding to each of the historical customer hierarchies, the first function relationship, the second function relationship, and the third function relationship are respectively constructed.
[0192] In combination with the second aspect, in some possible embodiments, the hierarchical lower bound parameter calculation module is further configured to:
[0193] According to the progressive relationship of the customer stratifications, the following steps are performed on each customer stratification layer by layer based on the expected number of users, the expected number of users meeting the standard, and the expected overall risk of each customer stratification:
[0194] The current stratification parameter corresponding to the current customer stratification is determined, and the current stratification parameter corresponding to the current customer stratification is calculated based on the comprehensive loss function. The gradient in the three direction angles;
[0195] Based on the gradient in the three direction angles corresponding to the current customer stratification, The parameter iteration formula is used for cyclic iteration.
[0196] The customer stratification parameter corresponding to the pre-set number of iterations is determined as the target stratification lower bound parameter.
[0197] In some possible embodiments, the user features include identity features, behavior features, consumption features, and interest and hobby features.
[0198] In the present application, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0199] The names of the messages or information exchanged between the multiple devices in the embodiments of the present application are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0200] In a third aspect, the exemplary embodiments of the present application also provide an electronic device, including at least one processor, and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to make the electronic device execute the method according to the embodiments of the present application when executed by the at least one processor.
[0201] The exemplary embodiments of the present application also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is used to make the computer execute the method according to the embodiments of the present application when executed by the processor of the computer.
[0202] The exemplary embodiments of the present application also provide a computer program product including a computer program, wherein the computer program is used to make the computer execute the method according to the embodiments of the present application when executed by the processor of the computer.
[0203] Reference Figure 8The present invention describes a structural block diagram of an electronic device 800 that can serve as a server or client of this application, which is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein.
[0204] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM 802) or a computer program loaded from a storage unit 808 into a random access memory (RAM 803). The RAM 803 may also store various programs and data required for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output interface (I / O interface 805) is also connected to the bus 804.
[0205] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. Input unit 806 can be any type of device capable of inputting information to electronic device 800. Input unit 806 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 808 may include, but is not limited to, disks and optical discs. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0206] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above. For example, in some embodiments, the aforementioned methods of clustering and stratifying target objects can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, portions or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the aforementioned methods of clustering and stratifying target objects by any suitable means, such as by means of firmware.
[0207] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, causes the functions / acts specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0208] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include, but are not limited to, an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0209] As used in this application, the terms "machine-readable medium" and "computer- readable medium" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0210] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0211] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0212] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
Claims
1. A method for grouping and hierarchically processing target objects, characterized in that, The method includes: The user characteristics of the user to be identified are obtained, and the user characteristics are input into a pre-trained hierarchical model. The hierarchical model determines the customer group and customer segment to which the user belongs based on the input user characteristics. Based on the customer segmentation results, statistical indicators corresponding to each customer segment are determined according to a preset statistical algorithm. The types of statistical indicators include: the number of users in the customer segment, the number of qualified users in the customer segment, and the overall risk of the customer segment. The customer segmentation is obtained by dividing the segmentation using different lower bound parameters, which are pre-determined according to the following steps: Based on the statistical indicators corresponding to each customer segment, a first functional relationship, a second functional relationship, and a third functional relationship are constructed respectively. The first functional relationship is the functional relationship between the user volume and the lower bound parameter of the segment, the second functional relationship is the functional relationship between the number of qualified users and the lower bound parameter of the segment, and the third functional relationship is the functional relationship between the overall risk and the lower bound parameter of the segment. Based on the expected number of users for each customer segment, the expected number of qualified users for each customer segment, and the expected overall risk for each customer segment, a user volume loss function, a qualified user loss function, and an overall risk loss function are constructed respectively; wherein, the user volume loss function is the loss function between the first functional relationship and the expected number of users, the qualified user loss function is the loss function between the second functional relationship and the expected number of qualified users, and the overall risk loss function is the loss function between the third functional relationship and the expected overall risk; A comprehensive loss function is constructed based on the user volume loss function, the qualified user loss function, and the overall risk loss function. A preset gradient descent algorithm is used to determine the lower bound parameter of each customer layer when the comprehensive loss function reaches convergence, which is the target lower bound parameter of the layer.
2. The method according to claim 1, characterized in that, The construction of a first functional relationship, a second functional relationship, and a third functional relationship based on the statistical indicators corresponding to each customer segment includes: Based on the statistical indicators corresponding to each customer segment, a target parameter matrix is constructed. Each column vector in the target parameter matrix represents the lower bound of each segment corresponding to each customer segment of the same customer group. Each row vector in the target parameter matrix represents the lower bound of each segment of the same customer group. Based on the target parameter matrix, the first functional relationship, the second functional relationship, and the third functional relationship are constructed respectively, using each row vector as the independent variable and the dimension corresponding to each statistical indicator as the dependent variable.
3. The method according to claim 1, characterized in that, The user volume loss function, the target user loss function, and the overall risk loss function are constructed based on the expected number of users in each customer segment, the expected number of users meeting the target in each customer segment, and the expected overall risk in each customer segment, respectively. include: Obtain the expected number of users, the expected number of users who meet the target, and the expected overall risk as provided by the user; The mean squared error loss function between the expected number of users and the first functional relationship , is determined to be the user-level loss function; The mean squared error loss function between the expected number of users and the second functional relationship. The loss function for the qualified users is determined to be... The mean squared error loss function between the expected overall risk and the third function relationship. This is determined to be the overall risk loss function.
4. The method according to claim 3, characterized in that, The construction of a comprehensive loss function based on the user volume loss function, the target user loss function, and the overall risk loss function includes: The comprehensive loss function Loss Determined based on the following formula: The method employs a preset gradient descent algorithm to determine the lower bound parameters of each customer stratum when the comprehensive loss function converges, which are then used as the target lower bound parameters. This includes: For the aforementioned comprehensive loss function, in Differentiate the functions in three directions to determine the comprehensive loss function at each angle. gradient in direction The comprehensive loss function is in gradient in direction And the comprehensive loss function in gradient in direction ; Based on the preset iteration step size α, the following parameter iteration formula is iterated: Until the comprehensive loss function converges, or until the number of iterations reaches the preset number of iterations; The comprehensive loss function converges, or the hierarchical lower bound parameter reaches the preset number of iterations. The lower bound parameter for the target layer.
5. The method according to claim 1, characterized in that, In the process of constructing the first functional relationship, the second functional relationship, and the third functional relationship based on the statistical indicators corresponding to each of the customer segments, the method further includes: Execute a preset data aggregation algorithm from the database to obtain the first sample data; Based on the first sample data, a preset prefix sum processing algorithm is executed to calculate the statistical indicators corresponding to each of the historical customer segments. Based on the statistical indicators corresponding to each of the historical customer segments, the first functional relationship, the second functional relationship, and the third functional relationship are constructed respectively.
6. The method according to claim 4, characterized in that, The method further includes: Based on the expected number of users, the expected number of users meeting the target, and the expected overall risk for each customer segment, the following steps are performed layer by layer for each customer segment according to the progressive relationship of the customer segments: Determine the current customer segment parameters corresponding to the current segment, and calculate the corresponding parameters for the current customer segment based on the comprehensive loss function. Gradient in three angular directions; Based on the current customer segmentation The gradients at the three directional angles are iterated cyclically using the aforementioned parameter iteration formula; The customer segmentation parameter corresponding to the preset number of iterations is determined as the target segmentation lower bound parameter.
7. The method according to claim 1, characterized in that, The user characteristics include: identity characteristics, behavioral characteristics, consumption characteristics, and interest and hobby characteristics.
8. An apparatus for grouping and hierarchically processing target objects, characterized in that, The device includes: A hierarchical processing module acquires user features of the user to be identified, inputs the user features into a pre-trained hierarchical model, and the hierarchical model determines the customer group and customer segment to which the user belongs based on the input user features, wherein the customer segment is obtained by dividing the segment using different lower bound parameters. The statistics module determines the statistical indicators corresponding to each customer segment based on the customer segmentation results and according to a preset statistical algorithm. The types of statistical indicators include: the user volume of the customer segment, the number of qualified users of the customer segment, and the overall risk of the customer segment. The hierarchical lower bound parameter calculation module is used to pre-determine the hierarchical lower bound parameter according to the following steps: Based on the statistical indicators corresponding to each customer segment, a first functional relationship, a second functional relationship, and a third functional relationship are constructed respectively. The first functional relationship is the functional relationship between the user volume and the lower bound parameter of the segment, the second functional relationship is the functional relationship between the number of qualified users and the lower bound parameter of the segment, and the third functional relationship is the functional relationship between the overall risk and the lower bound parameter of the segment. Based on the expected number of users for each customer segment, the expected number of qualified users for each customer segment, and the expected overall risk for each customer segment, a user volume loss function, a qualified user loss function, and an overall risk loss function are constructed respectively; wherein, the user volume loss function is the loss function between the first functional relationship and the expected number of users, the qualified user loss function is the loss function between the second functional relationship and the expected number of qualified users, and the overall risk loss function is the loss function between the third functional relationship and the expected overall risk; A comprehensive loss function is constructed based on the user volume loss function, the qualified user loss function, and the overall risk loss function. A preset gradient descent algorithm is used to determine the lower bound parameter of each customer layer when the comprehensive loss function reaches convergence, which is the target lower bound parameter of the layer.
9. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing a program; wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Data processing method and continuous risk control data discretization method and device
CN118260533A
Clustering model construction method based on causal inference and medical data processing method
WO2023050668A1