Risk user screening method and device
Through feature division and constraint optimization in genetic algorithms, the problems of accuracy and low efficiency in risk user screening in existing technologies are solved, and fast, accurate identification and balanced distribution of risk users are achieved.
Patent Information
- Application Number
- CN202410452238.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-10-21
AI Technical Summary
When screening risky users, existing technologies have difficulty in comprehensively considering the distribution inconsistency of various features and the unreasonable setting of weight values, resulting in low accuracy and efficiency of screening results.
A genetic algorithm is used to divide risk features into main features and auxiliary features, construct a fitness function and combine it with constraints of a specific value range to perform multiple iterative selection, crossover and mutation operations until the termination conditions are met, ensuring that the distribution of screened risky users is balanced among different categories.
It improves the accuracy and efficiency of risky user screening, avoids population evolution from falling into the Pareto set, and ensures the quality of screening results and compliance with business requirements.
Smart Images

Figure CN120822977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for screening risky users. Background Art
[0002] In actual scenarios, it is necessary to screen out a fixed number of risky users who meet the preset screening conditions from a large number of candidate users, for example, screening out 100 high-risk users from a user base of 10 million. In the prior art, screening is mainly performed through two methods: the first is to screen based on the numerical ranking of the feature values of the candidate users in multiple features, for example, selecting the 100 candidate users with the largest feature values of a certain feature as risky users. However, in specific scenarios, the feature value distribution of different users for each feature is often inconsistent. If the same user has a large feature value in one feature, the feature value in another feature is often small. Therefore, this method is difficult to obtain a comprehensive screening result that integrates all features. The second method is to set a corresponding weight value for each feature, calculate the weighted sum of the feature values of each candidate user as an evaluation score, and select the candidate user with the best evaluation score as a risky user. The disadvantage of this method is that there is a lack of a reasonable weight value formulation mechanism, and the results screened out with unreasonable weight values have low usability. Summary of the Invention
[0003] In view of this, an embodiment of the present invention provides a method and apparatus for screening risky users, which can accurately screen out a preset number of risky users from a large number of candidate users.
[0004] To achieve the above objective, according to one aspect of the present invention, a method for screening risky users is provided.
[0005] The risk user screening method of the embodiment of the present invention is used to select a preset number of risk users who meet preset screening conditions from multiple candidate users; the method includes: randomly selecting the preset number of candidate users from the multiple candidate users to form a sample, and taking each of the multiple samples formed as an individual; wherein each candidate user contained in each individual is taken as the gene of the individual, and each gene has a characteristic value under multiple preset risk characteristics; the multiple risk characteristics are arranged in descending order of importance, at least one risk characteristic in the front is determined as the main characteristic, and at least one risk characteristic other than the main characteristic is determined as an auxiliary characteristic; each auxiliary characteristic is determined The method comprises the following steps: constructing a fitness calculation function based on the main feature, determining the specific value range of each auxiliary feature as a constraint condition of the auxiliary feature, and combining the fitness calculation function and the constraint condition into a fitness function; performing a selection operation on the individuals according to the fitness function, performing a crossover operation on the selected multiple individuals to obtain new individuals, and performing a mutation operation on the new individuals to obtain the next generation of individuals; iteratively performing the selection operation, the crossover operation, and the mutation operation until a preset termination condition is met; determining a target individual from the individuals that meet the termination condition, and determining the candidate users included in the target individuals as the risky users.
[0006] Optionally, the selection operation is performed on the individuals according to the fitness function, including: for any individual, determining whether the characteristic value of each gene in the individual in each auxiliary feature meets the constraint condition of the auxiliary feature in the fitness function; if so, using the fitness calculation function to determine the fitness score of the individual; otherwise, setting the fitness score of the individual to a preset fitness minimum value; and performing selection based on the fitness score of each individual.
[0007] Optionally, performing a crossover operation on the selected multiple individuals to obtain new individuals includes: determining whether the same gene exists in each new individual obtained by the crossover operation; if so, removing the new individual.
[0008] Optionally, performing a mutation operation on the new individual to obtain the next generation of individuals includes: for each new individual on which the mutation operation will be performed and the gene to be mutated in the new individual, randomly selecting a gene from a mutation range set pre-established for the gene to be mutated to replace the gene to be mutated in the new individual; judging whether the same gene exists in the individual formed by the replacement; if so, removing the individual, and randomly selecting a gene from the mutation range set again to perform the replacement until an individual without the same gene is formed; wherein the multiple genes contained in the mutation range set belong to the same category as the gene to be mutated according to the specific dimension of the candidate user.
[0009] Optionally, determining the target individual from the individuals satisfying the termination condition includes: determining the fitness score of each individual satisfying the termination condition according to the fitness function, and determining the individual with the largest fitness score as the target individual.
[0010] Optionally, the number of the main features is one.
[0011] To achieve the above object, according to another aspect of the present invention, a device for screening risky users is provided.
[0012] The risk user screening device of the embodiment of the present invention is used to select a preset number of risk users who meet preset screening conditions from multiple candidate users; the device includes: a pre-processing unit, used to randomly select the preset number of candidate users from the multiple candidate users to form a sample, and each of the multiple samples formed is used as an individual; wherein each candidate user contained in each individual is used as the gene of the individual, and each gene has a characteristic value under multiple preset risk characteristics; a feature processing unit, used to arrange the multiple risk characteristics in descending order of importance, determine at least one risk characteristic in front as the main characteristic, and determine at least one risk characteristic other than the main characteristic as the auxiliary characteristic; determine the specific characteristics of each auxiliary characteristic value range; a fitness function construction unit, used to construct a fitness calculation function based on the main feature, determine the specific value range of each auxiliary feature as the constraint condition of the auxiliary feature, and combine the fitness calculation function and the constraint condition into a fitness function; an iterative operation unit, used to perform a selection operation on the individuals according to the fitness function, perform a crossover operation on the selected multiple individuals to obtain new individuals, and perform a mutation operation on the new individuals to obtain the next generation of individuals; iteratively perform the selection operation, the crossover operation and the mutation operation until a preset termination condition is met; a screening unit, used to determine a target individual from the individuals that meet the termination condition, and determine the candidate users contained in the target individuals as the risk users.
[0013] Optionally, the iterative operation unit is further used to: for any individual, determine whether the characteristic value of each gene in the individual in each auxiliary feature meets the constraint condition of the auxiliary feature in the fitness function; if so, use the fitness calculation function to determine the fitness score of the individual; otherwise, set the fitness score of the individual to a preset fitness minimum value; and perform selection based on the fitness score of each individual.
[0014] To achieve the above objective, according to another aspect of the present invention, an electronic device is provided.
[0015] An electronic device of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the risk user screening method provided by the present invention.
[0016] To achieve the above objective, according to another aspect of the present invention, a computer-readable storage medium is provided.
[0017] A computer-readable storage medium of the present invention stores a computer program, which, when executed by a processor, implements the risk user screening method provided by the present invention.
[0018] According to the technical solution of the present invention, the embodiments of the above invention have the following advantages or beneficial effects:
[0019] First, a preset number of candidate users are randomly selected from a pool of candidate users to form a sample. These samples are then treated as individuals, with each candidate user contained in each individual serving as its gene. Each gene has eigenvalues for multiple pre-set risk characteristics. Subsequently, multiple iterations of selection, crossover, and mutation are performed on these individuals until a termination condition is reached. Finally, a target individual is identified from among the individuals that meet the termination condition, and the candidate users contained in the target individual are identified as risky users. This transforms the risky user screening problem into a solvable form for a genetic algorithm, enabling rapid and accurate identification of risky users. Furthermore, during the genetic algorithm's selection operation, the risk characteristics of the gene are divided into primary and auxiliary features based on their importance. The auxiliary features are used only to form constraints in the fitness function through their specific value ranges. The actual fitness calculation function is constructed based on a small number of primary features. This approach prevents population evolution from falling into a Pareto set and helps improve the efficiency of risky user search. In addition, in the mutation operation of the genetic algorithm, a mutation range set for each gene is pre-configured. The set consists of genes of the same category divided according to the specific dimensions of the candidate user. Other genes in the set are randomly selected as mutation optional genes for the current gene, so that after the mutated gene replaces the current gene, it will not change the gene distribution of different categories in the population, thereby ensuring that the final risk users are distributed evenly among the categories.
[0020] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.
[0022] Figure 1This is a schematic diagram of the main steps of the risk user screening method according to an embodiment of the present invention;
[0023] Figure 2 1 is a schematic diagram of a genetic algorithm flow chart according to an embodiment of the present invention;
[0024] Figure 3 is a schematic diagram of the system architecture of the first embodiment of the present invention;
[0025] Figure 4 Schematic diagram of components of a risky user screening device according to an embodiment of the present invention;
[0026] Figure 5 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;
[0027] Figure 6 This is a schematic diagram of the structure of an electronic device used to implement the risky user screening method in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0029] It should be pointed out that, in the absence of conflict, the embodiments of the present invention and the technical features therein may be combined with each other.
[0030] Figure 1 2 is a schematic diagram of the main steps of the risk user screening method according to an embodiment of the present invention.
[0031] like Figure 1 As shown, the risk user screening method according to the embodiment of the present invention can be specifically performed according to the following steps:
[0032] Step S101: randomly selecting a preset number of candidate users from a plurality of candidate users to form a sample, and treating each of the plurality of samples as an individual.
[0033] The present invention is used to select a preset number of risk users who meet preset screening conditions from multiple candidate users, and can be executed by a server that executes a corresponding program. The above candidate users can be any entity or data user, and the risk users are candidate users that meet the screening conditions screened out from the initial candidate users, and the above screening conditions can be flexibly formulated according to the business scenario. The above candidate users have characteristic values under multiple preset risk characteristics. For example, in a risk user identification scenario related to claims, each user can have characteristic values of five risk characteristics, such as the total amount of claims, the total number of claims (i.e., the total number of orders involving claims), the payout ratio (the proportion of the number of orders involving claims in the order weight), the accident rate (the proportion of the number of orders involving accidents in the order weight), and the total amount of claims in the past month. These characteristic values can be continuous or discrete within a certain numerical range.
[0034] The above candidate users can be formed into different categories based on specific dimensions. Taking the specific dimension as geographical area as an example, all candidate users can be divided into multiple categories according to their geographical areas, such as candidate users in area 1, candidate users in area 2, candidate users in area 3, etc.
[0035] In the prior art, when selecting a fixed number (i.e., a preset number) of risky users from multiple candidate users, the lack of a reasonable mechanism for setting feature weights makes it impossible to calculate the weighted sum of each candidate user's multiple feature values as an evaluation score for screening. Furthermore, due to the inconsistency of the distribution of feature values across users, it is also impossible to identify risky users by numerically ranking the feature values of each feature. To address this issue, the present invention uses a genetic algorithm to identify a fixed number of risky users. First, the server needs to execute step S101 to model the risky user screening problem in a form that can be solved by a genetic algorithm.
[0036] In this step, the server randomly selects a preset number of candidate users from the above multiple candidate users to form a sample, and repeats this process to form multiple samples. It can be understood that the maximum number of samples is Here, N represents the total number of candidate users, and n represents the preset number. The server can then identify some or all of the samples as individuals, i.e., individuals in the genetic algorithm population. Each individual contains a preset number of candidate users, each of which is a gene for that individual. Obviously, each gene has characteristic values for the candidate user under the aforementioned multiple risk characteristics. The genetic algorithm process can then be executed on these individuals.
[0037] Step S102: Arrange the multiple risk features in descending order of importance, determine at least one risk feature in the front as a primary feature, and determine at least one risk feature other than the primary feature as an auxiliary feature; and determine a specific value range for each auxiliary feature.
[0038] The genetic algorithm process used in the present invention can be used to Figure 2 It can be obtained by improving the classical genetic algorithm process shown in the figure, or by improving the known differential evolution algorithm and parallel genetic algorithm. Figure 2 process as an example. Figure 2 In the process of generating the initial population, the fitness function is used to calculate the fitness scores of the individuals in the population. After that, selection, crossover, and mutation operations are performed to generate the next generation of population. In the selection operation, the corresponding selection probability can be determined based on the fitness scores of different individuals, and then individual selection is performed; the selected individuals are randomly paired to perform the crossover operation. During the crossover operation, each individual randomly generates a crossover point to perform gene exchange and generate new individuals; the new individuals undergo gene mutation with a certain probability, forming the next generation of population and individuals. After that, it is determined whether the termination condition is met. The termination condition can be whether the fitness score of the current individual reaches a threshold or whether the number of iterations (i.e., evolutionary generations) reaches a threshold. If the termination condition is not met, the process returns to the step of calculating the fitness and iteratively performs the above steps. If the termination condition is met, the search result is output.
[0039] In the technical solution of the present invention, the server can use a preset fitness function to perform a selection operation on individuals, perform a crossover operation on the selected multiple individuals to obtain new individuals, perform a mutation operation on the new individuals to obtain the next generation of individuals, and then iteratively perform the above selection operation, crossover operation and mutation operation until the preset termination condition is met.
[0040] In this step, the server can construct a fitness function based on the full risk characteristics of each gene (such as the five risk characteristics mentioned above) to directly calculate the fitness score of each individual. Preferably, before calculating the fitness score, the server can first normalize the eigenvalues of the same risk characteristic within the characteristic to avoid the influence of data scale differences on the calculation results. The above fitness function can be: for a certain individual, first add up all the eigenvalues of each gene to obtain the cumulative eigenvalues of each gene, and then add up the cumulative eigenvalues of each gene to obtain the fitness score of the individual. However, this method has the following defects, that is, the population evolution is prone to fall into the Pareto set, that is, the individuals selected by this method tend to have eigenvalues of each characteristic that are not prominent enough, but can meet the current conditions (that is, have a larger fitness score). This situation is unacceptable for practical applications, so it is necessary to optimize the fitness function and its search strategy.
[0041] As a preferred solution, the embodiment of the present invention performs the following optimization on the fitness function and the fitness score calculation strategy. First, the fitness function is set to include two parts: a fitness calculation function and at least one constraint condition. When calculating the fitness score of an individual, it is first determined whether the characteristic value of each gene in the individual meets the constraint condition. If it meets the constraint condition, the fitness calculation function is used to calculate its fitness score. If it does not meet the constraint condition, the fitness score is directly set to the fitness minimum value (such as zero or a value close to zero). The above-mentioned fitness calculation function and constraint condition are formulated as follows.
[0042] Preferably, the server first sorts the multiple risk features in descending order of importance. The importance of each risk feature can be quantitatively calculated according to a preset strategy. Next, the server identifies at least one of the aforementioned features as a primary feature and at least one feature other than the primary feature as an auxiliary feature. The server can also determine a specific value range for each auxiliary feature. The specific value range of an auxiliary feature represents the possible values of that auxiliary feature for a risky user. If any feature value of any gene in an individual is not within the specific value range of the corresponding feature, then that gene and that individual are not included in the final search results.
[0043] Step S103: constructing a fitness calculation function based on the main feature, determining a specific value range of each auxiliary feature as a constraint condition of the auxiliary feature, and combining the fitness calculation function and the constraint condition into a fitness function.
[0044] In this step, the server constructs a fitness calculation function based on the above main features, and determines the specific value range of each auxiliary feature as the constraint condition of the auxiliary feature, thereby forming a complete fitness function.
[0045] For example, among the five risk features mentioned above, namely, the total claim amount, total number of claims, payout ratio, accident rate, and total claim amount in the past month, the most important total claim amount can be determined as the main feature, and the total number of claims, payout ratio, accident rate, and total claim amount in the past month can be determined as auxiliary features. The specific value range of each auxiliary feature can also be determined, such as the specific value range of the total number of claims is (1000, +∞), and the specific value range of the payout ratio is (0.3, 1], etc. In the fitness function finally constructed, the fitness calculation function is: for an individual, the characteristic value of each gene in the total claim amount is added to obtain the individual's fitness score (the aforementioned characteristic value normalization can be performed in advance), and the constraint conditions are the specific value ranges of the four features, namely, the total number of claims, payout ratio, accident rate, and total claim amount in the past month.
[0046] Step S104: performing a selection operation on the individuals according to the fitness function, performing a crossover operation on the selected individuals to obtain new individuals, performing a mutation operation on the new individuals to obtain the next generation of individuals; iteratively performing the selection operation, the crossover operation and the mutation operation until a preset termination condition is met.
[0047] In an embodiment of the present invention, in the fitness score calculation based on the fitness function, for any individual, the server determines whether the characteristic value of each gene in each auxiliary feature within the individual satisfies the constraint conditions of the auxiliary feature in the fitness function; if so, the fitness score of the individual is determined using the above fitness calculation function; otherwise (i.e., at least one characteristic value of at least one gene in the individual does not satisfy the corresponding constraint conditions, i.e., the characteristic value is not within the specific value range of the corresponding feature), the fitness score of the individual is set to a preset fitness minimum value. Thereafter, the server can perform a selection operation based on the fitness score of each individual, and the execution method of the selection operation can be the same as the selection process of the known genetic algorithm.
[0048] Through the above steps, the number of genetic features used in the fitness function can be reduced to avoid the population evolution from falling into the Pareto set, thereby improving the search efficiency and accuracy of risk users. At the same time, some relatively low-importance features are used as auxiliary features to form pre-constraints, which can also reflect the screening and evaluation effects of these features on individuals in the fitness calculation process. In addition, the use of a specific value range of the auxiliary features is conducive to the formulation of screening rules in practical applications. It can be understood that the number of main features can be any number less than the initial number of features. Generally speaking, the smaller the number of main features, the better the effect of avoiding falling into the Pareto set. Therefore, as a better solution, the most important feature can be used as the main feature to construct the fitness calculation function, and other features can be used as auxiliary features to form constraints.
[0049] In specific applications, the multiple risky users ultimately found should all be different risky users, so the individuals in the genetic algorithm should not contain identical genes. Therefore, during the crossover operation in this embodiment of the present invention, the server can determine whether identical genes exist in each new individual obtained through the crossover operation. If so, the new individual is removed; if not, subsequent mutation operations are performed on the new individual. The remaining steps of the crossover operation can be the same as those of conventional genetic algorithms.
[0050] In the mutation operation of the embodiment of the present invention, the server first determines at least one new individual on which the mutation operation will be performed and the gene to be mutated in the new individual based on a preset mutation probability. Thereafter, a gene can be randomly selected from all genes except the gene to be mutated (i.e., all candidate users except the gene to be mutated) to replace the above gene to be mutated in the new individual to form the next generation of individuals.
[0051] The above mutation method has a flaw. As mentioned above, candidate users can form different categories based on specific dimensions, such as categories based on different geographical regions. During the individual selection and crossover process, due to the use of random calculation methods, the individual genes formed have an overall balance between different categories. However, the above mutation operation replaces a gene of a determined category with a gene of a random category, which destroys the original balance of gene distribution between categories. As a result, the risk users finally searched may not have a balance between categories, such as a balance between geographical regions. For example, the risk users searched are concentrated in Region 1 and Region 2, and there are no risk users in other regions, which cannot meet the requirements of actual applications.
[0052] In response to the above defects, the embodiment of the present invention performs the following optimization processing in the mutation operation. First, a variation range set is determined for each gene in advance. In practical applications, genes of the same category (i.e., candidate users) can be placed in the same set. Then, for any gene therein, the other genes in the set except the gene form its variation range set. In the mutation operation, for each new individual on which the mutation operation will be performed and the gene to be mutated in the new individual, the server randomly selects a gene from the variation range set pre-established for the gene to be mutated to replace the gene to be mutated in the new individual. Thereafter, the server determines whether the same gene exists in the individual formed by the above replacement; if so, the individual is removed, and a gene is randomly selected from the above variation range set to perform the above replacement again until an individual without the same gene is formed. Obviously, the multiple genes contained in the above variation range set and the gene to be mutated belong to the same category divided according to the specific dimension of the candidate user.
[0053] In this way, it can be ensured that the genes of the same individual before and after mutation belong to the same category, thereby maintaining the balanced distribution of the overall genes in different categories, and then making the risk users finally searched have a balanced distribution of different categories to meet actual application needs.
[0054] After the mutation operation, the next generation of individuals can be obtained. The server can iteratively perform the above selection operation, crossover operation and mutation operation until the preset termination condition is met. The termination condition can be set according to the actual scenario, for example, it can be set to whether the fitness score of the current individual reaches a threshold or the number of iterations reaches a threshold.
[0055] Step S105: determining a target individual from the individuals that meet the termination condition, and determining the candidate users included in the target individual as risky users.
[0056] In this step, the server can determine the fitness score of each individual that meets the termination condition according to the above fitness function, and determine the individual with the largest fitness score as the target individual. Finally, the candidate users included in the target individual are determined as risky users.
[0057] After the above steps, the risk user screening problem can be converted into a solvable form of genetic algorithm, and the rapid and accurate identification of risk users can be achieved based on the genetic algorithm. In addition, the population evolution will not fall into the Pareto set, the algorithm search efficiency is high, and the distribution of population genes in different categories will not change before and after the mutation, thereby ensuring that the final risk users are distributed evenly among the categories.
[0058] A first embodiment of the present invention will be described below.
[0059] In the express delivery business scenario, if the express delivery is damaged or lost, the express delivery service provider needs to process the claim for the user. However, in actual applications, some users will take advantage of the claim rules to maliciously defraud fees. Therefore, the service provider needs to screen out a preset number of high-risk users from a large number of users. For example, in a certain scenario, 100 high-risk users need to be screened out from 10 million users. This embodiment can search for the best 100 risky users based on the improved genetic algorithm taking into account multiple evaluation indicators (i.e., the aforementioned characteristics), so that the output results meet multiple business expectations at the same time. The system architecture of this embodiment is as follows: Figure 3 shown.
[0060] This embodiment is divided into three parts: data processing, strategy model and business system.
[0061] During the data processing phase, the service provider's server extracts historical claims data and related online data from the first database's users (i.e., candidate users) into a user candidate pool. This pool contains characteristic values for five risk characteristics: total claim amount, total number of claims, loss ratio, accident rate, and total claim amount in the past month. These users can be divided into multiple categories based on their geographic region. The server then randomly selects 100 users from the user candidate pool to form a sample. These samples are used as population individuals to perform model calculations based on the genetic algorithm.
[0062] In the strategy model stage, the server uses a genetic algorithm-based model to Figure 2The calculation process is shown. In actual applications, it was found that during the selection process, if the fitness function is constructed using five risk characteristics to directly calculate the fitness score of each individual, the population evolution will fall into the Pareto set, that is, the characteristic values are not prominent enough, but the fitness score is high. Therefore, the server sorts the characteristics, selecting the highly important total claim amount as the primary characteristic, and the other characteristics as auxiliary characteristics. The fitness calculation function is constructed based on the primary characteristics, and the specific value ranges of the auxiliary characteristics form the constraints. Because genes within the same individual must be non-reproducible, individuals with duplicate genes are removed during the crossover and mutation processes.
[0063] During the mutation process, a neighborhood search strategy needs to be implemented to meet the geographical balance of different genes. Specifically, when the server is mutating a certain gene, it can only select genes from the preset mutation range set of the gene. The genes in the above mutation range set correspond to the same geographical region as the gene to be mutated, thereby ensuring the geographical distribution balance of the gene, realizing the gene-directed mutation function, ensuring the overall controllability of the evolution direction, and ultimately improving the quality of search results and business interpretability. Other execution processes can be the same as existing genetic algorithm processes or any evolution templates such as awga, moea, nsga, pps, rvea, etc. The genetic algorithm model can periodically perform model updates based on updated data.
[0064] During the business system phase, the server can input the model calculation results into the second database and display them on the front end. The output result is the optimal combination of 100 risky users. Comparing the output results of the embodiment of the present invention with those of the prior art, it can be found that the present invention can significantly improve the accuracy of risky user screening.
[0065] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution of the present invention all comply with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken with respect to user personal information to prevent unauthorized access to user personal information data and maintain the security of user personal information, network security, and national security.
[0066] For ease of description, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should be aware that the present invention is not limited to the order of the actions described, and certain steps can actually be performed in other orders or simultaneously. In addition, those skilled in the art should also be aware that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily required to implement the present invention.
[0067] In order to better implement the above solutions of the embodiments of the present invention, relevant devices for implementing the above solutions are also provided below.
[0068] See also Figure 4 As shown, the risk user screening device 400 provided in an embodiment of the present invention is used to select a preset number of risk users who meet preset screening conditions from multiple candidate users; the device 400 includes: a preprocessing unit 401, a feature processing unit 402, a fitness function construction unit 403, an iterative operation unit 404 and a screening unit 405.
[0069] Among them, the pre-processing unit 401 is used to randomly select the preset number of candidate users from the multiple candidate users to form a sample, and each sample in the formed multiple samples is used as an individual; wherein each candidate user contained in each individual is used as the gene of the individual, and each gene has a characteristic value under multiple preset risk characteristics; the feature processing unit 402 is used to arrange the multiple risk characteristics in descending order of importance, determine at least one risk characteristic in front as the main characteristic, and determine at least one risk characteristic other than the main characteristic as the auxiliary characteristic; determine the specific value range of each auxiliary characteristic; the fitness function construction unit 403 is used to construct a fitness function based on the main characteristic A fitness calculation function is established, a specific value range of each auxiliary feature is determined as a constraint condition of the auxiliary feature, and the fitness calculation function and the constraint condition are combined into a fitness function; an iterative operation unit 404 is used to perform a selection operation on the individuals according to the fitness function, perform a crossover operation on the selected multiple individuals to obtain new individuals, and perform a mutation operation on the new individuals to obtain the next generation of individuals; the selection operation, the crossover operation and the mutation operation are iteratively performed until a preset termination condition is met; a screening unit 405 is used to determine a target individual from the individuals that meet the termination condition, and determine the candidate users included in the target individual as the risky users.
[0070] In an embodiment of the present invention, the iterative operation unit 404 can be further used to: for any individual, determine whether the characteristic value of each gene in each auxiliary feature in the individual meets the constraint condition of the auxiliary feature in the fitness function; if so, use the fitness calculation function to determine the fitness score of the individual; otherwise, set the fitness score of the individual to a preset fitness minimum value; and perform selection based on the fitness score of each individual.
[0071] In a specific application, the iterative operation unit 404 can be further used to determine whether the same gene exists in each new individual obtained by the crossover operation; if so, remove the new individual.
[0072] In practical applications, the iterative operation unit 404 can be further used to: for each new individual to be mutated and the gene to be mutated in the new individual, randomly select a gene from a mutation range set pre-established for the gene to be mutated to replace the gene to be mutated in the new individual; determine whether the same gene exists in the individual formed by the replacement; if so, remove the individual and randomly select a gene from the mutation range set to perform the replacement again until an individual without the same gene is formed; wherein, the multiple genes contained in the mutation range set belong to the same category as the gene to be mutated according to the specific dimension of the candidate user.
[0073] As a preferred solution, the screening unit 405 may be further configured to: determine the fitness score of each individual that meets the termination condition according to the fitness function, and determine the individual with the largest fitness score as the target individual.
[0074] In addition, in the embodiment of the present invention, the number of the main feature is one.
[0075] According to the technical solution of the embodiment of the present invention, the risk user screening problem is converted into a solvable form of a genetic algorithm, and the rapid and accurate identification of risk users is achieved based on the genetic algorithm. In addition, in the selection operation of the genetic algorithm, the risk characteristics of the gene are divided into main characteristics and auxiliary characteristics according to their importance. The auxiliary characteristics are only used to form the constraint conditions in the fitness function through their specific value ranges. The fitness calculation function that actually performs the calculation is constructed based on a small number of main characteristics. This method can prevent the population evolution from falling into the Pareto set, which helps to improve the search efficiency of risk users. In addition, in the mutation operation of the genetic algorithm, a mutation range set for each gene is pre-configured. The set is composed of genes of the same category divided according to the specific dimensions of the candidate user. The other genes in the set are randomly selected as the mutation optional genes of the current gene, so that after the mutation gene replaces the current gene, the gene distribution of different categories in the population will not be changed, thereby ensuring the final distribution balance of risk users among the categories.
[0076] Figure 5 An exemplary system architecture 500 is shown to which the risky user screening method or risky user screening apparatus according to an embodiment of the present invention may be applied.
[0077] like Figure 5As shown, system architecture 500 may include terminal devices 501, 502, and 503, a network 504, and a server 505 (this architecture is merely an example, and the components included in the specific architecture may be adjusted based on the specific application). Network 504 is used to provide a medium for communication links between terminal devices 501, 502, and 503 and server 505. Network 504 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0078] Users can use terminal devices 501, 502, 503 to interact with server 505 via network 504 to receive or send messages, etc. Various client applications can be installed on terminal devices 501, 502, 503, such as risky user identification applications (only as an example).
[0079] The terminal devices 501 , 502 , and 503 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0080] Server 505 may be a server that provides various services, such as a backend server (for example only) that supports risky user identification applications operated by users using terminal devices 501, 502, and 503. The backend server may process received risky user identification requests and provide feedback on the processing results (for example, the risky users identified—for example only) to terminal devices 501, 502, and 503.
[0081] It should be noted that the risky user screening method provided in the embodiment of the present invention is generally executed by the server 505 , and accordingly, the risky user screening device is generally provided in the server 505 .
[0082] It should be understood that Figure 5 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0083] The present invention also provides an electronic device. The electronic device in an embodiment of the present invention includes: one or more processors; and a storage device configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the risky user screening method provided by the present invention.
[0084] Reference below Figure 6 , which shows a schematic structural diagram of a computer system 600 of an electronic device suitable for implementing an embodiment of the present invention. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0085] like Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the computer system 600 are also stored in the RAM 603. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0086] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, and the like; an output section 607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or a modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 610 as needed, so that computer programs read therefrom can be installed in the storage section 608 as needed.
[0087] In particular, according to embodiments disclosed herein, the processes described in the main step diagrams above can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the main step diagrams. In the above embodiments, the computer program can be downloaded and installed from a network via the communication section 609 and / or installed from a removable medium 611. When the computer program is executed by the central processing unit 601, the above-described functions defined in the system of the present invention are performed.
[0088] It should be noted that the computer-readable medium described in the present invention may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.
[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0090] The units involved in the embodiments of the present invention may be implemented in software or hardware. The units described may also be provided in a processor. For example, they may be described as: a processor including a preprocessing unit, a feature processing unit, a fitness function construction unit, an iterative operation unit, and a screening unit. The names of these units do not, in some cases, limit the units themselves. For example, the preprocessing unit may also be described as a "unit that provides population individuals and genes to the iterative operation unit."
[0091] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the steps executed by the device include: randomly selecting a preset number of candidate users from multiple candidate users to form a sample, and treating each of the multiple samples formed as an individual; wherein each candidate user contained in each individual serves as the gene of the individual, and each gene has characteristic values under multiple preset risk characteristics; arranging the multiple risk characteristics in descending order of importance, determining at least one risk characteristic in front as the main characteristic, and determining at least one risk characteristic other than the main characteristic as an auxiliary characteristic; determining each auxiliary characteristic. A specific value range; constructing a fitness calculation function based on the main feature, determining the specific value range of each auxiliary feature as a constraint condition of the auxiliary feature, and combining the fitness calculation function and the constraint condition into a fitness function; performing a selection operation on the individuals according to the fitness function, performing a crossover operation on the selected multiple individuals to obtain new individuals, and performing a mutation operation on the new individuals to obtain the next generation of individuals; iteratively performing the selection operation, the crossover operation, and the mutation operation until a preset termination condition is met; determining a target individual from the individuals that meet the termination condition, and determining the candidate users included in the target individuals as the risk users.
[0092] In the technical solution of the embodiment of the present invention, the risk user screening problem is converted into a solvable form of a genetic algorithm, and the rapid and accurate identification of risk users is achieved based on the genetic algorithm. In addition, in the selection operation of the genetic algorithm, the risk characteristics of the gene are divided into main characteristics and auxiliary characteristics according to their importance. The auxiliary characteristics are only used to form the constraint conditions in the fitness function through their specific value ranges. The fitness calculation function that actually performs the calculation is constructed based on a small number of main characteristics. This method can prevent the population evolution from falling into the Pareto set, which helps to improve the search efficiency of risk users. In addition, in the mutation operation of the genetic algorithm, a mutation range set for each gene is pre-configured. The set is composed of genes of the same category divided according to the specific dimensions of the candidate user. The other genes in the set are randomly selected as the mutation optional genes of the current gene, so that after the mutation gene replaces the current gene, the gene distribution of different categories in the population will not be changed, thereby ensuring the final distribution balance of risk users among the categories.
[0093] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for screening risky users, characterized in that: The method is used to select a preset number of risky users who meet preset screening conditions from a plurality of candidate users; the method includes: Randomly selecting the preset number of candidate users from the plurality of candidate users to form a sample, and treating each of the plurality of samples as an individual; wherein each candidate user included in each individual is used as a gene of the individual, and each gene has characteristic values under a plurality of preset risk characteristics; Arrange the multiple risk features in descending order of importance, determine at least one risk feature at the top as a primary feature, and determine at least one risk feature other than the primary feature as an auxiliary feature; and determine a specific value range for each auxiliary feature; Constructing a fitness calculation function based on the main feature, determining a specific value range of each auxiliary feature as a constraint condition of the auxiliary feature, and combining the fitness calculation function and the constraint condition into a fitness function; Performing a selection operation on the individuals according to the fitness function, performing a crossover operation on the selected individuals to obtain new individuals, and performing a mutation operation on the new individuals to obtain next generation individuals; iteratively performing the selection operation, the crossover operation, and the mutation operation until a preset termination condition is met; A target individual is determined from the individuals that meet the termination condition, and candidate users included in the target individual are determined as the risky users.
2. The method according to claim 1, characterized in that The performing the selection operation on the individual according to the fitness function includes: For any individual, determine whether the characteristic value of each gene in each auxiliary feature of the individual satisfies the constraint condition of the auxiliary feature in the fitness function; If satisfied, the fitness calculation function is used to determine the fitness score of the individual; otherwise, the fitness score of the individual is set to a preset fitness minimum value; Selection is performed based on the fitness score of each individual.
3. The method according to claim 1, characterized in that The performing of a crossover operation on the selected multiple individuals to obtain a new individual includes: Determine whether the same gene exists in each new individual obtained by the crossover operation; if so, remove the new individual.
4. The method according to claim 1, wherein The performing of a mutation operation on the new individual to obtain the next generation individual includes: For each new individual to be mutated and the gene to be mutated in the new individual, randomly select a gene from a mutation range set pre-established for the gene to be mutated to replace the gene to be mutated in the new individual; Determine whether there are identical genes in the individuals formed by the replacement; if so, remove the individual and randomly select a gene from the variation range set to perform the replacement again until an individual without identical genes is formed; wherein, The multiple genes contained in the mutation range set and the genes to be mutated belong to the same category divided according to the specific dimension of the candidate user.
5. The method according to claim 1, wherein The determining of the target individual from the individuals satisfying the termination condition comprises: The fitness score of each individual that meets the termination condition is determined according to the fitness function, and the individual with the largest fitness score is determined as the target individual.
6. The method according to claim 1, characterized in that The number of the main features is one.
7. A risky user screening device, characterized in that: The device is used to select a preset number of risky users who meet preset screening conditions from a plurality of candidate users; the device comprises: a preprocessing unit, configured to randomly select a preset number of candidate users from the plurality of candidate users to form a sample, and treat each of the plurality of samples as an individual; wherein each candidate user included in each individual serves as a gene of the individual, and each gene has a characteristic value under a plurality of preset risk characteristics; a feature processing unit, configured to arrange the plurality of risk features in descending order of importance, determine at least one risk feature in the front as a primary feature, determine at least one risk feature other than the primary feature as an auxiliary feature, and determine a specific value range for each auxiliary feature; a fitness function construction unit, configured to construct a fitness calculation function based on the main feature, determine a specific value range of each auxiliary feature as a constraint condition of the auxiliary feature, and combine the fitness calculation function and the constraint condition into a fitness function; an iterative operation unit, configured to perform a selection operation on the individuals according to the fitness function, perform a crossover operation on the selected individuals to obtain new individuals, and perform a mutation operation on the new individuals to obtain next-generation individuals; and iteratively perform the selection operation, the crossover operation, and the mutation operation until a preset termination condition is satisfied; The screening unit is configured to determine target individuals from individuals that meet the termination condition, and determine candidate users included in the target individuals as risky users.
8. The device according to claim 7, characterized in that The iterative operation unit is further configured to: For any individual, determine whether the characteristic value of each gene in each auxiliary feature in the individual meets the constraint conditions of the auxiliary feature in the fitness function; if so, use the fitness calculation function to determine the fitness score of the individual; otherwise, set the fitness score of the individual to a preset fitness minimum value; and perform selection based on the fitness score of each individual.
9. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.