Similar object positioning method, device, equipment and medium based on data characteristics

By calculating the characteristic differences between seed sets and control sets, determining the target characteristics, and judging the similarity of the full amount of objects, the problem of inefficiency in financial institutions in money laundering risk analysis is solved, and the efficiency and accuracy of screening risk groups is improved.

CN114548196BActive Publication Date: 2025-09-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011357612.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-26
Publication Date
2025-09-02
Estimated Expiration
2040-11-26

AI Technical Summary

Technical Problem

In the prior art, when financial institutions conduct money laundering risk analysis, it is less efficient to determine the risk population based on data characteristics, requires multiple job cooperation and takes a long time.

Method used

By obtaining feature statistics of seed sets and control sets, calculating the degree of difference, determining the target characteristics, and judging the similarity of the entire object based on the target characteristic value, reducing the number of feature comparisons, and improving efficiency.

Benefits of technology

The similarity between the full object and the seed set is determined based on the statistical value on the target characteristics, and the efficiency and accuracy of determining the target full object that has similarity to the seed object is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548196B_ABST
    Figure CN114548196B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device, equipment and storage medium for locating similar objects based on data features, and relates to the field of data processing technology. The method includes: obtaining a first statistic corresponding to each feature of a seed set; obtaining a second statistic corresponding to each feature of a control set; obtaining the degree of difference between each feature corresponding to the seed set and the control set based on the first statistic and the second statistic corresponding to each feature; determining a target feature from each feature based on the degree of difference; obtaining a target feature value corresponding to the target feature of each full object in the full set; determining a target full object from each full object based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, wherein the target full object has similarity with the seed object in the seed set in terms of the target feature. The method improves the efficiency of determining a target full object that has similarity with the seed object based on data features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a method, device, electronic device, and computer-readable storage medium for locating similar objects based on data features. Background Art

[0002] When financial institutions perform anti-money laundering supervision and management duties and conduct money laundering risk analysis, they usually first analyze the characteristics of the customer population, then develop screening models or rules based on the extracted characteristics, and then use the established screening models or set screening rules to screen the risk population. The efficiency of finding risk populations is low.

[0003] As mentioned above, how to improve the efficiency of identifying risk groups based on data characteristics has become an urgent problem to be solved.

[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a method, apparatus, device and readable storage medium for locating similar objects based on data features, which at least to a certain extent improves the efficiency of determining the full set of target objects that are similar to the seed object based on data features.

[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0007] An embodiment of the present disclosure provides a method for locating similar objects based on data features, including: obtaining a first statistic corresponding to each feature of a seed set; obtaining a second statistic corresponding to each feature of a control set; obtaining a degree of difference between each feature corresponding to the seed set and the control set based on the first statistic and the second statistic corresponding to each feature; determining a target feature from each feature based on the degree of difference; obtaining a target feature value corresponding to the target feature of each full object in a full set; determining a target full object from each full object based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, the target full object having similarity with the seed object in the seed set in terms of the target feature.

[0008] An embodiment of the present disclosure provides a similar object positioning device based on data features, comprising: a seed feature statistics module, for obtaining a first statistic corresponding to each feature of a seed set; a control feature statistics module, for obtaining a second statistic corresponding to each feature of a control set; a feature difference acquisition module, for obtaining the degree of difference between each feature corresponding to the seed set and the control set based on the first statistic and the second statistic corresponding to each feature; a target feature determination module, for determining a target feature from each feature based on the degree of difference; a full feature acquisition module, for obtaining a target feature value corresponding to the target feature of each full object in the full set; and a similar object determination module, for determining a target full object from each full object based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, wherein the target full object has similarity with the seed object in the seed set in terms of the target feature.

[0009] According to one embodiment of the present disclosure, the similar object determination module is also used to determine the target full object from the various full objects based on the first statistic corresponding to the target feature of the seed set, the target feature value corresponding to the target feature of the various full objects, and the third statistic corresponding to the target feature of the full set.

[0010] According to one embodiment of the present disclosure, the first statistic corresponding to the target feature of the seed set includes the mean of the target feature values ​​corresponding to the target feature of the seed set, and the third statistic corresponding to the target feature of the full set includes the standard deviation of the target feature values ​​corresponding to the target feature of the full set; wherein the similar object determination module includes: a first feature mean difference acquisition module, used to obtain the difference between the target feature values ​​corresponding to each target feature of each full object and the mean of the target feature values ​​corresponding to each target feature of the seed set; a first feature ratio acquisition module, used to obtain the ratio of the difference corresponding to each target feature to the standard deviation of the target feature values ​​corresponding to the target feature of the full set; a similarity determination module, used to determine the degree of similarity between each full object and the seed set in the target feature based on the ratio corresponding to each target feature; a similar object selection module, used to select the target full object from each full object based on the degree of similarity between each full object and the seed set in the target feature.

[0011] According to one embodiment of the present disclosure, the similarity determination module includes: a feature weight acquisition module, used to obtain the weight parameters corresponding to each target feature; a similarity index calculation module, used to perform weighted summation based on the ratios corresponding to each target feature of each full object and their weight parameters, to obtain the similarity index of each full object; wherein, the smaller the similarity index of each full object, the greater the corresponding similarity.

[0012] According to an embodiment of the present disclosure, the first statistic corresponding to each feature of the seed set includes the mean of the eigenvalues ​​corresponding to each feature of the seed set and the standard deviation of the eigenvalues ​​corresponding to each feature of the seed set, and the first statistic corresponding to each feature of the control set includes the mean of the eigenvalues ​​corresponding to each feature of the control set and the standard deviation of the eigenvalues ​​corresponding to each feature of the control set; wherein the feature difference obtaining module includes: a second feature mean difference obtaining module, used to obtain the difference between the mean of the eigenvalues ​​corresponding to each feature of the control set and the mean of the eigenvalues ​​corresponding to each feature of the seed set; a feature standard deviation summing module, used to obtain the sum of the standard deviation of the eigenvalues ​​corresponding to each feature of the control set and the standard deviation of the eigenvalues ​​corresponding to each feature of the seed set; a second feature ratio obtaining module, used to obtain the ratio between the mean difference corresponding to each feature and the sum of the standard deviations corresponding to each feature; a difference degree determination module, used to determine the degree of difference between each feature corresponding to the seed set and the control set according to the ratio of each feature.

[0013] According to one embodiment of the present disclosure, the target feature determination module includes: a feature sorting module, which is used to sort each feature in order of difference from large to small to obtain a feature list; a target feature selection module, which is used to respond to a target feature selection instruction and determine the target feature from the feature list.

[0014] According to one embodiment of the present disclosure, the target feature selection module includes: a feature selection instruction response module, used to respond to the target feature selection instruction and obtain the feature corresponding to the target feature selection instruction; a feature threshold acquisition module, used to obtain the preset threshold of the feature corresponding to the target feature selection instruction; a target feature screening module, used to obtain the target feature based on the size relationship between the feature value corresponding to the feature corresponding to the target feature selection instruction and the preset threshold.

[0015] According to an embodiment of the present disclosure, the seed set includes a first number of the seed objects, and the full set includes a second number of the full objects, and the second number is greater than the first number; the device also includes: a full feature statistics module, used to obtain a third statistic corresponding to each feature of the full set; a full difference acquisition module, used to obtain the degree of difference between the seed set and the full set corresponding to each feature based on the first statistic and the third statistic corresponding to each feature; a difference threshold acquisition module, used to obtain the difference threshold of each feature; a control set acquisition module, used to obtain the control set as the full set when the degree of difference between the seed set and the full set corresponding to each feature is at least partially greater than the corresponding difference threshold.

[0016] According to one embodiment of the present disclosure, the control set acquisition module is also used to respond to a control object selection instruction and determine the control set from the full set when the degree of difference between the corresponding features of the seed set and the full set is less than the corresponding difference threshold.

[0017] An embodiment of the present disclosure provides an electronic device, comprising: a memory, a processor, and executable instructions stored in the memory and executable in the processor, wherein the processor implements any of the above methods when executing the executable instructions.

[0018] An embodiment of the present disclosure provides a computer-readable storage medium having computer-executable instructions stored thereon. When the executable instructions are executed by a processor, any of the above methods is implemented.

[0019] The present disclosure provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.

[0020] The embodiments of the present disclosure provide a similar object positioning method based on data features. The method obtains the degree of difference between the features corresponding to the seed set and the control set based on the first statistic corresponding to each feature of the seed set and the second statistic corresponding to each feature of the control set, and then determines the target feature from each feature based on the degree of difference, thereby reducing the number of features compared between the seed set and each full object. Then, based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, the target full object that has similarity with the seed object in the seed set in the target feature is determined from each full object. The similarity between the full object and the seed set can be judged on the target feature based on the statistic, thereby improving the efficiency and accuracy of determining the target full object that has similarity with the seed object based on data features.

[0021] It should be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other objects, features and advantages of the present disclosure will become more apparent by describing in detail example embodiments thereof with reference to the attached drawings.

[0023] Figure 1 A schematic diagram showing a system structure in an embodiment of the present disclosure.

[0024] Figure 2 A flowchart of a method for locating similar objects based on data features in an embodiment of the present disclosure is shown.

[0025] Figure 3 The figure is a flowchart of a method for obtaining a comparison set according to an exemplary embodiment.

[0026] Figure 4 The figure is a schematic diagram of an object selection interaction interface according to an exemplary embodiment.

[0027] Figure 5 Shown Figure 2 FIG. 5 is a schematic diagram of the processing process of step S206 in one embodiment.

[0028] Figure 6 Shown Figure 2 FIG. 5 is a schematic diagram of the processing process of step S208 in one embodiment.

[0029] Figure 7 The figure is a schematic diagram of a target feature selection interactive interface according to an exemplary embodiment.

[0030] Figure 8 Shown Figure 6 FIG. 5 is a schematic diagram of the processing process of step S2084 in one embodiment.

[0031] Figure 9 Shown Figure 2 FIG. 5 is a schematic diagram of the processing process of step S212 in one embodiment.

[0032] Figure 10 The figure is a schematic diagram of a target object selection interaction interface according to an exemplary embodiment.

[0033] Figure 11 Shown Figure 9 FIG. 5 is a schematic diagram of the processing process of step S2126 in one embodiment.

[0034] Figure 12 The figure is a schematic diagram of a similar population screening process according to an exemplary embodiment.

[0035] Figure 13 A block diagram of a similar object positioning device based on data features in an embodiment of the present disclosure is shown.

[0036] Figure 14 A block diagram of another similar object positioning device based on data features in an embodiment of the present disclosure is shown.

[0037] Figure 15 A schematic structural diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these examples are provided so that this disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures indicate identical or similar parts, and thus repeated descriptions thereof will be omitted.

[0039] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, devices, steps, etc. may be adopted. In other cases, well-known structures, methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0040] Furthermore, in the description of this disclosure, unless otherwise specified or limited, terms such as "connected" should be interpreted broadly. For example, they can mean electrically connected or capable of mutual communication; they can be directly connected or indirectly connected through an intermediary. "Multiple" means at least two, such as two or three, unless otherwise specified or limited. Those skilled in the art will understand the specific meanings of these terms in this disclosure based on the specific circumstances.

[0041] The abbreviations or self-defined nouns involved in this disclosure are explained below.

[0042] Anti-money laundering: refers to financial institutions controlling money laundering risks within the system through processes, rules, etc.

[0043] Review: In the prevention and control of money laundering risks, suspicious customers who have passed the rule audit need to be manually investigated to confirm whether they are real and suspicious enough, and then reported or not reported.

[0044] Audit: refers to the preliminary identification of suspicious money laundering clients through models or rules.

[0045] Characteristics: A certain attribute of a subject, such as "height", "age", "transaction amount in the last 7 days", etc. When certain characteristics or a combination of characteristics of a customer are different from those of a normal person, these characteristics can be called suspicious characteristics.

[0046] Feature configuration: refers to the personalized configuration of suspicious features related to money laundering risks through page interaction, so that they can be applied to risk prevention and control rules and audit people who meet the suspicious features.

[0047] Similar groups of people: refers to groups of people with certain similarities in characteristics, such as being in the same area, having similar ages, having similar transaction amounts, etc.

[0048] As mentioned above, when conducting money laundering risk analysis in related technologies, business personnel are often required to analyze cases and extract customer population characteristics, and then data analysts develop codes to implement feature data analysis and verify the differences in characteristics. Business personnel and data analysts need to work together to extract sensitive characteristics and screen corresponding populations that meet the characteristics. The entire process of finding risk feature populations involves multiple positions. The time required to analyze population characteristics and find similar populations is more than a week or even half a month, which is time-consuming and inefficient. Therefore, the present disclosure provides a similar object positioning method based on data features, which obtains the degree of difference between the corresponding features of the seed set and the control set based on the first statistic corresponding to each feature of the seed set and the second statistic corresponding to each feature of the control set, and then determines the target feature from each feature based on the degree of difference, and then determines the target full object from each full object that has similarity with the seed object in the seed set in the target feature based on the first statistic corresponding to each feature of the control set and the target feature value corresponding to the target feature of each full object, thereby improving the efficiency of determining the target object based on data features.

[0049] Figure 1 An exemplary system architecture 10 is shown to which the method and apparatus for locating similar objects based on data features of the present disclosure can be applied.

[0050] like Figure 1 As shown, system architecture 10 may include a terminal device 102, a network 104, a server 106, and a database 108. Terminal device 102 may be any electronic device with a display and input and output support, including but not limited to a smartphone, tablet computer, laptop computer, desktop computer, wearable device, virtual reality device, smart speaker, smart watch, smart home device, etc. Terminal device 102 and server 106 may be connected directly or indirectly via wired or wireless communication, which is not limited in this disclosure.

[0051] The network 104 is used to provide a medium for a communication link between the terminal device 102 and the server 106. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0052] The server 106 may be a standalone server, a server cluster or a distributed system consisting of multiple servers, or a cloud server providing cloud computing services. The database 108 may be a large database software located on a server or a small database software installed on a computer for storing data.

[0053] A user can use terminal device 102 to interact with server 106 and database 108 via network 104 to receive or send data, etc. For example, a user can use terminal device 102 to receive a list of population characteristics sorted by degree of difference from server 106 via network 104. Another example is a user can use terminal device 102 to retrieve a full list of populations from database 108 via network 104, and then use interactive software on terminal device 102 to send a selected control population to server 106 via network 104.

[0054] The server 106 may also receive data from or send data to the database 108 via the network 104. For example, the server 106 may be a backend processing server configured to obtain feature value data for each feature of the entire population from the database 108 via the network 104. For another example, the server 106 may be configured to obtain a user-selected target feature from the terminal device 102 via the network 104, and to obtain feature value data for the target feature of the seed population and the control population from the database 108.

[0055] It should be understood that Figure 1 The number of terminal devices, networks, servers and databases in the embodiment is merely illustrative. Any number of terminal devices, networks, servers and databases may be provided as required.

[0056] Figure 2 FIG. 1 is a flow chart showing a method for locating similar objects based on data features according to an exemplary embodiment. Figure 2 The method shown can be applied to a server of the above system, or to a terminal device of the above system, for example.

[0057] refer to Figure 2 , the method 20 provided in the embodiment of the present disclosure may include the following steps.

[0058] In step S202, a first statistic corresponding to each feature of the seed set is obtained.

[0059] In the disclosed embodiment, the seed set may be a set of objects that have been determined to have specific common characteristics, such as a set of risk groups in anti-money laundering analysis, etc. Some business-related attributes of the seed objects in the seed set can be used as features. For example, for attributes with specific values ​​such as transaction amount and quantity of goods in the last 20 days, their attribute values ​​can be used as corresponding feature values; for attributes without specific values ​​such as gender and region, they can be assigned values ​​according to category as corresponding feature values. For example, the feature value of the gender feature of "female" is 1, the feature value of the gender feature of "male" is 2, and so on. The first statistic of the feature refers to the indicator reflecting the central tendency of the set on the feature of this type, obtained by statistically analyzing the feature values ​​of all objects in the set for each type of feature. The first statistic can be the mean of the feature values, such as the arithmetic mean, weighted mean, etc., or it can be the variance, standard deviation, etc. of the feature values.

[0060] In step S204, a second statistic corresponding to each feature of the comparison set is obtained.

[0061] In the disclosed embodiment, the control set may be a collection of objects that may not necessarily share specific common characteristics, such as a general population. The characteristics of the control objects in the control set correspond one-to-one with the characteristics of the seed objects. A description of the characteristics and statistics can be found in step S202 and will not be repeated here.

[0062] In some embodiments, for example, when the feature differences between the seed set (such as the blacklist population) and the general population of the full set are not obvious, the whitelist population can be selected as the control set to specifically compare the main difference features of the two input populations. Figure 3 .

[0063] In step S206 , the degree of difference between the seed set and the control set corresponding to each feature is obtained based on the first statistic and the second statistic corresponding to each feature.

[0064] In the embodiment of the present disclosure, after obtaining the first statistic and the second statistic respectively reflecting the central tendency of each feature of the seed set and the control set, the first statistic and the second statistic can be compared to obtain the degree of difference of each feature of the two sets.

[0065] In some embodiments, for example, the first statistic and the second statistic may be the means of each feature of the seed set and the control set, respectively. The two means of each feature are subtracted, and the degree of difference is determined based on the size of the difference. A larger difference indicates a greater degree of difference.

[0066] In other embodiments, for example, the first statistic and the second statistic may include the mean and standard deviation of each feature of the seed set and the control set, respectively, and the degree of difference of each feature is calculated based on the mean and standard deviation. Figure 5 .

[0067] In step S208 , a target feature is determined from among the features based on the degree of difference.

[0068] In the disclosed embodiments, a seed set is a collection of objects sharing specific common characteristics. The target features to be determined are the common characteristics of the seed set obtained through analysis. The difference between the seed set and a collection of objects sharing no common characteristics can be reflected in these common characteristics. For example, in anti-money laundering analysis, the feature "transaction amount in the last 30 days" may be the feature that most distinguishes a risk group from the general population, and this feature can be determined as the target feature.

[0069] In some embodiments, for example, the features may be sorted from largest to smallest according to the degree of difference, and a number of features ranked at the top may be automatically selected as target features.

[0070] In other embodiments, for example, each feature can be sorted from largest to smallest according to the degree of difference and then displayed to the user through an interactive interface. The user can be a business person, who can select appropriate target features based on the degree of difference between the displayed seed set and the control set (general population) in each feature, as well as his or her own understanding of business experience, such as selecting suspicious features from the list recommended by the system. Figures 6 to 8 .

[0071] In step S210 , the target feature value corresponding to the target feature of each full object in the full set is obtained.

[0072] In the embodiment of the present disclosure, for a determined target feature, a feature value of each object in the full set to be screened is obtained, to be compared with the feature value of the feature in the seed set to determine similarity.

[0073] In step S212, a target full object is determined from each full object based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, and the target full object has similarity with the seed objects in the seed set in the target feature.

[0074] In the disclosed embodiment, since the target feature is a feature that has a large difference between the seed set screened in the previous step and the control set, the determined target full set also has a similar difference with the control set in this feature.

[0075] In some embodiments, for example, a target full object is determined from each full object based on the first statistic corresponding to the target feature of the seed set, the target feature value corresponding to the target feature of each full object, and the third statistic corresponding to the target feature of the full set. Figure 3 .

[0076] In other embodiments, for example, the first statistic corresponding to the target feature of the seed set may be the mean of the target feature of the seed set, and the target feature values ​​corresponding to the target feature of each full object are compared with the mean corresponding to the seed set. For example, a difference comparison may be performed. When the difference is less than a predetermined threshold, it can be considered that the full object and the seed set are similar in this feature, and the full object is determined as the target full object.

[0077] The embodiments of the present disclosure provide a similar object positioning method based on data features. The method obtains the degree of difference between the features corresponding to the seed set and the control set based on the first statistic corresponding to each feature of the seed set and the second statistic corresponding to each feature of the control set, and then determines the target feature from each feature based on the degree of difference, thereby reducing the number of features compared between the seed set and each full object. Then, based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, the target full object that has similarity with the seed object in the seed set in the target feature is determined from each full object. The similarity between the full object and the seed set can be judged on the target feature based on the statistic, thereby improving the efficiency and accuracy of determining the target full object that has similarity with the seed object based on data features.

[0078] Figure 3 FIG. 1 is a flow chart showing a method for obtaining a control set according to an exemplary embodiment. Figure 3 The method shown can be applied to a server of the above system, or to a terminal device of the above system, for example. Figure 3 The method shown can be used to determine the control objects in the control set before step S204.

[0079] refer to Figure 3 , the method 30 provided in the embodiment of the present disclosure may include the following steps.

[0080] In step S302, a third statistic corresponding to each feature of the full set is obtained.

[0081] In the disclosed embodiment, the objects in the full set can be all types of objects, for example, including risk groups (blacklisted groups), whitelisted groups, unknown groups to be screened, and so on. The features of the objects in the full set correspond one-to-one with the features of the seed objects. For a description of the features and statistics, refer to step S202 and will not be repeated here.

[0082] In step S304, the difference between the seed set and the full set corresponding to each feature is obtained based on the first statistic and the third statistic corresponding to each feature.

[0083] In the embodiment of the present disclosure, the specific calculation method of the difference degree can refer to step S206 and will not be repeated here.

[0084] In step S306, the difference threshold of each feature is obtained.

[0085] In the disclosed embodiment, a threshold may be set for the difference between each feature of the seed set and the full set. A single threshold may be set for the entire set, or thresholds may be set for each feature separately or in batches to measure whether the difference in features between the seed set and the full set is obvious.

[0086] In step S308 , when the degree of difference between the seed set and the full set corresponding to each feature is at least partially greater than the corresponding difference threshold, the control set is obtained as the full set.

[0087] In the embodiment of the present disclosure, the degree of difference between the seed set and the full set corresponding to each feature is at least partially greater than the corresponding difference threshold, indicating that the difference in the features is obvious enough to distinguish the more abnormal features to obtain the target features.

[0088] In step S310 , when the degree of difference between the seed set and the full set corresponding to each feature is less than the corresponding difference threshold, a control set is determined from the full set in response to a control object selection instruction.

[0089] In the embodiment of the present disclosure, when the degree of difference between the corresponding features of the seed set and the full set is less than the corresponding difference threshold, it can be considered that the characteristics of the blacklisted population in the seed set and the general population in the full set are not significantly different, and the whitelisted population can be selected from the full set as the control population.

[0090] In some embodiments, for example, in addition to selecting the whitelist population from the general population in the established full set, the whitelist population can also be manually input. Figure 4 FIG. 1 is a schematic diagram of an object selection interaction interface according to an exemplary embodiment. Figure 4 As shown, Figure 4At the top right is a clickable button, and below is an input box or a list of displayed objects. Since the number of subjects in the comparison population may be large, you can click the "Upload Population File" button to enter the control or seed population (comparison population) in file format; you can also enter the subject's identity (ID) in the input box below; or select the identity of the control or seed subject from the list below. After obtaining the input comparison population, based on the user IDs of the comparison population, the feature values ​​corresponding to these IDs are obtained from the database, and then statistics such as the mean and standard deviation of each feature of these users are calculated.

[0091] According to the method provided in the embodiments of the present disclosure, the control set compared with the seed set is screened by comparing the degree of difference in features with the seed set, thereby improving the accuracy of the determined target features.

[0092] Figure 5 Shown Figure 2 Schematic diagram of the processing process of step S206 in one embodiment. In one embodiment, the first statistic corresponding to each feature of the seed set includes the mean of the eigenvalues ​​corresponding to each feature of the seed set and the standard deviation of the eigenvalues ​​corresponding to each feature of the seed set, and the first statistic corresponding to each feature of the control set includes the mean of the eigenvalues ​​corresponding to each feature of the control set and the standard deviation of the eigenvalues ​​corresponding to each feature of the control set. Figure 5 As shown, in the embodiment of the present disclosure, the above step S206 may further include the following steps.

[0093] Step S2062: Obtain the difference between the mean of the feature values ​​corresponding to each feature of the control set and the mean of the feature values ​​corresponding to each feature of the seed set.

[0094] Step S2064: Obtain the sum of the standard deviation of the eigenvalues ​​corresponding to each feature of the control set and the standard deviation of the eigenvalues ​​corresponding to each feature of the seed set.

[0095] Step S2066: Obtain the ratio of the mean difference corresponding to each feature to the sum of the standard deviations corresponding to each feature.

[0096] Step S2068: Determine the degree of difference between the seed set and the control set corresponding to each feature based on the ratio of each feature.

[0097] In the embodiment of the present disclosure, the ratio of each feature is expressed as the difference Y, and the calculation formula of the difference Y is as follows:

[0098] Difference Y = | mean 对照 -mean 种子 | / (Standard Deviation 对照 +Standard Deviation 种子 ) (1)

[0099] From formula (1), we can see that the greater the difference in the mean between the two samples and the smaller the standard deviation, the greater the difference Y of the calculated results. The greater the difference Y, the greater the degree of difference.

[0100] According to the method provided in the embodiment of the present disclosure, the degree of difference of each feature is obtained by calculating the ratio of the difference between the mean of the seed set and the control set to the sum of the standard deviations, so that the characteristic features of the seed set can be accurately determined from the various features.

[0101] Figure 6 Shown Figure 2 Schematic diagram of the processing process of step S208 in one embodiment. Figure 6 As shown, in the embodiment of the present disclosure, the above step S208 may further include the following steps.

[0102] Step S2082: Sort the features in descending order of difference to obtain a feature list.

[0103] In the disclosed embodiment, the difference between each feature of the population in the input seed set and that of the general population can be calculated and analyzed, and then the feature list can be displayed to the user for selection after sorting by the degree of difference. Figure 7 FIG. 1 is a schematic diagram of an interactive interface for selecting target features according to an exemplary embodiment. Figure 7 As shown, the multiple features in the list displayed in the interactive interface are sorted from large to small according to the feature difference. The list can also display the main indicators of the control population (general population) and seed population (comparison population) under each feature, such as mean, standard deviation, etc.

[0104] Step S2084: respond to the target feature selection instruction and determine the target feature from the feature list. Figure 7 As shown, each feature in the list displayed in the interactive interface has a corresponding selection option. The user can check the corresponding box to indicate the selection. After selecting all target features, click the "Filter Population" button in the lower right corner, and the system will obtain the target features selected by the user.

[0105] According to the method provided in the embodiments of the present disclosure, after automatically calculating the feature differences between the seed population and the general population, feature recommendations are made based on the degree of difference, and then the target features are optimized based on the results of manual selection to screen the risk population in the entire population. While achieving high efficiency, it incorporates manual experience factors, enhances adaptability to different businesses, and improves the accuracy of screening people with similar features to the seed population.

[0106] Figure 8 Shown Figure 6Schematic diagram of the processing process of step S2084 in one embodiment. Figure 8 As shown, in the embodiment of the present disclosure, the above step S2084 may further include the following steps.

[0107] Step S20842: respond to the target feature selection instruction and obtain the feature corresponding to the target feature selection instruction.

[0108] Step S20844, obtaining the preset threshold of the feature corresponding to the target feature selection instruction.

[0109] Step S20846, obtaining the target feature according to the relationship between the feature value corresponding to the feature corresponding to the target feature selection instruction and the preset threshold.

[0110] In the embodiment of the present disclosure, after manually selecting a feature, additional artificial restrictions can be placed on the feature. In some embodiments, for example, the seed population can be restricted to meet a preset threshold on the feature, such as Figure 7 The feature of "recent transaction amount" has the largest feature difference and is ranked first in the list. When users select this feature as the target feature, they can also set a threshold of 1 million for it ( Figure 7 After the threshold is set, the feature will be pushed and displayed as the target feature only when the mean value of the feature in the seed set is greater than 1 million.

[0111] In other embodiments, for example, the threshold value set for the target feature can also be used to screen the target objects (risk groups) for the feature, for example, Figure 7 For the "recent transaction amount" feature, after setting a threshold of 1 million for it, the subsequent steps calculate the similarity between the recent transaction amount values ​​of each object in the full set and the seed set, and sort them. The values ​​can be compared with the set threshold. Only when the recent transaction amount is greater than 1 million will it be pushed to the front end and displayed as a target risk object.

[0112] According to the method provided in the embodiments of the present disclosure, by automatically calculating the feature differences between the seed population and the general population, recommending features based on the degree of difference, and then optimizing the risk population screening plan based on the set feature threshold, the efficiency and accuracy of screening people with similar features to the seed population are improved.

[0113] Figure 9 Shown Figure 2 Schematic diagram of the processing process of step S212 in one embodiment. In one embodiment, the first statistic corresponding to the target feature of the seed set includes the mean of the target feature value corresponding to the target feature of the seed set, and the third statistic corresponding to the target feature of the full set includes the standard deviation of the target feature value corresponding to the target feature of the full set. Figure 9As shown, in the embodiments of the present disclosure, the above step S212 may further include the following steps.

[0114] Step S2122: Obtain the difference between the target feature values corresponding to each target feature of each full - scale object and the mean of the target feature values corresponding to each target feature of the seed set.

[0115] Step S2124: Obtain the ratio of the difference corresponding to each target feature to the standard deviation of the target feature values corresponding to the target features of the full - scale set.

[0116] Step S2126: Determine the similarity degree between each full - scale object and the seed set in terms of the target features according to the ratios corresponding to each target feature.

[0117] In some embodiments, for example, if there are k target features, for a full - scale object n, the similarity index Sn between it and the seed set in the k target features can be calculated by the following formula:

[0118] Sn = ∑ k (|feature value nj - mean of the full - scale set feature j| / standard deviation of the full - scale set feature j) (2) In the formula, 0 < j ≤ k, and j is a positive integer. Formula (2) means that for each full - scale object n, the differences between the feature means of all selected k target features and the corresponding feature values of this object are divided by the standard deviation of the corresponding feature in the full - scale set, and then accumulated to obtain the similarity index Sn. The larger the similarity index Sn, the greater the difference between the full - scale object and the seed population within the selected features, that is, the smaller the similarity degree.

[0119] Step S2128: Select target full - scale objects from each full - scale object according to the similarity degree between each full - scale object and the seed set in terms of the target features.

[0120] In the embodiments of the present disclosure, the first predetermined number of accounts with the smallest similarity index can be selected as target full - scale objects and displayed to the front - end for users to view and analyze, so as to further screen similar populations. Figure 10 It is a schematic diagram of an interactive interface for selecting target objects shown according to an exemplary embodiment. As Figure 10 shown, in the account list sorted by similarity index, in addition to displaying the core index of similarity, other index information of the target account can also be displayed, such as basic attributes or risk data such as account name, registration time, transaction amount, etc. According to the list of similar populations recommended by the system, business personnel can manually query and analyze the risks of each account one by one, or analyze them on other systems by means of export, call, etc.

[0121] According to the method provided in the embodiments of the present disclosure, the degree of similarity between the target feature of all objects to be screened and the seed set in the target feature is calculated by taking the difference between the feature value of the target feature of the full set of objects to be screened and the mean value of the feature of the seed set and then comparing it with the standard deviation of the feature of the seed set to screen the target objects, thereby improving the accuracy of screening people with similar characteristics to the seed population.

[0122] Figure 11 Shown Figure 9 Schematic diagram of the processing process of step S2126 in one embodiment. Figure 11 As shown, in the embodiment of the present disclosure, the above step S2126 may further include the following steps.

[0123] Step S21262, obtain the weight parameters corresponding to each target feature.

[0124] Step S21264, performing weighted summation based on the ratios and weight parameters corresponding to the target features of the full objects, to obtain similarity indices of the full objects; wherein, the smaller the similarity indices of the full objects, the greater the corresponding similarity.

[0125] In the disclosed embodiments, when determining the target feature, the weight of the selected feature in the similarity algorithm can be manually set. For example, if a business operator places more importance on the feature of recent transaction amount, the weight of this feature can be increased when determining the target feature based on the difference. This will increase the weight of this feature in the similarity calculation when calculating the similarity index.

[0126] According to the method provided in the embodiment of the present disclosure, the weight of the target feature is taken into account when calculating the similarity between the target feature of all objects and the seed set to perform similarity calculation, thereby improving the accuracy of screening people with similar characteristics to the seed population.

[0127] Figure 12 FIG. 1 is a schematic diagram of a similar population screening process according to an exemplary embodiment. Figure 12 The system first obtains the feature value data (122) of all users in the population, such as the age, gender, transaction amount, number of products, and other feature data of all users in the banking system covered by the system. The similar population screening process may include the following steps.

[0128] Step S1202: The system front desk obtains the input seed population users.

[0129] Step S1204 , obtaining the feature values ​​of each feature corresponding to the seed population users' IDs according to the IDs of the seed population users, and calculating the mean and standard deviation of the seed users.

[0130] Step S1206 : Based on the feature values ​​of all users in the population, the mean and standard deviation of each feature of all users in the population (ordinary users) are calculated.

[0131] Step S1208: Calculate the difference of the seed population in all features using the feature means and standard deviations of the seed population and the full population.

[0132] Step S1210, sort by difference value and display on the foreground, manually select and confirm features with large differences, and obtain a difference feature set (124).

[0133] Step S1212, based on the manually selected features, as well as the mean of the seed population, the variance or standard deviation of the full population, etc., screen the users who best meet these features from the full population, that is, the users who have the smallest difference between the seed population and the full population in these features.

[0134] Step S1212: Select the first predetermined number of users with the smallest differences and display them to the front desk for staff to review and analyze for risk detection.

[0135] According to the method provided by the present disclosure, based on the automatic cleaning and comparison of data, business personnel can quickly complete population characteristic analysis and similar population screening on their own, which greatly improves the efficiency of screening similar risk populations. The analysis and screening operations can generally be completed within several hours.

[0136] Figure 13 FIG. 1 is a block diagram of a device for locating similar objects based on data features according to an exemplary embodiment. Figure 13 The device shown can be applied to the server side of the above system, for example, and can also be applied to the terminal device of the above system.

[0137] refer to Figure 13 The device 130 provided in the embodiment of the present disclosure may include a seed feature statistics module 1302, a control feature statistics module 1304, a feature difference acquisition module 1306, a target feature determination module 1308, a full feature acquisition module 1310 and a similar object determination module 1312.

[0138] The seed feature statistics module 1302 may be configured to obtain a first statistic corresponding to each feature of the seed set.

[0139] The control feature statistics module 1304 may be configured to obtain a second statistic corresponding to each feature of the control set.

[0140] The feature difference obtaining module 1306 may be configured to obtain the degree of difference between the seed set and the control set corresponding to each feature based on the first statistic and the second statistic corresponding to each feature.

[0141] The target feature determination module 1308 may be configured to determine a target feature from among the features based on the degree of difference.

[0142] The full feature acquisition module 1310 may be used to acquire target feature values ​​corresponding to target features of all full objects in the full set.

[0143] The similar object determination module 1312 can be used to determine a target full object from each full object based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, where the target full object has similarity with the seed object in the seed set in the target feature.

[0144] Figure 14 FIG. 1 is a block diagram of a device for locating similar objects based on data features according to an exemplary embodiment. Figure 14 The device shown can be applied to the server side of the above system, for example, and can also be applied to the terminal device of the above system.

[0145] refer to Figure 14 The device 140 provided in the embodiment of the present disclosure may include a seed feature statistics module 1402, a full feature statistics module 14032, a full difference acquisition module 14034, a difference threshold acquisition module 14036, a control set acquisition module 14038, a control feature statistics module 1404, a feature difference acquisition module 1406, a target feature determination module 1408, a full feature acquisition module 1410 and a similar object determination module 1412, wherein the feature difference acquisition module 1406 may include a second feature mean difference acquisition module 14062, a feature standard deviation summation module 14064, a second feature ratio acquisition module 14066 and a difference degree threshold acquisition module 1412. The target feature determination module 1408 may include a feature sorting module 14082 and a target feature selection module 14084. The target feature selection module 14084 may include a feature threshold acquisition module 140844 and a target feature screening module 140846. The similar object determination module 1412 may include a first feature mean difference acquisition module 14122, a first feature ratio acquisition module 14124, a similarity determination module 14126 and a similar object selection module 14128. The similarity determination module 14126 may include a feature weight acquisition module 141262 and a similarity index calculation module 141264.

[0146] The seed feature statistics module 1402 may be configured to obtain a first statistic corresponding to each feature of the seed set. The first statistic corresponding to each feature of the seed set includes a mean of the eigenvalues ​​corresponding to each feature of the seed set and a standard deviation of the eigenvalues ​​corresponding to each feature of the seed set. The seed set includes a first number of seed objects.

[0147] The full feature statistics module 14032 may be used to obtain a third statistic corresponding to each feature of the full set.

[0148] The full difference obtaining module 14034 may be used to obtain the degree of difference between the seed set and the full set corresponding to each feature based on the first statistic and the third statistic corresponding to each feature.

[0149] The difference threshold acquisition module 14036 may be used to obtain the difference threshold of each feature.

[0150] The control set obtaining module 14038 may be configured to obtain the control set as the full set when the degree of difference between the seed set and the full set corresponding to each feature is at least partially greater than a corresponding difference threshold.

[0151] The control set obtaining module 14038 may also be configured to determine a control set from the full set in response to a control object selection instruction when the degree of difference between each feature corresponding to the seed set and the full set is less than a corresponding difference threshold.

[0152] The control feature statistics module 1404 can be used to obtain the second statistics corresponding to each feature of the control set. The first statistics corresponding to each feature of the control set include the mean of the feature value corresponding to each feature of the control set and the standard deviation of the feature value corresponding to each feature of the control set.

[0153] The feature difference obtaining module 1406 may be configured to obtain the degree of difference between the seed set and the control set corresponding to each feature based on the first statistic and the second statistic corresponding to each feature.

[0154] The second feature mean difference obtaining module 14062 may be used to obtain the difference between the mean of the feature values ​​corresponding to each feature of the control set and the mean of the feature values ​​corresponding to each feature of the seed set.

[0155] The feature standard deviation summing module 14064 may be used to obtain the sum of the standard deviation of the feature value corresponding to each feature of the control set and the standard deviation of the feature value corresponding to each feature of the seed set.

[0156] The second feature ratio obtaining module 14066 may be used to obtain the ratio between the mean difference corresponding to each feature and the sum of the standard deviations corresponding to each feature.

[0157] The difference degree determination module 14068 may be used to determine the difference degree of each feature between the seed set and the control set based on the ratio of each feature.

[0158] The target feature determination module 1408 may be configured to determine a target feature from among the features based on the degree of difference.

[0159] The feature sorting module 14082 may be used to sort the features in descending order of difference to obtain a feature list.

[0160] The target feature selection module 14084 may be configured to respond to a target feature selection instruction and determine a target feature from a feature list.

[0161] The feature selection instruction response module 140842 can be used to respond to the target feature selection instruction and obtain the features corresponding to the target feature selection instruction.

[0162] The feature threshold acquisition module 140844 may be used to obtain a preset threshold of a feature corresponding to a target feature selection instruction.

[0163] The target feature screening module 140846 may be used to obtain the target feature based on a relationship between a feature value corresponding to the feature corresponding to the target feature selection instruction and a preset threshold.

[0164] The full feature acquisition module 1410 may be configured to acquire target feature values ​​corresponding to target features of all full objects in the full set.

[0165] The similar object determination module 1412 may be configured to determine a target full object from the full objects based on a first statistic corresponding to a target feature of the seed set and a target feature value corresponding to the target feature of each full object, wherein the target full object has similarity with the seed objects in the seed set in terms of the target feature. The full set includes a second number of full objects, the second number being greater than the first number.

[0166] The similar object determination module 1412 may also be configured to determine a target full object from each full set of objects based on a first statistic corresponding to the target feature of the seed set, target feature values ​​corresponding to the target feature of each full set of objects, and a third statistic corresponding to the target feature of the full set. The first statistic corresponding to the target feature of the seed set includes the mean of the target feature values ​​corresponding to the target feature of the seed set, and the third statistic corresponding to the target feature of the full set includes the standard deviation of the target feature values ​​corresponding to the target feature of the full set.

[0167] The first feature mean difference obtaining module 14122 may be used to obtain the difference between the target feature value corresponding to each target feature of each full object and the mean of the target feature value corresponding to each target feature of the seed set.

[0168] The first feature ratio obtaining module 14124 may be used to obtain the ratio of the difference corresponding to each target feature to the standard deviation of the target feature value corresponding to the target feature of the entire set.

[0169] The similarity determination module 14126 may be used to determine the similarity between each full set of objects and the seed set in terms of the target features based on the ratios corresponding to the target features.

[0170] The feature weight acquisition module 141262 can be used to obtain the weight parameters corresponding to each target feature.

[0171] The similarity index calculation module 141264 can be used to perform weighted summation based on the ratios corresponding to the target features of each full object and their weight parameters to obtain the similarity index of each full object; wherein, the smaller the similarity index of each full object, the greater the corresponding similarity.

[0172] The similar object selection module 14128 may be configured to select a target full object from the full objects based on the degree of similarity between the full objects and the seed set in terms of target features.

[0173] The specific implementation of each module in the device provided by the embodiment of the present disclosure can refer to the content of the above method and will not be repeated here.

[0174] Figure 15 FIG. 1 shows a schematic diagram of the structure of an electronic device in an embodiment of the present disclosure. It should be noted that: Figure 15 The device shown is only an example of a computer system and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0175] like Figure 15 As shown, device 1500 includes a central processing unit (CPU) 1501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1502 or a program loaded from a storage portion 1508 into a random access memory (RAM) 1503. Various programs and data required for the operation of device 1500 are also stored in RAM 1503. CPU 1501, ROM 1502, and RAM 1503 are connected to each other via a bus 1504. An input / output (I / O) interface 1505 is also connected to bus 1504.

[0176] The following components are connected to the I / O interface 1505: an input section 1506 including a keyboard, a mouse, and the like; an output section 1507 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1508 including a hard disk; and a communication section 1509 including a network interface card such as a LAN card or a modem. The communication section 1509 performs communication processing via a network such as the Internet. A drive 1510 is also connected to the I / O interface 1505 as needed. Removable media 1511, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1510 as needed, so that computer programs read therefrom can be installed into the storage section 1508 as needed.

[0177] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1509, and / or installed from a removable medium 1511. When the computer program is executed by the central processing unit (CPU) 1501, the above-mentioned functions defined in the system of the present disclosure are performed.

[0178] It should be noted that the computer-readable medium described in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0179] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0180] The modules involved in the embodiments described in the present disclosure may be implemented in software or in hardware. The modules described may also be provided in a processor. For example, they may be described as follows: a processor including a seed feature statistics module, a feature difference acquisition module, a target feature determination module, a full feature acquisition module, and a similar object determination module. The names of these modules do not, in some cases, constitute a limitation on the modules themselves. For example, the seed feature statistics module may also be described as a "module for obtaining the first statistic corresponding to each feature of the seed set."

[0181] According to one aspect of the present disclosure, the present disclosure also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: obtaining a first statistic corresponding to each feature of the seed set; obtaining a second statistic corresponding to each feature of the control set; obtaining the degree of difference between each feature corresponding to the seed set and the control set based on the first statistic and the second statistic corresponding to each feature; determining a target feature from each feature based on the degree of difference; obtaining a target feature value corresponding to the target feature of each full object in the full set; determining a target full object from each full object based on the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, and the target full object has similarity with the seed object in the seed set in terms of the target feature.

[0182] According to one aspect of the present disclosure, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.

[0183] While the exemplary embodiments of the present disclosure have been specifically illustrated and described above, it should be understood that the present disclosure is not limited to the detailed structures, configurations, or implementations described herein; rather, the present disclosure is intended to encompass various modifications and equivalent configurations within the spirit and scope of the appended claims.

Claims

1. A similar object positioning method based on data features, characterized in that: Applied to a system comprising a terminal device, a server, and a database, the method comprises: Obtaining a first statistic corresponding to each feature of a seed set, where the seed set is a set of people having specific common features, where the features include at least one of transaction amount, number of goods, gender, and region; Obtaining a second statistic corresponding to each feature of a control set, where the control set is a set of people who are uncertain about having a specific common feature; Comparing the first statistic and the second statistic corresponding to each feature to obtain the degree of difference between the seed set and the control set corresponding to each feature; Sort the features in descending order of difference to obtain a feature list; In response to a target feature selection instruction, determining a target feature from the feature list; Obtaining a target feature value corresponding to the target feature of each full set of objects in a full set, where the full set is a set of people from which to be screened for similarities with the people in the seed set in terms of the target feature; According to the first statistic corresponding to the target feature of the seed set and the target feature value corresponding to the target feature of each full object, a target full object is determined from the various full objects, and the target full object is a group of people who have similarities with the people in the seed set in the target feature.

2. The method according to claim 1, characterized in that Determining a target full object from the full objects according to the first statistic corresponding to the target feature of the seed set and the target feature values ​​corresponding to the target features of the full objects includes: The target full object is determined from the full objects according to the first statistic corresponding to the target feature of the seed set, the target feature value corresponding to the target feature of the full objects, and the third statistic corresponding to the target feature of the full set.

3. The method according to claim 2, characterized in that The first statistic corresponding to the target feature of the seed set includes the mean of the target feature values ​​corresponding to the target feature of the seed set, and the third statistic corresponding to the target feature of the full set includes the standard deviation of the target feature values ​​corresponding to the target feature of the full set; The determining the target full object from the respective full objects according to the first statistic corresponding to the target feature of the seed set, the target feature value corresponding to the target feature of each full object, and the third statistic corresponding to the target feature of the full set includes: Obtaining the difference between the target feature value corresponding to each target feature of each full object and the mean of the target feature value corresponding to each target feature of the seed set; Obtaining a ratio of a difference value corresponding to each target feature to a standard deviation of a target feature value corresponding to the target feature of the entire set; Determining the similarity between the target features and the seed set based on the ratios corresponding to the target features; The target full object is selected from the full objects according to the similarity between the full objects and the seed set in terms of the target feature.

4. The method according to claim 3, characterized in that Determining the similarity between each full set of objects and the seed set in terms of the target features based on the ratios corresponding to the target features includes: Get the weight parameters corresponding to each target feature; The similarity index of each full object is obtained by performing weighted summation based on the ratios and weight parameters corresponding to each target feature of each full object; Among them, the smaller the similarity index of each full object, the greater the corresponding similarity.

5. The method according to claim 1, wherein The first statistic corresponding to each feature of the seed set includes a mean of the eigenvalues ​​corresponding to each feature of the seed set and a standard deviation of the eigenvalues ​​corresponding to each feature of the seed set; the first statistic corresponding to each feature of the control set includes a mean of the eigenvalues ​​corresponding to each feature of the control set and a standard deviation of the eigenvalues ​​corresponding to each feature of the control set; The obtaining, based on the first statistic and the second statistic corresponding to each feature, the degree of difference between the seed set and the control set corresponding to each feature includes: Obtaining the difference between the mean of the feature values ​​corresponding to each feature of the control set and the mean of the feature values ​​corresponding to each feature of the seed set; Obtaining the sum of the standard deviation of the characteristic values ​​corresponding to each characteristic of the control set and the standard deviation of the characteristic values ​​corresponding to each characteristic of the seed set; Get the ratio between the mean difference corresponding to each feature and the sum of the standard deviations corresponding to each feature; The degree of difference between the seed set and the control set corresponding to each feature is determined based on the ratio of each feature.

6. The method according to claim 1, wherein In response to the selection instruction, determining the target feature from the feature list includes: In response to the target feature selection instruction, obtaining a feature corresponding to the target feature selection instruction; Obtaining a preset threshold value of a feature corresponding to the target feature selection instruction; The target feature is obtained according to the relationship between the feature value corresponding to the feature corresponding to the target feature selection instruction and the preset threshold.

7. The method according to any one of claims 1 to 5, characterized in that The seed set includes a first number of seed objects, and the full set includes a second number of the full objects, where the second number is greater than the first number; The method further comprises: Obtaining a third statistic corresponding to each feature of the full set; Obtaining, based on the first statistic and the third statistic corresponding to each feature, a degree of difference between the seed set and the full set corresponding to each feature; Get the difference threshold of each feature; When the degree of difference between the seed set and the full set corresponding to each feature is at least partially greater than the corresponding difference threshold, the control set is obtained as the full set.

8. The method according to claim 7, characterized in that When the degree of difference between the seed set and the full set corresponding to each feature is less than the corresponding difference threshold, the control set is determined from the full set in response to a control object selection instruction.

9. A similar object positioning device based on data features, characterized in that: Applied to a system comprising a terminal device, a server and a database, the apparatus comprises: A seed feature statistics module is configured to obtain a first statistic corresponding to each feature of a seed set, wherein the seed set is a set of people having specific common features, wherein the features include at least one of transaction amount, number of goods, gender, and region; A control feature statistics module, configured to obtain a second statistic corresponding to each feature of a control set, wherein the control set is a set of people with uncertain common characteristics; a feature difference obtaining module, configured to compare the first statistic and the second statistic corresponding to each feature to obtain the degree of difference between the seed set and the control set corresponding to each feature; The feature sorting module is used to sort the features in descending order of difference to obtain a feature list; a target feature selection module, configured to respond to a target feature selection instruction and determine a target feature from the feature list; a full feature acquisition module, configured to obtain a target feature value corresponding to the target feature of each full object in a full set, wherein the full set is a set of people to be screened for similarities with the people in the seed set in terms of the target feature; A similar object determination module is used to determine a target full object from the various full objects based on a first statistic corresponding to the target feature of the seed set and a target feature value corresponding to the target feature of each full object, wherein the target full object is a group of people who have similarities with the people in the seed set in the target feature.

10. An electronic device comprising: A memory, a processor, and executable instructions stored in the memory and executable in the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the executable instructions.

11. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Content recommendation method and device, training method and device, equipment and storage medium

    CN110162703A

  • Generalization model training method and device based on artificial intelligence

    CN111143684A