Key feature determination method and apparatus, computer device, medium, and program product

By using nonparametric validation and factor analysis to determine the key features of the target object set, the problem of low accuracy in existing technologies is solved, and more accurate feature analysis and improved user experience are achieved.

CN116662664BActive Publication Date: 2026-05-01CHINA CONSTRUCTION BANK +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA CONSTRUCTION BANK
Filing Date
2023-06-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of user key feature analysis is low, which makes it impossible to effectively improve the user experience.

Method used

By acquiring the target object set and the preset feature pool, the control object set is determined, and the feature difference value is calculated based on the nonparametric verification method. Key features are identified using factor analysis and difference rules, redundant features are eliminated, and frequency statistics are generated.

Benefits of technology

Accurately identifying the key features of a target object set improves the accuracy and efficiency of feature analysis, and better reflects the characteristics of the object set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116662664B_ABST
    Figure CN116662664B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of big data data processing, and discloses a key feature determination method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: obtaining a target object set and a preset feature pool; determining a control object set; performing non-parametric verification on the target object set and the control object set based on each feature in the preset feature pool to obtain a difference value of each feature; and determining a key feature of the target object set according to the difference value of each feature in the preset feature pool. The whole scheme determines the control object set as a reference of the target object set, then performs non-parametric verification on the two sets based on the features in the preset feature pool, can accurately determine the difference value of each feature in the two sets, and can determine the key feature belonging to the target object set based on the difference value of the feature. The determined key feature can accurately reflect the characteristics of the target object set.
Need to check novelty before this filing date? Find Prior Art

Description

Key feature determination methods, apparatus, computer equipment, media and program products Technical Field

[0001] This application relates to the field of big data data analysis technology, and in particular to a method, apparatus, computer equipment, storage medium and computer program product for determining key features. Background Technology

[0002] With the development of internet technology, a plethora of applications have emerged across various industries, greatly improving users' efficiency and bringing them immense convenience. However, because different users have different needs when using these applications, application providers need to provide the services users require to enhance the user experience.

[0003] Currently, the main approach is to analyze key user characteristics to identify multiple object sets, and then extract corresponding services from different object sets.

[0004] However, in the current process of determining key features of object sets, feature analysis is mainly performed manually, which results in low accuracy of key feature extraction. Summary of the Invention

[0005] Therefore, it is necessary to provide a precise method, apparatus, computer device, computer-readable storage medium, and computer program product for determining key features in response to the above-mentioned technical problems.

[0006] Firstly, this application provides a method for determining key features. The method includes:

[0007] Obtain a set of target objects and a preset feature pool, wherein the preset feature pool contains at least two object features;

[0008] Determine a set of reference objects, wherein the number of objects in the set of reference objects is the same as the number of objects in the target set;

[0009] Based on each feature in the preset feature pool, nonparametric verification is performed on the target object set and the control object set to obtain the difference value of each feature.

[0010] Based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0011] In one embodiment, the step of performing nonparametric verification on the target object set and the control object set based on each feature in the preset feature pool to obtain the difference value of each feature includes:

[0012] Construct a first feature vector and a second feature vector for each feature in the preset feature pool, wherein the first feature vector is the feature vector of each feature in the target object set, and the second feature vector is the feature vector of each feature in the reference object set;

[0013] Nonparametric verification is performed on the first feature vector and the second feature vector to obtain the difference value of each feature.

[0014] In one embodiment, determining the key features of the target object set based on the difference value of each feature in the preset feature pool includes:

[0015] Based on the comparison of the difference value of each feature in the preset feature pool, the features that satisfy the preset difference rule are determined, and the key feature set is obtained.

[0016] Factor analysis was performed on the set of key features to obtain the factor analysis results;

[0017] Based on the factor analysis results, the set of key features is divided into multiple feature subsets;

[0018] Based on multiple subsets of features, the key features of the target object set are determined.

[0019] In one embodiment, the step of performing factor analysis on the set of key features to obtain factor analysis results includes:

[0020] Factor analysis was performed based on the set of key features to obtain the common factor contribution value and the variance contribution rate.

[0021] Based on the common factor contribution value and variance contribution rate, the number of key feature sets is determined.

[0022] In one embodiment, determining the set of control subjects includes:

[0023] Based on preset sampling rules and a preset number of objects, the object set is sampled a preset number of times to obtain a reference object set, which is the object set corresponding to the target business.

[0024] In one embodiment, the method further includes:

[0025] Frequency statistics are performed based on the key features of the target object set to obtain the first distribution frequency of the key features on the target object set and the second distribution frequency of the key features on the control object set;

[0026] Based on the first and second distribution frequencies, a key feature frequency statistics chart is generated and pushed.

[0027] Secondly, this application also provides a key feature determination apparatus. The apparatus includes:

[0028] The acquisition module is used to acquire a set of target objects and a preset feature pool, wherein the preset feature pool contains at least two object features;

[0029] The reference object determination module is used to determine a reference object set, wherein the number of objects in the reference object set is consistent with the number of objects in the target object set;

[0030] The verification module is used to perform nonparametric verification on the target object set and the control object set based on each feature in the preset feature pool, and to obtain the difference value of each feature.

[0031] The key feature determination module is used to determine the key features of the target object set based on the difference value of each feature in the preset feature pool.

[0032] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0033] Obtain a set of target objects and a preset feature pool, wherein the preset feature pool contains at least two object features;

[0034] Determine a set of reference objects, wherein the number of objects in the set of reference objects is the same as the number of objects in the target set;

[0035] Based on each feature in the preset feature pool, nonparametric verification is performed on the target object set and the control object set to obtain the difference value of each feature.

[0036] Based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0037] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0038] Obtain a set of target objects and a preset feature pool, wherein the preset feature pool contains at least two object features;

[0039] Determine a set of reference objects, wherein the number of objects in the set of reference objects is the same as the number of objects in the target set;

[0040] Based on each feature in the preset feature pool, nonparametric verification is performed on the target object set and the control object set to obtain the difference value of each feature.

[0041] Based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0042] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0043] Obtain a set of target objects and a preset feature pool, wherein the preset feature pool contains at least two object features;

[0044] Determine a set of reference objects, wherein the number of objects in the set of reference objects is the same as the number of objects in the target set;

[0045] Based on each feature in the preset feature pool, nonparametric verification is performed on the target object set and the control object set to obtain the difference value of each feature.

[0046] Based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0047] The aforementioned key feature determination method, apparatus, computer equipment, storage medium, and computer program product acquire a target object set and a preset feature pool, wherein the preset feature pool contains at least two object features; determine a reference object set, the number of objects in the reference object set being consistent with the number of objects in the target object set; perform non-parametric verification on the target object set and the reference object set based on each feature in the preset feature pool to obtain the difference value of each feature; and determine the key features of the target object set based on the difference value of each feature in the preset feature pool. The entire scheme, by determining a reference object set as a reference to the target object set, and then performing non-parametric verification on the two sets based on the features in the preset feature pool, can accurately determine the difference value of each feature in the two sets. Furthermore, based on the feature difference value, the key features belonging to the target object set can be determined, and the determined key features can accurately reflect the characteristics of the target object set. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 is an application environment diagram of a key feature determination method in one embodiment;

[0050] Figure 2 is a flowchart illustrating a key feature determination method in one embodiment;

[0051] Figure 3 is a flowchart illustrating the key feature determination method in another embodiment;

[0052] Figure 4 is a flowchart illustrating the key feature determination method in another embodiment;

[0053] Figure 5 is a structural block diagram of a key feature determination device in one embodiment;

[0054] Figure 6 is an internal structure diagram of a computer device in one embodiment. Detailed Implementation

[0055] It should be noted that the user information (including but not limited to user device information and user personal information during the key feature extraction process) and data (including but not limited to data used for analysis, stored data, and displayed data during the key feature extraction process) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws and regulations of the relevant countries and regions.

[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0057] A customer group, or collection of customers, typically shares certain characteristics, such as a male customer group or a customer group whose registered address is in Beijing. Extracting key features of a customer group is crucial during application service iteration and updates; however, it suffers from low accuracy. What constitutes a key feature of a customer group, and how to find it, are the two core aspects of key feature extraction.

[0058] First, what are the key characteristics of a customer group? For example, in a clothing store in a shopping mall, 70% of the daily customer traffic is female and 30% is male. Since the majority of visitors to this clothing store are female, is "female" the key characteristic of its visitors? This question needs further examination. If 80% of the mall's visitors are female and only 20% are male, then comparing the clothing store's visitor group with the mall's visitor group, it seems that "female" cannot be considered a key characteristic.

[0059] Key characteristics are a relative concept. To identify the key characteristics of a customer group, it is essential to compare it with the overall customer base to determine the significant features that the customer group possesses relative to the total customer base. For example, if a company has 1,000 employees, 999 of whom have bachelor's degrees, and only one employee has a doctorate, then this doctoral employee can be considered a customer group (this customer group is very small, consisting of only one person), and "doctorate" would be its key characteristic.

[0060] Therefore, a customer group can have multiple characteristics, but the key characteristic is that the group is significantly different from the general population.

[0061] So how do we determine the key characteristics of the target customer group (the set of target objects)? That is, given a set of candidate features, how do we identify the features that differ significantly from the overall population? This application, based on the idea of ​​comparison, first determines a "feature pool" for identifying key features. The "feature pool" is a set of candidate features, such as 100 tags including "age," "gender," "time elapsed since the customer's first account opening," and "customer's investment and financial preferences." The feature pool is usually pre-defined according to the business application scenario. Key features are determined by comparing the characteristics of the target customer group with those of the overall customer base.

[0062] The key feature determination method provided in this application embodiment can be applied to the application environment shown in Figure 1. In this environment, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on other network servers.

[0063] In one possible implementation, terminal 102 responds to a key feature determination instruction, acquires a target object set and a preset feature pool, wherein the target object set and the preset feature pool are stored in server 104. The preset feature pool contains at least two object features; a reference object set is determined, the number of objects in the reference object set being consistent with the number of objects in the target object set; based on each feature in the preset feature pool, a non-parametric verification is performed on the target object set and the reference object set to obtain the difference value of each feature; and based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0064] In another possible implementation, server 104 acquires a target object set and a preset feature pool, the preset feature pool containing at least two object features; determines a reference object set, the number of objects in the reference object set being the same as the number of objects in the target object set; performs non-parametric validation on the target object set and the reference object set based on each feature in the preset feature pool, obtaining the difference value of each feature; and determines the key features of the target object set based on the difference value of each feature in the preset feature pool. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0065] In one embodiment, as shown in Figure 2, a key feature determination method is provided. Taking the application of this method to terminal 102 in Figure 1 as an example, the method includes the following steps:

[0066] Step 200: Obtain the target object set and the preset feature pool.

[0067] The target object set refers to the group from which key features need to be extracted. The target object set can be a set of users from different fields and different service objects. For example, it can be a set of users of a certain business in the financial field, or a set of users of a certain store in a shopping mall. Different target object sets can be determined according to the feature analysis requirements.

[0068] The preset feature pool contains features related to the target object in different dimensions, including basic object features, consumption features, and deposit features. Basic object features may include fundamental attributes such as gender, age, and education level. Consumption features may include consumption-related attributes such as consumption frequency, consumption type, and consumption amount. Deposit features may include deposit amount, deposit term, and number of deposits. The preset feature pool contains at least two object features. The preset feature pool can be obtained by the user selecting the feature pool corresponding to the specific business scenario from multiple feature pools, or by the user inputting multiple features according to the business scenario; this embodiment does not impose such limitations.

[0069] Specifically, the user selects the target object set and the preset feature pool corresponding to the target service on the terminal, triggering the key feature determination operation. The terminal listens for and responds to the key feature determination operation to obtain the target object set and the preset feature pool.

[0070] Step 400: Determine the set of reference objects.

[0071] The number of objects in the reference set is the same as the number of objects in the target set. The reference set is extracted from the overall customer base of the target business and serves as a reference for the target set, used for comparative analysis to determine key characteristics. The objects in the target set and the reference set can be partially the same or completely different.

[0072] Specifically, based on the target service, the terminal obtains all users of the target service, i.e. the total customers of the target service, and extracts users from all users whose number matches the number of objects in the target object set to obtain the reference object set.

[0073] Step 600: Based on each feature in the preset feature pool, perform nonparametric verification on the target object set and the control object set to obtain the difference value of each feature.

[0074] Nonparametric validation involves determining whether the population belongs to a certain theoretical distribution based on information from a set of samples, without knowing the overall population distribution; or whether two independent samples belong to the same distribution.

[0075] Specifically, the terminal performs distribution verification on each feature in the preset feature pool on the target object set and the control object set to determine whether the target object set and the control object set meet the theoretical distribution of the feature, and obtains the difference value of the feature on the target object set and the control object set.

[0076] Step 800: Determine the key features of the target object set based on the difference value of each feature in the preset feature pool.

[0077] Specifically, the terminal selects features that satisfy a preset difference rule based on the difference value of each feature in a preset feature pool, thereby obtaining the key features of the target object set. The preset difference rule can be a difference value threshold, where features with difference values ​​greater than the threshold are selected as key features.

[0078] The aforementioned method for determining key features involves: acquiring a target object set and a preset feature pool, where the preset feature pool contains at least two object features; determining a reference object set, where the number of objects in the reference object set is the same as the number of objects in the target object set; performing nonparametric verification on the target object set and the reference object set based on each feature in the preset feature pool to obtain the difference value of each feature; and determining the key features of the target object set based on the difference values ​​of each feature in the preset feature pool. The entire scheme accurately determines the difference value of each feature in the two sets by determining a reference object set with the target object set as a reference, and then performing nonparametric verification on the two sets based on the features in the preset feature pool. Based on these difference values, the key features belonging to the target object set can be determined, and the determined key features accurately reflect the characteristics of the target object set.

[0079] In some optional embodiments, as shown in Figure 3, based on each feature in a preset feature pool, a non-parametric verification is performed on the target object set and the control object set to obtain the difference value of each feature, including:

[0080] Step 620: Construct the first feature vector and the second feature vector for each feature in the preset feature pool.

[0081] Step 640: Perform nonparametric verification on the first feature vector and the second feature vector to obtain the difference value of each feature.

[0082] Wherein, the first feature vector is the feature vector of each feature in the target object set, and the second feature vector is the feature vector of each feature in the control object set.

[0083] Specifically, for each feature in the preset feature pool, the terminal constructs two feature vectors: a first feature vector of the feature on the target object set and a second feature vector of the feature on the reference object set.

[0084] Furthermore, for the preset feature pool F = {f1, f2, f3, ..., f...} n Each feature f in} i f i ∈F, construct two eigenvectors respectively:

[0085] First eigenvector V i ={v 1i ,v 2i ,v 3i ,…v ji ,…,v mi}, where i represents the i-th feature f in the feature pool F. i m represents the number of customers in the target customer group, v ji This indicates that the j-th customer in the target customer group has the following characteristics fi The value of V i For the target object set (target customer group), the feature f i The eigenvectors formed by these vectors;

[0086] Second eigenvector V′ i ={v′ 1i ,v′ 2i ,v′ 3i ,…,v′ ji ,…,v′ mi}, where i represents the i-th feature f in the feature pool F. i m represents the number of customers in the control group, v′ ji This indicates that the j-th customer in the comparison customer group has characteristics f. i The value of V i ′ represents the set of control subjects (control group), defined by feature f. i The eigenvectors formed by these vectors.

[0087] Next, the terminal uses the Mann-Whitney U test method to analyze the first feature vector V. i The second eigenvector V i Perform a two-sample nonparametric test to obtain the results of the two samples on feature f. i The statistical distribution formed above is used to obtain the significance level value p of the hypothesis test for this feature on the two sets. i ;

[0088] Finally, for each feature in the preset feature pool, perform the above nonparametric test to obtain the significance level value of the feature hypothesis test for each feature. Collect the significance level values ​​of the hypothesis tests for all features to obtain a list of validation results: P = {p1, p2, p3, ..., p...} j ,…,p n}, where p j This represents the significance level value of the hypothesis test for each feature in the feature set, where n is the total number of features.

[0089] In some optional embodiments, as shown in Figure 4, the key features of the target object set are determined based on the difference value of each feature in a preset feature pool, including:

[0090] Step 820: Based on the difference value of each feature in the preset feature pool, compare the differences to determine the features that satisfy the preset difference rules, and obtain the key feature set;

[0091] Step 840: Perform factor analysis on the set of key features to obtain the factor analysis results;

[0092] Step 860: Based on the factor analysis results, the key feature set is divided into multiple feature subsets;

[0093] Step 880: Determine the key features of the target object set based on multiple feature subsets.

[0094] The preset difference rule can also select the top k features with the largest differences. Factor analysis results include correlation analysis results.

[0095] Specifically, the terminal sorts each feature difference value in the preset feature pool in descending order, and then selects the top k features with the largest difference values ​​to obtain the key feature set.

[0096] Because the preset feature pool contains redundant features, such as the purchase frequency of the current month and the purchase frequency of the previous month, it is necessary to remove highly correlated features to reduce redundant key features. The terminal performs factor analysis on the key feature set to obtain correlation analysis results. Then, features with a correlation greater than a preset correlation threshold are divided into feature subsets, resulting in multiple feature subsets. The correlation between different feature subsets is weak, while the correlation within each feature subset is high. Finally, each feature subset is treated as a topic, and multiple feature subsets are used as key features of the target object.

[0097] This embodiment uses factor analysis to divide redundant features into a feature subset, thereby improving the accuracy of key feature analysis results.

[0098] In some optional embodiments, factor analysis is performed on the key feature set to obtain factor analysis results, including: performing factor analysis based on the key feature set to obtain common factor contribution values ​​and variance contribution rates; and determining the number of key feature sets based on the common factor contribution values ​​and variance contribution rates.

[0099] Specifically, the terminal calculates the common factor contribution value and variance contribution rate of the features in the key feature set. Then, based on the common factor contribution value and variance contribution rate, it draws a scree plot of the number of factors and variance contribution rate. Based on the elbow rule, it determines the inflection point of the scree plot of the number of factors and variance contribution rate. The number of factors corresponding to the inflection point is taken as the number of key feature sets. Based on the factor dimensionality reduction method, the key feature set is divided into multiple feature subsets corresponding to the number of key feature sets.

[0100] Furthermore, for the key feature set F′={f1′,f2′,f3′,…,f′ kThe common factor contribution value and variance contribution rate are calculated. Based on the scree plot of the number of factors and variance contribution rate, the key feature set includes k features, which can be divided into t key feature subsets. During feature subset partitioning, the factor loading matrix is ​​calculated, which is a k×t two-dimensional matrix. Each row in the loading matrix represents a feature, and each column represents a common factor. The value in each box is the factor loading value, representing the correlation coefficient between the original feature and the common factor. To visually display the correlation coefficient results, a background color is set for the boxes; the larger the factor loading value, the darker the background color of the box. The k key features are grouped into t subsets. For each common factor (each column in the figure), the original features with factor loading values ​​greater than 0.5 are assigned to the j-th feature subset, ultimately obtaining the key feature subsets. Each feature subset is considered a "topic": F′={F1′∪F2′∪…∪F j ′…∪F t ′}, where t is the number of key feature subsets, F j Let j represent the j-th feature subset, where j∈(1,t).

[0101] In some optional embodiments, determining the set of reference objects includes: sampling the set of objects a preset number of times based on preset sampling rules and preset number of objects to obtain the set of reference objects, wherein the set of objects is the set of objects corresponding to the target business.

[0102] The preset sampling rules include random sampling rules and fixed sampling rules. Random sampling rules involve randomly selecting customers from the total customer base as control subjects, with the number of control subjects matching the number of objects in the target object set. Fixed sampling rules can select customers at fixed locations, for example, selecting customers at odd-numbered locations from the total customer base until the number of control subjects reaches the number of objects in the target object set.

[0103] Specifically, the terminal performs random sampling on the total number of objects in the target object set based on random sampling rules to obtain a control object set.

[0104] In some optional embodiments, the above method further includes: performing frequency statistics based on key features of the target object set to obtain a first distribution frequency of key features on the target object set and a second distribution frequency of key features on the control object set; and generating and pushing a key feature frequency statistics chart based on the first distribution frequency and the second distribution frequency.

[0105] Specifically, the terminal calculates the distribution frequency of key features based on two sets, obtaining the statistical value of the distribution frequency of the target object set on each key feature, i.e., the first distribution frequency, and the statistical value of the distribution frequency of the control object set on each key feature, i.e., the second distribution frequency. The first distribution frequency of each key feature is aggregated to obtain a first distribution frequency list, and the second distribution frequency of each key feature is aggregated to obtain a second distribution frequency list. Finally, the first and second distribution frequency lists are combined to obtain a distribution frequency statistics list.

[0106] Furthermore, for each feature f′ in the feature subset F′ j Calculate the frequency distribution statistics lists for the target object set (target customer group) and the control object set (comparison customer group) respectively:

[0107] First distribution frequency list D j ={d 1j ,d 2j ,d 3j ,…,d ij ,…d pj}, where j represents the j-th feature, p represents the number of bins for feature j, and d ij D represents the frequency distribution value of feature j in the i-th bin. j This represents a list of frequency statistics of the target customer group on feature j.

[0108] Second distribution frequency list D′ j ={d′ 1j ,d′ 2j ,d′ 3j ,…,d′ ij ,…,d′ pj}, where j represents the j-th feature, p represents the number of bins for feature j, and d′ ij D′ represents the frequency distribution value of feature j in the i-th bin. j This represents a statistical list of the distribution frequency of the target customer group on feature j; based on the statistical list of distribution frequency, a frequency statistical chart of key features is drawn, that is, the frequency histogram of each key feature on the target object set and the control object set is obtained, and the frequency statistical chart of key features is pushed.

[0109] In this embodiment, key features are displayed intuitively using frequency histograms, improving the efficiency of user analysis of key features and the efficiency of generating analysis conclusions.

[0110] To facilitate understanding of the technical solutions provided in the embodiments of this application, the key feature determination method provided in the embodiments of this application will be briefly described with a complete key feature determination process:

[0111] (1) Obtain the target customer group G = {c1, c2, c3, ..., cm}, where m is the number of customers in the target customer group.

[0112] (2) Obtain the preset feature pool F = {f1, f2, f3, ..., f n}, where n represents the number of features.

[0113] (3) Randomly sample from the population to form a control group. The control group is represented as: G′={c′1,c′2,c′3,…,c′ m}, where m is the number of customers in the control group, which is consistent with the number of customers in the target group;

[0114] (4) For the feature pool F = {f1, f2, f3, ..., f n Each feature f in} i f i ∈F, construct two eigenvectors respectively:

[0115] V i ={v 1i ,v 2i ,v 3i ,…v ji ,…,v mi}, where i represents the i-th feature f in the feature pool F. i m represents the number of customers in the target customer group, v ji This indicates that the j-th customer in the target customer group has the following characteristics f i The value of V i For the target object set (target customer group), the feature f i The eigenvectors formed by these vectors;

[0116] V′ i ={v′ 1i ,v′ 2i ,v′ 3i ,…,v′ ji ,…,v′ mi}, where i represents the i-th feature f in the feature pool F. i m represents the number of customers in the control group, v′ ji This indicates that the j-th customer in the comparison customer group has characteristics f. i The value of V′ i For the set of control subjects (control group), the feature f i The eigenvectors formed by these vectors.

[0117] (5) Using the Mann-Whitney U test, the eigenvector V i V′ i Perform nonparametric tests on the two customer groups to obtain the results of feature f between the two customer groups. iThe statistical distribution formed above, and its hypothesis test significance level p. i For each feature in the feature pool, perform a nonparametric test to obtain a list: P = {p1, p2, p3, ..., p j ,…,p n}, where p j This represents the significance level value of the hypothesis test for each feature in the feature set, where n is the total number of features.

[0118] (6) For the difference measurement index P = {p1, p2, p3, ..., p4} obtained in step S4, j ,…,p n Sort the data by value from highest to lowest, and take the top K indicators to obtain the set of key features with significant differences: F′={f1′,f2′,f3′,…,f′ k}

[0119] (7) For the key feature set F′={f1′,f2′,f3′,…,f′ k} Calculate the common factor contribution value and variance contribution rate. Based on the scree plot of factor number - variance contribution rate, the key feature set includes k features, which can be divided into t key feature subsets.

[0120] (8) For each feature within the theme, perform frequency statistics description, draw frequency statistics chart, and push frequency histogram of the distribution difference between the target customer group and the control customer group on the key features.

[0121] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0122] Based on the same inventive concept, this application also provides a key feature determination apparatus for implementing the key feature determination method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more key feature determination apparatus embodiments provided below can be found in the limitations of the key feature determination method described above, and will not be repeated here.

[0123] In one embodiment, as shown in FIG5, a key feature determination device is provided, including: an acquisition module 502, a reference object determination module 504, a verification module 506, and a key feature determination module 508, wherein:

[0124] The acquisition module 502 is used to acquire a set of target objects and a preset feature pool, wherein the preset feature pool contains at least two object features.

[0125] The reference object determination module 504 is used to determine the reference object set, wherein the number of objects in the reference object set is consistent with the number of objects in the target object set;

[0126] The verification module 506 is used to perform nonparametric verification on the target object set and the control object set based on each feature in the preset feature pool, and to obtain the difference value of each feature.

[0127] The key feature determination module 508 is used to determine the key features of the target object set based on the difference value of each feature in the preset feature pool.

[0128] In an optional embodiment, the verification module 506 is further configured to construct a first feature vector and a second feature vector for each feature in a preset feature pool, wherein the first feature vector is the feature vector of each feature on the target object set, and the second feature vector is the feature vector of each feature on the reference object set; and to perform nonparametric verification on the first feature vector and the second feature vector to obtain the difference value of each feature.

[0129] In an optional embodiment, the key feature determination module 508 is further configured to compare the difference values ​​of each feature in the preset feature pool, determine the features that satisfy the preset difference rules, and obtain a key feature set; perform factor analysis on the key feature set to obtain factor analysis results; divide the key feature set into multiple feature subsets based on the factor analysis results; and determine the key features of the target object set based on the multiple feature subsets.

[0130] In an optional embodiment, the key feature determination module 508 is further configured to perform factor analysis based on the key feature set to obtain the common factor contribution value and the variance contribution rate; and to determine the number of key feature sets based on the common factor contribution value and the variance contribution rate.

[0131] In an optional embodiment, the reference object determination module 504 is further configured to sample the object set a preset number of times based on a preset sampling rule and a preset number of objects to obtain a reference object set, wherein the object set is the object set corresponding to the target business.

[0132] In an optional embodiment, the key feature determination device further includes a statistics module, which performs frequency statistics based on the key features of the target object set to obtain a first distribution frequency of the key features on the target object set and a second distribution frequency of the key features on the control object set; and generates and pushes a key feature frequency statistics chart based on the first distribution frequency and the second distribution frequency.

[0133] The modules in the aforementioned key feature determination device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0134] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as shown in Figure 6. The computer device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used for wired or wireless communication with an external terminal. Wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a key feature determination method. The display screen of the computer device may be a liquid crystal display (LCD) or an e-ink display. The input device of the computer device may be a touch layer covering the display screen, or buttons, a trackball, or a touchpad located on the casing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0135] Those skilled in the art will understand that the structure shown in Figure 6 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0136] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0137] Obtain the target object set and the preset feature pool. The preset feature pool contains at least two object features.

[0138] Determine the reference set, ensuring that the number of objects in the reference set is the same as the number of objects in the target set;

[0139] Based on each feature in the preset feature pool, nonparametric verification is performed on the target object set and the control object set to obtain the difference value of each feature;

[0140] Based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0141] In one embodiment, when the processor executes the computer program, it further performs the following steps: based on each feature in the preset feature pool, performs nonparametric verification on the target object set and the reference object set to obtain the difference value of each feature, including: constructing a first feature vector and a second feature vector for each feature in the preset feature pool, wherein the first feature vector is the feature vector of each feature on the target object set, and the second feature vector is the feature vector of each feature on the reference object set; and performs nonparametric verification on the first feature vector and the second feature vector to obtain the difference value of each feature.

[0142] In one embodiment, when the processor executes the computer program, it further implements the following steps: determining the key features of the target object set based on the difference value of each feature in the preset feature pool, including: comparing the difference values ​​of each feature in the preset feature pool to determine the features that satisfy the preset difference rules, and obtaining a key feature set; performing factor analysis on the key feature set to obtain factor analysis results; dividing the key feature set into multiple feature subsets based on the factor analysis results; and determining the key features of the target object set based on the multiple feature subsets.

[0143] In one embodiment, when the processor executes the computer program, it further performs the following steps: performing factor analysis on the key feature set to obtain factor analysis results, including: performing factor analysis based on the key feature set to obtain common factor contribution values ​​and variance contribution rates; and determining the number of key feature sets based on the common factor contribution values ​​and variance contribution rates.

[0144] In one embodiment, when the processor executes the computer program, it further implements the following steps: determining a set of reference objects, including: sampling the set of objects a preset number of times based on a preset sampling rule and a preset number of objects to obtain a set of reference objects, wherein the set of objects is the set of objects corresponding to the target business.

[0145] In one embodiment, when the processor executes the computer program, it further performs the following steps: performing frequency statistics based on key features of the target object set to obtain a first distribution frequency of the key features on the target object set and a second distribution frequency of the key features on the control object set; and generating and pushing a key feature frequency statistics chart based on the first distribution frequency and the second distribution frequency.

[0146] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0147] Obtain the target object set and the preset feature pool. The preset feature pool contains at least two object features.

[0148] Determine the reference set, ensuring that the number of objects in the reference set is the same as the number of objects in the target set;

[0149] Based on each feature in the preset feature pool, nonparametric verification is performed on the target object set and the control object set to obtain the difference value of each feature;

[0150] Based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0151] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: performing nonparametric verification on the target object set and the control object set based on each feature in the preset feature pool to obtain the difference value of each feature, including: constructing a first feature vector and a second feature vector for each feature in the preset feature pool, wherein the first feature vector is the feature vector of each feature on the target object set, and the second feature vector is the feature vector of each feature on the control object set; performing nonparametric verification on the first feature vector and the second feature vector to obtain the difference value of each feature.

[0152] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: determining the key features of the target object set based on the difference value of each feature in the preset feature pool, including: comparing the difference values ​​of each feature in the preset feature pool to determine the features that satisfy the preset difference rules, and obtaining a key feature set; performing factor analysis on the key feature set to obtain factor analysis results; dividing the key feature set into multiple feature subsets based on the factor analysis results; and determining the key features of the target object set based on the multiple feature subsets.

[0153] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: performing factor analysis on the key feature set to obtain factor analysis results, including: performing factor analysis based on the key feature set to obtain common factor contribution values ​​and variance contribution rates; and determining the number of key feature sets based on the common factor contribution values ​​and variance contribution rates.

[0154] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: determining a set of reference objects, including: sampling the set of objects a preset number of times based on a preset sampling rule and a preset number of objects to obtain a set of reference objects, wherein the set of objects is the set of objects corresponding to the target business.

[0155] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: performing frequency statistics based on the key features of the target object set to obtain a first distribution frequency of the key features on the target object set and a second distribution frequency of the key features on the control object set; and generating and pushing a key feature frequency statistics chart based on the first distribution frequency and the second distribution frequency.

[0156] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0157] Obtain the target object set and the preset feature pool. The preset feature pool contains at least two object features.

[0158] Determine the reference set, ensuring that the number of objects in the reference set is the same as the number of objects in the target set;

[0159] Based on each feature in the preset feature pool, nonparametric verification is performed on the target object set and the control object set to obtain the difference value of each feature;

[0160] Based on the difference value of each feature in the preset feature pool, the key features of the target object set are determined.

[0161] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: performing nonparametric verification on the target object set and the control object set based on each feature in the preset feature pool to obtain the difference value of each feature, including: constructing a first feature vector and a second feature vector for each feature in the preset feature pool, wherein the first feature vector is the feature vector of each feature on the target object set, and the second feature vector is the feature vector of each feature on the control object set; performing nonparametric verification on the first feature vector and the second feature vector to obtain the difference value of each feature.

[0162] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: determining the key features of the target object set based on the difference value of each feature in the preset feature pool, including: comparing the difference values ​​of each feature in the preset feature pool to determine the features that satisfy the preset difference rules, and obtaining a key feature set; performing factor analysis on the key feature set to obtain factor analysis results; dividing the key feature set into multiple feature subsets based on the factor analysis results; and determining the key features of the target object set based on the multiple feature subsets.

[0163] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: performing factor analysis on the key feature set to obtain factor analysis results, including: performing factor analysis based on the key feature set to obtain common factor contribution values ​​and variance contribution rates; and determining the number of key feature sets based on the common factor contribution values ​​and variance contribution rates.

[0164] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: determining a set of reference objects, including: sampling the set of objects a preset number of times based on a preset sampling rule and a preset number of objects to obtain a set of reference objects, wherein the set of objects is the set of objects corresponding to the target business.

[0165] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: performing frequency statistics based on the key features of the target object set to obtain a first distribution frequency of the key features on the target object set and a second distribution frequency of the key features on the control object set; and generating and pushing a key feature frequency statistics chart based on the first distribution frequency and the second distribution frequency.

[0166] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0167] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0168] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for determining key features, characterized in that, Applied to a terminal or server, the method includes: acquiring a target object set and a preset feature pool, wherein the preset feature pool contains at least two object features, the target object set is a user set corresponding to a target business, and the preset feature pool includes basic object features, object consumption features, and object deposit features; determining a reference object set, wherein the number of objects in the reference object set is consistent with the number of objects in the target object set, and the reference object set is extracted from the total customers of the target business; performing non-parametric validation on the target object set and the reference object set based on each feature in the preset feature pool to obtain the difference value of each feature; determining key features of the target object set based on the difference value of each feature in the preset feature pool, wherein the key features are used to reflect the characteristics of the target object set relative to the total customers; and performing frequency statistics based on the key features of the target object set to obtain a first distribution frequency of the key features on the target object set and a second distribution frequency of the key features on the reference object set. Based on the first and second distribution frequencies, a key feature frequency statistical chart is generated and pushed, wherein the key feature frequency statistical chart displays the difference in distribution frequency of the key features in the target object set and the control object set in the form of a histogram; wherein, determining the key features of the target object set based on the difference value of each feature in the preset feature pool includes: comparing the difference values ​​of each feature in the preset feature pool to determine features that satisfy preset difference rules, thereby obtaining a key feature set; performing factor analysis on the key feature set to obtain factor analysis results; dividing the key feature set into multiple feature subsets based on the factor analysis results; determining the key features of the target object set based on the multiple feature subsets; wherein, performing factor analysis on the key feature set to obtain factor analysis results includes: performing factor analysis on the key feature set to obtain common factor contribution values ​​and variance contribution rates; determining the number of key feature sets based on the common factor contribution values ​​and variance contribution rates.

2. The method according to claim 1, characterized in that, The step of performing nonparametric verification on the target object set and the control object set based on each feature in the preset feature pool to obtain the difference value of each feature includes: constructing a first feature vector and a second feature vector for each feature in the preset feature pool, wherein the first feature vector is the feature vector of each feature on the target object set, and the second feature vector is the feature vector of each feature on the control object set; and performing nonparametric verification on the first feature vector and the second feature vector to obtain the difference value of each feature.

3. The method according to any one of claims 1-2, characterized in that, The determination of the reference object set includes: sampling the object set a preset number of times based on preset sampling rules and preset object number to obtain the reference object set, wherein the object set is the object set corresponding to the target business.

4. A key feature determination device, characterized in that, The device, applied to a terminal or server, includes: an acquisition module for acquiring a target object set and a preset feature pool, wherein the preset feature pool contains at least two object features, the target object set is a user set corresponding to a target business, and the preset feature pool includes basic object features, object consumption features, and object deposit features; a reference object determination module for determining a reference object set, wherein the number of objects in the reference object set is consistent with the number of objects in the target object set, and the reference object set is extracted from the total customers of the target business; a verification module for performing non-parametric verification on the target object set and the reference object set based on each feature in the preset feature pool to obtain the difference value of each feature; and a key feature determination module for determining key features of the target object set based on the difference value of each feature in the preset feature pool, wherein the key features reflect the characteristics of the target object set relative to the total customers; the key feature determination device further includes: frequency determination based on the key features of the target object set. The system statistically analyzes and obtains a first distribution frequency of the key feature in the target object set and a second distribution frequency of the key feature in the control object set. Based on the first and second distribution frequencies, a key feature frequency statistical chart is generated and pushed, wherein the key feature frequency statistical chart displays the difference in distribution frequency of the key feature in the target object set and the control object set in the form of a histogram. The key feature determination module is further configured to: compare the difference values ​​of each feature in the preset feature pool to determine the features that satisfy the preset difference rules, thereby obtaining a key feature set; perform factor analysis on the key feature set to obtain factor analysis results; divide the key feature set into multiple feature subsets based on the factor analysis results; determine the key features of the target object set based on the multiple feature subsets; and determine the key features of the target object set based on the multiple feature subsets. The key feature determination module is further configured to: perform factor analysis on the key feature set to obtain the common factor contribution value and variance contribution rate; and determine the number of key feature sets based on the common factor contribution value and variance contribution rate.

5. The apparatus according to claim 4, characterized in that, The verification module is further configured to: construct a first feature vector and a second feature vector for each feature in the preset feature pool, wherein the first feature vector is the feature vector of each feature in the target object set, and the second feature vector is the feature vector of each feature in the reference object set; and perform nonparametric verification on the first feature vector and the second feature vector to obtain the difference value of each feature.

6. The apparatus according to any one of claims 4-5, wherein the reference object determination module is further configured to: sample the object set a preset number of times based on a preset sampling rule and a preset number of objects to obtain a reference object set, wherein the object set is the object set corresponding to the target business.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Important feature determination method and device, equipment and storage medium

    CN110796492A

  • Target user positioning method and device and related equipment

    CN113570404A