Customer data dimension reduction method and device

By using the combination methods of sample pairing, intra-class graphs, inter-class graphs, and sparse self-representation models in customer data feature selection, the limitations of small-marked samples and data graph structure utilization are solved, and efficient and accurate customer data feature selection is achieved.

CN120086559APending Publication Date: 2025-06-03中国工商银行股份有限公司湖南省分行
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210184.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When existing customer data feature selection technology faces the scarcity of labeled samples and the limitations of data graph structure utilization, it is difficult to effectively solve the problem of small labeled samples, and it is difficult to accurately capture the distribution information of complex data.

Method used

By pairing multiple samples in the data sample set, the intra-class and inter-class graphs are determined, and the target model is constructed by combining the sparse self-representation model, the importance indicators of each feature contained in the sample are determined, and the target characteristics whose importance indicators meet the conditions are selected to construct a subset of target features.

Benefits of technology

It realizes effective customer data feature selection in small label sample scenarios, makes full use of label information, can better capture the distribution information of the data and the manifold structure of complex data, and improves the efficiency and accuracy of feature selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086559A_ABST
    Figure CN120086559A_ABST
Patent Text Reader

Abstract

The invention discloses a customer data dimension reduction method and device, and the method comprises the steps: carrying out the pairing of a plurality of samples in a data sample set, and obtaining a plurality of sample pairs; determining intra-class graphs and inter-class graphs based on the plurality of sample pairs; based on the intra-class graph and the inter-class graph, combining a sparse self-representation model of the data sample set to construct a target model; based on the optimal value of the sparse coefficient matrix in the target model, determining the importance index of each feature contained in the sample; selecting target features with importance indexes conforming to conditions from the features contained in the samples, and constructing a target feature subset; according to the method, corresponding intra-class diagrams and inter-class diagrams are established according to similarity and dissimilarity between samples, a sparse self-representation model of a data sample set is transformed by utilizing the intra-class diagrams and the inter-class diagrams to obtain a target model, and finally feature selection of customer data is effectively completed by utilizing the target model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a method and device for reducing the dimension of customer data. Background Art

[0002] Customer data is the basis for banks to understand customers, provide personalized services, evaluate credit risks, and design financial products. By analyzing customer data, credit scores and risk assessments of customers can be made more accurately, so as to provide financial products and services that better meet customer needs. With the development of big data technology, the amount of data that banks can collect and process has increased significantly. While providing rich information, it also brings problems such as a huge amount of data and a high degree of information redundancy, posing higher requirements for data processing. Banks can use technologies such as data mining and machine learning to reduce the data dimension, so as to extract valuable information from the vast amount of data, improve the efficiency and accuracy of data processing, especially in the fields of customer data analysis and precision marketing. For example, it can increase the classification accuracy during customer segmentation and better provide personalized services for customers.

[0003] Currently, customer data dimension reduction technologies can be divided into two categories: feature extraction and feature selection. Compared with feature extraction, feature selection is usually more efficient. Without complex transformations and constructing new feature spaces, it directly selects subsets from the original features, retains the interpretability of the data, and is more easily directly associated with business problems. Most existing feature selection models are supervised and unsupervised models. Supervised feature selection technologies, such as random forests and model-based feature selection, identify key features by evaluating the contribution of features to the model's prediction ability. Unsupervised feature selection technologies, such as principal component analysis and correlation-based feature selection, do not rely on label information, but identify important features by analyzing the statistical relationships between features. In customer data analysis, feature selection technologies not only improve the efficiency of data processing and the prediction accuracy of the model, but also enhance the bank's understanding of customer behavior, providing more targeted marketing and service strategies for the bank.

[0004] In customer data feature selection, there is usually a problem of scarce labeled samples and abundant unlabeled samples. Supervised methods rely on a large amount of labeled data to train the model, while unsupervised methods may not be able to fully utilize the label information in the data. Both cannot well solve the "small labeled sample" problem. Moreover, the current data feature selection technologies have certain limitations in the utilization of data graph structures. For example, they often rely on a single graph regularization method and cannot fully capture the distribution information of samples. And most are only applicable to the Euclidean domain. In the Euclidean space, the distance and similarity of data are defined by simple geometric relationships, making it difficult to accurately represent relatively complex data.

[0005] How to achieve effective customer data feature selection has become a research hotspot in the banking and finance industry. Summary of the Invention

[0006] This application provides a method and device for dimensionality reduction of customer data, aiming to achieve effective selection of customer data features.

[0007] To achieve the above object, this application provides the following technical solutions:

[0008] A method for dimensionality reduction of customer data includes:

[0009] Pairing multiple samples in a data sample set to obtain multiple sample pairs; the samples are obtained by preprocessing based on corresponding customer data;

[0010] Based on multiple sample pairs, determine an intra-class graph and an inter-class graph; each element value in the matrix corresponding to the intra-class graph represents the similarity between two samples in the corresponding sample pair; each element value in the matrix corresponding to the inter-class graph represents the dissimilarity between two samples in the corresponding sample pair;

[0011] Based on the intra-class graph and the inter-class graph, combined with the sparse self-representation model of the data sample set, construct a target model;

[0012] Based on the optimal value of the sparse coefficient matrix in the target model, determine the importance index of each feature included in the sample;

[0013] Select target features whose importance indexes meet the conditions from each feature included in the sample, and construct a target feature subset.

[0014] Optionally, based on multiple sample pairs, determining an intra-class graph and an inter-class graph includes:

[0015] Calculate the similarity of multiple sample pairs based on a similarity algorithm;

[0016] Based on the similarity of multiple sample pairs, establish a sample similarity matrix;

[0017] Calculate the dissimilarity of multiple sample pairs based on an asymmetric metric algorithm;

[0018] Based on the dissimilarity of multiple sample pairs, establish a sample dissimilarity matrix;

[0019] Normalize the sample similarity matrix and the sample dissimilarity matrix;

[0020] Based on the label information carried by some samples in the data sample set, correct the normalized sample similarity matrix to obtain an intra-class graph;

[0021] Based on the label information carried by the partial samples, correct the dissimilarity matrix between the normalized samples to obtain an inter-class graph.

[0022] Optionally, based on the label information carried by partial samples in the data sample set, correct the similarity matrix between the normalized samples to obtain an intra-class graph, including:

[0023] Determine the type of the label information carried by partial samples in the data sample set;

[0024] According to the first preset step, correct each element in the similarity matrix between the normalized samples to obtain an intra-class graph; the first preset step is: if the two samples corresponding to the element are both the partial samples and the types of the label information carried by the two samples are the same, update the value of the element to a first numerical value; if the two samples corresponding to the element are both the partial samples and the types of the label information carried by the two samples are different, update the value of the element to a second numerical value.

[0025] Optionally, based on the label information carried by the partial samples, correct the dissimilarity matrix between the normalized samples to obtain an inter-class graph, including:

[0026] Determine the type of the label information carried by the partial samples;

[0027] According to the second preset step, correct each element in the dissimilarity matrix between the normalized samples to obtain an inter-class graph; the second preset step is: if the two samples corresponding to the element are both the partial samples and the types of the label information carried by the two samples are the same, update the value of the element to a second numerical value; if the two samples corresponding to the element are both the partial samples and the types of the label information carried by the two samples are different, update the value of the element to a first numerical value.

[0028] Optionally, based on the intra-class graph and the inter-class graph, combine with the sparse self-representation model of the data sample set to construct a target model, including:

[0029] Determine the sparse self-representation model of the data sample set; the sparse self-representation model includes a sparse coefficient matrix and a dictionary matrix constructed based on multiple samples; the value of each element in the sparse coefficient matrix represents the linear representation relationship between two features in the corresponding feature pair; each element in the dictionary matrix represents each feature of the sample;

[0030] On the basis of the sparse self-representation model, use the intra-class graph to weight the dictionary matrix to obtain a graph sparse self-representation model;

[0031] Determine a corresponding graph regularization term based on the within-class graph and the between-class graph;

[0032] Add the graph regularization term to the graph sparse self-representation model to obtain a target model.

[0033] Optionally, based on the optimal value of the sparse coefficient matrix in the target model, determine the importance index of each feature included in the sample, including:

[0034] Solve the target model to obtain the optimal value of the sparse coefficient matrix;

[0035] Based on the Euclidean norm of each row vector in the optimal value of the sparse coefficient matrix, determine the importance index of each feature included in the sample.

[0036] A customer data dimensionality reduction device, comprising:

[0037] A sample pairing unit, configured to pair multiple samples in a data sample set to obtain multiple sample pairs; the samples are obtained by preprocessing based on corresponding customer data;

[0038] A matrix determination unit, configured to determine a within-class graph and a between-class graph based on multiple sample pairs; the value of each element in the matrix corresponding to the within-class graph represents the similarity between two samples in the corresponding sample pair; the value of each element in the matrix corresponding to the between-class graph represents the dissimilarity between two samples in the corresponding sample pair;

[0039] A model construction unit, configured to construct a target model based on the within-class graph and the between-class graph, in combination with the sparse self-representation model of the data sample set;

[0040] An index determination unit, configured to determine the importance index of each feature included in the sample based on the optimal value of the sparse coefficient matrix in the target model;

[0041] A feature selection unit, configured to select target features whose importance index meets the conditions from each feature included in the sample, and construct a target feature subset.

[0042] A storage medium, the storage medium includes a stored program, wherein the program, when executed by a processor, executes the customer data dimensionality reduction method.

[0043] An electronic device, comprising: a processor, a memory, and a bus; the processor is connected to the memory through the bus;

[0044] The memory is used to store a program, and the processor is used to run the program, wherein the program, when executed by the processor, executes the customer data dimensionality reduction method.

[0045] The technical solution provided by this application pairs multiple samples in a data sample set to obtain multiple sample pairs. Based on the multiple sample pairs, an intra-class graph and an inter-class graph are determined. Based on the intra-class graph and the inter-class graph, combined with the sparse self-representation model of the data sample set, a target model is constructed. Based on the optimal value of the sparse coefficient matrix in the target model, the importance index of each feature included in the sample is determined. From each feature included in the sample, target features whose importance index meets the conditions are selected to construct a target feature subset. This application establishes corresponding intra-class and inter-class graphs according to the similarity and dissimilarity between samples, and uses the intra-class and inter-class graphs to transform the sparse self-representation model of the data sample set to obtain a target model. Finally, the target model is used to effectively complete the feature selection of customer data. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a schematic flowchart of a method for reducing the dimension of customer data provided by an embodiment of this application;

[0048] Figure 2 It is a schematic flowchart of another method for reducing the dimension of customer data provided by an embodiment of this application;

[0049] Figure 3 It is a schematic flowchart of another method for reducing the dimension of customer data provided by an embodiment of this application;

[0050] Figure 4 It is a schematic architecture diagram of a device for reducing the dimension of customer data provided by an embodiment of this application. Detailed Embodiments

[0051] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0052] In this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0053] As Figure 1 shown, it is a schematic flowchart of a method for dimensionality reduction of customer data provided by an embodiment of this application, including the following steps.

[0054] S101: Pair multiple samples in a data sample set to obtain multiple sample pairs.

[0055] Among them, the samples are obtained by preprocessing based on the corresponding customer data.

[0056] In some examples, the customer data collected from a relevant business system can be preprocessed to obtain corresponding samples. The samples usually contain multiple features, and the multiple features can be determined based on various attributes in the customer data, such as the basic information of the customer, transaction records, account balance, credit score, etc.

[0057] In possible implementation manners, the preprocessing process of customer data includes but is not limited to operations such as filling missing values, identifying outliers, removing duplicate values, data variable transformation, and data standardization. Generally speaking, the missing values in customer data can be processed by filling (such as filling with mean or median) or deleting the records with missing values. The outliers in customer data can be identified and removed by statistical methods (such as the Z-score method) or based on business rules. The duplicate values in customer data can be de-duplicated. In addition, the categorical feature variables in customer data need to be converted into numerical variables, and common methods include one-hot encoding, label encoding, etc. After determining each numerical data in the customer data, it is also necessary to standardize or normalize each numerical data so that each numerical data has a unified quantization scale.

[0058] In some examples, the data sample set can be expressed as , represents the total number of samples, represents the total number of features, and the th sample can be expressed as , Representative sample The value of the th feature, represents a matrix of

[0059] S102: Determine the within-class graph and the between-class graph based on multiple sample pairs.

[0060] Among them, the value of each element in the matrix corresponding to the within-class graph represents the similarity between two samples in the corresponding sample pair, and the value of each element in the matrix corresponding to the between-class graph represents the dissimilarity between two samples in the corresponding sample pair.

[0061] It should be noted that the within-class graph and the between-class graph jointly reveal the manifold structure of customer data. Generally speaking, a graph is a data structure and model, and the simplest and most effective way to store a graph in a computer is a matrix.

[0062] In some examples, similarity is a numerical measure of the degree of similarity between two samples. Usually, similarity is non-negative and takes values in the range [0, 1]. Dissimilarity is a numerical measure of the degree of difference between two samples. Usually, dissimilarity is non-negative and takes values in the range [0, 1].

[0063] Optionally, for the implementation process of determining the within-class graph and the between-class graph based on multiple sample pairs, reference can be made to Figure 2 the steps shown and the corresponding explanatory notes.

[0064] S103: Based on the within-class graph and the between-class graph, combined with the sparse self-representation model of the data sample set, construct a target model.

[0065] Among them, the target model can be regarded as a model obtained by transforming the sparse self-representation model of the data sample set using the within-class graph and the between-class graph.

[0066] Optionally, for the implementation process of constructing the target model based on the within-class graph and the between-class graph, combined with the sparse self-representation model of the data sample set, reference can be made to Figure 3 the steps shown and the corresponding explanatory notes.

[0067] S104: Based on the optimal value of the sparse coefficient matrix in the target model, determine the importance index of each feature included in the sample.

[0068] Among them, after determining the target model, appropriate evaluation metrics (such as accuracy, precision, recall, F1 score, etc.) can be selected as the criteria for measuring the performance of the target model to determine whether the target model determined based on multiple samples reaches the best performance.

[0069] In some examples, the k-fold cross-validation method can be used to implement the steps shown in S102 - S104 on the data training set to obtain the optimal value of the sparse coefficient matrix. In addition, according to the evaluation metrics, the preset threshold can be adjusted for the feature subset, such as increasing the number of features of the sample or reducing the number of features of the sample, and re-performing feature selection and evaluation until the target model reaches the best performance.

[0070] Optionally, the implementation process of determining the importance index of each feature included in the sample based on the optimal value of the sparse coefficient matrix in the target model can be: solving the target model to obtain the optimal value of the sparse coefficient matrix; determining the importance index of each feature included in the sample based on the Euclidean norm of each row vector in the optimal value of the sparse coefficient matrix.

[0071] It should be noted that the row vector is composed of the values of each element in the same row of the sparse coefficient matrix. Each row vector corresponds to a feature respectively. The Euclidean norm of the row vector can be regarded as the importance index of the corresponding feature, and the importance index is used to characterize the importance of the feature.

[0072] S105: Select target features whose importance indices meet the conditions from each feature included in the sample to construct a target feature subset.

[0073] Among them, after determining the importance index of each feature, an effective customer data feature selection can be achieved by selecting target features whose importance indices meet the conditions to construct a target feature subset.

[0074] In a possible implementation manner, the target model can be integrated into the business platform to support the customer data feature dimensionality reduction function of the business platform.

[0075] Compared with supervised models and unsupervised models, semi-supervised models (i.e., target models) can better handle the scenario of small labeled samples and make full use of label information. In data dimensionality reduction technology, sparse self-representation is a method that can effectively solve the problem of complex and redundant data. This theory shows that any feature of the original data can be expressed as a linear combination of representative features. By introducing an intra-class graph to weight the samples in the sparse self-representation model of the data sample set, the neighbor information of the sample nodes can be utilized, the distribution characteristics of the samples can be retained, and at the same time, the processing scope of the data can be extended to the non-Euclidean domain, and the sample distribution can be explored more effectively. In addition, adding a graph regularization term to the graph sparse self-representation model can more accurately reveal the local manifold structure of customer data and maintain the intra-class similarity and inter-class dissimilarity.

[0076] According to the similarities and dissimilarities among samples, the within-class graph and between-class graph are established for the process shown in S101-S105 above, and the sparse self-representation model of the data sample set is transformed using the within-class graph and between-class graph to obtain the target model. Finally, the target model is used to effectively complete the feature selection of customer data.

[0077] As Figure 2 shown, it is a schematic flowchart of another customer data dimensionality reduction method provided by an embodiment of the present application, including the following steps.

[0078] S201: Calculate the similarity of multiple sample pairs based on the similarity algorithm.

[0079] Among them, there are multiple algorithms to calculate the similarity measure between two samples in a sample pair. The multiple similarity algorithms include but are not limited to the cosine similarity algorithm, Gaussian kernel weighted algorithm, etc.

[0080] In some examples, multiple similarity algorithms can be selected or comprehensively used according to data characteristics and business requirements to calculate, so as to find a similarity measure suitable for a specific data set and problem.

[0081] S202: Based on the similarity of multiple sample pairs, establish a similarity matrix between samples.

[0082] Among them, each element in the similarity matrix between samples represents the corresponding sample pair, and the value of each element represents the similarity of the corresponding sample pair.

[0083] S203: Calculate the dissimilarity of multiple sample pairs based on the asymmetric metric algorithm.

[0084] Among them, the asymmetric metric algorithm can be the KL divergence. In the process of using the KL divergence to calculate the dissimilarity between two samples in a sample pair, since the KL divergence is not symmetric, the symmetry of the dissimilarity measure needs to be achieved by taking the average of the KL divergences between two samples.

[0085] In some examples, the KL divergence is defined as , represents the th sample after normalization of the probability distribution, represents the th sample after normalization of the probability distribution.

[0086] S204: Based on the dissimilarity of multiple sample pairs, establish a dissimilarity matrix between samples.

[0087] Among them, each element in the dissimilarity matrix between samples represents the corresponding sample pair, and the value of each element represents the dissimilarity of the corresponding sample pair.

[0088] In some examples, the inter-sample dissimilarity matrix the element value at the row and the column is shown by formula (1).

[0089] (1)

[0090] S205: Normalize the inter-sample similarity matrix and the inter-sample dissimilarity matrix.

[0091] Among them, normalizing the inter-sample similarity matrix and the inter-sample dissimilarity matrix can align the data scale and ensure that a standard data matrix is obtained.

[0092] S206: Based on the label information carried by some samples in the data sample set, correct the normalized inter-sample similarity matrix to obtain an intra-class graph.

[0093] Among them, the part of the samples carrying label information in the data sample set can be understood as small labeled samples. Using the label information carried by some samples, the values of the elements in the normalized inter-sample similarity matrix can be corrected, thereby realizing the full utilization of limited label information.

[0094] Optionally, based on the label information carried by some samples in the data sample set, the process of correcting the normalized inter-sample similarity matrix to obtain an intra-class graph can be: determining the type of the label information carried by some samples in the data sample set; according to the first preset step, correcting each element in the normalized inter-sample similarity matrix to obtain an intra-class graph.

[0095] The first preset step is: if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are the same, then update the value of the element to the first numerical value; if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are different, then update the value of the element to the second numerical value.

[0096] In some examples, the first numerical value can be set to 1 and the second numerical value can be set to 0.

[0097] S207: Based on the label information carried by some samples, correct the normalized inter-sample dissimilarity matrix to obtain an inter-class graph.

[0098] Among them, using the label information carried by some samples, the values of the elements in the normalized inter-sample dissimilarity matrix can be corrected, thereby realizing the full utilization of limited label information.

[0099] Optionally, based on the label information carried by some samples, the dissimilarity matrix between the normalized samples is corrected to obtain the implementation process of the intra-class graph, which can be: determining the type of the label information carried by some samples; according to the second preset step, correcting each element in the dissimilarity matrix between the normalized samples to obtain the intra-class graph.

[0100] The second preset step is: if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are the same, then update the value of the element to the second value; if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are different, then update the value of the element to the first value.

[0101] The processes shown in S201-S207 above can determine the corresponding intra-class graph based on the similarity of multiple sample pairs and the label information of some samples, and determine the corresponding inter-class graph based on the dissimilarity of multiple sample pairs and the label information of some samples, so as to effectively construct the manifold structure of customer data.

[0102] As Figure 3 shown, it is a schematic flowchart of another customer data dimensionality reduction method provided by an embodiment of the present application, including the following steps.

[0103] S301: Determine the sparse self-representation model of the data sample set.

[0104] Among them, the sparse self-representation model includes a sparse coefficient matrix and a dictionary matrix constructed based on multiple samples. The value of each element in the sparse coefficient matrix represents the linear representation relationship between two features in the corresponding feature pair, and each element in the dictionary matrix represents each feature of the sample.

[0105] In some examples, the sparse self-representation model can be defined as shown in formula (2).

[0106] (2)

[0107] In formula (2), represents the dictionary matrix composed of multiple samples, represents the sparse coefficient matrix and , represents the regularization parameter. Since the samples have the characteristics of high redundancy and strong mutual correlation, the corresponding dictionary matrix is constructed based on multiple samples.

[0108] S302: On the basis of the sparse self-representation model, use the intra-class graph to weight the dictionary matrix to obtain the graph sparse self-representation model.

[0109] Among them, by using the within-class graph to weight the dictionary matrix, the within-class graph structure can be introduced on the basis of the sparse self-representation model to combine the geometric distribution information of the samples.

[0110] In some examples, the expression of the graph sparse self-representation model can be seen in Formula (3).

[0111] (3)

[0112] In Formula (3), represents the matrix corresponding to the within-class graph.

[0113] S303: Determine the corresponding graph regularization term based on the within-class graph and the between-class graph.

[0114] Among them, the graph Laplacian regularization method can be used to construct the regularization term of the samples. Considering the bi-graph structure of the samples (i.e., the within-class graph and the between-class graph), the corresponding graph regularization term can be constructed so that after the samples are represented by the sparse coefficient matrix, the original similarity and dissimilarity can still be maintained.

[0115] In some examples, the expression of the graph regularization term can be seen in Formula (4).

[0116] (4)

[0117] In Formula (4), represents the element value of the th row and the th column in the matrix corresponding to the within-class graph, represents the element value of the th row and the th column in the matrix corresponding to the between-class graph.

[0118] S304: Add the graph regularization term to the graph sparse self-representation model to obtain the target model.

[0119] Among them, the expression of the target model can be seen in Formula (5).

[0120]

[0121] In Formula (5), and are both regularization parameters.

[0122] The above process shown in S301-S304 can use the within-class graph and the between-class graph to transform the sparse self-representation model of the data sample set to obtain the target model.

[0123] Such as Figure 4As shown in the figure, it is a schematic architecture diagram of a customer data dimensionality reduction device provided by an embodiment of the present application, including the following units.

[0124] The sample pairing unit 100 is used to pair multiple samples in the data sample set to obtain multiple sample pairs; the samples are obtained by preprocessing based on the corresponding customer data.

[0125] The matrix determination unit 200 is used to determine the within-class graph and the between-class graph based on multiple sample pairs; the value of each element in the matrix corresponding to the within-class graph represents the similarity between the two samples in the corresponding sample pair; the value of each element in the matrix corresponding to the between-class graph represents the dissimilarity between the two samples in the corresponding sample pair.

[0126] Optionally, the matrix determination unit 200 is specifically used for: calculating the similarity of multiple sample pairs based on the similarity algorithm; establishing a sample similarity matrix based on the similarity of multiple sample pairs; calculating the dissimilarity of multiple sample pairs based on the asymmetric metric algorithm; establishing a sample dissimilarity matrix based on the dissimilarity of multiple sample pairs; normalizing the sample similarity matrix and the sample dissimilarity matrix; correcting the normalized sample similarity matrix based on the label information carried by some samples in the data sample set to obtain the within-class graph; correcting the normalized sample dissimilarity matrix based on the label information carried by some samples to obtain the between-class graph.

[0127] Optionally, the matrix determination unit 200 is specifically used for: determining the type of the label information carried by some samples in the data sample set; correcting each element in the normalized sample similarity matrix according to the first preset step to obtain the within-class graph; the first preset step is: if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are the same, then update the value of the element to the first value; if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are different, then update the value of the element to the second value.

[0128] Optionally, the matrix determination unit 200 is specifically used for: determining the type of the label information carried by some samples; correcting each element in the normalized sample dissimilarity matrix according to the second preset step to obtain the between-class graph; the second preset step is: if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are the same, then update the value of the element to the second value; if the two samples corresponding to the element are both some samples and the types of the label information carried by the two samples are different, then update the value of the element to the first value.

[0129] The model construction unit 300 is used to construct a target model based on the within-class graph and the between-class graph in combination with the sparse self-representation model of the data sample set.

[0130] Optionally, the model construction unit 300 is specifically configured to: determine a sparse self-representation model of the data sample set; the sparse self-representation model includes a sparse coefficient matrix and a dictionary matrix constructed based on multiple samples; the values of the elements in the sparse coefficient matrix represent the linear representation relationship between two features in the corresponding feature pair; each element in the dictionary matrix represents each feature of the sample; based on the sparse self-representation model, use the intra-class graph to weight the dictionary matrix to obtain a graph sparse self-representation model; determine the corresponding graph regularization term based on the intra-class graph and the inter-class graph; add the graph regularization term to the graph sparse self-representation model to obtain the target model.

[0131] The index determination unit 400 is configured to determine the importance index of each feature included in the sample based on the optimal value of the sparse coefficient matrix in the target model.

[0132] Optionally, the index determination unit 400 is specifically configured to: solve the target model to obtain the optimal value of the sparse coefficient matrix; determine the importance index of each feature included in the sample based on the Euclidean norm of each row vector in the optimal value of the sparse coefficient matrix.

[0133] The feature selection unit 500 is configured to select target features whose importance indexes meet the conditions from each feature included in the sample, and construct a target feature subset.

[0134] Each of the above units establishes corresponding intra-class graphs and inter-class graphs according to the similarity and dissimilarity between samples, and uses the intra-class graphs and inter-class graphs to transform the sparse self-representation model of the data sample set to obtain the target model, and finally uses the target model to effectively complete the feature selection of customer data.

[0135] The present application also provides a computer-readable storage medium, and the computer-readable storage medium includes a stored program, wherein the program executes the customer data dimensionality reduction method provided by the present application.

[0136] The present application also provides an electronic device, including: a processor, a memory, and a bus. The processor is connected to the memory through the bus, the memory is used to store a program, and the processor is used to run the program, wherein when the program runs, it executes the customer data dimensionality reduction method provided by the present application.

[0137] In addition, the functions described above in the embodiments of the present application can be at least partially executed by one or more hardware logic components. For example, without limitation, the exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0138] Although several specific implementation details are included in the above description, these should not be construed as limiting the scope of the present application. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0139] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.

Claims

1. A customer data dimensionality reduction method, characterized in that: include: Pair multiple samples in the data sample set to obtain multiple sample pairs; The sample is obtained by preprocessing the corresponding customer data; Based on the plurality of sample pairs, an intra-class graph and an inter-class graph are determined; the value of each element in the matrix corresponding to the intra-class graph represents the similarity between two samples in the corresponding sample pair; the value of each element in the matrix corresponding to the inter-class graph represents the dissimilarity between two samples in the corresponding sample pair; Based on the intra-class graph and the inter-class graph, and in combination with the sparse self-representation model of the data sample set, construct a target model; Determining the importance index of each feature included in the sample based on the optimal value of the sparse coefficient matrix in the target model; From the various features included in the sample, target features whose importance indicators meet the conditions are selected to construct a target feature subset.

2. The method according to claim 1, characterized in that Based on the plurality of sample pairs, determining an intra-class graph and an inter-class graph includes: Based on the similarity algorithm, calculate the similarity of multiple sample pairs; Based on the similarities of the plurality of sample pairs, establishing a similarity matrix between samples; Based on an asymmetric metric algorithm, calculating the dissimilarity of the plurality of sample pairs; Based on the dissimilarity of the plurality of sample pairs, establishing an inter-sample dissimilarity matrix; Normalizing the inter-sample similarity matrix and the inter-sample dissimilarity matrix; Based on the label information carried by some samples in the data sample set, the normalized similarity matrix between samples is corrected to obtain an intra-class graph; Based on the label information carried by the partial samples, the normalized inter-sample dissimilarity matrix is ​​modified to obtain an inter-class graph.

3. The method according to claim 2, characterized in that Based on the label information carried by some samples in the data sample set, the normalized similarity matrix between samples is modified to obtain an intra-class graph, including: Determining the type of label information carried by some samples in the data sample set; According to the first preset step, each element in the normalized similarity matrix between samples is corrected to obtain an intra-class graph; the first preset step is: if the two samples corresponding to the element are both the partial samples and the types of label information carried by the two samples are the same, then the value of the element is updated to a first value; if the two samples corresponding to the element are both the partial samples and the types of label information carried by the two samples are different, then the value of the element is updated to a second value.

4. The method according to claim 2, characterized in that: Based on the label information carried by the partial samples, the normalized inter-sample dissimilarity matrix is ​​modified to obtain an inter-class graph, including: Determining the type of label information carried by the portion of samples; According to the second preset step, each element in the normalized inter-sample dissimilarity matrix is ​​corrected to obtain an inter-class graph; the second preset step is: if the two samples corresponding to the element are both the partial samples and the types of label information carried by the two samples are the same, then the value of the element is updated to the second value; if the two samples corresponding to the element are both the partial samples and the types of label information carried by the two samples are different, then the value of the element is updated to the first value.

5. The method according to claim 1, characterized in that Based on the intra-class graph and the inter-class graph, and in combination with the sparse self-representation model of the data sample set, a target model is constructed, including: Determine a sparse self-representation model of the data sample set; the sparse self-representation model includes a sparse coefficient matrix and a dictionary matrix constructed based on multiple samples; the value of each element in the sparse coefficient matrix represents a linear representation relationship between two features in a corresponding feature pair; each element in the dictionary matrix represents each feature of the sample; Based on the sparse self-representation model, the dictionary matrix is ​​weighted by using the intra-class graph to obtain a graph sparse self-representation model; Based on the intra-class graph and the inter-class graph, determining a corresponding graph regularization term; The graph regularization term is added to the graph sparse self-representation model to obtain a target model.

6. The method according to claim 1, characterized in that Based on the optimal value of the sparse coefficient matrix in the target model, determining the importance index of each feature included in the sample, including: Solving the target model to obtain an optimal value of a sparse coefficient matrix; Based on the Euclidean norm of each row vector in the optimal value of the sparse coefficient matrix, an importance index of each feature included in the sample is determined.

7. A customer data dimension reduction device, characterized in that: include: A sample pairing unit, used to pair multiple samples in a data sample set to obtain multiple sample pairs; The sample is obtained by preprocessing the corresponding customer data; A matrix determination unit, used to determine an intra-class graph and an inter-class graph based on the plurality of sample pairs; the value of each element in the matrix corresponding to the intra-class graph represents the similarity between two samples in the corresponding sample pair; the value of each element in the matrix corresponding to the inter-class graph represents the dissimilarity between two samples in the corresponding sample pair; A model building unit, configured to build a target model based on the intra-class graph and the inter-class graph in combination with a sparse self-representation model of the data sample set; An index determination unit, used to determine the importance index of each feature included in the sample based on the optimal value of the sparse coefficient matrix in the target model; The feature selection unit is used to select target features whose importance index meets the conditions from the various features included in the sample and construct a target feature subset.

8. The device according to claim 7, characterized in that The matrix determination unit is specifically used for: Based on the similarity algorithm, calculate the similarity of multiple sample pairs; Based on the similarities of the plurality of sample pairs, establishing a similarity matrix between samples; Based on an asymmetric metric algorithm, calculating the dissimilarity of the plurality of sample pairs; Based on the dissimilarity of the plurality of sample pairs, establishing an inter-sample dissimilarity matrix; Normalizing the inter-sample similarity matrix and the inter-sample dissimilarity matrix; Based on the label information carried by some samples in the data sample set, the normalized similarity matrix between samples is corrected to obtain an intra-class graph; Based on the label information carried by the partial samples, the normalized inter-sample dissimilarity matrix is ​​modified to obtain an inter-class graph.

9. A storage medium, characterized in that: The storage medium includes a stored program, wherein the program, when executed by a processor, executes the customer data dimensionality reduction method according to any one of claims 1 to 6.

10. An electronic device, characterized in that: include: processor, memory, and bus; The processor is connected to the memory via the bus; The memory is used to store programs, and the processor is used to run programs, wherein the program, when run by the processor, executes the customer data dimensionality reduction method described in any one of claims 1-6.