Bank merchant value evaluation method and device based on big data, medium and equipment

By constructing a weighted Euclidean distance and density clustering method and integrating explicit and implicit features, the problems of parameter sensitivity and complex distribution in bank merchant value evaluation are solved, and intelligent classification and precise marketing of merchant value are achieved.

CN120725530APending Publication Date: 2025-09-30深圳优讯科技有限公司

Patent Information

Application Number
CN202510894990.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies cannot effectively deal with heterogeneous data, implicit features and data noise in the evaluation of bank customer and merchant value, and lack dynamic weight modeling capabilities, resulting in the inability to achieve intelligent identification and behavior prediction of high-value merchants.

Method used

A big data-based bank merchant value evaluation method is adopted. By integrating explicit and implicit features, a weighted Euclidean distance calculation feature matrix is ​​constructed. Combined with density clustering and neighborhood query, clustering parameters are automatically optimized to achieve intelligent classification of merchant value.

Benefits of technology

It improves the accuracy and business adaptability of merchant value evaluation, overcomes the parameter sensitivity and complex distribution problems of traditional algorithms in big data environments, and realizes precise marketing and risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725530A_ABST
    Figure CN120725530A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, in particular to a bank merchant value evaluation method and device based on big data, a medium and equipment, and the method comprises the steps: collecting bank customer original data based on the big data; missing values are filled through preprocessing, repeated data are removed, and classification variables are numeralized: binary variables are kept in an original state, multi-valued class variables are subjected to one-hot coding, and ordered class variables are subjected to label coding; performing Z-score standardization on all numerical value characteristics; the dominant and implicit indexes are fused to construct a feature vector, the implicit indexes are obtained through customer behavior data quantification, and differences between samples are measured through a weighted Euclidean distance; initial clustering is carried out on the features, samples are classified into the nearest center point cluster, and iteration updating is carried out until convergence; and performing density clustering in each initial cluster, determining a core point, a boundary point and a noise point by calculating an average distance and a nearest neighbor distance in the cluster, and performing repeated expansion to form a stable sub-cluster. And finally, combining the initial cluster with a density division result, if the initial cluster is not subdivided, reserving the original cluster, and if the initial cluster is subdivided, removing noise points by taking a sub-cluster as a criterion, thereby obtaining a final value label of the customer. According to the method, the interpretability and the calculation efficiency are considered while the model precision is ensured, and the technical capabilities of banks in the aspects of precision marketing, risk control management and merchant asset value mining are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a method, device, medium and equipment for evaluating the value of bank merchants based on big data. Background Art

[0002] With the rapid development of "Internet Plus" and big data technologies, the banking industry is gradually shifting to a data-driven intelligent management model for customer engagement and business innovation. Massive amounts of customer and merchant data are continuously accumulating through multiple online and offline channels. This data encompasses multiple dimensions, including basic customer attributes, transaction behaviors, asset status, risk preferences, and interactive feedback. These data exhibit typical "big data" characteristics, including large volumes, multiple dimensions, rapid updates, and complex structures. Unlocking the potential value of this information to implement precision marketing, risk control, and customer / merchant value discovery has become a core challenge in commercial banks' digital transformation.

[0003] Existing technologies for bank customer segmentation and merchant value assessment often rely on static rule-based models or single clustering methods, such as the traditional K-Means algorithm and simple rule-based classification. These methods are unable to effectively address heterogeneous data, hidden features, and data noise. Especially in the context of large-scale customer or merchant data, key technical bottlenecks hinder the implementation of existing methods include how to weight features, automatically determine cluster boundaries and structure, and extract business-oriented sub-categorization labels from clustering results.

[0004] In terms of merchant value evaluation, although some studies have introduced multi-source data for label construction and grading, most of them rely on manual rules and static weights, lack dynamic training and multi-model fusion mechanisms, and it is difficult to achieve intelligent identification and behavior prediction of high-value merchants.

[0005] Therefore, there is an urgent need for a customer or merchant value analysis method that integrates explicit and implicit features, has dynamic weight modeling capabilities, can automatically optimize clustering parameters, and is suitable for big data scenarios. Summary of the Invention

[0006] In view of the above technical problems, the present invention provides a bank merchant value evaluation method, device, medium and equipment based on big data, which takes into account interpretability and computational efficiency while ensuring model accuracy, thereby improving the bank's technical capabilities in precision marketing, risk management and merchant asset value mining.

[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0008] According to one aspect of the present invention, a bank merchant value evaluation method based on big data is proposed, characterized in that the method includes: Collecting an original data set of bank customers from big data, wherein the original data set includes explicit indicators and implicit indicators; Preprocessing the original dataset includes filling in missing values ​​and removing duplicate data, converting categorical variables in the original dataset into numerical form, maintaining binary features in the original dataset in binary form, using one-hot encoding for multi-valued categorical features, and using label encoding for ordered categorical features; and applying Z-score normalization to all numerical features to eliminate dimensionality effects; The explicit indicators and the implicit indicators are jointly constructed into a unified feature vector, wherein the implicit indicators are quantitatively represented by analyzing customer behavior data. The explicit indicators and the implicit indicators are combined to form a feature matrix, each feature in the feature matrix is ​​assigned a weight, and the distance between sample points is calculated using weighted Euclidean distance to reflect their importance; Taking the feature matrix as input, randomly select K sample points as initial cluster centers, calculate the weighted Euclidean distance from each sample point to each cluster center, assign the sample point to the cluster center with the smallest distance, recalculate the weighted mean of the samples within each cluster and update the cluster center, and repeat the assignment and update steps until the preset maximum number of iterations is reached or the cluster centers converge and remain unchanged, thus obtaining K initial clusters; For each of the initial clusters, a neighborhood radius and a density threshold are determined to perform density clustering. During the execution, the cluster with the most samples is selected, the average distance from each point in the cluster with the most samples to the cluster center is calculated, and the nearest neighbor distance of the sample point within the average distance is obtained; the point density within the cluster is calculated based on the average distance and the nearest neighbor distance; then a neighborhood query is performed on each sample point in the cluster. During the query, a set of neighborhood points whose distance from the point does not exceed the average distance is calculated; if the number of points in the neighborhood of a point is not less than the nearest neighbor distance, the point is marked as a core point; starting from each core point, all points in its neighborhood are grouped into the same cluster; points that fall within the neighborhood of the core point but do not reach the nearest neighbor distance are marked as boundary points and assigned to the cluster closest to the core point; the remaining points that are not included in any cluster are marked as noise; the neighborhood query and cluster expansion operations are repeated until the cluster structure is stable and the final cluster division is obtained; The initial cluster is merged with the cluster division to obtain the final value label of each sample. During the merging, if no new sub-cluster is subdivided within the initial cluster and no noise is marked, the corresponding initial cluster is retained as the final cluster; if multiple sub-clusters are identified within the initial cluster, these sub-clusters are used as new final clusters; for the part marked as noise, it is removed from the initial cluster.

[0009] Furthermore, the preprocessing specifically includes: Fill missing values ​​with the mean of the feature column; For duplicate records, only one is retained; Encode non-numeric categorical features, such as using one-hot encoding for multi-valued categorical features, label encoding for ordered categorical features, and keeping the original 0 / 1 form for binary categorical features.

[0010] Furthermore, the explicit indicators include basic attribute characteristics of the customer, which include operating time, asset balance, and industry type; the implicit indicators include preference information extracted from the customer's historical behavior or feedback data, which includes the customer's interest score in a specific financial product, and the explicit indicators and the implicit indicators are formed into a unified feature vector.

[0011] Furthermore, the Z-score normalization process includes subtracting the mean from the eigenvalue and dividing it by the standard deviation to eliminate the influence of different dimensions, and the normalized eigenvalues ​​are used to form the feature matrix.

[0012] Furthermore, after the preprocessing is performed on the original data set, principal component analysis is performed on the original data set to reduce the dimension of the feature matrix.

[0013] According to a second aspect of the present disclosure, a bank merchant value evaluation device based on big data is provided, comprising: A collection module, used to collect original data sets of bank customers from big data, wherein the original data sets include explicit indicators and implicit indicators; a preprocessing module for preprocessing the original dataset, the preprocessing including filling missing values ​​and removing duplicate data, converting categorical variables in the original dataset into numerical form, maintaining binary form for binary features in the original dataset, using one-hot encoding for multi-valued categorical features, and using label encoding for ordered categorical features; and applying Z-score normalization to all numerical features to eliminate dimensionality effects; An indicator construction module is used to jointly construct the explicit indicators and the implicit indicators into a unified feature vector, wherein the implicit indicators are quantitatively represented by analyzing customer behavior data, and the explicit indicators and the implicit indicators are combined to form a feature matrix. A weight is assigned to each feature in the feature matrix, and the distance between sample points is calculated using weighted Euclidean distance to reflect their importance. A preliminary clustering module is used to randomly select K sample points as initial cluster centers using the feature matrix as input, calculate the weighted Euclidean distance from each sample point to each cluster center, assign the sample point to the cluster center with the smallest distance, recalculate the weighted mean of the samples within each cluster and update the cluster center, and repeat the assignment and update steps until the preset maximum number of iterations is reached or the cluster centers converge and remain unchanged, thereby obtaining K initial clusters; The density clustering module is used to determine the neighborhood radius and density threshold for each of the initial clusters to perform density clustering. During the execution, the cluster with the most samples is selected, the average distance of each point in the cluster with the most samples to the cluster center is calculated, and the nearest neighbor distance of the sample point within the average distance is obtained; the point density within the cluster is calculated based on the average distance and the nearest neighbor distance; then a neighborhood query is performed on each sample point in the cluster. During the query, the set of neighborhood points whose distance from the point does not exceed the average distance is calculated; if the number of points in the neighborhood of a point is not less than the nearest neighbor distance, the point is marked as a core point; starting from each core point, all points in its neighborhood are grouped into the same cluster; points that fall within the neighborhood of the core point but do not reach the nearest neighbor distance are marked as boundary points and assigned to the cluster closest to the core point; the remaining points not included in any cluster are marked as noise; the neighborhood query and cluster expansion operations are repeated until the cluster structure is stable and the final cluster division is obtained; The result fusion module is used to merge the initial cluster with the cluster division to obtain the final value label of each sample. During the merging, if no new sub-cluster is subdivided within the initial cluster and no noise is marked, the corresponding initial cluster is retained as the final cluster; if multiple sub-clusters are identified within the initial cluster, these sub-clusters are used as new final clusters; for the part marked as noise, it is removed from the initial cluster.

[0014] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the above-mentioned big data-based bank merchant value evaluation method.

[0015] According to a fourth aspect of the present disclosure, a bank merchant value evaluation device based on big data is provided, comprising: a processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions, when executed, enable the processor to execute the above-mentioned bank merchant value evaluation method based on big data.

[0016] The technical solution disclosed in this disclosure has the following beneficial effects: This method constructs a feature matrix by fusing explicit attributes and implicit preferences, and introduces feature weighting and dynamic parameter mechanisms to achieve intelligent customer or merchant clustering and value classification suitable for big data environments, significantly improving the accuracy of customer segmentation and business adaptability. It also automatically selects neighborhood radius and density thresholds based on local cluster structures, overcoming the technical bottlenecks of traditional algorithms, which are sensitive to parameters and unable to handle complex distributions and high-dimensional heterogeneous data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flow chart of a bank merchant value evaluation method based on big data in an embodiment of this specification; Figure 2 This is a structural block diagram of a bank merchant value evaluation device based on big data in an embodiment of this specification; Figure 3 This is a terminal device for implementing a bank merchant value evaluation method based on big data in the embodiments of this specification; Figure 4 This is a computer-readable storage medium storing a bank merchant value evaluation method based on big data in an embodiment of this specification. DETAILED DESCRIPTION

[0018] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0019] The accompanying drawings are merely schematic illustrations of the present disclosure. Identical reference numerals in the drawings denote identical or similar components, and thus their repeated description will be omitted. Some of the blocks shown in the accompanying drawings represent functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0020] The present invention provides a method for evaluating the value of bank merchants based on big data. Figure 1 The figure shows a flow chart of a method for evaluating the value of a bank merchant based on big data, provided by an embodiment of the present invention. The method can be applied to electronic devices such as personal computers and servers. The method can be performed by a device, which can be implemented by software and / or hardware. The method can specifically include the following steps S101-S102: In step S101 , an original data set of bank customers is collected from big data, where the original data set includes explicit indicators and implicit indicators.

[0021] The original dataset, sourced from internal banking systems and external data platforms, covers a large number of merchant behavioral data, transaction data, asset data, and external qualification information in banking operations. It possesses the characteristics of large data volume, wide dimensionality, and high update frequency. The dataset was collected through multiple channels, including bank data warehouses, acquiring systems, wealth management and credit systems, industrial and commercial data interfaces, judicial and blacklist databases, and more.

[0022] Explicit indicators refer to objective merchant attributes that can be directly captured and structured. These primarily include: merchant industry type, years of operation, registered capital, industrial and commercial credit rating, account balance, historical acquiring transaction amount and number, availability of wealth management products, loan products, default history, account activity frequency, and transaction peak distribution. Implicit indicators, on the other hand, are characteristics indirectly inferred through analysis of merchant behavior patterns and external performance. These include, but are not limited to, merchant financial preferences (e.g., preference for deposits or wealth management), responsiveness to marketing campaigns, stability of cash flow characteristics, acceptance of risky products, platform activity, cyclical fluctuations, external review and reputation, and relevance to blacklists. These indicators can be modeled based on multiple dimensions using machine learning models or labeling systems to characterize merchants' potential value and behavioral tendencies.

[0023] In step S102, the original data set is preprocessed, and the preprocessing includes filling in missing values ​​and deduplicating duplicate data, converting the categorical variables of the original data set into numerical form, and maintaining the binary form of binary features in the original data set, using one-hot encoding for multi-valued categorical features, and using label encoding for ordered categorical features; applying Z-score standardization to all numerical features to eliminate dimensionality effects.

[0024] The original merchant data contains some abnormal or missing fields, such as missing business registration information, blank industry types, inconsistent binary variables, and unusual transaction records. To ensure the accuracy of subsequent modeling, the following cleaning operations are performed. Cleaning includes missing value processing, duplicate value processing, and encoding conversion. For missing value processing, for example, if some missing values ​​exist in fields such as "Last Transaction Amount" or "Average Monthly Number of Transactions," to prevent bias in model predictions, the missing values ​​are filled using the mean of the non-missing portion of the field. For duplicate value processing, if duplicate records are found for some merchants during data structure analysis, only the unique identifier for each merchant is retained to eliminate weight bias caused by duplication. For encoding conversion, fields such as "Industry Type," "Business Model," "Loan Holder," and "Financial Management Purchaser" are represented as text and need to be converted to numeric values. Specific methods include using one-hot encoding for categorical fields (such as industry / risk level) and label encoding for ordinal fields (such as credit rating) to facilitate subsequent feature modeling.

[0025] To speed up model convergence and prevent high-dimensional variables from dominating the learning process, numerical variables need to be standardized. For binary variables (such as whether a merchant is a key merchant or whether a credit product is available), one-hot encoding is used; for multi-valued categorical variables (such as industry and qualification level), Z-score standardization is used.

[0026] In this embodiment, Z-score normalization is to subtract the mean of each feature value from the feature value and then divide it by its standard deviation, so that all features are modeled at the same scale, which helps to alleviate the problem of inconsistent dimensions and improve the comparability of each feature in the clustering algorithm. Specifically, its formula can be expressed as: ; ; Among them, x is the merchant characteristic value before normalization, is the standardized value, is the mean, is the standard deviation. After standardization, the mean of each field is normalized to 0 and the standard deviation is normalized to 1, thereby eliminating the impact of different indicators on the order of magnitude and improving the overall stability and generalization ability of subsequent clustering modeling.

[0027] In step S103, the explicit indicators and the implicit indicators are jointly constructed into a unified feature vector, wherein the implicit indicators are quantitatively represented by analyzing customer behavior data, the explicit indicators and the implicit indicators are combined to form a feature matrix, a weight is assigned to each feature in the feature matrix, and the distance between sample points is calculated using weighted Euclidean distance to reflect their importance.

[0028] After data preprocessing, a comprehensive characteristic indicator system is constructed based on the multidimensional attributes of merchants. The extracted merchant attributes are divided into two categories: explicit and implicit attributes. Explicit attributes refer to static attributes of merchants that can be directly observed or systematically recorded, including industry category, registered capital, years of operation, legal entity type, registered location (city or region), historical average account balance, loan limit, wealth management product holdings, and credit rating. These attributes are widely used to characterize merchants' basic qualifications and the extent of banking coverage. For explicit indicators, industry categories are encoded using one-hot encoding, such as "retail," "manufacturing," and "catering services," which are encoded as [1, 0, 0], [0, 1, 0], and [0, 0, 1], respectively. Ordinal variables such as educational background and credit rating are encoded using label encoding. For example, "high," "medium," and "low" are encoded as [3, 2, 1].

[0029] Implicit attributes refer to potential merchant behavioral characteristics that cannot be directly observed but can be inferred through behavioral analysis or derivative calculations. These include: transaction activity (e.g., average daily transaction volume, seasonal transaction peaks), account liquidity volatility index, fund allocation cycle patterns, marketing response (e.g., probability of product purchase after SMS / telemarketing), product portfolio preferences (e.g., preference for short-term vs. long-term financial products), merchant acceptance of risky products, and behavioral consistency indicators over multiple time periods. Implicit indicators can be mapped to a 1-5 scale, representing merchant activity or propensity. All features are integrated into a feature matrix, which is then used in subsequent algorithms through weighted clustering modeling.

[0030] In addition, different features have different effects on clustering. This embodiment introduces a weight factor to adjust the Euclidean distance calculation formula so that the distance calculation result more accurately reflects the similarity of merchants. The weighted distance between any two samples is defined as: ; in, and It's a different sample. is the weight factor.

[0031] In step S104, the feature matrix is ​​used as input, K sample points are randomly selected as initial cluster centers, the weighted Euclidean distance of each sample point to each cluster center is calculated, the sample point is assigned to the cluster center with the smallest distance, the weighted mean of the samples within each cluster is recalculated and the cluster center is updated, and the assignment and update steps are repeated until the preset maximum number of iterations is reached or the cluster centers converge and remain unchanged, thereby obtaining K initial clusters.

[0032] Specifically, the process of calculating K initial clusters can be decomposed into: Randomly select data points from the dataset as the initial centroid, expressed as: ; Calculate the weighted Euclidean distance from each merchant point to each center point, and store the merchant data points Assign to the nearest central cluster In the formula, it is expressed as: ; For each , calculate the mean of the data points in the cluster and use this mean as the new centroid: ; Repeat the above steps until the maximum number of iterations is reached, and the dataset is finally divided into K initial clusters.

[0033] In step S105, for each of the initial clusters, a neighborhood radius and a density threshold are determined to perform density clustering. During the execution, the cluster with the most samples is selected, the average distance from each point in the cluster with the most samples to the cluster center is calculated, and the nearest neighbor distance of the sample point within the average distance is obtained; the point density within the cluster is calculated based on the average distance and the nearest neighbor distance; then a neighborhood query is performed on each sample point in the cluster, and during the query, a set of neighborhood points whose distance from the point does not exceed the average distance is calculated; if the number of points in the neighborhood of a point is not less than the nearest neighbor distance, the point is marked as a core point; starting from each core point, all points in its neighborhood are grouped into the same cluster; points that fall within the neighborhood of the core point but do not reach the nearest neighbor distance are marked as boundary points and assigned to the cluster closest to the core point; the remaining points not included in any cluster are marked as noise; the neighborhood query and cluster expansion operations are repeated until the cluster structure is stable and the final cluster division is obtained.

[0034] Specifically, after the region is divided into K initial clusters, each cluster needs to be further subdivided. The decomposition steps include domain division, core point identification, cluster expansion, and noisy markup. Specifically: Neighborhood partitioning: clustering Each point , calculate the point set in its neighborhood, that is: ; Core point recognition: If a point There are at least MinPts points in its neighborhood, then It is considered as the core point; Cluster expansion: If is the core point, then and all points in its neighborhood are marked as belonging to the same cluster. If a point is a border point, it will be assigned to the cluster where the nearest core point is located; Noise point marking: If a point is neither a core point nor a boundary point, it is marked as noise; Select the cluster with the largest number of points in the initial clusters, calculate the average distance as Eps; take the nearest neighbor distance as MinPts, and calculate the density value: ; If the number of clusters remains unchanged after 5 consecutive iterations, the current clustering structure is considered stable.

[0035] In step S106, the initial cluster is merged with the cluster division to obtain the final value label of each sample. During the merging, if no new sub-cluster is subdivided within the initial cluster and no noise is marked, the corresponding initial cluster is retained as the final cluster; if multiple sub-clusters are identified within the initial cluster, these sub-clusters are used as new final clusters; for the part marked as noise, it is removed from the initial cluster.

[0036] As an explanation, there are three main situations when merging: The initial cluster is not subdivided, and no noise points are identified: This indicates that all merchants in the initial cluster are evenly distributed in the local density structure, requiring no further splitting, and there are no behavioral anomalies or edge cases. Therefore, this initial cluster is retained as the final merchant classification cluster, and all merchants within it are assigned a unified cluster label and treated as a single value group.

[0037] The initial cluster is subdivided into multiple subclusters: If multiple substructures with different densities (i.e., multiple subclusters) are identified within the initial cluster, this indicates that the initial clustering is too coarse and that significant behavioral differences still exist between merchants. In this case, the final merchant clusters should be redefined based on these subclusters. Each subcluster corresponds to an independent merchant value group, and the system will assign it a unique cluster label.

[0038] Some merchants are identified as "noise points": If certain merchant samples do not belong to any sub-cluster based on density assessment and do not meet the minimum density requirements (e.g., too few transaction records, unusual behavior, etc.), these samples are labeled as "noise points." These noisy merchants typically represent customers with unclear value, unstable data, or potential risks. The system will remove them from the original cluster and not assign a value label in the final clustering results, allowing for subsequent risk control, data re-entry, or strategic exclusion.

[0039] In one embodiment, the preprocessing specifically includes: filling missing values ​​with the mean of the feature column; retaining only one duplicate record; encoding non-numeric categorical features, such as using one-hot encoding for multi-valued categorical features, using label encoding for ordered categorical features, and maintaining the original 0 / 1 format for binary categorical features.

[0040] In one embodiment, the explicit indicators include basic attribute characteristics of the customer, which include operating time, asset balance, and industry type; the implicit indicators include preference information extracted from the customer's historical behavior or feedback data, which includes the customer's interest score for a specific financial product, and the explicit indicators and the implicit indicators are formed into a unified feature vector.

[0041] In one embodiment, after the preprocessing is performed on the original data set, principal component analysis is further performed on the original data set to reduce the dimension of the feature matrix.

[0042] Based on the same idea, Figure 2 As shown, a bank merchant value evaluation device based on big data is provided, comprising: The acquisition module 201 is used to collect the original data set of bank customers from big data, wherein the original data set includes explicit indicators and implicit indicators; A preprocessing module 202 is configured to preprocess the original dataset, including filling in missing values ​​and removing duplicate data, converting categorical variables in the original dataset into numerical form, maintaining binary form for binary features in the original dataset, using one-hot encoding for multi-valued categorical features, and using label encoding for ordered categorical features; and applying Z-score normalization to all numerical features to eliminate dimensionality effects. An indicator construction module 203 is configured to jointly construct the explicit indicator and the implicit indicator into a unified feature vector, wherein the implicit indicator is quantitatively represented by analyzing customer behavior data, and the explicit indicator and the implicit indicator are combined to form a feature matrix. A weight is assigned to each feature in the feature matrix, and the distance between sample points is calculated using weighted Euclidean distance to reflect their importance. The preliminary clustering module 204 is configured to take the feature matrix as input, randomly select K sample points as initial cluster centers, calculate the weighted Euclidean distance between each sample point and each cluster center, assign the sample point to the cluster center with the smallest distance, recalculate the weighted mean of the samples within each cluster and update the cluster center, and repeat the assigning and updating steps until a preset maximum number of iterations is reached or the cluster centers converge and remain unchanged, thereby obtaining K initial clusters. The density clustering module 205 is used to determine the neighborhood radius and density threshold for each of the initial clusters to perform density clustering. During the execution, the cluster with the most samples is selected, the average distance of each point in the cluster with the most samples to the cluster center is calculated, and the nearest neighbor distance of the sample point within the average distance is obtained; the point density within the cluster is calculated based on the average distance and the nearest neighbor distance; then, a neighborhood query is performed on each sample point in the cluster. During the query, the set of neighborhood points whose distance from the point does not exceed the average distance is calculated; if the number of points in the neighborhood of a point is not less than the nearest neighbor distance, the point is marked as a core point; starting from each core point, all points in its neighborhood are grouped into the same cluster; points that fall within the neighborhood of the core point but do not reach the nearest neighbor distance are marked as boundary points and assigned to the cluster closest to the core point; the remaining points not included in any cluster are marked as noise; the neighborhood query and cluster expansion operations are repeated until the cluster structure is stable and the final cluster division is obtained; The result fusion module 206 is used to merge the initial cluster with the cluster division to obtain the final value label of each sample. During the merging, if no new sub-cluster is subdivided within the initial cluster and no noise is marked, the corresponding initial cluster is retained as the final cluster; if multiple sub-clusters are identified within the initial cluster, these sub-clusters are used as new final clusters; for the part marked as noise, it is removed from the initial cluster.

[0043] This device constructs a feature matrix by fusing explicit attributes and implicit preferences, and introduces feature weighting and dynamic parameter mechanisms to achieve intelligent customer or merchant clustering and value classification suitable for big data environments, significantly improving the accuracy of customer segmentation and business adaptability. It automatically selects neighborhood radius and density thresholds based on the local cluster structure, overcoming the technical bottlenecks of traditional algorithms that are sensitive to parameters and unable to cope with complex distributions and high-dimensional heterogeneous data.

[0044] The specific details of each module / unit in the above device have been described in detail in the implementation method part. For undisclosed details, please refer to the implementation method part, and thus will not be repeated here.

[0045] Based on the same idea, the embodiment of this specification also provides a bank merchant value evaluation device based on big data, such as Figure 3 shown.

[0046] The bank merchant value evaluation device based on big data can be the terminal device or server provided in the above embodiment.

[0047] Big data-based bank merchant value assessment devices can vary significantly depending on their configuration or performance. They may include one or more processors 301, memory 302, and a bus. Memory 302 may store one or more applications or data. Memory 302 may include readable media in the form of volatile storage units, such as random access memory (RAM) and / or cache memory, such as plug-in removable hard drives, Smart Media Cards (SMCs), Secure Digital (SD) cards, or Flash Cards found on electronic devices. It may also include read-only storage units. Applications stored in memory 302 may include one or more program modules (not shown). Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Furthermore, processor 301 may be configured to communicate with memory 302 to execute a series of computer-executable instructions stored in memory 302 on the big data-based bank merchant value assessment device. The big data-based bank merchant value assessment device may also include one or more power supplies 303, one or more wired or wireless network interfaces 304, one or more I / O interfaces (input and output interfaces) 305, and one or more external devices 306 (e.g., a keyboard). It may also communicate with one or more devices that enable a user to interact with the device, and / or with any device that enables the device to communicate with one or more other computing devices (e.g., a router, a network switch, etc.). Such communication may be performed via the I / O interface 305. Furthermore, the device may also communicate with one or more networks (e.g., a local area network (LAN)) via the wired or wireless interface 304.

[0048] Figure 3 Only the bank merchant value evaluation device based on big data with components is shown. It can be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the bank merchant value evaluation device based on big data, and may include fewer or more components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0049] Specifically, in this embodiment, the bank merchant value evaluation device based on big data includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions in the bank merchant value evaluation device based on big data, and the one or more programs are configured to be executed by one or more processors, including computer-executable instructions for performing the following: Collecting an original data set of bank customers from big data, wherein the original data set includes explicit indicators and implicit indicators; Preprocessing the original dataset includes filling in missing values ​​and removing duplicate data, converting categorical variables in the original dataset into numerical form, maintaining binary features in the original dataset in binary form, using one-hot encoding for multi-valued categorical features, and using label encoding for ordered categorical features; and applying Z-score normalization to all numerical features to eliminate dimensionality effects; The explicit indicators and the implicit indicators are jointly constructed into a unified feature vector, wherein the implicit indicators are quantitatively represented by analyzing customer behavior data. The explicit indicators and the implicit indicators are combined to form a feature matrix, each feature in the feature matrix is ​​assigned a weight, and the distance between sample points is calculated using weighted Euclidean distance to reflect their importance; Taking the feature matrix as input, randomly select K sample points as initial cluster centers, calculate the weighted Euclidean distance from each sample point to each cluster center, assign the sample point to the cluster center with the smallest distance, recalculate the weighted mean of the samples within each cluster and update the cluster center, and repeat the assignment and update steps until the preset maximum number of iterations is reached or the cluster centers converge and remain unchanged, thus obtaining K initial clusters; For each of the initial clusters, a neighborhood radius and a density threshold are determined to perform density clustering. During the execution, the cluster with the most samples is selected, the average distance from each point in the cluster with the most samples to the cluster center is calculated, and the nearest neighbor distance of the sample point within the average distance is obtained; the point density within the cluster is calculated based on the average distance and the nearest neighbor distance; then a neighborhood query is performed on each sample point in the cluster. During the query, a set of neighborhood points whose distance from the point does not exceed the average distance is calculated; if the number of points in the neighborhood of a point is not less than the nearest neighbor distance, the point is marked as a core point; starting from each core point, all points in its neighborhood are grouped into the same cluster; points that fall within the neighborhood of the core point but do not reach the nearest neighbor distance are marked as boundary points and assigned to the cluster closest to the core point; the remaining points that are not included in any cluster are marked as noise; the neighborhood query and cluster expansion operations are repeated until the cluster structure is stable and the final cluster division is obtained; The initial cluster is merged with the cluster division to obtain the final value label of each sample. During the merging, if no new sub-cluster is subdivided within the initial cluster and no noise is marked, the corresponding initial cluster is retained as the final cluster; if multiple sub-clusters are identified within the initial cluster, these sub-clusters are used as new final clusters; for the part marked as noise, it is removed from the initial cluster.

[0050] Based on the same idea, the exemplary embodiments of the present invention further provide a computer-readable storage medium storing a program product capable of implementing the methods described above. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product comprising program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps described in the "Exemplary Methods" section above according to various exemplary embodiments of the present disclosure.

[0051] refer to Figure 4 As shown, a program 700 for implementing the above method according to an exemplary embodiment of the present disclosure is described. The program 700 may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0052] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0053] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0054] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, CSS, HTML, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0055] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the exemplary embodiment of the present disclosure.

[0056] Furthermore, the figures above are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the figures above do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0057] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0058] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and embodiments are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0059] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A bank merchant value evaluation method based on big data, characterized in that: The method comprises: Collecting an original data set of bank customers from big data, wherein the original data set includes explicit indicators and implicit indicators; Preprocessing the original dataset includes filling in missing values ​​and removing duplicate data, converting categorical variables in the original dataset into numerical form, maintaining binary features in the original dataset in binary form, using one-hot encoding for multi-valued categorical features, and using label encoding for ordered categorical features; and applying Z-score normalization to all numerical features to eliminate dimensionality effects; The explicit indicators and the implicit indicators are jointly constructed into a unified feature vector, wherein the implicit indicators are quantitatively represented by analyzing customer behavior data. The explicit indicators and the implicit indicators are combined to form a feature matrix, each feature in the feature matrix is ​​assigned a weight, and the distance between sample points is calculated using weighted Euclidean distance to reflect their importance; Taking the feature matrix as input, randomly select K sample points as initial cluster centers, calculate the weighted Euclidean distance from each sample point to each cluster center, assign the sample point to the cluster center with the smallest distance, recalculate the weighted mean of the samples within each cluster and update the cluster center, and repeat the assignment and update steps until the preset maximum number of iterations is reached or the cluster centers converge and remain unchanged, thus obtaining K initial clusters; For each of the initial clusters, a neighborhood radius and a density threshold are determined to perform density clustering. During the execution, the cluster with the most samples is selected, the average distance from each point in the cluster with the most samples to the cluster center is calculated, and the nearest neighbor distance of the sample point within the average distance is obtained; the point density within the cluster is calculated based on the average distance and the nearest neighbor distance; then a neighborhood query is performed on each sample point in the cluster. During the query, a set of neighborhood points whose distance from the point does not exceed the average distance is calculated; if the number of points in the neighborhood of a point is not less than the nearest neighbor distance, the point is marked as a core point; starting from each core point, all points in its neighborhood are grouped into the same cluster; points that fall within the neighborhood of the core point but do not reach the nearest neighbor distance are marked as boundary points and assigned to the cluster closest to the core point; the remaining points that are not included in any cluster are marked as noise; the neighborhood query and cluster expansion operations are repeated until the cluster structure is stable and the final cluster division is obtained; The initial cluster is merged with the cluster division to obtain the final value label of each sample. During the merging, if no new sub-cluster is subdivided within the initial cluster and no noise is marked, the corresponding initial cluster is retained as the final cluster; if multiple sub-clusters are identified within the initial cluster, these sub-clusters are used as new final clusters; for the part marked as noise, it is removed from the initial cluster.

2. The bank merchant value evaluation method based on big data according to claim 1 is characterized in that: The pretreatment specifically includes: Fill missing values ​​with the mean of the feature column; For duplicate records, only one is retained; Encode non-numeric categorical features, such as using one-hot encoding for multi-valued categorical features, label encoding for ordered categorical features, and keeping the original 0 / 1 form for binary categorical features.

3. The bank merchant value evaluation method based on big data according to claim 1 is characterized in that: The explicit indicators include basic customer attribute characteristics, which include operating time, asset balance, and industry type. The implicit indicators include preference information extracted from customer historical behavior or feedback data, which includes the customer's interest score in specific financial products. The explicit indicators and the implicit indicators are formed into a unified feature vector.

4. The bank merchant value evaluation method based on big data according to claim 1 is characterized in that: The Z-score normalization process includes subtracting the mean from the eigenvalue and dividing it by the standard deviation to eliminate the influence of different dimensions, and the normalized eigenvalues ​​are used to form the feature matrix.

5. The bank merchant value evaluation method based on big data according to claim 1 is characterized in that: After the preprocessing is performed on the original data set, principal component analysis is further performed on the original data set to reduce the dimension of the feature matrix.

6. A bank merchant value evaluation device based on big data, characterized in that: include: A collection module, used to collect original data sets of bank customers from big data, wherein the original data sets include explicit indicators and implicit indicators; a preprocessing module for preprocessing the original dataset, the preprocessing including filling missing values ​​and removing duplicate data, converting categorical variables in the original dataset into numerical form, maintaining binary form for binary features in the original dataset, using one-hot encoding for multi-valued categorical features, and using label encoding for ordered categorical features; and applying Z-score normalization to all numerical features to eliminate dimensionality effects; An indicator construction module is used to jointly construct the explicit indicators and the implicit indicators into a unified feature vector, wherein the implicit indicators are quantitatively represented by analyzing customer behavior data, and the explicit indicators and the implicit indicators are combined to form a feature matrix. A weight is assigned to each feature in the feature matrix, and the distance between sample points is calculated using weighted Euclidean distance to reflect their importance. A preliminary clustering module is used to randomly select K sample points as initial cluster centers using the feature matrix as input, calculate the weighted Euclidean distance from each sample point to each cluster center, assign the sample point to the cluster center with the smallest distance, recalculate the weighted mean of the samples within each cluster and update the cluster center, and repeat the assignment and update steps until the preset maximum number of iterations is reached or the cluster centers converge and remain unchanged, thereby obtaining K initial clusters; The density clustering module is used to determine the neighborhood radius and density threshold for each of the initial clusters to perform density clustering. During the execution, the cluster with the most samples is selected, the average distance of each point in the cluster with the most samples to the cluster center is calculated, and the nearest neighbor distance of the sample point within the average distance is obtained; the point density within the cluster is calculated based on the average distance and the nearest neighbor distance; then a neighborhood query is performed on each sample point in the cluster. During the query, the set of neighborhood points whose distance from the point does not exceed the average distance is calculated; if the number of points in the neighborhood of a point is not less than the nearest neighbor distance, the point is marked as a core point; starting from each core point, all points in its neighborhood are grouped into the same cluster; points that fall within the neighborhood of the core point but do not reach the nearest neighbor distance are marked as boundary points and assigned to the cluster closest to the core point; the remaining points not assigned to any cluster are marked as noise; Repeat the neighborhood query and cluster expansion operations until the cluster structure is stable and the final cluster division is obtained; The result fusion module is used to merge the initial cluster with the cluster division to obtain the final value label of each sample. During the merging, if no new sub-cluster is subdivided within the initial cluster and no noise is marked, the corresponding initial cluster is retained as the final cluster; if multiple sub-clusters are identified within the initial cluster, these sub-clusters are used as new final clusters; for the part marked as noise, it is removed from the initial cluster.

7. A bank merchant value evaluation device based on big data, characterized in that: The device comprises: A processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to execute the bank merchant value evaluation method based on big data as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the bank merchant value evaluation method based on big data as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Customer segmentation method and device based on cluster analysis

    CN108734217A

  • A K-nearest neighbor and multi-class merge based density peak clustering method and image segmentation system

    CN109409400A

  • Data classification method based on data-in-data table

    CN118839191A

Cited By

  • Industrial water quota system evaluation method and device, electronic equipment and storage medium

    CN122047946A