Merchant risk identification method and system based on multi-dimensional feature quantification

By employing a multi-dimensional feature quantification method for merchant risk identification, combined with geospatial and machine learning technologies, merchant risks are identified and assessed, solving the regulatory challenges of cross-regional violations and achieving efficient and accurate risk identification and supervision.

CN121073232APending Publication Date: 2025-12-05CHINA NATIONAL DIGITAL SECURITY TECHNOLOGY (ZHEJIANG) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511637415.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively identify and regulate complex violations such as cross-regional sales of cigarettes, resulting in high underreporting rates, delayed regulatory timeliness, and data silos. They cannot meet the needs of dynamic monitoring and networked collaborative investigations, leading to a continuous increase in illegal operations and violations.

Method used

By acquiring multidimensional merchant data and preprocessing it, geospatial relationships and machine learning methods are used to assess the static risk and regional deviation risk of merchants. The network risk impact value is calculated by combining the random walk algorithm, generating a comprehensive risk score, and then visualized.

Benefits of technology

It has enabled automated and accurate identification and assessment of merchant risks, provided scientific decision support, improved regulatory efficiency and accuracy, and formed a closed-loop regulatory system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073232A_ABST
    Figure CN121073232A_ABST
Patent Text Reader

Abstract

The invention provides a merchant risk identification method and system based on multi-dimensional feature quantization, and the method comprises the steps: carrying out the target processing of each merchant, and carrying out the target processing of each merchant according to the merchant operation data of the merchant and the respective merchant operation data of a plurality of merchants, obtaining a scoring result of each static risk and the regional deviation risk of the merchant; calculating a network risk influence value of the merchant according to a random walk algorithm and the respective geographic coordinates of the plurality of merchants; inputting the scoring result of the static risk of the commercial tenant into a scoring card model to obtain a comprehensive risk score of the commercial tenant; and performing weighted summation according to the scoring result of the regional deviation risk, the network risk influence value, the comprehensive risk score and the scoring result of the static risk to obtain a final risk score of the merchant. Therefore, the efficiency and accuracy of identifying the risk of the commercial tenant can be improved by performing multi-dimensional feature quantification on the commercial tenant operation data of the commercial tenant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of merchant risk management technology, and in particular to a merchant risk identification method and system based on multidimensional feature quantification. Background Technology

[0002] The sale of special categories of goods, such as tobacco and pharmaceuticals, requires specialized regulation. In the tobacco monopoly regulatory market, current daily supervision of retailers primarily relies on industry-owned data (such as data on tobacco operations, marketing, violations, and licensing). However, for complex violations such as the cross-regional sale of cigarettes, the traditional regulatory model suffers from significant shortcomings due to a lack of social data resources: First, there are blind spots in the network of connections; for violations involving multiple stores by one person or cross-regional group operations, the lack of equity penetration and spatiotemporal correlation analysis capabilities results in an underreporting rate of over 30% for cross-regional violations, according to industry sampling statistics. Second, regulatory timeliness is lagging; relying on manual on-site verification, the average investigation cycle for cross-regional cases exceeds 30 days. Third, there is the problem of data silos; business license information is separated from the authorization and licensing system, making it impossible to verify in real time anomalies such as "inconsistent licenses and permits" or "inconsistent identification between person and document."

[0003] Existing technologies cannot meet core regulatory needs such as "dynamic monitoring" and "networked collaborative investigation," leading to a continued rapid increase in cross-regional illegal operations and violations, which seriously endangers market order and public interests. Summary of the Invention

[0004] This application provides a merchant risk identification method and system based on multi-dimensional feature quantification, which can quantify and detect merchant risks from multiple dimensions, thereby improving the efficiency and accuracy of merchant risk identification.

[0005] Firstly, embodiments of this application provide a merchant risk identification method based on multi-dimensional feature quantification. The method includes: acquiring merchant data and business registration data of multiple merchants; processing the merchant data and business registration data to obtain merchant operating data; and performing target processing on each of the multiple merchants, the target processing including: scoring various static risks and regional deviation risks of the merchant based on its operating data and the individual operating data of the multiple merchants, obtaining respective scoring results; the static risks include at least one of the following: cross-provincial risk of legal entity, license expiration risk, and historical penalty risk; regional deviation risk... The system measures the degree of deviation between a merchant's sales performance in a specific region and the average sales performance in that region. The static risk score of the merchant is input into a scoring card model to obtain the merchant's comprehensive risk score. The scoring card model is obtained through supervised learning based on the static risk scores of multiple historical merchants. Using a random walk algorithm and the geographical coordinates of multiple merchants, the network risk impact value of the merchant's risk propagation to other merchants is calculated. Finally, the merchant's final risk score is obtained by weighted summation of the regional deviation risk score, the network risk impact value, the comprehensive risk score, and the static risk score.

[0006] Therefore, this solution first automates and efficiently acquires merchant and business registration data from multiple merchants, and preprocesses it to construct a comprehensive, multi-dimensional data source for assessing the risk status of each merchant. Based on this multi-dimensional data, the solution comprehensively calculates the risk score of tobacco merchants by integrating geospatial relationships, historical behavioral data, and machine learning methods. This includes: scoring each merchant's static risks and regional deviation risks based on historical behavioral data, generating detailed scoring results; calculating the network risk impact value of each merchant by integrating the geospatial relationships of multiple merchants; and comprehensively processing the scoring results of each static risk using a machine learning model to obtain a comprehensive risk score for each merchant. Finally, by weighting and summing the scoring results of each static risk, the scoring results of regional deviation risk, the network risk impact value, and the comprehensive risk score, the final risk score of the merchant is obtained and visualized. Through this automated risk perception mechanism, this solution can effectively identify and assess the operational risks of merchants, providing a scientific basis and decision support for timely risk prevention and mitigation measures.

[0007] Secondly, embodiments of this application provide a merchant risk identification system based on multi-dimensional feature quantization. The system includes: an acquisition module for acquiring merchant data and business registration data from multiple merchants; a preprocessing module for processing the merchant data and business registration data to obtain merchant operating data; the merchant operating data includes the merchant's business license information, operating permit information, administrative penalty code, and unified social credit code; and a risk identification module for performing target processing on each merchant among the multiple merchants. Target processing includes: scoring the merchant's various static risks and regional deviation risks based on the merchant's operating data and the individual operating data of the multiple merchants, obtaining their respective score results; the various static risks include the following... At least one of the following: cross-provincial risk for legal entities, license expiration risk, and historical penalty risk; regional deviation risk is the degree of deviation between a merchant's sales performance in a specific region and its average sales performance in that region; the merchant's static risk score is input into the scoring card model to obtain the merchant's comprehensive risk score; the scoring card model is obtained through supervised learning based on the scores of various static risks of multiple historical merchants; the network risk impact value of the merchant's risk propagation to other merchants is calculated based on the random walk algorithm and the geographical coordinates of multiple merchants; the final risk score of the merchant is obtained by weighted summation of the regional deviation risk score, the network risk impact value, the comprehensive risk score, and the static risk score.

[0008] Thirdly, embodiments of this application provide a computer storage medium storing a computer program thereon, which, when executed in a computer, causes the computer to perform the method described in the first aspect or any possible implementation of the first aspect.

[0009] Fourthly, embodiments of this application provide a computer program product containing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect or any possible implementation thereof.

[0010] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1A framework diagram of a merchant risk identification system based on multidimensional feature quantization provided for embodiments of this application; Figure 2 A flowchart illustrating a merchant risk identification method based on multidimensional feature quantization, provided for embodiments of this application; Figure 3 A flowchart illustrating a training method for a scorecard model provided in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings.

[0014] In the description of the embodiments in this application, any embodiment or design that is “exemplary,” “for example,” or “by way of example” should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as “exemplary,” “for example,” or “by way of example” is intended to present the relevant concepts in a concrete manner.

[0015] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized. "Multiple" can refer to one or more, where multiple means two or more.

[0016] This application proposes an innovative solution to the problems mentioned in the background art. The solution first automates and efficiently acquires merchant data and business registration data from multiple merchants, and preprocesses this data to construct a comprehensive multi-dimensional data source for assessing the risk status of each merchant. Based on this multi-dimensional data, the solution comprehensively calculates the risk score of tobacco merchants by integrating geospatial relationships, historical behavioral data, and machine learning methods. This includes: scoring each merchant's static risks and regional deviation risks based on historical behavioral data, generating detailed scoring results; calculating the network risk impact value of each merchant by integrating the geospatial relationships of multiple merchants; and comprehensively processing the scoring results of each static risk using a machine learning model to obtain a comprehensive risk score for each merchant. Finally, the final risk score of the merchant is obtained and visualized by weighted summing of the scoring results of each static risk, the scoring results of regional deviation risk, the network risk impact value, and the comprehensive risk score. Through this automated risk perception mechanism, this solution can effectively identify and assess the operational risks of merchants, providing a scientific basis and decision support for timely risk prevention and mitigation measures.

[0017] For example, Figure 1 The image below is a framework diagram of a merchant risk identification system based on multi-dimensional feature quantization, provided in an embodiment of this application. Figure 1 As shown, the system 100 can be deployed on any computing unit, server, device, device cluster, etc. that has any computing and processing capabilities.

[0018] System 100 includes an acquisition module 10, a preprocessing module 20, a risk identification module 30, and a risk visualization module 40. The risk identification module 30 includes a static risk processing module 31 and a dynamic risk processing module 32. The dynamic risk processing module 32 includes a regional deviation assessment module 321, a network risk assessment module 322, and a comprehensive risk assessment module 323.

[0019] Module 10 is used to automatically synchronize and efficiently acquire the latest tobacco merchant data and business registration data. This module establishes a stable data connection channel with relevant units to obtain authoritative and accurate tobacco merchant data sources in real time. Based on this, and according to pre-defined and rigorously validated tobacco data field standards, it acquires the corresponding latest business registration data.

[0020] The preprocessing module 20 utilizes advanced data integration technology to deeply clean data from different data sources, removing duplicate, erroneous, or incomplete data records to ensure data accuracy and consistency. Simultaneously, it precisely matches and aligns tobacco merchant data with industrial and commercial operation data across key fields, establishing a close and logically rigorous relationship between the two.

[0021] In addition, this module also features a scheduled task mechanism. Through periodic secondary processing of the acquired raw merchant and business registration data, a series of key multidimensional data points are extracted and generated based on business logic. This multidimensional data will provide rich, comprehensive, and structured foundational data resources for subsequent in-depth data analysis and decision support, ensuring the efficient and stable operation of other functional modules and providing solid data support and assurance for the entire system's business processes.

[0022] The risk identification module 30 is used to comprehensively calculate the risk score of tobacco merchants based on these multidimensional data by integrating geospatial relationships, historical behavioral data and machine learning methods.

[0023] The static risk processing module 31 is used to score various static risks of each merchant based on historical behavioral data from multidimensional data, generating detailed scoring results. This module uses advanced multi-source data fusion technology to deeply integrate and correlate massive amounts of heterogeneous data. Based on key data dimensions such as national business licenses (covering information on legal persons, shareholders, and branches), tobacco monopoly license details, administrative penalty records, and enterprise unified social credit codes, this module scores various static risks of each merchant, generating detailed scoring results. Each static risk includes at least one of the following: legal person cross-province risk (legal_risk), license expiration risk (expiry_risk), and historical penalty risk (penalty_risk).

[0024] The dynamic risk processing module 32 is used to obtain the feature values ​​of dynamic risk characteristics for each merchant by integrating geospatial relationships, historical behavioral data, and machine learning methods. This module uses intelligent algorithms to deeply mine the potential correlations between different data sources, achieving a comprehensive mapping from individual information to group relationships, providing multi-faceted and dynamic perspectives for risk assessment. In particular, cross-validation of merchants' cross-regional business activities, tobacco license compliance, corporate credit status, and historical penalty records can accurately identify potential chain risk nodes, forming a risk relationship graph and providing data-driven intelligent support for subsequent regulatory decisions. Dynamic risk characteristics include regional deviation risk, network risk impact value, and comprehensive risk score.

[0025] The regional deviation assessment module 321 is used to score the regional deviation risk of each merchant based on historical behavioral data in multidimensional data, generating detailed scoring results. Various static risk characteristics and regional deviation risk characteristics constitute the basic risk profile of the merchant, providing a solid foundation for further risk assessment and management.

[0026] The network risk assessment module 322 is used to construct a relationship graph by integrating geospatial relationships, and to use the relationships in the graph to predict how a merchant's risk will affect other merchants. A relationship graph is a database for storing and querying graph-structured data, which can be used to represent relationships between merchants, such as transaction relationships, geographical location relationships, etc. Predicting network risk involves building a model to analyze and simulate how risk propagates within a merchant network.

[0027] The comprehensive risk assessment module 323 is used to calculate a merchant's comprehensive risk using machine learning methods and the scoring results of various static risks. A scoring card model is built using machine learning methods; this is a common credit scoring tool that can be used to calculate a merchant's comprehensive risk score based on various static risk characteristics (such as historical penalties, license expiration, etc.).

[0028] The risk visualization module 40 integrates different risk scores and model outputs to obtain a comprehensive score. It also visualizes the scoring results, making the risk assessment more intuitive and understandable. These risk assessment results help evaluate the overall risk of tobacco merchants. This will help decision-makers understand which merchants may face higher risks and take appropriate measures.

[0029] Therefore, through the risk assessment system 1000 described above, this solution aims to employ a data-driven strategy to conduct a comprehensive and detailed assessment of the operational risks of tobacco merchants. This method utilizes advanced data processing and analysis technologies to ensure the accuracy and foresight of the risk assessment, thereby providing a solid foundation for merchants' risk management and decision-making.

[0030] Based on the above, this application provides a detailed description of a merchant risk identification method based on multidimensional feature quantification.

[0031] For example, Figure 2 A flowchart illustrating a merchant risk identification method based on multidimensional feature quantization, provided for an embodiment of this application.

[0032] Step S201: Obtain merchant data and business registration data from multiple merchants.

[0033] For example, using Figure 1 The acquisition module 10 shown acquires merchant data and business registration data from multiple merchants through timed or user-triggered methods. These multiple merchants are the target merchants for whom risk assessment is desired.

[0034] Merchant data includes tobacco business data. The latest tobacco business data from multiple merchants is obtained through encrypted channels, and the data undergoes preliminary verification before import. If tobacco data is not present in the database, it is written; if relevant tobacco merchant information already exists, it is updated to ensure the tobacco data is up-to-date.

[0035] Business registration data includes merchants' business registration information. After completing the tobacco data writing operation, the system will call the business registration data interface to obtain the merchants' business registration information based on the preset relevant fields. The system will comprehensively query business registration information related to the merchants based on the existing basic fields, including but not limited to the merchant's registered address, the legal representative's registered address, the precise latitude and longitude coordinates of the merchant's address, the amount of the enterprise's registered capital, the validity period of the business license, and whether the merchant has any illegal records, etc.

[0036] Step S202: Process the merchant data and business registration data to obtain merchant operating data.

[0037] For example, the acquired merchant data and business registration data are preprocessed to obtain merchant operation data for multiple merchants. This process involves integrating business license data with multi-source data such as tobacco retail merchant data, administrative permits, and administrative penalties to construct a multi-dimensional database, providing a more comprehensive and accurate data foundation for tobacco market supervision. Merchant operation data includes, but is not limited to, information such as business license information, operating permit information, administrative penalty codes, and unified social credit codes of enterprises.

[0038] Before integrating tobacco and business data, define the fields of the processed data model, including key information fields of tobacco and business data, such as: unified social credit code, tobacco license number, store address and registered address, etc.

[0039] In the process of integrating tobacco and industrial and commercial data, the system first needs to preprocess heterogeneous data from different data sources. This includes unifying the data format to ensure that all data follows the same format standard; converting the data type so that similar data from different data sources can be correctly identified and processed; and standardizing the data encoding to eliminate data inconsistencies caused by encoding differences.

[0040] These preprocessing steps ensure data consistency and compatibility. Next, by establishing field mapping relationships and data association rules, tobacco data and industrial and commercial data can be logically linked, achieving deep data integration. This step not only improves data usability but also provides a solid foundation for subsequent analysis and regulation.

[0041] After integration, the system's scheduled tasks will process and assign values ​​to preset fields based on the integrated data. For example, it may determine whether a merchant is operating across regions based on the merchant's address and the legal representative's registered address; or it may count whether a legal representative holds multiple tobacco licenses and whether the licenses are being used for cross-regional operations.

[0042] For each of the multiple merchants, target processing is performed. Target processing includes: Step S203, based on the merchant's business data and the individual business data of the multiple merchants, scoring the merchant's various static risks and regional deviation risks, and obtaining their respective scoring results; each static risk includes at least one of the following: legal entity cross-province risk, license expiration risk, and historical penalty risk; regional deviation risk is the degree of deviation between the merchant's sales performance in a specific region and its average sales performance in the region.

[0043] For example, three types of static risk characteristics, as well as regional deviation risk characteristics, are extracted for each merchant:

[0044] The cross-provincial risk characteristic for legal entities has a first preset value of a non-zero cross-provincial risk score, indicating that the merchant's registered address and the legal entity's registered address are in different provinces; and a cross-provincial risk score of zero indicates that the merchant's registered address and the legal entity's registered address are in the same province. For example, if the merchant's registered address and the legal entity's registered address are not in the same province, the score is assigned to 0.8; otherwise, it is assigned to 0.

[0045] The license expiration risk characteristic includes a second preset value where the license expiration risk score is non-zero, indicating that the merchant's license expiration time is no greater than a preset threshold; and a license expiration risk score of zero, indicating that the merchant's license expiration time is greater than the preset threshold. For example, if the license expires within the next 180 days, the score is set to 0.7; otherwise, it is 0.

[0046] Historical penalty risk characteristics are determined by normalizing a merchant's historical penalty count based on the maximum number of penalties among multiple merchants, resulting in a historical penalty risk score for that merchant. The formula for calculating the normalized number of historical penalties for a merchant is as follows:

[0047] (1)

[0048] In formula (1), Indicates the first among multiple users The historical penalty risk score for each merchant. Indicates the first The number of historical penalties for each merchant This indicates the maximum number of penalties imposed on a merchant among multiple users.

[0049] Regional deviation risk characteristic calculation: Regional deviation risk refers to the degree of deviation between a merchant's sales performance in a specific region and the average sales performance in the same region. This deviation is quantified by comparing the merchant's sales share with the regional average sales share, in order to predict the market risks that the merchant may face in the future.

[0050] For example, the product categories sold by multiple merchants constitute a category set. Based on the merchant's sales percentage for each brand within the category set, and the average sales percentage of each brand within the brand set in the merchant's region, a regional deviation risk score is obtained for that merchant. The merchant's regional deviation risk is then calculated. The "regional deviation" index The calculation formula is:

[0051] (2)

[0052] In formula (2), Indicates the first among multiple users Individual merchants in product categories Sales share of the top Indicates product category The average sales percentage in the region where this merchant is located. This is for smoothing terms to prevent division by zero errors.

[0053] These four characteristics constitute the basic risk profile of merchants, providing a solid foundation for further risk assessment and management.

[0054] Step S204: Input the static risk score of the merchant into the scoring card model to obtain the comprehensive risk score of the merchant; the scoring card model is obtained through supervised learning based on the static risk scores of multiple historical merchants.

[0055] For example, based on a basic risk profile, one or more risk scoring models can be built. These models can be statistical models, such as logistic regression, or machine learning models, such as random forests or neural networks. Model construction needs to consider the following factors: Feature selection: Selecting features most relevant to risk. Model training: Training the model using historical data to predict the risk level of merchants. Model validation: Validating the model's accuracy and generalization ability through methods such as cross-validation.

[0056] In one implementation, a scorecard model is obtained through supervised learning based on the scoring results of various static risks from multiple historical merchants. The parameter tuning and training process of the logistic regression-based scorecard model can be performed through the following steps:

[0057] The first step is to collect original characteristics that reflect the risks of merchants.

[0058] For example, the scorecard model needs to be based on raw features that can reflect merchant risk. The features collected in this embodiment include the three types of static risk features mentioned above:

[0059] Risk of legal entities operating across provinces: Risk score related to the geographical dispersion of legal entities (e.g., cross-provincial operations).

[0060] License expiration risk: Whether the license held by the merchant is abnormal (e.g., expired, revoked, etc.).

[0061] Historical penalty risk: The number of times or severity of administrative penalties a merchant has received in the past (a continuous value, such as a risk value standardized from 0 to 1).

[0062] The scoring results of the above three types of static risk characteristics have been given in step S203.

[0063] The second step is to define high-risk merchant labels.

[0064] For example, to enable the scorecard model to distinguish between "high-risk" and "non-high-risk" merchants during training and inference, a clear target variable (label) needs to be defined. For each historical merchant, a risk / non-risk label is assigned based on the scoring results of the target items in each static risk category. In this embodiment: the target item is "historical penalty risk." If a merchant has more than 2 historical penalties, it is marked as a high-risk merchant (label: Otherwise, mark it as a non-high-risk merchant (tag: ).

[0065] This step is the foundation of supervised learning. Merchants are divided into two categories by labels. The pointscard model can learn the relationship between features and labels and use this knowledge to reason and predict new data, thereby achieving effective assessment and management of merchant risks.

[0066] The third step is to discretize and bin the continuous features.

[0067] For example, traditional models such as logistic regression are less stable for continuous features (e.g., sensitive to outliers and the influence of units), so continuous features (e.g., the scoring results of "historical penalty risk") need to be discretized into several intervals (bins).

[0068] Discretization binning is performed based on the static risk scores of multiple historical merchants to obtain the evidence weight value shared by the static risks of multiple historical merchants. Discretization binning includes dividing the static risk scores of multiple historical merchants into several discrete intervals, with each interval corresponding to a bin.

[0069] In this embodiment, the target item is "historical penalty risk". Assuming the score for "historical penalty risk" is a standardized continuous value of 0-1 (0 representing no risk, 1 representing extremely high risk), it is divided into three equally spaced bins: Bin 1: (Low risk), Container 2: (Medium risk), Container 3: (High risk). After binning, the score result of each "historical penalty risk" will be mapped to the corresponding bin (e.g., 0.75 belongs to bin 3).

[0070] The fourth step is to calculate the evidence weight value for each bin.

[0071] For example, the Weight of Evidence (WOE) is used to quantify the contribution of each bin to the "high risk" label, and the formula is as follows:

[0072] (3)

[0073] In equation (3), Indicates the first The evidence weight values ​​in each bin.

[0074] Key meanings: WOE > 0: The proportion of high-risk items in this container is higher than the overall average (higher risk); WOE < 0: The proportion of high-risk items in this container is lower than the overall average (lower risk); The larger the absolute value of WOE, the more significant the risk difference.

[0075] For example, the implementation of this process includes:

[0076] First, calculate the percentage of high-risk merchants within each container.

[0077] The percentage of high-risk merchants within the sub-container is as follows: ,

[0078] But more importantly, the proportion of high-risk merchants within each sub-container out of all high-risk merchants in the overall market is: .

[0079] Secondly, calculate the percentage of non-high-risk merchants within each bin.

[0080] The proportion of non-high-risk merchants in the sub-boxes out of the total number of non-high-risk merchants is: .

[0081] Furthermore, calculate the ratio between the two.

[0082] The core of WOE (Women of Risk) is comparing the ratio of "the relative proportion of high-risk items within a container" to "the relative proportion of non-high-risk items within a container": .

[0083] Finally, the natural logarithm is taken to obtain the static risk characteristics of multiple historical merchants: the WOE value shared by "historical penalty risk". .

[0084] It is understandable that the WOE value shared by multiple historical merchants for static risk includes the WOE values ​​corresponding to each of the several sub-boxes.

[0085] The fifth step is to convert the original feature values ​​into WOE values.

[0086] For example, the discretized bins are themselves categorical variables, and the original values ​​need to be replaced with WOE values ​​so that the model can directly utilize the risk information of the bins.

[0087] The WOE value corresponding to the bin in which the static risk score of a historical merchant falls is used as the WOE value of the static risk of that historical merchant.

[0088] For example, if a merchant's "historical penalty risk" score is 0.75, it belongs to bin 3 ( Then the WOE value of sub-box 3 will be used as the WOE value of the "historical penalty risk" characteristic of this historical merchant.

[0089] Step 6: Establish a logistic regression model.

[0090] For example, logistic regression uses a linear combination of the historical data of a merchant. Each static feature corresponds to Value, combined with model coefficients ( This outputs the log-odds probability of a "high risk" condition. The formula for calculating the log-odds probability is:

[0091] (4)

[0092] in, This indicates the probability that a particular historical merchant is high-risk. This represents the intercept term (benchmark risk). This represents the coefficient of the WOE value corresponding to the scoring results of each characteristic of a historical merchant (reflecting the marginal impact of binning on risk).

[0093] Step 7: Substitute the coefficients to calculate the comprehensive risk score.

[0094] For example, suppose the initial parameters of the model have been trained using historical merchant data (example values):

[0095] (intercept); (WOE coefficient of compartment 1); (WOE coefficient of compartment 2);

[0096] (WOE coefficient of compartment 3).

[0097] If the WOE values ​​corresponding to the three static characteristics of a historical merchant are 0.98, 0.85, and 1.3863 respectively, then:

[0098] ,

[0099] Calculation process: , , .

[0100] sum: (Approximately 3.262).

[0101] Therefore, based on the WOE values ​​corresponding to each static risk of this historical merchant and the initial / current parameters of the pointscard model, the comprehensive risk score of this historical merchant is obtained. At this point, the z-value of the comprehensive risk score is the logarithmic probability.

[0102] Based on the risk / non-risk labels and comprehensive risk scores of multiple historical merchants, the scoring card model is trained with parameter tuning to improve the model's inference accuracy.

[0103] In one implementation, the process of obtaining a comprehensive risk score using a logistic regression-based scorecard model can be carried out through the following steps:

[0104] By using the business operation data of multiple merchants, the various static risks of the merchant are scored separately, and the scores are obtained for each.

[0105] For each merchant, the static risk score of that merchant is input into the scoring card model to obtain the log odds corresponding to the merchant's comprehensive risk score.

[0106] The logarithmic odds corresponding to the comprehensive risk score are converted into a probability using the Sigmoid function, and this probability is used as the merchant's final comprehensive risk score.

[0107] For example, the logarithmic odds (z-value) needs to be converted into a probability value between 0 and 1 using the Sigmoid function. The formula is:

[0108] (5)

[0109] Example calculation: when hour, .

[0110] The formula for converting to a percentage system is as follows:

[0111] (6)

[0112] According to formula (6), the final percentage is 96.3%, which means the merchant... The probability of being high-risk Approximately 96.3%.

[0113] Step S205: Based on the random walk algorithm and the geographical coordinates of each merchant, calculate the network risk impact value of the merchant spreading the risk to other merchants among the multiple merchants.

[0114] For example, among the multiple merchants, there is at least one high-risk merchant with a historical number of penalties exceeding a preset number; based on the random walk algorithm and the geographical coordinates of each of the multiple merchants, the network risk impact value of this merchant spreading risk to other merchants among the multiple merchants is calculated, including:

[0115] A relationship graph is established by treating multiple merchants as nodes and their geographical coordinates as attributes of the corresponding nodes. Starting from any one of the at least one high-risk merchants, a random walk algorithm is used to perform a predetermined number of random walks in the relationship graph to obtain the number of visits to that merchant. If that merchant is one of the at least one high-risk merchants, the number of visits to that merchant is amplified by a predetermined factor to obtain an updated number of visits. The number of visits to that merchant is normalized based on the maximum number of visits among the multiple merchants to obtain the network risk impact value of that merchant.

[0116] In one implementation, during the relationship graph construction phase, the geographical distance between any two merchants is calculated using GeoPy, based on their geographical coordinates. If the distance is no more than 100 kilometers, a "NEARBY" relationship is established between the two merchants in Neo4j, thus forming a spatial graph structure. This structure is used to model potential risk propagation chains.

[0117] In the risk propagation calculation phase, a random walk algorithm is introduced. Starting from a designated high-risk merchant, a multi-step (e.g., 500 steps) random walk is performed, recording the cumulative risk count of visited nodes at each step. During the walk, if the current node is a key merchant (e.g., with more than 3 historical penalties), its cumulative risk value is multiplied by 1.5, indicating that its risk has stronger propagation potential. Finally, the visit counts of all nodes are normalized to generate the network risk value for each merchant, i.e.:

[0118] (7)

[0119] in, For merchants Total number of visits This represents the maximum number of visits for a merchant among multiple merchants.

[0120] Step S206: The merchant's final risk score is obtained by weighted summation of the regional deviation risk score, network risk impact value, comprehensive risk score, and static risk score.

[0121] All risk sources are weighted and aggregated to calculate the merchant. Final risk score :

[0122] (8)

[0123] Weight example: All sub-risk indicators have been normalized to The interval ensures additivity.

[0124] For example, the underlying risk The result of a weighted calculation of the three static indicators:

[0125] .

[0126] in, , Merchants The scoring results for the legal entity's cross-provincial risks, license expiration risks, and historical penalty risks. It is understandable that the formula... , , The weighting coefficients can be adjusted arbitrarily.

[0127] Finally, the model maps the comprehensive risks onto a spatial heat map through a geographic visualization module, making the geographical location, frequency of penalties, and risk level of merchants clear at a glance, providing tobacco regulatory authorities with an intuitive basis for identifying key merchants and making inspection decisions.

[0128] For example, the risk visualization module 40, with its powerful data visualization capabilities, provides users with an intuitive and highly interactive analysis tool. Through visualized reports, users can use map drill-down functionality to explore in-depth details of geographic spatial distribution, use line trend charts to clearly present the dynamic changes of data over time, and use pie charts to intuitively display the proportion and structure of data. These interactive visualization methods help users analyze current tobacco merchant data from multiple dimensions, thereby gaining a more comprehensive understanding of the business logic and potential risks behind the data.

[0129] In summary, this application proposes a solution that integrates business license data, tobacco monopoly license data, and administrative penalty data to construct a multi-level merchant relationship network centered on geographical association and spatiotemporal clustering; it achieves intelligent identification of violations such as cross-regional tobacco trafficking, unlicensed operation, and discrepancies between licenses and permits based on dynamic evidence weight rules; and it displays regional risk heat maps through visualization functions, forming a closed-loop regulatory system from single-point merchant inspection to precise crackdown on related networks.

[0130] For example, Figure 3 This is a flowchart illustrating a training method for a scorecard model provided in an embodiment of this application. Figure 3 As shown, this training method can be implemented through the following steps:

[0131] Step S301: Obtain the scoring results of various static risks for multiple historical merchants; each static risk includes at least one of the following: cross-province risk of legal entity, license expiration risk, and historical penalty risk;

[0132] For each historical merchant, in step S302, based on the scoring results of the target items in each static risk, a risk / non-risk label is assigned to the historical merchant.

[0133] For each static risk of this historical merchant,

[0134] Discretization binning is performed based on the scoring results of this static risk from multiple historical merchants to obtain the evidence weight value shared by multiple historical merchants for this static risk. Discretization binning includes dividing the scoring results of this static risk from multiple historical merchants into several discrete intervals, with each interval corresponding to a bin.

[0135] Based on the evidence weight value shared by multiple historical merchants for this static risk, the evidence weight value corresponding to this historical merchant for this static risk is obtained.

[0136] For example, the evidentiary weight value corresponding to the bin in which the score result of this static risk of the historical merchant falls is used as the evidentiary weight value of this static risk of the historical merchant.

[0137] Step S303: Based on the evidence weight values ​​corresponding to the various static risks of the historical merchant and the initial parameters of the points card model, obtain the comprehensive risk score of the historical merchant.

[0138] Step S304: Based on the risk / non-risk labels and comprehensive risk scores of multiple historical merchants, the scoring card model is trained and its parameters are adjusted.

[0139] The model parameter tuning and training process in this training method can be found in the description of step S204, and will not be repeated here.

[0140] Finally, by tuning and training the scoring card model, the trained model was deployed to the production environment for risk scoring of new merchant data.

[0141] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. Furthermore, in some possible implementations, each step in the above embodiments may be selectively executed according to actual circumstances; it may be partially or fully executed, without limitation here. Additionally, all or part of any feature in the above embodiments can be freely and arbitrarily combined without contradiction. The combined technical solutions are also within the scope of this application.

[0142] Based on the methods in the above embodiments, this application provides an electronic device. The electronic device may include: at least one memory for storing a program; and at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor executes the methods described in the above embodiments. Exemplarily, the electronic device may be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, server, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), or artificial intelligence (AI) device. This application does not impose any special limitations on the specific type of the electronic device.

[0143] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0144] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. It should be understood that in the embodiments of this application, the order of the process numbers does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0145] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this application.

Claims

1. A method for identifying risks of a merchant based on multi-dimensional feature quantization, characterized in that, The method comprises: obtaining merchant data and industry and commerce data of a plurality of merchants; processing the merchant data and the industry and commerce data to obtain merchant operation data; performing target processing on each merchant in the plurality of merchants, the target processing comprising: scoring each static risk and regional deviation risk of the merchant according to the merchant operation data of the merchant and the merchant operation data of each of the plurality of merchants, to obtain respective scoring results; the static risks comprise at least one of the following: cross-provincial risk of a legal person, license expiration risk, and historical punishment risk; the regional deviation risk is a degree of deviation between sales performance of the merchant in a specific region and average sales performance in the region where the merchant is located; inputting the scoring results of the static risks of the merchant into a scoring card model to obtain a comprehensive risk score of the merchant; the scoring card model is obtained based on supervised learning of scoring results of static risks of a plurality of historical merchants; calculating a network risk influence value of the merchant to other merchants in the plurality of merchants according to a random walk algorithm and geographical coordinates of each of the plurality of merchants; performing weighted summation on the scoring results of the regional deviation risks, the network risk influence value, the comprehensive risk score, and the scoring results of the static risks to obtain a final risk score of the merchant.

2. The method of claim 1, wherein, the cross-provincial risk score of the legal person is a first preset non-zero value, indicating that the registered address of the merchant and the registered address of the legal person are in different provinces; and the cross-provincial risk score of the legal person is zero, indicating that the registered address of the merchant and the registered address of the legal person are in the same province.

3. The method of claim 1, wherein, the license expiration risk score is a second preset non-zero value, indicating that the expiration time of the license of the merchant is not greater than a preset threshold; and the license expiration risk score is zero, indicating that the expiration time of the license of the merchant is greater than the preset threshold.

4. The method of claim 1, wherein, the scoring of each static risk and regional deviation risk of the merchant using the merchant operation data of the plurality of merchants comprises: normalizing the number of historical punishments of a merchant in the plurality of merchants according to the maximum number of punishments of the merchant to obtain a score of the historical punishment risk of the merchant.

5. The method of claim 1, wherein a product category structure of the plurality of merchants forms a product category set; the scoring of each static risk and regional deviation risk of the merchant using the merchant operation data of the plurality of merchants comprises: obtaining a score of the regional deviation risk of the merchant according to a sales proportion of the merchant on each brand in the product category set and an average sales proportion of each brand in the brand set in a region where the merchant is located.

6. The method of claim 1, wherein the plurality of merchants includes at least one high-risk merchant with a number of historical punishments exceeding a preset number; the calculation of a network risk influence value of the merchant to other merchants in the plurality of merchants according to a random walk algorithm and geographical coordinates of each of the plurality of merchants comprises: establishing a relationship graph by taking the plurality of merchants as nodes and taking geographical coordinates of each merchant as attributes of the corresponding node; and starting from any of the at least one high-risk merchant, performing random walk in the relationship graph by using a random walk algorithm for a preset number of steps to obtain a visit frequency of the merchant; when the merchant is one of the at least one high-risk merchant, amplifying the visit frequency of the merchant by a preset multiple to obtain an updated visit frequency; normalizing the visit frequency of the merchant according to a maximum visit frequency of the merchants in the plurality of merchants to obtain a network risk influence value of the merchant.

7. The method of claim 1, wherein the inputting the score result of the static risk of the merchant into the scoring card model to obtain a comprehensive risk score of the merchant comprises: inputting evidence weight values of each static risk of the merchant into the scoring card model to output a log odds of the merchant being a high-risk merchant; calculating the log odds by using a Sigmoid function to obtain a comprehensive risk score of the merchant.

8. A method for training a scorecard model, the method comprising: The method comprises: obtaining score results of each static risk of a plurality of historical merchants; the static risks comprise at least one of the following: a cross-provincial risk of a legal person, an expiration risk of a license, and a historical penalty risk; for each historical merchant, assigning a risk / non-risk label to the historical merchant based on a score result of a target static risk in the static risks; for each static risk of the historical merchant, performing discretization binning processing on the score results of the static risk of the plurality of historical merchants to obtain evidence weight values commonly used by the static risk of the plurality of historical merchants; the discretization binning processing comprises dividing the score results of the static risk of the plurality of historical merchants into a plurality of discrete intervals, each interval corresponding to a bin; obtaining an evidence weight value corresponding to the static risk of the historical merchant according to the evidence weight values commonly used by the static risk of the plurality of historical merchants; obtaining a comprehensive risk score of the historical merchant based on the evidence weight values respectively corresponding to the static risks of the historical merchant and initial parameters of the scoring card model; training the scoring card model by adjusting parameters according to the risk / non-risk labels and the comprehensive risk scores of the plurality of historical merchants.

9. The method of claim 8, wherein the evidence weight values commonly used by the static risk of the plurality of historical merchants comprise evidence weight values respectively corresponding to the plurality of bins; the obtaining of the evidence weight value of the static risk of the historical merchant according to the evidence weight values commonly used by the static risk of the plurality of historical merchants comprises: taking the evidence weight value corresponding to the bin into which the score result of the static risk of the historical merchant falls as the evidence weight value of the static risk of the historical merchant. 10.A merchant risk identification system based on multi-dimensional feature quantization, characterized in that, The system comprises: an acquisition module configured to acquire merchant data and business data of a plurality of merchants; a preprocessing module configured to process the merchant data and the business data to obtain merchant operation data; the merchant operation data comprises business license information, business license information, administrative penalty codes, and enterprise unified social credit codes of the merchants; A risk identification module is configured to perform target processing on each of the plurality of merchants, the target processing comprising: scoring each of the static risks and the regional deviation risk of the merchant based on the merchant operating data of the merchant and the merchant operating data of each of the plurality of merchants, respectively, to obtain respective scoring results; the static risks comprise at least one of the following: a legal person cross-provincial risk, a license expiration risk, and a historical punishment risk; the regional deviation risk is a degree of deviation between a sales performance of the merchant in a specific region and an average sales performance in the region where the merchant is located; inputting the scoring results of the static risks of the merchant into a scoring card model to obtain a comprehensive risk score of the merchant; the scoring card model is obtained based on supervised learning of scoring results of static risks of a plurality of historical merchants; calculating, according to a random walk algorithm and geographical coordinates of each of the plurality of merchants, a network risk influence value of the merchant to other merchants in the plurality of merchants; performing weighted summation on the scoring results of the regional deviation risk, the network risk influence value, the comprehensive risk score, and the scoring results of the static risks to obtain a final risk score of the merchant.

Citation Information

Patent Citations

  • Merchant dynamic management and control method and device, server and readable storage medium

    CN110675029A

  • Operation risk identification method

    CN113962514A

  • Cigarette retail customer classification method based on big data thinking

    CN115587820A

  • Merchant fraudulent transaction detection method based on graph

    CN116245543A

  • Credit evaluation model construction method and credit evaluation method

    CN120782540A