Data processing method and device for a risk control system, and electronic device

By improving the data processing methods of the risk control system and using a pre-defined dimensionality reduction algorithm to determine the similarity of the data set, the challenges caused by the direct input of massive amounts of data were solved, and effective data identification and sharing were achieved, thereby improving the efficiency and cost-effectiveness of the risk control system.

CN112907065BActive Publication Date: 2026-05-05WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WEBANK (CHINA)
Filing Date
2021-02-10
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Using massive amounts of data from various data sources directly as the source data for the risk control system in existing technologies will pose greater challenges to the data processing capabilities, maintenance costs, timeliness, and effectiveness of the risk control system, and may even have serious negative impacts.

Method used

By acquiring the first data from different target data sources, determining its data type and quantifying it, and using a preset dimensionality reduction algorithm to determine the similarity between each target data set and the preset data set, the source data is determined based on the similarity. Classification, labeling, and storage operations are then performed to reduce redundancy and form a standardized data pool.

Benefits of technology

It reduces the data processing workload of the risk control system when implementing risk control strategies, ensures the effectiveness and timeliness of risk control, reduces maintenance costs, and promotes the effective accumulation and sharing of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112907065B_ABST
    Figure CN112907065B_ABST
Patent Text Reader

Abstract

This application provides a data processing method, apparatus, and electronic device for a risk control system. The data processing method first acquires first data about multiple risk control objects from different target data sources. Then, it determines the data type of the first data, quantifies it according to the data type, and determines the second data to obtain the target data set. Finally, it determines the similarity between each target data set and a preset data set using a pre-defined dimensionality reduction algorithm. This allows the risk control system to determine the source data from the target data sources based on the similarity scores. It effectively identifies data redundancy from each target data source based on the similarity scores, thus determining the source data required by the system. This reduces the workload of processing massive amounts of data when implementing risk control strategies, ensures the effectiveness and timeliness of risk control, effectively reduces maintenance costs, and facilitates data sharing mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial technology (Fintech), and in particular to a data processing method, apparatus and electronic device for a risk control system. Background Technology

[0002] With the rapid development of computer and internet technologies, Fintech, as a product of the deep integration of finance and technology, is currently becoming a hot topic for innovation and development in the financial industry. Furthermore, micro and small enterprises now account for over 70% of all enterprises, leading to a surge in their numbers. For financial institutions, this necessitates building corresponding risk control systems based on big data for each micro and small enterprise, and implementing appropriate risk control strategies through these systems.

[0003] However, with the rapid development of big data technology, the amount of data from various data sources used to characterize the various features of micro and small enterprises is growing exponentially. To fully reach these enterprises, risk control systems typically receive data on various features to ensure effective risk control for micro and small enterprises. However, due to the sheer volume and complexity of this massive amount of data, with varying quality and value density, directly using data from various data sources as the input to the risk control system in current applications poses greater challenges to its data processing capabilities, maintenance costs, timeliness, strategies, and effectiveness. It can even lead to more serious negative impacts. For example, redundant data from data sources may exist, but the risk control system's inability to identify this can result in redundant data processing. Furthermore, the system may be unable to retain effective data, thus hindering the formation of a data sharing mechanism.

[0004] It is evident that a solution is urgently needed to overcome the various problems that exist in existing technologies that directly use massive amounts of data from various data sources as the source data for risk control systems. Summary of the Invention

[0005] This application provides a data processing method, apparatus, and electronic device for a risk control system, which addresses the technical problem that directly using massive amounts of data from various data sources as the source data for a risk control system can pose greater challenges or even cause serious negative impacts on the system.

[0006] Firstly, this application provides a data processing method for a risk control system, including:

[0007] Acquire first data about multiple risk control objects from different target data sources, wherein the first data is used to characterize the preset feature information of the risk control objects;

[0008] The data type of the first data is determined so as to determine the second data based on the data type. The second data is used to represent the first data after it has been numericalized. The data type includes numeric type and non-numeric type.

[0009] The similarity between the second data corresponding to each target data source and the data of the preset target dimension indicator is determined according to the preset dimensionality reduction algorithm, so that the risk control system can determine the source data based on the similarity.

[0010] In one possible design, after determining the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm, the method further includes:

[0011] The first data provided by the target data source is classified according to the similarity and preset business requirements, and the classified first data is generated into different data regions according to preset logical rules, so as to determine the usage rights and scope of the first data through the different data regions; and / or

[0012] The target data sources are labeled according to the similarity and preset labeling rules to standardize the recording of each target data source; and / or

[0013] Based on the similarity, the post source data that satisfies the preset business scenario is determined from the first data provided by the target data source.

[0014] In one possible design, after determining the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm, the method further includes:

[0015] Based on the similarity, the first data provided by the target data source is stored or deleted to form a standardized data pool.

[0016] In one possible design, determining the data type of the first data to determine the second data based on the data type includes:

[0017] If the data type of the first data is the numeric type, then the first data itself is determined as the second data;

[0018] If the data type of the first data is the non-numerical type, then the first data is processed according to a preset quantization algorithm to obtain the second data.

[0019] In one possible design, the step of processing the first data according to a preset quantization algorithm to obtain the second data includes:

[0020] If the first data of the non-numerical type is represented in non-encoded text, then the second data is obtained by processing the first data according to a preset word segmentation algorithm, wherein the preset quantization algorithm includes the preset word segmentation algorithm;

[0021] If the first data of the non-numerical type is represented by characters or coded text, then the second data is obtained by processing the first data according to a preset encoding rule. The preset encoding rule contains a unique correspondence between the characters or coded text and the second data. The preset quantization algorithm includes the preset encoding rule.

[0022] In one possible design, the step of processing the first data according to a preset word segmentation algorithm to obtain the second data includes:

[0023] The first data and the data of the preset target dimension indicator are segmented according to the preset word segmentation algorithm to obtain the target word segmentation number and the target word segmentation number respectively.

[0024] Determine the ratio between the number of target word segments and the number of target word segments, so that the ratio is determined as the second data corresponding to the first data;

[0025] Wherein, the target word count is used to count the words contained in the first data, and the target word count is used to count the words contained in the data indicated by the preset target dimension.

[0026] In one possible design, determining the similarity between the second data corresponding to each target data source and the data of the preset target dimension indicator according to a preset dimensionality reduction algorithm includes:

[0027] Obtain the first expected value of each target data set and the second expected value of the preset data set;

[0028] Obtain the first standard deviation of each target data set and the second standard deviation of the preset data set;

[0029] Based on the preset dimensionality reduction algorithm, each first expected value, each first standard deviation, the second expected value, and the second standard deviation, the distance data between each first standard deviation and the second standard deviation is determined, so that each distance data is determined as the similarity between the corresponding target data set and the preset data set.

[0030] In one possible design, after determining the distance data as the similarity between the target data set and the preset data set, the method further includes:

[0031] If the similarity is zero, then it is determined that the first data provided by the current target data source is completely redundant with the data of the preset target dimension indicator;

[0032] If the similarity is not zero and is less than the preset redundancy threshold, then the data redundancy between the first data provided by the current target data source and the preset target dimension indicator is determined.

[0033] If the similarity is greater than a preset redundancy threshold, it is determined that the first data provided by the current target data source is not related to the data of the preset target dimension indicator.

[0034] In one possible design, before determining the similarity between each target dataset and the preset dataset according to the preset dimensionality reduction algorithm, the following is also included:

[0035] An improvement is made to the preset random neighborhood filling algorithm to adjust the variance of the normal distribution centered on the sample data in the preset random neighborhood filling algorithm to the variance of the corresponding data set, thus obtaining the preset dimensionality reduction algorithm.

[0036] In one possible design, the preset feature information includes at least one of the following: business registration information, tax information, credit information, judicial information, asset information, operating income information, and suspicious information related to illegal fund transfers of the risk control object.

[0037] Secondly, this application provides a data processing device for a risk control system, comprising:

[0038] The acquisition module is used to acquire first data about multiple risk control objects provided by different target data sources. The first data is used to characterize the preset feature information of the risk control objects.

[0039] The first processing module is used to determine the data type of the first data, and to determine the second data based on the data type to obtain a target data set. The second data is used to represent the first data after numericalization. The data type includes numerical type and non-numerical type.

[0040] The second processing module is used to determine the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm, so that the risk control system can determine the source data based on the similarity. The preset data set includes data of preset target dimensionality indicators.

[0041] In one possible design, the data processing device for the risk control system further includes: a third processing module; the third processing module is used for:

[0042] The first data provided by the target data source is classified according to the similarity and preset business requirements, and the classified first data is generated into different data regions according to preset logical rules, so as to determine the usage rights and scope of the first data through the different data regions; and / or

[0043] The target data sources are labeled according to the similarity and preset labeling rules to standardize the recording of each target data source; and / or

[0044] Based on the similarity, the post source data that satisfies the preset business scenario is determined from the first data provided by the target data source.

[0045] In one possible design, the data processing device for the risk control system further includes: a fourth processing module; the fourth processing module is used for:

[0046] Based on the similarity, the first data provided by the target data source is stored or deleted to form a standardized data pool.

[0047] In one possible design, the first processing module is specifically used for:

[0048] If the data type of the first data is the numeric type, then the first data itself is determined as the second data;

[0049] If the data type of the first data is the non-numerical type, then the first data is processed according to a preset quantization algorithm to obtain the second data.

[0050] In one possible design, the first processing module further includes:

[0051] The first quantization processing unit is configured to process the first data according to a preset word segmentation algorithm to obtain the second data if the first data of the non-numerical type is represented in non-encoded text, wherein the preset quantization algorithm includes the preset word segmentation algorithm.

[0052] The second quantization processing unit is configured to process the first data according to a preset encoding rule to obtain the second data if the first data of the non-numerical type is represented by characters or coded text. The preset encoding rule contains a unique correspondence between the characters or coded text and the second data, and the preset quantization algorithm includes the preset encoding rule.

[0053] In one possible design, the first quantization processing unit is specifically used for:

[0054] The first data and the data of the preset target dimension indicator are segmented according to the preset word segmentation algorithm to obtain target word segmentation and target word segmentation respectively.

[0055] The proportion of the target word segment in the target word segment is determined, and the proportion is used as the second data corresponding to the first data.

[0056] In one possible design, the second processing module is specifically used for:

[0057] Obtain the first expected value of each target data set and the second expected value of the preset data set;

[0058] Obtain the first standard deviation of each target data set and the second standard deviation of the preset data set;

[0059] Based on the preset dimensionality reduction algorithm, each first expected value, each first standard deviation, the second expected value, and the second standard deviation, the distance data between each first standard deviation and the second standard deviation is determined, so that each distance data is determined as the similarity between the corresponding target data set and the preset data set.

[0060] In one possible design, the data processing device for the risk control system further includes: a fifth processing module; the fifth processing module is used for:

[0061] If the similarity is zero, then it is determined that the first data provided by the current target data source is completely redundant with the data of the preset target dimension indicator;

[0062] If the similarity is not zero and is less than the preset redundancy threshold, then the data redundancy between the first data provided by the current target data source and the preset target dimension indicator is determined.

[0063] If the similarity is greater than a preset redundancy threshold, it is determined that the first data provided by the current target data source is not related to the data of the preset target dimension indicator.

[0064] In one possible design, the data processing device for the risk control system further includes: a sixth processing module; the sixth processing module is used for:

[0065] An improvement is made to the preset random neighborhood filling algorithm to adjust the variance of the normal distribution centered on the sample data in the preset random neighborhood filling algorithm to the variance of the corresponding data set, thereby obtaining the preset dimensionality reduction algorithm. The corresponding data set includes each target data set and the preset data set.

[0066] In one possible design, the preset feature information includes at least one of the following: business registration information, tax information, credit information, judicial information, asset information, operating income information, and suspicious information related to illegal fund transfers of the risk control object.

[0067] Thirdly, this application provides an electronic device, comprising:

[0068] Processor; and

[0069] Memory for storing the computer program of the processor;

[0070] The processor is configured to execute any of the possible data processing methods for a risk control system provided in the first aspect by executing the computer program.

[0071] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, the computer program being used to execute any of the possible data processing methods for a risk control system provided in the first aspect.

[0072] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the possible data processing methods for a risk control system provided in the first aspect.

[0073] This application provides a data processing method, apparatus, and electronic device for a risk control system. The data processing method first acquires first data about multiple risk control objects from different target data sources, where the first data represents preset characteristic information of the risk control objects. Then, it determines the data type of the first data and determines second data based on the determined data type. The second data is used to quantify the first data, and the data type includes both numerical and non-numerical types. Finally, it determines the similarity between the second data corresponding to each target data source and the data of the preset target dimension indicator based on a preset dimensionality reduction algorithm. This allows the risk control system to determine the source data based on the obtained similarities, enabling the risk control system to effectively identify data redundancy provided by each target data source based on similarity, thereby determining the source data it needs. This not only reduces the processing workload of massive amounts of data when implementing risk control strategies, but also ensures the effectiveness and timeliness of risk control, effectively reduces the maintenance cost of the risk control system, and facilitates the effective accumulation of data to form a data sharing mechanism. Attached Figure Description

[0074] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and application, and together with the description serve to explain the principles of this disclosure and application.

[0075] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application;

[0076] Figure 2 A flowchart illustrating a data processing method provided in an embodiment of this application;

[0077] Figure 3 A flowchart illustrating another data processing method provided in an embodiment of this application;

[0078] Figure 4 A flowchart illustrating another data processing method provided in an embodiment of this application;

[0079] Figure 5 A flowchart illustrating another data processing method provided in an embodiment of this application;

[0080] Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0081] Figure 7 This is a schematic diagram of another data processing apparatus provided in an embodiment of this application;

[0082] Figure 8 This is a schematic diagram of the structure of a processing module provided in an embodiment of this application;

[0083] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0084] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods and apparatus consistent with some aspects of this application as detailed in the appended claims.

[0085] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0086] Currently, with the rapid development of big data technology, the amount of data from various data sources that can characterize various features of micro and small enterprises is growing exponentially. To fully reach these enterprises, financial industry risk control systems typically receive as much feature information as possible from data sources to ensure effective risk control for micro and small enterprises. However, this massive amount of data is diverse and varied, with inconsistent data quality and value density. Directly using data from various data sources as the input to the risk control system in current technologies poses greater challenges to the system's data processing capabilities, maintenance costs, timeliness, strategies, and effectiveness, and may even lead to more serious negative impacts. For example, it may fail to identify redundancy in the data provided by the data sources, resulting in redundant data processing, or the risk control system may be unable to retain effective data and establish a data sharing mechanism. Given these problems existing in the risk control systems of existing financial institutions when implementing risk control strategies based on big data, an effective solution is urgently needed to overcome these issues.

[0087] This application provides a data processing method, apparatus, and electronic device for a risk control system. The inventive concept of the data processing method for a risk control system provided in this application is as follows: before the source data provided by each target data source to the risk control system, the first data provided by each target data source is processed accordingly to obtain the similarity between the data set composed of the first data provided by each target data source and the data set composed of preset target dimension indicators. This allows the risk control system to determine its required source data based on the similarity. The similarity can reflect the redundancy between the first data provided by the target data source and the data of the preset target dimension indicators. Therefore, the risk control system can effectively identify the redundancy between the first data provided by the target data source and the data of the preset target dimension indicators, thereby determining its required source data. This differs from the prior art where massive amounts of data are directly input into the risk control system as source data. This not only reduces the processing workload of the risk control system when implementing risk control strategies, but also ensures the effectiveness and timeliness of risk control, thereby effectively reducing the maintenance cost of the risk control system and facilitating the effective accumulation of data to form a data sharing mechanism.

[0088] The following describes exemplary application scenarios of the embodiments of this application.

[0089] Figure 1 This is a schematic diagram of an application scenario provided in an embodiment of this application, such as... Figure 1 As shown, the network serves as a medium for providing a communication link between terminal device 11 and server 12. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables. Terminal device 11 and server 12 can interact through the network to receive or send messages. Terminal device 11 can be any terminal capable of obtaining various preset characteristic information representing the risk control object, such as business registration information, tax information, credit information, judicial information, asset information, operating income information, and suspicious information regarding illegal fund transfers. Terminal device 11 can be configured at the risk control object itself or at a third-party institution, etc., and this embodiment does not limit this. The function of terminal device 11 is to provide first data. Server 12 is an electronic device corresponding to the data processing device for the risk control system that can execute the data processing method for the risk control system provided in the embodiments of this application. Server 12 can be configured in the corresponding equipment of the financial institution to which the risk control system belongs. Server 12 obtains first data about multiple risk control objects provided by different target data sources from terminal device 11, and then executes the data processing method for the risk control system provided in the embodiments of this application to complete the determination of the source data, so as to facilitate the risk control system to implement risk control strategies for the risk control objects.

[0090] It should be noted that the embodiments of this application do not limit the type of terminal device 11 described above. For example, terminal device 11 can be a computer, smartphone, smart glasses, smart bracelet, smartwatch, tablet computer, etc. Figure 1 The terminal device 11 is illustrated as a computer. The server 12 can also be a server cluster, but this embodiment does not limit this.

[0091] It should be noted that the above application scenarios are merely illustrative, and the data processing methods for risk control systems provided in this application include, but are not limited to, the above application scenarios.

[0092] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0093] Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 2 As shown, the data processing method for a risk control system provided in this embodiment includes:

[0094] S101: Obtain first data about multiple risk control objects from different target data sources.

[0095] The first data is used to characterize the preset features of the risk control object.

[0096] The risk control target refers to the object on which a financial institution implements risk control strategies using its risk control system, such as the various micro and small enterprises for which the financial institution implements risk control. The first data refers to data that can characterize the preset feature information of the risk control target. This preset feature information can be one or more of the following: business registration information, tax information, credit information, legal information, asset information, operating income information, and information indicating suspicious illegal fund transfers. Therefore, the first data can be various data representing the preset feature information of each micro and small enterprise, such as business registration information, tax information, credit information, legal information, and asset information. It should be noted that the preset feature information of the risk control target includes, but is not limited to, the information listed above. It can also include other information required by the financial institution when implementing risk control strategies on the risk control target; however, this embodiment does not limit this.

[0097] The target data source refers to various data sources that can provide first-hand data, such as official tax platforms, business registration platforms, or other big data providers and other third-party organizations. There are no restrictions on the nature of the third-party organizations.

[0098] Furthermore, the aforementioned first data refers to data from different target data sources regarding multiple risk control objects, not data for a single risk control object. This step involves obtaining the first data from different target data sources regarding multiple risk control objects. The acquisition method can be downloading from different target data sources, or the data being provided by the risk control object itself, etc. This embodiment does not limit the specific method used.

[0099] S102: Determine the data type of the first data, and determine the second data based on the data type to obtain the target data set.

[0100] The second data is used to quantify the first data, and the data types include numeric and non-numeric types.

[0101] The initial data provided by different target data sources may be the same preset feature information for the risk control object, or it may be different preset feature information for the risk control object. Furthermore, due to individual differences between different target data sources, the data type of the initial data may be numeric or non-numeric. In other words, the initial data may be represented by numbers or by non-numeric forms, such as by letters or Chinese characters, etc.

[0102] Therefore, after obtaining the first data about multiple risk control objects from different target data sources, it is necessary to determine the data type of the first data in order to further quantify the first data according to the data type, obtain the corresponding second data, and determine the set of the second data obtained from each target data source as the target data set, thereby obtaining each target data set, where the second data is used to represent the quantified first data.

[0103] In one possible design, step S102 could be implemented as follows: Figure 3 As shown. Figure 3 This is a flowchart illustrating another data processing method provided in an embodiment of this application. Figure 3 As shown, in the data processing method for a risk control system provided in this embodiment, determining the data type of the first data to determine the second data based on the data type includes:

[0104] S201: Determine the data type of the first data.

[0105] The data type of the first data provided by different target data sources is determined, including numeric and non-numeric types. If the determination result is numeric, proceed to step S202. If the determination result is non-numeric, proceed to step S203.

[0106] S202: If the data type of the first data is numeric, then the first data itself is determined as the second data.

[0107] If the first data is represented in numerical form, it indicates that the data type of the first data is numeric, and therefore the first data itself is identified as the corresponding second data. In other words, the first data itself is retained, and it is not converted into a numerical value.

[0108] Typically, some pre-defined characteristic information, such as operating revenue, is directly represented by its numerical value. However, some pre-defined characteristic information is also represented by numbers, but these numbers may represent corresponding levels. For example, tax information usually uses numbers like "1, 2, 3, 4, ..." to represent the corresponding levels. When the data type is numeric, regardless of whether the first data is the numerical value corresponding to the pre-defined characteristic information itself or other numerical values ​​with a certain mapping relationship, if the first data is represented in numerical form, then the first data is not numericalized; instead, the first data itself is determined as the second data.

[0109] S203: If the data type of the first data is a non-numeric type, then the first data is processed according to the preset quantization algorithm to obtain the second data.

[0110] If the first data is not represented in numerical form, it indicates that the data type of the first data is non-numerical. In this case, the first data needs to be quantized according to the preset quantization algorithm to obtain the corresponding second data.

[0111] For example, the first data provided by some target data sources may be represented in the form of "A, B, C, ...", "I, II, III, IV, ...", "One, Two, Three, Four, ...", etc. The data corresponding to the first data may also need to be expressed through text description. Therefore, it is necessary to process the first data, which is of non-numerical type, according to the preset quantization algorithm to convert it into numerical data, and then obtain the corresponding second data.

[0112] S103: Determine the similarity between each target dataset and the preset dataset according to the preset dimensionality reduction algorithm, so that the risk control system can determine the source data based on the similarity.

[0113] The preset data set includes data on the dimensions and indicators of the preset targets.

[0114] After quantifying the first data from different target data sources to obtain second data and forming target data sets, a pre-defined dimensionality reduction algorithm is used to determine the similarity between each target data set and a pre-defined data set. This similarity characterizes the redundancy between the second data in the target data set and the data in the pre-defined target dimension indicator in the pre-defined data set; for example, a lower similarity indicates higher redundancy. This allows the risk control system to determine its required source data from the target data sets based on similarity. Therefore, the source data input to the risk control system is obtained from the target data set based on similarity determination, rather than directly using the first data from the target data sources as source data input. Thus, inputting source data determined based on similarity into the risk control system reduces the data processing workload, thereby ensuring the effectiveness and timeliness of risk control and reducing maintenance costs.

[0115] It should be noted that the preset data set includes data on preset target dimension indicators. These preset target dimension indicator data refer to the default reference data of the risk control system. This reference data is also used to characterize the preset feature information of the risk control object. Unlike the first data, this reference data has been recognized by the risk control system; for example, it can be used as reliable source data. Furthermore, the data type of the preset target dimension indicator data is numeric. The preset target dimension indicator data can also be second data obtained by numerically converting the first data provided by the target data source; this embodiment does not limit this.

[0116] The preset dimensionality reduction algorithm can be any algorithm used to determine the similarity between the variances of the corresponding datasets. For example, the preset random neighborhood embedding algorithm can be improved by adjusting the variance of the normal distribution centered on the sample data in the preset random neighborhood embedding algorithm to the variance corresponding to the standard deviation of the dataset composed of the second data after the first data representing a preset feature information is numerically converted, thereby obtaining the preset dimensionality reduction algorithm.

[0117] A pre-defined random neighborhood filling algorithm, such as the SNE (Stochastic Neighbor Embedding) algorithm, can be represented by the following expression (1):

[0118] (1)

[0119] in, and yes Surrounding sample points, It is a point The variance of a normally distributed system centered at 0. This indicates the expectation of an expression. , , Used to characterize each sample point in the sample data.

[0120] It should be noted that the core idea of ​​the pre-defined random neighborhood embedding algorithm is to convert the shortest distance between sample points into the similarity probability of these sample points. That is, if the shortest distance between sample points is... Under a Gaussian (normal) distribution centered at the nearest digit, if the neighborhood is selected in proportion to the probability density of the neighborhood, then... Will choose The conditional probability of being its neighbor.

[0121] For the second data obtained by quantifying the first data from multiple risk control objects provided by the target data source, the second data of each risk control object constitutes the target data set for a preset feature information. In other words, the target data set is formed by quantifying the first data from each target data source for the same preset feature information of each risk control object. The standard deviation between similar data sets is only affected by the magnitude of their expected values. If the magnitude of the expected value is removed, the standard deviations of similar data sets will also be similar. In other words, the similarity between data sets can be characterized by the distance between the ratios of their standard deviations. Determining the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm can be understood as determining the similarity between the data set composed of the second data obtained from the first data for the same preset feature information from each target data source and the data set composed of the data with the preset target dimension indicator.

[0122] Therefore, the variance of the normal distribution centered on the sample data in expression (1) can be adjusted to the variance of the corresponding data set, resulting in the expression (2) representing the preset dimensionality reduction algorithm:

[0123] (2)

[0124] in, X , Y These represent corresponding datasets, where data can be numerical data such as customer credit ratings and asset quality. For data sets X The standard deviation of all data in the dataset. For data sets X The expectation of all data in the middle, For data sets Y The set consisting of all data sets except those mentioned above. For set standard deviation For set Expectations ABS This indicates the operation of finding the absolute value.

[0125] The corresponding data set in relation (2) includes each target data set and the preset data set. Therefore, if the above... X The data set represented is considered as the preset data set. Y If the data set represented is considered as the target data set, then... The standard deviation of all second data in the preset dataset. The expected value of all second data in the preset dataset. To exclude the target data set Y The set consisting of all target data sets other than those mentioned above. For set standard deviation For set The expected value. Therefore, the similarity between each target data set and the preset data set can be determined by the relational expression of the preset dimensionality reduction algorithm, which is represented by expression (2).

[0126] After determining the similarity between each target dataset and the preset dataset, the similarity can characterize the redundancy between the first data provided by each target data source and the data of the preset target dimension indicator for the same preset feature information. For example, the lower the similarity, the higher the redundancy. This allows the risk control system to determine the required source data based on the similarity.

[0127] It should be noted that, according to the preset dimensionality reduction algorithm, the similarity between any two target data sets can also be determined. That is, when applying expression (2), the corresponding data of the preset data set Y is replaced with the corresponding data of another target data set, and the similarity between any two target data sets can be obtained. Based on this similarity, the redundancy between the first data provided by different target data sources for the same preset feature information can be fed back. For example, if the first data provided by the two target data sources are completely redundant, then one of the two target data sources can be selected to determine the source data.

[0128] Furthermore, in one possible design, after step S103, the data processing method for a risk control system provided in this application embodiment further includes performing one or more operations as shown below on the first data provided by the target data source to achieve the purpose of a data view.

[0129] For example, the first data provided by the target data source can be classified according to similarity and preset business requirements, and the classified first data can be generated into different data areas according to preset logical rules. The usage rights and scope of the first data can be determined through different data areas, so as to achieve the purpose of standardizing the use of the first data provided by the target data source.

[0130] For example, target data sources can be labeled based on similarity and preset labeling rules to standardize the recording of each target data source. The preset labeling rules can be set according to actual working conditions, and this embodiment does not limit this. For instance, target data sources (such as tax bureaus, industry and commerce bureaus, public security bureaus, courts, health organizations, etc.) can be labeled according to preset labeling rules based on the order of similarity or specific numerical values ​​to form standardized records. Furthermore, for some target data sources where the primary data may be composite data, such as monthly summaries of financial report A and account B, the corresponding target data sources can be labeled according to preset labeling rules based on the order of similarity or specific numerical values ​​combined with the specific content of this composite data.

[0131] For example, similarity can be used to determine the source data that meets the preset business scenario from the first data provided by the target data source, thus avoiding the business risks caused by differences between highly similar data. For instance, if the preset feature information of the risk control object is registration information, both the industrial and commercial data source and the tax data source can provide registration information. There may be a high degree of similarity between these two data sources. However, since the update frequency of the industrial and commercial data source is generally higher than that of the tax data source, the first data provided by the industrial and commercial data source can be used for this preset business scenario of customer identification.

[0132] Optionally, after step S103, the first data provided by the target data source can be stored or deleted based on the similarity to form a standardized data pool.

[0133] For example, in existing technologies, the first data provided by the target data source, when used as the source data in a risk control system to achieve risk control objectives, is usually discarded directly along with the output of the risk control results during the implementation of the risk control strategy, without ever considering whether the source data has any storage value. However, the data processing method for a risk control system provided in this application, after obtaining the similarity between the target data set and the preset data set, further performs storage or deletion operations on the first data provided by the target data source based on each similarity. For instance, it can determine the first data provided by the target data source corresponding to the target data set that is unrelated to the preset data set based on the similarity. Unrelatedness to the preset data set indicates that the first data provided by the target data source corresponding to that target data set is irreplaceable; therefore, the first data provided by the target data source can be stored. Conversely, by comparing the similarity with a preset redundancy threshold, the redundancy between the target data set and the preset data set is determined. The first data provided by the target data source corresponding to the target data set that is completely redundant with the preset data set can be deleted. Furthermore, the first data provided by the target data source corresponding to multiple completely redundant target data sets can be deleted based on actual operating conditions using the similarity. Through the aforementioned storage or deletion operations, the first data with storage value is stored, or the first data that is completely redundant with the preset target dimension indicators is deleted, so as to retain the first data with corresponding data value, form a standardized data pool, accumulate valuable first data, enrich the source data required by the risk control system, and thus realize the data sharing mechanism.

[0134] The data processing method for a risk control system provided in this application first acquires first data about multiple risk control objects from different target data sources. Then, it determines the data type of the first data and determines second data based on the determined data type to obtain a target data set. The second data is used to quantify the first data, and the data type includes both numeric and non-numeric types. Finally, it determines the similarity between each target data set and a preset data set using a preset dimensionality reduction algorithm. This allows the risk control system to determine the source data based on the obtained similarities. The method effectively identifies data redundancy from each target data source based on similarity, thereby determining the required source data. This not only reduces the processing workload of massive amounts of data when implementing risk control strategies but also ensures the effectiveness and timeliness of the risk control system. Furthermore, it effectively reduces the maintenance cost of the risk control system and facilitates the effective accumulation of data to form a data sharing mechanism.

[0135] In one possible design, step S103 of the above embodiment may be implemented as follows: Figure 4 As shown, Figure 4This is a flowchart illustrating another data processing method provided in an embodiment of this application. Figure 4 As shown in the figure, the method provided in this embodiment for determining the similarity between each target dataset and a preset dataset based on a preset dimensionality reduction algorithm includes:

[0136] S1031: Obtain the first expected value of each target data set and the second expected value of the preset data set.

[0137] For each target dataset, a first expected value for that target dataset is obtained using the second data within that dataset. Correspondingly, a second expected value for the preset dataset is obtained using data from preset target dimension indicators within the preset dataset.

[0138] For example, the first expected value of each target data set and the second expected value of the preset data set can be obtained through the expression (3) shown below:

[0139] (3)

[0140] in, Data set for calculating expected value X The Middle n The corresponding values ​​of each data point The first data set for calculating the expected value n The weight of each data point relative to all data points in the dataset. For data sets X The expectation of all data in the dataset.

[0141] The first expected value of each target data set and the second expected value of the preset data set can be obtained through expression (3). The data set in expression (3) X Each can be viewed as a target data set and a preset data set, respectively, to obtain the first expected value and the second expected value.

[0142] S1032: Obtain the first standard deviation of each target data set and the second standard deviation of the preset data set.

[0143] After obtaining the first expected value of each target data set and the second expected value of the preset data set, the first standard deviation of each target data set and the second standard deviation of the preset data set are further obtained.

[0144] For example, the first standard deviation of each target data set and the second standard deviation of the preset data set can be obtained by expression (4):

[0145] (4)

[0146] Among the symbols This indicates that the expectation of the expression is being calculated. For data sets X The standard deviation of all data in the dataset.

[0147] The first standard deviation of each target data set and the second standard deviation of the preset data set can be obtained through expression (4). The data set in expression (4) X Each can be viewed as a target data set and a preset data set, respectively, to obtain the first standard deviation and the second standard deviation.

[0148] S1033: Determine the distance data between each first standard deviation and the second standard deviation based on the preset dimensionality reduction algorithm, each first expected value, each first standard deviation, the second expected value, and the second standard deviation, so as to determine each distance data as the similarity between the corresponding target data set and the preset data set.

[0149] After obtaining the first expected value and first standard deviation of each target data set, and the second expected value and second standard deviation of the preset data set, the distance data between each first standard deviation and second standard deviation can be determined by further combining the preset dimensionality reduction algorithm, so as to determine the obtained distance data as the similarity between the corresponding target data set and the preset data set.

[0150] For example, by substituting the first expected value, the first standard deviation, the second expected value, and the second standard deviation obtained by the above expressions (3) and (4) into the above expression (2), the distance data between each first standard deviation and the second standard deviation can be obtained in sequence. This distance data represents the similarity between the corresponding target data set and the preset data set.

[0151] It is understandable that if the expected value and standard deviation of any two target data sets are obtained and substituted into the preset dimensionality reduction algorithm represented by expression (2), the distance data between the standard deviations of the two target data sets can be obtained. The similarity between the two target data sets can be represented by the distance data.

[0152] Table 1 below lists sample points of the first data for 14 risk control targets. These first data represent four preset characteristics of these 14 risk control targets: tax bureau asset quality, business asset quality, tax registration, and asset size.

[0153] Table 1

[0154]

[0155] As shown in Table 1, only the data type of the first data corresponding to the quality of industrial and commercial assets is non-numerical. The data types of the other first data are shown in Table 1. The table also shows the second data obtained after processing the first data corresponding to the quality of industrial and commercial assets according to the data type according to the preset quantification algorithm. Table 1 shows the corresponding second data (industrial and commercial quantification method 1 to industrial and commercial quantification method 3) obtained by processing the first data using three different preset coding rules in the preset quantification algorithm.

[0156] The sample points shown in Table 1 ( i 1~ i 14 Tax Bureau Asset Quality X i Industrial and commercial asset quality Y i-1 ~ Y i-3 Tax Class Z i and asset size W i Each set of data corresponds to a specific data set. Assuming we take the data set corresponding to the tax bureau's asset quality as the preset data set, and the data sets corresponding to the industrial and commercial asset quality, tax registration, and asset size as the target data sets, then based on the data provided in Table 1 and the preset dimensionality reduction algorithm, the corresponding similarity can be determined.

[0157] The first expectation of each hypothesized target data set and the second expectation of the preset data set are further obtained as shown in Table 2 below. In Table 2, ... This is used to represent the data sets in Table 2, which is equivalent to the data set in relation (2). X .

[0158] Table 2

[0159]

[0160] Based on the first and second expected values ​​obtained in Table 2, the first standard deviation of the assumed target data set and the second standard deviation of the preset data set are further obtained as shown in Table 3 (Table 3 is expressed in the form of variance).

[0161] Table 3

[0162]

[0163] Based on the values ​​obtained in Tables 2 and 3, the intermediate values ​​required for the preset dimensionality reduction algorithm are further determined as shown in Table 4.

[0164] Table 4

[0165]

[0166] Among them, the asset quality of the tax bureau in Table 4 X i The median value of the corresponding dataset Equivalent to relation (2) Industrial and commercial asset quality Y i-1 ~ Y i-3 Tax Class Z i and asset size W i The median value of the corresponding dataset The equivalent relation is in (2) As shown in Table 4, the tax bureau's asset quality X i Industrial and commercial asset quality Y i-1 ~ Y i-3 Tax Class Z i and asset size W i The corresponding datasets determine the median similarity value using a preset dimensionality reduction algorithm represented by relation (2). They are completely identical, therefore the similarity between the data sets obtained through relation (2) is: ) )=0, and and The differences are quite large.

[0167] Additionally, it should be noted that the sample points in Tables 2 to 4 above... i 1~ i 14 The corresponding values ​​are all used to determine the expected value, variance, and... of the corresponding data set. The corresponding intermediate values ​​in the process.

[0168] Optionally, after determining the distance data as the similarity between the corresponding target data set and the preset data set, the redundancy between the target data set and the preset data set can also be determined based on the similarity.

[0169] For example, if the similarity is zero, it is determined that the first data provided by the current target data source is completely redundant with the data of the preset target dimension indicator.

[0170] If the similarity is not zero and is less than the preset redundancy threshold, then the data redundancy between the first data provided by the current target data source and the preset target dimension indicator is determined.

[0171] If the similarity is greater than the preset redundancy threshold, it is determined that the first data provided by the current target data source is not related to the data of the preset target dimension indicator.

[0172] Specifically, if the similarity between the target dataset and the preset dataset is zero, it indicates that the first data provided by the current target data source corresponding to the target dataset is completely redundant with the data of the preset target dimension indicator corresponding to the preset dataset. If the similarity between the target dataset and the preset dataset is not zero and is less than the preset redundancy threshold, it indicates that the first data provided by the current target data source corresponding to the target dataset is redundant with the data of the preset target dimension indicator corresponding to the preset dataset, but not completely redundant. The risk control system can further decide whether to use the first data provided by the current target data source as the source data. If the similarity between the target dataset and the preset dataset is greater than the preset redundancy threshold, it indicates that the first data provided by the current target data source corresponding to the target dataset is not redundant with the data of the preset target dimension indicator corresponding to the preset dataset. These data are irreplaceable, and the risk control system can further decide whether to use the first data provided by the current target data source as the source data or to store the first data, etc.

[0173] It should be noted that as big data acquisition channels become more numerous and cost-effective, different target data sources often have certain data quality issues, such as missing data, loss of precision, and mixing of half-width and full-width characters. The level of data quality will further affect the value corresponding to the similarity score. After extensive data training, the similarity between datasets and the data quality ε of the target dataset roughly conform to a normal distribution. Based on this, it is generally considered that when the similarity is less than or equal to 0.47, the target dataset involved in that similarity score is approximately redundant with the preset dataset. Therefore, 0.47 can be set as the preset redundancy threshold, and the similarity is compared with the preset redundancy threshold to obtain the redundancy situation between the target dataset and the preset dataset. It should be noted that the preset redundancy threshold described above is only exemplary, and its value can also be set according to actual working conditions. This embodiment does not limit this.

[0174] In one possible design, step S203, which quantizes the first data according to a preset quantization algorithm, may be implemented as follows: Figure 5 As shown. Figure 5 This is a flowchart illustrating another data processing method provided in an embodiment of this application. As shown in Figure 5, the method provided in this embodiment processes first data according to a preset quantization algorithm to obtain second data, including:

[0175] S2031: Determine the representation of the first data that is not a numeric type.

[0176] The representation of the first data of non-numeric type is determined. If the first data of non-numeric type is represented by non-encoded text, then step S2032 is executed. If the first data of non-numeric type is represented by characters or encoded text, then step S2033 is executed.

[0177] S2032: If the first data, which is not a numerical type, is represented by non-encoded text, then the first data is processed according to the preset word segmentation algorithm to obtain the second data.

[0178] Among them, the preset quantization algorithm includes the preset word segmentation algorithm.

[0179] If the first data, which is not a numerical type, is represented as plain text without encoding, then the second data is obtained by processing the first data according to the preset word segmentation algorithm in the preset quantization algorithm.

[0180] For example, firstly, the first data and the data with preset target dimension indicators can be segmented separately according to a preset segmentation algorithm to obtain the number of words contained in the first data (target words) and the number of words contained in the data with preset target dimension indicators (target words). The preset segmentation algorithm can be an NLP (Natural Language Processing) algorithm, which performs lexical, syntactic, or semantic analysis on the information represented by pure text. Therefore, the preset segmentation algorithm can segment the first data and the data with preset target dimension indicators to obtain target words and target words accordingly. Then, the target words are compared with the target words to determine the proportion of target words in the target words; in other words, the proportion of target words contained in the target words is determined. This proportion is then used as the second data corresponding to the first data, completing the numericalization of the non-encoded pure text representation of the first data.

[0181] Table 5 below lists the cases where the first data of non-numerical type is represented as text in non-encoded form, and the corresponding second data is obtained by numericalizing each first data through a preset word segmentation algorithm.

[0182] Table 5

[0183]

[0184] As shown in Table 5, the preset word segmentation algorithm splits the first data "Xinglong Road C01, City B, Province A" for word segmentation processing, obtaining the target word segmentation "A, B, C01". At the same time, it splits the data of the preset subject dimension indication "AB" for word segmentation processing, obtaining the subject word segmentation "A, B". Then, it determines the proportion of the subject word segmentation "A, B" in the target word segmentation "A, B, C01", obtaining the proportion of "2 / 3", and thus determines "2 / 3" as the second data of the first data "Xinglong Road C01, City B, Province A".

[0185] S2033: If the first data of non-numerical type is represented by characters or coded words, the first data is processed according to the preset coding rule to obtain the second data.

[0186] Among them, the preset coding rule includes the unique correspondence between characters or coded words and the second data, and the preset quantization algorithm includes the preset coding rule.

[0187] If the first data of non-numerical type is represented by characters or coded words, such as represented by characters like "A, B, C,...", "I, II, III, IV,...", or coded words like "one, two, three, four,...", then the first data is processed through the preset coding rule, where the preset coding rule includes the unique correspondence between characters or coded words and the second data. In other words, the preset quantization algorithm can be set as the preset coding rule, and this preset coding rule ensures that there is a unique correspondence between a character or a coded word and an exact numerical value. Furthermore, the first data can be processed through the preset coding rule to obtain the value mapped by the preset coding rule to represent the first data, that is, to obtain the second data.

[0188] Table 6 listed below shows the situation where when the first data of non-numerical type is represented by characters or coded words, the first data is numerically converted through the preset coding rule to obtain the corresponding second data.

[0189] Table 6

[0190]

[0191] Table 6 lists the corresponding second data obtained by processing two types of first data through three preset coding rules. It should be noted that the specific content of the preset coding rule in this embodiment is not limited, as long as it has the single mapping principle, that is, the preset coding rule can ensure that there is a unique correspondence between a character or a coded word and a determined numerical value.

[0192] The data processing method for a risk control system provided in this application, after obtaining first data about multiple risk control objects from different target data sources, needs to determine the data type of the first data, which may be numeric or non-numeric, in order to quantify the first data according to the data type and obtain the corresponding second data. When the data type of the first data is numeric, it is directly determined as the second data. When the data type of the first data is non-numeric, it is processed according to a preset quantization algorithm to obtain the second data. Specifically, when the non-numeric first data is represented as plain text without encoding, it is segmented using a preset word segmentation algorithm, and the second data is obtained based on the proportion of the target word corresponding to the data with the preset target dimension indicator in the target word corresponding to the first data. When the non-numeric first data is represented as characters or encoded text, the preset quantization algorithm is set to a preset encoding rule, and the first data is quantified according to the correspondence represented by the preset encoding rule to obtain the corresponding second data. This completes the numerical processing of the first data to obtain the corresponding second data, which is then further processed in the data processing method provided in the embodiments of this application according to a preset dimensionality reduction algorithm.

[0193] The following are embodiments of the apparatus described in this application, which can be used to execute the corresponding method embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the corresponding method embodiments of this application.

[0194] Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 6 As shown, the data processing device 600 for a risk control system provided in this embodiment includes:

[0195] The acquisition module 601 is used to acquire first data about multiple risk control objects provided by different target data sources. The first data is used to characterize the preset feature information of the risk control objects.

[0196] The first processing module 602 is used to determine the data type of the first data, so as to determine the second data according to the data type and obtain the target data set. The second data is used to represent the first data after numericalization. The data type includes numeric type and non-numeric type.

[0197] The second processing module 603 is used to determine the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm, so that the risk control system can determine the source data based on the similarity. The preset data set includes data of the preset target dimensional indicators.

[0198] exist Figure 6 Based on the illustrated embodiment, Figure 7 This is a schematic diagram of another data processing apparatus provided in an embodiment of this application. Figure 7 As shown, the data processing device 600 for a risk control system provided in this embodiment further includes a third processing module 604. This third processing module 604 is used for:

[0199] The first data provided by the target data source is classified according to similarity and preset business requirements. The classified first data is then divided into different data regions according to preset logical rules to determine the usage permissions and scope of the first data; and / or

[0200] The target data sources are labeled according to similarity and preset labeling rules to standardize the recording of each target data source; and / or

[0201] Based on similarity, determine the source data that meets the preset business scenario from the first data provided by the target data source.

[0202] In one possible design, the data processing device 600 for the risk control system further includes: a fourth processing module, which is used for:

[0203] Based on similarity, the first data provided by the target data source is stored or deleted to form a standardized data pool.

[0204] In one possible design, the first processing module 602 is specifically used for:

[0205] If the data type of the first data is numeric, then the first data itself will be identified as the second data.

[0206] If the data type of the first data is non-numeric, then the first data is processed according to the preset quantization algorithm to obtain the second data.

[0207] Based on the above embodiments, Figure 8 This is a schematic diagram of the structure of a processing module provided in an embodiment of this application. Figure 8 As shown, the first processing module 602 provided in this embodiment further includes:

[0208] The first quantization processing unit 6021 is used to process the first data according to a preset word segmentation algorithm to obtain the second data if the first data, which is not a numerical type, is represented by non-encoded text. The preset quantization algorithm includes the preset word segmentation algorithm.

[0209] The second quantization processing unit 6022 is used to process the first data according to a preset encoding rule to obtain the second data if the first data, which is not a numerical type, is represented by characters or coded text. The preset encoding rule includes a unique correspondence between the characters or coded text and the second data, and the preset quantization algorithm includes the preset encoding rule.

[0210] In one possible design, the first quantization processing unit 6021 is specifically used for:

[0211] The first data and the data of the preset target dimension indicators are segmented according to the preset word segmentation algorithm to obtain the target word segmentation and the target word segmentation respectively.

[0212] Determine the proportion of the target word segment in the target word segment, and use this proportion as the second data corresponding to the first data.

[0213] In one possible design, the second processing module 603 is specifically used for:

[0214] Obtain the first expected value of each target data set and the second expected value of the preset data set;

[0215] Obtain the first standard deviation of each target dataset and the second standard deviation of a preset dataset;

[0216] Based on the preset dimensionality reduction algorithm, each first expected value, each first standard deviation, the second expected value, and the second standard deviation, the distance data between each first standard deviation and the second standard deviation is determined, so that each distance data is determined as the similarity between the corresponding target data set and the preset data set.

[0217] In one possible design, the data processing device 600 for the risk control system further includes a fifth processing module. This fifth processing module is used for:

[0218] If the similarity is zero, it is determined that the first data provided by the current target data source is completely redundant with the data of the preset target dimension indicator;

[0219] If the similarity is not zero and is less than the preset redundancy threshold, then the data redundancy between the first data provided by the current target data source and the preset target dimension indicator is determined.

[0220] If the similarity is greater than the preset redundancy threshold, it is determined that the first data provided by the current target data source is not related to the data of the preset target dimension indicator.

[0221] In one possible design, the data processing device 600 for the risk control system further includes a sixth processing module. This sixth processing module is used for:

[0222] An improvement is made to the preset random neighborhood filling algorithm to adjust the variance of the normal distribution centered on the sample data in the preset random neighborhood filling algorithm to the variance of the corresponding data set, thus obtaining a preset dimensionality reduction algorithm. The corresponding data set includes each target data set and the preset data set.

[0223] In one possible design, the preset feature information includes at least one of the following: business registration information, tax information, credit information, judicial information, asset information, operating income information, and suspicious information related to illegal fund transfers of the risk control object.

[0224] It is worth noting that the data processing device for entering the risk control system provided in the above embodiments can be used to execute each step in the data processing method for the risk control system provided in any of the above embodiments. The specific implementation methods and technical effects are similar, and will not be repeated here.

[0225] The device embodiments provided in this application are merely illustrative, and the module division is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system. The coupling between the modules can be achieved through some interfaces, which are usually electrical communication interfaces, but mechanical interfaces or other forms of interfaces are also possible. Therefore, the modules described as separate components may or may not be physically separated; they may be located in one place or distributed in different locations on the same or different devices.

[0226] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 9 As shown, the electronic device 700 may include at least one processor 701 and a memory 702. Figure 9 Let's take a processor as an example.

[0227] The memory 702 is used to store the program of the processor 701. Specifically, the program may include program code, which includes computer operation instructions.

[0228] The memory 702 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0229] The processor 701 is configured to execute the computer program stored in the memory 702 to implement the steps in the data processing method for the risk control system in the above method embodiments.

[0230] The processor 701 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0231] Optionally, the memory 702 can be either standalone or integrated with the processor 701. When the memory 702 is a device independent of the processor 701, the electronic device 700 may further include:

[0232] Bus 703 is used to connect processor 701 and memory 702. The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc., but this does not mean there is only one bus or one type of bus.

[0233] Optionally, in a specific implementation, if the memory 702 and the processor 701 are integrated on a single chip, the memory 702 and the processor 701 can communicate through an internal interface.

[0234] This application also provides a computer-readable storage medium, which may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Specifically, the computer-readable storage medium stores a computer program. When at least one processor of the aforementioned electronic device executes the computer program, the electronic device performs each step of the data processing method for the risk control system provided in the various embodiments described above.

[0235] This application also provides a computer program product comprising a computer program stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the device to perform the various steps of the data processing method for a risk control system provided in the various embodiments described above.

[0236] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.

[0237] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A data processing method for a risk control system, characterized in that, include: Acquire first data about multiple risk control objects from different target data sources, wherein the first data is used to characterize the preset feature information of the risk control objects; The data type of the first data is determined, and the second data is determined based on the data type to obtain the target data set. The second data is used to represent the first data after numericalization. The data type includes numerical type and non-numerical type. The step of determining the data type of the first data to determine the second data based on the data type includes: if the data type of the first data is the numeric type, then the first data itself is determined as the second data; if the data type of the first data is the non-numeric type, and the non-numeric first data is represented by non-encoded text, then the first data is processed according to a preset word segmentation algorithm to obtain the second data, the preset quantization algorithm including the preset word segmentation algorithm; if the data type of the first data is the non-numeric type, and the non-numeric first data is represented by characters or encoded text, then the first data is processed according to a preset encoding rule to obtain the second data, the preset encoding rule including a unique correspondence between the characters or encoded text and the second data, the preset quantization algorithm including the preset encoding rule; The similarity between each target data set and the preset data set is determined according to the preset dimensionality reduction algorithm, so that the risk control system can determine the source data based on the similarity. The preset data set includes data of the preset target dimensionality indicators. Before determining the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm, the method further includes: improving the preset random neighborhood embedding algorithm to adjust the variance of the normal distribution centered on the sample data in the preset random neighborhood filling algorithm to the variance of the corresponding data set, thereby obtaining the preset dimensionality reduction algorithm, wherein the corresponding data set includes each target data set and the preset data set. The step of determining the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm includes: obtaining a first expected value for each target data set and a second expected value for the preset data set; obtaining a first standard deviation for each target data set and a second standard deviation for the preset data set; and determining the distance data between each first standard deviation and the second standard deviation according to the preset dimensionality reduction algorithm, each first expected value, each first standard deviation, the second expected value, and the second standard deviation, so as to determine each distance data as the similarity between the corresponding target data set and the preset data set.

2. The data processing method for a risk control system according to claim 1, characterized in that, After determining the similarity between each target dataset and the preset dataset using a preset dimensionality reduction algorithm, the process further includes: The first data provided by the target data source is classified according to the similarity and preset business requirements, and the classified first data is generated into different data regions according to preset logical rules, so as to determine the usage rights and scope of the first data through the different data regions; and / or The target data sources are labeled according to the similarity and preset labeling rules to standardize the recording of each target data source; and / or Based on the similarity, the post source data that satisfies the preset business scenario is determined from the first data provided by the target data source.

3. The data processing method for a risk control system according to claim 1, characterized in that, After determining the similarity between each target dataset and the preset dataset using a preset dimensionality reduction algorithm, the process further includes: Based on the similarity, the first data provided by the target data source is stored or deleted to form a standardized data pool.

4. The data processing method for a risk control system according to claim 1, characterized in that, The step of processing the first data according to a preset word segmentation algorithm to obtain the second data includes: The first data and the data of the preset target dimension indicator are segmented according to the preset word segmentation algorithm to obtain target word segmentation and target word segmentation respectively. The proportion of the target word segment in the target word segment is determined, and the proportion is used as the second data corresponding to the first data.

5. The data processing method for a risk control system according to claim 1, characterized in that, After determining each distance data as the similarity between the corresponding target data set and the preset data set, the method further includes: If the similarity is zero, then it is determined that the first data provided by the current target data source is completely redundant with the data of the preset target dimension indicator; If the similarity is not zero and is less than the preset redundancy threshold, then the data redundancy between the first data provided by the current target data source and the preset target dimension indicator is determined. If the similarity is greater than a preset redundancy threshold, it is determined that the first data provided by the current target data source is not related to the data of the preset target dimension indicator.

6. The data processing method for a risk control system according to any one of claims 1-3, characterized in that, The preset feature information includes at least one of the following: business registration information, tax information, credit information, judicial information, asset information, operating income information, and suspicious information related to illegal fund transfers of the risk control object.

7. A data processing device for a risk control system, characterized in that, include: The acquisition module is used to acquire first data about multiple risk control objects provided by different target data sources. The first data is used to characterize the preset feature information of the risk control objects. The first processing module is used to determine the data type of the first data, and to determine the second data based on the data type to obtain a target data set. The second data is used to represent the first data after numericalization. The data type includes numerical type and non-numerical type. The first processing module is specifically configured to: if the data type of the first data is the numeric type, then determine the first data itself as the second data; if the data type of the first data is the non-numeric type, and the non-numeric first data is represented by non-encoded text, then process the first data according to a preset word segmentation algorithm to obtain the second data, wherein the preset quantization algorithm includes the preset word segmentation algorithm; If the data type of the first data is the non-numeric type, and the first data of the non-numeric type is represented by characters or coded text, then the first data is processed according to a preset encoding rule to obtain the second data. The preset encoding rule includes a unique correspondence between the characters or coded text and the second data. The preset quantization algorithm includes the preset encoding rule. The second processing module is used to determine the similarity between each target data set and the preset data set according to the preset dimensionality reduction algorithm, so that the risk control system can determine the source data based on the similarity. The preset data set includes data of the preset target dimensionality indicators. The second processing module is further used to improve the preset random neighborhood embedding algorithm, so as to adjust the variance of the normal distribution centered on the sample data in the preset random neighborhood filling algorithm to the variance of the corresponding data set, thereby obtaining the preset dimensionality reduction algorithm, wherein the corresponding data set includes each target data set and the preset data set; The second processing module is specifically used to obtain the first expected value of each target data set and the second expected value of the preset data set; obtain the first standard deviation of each target data set and the second standard deviation of the preset data set; and determine the distance data between each first standard deviation and the second standard deviation according to the preset dimensionality reduction algorithm, each first expected value, each first standard deviation, the second expected value, and the second standard deviation, so as to determine each distance data as the similarity between the corresponding target data set and the preset data set.

8. An electronic device, characterized in that, include: processor; as well as Memory for storing the computer program of the processor; The processor is configured to execute the data processing method for a risk control system according to any one of claims 1 to 6 by executing the computer program.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data processing method for a risk control system as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the data processing method for a risk control system as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Financial data risk control method and device based on artificial intelligence

    CN108985583A

  • Neural network model training method and device for weak annotation data

    CN110070183A