Data source evaluation method, device, electronic device and storage medium

By calculating the historical and online service data of the target data source, the problem of excessive resource consumption in the existing technology is solved, and efficient and accurate evaluation of TB/PB-level financial data is achieved.

CN115034659BActive Publication Date: 2025-08-22DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210749526.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-08-22
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

The existing data value assessment model consumes too much resources when processing TB or even PB level financial data, and lacks forward-lookingness, so it cannot effectively evaluate credit risks.

Method used

By calculating the index of the historical service data of the target data source, the statistical results of the first indicator are obtained, and based on this, whether to conduct online evaluation is determined; if the conditions are met, the index calculation of the online service data is obtained, and the statistical results of the second indicator are finally determined, and the value evaluation results of the data source are finally determined.

Benefits of technology

Reduces the consumption of resource evaluation for data sources that do not meet the online evaluation conditions, and improves the accuracy and efficiency of evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115034659B_ABST
    Figure CN115034659B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data source evaluation method, device, electronic device, and storage medium, which belong to the field of big data offline computing. The method includes: obtaining historical service data of a target data source to be evaluated; performing index calculation processing on the historical service data to obtain a first index statistical result of the target data source; when determining to perform an online evaluation based on the first index statistical result, obtaining online service data of the target data source; performing index calculation processing on the online service data to obtain a second index statistical result of the target data source; and determining a value evaluation result of the target data source based on the first index statistical result and / or the second index statistical result. The use of the present disclosure can reduce resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data offline computing, and in particular to a data source evaluation method, device, electronic device and storage medium. Background Art

[0002] The essence of data value assessment lies in the analysis, calculation, and storage of data, involving technologies related to big data computing. With the rapid development of the information industry, the number of users, network services, and computing devices are growing exponentially, directly impacting the explosive growth of "data traffic." Faced with the generation of terabytes, petabytes, and even larger amounts of data, the term "big data" is rapidly gaining popularity in the data computing, data mining, and analysis industries. To handle massive amounts of service data, various data processing architecture models have been introduced in various fields. With the continuous iteration and development of products, distributed computing and storage data processing frameworks are gaining favor and adoption among a growing number of internet companies.

[0003] Currently, big data technology is experiencing significant development in data value assessment models for the financial sector. The application of big data technology in the financial industry has improved resource allocation efficiency and strengthened risk management capabilities. Traditionally, default risk assessment relies on analyzing historical transaction and credit data, a method that lacks foresight. However, using big data technology to assess credit risk is more factual. Furthermore, the financial sector currently has numerous external service data sources with complex and diverse formats. Furthermore, due to the unique nature of the financial sector, data storage may be required for three years or even longer, resulting in data volumes that can easily reach terabytes or even petabytes.

[0004] Existing data value assessment models meet business needs when applied to small and medium-sized data sets. However, when applied to the financial industry, which needs to process TB or even PB-level data volumes, they expose many drawbacks. For example, ultra-large-scale data sets require ultra-strong computing power, and the foundation of computing power is the support of physical resources. Therefore, when facing ultra-large-scale data sets, the current data value assessment model will consume too many resources. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a data source evaluation method, device, electronic device, and storage medium to solve the problem of excessive resource consumption.

[0006] According to one aspect of the present disclosure, a data source evaluation method is provided, the method comprising:

[0007] Obtain historical service data of the target data source to be evaluated;

[0008] Performing indicator calculation processing on the historical service data to obtain a first indicator statistical result of the target data source;

[0009] When it is determined to perform online evaluation based on the statistical result of the first indicator, obtaining online service data of the target data source;

[0010] Performing indicator calculation processing on the online service data to obtain a second indicator statistical result of the target data source;

[0011] Based on the first indicator statistical result and / or the second indicator statistical result, a value assessment result of the target data source is determined.

[0012] According to another aspect of the present disclosure, a data source evaluation device is provided, comprising:

[0013] A first acquisition module is used to acquire historical service data of a target data source to be evaluated;

[0014] An indicator calculation module, configured to perform indicator calculation processing on the historical service data to obtain a first indicator statistical result of the target data source;

[0015] A second acquisition module is configured to acquire online service data of the target data source when it is determined to perform online evaluation based on the statistical result of the first indicator;

[0016] The indicator calculation module is further configured to perform indicator calculation processing on the online service data to obtain a second indicator statistical result of the target data source;

[0017] An evaluation module is used to determine a value evaluation result of the target data source based on the first indicator statistical result and / or the second indicator statistical result.

[0018] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0019] processor; and

[0020] Memory for storing programs,

[0021] The program includes instructions, which, when executed by the processor, cause the processor to execute the data source evaluation method.

[0022] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the data source evaluation method.

[0023] In the present disclosure, by performing indicator calculation processing on the historical service data of the target data source, a first indicator statistical result is obtained, and based on the first indicator statistical result, it is determined whether to perform an online evaluation on the target data source. If it is determined that the online evaluation should continue, the indicator calculation processing is performed on the online service data of the target data source to obtain a second indicator statistical result. Ultimately, the value evaluation result of the target data source can be determined based on the first indicator statistical result and / or the second indicator statistical result. Since the conditions for online evaluation are set, it is possible to avoid evaluating data sources that do not meet the online evaluation conditions, reduce resource consumption, and also improve the accuracy of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0025] Figure 1 A flow chart of a data source evaluation method according to an exemplary embodiment of the present disclosure is shown;

[0026] Figure 2 A schematic diagram of customer dimensions provided according to an exemplary embodiment of the present disclosure is shown;

[0027] Figure 3 A flow chart of a method for constructing data to be evaluated for a first customer group according to an exemplary embodiment of the present disclosure is shown;

[0028] Figure 4 A flow chart of a method for constructing data to be evaluated for a second customer group according to an exemplary embodiment of the present disclosure is shown;

[0029] Figure 5 A schematic block diagram of a data source evaluation device provided according to an exemplary embodiment of the present disclosure is shown;

[0030] Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0031] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0032] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0033] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0034] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0035] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0036] This disclosure provides a data source evaluation method, which is applicable to a data source value assessment system and can be performed by a terminal, a server, and / or other devices with processing capabilities. The methods provided in the embodiments of this disclosure can be performed by any of the aforementioned devices, or by multiple devices working together, without limitation in this disclosure.

[0037] The following will refer to Figure 1 The data source evaluation method flow chart is shown to introduce the method. The method includes the following steps 101-105.

[0038] Step 101: Obtain historical service data of a target data source to be evaluated.

[0039] In one possible implementation, a financial services institution can provide services to clients through multiple external services, and these multiple external services can be connected as data sources to the financial services institution's data source value assessment system. When performing a value assessment on any of these data sources, that data source can be used as the target data source, and service data triggered by the target data source during a historical time period, or offline sample service data, can be uploaded to the data warehouse. For ease of description, this service data will be referred to as historical service data. For example, offline sample service data can refer to service data uploaded by the institution providing the external service after completing the service offline.

[0040] As an example, the above service data is the interface log data that has been cleaned and dropped into the Hive data warehouse (a data warehouse tool based on Hadoop). The data source value assessment system can use offline batch processing Spark (a fast and general computing engine designed for large-scale data processing) to read the service data in Hive.

[0041] Step 102: perform index calculation processing on the historical service data to obtain a first index statistical result of the target data source.

[0042] In one possible implementation, evaluation indicators can be pre-set. Historical service data can then be parsed to obtain the data required to calculate the evaluation indicators. The obtained data is then statistically analyzed to obtain the corresponding statistical results for each evaluation indicator. The statistical results for each evaluation indicator are then integrated to form the first indicator statistical results for the target data source. The first indicator statistical results for the target data source are persisted and written to the data warehouse.

[0043] Step 103: When it is determined to perform online evaluation based on the statistical result of the first indicator, online service data of the target data source is obtained.

[0044] In one possible implementation, the statistical results of the first indicator can be analyzed to determine whether online evaluation is worthwhile. For example, if the analysis of the statistical results of the first indicator indicates that the data from the target data source has a good evaluation effect, then online evaluation can be performed. If the analysis of the statistical results of the first indicator indicates that the data from the target data source has a poor evaluation effect, then online evaluation is not performed.

[0045] Online service data can include online tracking data or online usage data. In some application scenarios, when an external service is not yet online, you can test the service, trigger tracking traffic, and collect corresponding service data. This online tracking data can then be used for metric calculations. When the external service is online, the tracking traffic becomes online traffic, and the collected data is online usage data, which is then used for metric calculations.

[0046] Step 104: perform index calculation processing on the online service data to obtain a second index statistical result of the target data source.

[0047] In a possible implementation, the online service data may be an access service data stream, which includes a plurality of messages, and each message may be parsed to obtain various data required for calculating the evaluation index.

[0048] Optionally, the message formats of the data sources of multiple external services are complex and diverse. The method of parsing based on the message format is only applicable to specific messages, has poor portability, and is not compatible with adapting and analyzing messages from multiple data sources. Therefore, this application provides a message parsing method applicable to the data sources of multiple external services. Accordingly, the processing of the above step 104 can be as follows:

[0049] Use pre-set regular expression parsing rules to parse online service data and obtain the target message value of the target variable.

[0050] Based on the target message value of the target variable of the online service data, an indicator calculation process is performed on the online service data to obtain a second indicator statistical result of the target data source.

[0051] The target message value is a message value common to multiple data sources, and the target data source is any one of the multiple data sources.

[0052] In one possible implementation, message values ​​from multiple data sources can be analyzed to identify common message value types. For example, these common values ​​can be natural numbers from 0 to 9, with the exception values ​​-1 and -9999, and discrete message values ​​can be enumerated values ​​from [high, medium, low], [A-Z], and [1st, 2nd, ..., 9th].

[0053] Based on the message value type determined above, a corresponding regular parsing rule can be constructed. Then, the message is parsed according to the regular parsing rule to obtain the target message value of the target variable.

[0054] As an example, the numerical message parsing expression is f(x) = "[0-9]+|-1|-9999".r, and the discrete message parsing expression is g(x) = "[AZ]|high|medium|low|first gear|second gear|third gear|fourth gear|fifth gear|sixth gear|seventh gear|eighth gear|ninth gear".r, where the parsing expressions f(x) and g(x) are both regular parsing rules.

[0055] Let the message to be parsed be Data and the target variable to be parsed be indexA.

[0056] If the message value corresponding to indexA is a numeric value, the data parsing model formula can be as follows (1):

[0057]

[0058] If the message value corresponding to indexA is discrete, the data parsing model formula can be as follows (2):

[0059]

[0060] The above formulas (1) and (2) are explained as follows: First, locate whether indexA is contained in the message, that is, whether Data(indexA) is greater than 0. If it does not exist, that is, the message does not contain the variable to be parsed, then the message value of indexA is determined to be "NULL"; if it exists, then regular matching is performed according to the regular parsing rule f(x) or g(x) to obtain the message value of indexA. The above is only for the case where the message to be parsed has only one numerical or discrete variable. If there are multiple variables, the regular parsing rule can be used to match the value in the message Data that is closest to the indexA variable as the message value of indexA. For example, if there are two variables indexA and indexB in the message Data, the message value of indexA is null, and indexB has a message value. If the value that is closest to the indexA variable is not used as the message value of indexA, the message value of indexA may be incorrectly assigned to the message value of indexB; if the value that is closest to the indexA variable is used as the message value of indexA, the message value of indexA can be correctly assigned to null.

[0061] The above data parsing model can be registered as a Spark UDF (user-defined function) and used to parse messages when accessing data from the Hive data warehouse. After parsing the message values ​​for each target variable, the parsing results can be partitioned and stored in the Hive data warehouse based on the data source interface name and the time the interface was accessed.

[0062] After the analysis is completed, the various data required to calculate the evaluation indicators can be obtained, and then the obtained data can be statistically analyzed to obtain the indicator statistical results corresponding to each evaluation indicator.

[0063] Optionally, to improve the accuracy of the assessment, this embodiment further granularizes the customer dimension and performs indicator calculation processing on each dimension separately. Based on this, the indicator calculation processing may include: categorizing the service data to be processed into various dimensions according to pre-set dimensional division rules; and performing indicator calculation processing on the service data in each dimension separately.

[0064] The service data to be processed may refer to the above-mentioned historical service data or online service data.

[0065] Reference Figure 2 The customer dimension diagram shown in the figure shows that for the first dimension, customers can be divided into two dimensions: new customer groups and old customer groups based on the time of application for credit; for the second dimension, for each customer group, customer labels can be subdivided into corporate customer labels and non-corporate customer labels; for the third dimension, customer quality can be divided into multiple levels, among which customer quality can be determined based on customer characteristic information, which may include education level, overdue information, age, company information, etc.

[0066] Optionally, this embodiment provides different data acquisition methods for the new customer group and the old customer group. Accordingly, the processing of step 104 may be as follows:

[0067] Based on the access time information of the access service data flow, respectively determining a first customer group data flow and a second customer group data flow in the online service data;

[0068] Obtaining first user attribute data and first loan receipt data corresponding to the first customer group, and combining them with the first customer group data stream to determine first service data corresponding to the first customer group;

[0069] Obtaining second user attribute data, second loan note data, and post-loan transaction data corresponding to the second customer group, and combining them with the second customer group data stream to determine second service data corresponding to the second customer group;

[0070] According to the pre-set evaluation index, the first service data and the second service data are subjected to index calculation processing to obtain the statistical result of the evaluation index as the second index statistical result of the target data source.

[0071] In one possible implementation, based on access time information, access service data flows within a more distant time range than the current time can be considered as first customer group data flows, while access service data flows within a more recent time range can be considered as second customer group data flows. In other words, the first customer group data flows can correspond to existing customer data flows, while the second customer group data flows can correspond to new customer data flows. For example, the first time range can be from January 1, 2019, to the end of the previous month, and the second time range can be today. This embodiment does not limit the specific time ranges.

[0072] The following introduces the processing of the first customer group and the second customer group respectively.

[0073] The first customer group may correspond to the above-mentioned old customer group, including multiple first customers, that is, customers who have applied for credit for a longer time.

[0074] like Figure 3 The flowchart of the method for constructing the data to be evaluated of the first customer group is shown, and the method includes the following steps 301-303.

[0075] Step 301: Obtain pre-stored user attribute information and loan note information, where the user attribute information includes a customer identifier of each first customer;

[0076] Step 302: Based on the customer identifier of each first customer, obtain first user attribute data of each first customer from the user attribute information, obtain first loan note data of each first customer from the loan note information, and obtain first customer data of each first customer from the first customer group data stream.

[0077] Step 303: Construct first service data corresponding to the first customer group based on the first user attribute data, the first loan note data, and the first customer data of each first customer.

[0078] In one possible implementation, the pre-stored user attribute information may include attribute information for each customer. This attribute information includes at least a customer identifier (e.g., customer ID) and the time of application for credit. This allows the customer to be determined as a new or existing customer based on the time of application. When constructing the data to be evaluated for existing customers, first user attribute data corresponding to the existing customer can be obtained from the pre-stored user attribute information. This first user attribute data can also be used to determine the customer dimension to which the customer belongs.

[0079] Furthermore, pre-stored IOU information, such as an original IOU table, can be obtained. The IOU information also includes the customer ID of the corresponding customer. Therefore, when constructing the data to be evaluated for an existing customer, the customer ID of the existing customer can be used to search for the first IOU data corresponding to the existing customer and perform an association.

[0080] Similarly, the message of the data source may also carry a customer identifier for triggering access, so that the first customer data corresponding to the old customer is obtained from the first customer group data stream and associated based on the customer identifier of the old customer.

[0081] In summary, the first user attribute data, first loan note data and first customer data of each old customer (i.e., the first customer) can be obtained. After sorting, for example, data boxing and / or classification according to the above dimensions, the first service data corresponding to the first customer group is obtained. The first service data is the data to be evaluated for the first customer group.

[0082] The second customer group may correspond to the above-mentioned new customer group, including a plurality of second customers, ie customers who have applied for credit for a shorter time.

[0083] like Figure 4 The flowchart of the method for constructing the data to be evaluated of the second customer group is shown, and the method includes the following steps 401-404.

[0084] Step 401: Obtain pre-stored loan information and post-loan transaction information;

[0085] Step 402, determining the customer ID of each second customer based on the pre-stored loan note information and post-loan transaction information;

[0086] Step 403: Based on the customer identifier of each second customer, obtain second user attribute data of each second customer from the pre-stored user attribute information, obtain second IOU data of each second customer from the IOU information, obtain post-loan transaction data of each second customer from the post-loan transaction information, and obtain second customer data of each second customer from the second customer group data stream.

[0087] Step 404 : constructing second service data corresponding to the second customer group based on the second user attribute data, second loan note data, post-loan transaction data and second customer data of each second customer.

[0088] In a possible implementation, pre-stored loan note information and post-loan transaction information may be obtained, and the loan note information and post-loan transaction information also include a customer identifier of the corresponding customer.

[0089] Optionally, the processing of the above step 402 can be as follows: according to a preset time slicing rule, obtain the intermediate promissory note information and intermediate post-loan transaction information from the pre-stored promissory note information and post-loan transaction information; treat each customer in the intermediate promissory note information and intermediate post-loan transaction information as a second customer, and determine the customer identification of each second customer.

[0090] For example, you can extract the most recent day's IOU information and post-loan transaction information by time slice, using this as intermediate IOU and post-loan transaction information. Compared to the full set of IOU and post-loan transaction information, the data size of these intermediate IOU and post-loan transaction information is significantly reduced, for example, from terabytes to gigabytes. This overall data reduction is an order of magnitude, improving task execution time, reducing the resources required, and lowering the risk of system memory overflow.

[0091] Moreover, since the extracted time slices are relatively recent, the customers involved can be considered as new customers. Therefore, each customer in the intermediate loan note information and intermediate post-loan flow information can be regarded as a second customer, and the customer identification of each second customer can be determined.

[0092] Furthermore, based on the determined second customer identifier, the second IOU data for each second customer can be obtained from the intermediate IOU information, the post-loan transaction data for each second customer can be obtained from the intermediate post-loan transaction information, the second user attribute data for each second customer can be obtained from the pre-stored user attribute information, and the second customer data for each second customer can be obtained from the second customer group data stream. The specific processing is similar to that for the first customer group described above and will not be repeated here.

[0093] After organizing the second user attribute data, second loan note data, post-loan transaction data and second customer data of each second customer, for example, after data binning and / or classification according to the above dimensions, the second service data corresponding to the second customer group is obtained, and the second service data is the data to be evaluated for the second customer group.

[0094] Optionally, before calculating the metrics for the data to be evaluated, you can perform data binning. For discrete data, each discrete message value is a bin. For example, the data corresponding to "level 1" is a bin. For data types, you can use any of the following binning methods: equidistant binning, equal frequency binning, and custom binning.

[0095] As an example, let the number of bins be box_num and the variable to be evaluated be indexA.

[0096] Equally spaced bins mean that the distance between the upper and lower limits of each bin is equal.

[0097] The maximum value of the calculated variable indexA is max, and the minimum value is min. The step frequency Step calculation formula of the equally spaced bins is as follows (3):

[0098]

[0099] After the step frequency calculation of the equally spaced bins is completed, the interval boundary value box_region of the data division can be calculated by the following formula (4):

[0100] box_region=min+(1→box_num)×Step (4)

[0101] Equal-frequency binning means that the amount of data in each bin is equal.

[0102] First, calculate the step frequency Step of the equal-frequency division quantile. The calculation formula is as follows (5):

[0103]

[0104] Then, the percentil_approx quantile function in the Hive data warehouse is used to divide the data into the interval boundary value set box_region_list. The calculation formula is as follows (6):

[0105] box_region_list=percentil_approx(indexA,(1→box_num)×Step) (6)

[0106] Custom binning divides data into bins based on custom interval boundary values.

[0107] After completing the binning of the data, you can aggregate the data based on the binning results and calculate the indicator distribution of each evaluation indicator in different bins.

[0108] Optionally, the pre-set evaluation indicators include any one or more of the following: number of applicants, total number of people checked, check rate, number of check calculations, number of people approved for credit, average amount of credit, number of credit users, number of overdue people within the first time range, overdue amount within the first time range, number of overdue people within the second time range, overdue amount within the second time range, amount of new IOUs for old customers, and amount of data in boxes.

[0109] Here, "find" refers to finding the target variable. For example, when the target variable indexA is found in the message corresponding to customer A, it is recorded as "find".

[0110] The first time range is, for example, 0-30 days, the overdue amount within the first time range is, for example, the cumulative amount of IOUs overdue for 0-30 days, and the number of overdue people within the first time range is, for example, the cumulative number of IOUs overdue for 0-30 days.

[0111] The second time range is, for example, 30-60 days, the overdue amount in the second time range is, for example, the cumulative amount of IOUs overdue for 30-60 days, and the number of overdue people in the second time range is, for example, the cumulative number of IOUs overdue for 30-60 days.

[0112] The remaining evaluation indicators can refer to existing definitions and will not be listed one by one in this embodiment.

[0113] Step 105: Determine the value assessment result of the target data source based on the first indicator statistical result and / or the second indicator statistical result.

[0114] In one possible implementation, after determining the statistical results of the first indicator and / or the second indicator, these results can be persisted and written to a data warehouse. Furthermore, analysis and calculations can be performed based on the statistical results of the first indicator and / or the second indicator. For example, a weighted summation of the statistical results of each evaluation indicator can be performed to obtain a value assessment result for the target data source, and the value assessment result can be written to the data warehouse. This embodiment does not limit the specific analysis and calculation.

[0115] This embodiment does not limit the specific evaluation purpose. For example, it can be used to evaluate whether the target data source has a default risk. For another example, it can be used to evaluate whether the target data source has value to the financial service institution.

[0116] Optionally, the above data source evaluation method can adopt a distributed computing framework. During the task calculation process, intermediate data can be cached. The intermediate data can refer to repeatedly called data, so as to reduce the operation of re-reading data and improve the operation efficiency of the task to a certain extent.

[0117] Optionally, you can adjust the task concurrency based on the computing resources to improve the task's running efficiency.

[0118] Optionally, a report display may be performed based on the data stored in the data warehouse, such as the statistical results of the first indicator, the statistical results of the second indicator, and / or the value assessment results. For example, data from the Hive data warehouse may be synchronized to Green Plum (a relational distributed database) for report display.

[0119] In this embodiment, a first indicator statistical result is obtained by performing indicator calculation processing on the historical service data of the target data source. Based on the first indicator statistical result, a determination is made as to whether to perform an online evaluation on the target data source. If the determination is to continue the online evaluation, an indicator calculation processing is performed on the online service data of the target data source to obtain a second indicator statistical result. Ultimately, the value assessment result of the target data source can be determined based on the first indicator statistical result and / or the second indicator statistical result. By setting the conditions for online evaluation, it is possible to avoid evaluating data sources that do not meet the online evaluation conditions, thereby reducing resource consumption and improving the accuracy of the evaluation.

[0120] The embodiment of the present disclosure provides a data source evaluation device, which is used to implement the above data source evaluation method. Figure 5 As shown, the data source evaluation device 500 includes: a first acquisition module 501 , an indicator calculation module 502 , a second acquisition module 503 , and an evaluation module 504 .

[0121] The first acquisition module 501 is used to acquire historical service data of the target data source to be evaluated;

[0122] An indicator calculation module 502 is configured to perform indicator calculation processing on the historical service data to obtain a first indicator statistical result of the target data source;

[0123] A second acquisition module 503 is configured to acquire online service data of the target data source when it is determined to perform online evaluation based on the first indicator statistical result;

[0124] The indicator calculation module 502 is further configured to perform indicator calculation processing on the online service data to obtain a second indicator statistical result of the target data source;

[0125] The evaluation module 504 is configured to determine a value evaluation result of the target data source based on the first indicator statistical result and / or the second indicator statistical result.

[0126] Optionally, the online service data is an access service data stream;

[0127] The indicator calculation module 502 is used to:

[0128] determining a first customer group data flow and a second customer group data flow in the online service data based on access time information of the access service data flow;

[0129] Obtaining first user attribute data and first loan receipt data corresponding to a first customer group, and combining them with the first customer group data stream to determine first service data corresponding to the first customer group;

[0130] Obtaining second user attribute data, second loan note data, and post-loan transaction data corresponding to the second customer group, and combining these with the second customer group data stream to determine second service data corresponding to the second customer group;

[0131] According to the pre-set evaluation index, the first service data and the second service data are subjected to index calculation processing to obtain the statistical result of the evaluation index as the second index statistical result of the target data source.

[0132] Optionally, the first customer group includes multiple first customers;

[0133] The indicator calculation module 502 is further configured to:

[0134] Acquiring pre-stored user attribute information and loan note information, wherein the user attribute information includes a customer identifier of each first customer;

[0135] Based on the customer identifier of each first customer, obtaining first user attribute data of each first customer from the user attribute information, obtaining first loan note data of each first customer from the loan note information, and obtaining first customer data of each first customer from the first customer group data stream;

[0136] Based on the first user attribute data, the first loan note data and the first customer data of each first customer, first service data corresponding to the first customer group is constructed.

[0137] Optionally, the second customer group includes multiple second customers;

[0138] The indicator calculation module 502 is further configured to:

[0139] Obtain pre-stored loan information and post-loan transaction information;

[0140] Determining a customer identification of each second customer based on the pre-stored loan note information and post-loan transaction information;

[0141] Based on the customer identifier of each second customer, obtaining second user attribute data of each second customer from the pre-stored user attribute information, obtaining second IOU data of each second customer from the IOU information, obtaining post-loan transaction data of each second customer from the post-loan transaction information, and obtaining second customer data of each second customer from the second customer group data stream;

[0142] Based on the second user attribute data, second loan note data, post-loan transaction data and second customer data of each second customer, second service data corresponding to the second customer group is constructed.

[0143] Optionally, the indicator calculation module 502 is further configured to:

[0144] According to a preset time slicing rule, obtaining intermediate IOU information and intermediate post-loan transaction information from the pre-stored IOU information and post-loan transaction information;

[0145] Taking each customer in the intermediate loan note information and the intermediate post-loan transaction information as a second customer, and determining a customer identifier of each second customer;

[0146] The acquiring the second IOU data of each second customer from the IOU information includes: acquiring the second IOU data of each second customer from the intermediate IOU information;

[0147] The obtaining of the post-loan transaction data of each second customer from the post-loan transaction information includes: obtaining the post-loan transaction data of each second customer from the intermediate post-loan transaction information.

[0148] Optionally, the indicator calculation module 502 is further configured to:

[0149] Parsing the online service data using pre-set regular expression parsing rules to obtain a target message value of a target variable, wherein the target message value is a message value common to multiple data sources, and the target data source is any one of the multiple data sources;

[0150] Based on the target message value of the target variable of the online service data, an indicator calculation process is performed on the online service data to obtain a second indicator statistical result of the target data source.

[0151] Optionally, the indicator calculation module 502 is further configured to:

[0152] Classifying the service data to be processed into various dimensions according to a preset dimension division rule, wherein the service data to be processed refers to the historical service data or the online service data;

[0153] Indicator calculation and processing are performed on the service data of each dimension respectively.

[0154] Optionally, the pre-set evaluation indicators include any one or more of the following: number of applicants, total number of people checked, check rate, number of check calculations, number of people approved for credit, average amount of credit, number of credit users, number of overdue people within the first time range, overdue amount within the first time range, number of overdue people within the second time range, overdue amount within the second time range, amount of new IOUs for old customers, and amount of data in boxes.

[0155] In this embodiment, a first indicator statistical result is obtained by performing indicator calculation processing on the historical service data of the target data source. Based on the first indicator statistical result, a determination is made as to whether to perform an online evaluation on the target data source. If the determination is to continue the online evaluation, an indicator calculation processing is performed on the online service data of the target data source to obtain a second indicator statistical result. Ultimately, the value assessment result of the target data source can be determined based on the first indicator statistical result and / or the second indicator statistical result. By setting the conditions for online evaluation, it is possible to avoid evaluating data sources that do not meet the online evaluation conditions, thereby reducing resource consumption and improving the accuracy of the evaluation.

[0156] The exemplary embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being configured to cause the electronic device to perform a method according to an exemplary embodiment of the present disclosure when executed by the at least one processor.

[0157] Exemplary embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present disclosure.

[0158] Exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to perform the method according to the embodiment of the present disclosure.

[0159] refer to Figure 6 , a block diagram of an electronic device 600 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0160] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0161] Multiple components within electronic device 600 are connected to I / O interface 605, including an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. Input unit 606 can be any type of device capable of inputting information into electronic device 600. Input unit 606 can receive digital or textual input and generate key input signals related to user settings and / or function control of the electronic device. Output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 608 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth device, a WiFi device, a Wi-Fi device, a cellular communication device, and / or the like.

[0162] The computing unit 601 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above. For example, in some embodiments, the above-mentioned data source evaluation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to perform the above-mentioned data source evaluation method by any other appropriate means (e.g., by means of firmware).

[0163] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0164] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0165] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0167] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0168] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. A data source evaluation method, characterized in that: The method comprises: Obtain historical service data of the target data source to be evaluated; Performing indicator calculation processing on the historical service data to obtain a first indicator statistical result of the target data source; When determining to perform online evaluation based on the statistical result of the first indicator, obtaining online service data of the target data source, the online service data is an access service data flow, and the access service data flow includes multiple messages; Performing indicator calculation processing on the online service data to obtain a second indicator statistical result of the target data source; Determining a value assessment result of the target data source based on the first indicator statistical result and / or the second indicator statistical result; The performing of indicator calculation processing on the online service data to obtain a second indicator statistical result of the target data source includes: Performing message parsing on the online service data using pre-set regular parsing rules to obtain a target message value of a target variable, wherein the target message value is a message value common to multiple data sources, and the target data source is any one of the multiple data sources, wherein different message value types correspond to different regular parsing rules; Based on the target message value of the target variable of the online service data, an indicator calculation process is performed on the online service data to obtain a second indicator statistical result of the target data source.

2. The method according to claim 1, characterized in that The performing indicator calculation processing on the online service data to obtain a second indicator statistical result of the target data source includes: determining a first customer group data flow and a second customer group data flow in the online service data based on access time information of the access service data flow; Obtaining first user attribute data and first loan receipt data corresponding to a first customer group, and combining them with the first customer group data stream to determine first service data corresponding to the first customer group; Obtaining second user attribute data, second loan note data, and post-loan transaction data corresponding to the second customer group, and combining these with the second customer group data stream to determine second service data corresponding to the second customer group; According to the pre-set evaluation index, the first service data and the second service data are subjected to index calculation processing to obtain the statistical result of the evaluation index as the second index statistical result of the target data source.

3. The method according to claim 2, characterized in that The first customer group includes a plurality of first customers; The obtaining of first user attribute data and first loan receipt data corresponding to the first customer group and combining the data with the first customer group data stream to determine first service data corresponding to the first customer group includes: Acquiring pre-stored user attribute information and loan note information, wherein the user attribute information includes a customer identifier of each first customer; Based on the customer identifier of each first customer, obtaining first user attribute data of each first customer from the user attribute information, obtaining first loan note data of each first customer from the loan note information, and obtaining first customer data of each first customer from the first customer group data stream; Based on the first user attribute data, the first loan note data and the first customer data of each first customer, first service data corresponding to the first customer group is constructed.

4. The method according to claim 2, characterized in that The second customer group includes a plurality of second customers; The obtaining of the second user attribute data, the second loan note data, and the post-loan transaction data corresponding to the second customer group, and combining the data with the second customer group data stream to determine the second service data corresponding to the second customer group, includes: Obtain pre-stored loan information and post-loan transaction information; Determining a customer identification of each second customer based on the pre-stored loan note information and post-loan transaction information; Based on the customer identifier of each second customer, obtaining second user attribute data of each second customer from the pre-stored user attribute information, obtaining second IOU data of each second customer from the IOU information, obtaining post-loan transaction data of each second customer from the post-loan transaction information, and obtaining second customer data of each second customer from the second customer group data stream; Based on the second user attribute data, second loan note data, post-loan transaction data and second customer data of each second customer, second service data corresponding to the second customer group is constructed.

5. The method according to claim 4, characterized in that The step of determining the customer identification of each second customer based on the pre-stored loan note information and post-loan transaction information includes: According to a preset time slicing rule, obtaining intermediate IOU information and intermediate post-loan transaction information from the pre-stored IOU information and post-loan transaction information; Taking each customer in the intermediate loan note information and the intermediate post-loan transaction information as a second customer, and determining a customer identifier of each second customer; The acquiring the second IOU data of each second customer from the IOU information includes: acquiring the second IOU data of each second customer from the intermediate IOU information; The obtaining of the post-loan transaction data of each second customer from the post-loan transaction information includes: obtaining the post-loan transaction data of each second customer from the intermediate post-loan transaction information.

6. The method according to claim 1, characterized in that The indicator calculation process includes: Classifying the service data to be processed into various dimensions according to a preset dimension division rule, wherein the service data to be processed refers to the historical service data or the online service data; Indicator calculation and processing are performed on the service data of each dimension respectively.

7. The method according to any one of claims 1 to 6, characterized in that The pre-set evaluation indicators include any one or more of the following: number of applicants, total number of people found, search rate, number of search calculations, number of people approved for credit, average amount of credit, number of credit users, number of overdue payments within the first time range, amount of overdue payments within the first time range, number of overdue payments within the second time range, amount of overdue payments within the second time range, amount of new IOUs for old customers, and amount of data in each box.

8. A data source evaluation device, characterized in that: The device comprises: A first acquisition module is used to acquire historical service data of a target data source to be evaluated; An indicator calculation module, configured to perform indicator calculation processing on the historical service data to obtain a first indicator statistical result of the target data source; a second acquisition module, configured to acquire online service data of the target data source when determining to perform online evaluation based on the statistical result of the first indicator, the online service data being an access service data flow including a plurality of messages; The indicator calculation module is further configured to perform indicator calculation processing on the online service data to obtain a second indicator statistical result of the target data source; an evaluation module, configured to determine a value evaluation result of the target data source based on the first indicator statistical result and / or the second indicator statistical result; The indicator calculation module is further configured to perform message parsing on the online service data using preset regular parsing rules to obtain a target message value of a target variable, where the target message value is a message value common to multiple data sources, and the target data source is any one of the multiple data sources, wherein different message value types correspond to different regular parsing rules; and perform indicator calculation processing on the online service data based on the target message value of the target variable of the online service data to obtain a second indicator statistical result of the target data source.

9. An electronic device comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Simulation-based data risk control value evaluation method and device, equipment and medium

    CN111127195A