A hierarchical computing method for solving real-time derivation of complex variables of large-scale data
By employing a hierarchical computing approach, credit feature data is divided into static and dynamic layers. It is then processed using offline and real-time computing engines, resolving the contradiction between timeliness and complexity in large-scale data processing under high-performance real-time scenarios and achieving efficient and accurate credit feature calculation.
Patent Information
- Application Number
- CN202511113469.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Existing technologies struggle to simultaneously satisfy the contradiction between timeliness, complexity, and computational resource consumption in large-scale data processing under high-performance real-time scenarios, leading to difficulties in feature update processing for both offline and real-time computing.
A hierarchical computing approach is adopted, which decomposes features into static feature layers and dynamic feature layers. An offline distributed engine is used to process historical data and an online real-time computing engine is used to process incremental data. Data alignment is achieved through logical clocks and version snapshots, thereby reducing the computational load.
It improves the efficiency and accuracy of real-time computing, avoids the problem of low data processing efficiency, and ensures the reliability and accuracy of credit feature data.
Smart Images

Figure CN120634714B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of data processing, and particularly relates to a hierarchical computing method for solving real-time derivation of complex variables of large-scale data. BACKGROUND
[0002] The current real-time feature computing field faces the contradiction between scale, timeliness and complexity, which is difficult to reconcile, and presents the dilemma of taking two of the three, which seriously hinders the realization of the high-value complex feature variable derivation demand in the high-performance real-time scene.
[0003] 1) Between data scale and timeliness: traditional offline computing can process PB-level historical data, but there is a delay of hours, which cannot meet the high real-time scene (such as millisecond-level response) of risk control and the like;
[0004] Although pure real-time computing can respond with low delay, when facing multi-source heterogeneous data (three-party, log, database, etc.), massive data transmission and semi-structured analysis consume a large amount of resources, for example, single feature calculation needs to traverse massive user behavior and credit data, resulting in response time exceeding the threshold, that is, it is extremely difficult to achieve both data scale and timeliness.
[0005] 2) Between complexity and computing timeliness: the business feature computing logic is increasingly complex (such as user 30-day cross-transaction frequency + geographic location aggregation), and traditional real-time computing needs to process full-amount data one by one in the real-time link, and the computing time grows exponentially;
[0006] 3) Between data scale and complexity: when the data scale is large, it is difficult to handle complex business logic, whether using traditional offline computing or real-time computing. If traditional offline computing is used, although it can handle large-scale data, its timeliness is poor and it is difficult to meet the complex and changing real-time business needs; if real-time computing is used, it is difficult to efficiently handle complex logic in the face of massive data, and resource consumption is huge, so in actual application, only two of the three of scale, timeliness and complexity can be chosen, and all three conditions cannot be met at the same time.
[0007] To solve the above technical problems, the existing technical solutions often combine offline computing and real-time computing to perform computing and processing, thereby ensuring the real-time of computing and processing, but the above technical solutions have the following technical defects:
[0008] The features of offline computing and real-time computing in the existing technical solutions are often fixed, which increases the proportion of data volume of offline computing, and gradually increases the influence of the result of computing and processing by offline computing, which makes it an urgent technical problem to dynamically update the features of offline computing and real-time computing.
[0009] To solve the above technical problems, the application provides a hierarchical computing method for real-time derivation of complex variables of large-scale data. SUMMARY
[0010] To achieve the object of the application, the application adopts the following technical solutions:
[0011] Specifically, the application provides a hierarchical computing method for real-time derivation of complex variables of large-scale data, which specifically comprises:
[0012] S1 determining real-time optimization services in the credit services in the case of data heterogeneity of data sources of credit users involved in a credit model of the credit services, and determining real-time computing users in the credit users based on credit characteristic data of the credit users in the real-time optimization services;
[0013] S2 if the number of real-time computing users in the credit users meets the requirements, regarding the credit users except the real-time computing users as other users, determining offline data volume of the other users in credit characteristics based on credit characteristic data of the other users, and determining credit characteristics for real-time computing based on the offline data volume;
[0014] S3 determining fallback of the credit characteristics for real-time computing to credit characteristics for offline computing according to distribution data of the credit characteristic data of the real-time computing users in different data sources.
[0015] The application has the following beneficial effects:
[0016] Based on the credit characteristic data of the credit users in the real-time optimization services, the real-time computing users in the credit users are determined, so that the identification processing of the credit users with less number of data sources involved in the credit characteristic data and smaller data volume of the credit characteristic data is realized, and the credit characteristics of the credit users are processed in the real-time computing mode, so that the reliability and accuracy of the data processing of the credit characteristics of the real-time computing users are ensured, and the technical problems of the occurrence of data processing problems and low efficiency of data processing when offline computing and real-time computing are used to process data in different computing models are avoided.
[0017] According to the distribution data of the credit characteristic data of the real-time computing users in different data sources, the fallback of the credit characteristics for real-time computing to the credit characteristics for offline computing is determined, so that the technical problem of slow computing and processing efficiency of the credit characteristics for real-time computing caused by slow computing and processing efficiency of the credit characteristics of the real-time computing users due to the increase of the data volume of the real-time computing users is avoided, and the computing and processing efficiency and accuracy of the real-time computing users and the credit characteristics for real-time computing are further improved by the fallback of the credit characteristics for real-time computing to offline computing.
[0018] Further, the data source involved in the real-time calculation includes a data source of user information of the credit user, and can understand that a plurality of third-party data sources containing credit data of the credit user, in a possible embodiment, include a people's bank data source, a shopping platform, a market credit investigation agency, a communication operator and an education verification platform.
[0019] Further, the data heterogeneity is determined according to the data format of different data sources, and specifically includes a data source of a data format that needs to be converted to meet the data calculation processing requirement.
[0020] It can be understood that the credit business is divided according to the customer type of the credit customer, and in a possible embodiment, specifically includes enterprise customers and individual customers, and the corresponding credit model is a mathematical model used for credit risk assessment of the credit customer of the credit business.
[0021] Further, the method for determining the real-time optimization business in the credit business is:
[0022] According to the data heterogeneity of the data source of the credit user involved in the credit model of the credit business, the deviation of the data format of different data sources is determined;
[0023] Based on the deviation of the data format of different data sources, the data source that needs to be converted in data format for data processing of the credit model is determined;
[0024] According to the deviation of the data format of different data sources that need to be converted, it is determined whether the credit business is a real-time optimization business.
[0025] Further, the method for determining the fallback of the real-time calculated credit feature to the offline calculated credit feature is:
[0026] According to the distribution data of the credit feature data of the real-time calculation user in different data sources, the data amount of the credit feature data of the real-time calculation user in different data sources is determined;
[0027] Based on the change of the data amount of the credit feature data in different data sources, the data amount change data source of the credit feature data of different real-time calculation users is determined;
[0028] Through the data amount change data source of the credit feature data of different real-time calculation users, the fallback of the real-time calculated credit feature to the offline calculated credit feature is determined.
[0029] Other features and advantages will be set forth in the following description, and the objects and other advantages of the present application will be achieved and obtained by the structure particularly pointed out in the description and the drawings.
[0030] In order to make the above-mentioned objects, features and advantages of the present application more apparent, the following preferred embodiments are specifically described below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0031] The above and other features and advantages of the present application will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings.
[0032] Figure 1 is a flow chart of a hierarchical computing method for solving real-time derivation of complex variables of large-scale data;
[0033] Figure 2 is a flow chart of a method for determining real-time optimization business in a credit business;
[0034] Figure 3 is a flow chart of a method for determining real-time calculation of users in a credit user;
[0035] Figure 4 is a flow chart of a method for determining credit features for real-time calculation in a credit feature;
[0036] Figure 5 is a flow chart of a method for determining rollback of real-time calculation credit features to offline calculation credit features. DETAILED DESCRIPTION
[0037] In order to make the technical solutions in the specification better understood by those skilled in the art, the technical solutions in the specification will be described clearly and completely below in conjunction with the drawings in the embodiments of the specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all the embodiments. Based on the embodiments of the specification, all other embodiments obtained by those skilled in the art without creative labor should be within the protection scope of the specification.
[0038] The current real-time feature calculation field faces the contradiction between "scale, timeliness, and complexity" which is difficult to reconcile, and presents a dilemma that only two of the three can be taken, which seriously hinders the realization of the high-value complex feature variable derivation demand in high-performance real-time scenarios:
[0039] 1) Between data scale and timeliness: traditional offline calculation can process PB-level historical data, but there is a delay of hours, which cannot meet the high real-time scene (such as millisecond-level response) of risk control, etc.
[0040] Pure real-time computing can respond with low latency, but when faced with multi-source heterogeneous data (pedestrian, hundred lines, etc. Three parties, logs, databases), massive data transmission and semi-structured analysis consume a lot of resources, such as single feature calculation needs to traverse massive user behavior, credit data, resulting in response time exceeding threshold, that is, it is extremely difficult to achieve both data size and timeliness.
[0041] 2) Complexity and computing timeliness: Business feature calculation logic is becoming more complex (such as user 30-day cross-transaction frequency + geographic location aggregation), and traditional streaming computing needs to process full data one by one in real-time link, and the time-consuming of calculation grows exponentially;
[0042] 3) Data size and complexity: When the data size is large, it is difficult to handle complex business logic, whether using traditional offline computing or real-time computing. If you use traditional offline computing, although it can handle large-scale data, its timeliness is poor and it is difficult to meet the complex and changing real-time business needs; If you use real-time computing, it is difficult to efficiently handle complex logic when faced with massive data, and resource consumption is huge.
[0043] Therefore, in practical applications, we often have to choose two of the three: size, timeliness, and complexity, and it is difficult to meet all three conditions.
[0044] 1) Hierarchical-division-of-labor-staged computing architecture (breakthrough design)
[0045] Core innovation: Decompose features into static feature layer (offline pre-computation) and dynamic feature layer (online lightweight computation), and realize "complex logic decoupling" through computing granularity staging.
[0046] Static layer: Offline distributed engine (Spark / Hive) processes historical full data, generates high-complexity intermediate features (such as user historical borrowing behavior data), and stores them in columnar database (HBase) and establishes secondary index;
[0047] Dynamic layer: Online real-time computing engine only processes incremental data and state updates, for example, "user historical loan application number" is decomposed into "offline pre-stored T-1 data loan application number" + "real-time cumulative loan application number today";
[0048] Traffic peak automatic switching computing mode, for example, in the low peak period, some real-time tasks are returned to offline batch processing to reduce cluster load.
[0049] 2) Spatiotemporal folding consistency guarantee mechanism
[0050] Core innovation: Realize offline-online data automatic alignment through logical clock + version snapshot, ensure seamless splicing of offline pre-computation and real-time incremental data, and completely eliminate version conflict and data loss risk.
[0051] Embodiment 1
[0052] As Figure 1 indicated, the application provides a hierarchical computing method for solving real-time derivation of complex variables of large-scale data, specifically including:
[0053] S1 determines the real-time optimization business in the credit business in the light of the data heterogeneity of the data source of the credit model of the credit business, and determines the real-time computing user in the credit user based on the credit feature data of the credit user in the real-time optimization business.
[0054] Further, the data source involved in the real-time computing includes the data source of the user information of the credit user, and it can be understood that it includes a plurality of third-party data sources containing credit data of the credit user, including the data source of the People's Bank of China, shopping platforms, market credit investigation institutions, communication operators and education verification platforms in a possible embodiment.
[0055] Further, the data heterogeneity is determined according to the data format of different data sources, specifically including the data source of the data format which needs to be converted to meet the data format requirement of data computing processing.
[0056] It can be understood that the credit business is divided according to the customer type of the credit customer, and in a possible embodiment, it specifically includes enterprise customers and individual customers, and the corresponding credit model is a mathematical model for credit risk assessment of the credit customer of the credit business.
[0057] Specifically, as Figure 2 indicated, the method for determining the real-time optimization business in the credit business is:
[0058] Determine the deviation of the data format of different data sources in the light of the data heterogeneity of the data source of the credit model of the credit business;
[0059] Determine the data source that needs data format conversion for data processing of the credit model based on the deviation of the data format of different data sources;
[0060] Determine whether the credit business is a real-time optimization business in the light of the deviation of the data format of different data sources that need data format conversion.
[0061] It can be understood that when the data source requiring data format conversion does not exist multiple data formats and the number of data sources requiring data format conversion meets the requirements, in a possible embodiment, the data formats of the data sources requiring data format conversion are all the same and the number of data sources requiring data format conversion is less than 3, then the data processing difficulty of the credit business is smaller, so the credit business is taken as a real-time optimization business, that is, the online calculation mode is used for part of the credit users.
[0062] In addition, it should be noted that if the data formats of the data sources requiring data format conversion are not all the same or the number of data sources requiring data format conversion is not less than 3, then the data processing difficulty of the credit business is greater, so the offline calculation and real-time calculation are combined to process the data of the credit model.
[0063] Optionally, the method for determining the real-time optimization business in the credit business comprises:
[0064] The data of the credit model of the credit business involves data heterogeneity of the data source of the credit user, and the data source requiring data format conversion is determined for data processing of the credit model;
[0065] The number of data sources requiring data format conversion is determined according to different data format deviations of the data sources requiring data format conversion.
[0066] The number of data sources requiring data format conversion is determined according to different data format deviations of the data sources requiring data format conversion.
[0067] It should be noted that when either the number of data formats or the number of data sources requiring data format conversion does not meet the requirements, that is, the number is greater than the preset threshold, then it is determined that the credit business is not a real-time optimization business.
[0068] Specifically, as shown in the method for determining the real-time calculation user in the credit user comprises: Figure 3
[0069] The data amount of the credit data in different data sources is determined based on the credit feature data of the credit user in the real-time optimization business.
[0070] The data source existing the credit feature is determined according to the data amount of the credit data in different data sources.
[0071] The credit user is determined to be a real-time calculation user according to the data amount of the credit data of the data source existing the credit feature of the credit user.
[0072] It should be noted that the credit features include professional features, identity features, marital status, social security features, real estate holding features, and historical loan features.
[0073] It can be understood that, according to the data amount of the credit data of the data source existing the credit features of the credit user, whether the credit user is a real-time calculation user is determined, specifically including:
[0074] According to the data amount of the credit data of the data source existing the credit features of the credit user, the data source existing the credit features is determined.
[0075] When the data processing time length of the total data amount of the credit data of all the data sources existing the credit features of the credit user does not meet the requirement, it is determined that the credit user does not belong to the real-time calculation user.
[0076] In one possible embodiment, the total data amount of the credit data of all the data sources existing the credit features of the credit user is taken as the credit analysis data amount, wherein the data processing time length under the credit analysis data amount is determined according to the historical average time length of the feature calculation and the model result output according to the online model, and when the historical average time length is more than 20 seconds, it is determined that the credit user does not belong to the real-time calculation user.
[0077] Optionally, the method for determining the real-time calculation user in the credit user is:
[0078] Based on the credit feature data of the credit user in the real-time optimization business, the data amount of the credit data in different data sources is determined.
[0079] The data source existing the credit features is determined according to the data amount of the credit data in different data sources.
[0080] According to the number of the data sources existing the credit features of the credit user and the data amount of the credit data of different data sources, whether the credit user is a real-time calculation user is determined.
[0081] It can be understood that, when the credit user has a data source with a data amount of credit data greater than a preset data amount value or the number of the data sources existing the credit features of the credit user is greater than a preset number value, it is determined that the credit user does not belong to the real-time calculation user.
[0082] It should be noted that, when not belonging to the real-time calculation user, the data processing of the credit model is performed in a combination of offline calculation and real-time calculation, and when belonging to the real-time calculation user, the data processing of the credit model from different data sources is performed in real-time calculation.
[0083] S2If the number of real-time computing users in the credit user meets the requirement, the credit user of the real-time computing user is removed as the other user, the offline data amount of the other user in the credit feature is determined based on the credit feature data of the other user, and the credit feature in which real-time computing is performed is determined based on the offline data amount.
[0084] Further, when the proportion of the number of credit services of the real-time computing user in the credit user under the credit service type is less than 0.05, it is determined that the number of real-time computing users in the credit user meets the requirement.
[0085] It should be noted that when the number of real-time computing users in the credit user does not meet the requirement, the combination of offline computing and real-time computing is used in all credit features for data processing of the credit model.
[0086] Among them, real-time computing is used to update the credit data of different credit features of different data sources on the same day, offline computing is used to calculate the credit features except the current updated credit data, the credit features are obtained, the credit features are used as the input of the credit model, and the credit risk is identified and processed.
[0087] Specifically, as shown in Figure 4 The method for determining the credit feature in which real-time computing is performed in the credit feature is:
[0088] The average value of the offline data amount of the other user in the credit feature is determined based on the offline data amount of the other user in the credit feature.
[0089] Based on the average value of the offline data amount of different other users in the credit feature, it is determined whether the credit feature is a credit feature in which real-time computing is performed.
[0090] It should be noted that the offline data amount is the data amount stored in the offline database of the credit feature of the other user.
[0091] It can be understood that based on the average value of the offline data amount of different other users in the credit feature, it is determined whether the credit feature is a credit feature in which real-time computing is performed, which specifically includes:
[0092] The average value of the offline data amount of the different other users in the credit feature is taken as the credit feature data amount.
[0093] The credit feature in which real-time computing is performed is determined as the credit feature whose number is greater than the credit feature data amount of the credit feature.
[0094] In a possible embodiment, if the credit characteristic data amount of two-thirds of the credit characteristics is more than the credit characteristic data amount of the credit characteristic, the credit characteristic is determined as the credit characteristic for real-time calculation, that is, one-third of the credit characteristics with the smallest credit characteristic data amount is taken as the credit characteristic for real-time calculation.
[0095] S3 determines the fallback of the real-time calculated credit characteristic to the offline calculated credit characteristic according to the distribution data of the credit characteristic data of the real-time calculation user in different data sources.
[0096] Specifically, as shown in Figure 5 the method for determining the fallback of the real-time calculated credit characteristic to the offline calculated credit characteristic is:
[0097] According to the distribution data of the credit characteristic data of the real-time calculation user in different data sources, the data amount of the credit characteristic data of the real-time calculation user in different data sources is determined.
[0098] Based on the change of the data amount of the credit characteristic data in different data sources, the data amount change data source of the credit characteristic data of different real-time calculation users is determined.
[0099] In a possible embodiment, the data source with a data amount change rate greater than 20% is taken as the data amount change data source, that is, the increase amount of the credit characteristic data in the data source of the real-time calculation user from the start of the real-time calculation processing to the present time is compared with the data amount of the data source from the start of the real-time calculation processing to determine the change rate.
[0100] The fallback of the real-time calculated credit characteristic to the offline calculated credit characteristic is determined through the data amount change data source of the credit characteristic data of different real-time calculation users.
[0101] It should be noted that the change of the data amount of the credit characteristic data in different data sources is determined according to the increase amount of the credit characteristic data in the data source of the real-time calculation user from the start of the real-time calculation processing to the present time.
[0102] Specifically, if the sum of the data amount of the credit characteristic data of the real-time calculation user in different data sources is small, in a possible embodiment, the proportion of the sum of the data amount of the credit characteristic data of the real-time calculation user in different data sources in the credit users in the credit business type is taken as the data amount proportion in different data sources.
[0103] In a possible embodiment, if the proportion of the data quantity in different data sources all meets the requirement, it is determined that the credit features calculated in real time do not need to be rolled back to the credit features calculated offline, wherein if the proportion of the data quantity in different data sources is all less than 0.05, it is determined that the proportion of the data quantity in different data sources all meets the requirement.
[0104] In addition, it is to be noted that if there is a data source whose proportion of the data quantity does not meet the requirement, at this time if the proportion of the data quantity in different data sources all does not meet the requirement, all the credit features calculated in real time are rolled back to the credit features calculated offline.
[0105] If there is a data source whose proportion of the data quantity meets the requirement, at this time if there is no data source with data quantity variation for different real-time computing users, it is determined that the credit features calculated in real time do not need to be rolled back to the credit features calculated offline.
[0106] Further, if there is a real-time computing user with a data source with data quantity variation, the number of real-time computing users with a data source with data quantity variation is determined, and when the number of real-time computing users with a data source with data quantity variation does not meet the requirement, in a possible embodiment, if the proportion of the credit users of the real-time computing users with a data source with data quantity variation in the credit business type is greater than 0.03, it is determined that all the credit features calculated in real time need to be rolled back to the credit features calculated offline.
[0107] If the number of real-time computing users with a data source with data quantity variation meets the requirement, the proportion of the credit features calculated in real time that are rolled back to the credit features calculated offline is determined according to the proportion of the credit users of the real-time computing users with a data source with data quantity variation in the credit business type, and the total amount of the credit data of the credit features calculated in real time in different data sources is used as a basis to determine the credit features calculated in real time that are rolled back to the credit features calculated offline.
[0108] In a possible embodiment, the proportion of the credit features calculated in real time that are rolled back to the credit features calculated offline is determined according to the product of the proportion of the credit users of the real-time computing users with a data source with data quantity variation in the credit business type and 10.
[0109] In addition, the total amount of the credit data of the credit features calculated in real time in different data sources is used as a basis to determine the credit features calculated in real time that are rolled back to the credit features calculated offline, specifically including:
[0110] The number of the credit features calculated in real time that are rolled back to the credit features calculated offline is determined according to the product of the proportion and the number of the credit features calculated in real time.
[0111] Determine the offline credit features to which the real-time calculated credit features are to be rolled back based on the number of real-time calculation users of the different data sources and the total amount of credit data in the different data sources.
[0112] Optionally, the method for determining the offline credit features to which the real-time calculated credit features are to be rolled back comprises:
[0113] Determine the amount of credit data of the real-time calculation user in the different data sources based on the distribution data of the real-time calculated credit data of the user in the different data sources.
[0114] Determine the data amount variation data source of the credit data of the different real-time calculation users based on the variation of the amount of credit data in the different data sources.
[0115] Determine the offline credit features to which the real-time calculated credit features are to be rolled back based on the number of real-time calculation users of the different data sources belonging to the data amount variation data source.
[0116] Further, when the different data sources do not belong to the data amount variation data source for the different real-time calculation users, it is determined that the real-time calculated credit features do not need to be rolled back to offline calculation.
[0117] In addition, it can be understood that when there is a data amount variation data source, the interference weight value of the different data sources is determined by the ratio of the number of real-time calculation users of the different data sources belonging to the data amount variation data source, i.e., the ratio in all real-time calculation users.
[0118] When the sum of the interference weight values of the different data sources is greater than a preset interference threshold, in a possible embodiment, when the sum of the interference weight values of the different data sources is greater than 0.6 times the product of the number of data sources, all real-time calculated credit features are rolled back to offline calculation, and when the sum of the interference weight values of the different data sources is not greater than the preset interference threshold, but within a certain interval, in a possible embodiment, between 0.2 times the product of the number of data sources and 0.6 times the product of the number of data sources, the rollback ratio is determined by the product of the average value of the interference weight values of the different data sources and a preset proportion factor, wherein the value of the preset proportion factor is one-half, and the offline credit features to which the real-time calculated credit features are to be rolled back are determined based on the rollback ratio and the total amount of credit data in the different data sources from large to small.
[0119] When not within a certain interval, it is determined that the real-time calculated credit features do not need to be rolled back to offline calculation.
[0120] The various embodiments in this specification describe the application in progressive stages. Each stage builds on the previous stages, and each stage can be described in the context of similar or identical stages in other embodiments. Each stage is intended to highlight differences between that stage and other stages. In particular, the device, apparatus, and non-transitory computer storage medium embodiments are described relatively simply because they are substantially similar to the method embodiments. The relevant portions of the method embodiments are referenced.
[0121] The above description only illustrates certain embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order and still accomplish the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.
[0122] The above description only illustrates certain embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order and still accomplish the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.
[0122] The above description only illustrates certain embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order and still accomplish the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.
Claims
1. A hierarchical computing method for solving real-time derivation of complex variables from large-scale data, characterized in that, Specifically comprising: The credit model of the credit business involves data heterogeneity of data sources of credit users, determines real-time optimization business in the credit business, determines real-time calculation users among credit users in the real-time optimization business based on credit characteristic data of the credit users in the real-time optimization business; If the number of real-time calculation users among credit users in the real-time optimization business meets the requirement, the credit users except the real-time calculation users are regarded as other users, the offline data amount of the other users in credit characteristics is determined based on credit characteristic data of the other users, and credit characteristics for real-time calculation of the other users are determined based on the offline data amount; The distribution data of credit characteristic data of the real-time calculation users in different data sources is determined, and the real-time calculation credit characteristic is determined to fall back to offline calculation credit characteristic; The offline data amount is the data amount stored in the offline database of the credit characteristics of the other users; The method for determining the real-time calculation users among credit users in the real-time optimization business comprises: The data amount of credit data in different data sources is determined based on credit characteristic data of credit users in the real-time optimization business; The data source of credit characteristics existing in the real-time optimization business is determined based on the data amount of credit data of the data source of credit characteristics existing in the real-time optimization business. The data source involved in real-time calculation includes the data source of user information of credit users involved in the credit model.
2. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 1, wherein, The data heterogeneity is determined according to the data format of different data sources.
3. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 1, wherein, The credit business is divided according to the customer type of credit customers.
4. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 1, wherein, The method for determining the real-time optimization business in the credit business comprises:
5. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 1, wherein, The deviation of data formats of different data sources is determined based on data heterogeneity of data sources of credit users involved in the credit model of the credit business; The data source requiring data format conversion for data processing of the credit model is determined based on the deviation of data formats of different data sources; Whether the credit business is real-time optimization business is determined based on the deviation of data formats of different data sources requiring data format conversion. When there is only one data format in the data source requiring data format conversion and the number of data sources requiring data format conversion meets the requirement, the credit business is regarded as real-time optimization business.
6. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 5, wherein, The credit characteristics of the credit users include professional characteristics, identity characteristics, marital status, social security characteristics, real estate holding characteristics and historical loan characteristics.
7. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 1, wherein, When the number of real-time calculation users among credit users in the real-time optimization business does not meet the requirement, offline calculation and real-time calculation are combined to process data of the credit model in all credit characteristics in the real-time optimization business.
8. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 1, wherein, The method for determining the credit characteristics for real-time calculation of the other users comprises:
9. The hierarchical computing method for solving real-time complex variable derivation of large-scale data according to claim 1, wherein, The average value of offline data amount of the other users in credit characteristics is determined based on the offline data amount of the other users in credit characteristics. determining whether credit features of different other users are credit features for which to perform real-time calculations based on an average of an offline amount of data in the credit features of the other users.
Citation Information
Patent Citations
Data processing method and device
CN115797056A
Data processing method and device, storage medium and computing equipment
CN116185977A