Multi-dimensional feature derivation method and device based on master model
By using a multi-dimensional feature derivation method based on the master model, the problem of multi-dimensional reflection and automation of feature extraction from time series data is solved, achieving flexible and robust feature computation that is applicable to various business scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for time series data feature extraction suffer from problems such as single feature dimensions, low automation, low computational efficiency, poor scalability, and insufficient robustness, making it difficult to meet the needs of different business scenarios.
A multi-dimensional feature derivation method based on the master model is adopted. The master model predicts the user's target evaluation results within a preset time period. By combining multiple feature dimensions and feature functions, the derived features are calculated and merged to form a feature matrix, which supports multi-dimensional features to reflect the complex characteristics of time series data.
It enables automated and flexible feature derivation of time series data, enhances feature robustness and computational efficiency, adapts to the needs of different business scenarios, and improves the accuracy and applicability of feature calculation.
Smart Images

Figure CN121808244A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of data analysis and machine learning technology, and more specifically, to a method and apparatus for multidimensional feature derivation based on a master model. Background Technology
[0002] In the fields of data analysis and machine learning, feature engineering is a key step in building high-performance models. For feature extraction from time series data, existing technologies mainly employ the following methods: 1. Manual Feature Engineering: Relying on the expertise and experience of data analysts, features are designed and calculated manually.
[0003] 2. Simple statistical characteristics: Calculate the basic statistics of the data, such as mean, variance, maximum value, minimum value, etc.
[0004] 3. Traditional time series characteristics: such as moving average, exponential smoothing, differencing, etc.
[0005] 4. Automatic feature extraction based on deep learning: Automatically learns data features through models such as neural networks.
[0006] The aforementioned prior art has the following disadvantages: 1. Single feature dimension: Features extracted by traditional methods are difficult to fully reflect the complex characteristics of time series and cannot simultaneously capture information from multiple dimensions such as momentum, trend and fluctuation.
[0007] 2. Low degree of automation: Manual feature engineering is time-consuming and labor-intensive, and relies on personal experience, making it difficult to cover all possible effective features.
[0008] 3. Low computational efficiency: When processing large-scale datasets, traditional serial computing methods cannot meet real-time requirements.
[0009] 4. Poor scalability: Existing methods are usually designed for specific scenarios and are difficult to apply flexibly to different business needs.
[0010] 5. Lack of a systematic feature system: Existing technologies lack a systematic classification and unified management of features, which is not conducive to feature reuse and maintenance.
[0011] 6. Insufficient ability to handle complex data: Existing technologies are not robust enough when faced with missing data, outliers, etc. Summary of the Invention
[0012] In view of this, this application provides a multi-dimensional feature derivation method and apparatus based on a master model, which can characterize the multi-dimensional derived features of the time series data in a preset time period from multiple feature dimensions, with the master model as the anchor, for the time series data corresponding to the user in the target business scenario. This is conducive to more comprehensively reflecting the different characteristics of the time series data, realizing the automatic derivation of features, and the parameter configuration is more flexible, which can support the actual needs of different business scenarios, and also helps to enhance the robustness of feature calculation.
[0013] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.
[0014] In a first aspect, embodiments of this application provide a multidimensional feature derivation method based on a master model, the multidimensional feature derivation method comprising: The user's historical behavior data in the target business scenario is input into the main model. The main model then predicts the target evaluation result of the user at each time node within a preset time period, and outputs the time series data of the user within the preset time period. The target evaluation result is determined based on the target business scenario. For multiple preset feature dimensions, based on multiple feature functions corresponding to each feature dimension, multiple derived features corresponding to the time series data under each feature dimension are calculated respectively. The multiple derived features corresponding to each feature dimension of the time series data are merged to obtain the feature matrix corresponding to the user in the target business scenario.
[0015] Secondly, embodiments of this application provide a multi-dimensional feature derivation device based on a master model, the multi-dimensional feature derivation device comprising: The data input module is used to input the user's historical behavior data in the target business scenario into the main model, and through the main model, predict the target evaluation result of the user at each time node within a preset time period, and output the time series data of the user within the preset time period; wherein, the target evaluation result is determined according to the target business scenario; The feature derivation module is used to calculate, based on multiple feature functions corresponding to each feature dimension, multiple derived features of the time series data under each feature dimension, for multiple preset feature dimensions. The result output module is used to merge multiple derived features corresponding to each feature dimension of the time series data to obtain the feature matrix corresponding to the user in the target business scenario.
[0016] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described custom task jump method.
[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described custom task jump method.
[0018] The technical solutions provided by the embodiments of this application may include the following beneficial effects: This application provides a method and apparatus for multidimensional feature derivation based on a master model. It can characterize the multidimensional derived features of the time series data in a preset time period from multiple feature dimensions, using the master model as the anchor, for the time series data corresponding to the user in the target business scenario. This is conducive to more comprehensively reflecting the different characteristics of the time series data, realizing the automatic derivation of features, and making the parameter configuration more flexible. It can support the actual needs of different business scenarios and also enhance the robustness of feature calculation. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating a multi-dimensional feature derivation method based on a master model provided in an embodiment of this application is shown. Figure 2 This illustration shows a schematic diagram of a multi-dimensional feature derivation device based on a master model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device 300 provided in an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0022] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0023] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0024] To facilitate understanding of the embodiments of this application, a multidimensional feature derivation method and apparatus based on a master model provided in the embodiments of this application will be described in detail below.
[0025] Reference Figure 1 As shown, Figure 1 The diagram illustrates a flowchart of a multidimensional feature derivation method based on a master model provided in an embodiment of this application. The multidimensional feature derivation method includes steps S101-S103; specifically: S101, input the user's historical behavior data in the target business scenario into the main model, and use the main model to predict the target evaluation result of the user at each time node within the preset time period, and output the time series data of the user within the preset time period.
[0026] Here, the target business scenarios include, but are not limited to: financial risk control scenarios, user consumption amount analysis scenarios, user credit card limit assessment scenarios, etc. The specific scenario type of the target business scenario can be flexibly adjusted according to actual business needs, and this application embodiment does not impose any limitations on it.
[0027] Specifically, the aforementioned historical behavior data can be determined based on the user's historical behavior related to the target business scenario at each time node within a preset time period. The preset time period can be flexibly adjusted according to the actual characteristic analysis needs of the user within the target business scenario. For example, the preset time period could be the most recent six months or the most recent year. The time node represents the length of a unit of time within the preset time period. For example, if the preset time period is the most recent six months, then the time node could represent each day within the most recent six months. This application embodiment does not impose any mandatory limitations on the specific time length represented by the preset time period and the time node.
[0028] For example, if the target business scenario is a user spending analysis scenario, the user's historical behavior data in this target business scenario can be the user's spending records at each time point within a preset time period (e.g., the user's daily spending amount in the last six months); if the target business scenario is a user credit card limit assessment scenario, the user's historical behavior data in this target business scenario can be the user's credit card usage records at each time point within a preset time period (e.g., the user's daily credit card usage records in the last six months).
[0029] Here, the main model can be a time series model capable of generating time series data, wherein the target evaluation result is determined based on the target business scenario.
[0030] For example, if the target business scenario is a credit assessment scenario, the target assessment result can be the user's credit score. In this case, taking the preset time period as the last six months as an example, the user's historical behavior data in the credit assessment scenario (such as the user's repayment records, overdue records and other historical credit performance data in the last six months) can be input into the main model. The main model can then predict the user's credit score (i.e., the target assessment result) for each day in the last six months (i.e., each time point in the preset time period) and output the user's credit score results for each day in the last six months as the aforementioned time series data.
[0031] For example, if the target business scenario is a credit card limit assessment scenario, the target assessment result can be the user's credit card limit assessment result. In this case, taking the preset time period as the last six months as an example, the user's historical behavior data in the credit card limit assessment scenario (such as the user's credit card consumption records, income data, expenditure data, etc. in the last six months) can be input into the main model. The main model can then predict the user's credit card limit (i.e., the target assessment result) for each day in the last six months (i.e., each time point in the preset time period) and output the user's credit card limit assessment result for each day in the last six months as the aforementioned time series data.
[0032] It should be noted that in this embodiment, each user can correspond to an ID (identity document) account, so that the application device (i.e., the application device based on the multi-dimensional feature derivation method of the master model provided in this embodiment) can distinguish the historical behavior data of different users in the same target business scenario according to different ID accounts. In this embodiment, the processing method of the above-mentioned historical behavior data corresponding to each user ID is the same, and the above-mentioned historical behavior data corresponding to multiple user IDs can be processed in parallel. This embodiment does not limit the number of users involved in actual applications.
[0033] Here, after obtaining the user's time series data within a preset time period through the output of the main model, before executing step S102, as an optional embodiment, the time series data can be preprocessed according to the method shown in steps a1-a2 below. This is to further improve the feature accuracy of the final derived features when calculating derived features from the preprocessed time series data. Specifically: Step a1: In response to the presence of a target value in the time series data, calculate the target prediction value of the target evaluation result corresponding to the target time node in the time series data based on the predicted values of the target evaluation results corresponding to other time nodes in the time series data, and use the target prediction value as the filling value corresponding to the target value to obtain the filled time series data.
[0034] Here, the target value is an outlier or missing value in the time series data, the target time node represents the time node corresponding to the target value within the preset time period, and other time nodes are time nodes other than the target time node within the preset time period.
[0035] For example, if the time series data consists of a user's credit score results for each day within the last six months (i.e., the preset time period), and the credit score result for the 6th day of the 2nd month is missing from the time series data, then the target time point can be determined to be the 6th day of the 2nd month, and the target value is the missing credit score result for the target time point (i.e., the target value is a missing value at this time).
[0036] For example, if the time series data is the user's credit score results for each day within the last six months (i.e., the preset time period), and the credit score result on the 7th day of the 3rd month in the time series data is much greater than or much less than the average of the credit score results at other time points in the time series data, then the target time point can be determined to be the 7th day of the 3rd month, and the target value is the abnormal credit score result corresponding to the target time point (i.e., the target value is an outlier at this time).
[0037] Specifically, in step a1, the above target predicted value can be calculated using calculation methods such as mean calculation, median calculation, and interpolation calculation. If the target value is a missing value, the calculated target predicted value can be used to fill the missing value. If the target value is an outlier, the calculated target predicted value can be used to replace the outlier, thereby obtaining complete and accurate time series data.
[0038] Step a2: In response to the mismatch between the dimension of the time series data and the target dimension corresponding to the target business scenario, the time series data is mapped to the value range corresponding to the target dimension to obtain time series data that matches the target dimension.
[0039] Here, in order to ensure that the time series data can meet the calculation requirements derived from the subsequent multi-dimensional features, it is necessary to standardize the time series data, that is, to ensure that the dimensions of the time series data match the target dimensions corresponding to the target business scenario.
[0040] For example, taking a credit assessment scenario as an example, the credit score that users can see on the credit assessment page is usually in the range of 300-850. Since the main model predicts the credit assessment result of the user at each time point within a preset time period based on the user's credit performance data (i.e., historical behavior data), the credit assessment result directly output by the main model at each time point is a value between 0 and 1 (equivalent to the probability value of the user's creditworthiness). At this time, since the dimension of the time series data directly output by the main model does not match the target dimension corresponding to the credit assessment scenario, data standardization can be used to map the value in the time series data from the range of 0-1 to the range of 300-850, thereby obtaining time series data that matches the target dimension.
[0041] It should be noted that, in this embodiment of the application, when there are multiple users, it is also necessary to ensure that the time series data of all user IDs are aligned in the time dimension.
[0042] S102, for multiple preset feature dimensions, based on multiple feature functions corresponding to each feature dimension, calculate multiple derived features of the time series data under each feature dimension.
[0043] In this embodiment, the preset multiple feature dimensions may include, but are not limited to, momentum feature dimension, trend feature dimension, and volatility feature dimension. Each feature dimension corresponds to multiple feature functions. By calculating the feature functions, multiple derived features of the time series data under each feature dimension can be obtained. This allows for a more comprehensive reflection of the characteristics of the time series data under different feature dimensions based on a systematic multidimensional feature system, so as to provide rich feature inputs for various machine learning models.
[0044] It should be noted that, compared with existing technologies that directly input time series data into deep learning models such as LSTM and GRU and automatically extract features from the time series data through deep learning models, the method of performing multi-dimensional feature derivation calculation on time series data based on a systematic multi-dimensional feature system (i.e., multiple feature dimensions and multiple feature functions corresponding to each feature dimension) in the embodiments of this application is beneficial for capturing more complex nonlinear relationships in time series data.
[0045] Here, the aforementioned momentum feature dimension is mainly used to reflect the short-term abnormal change pattern of time series data within a preset time period; wherein, for the aforementioned momentum feature dimension, the application device can calculate multiple momentum derived features of the time series data under the momentum feature dimension according to multiple feature functions corresponding to the momentum feature dimension; wherein, the multiple momentum derived features are used to reflect the direction and speed of change of the time series data within the preset time period.
[0046] Specifically, as an optional embodiment, under the aforementioned momentum feature dimension, 12 derived features of the time series data under the aforementioned momentum feature dimension can be calculated based on the 12 feature functions shown below: 1. Num function: Calculates the number of months in the most recent p months where the variable value (i.e., the data value corresponding to each time node in the time series data) is greater than a threshold, reflecting the activity level of the variable.
[0047] Calculation method: Perform threshold judgment on the data within the specified time window (i.e., the most recent p months) and count the number of months that meet the conditions.
[0048] Mathematical formula:
[0049] Where I(·) is an indicator function, which is 1 when the condition is met and 0 otherwise; This represents the variable value for month i (i.e., the data corresponding to month i in the time series data), and threshold is the set threshold value.
[0050] 2. Nmz function: Calculates the number of months in the most recent p months where the variable value (i.e., the data value corresponding to each time node in the time series data) is less than or equal to a threshold, reflecting the degree of inactivity of the variable.
[0051] Calculation method: Perform threshold judgment on the data within the specified time window (i.e., the most recent p months) and count the number of months that do not meet the conditions.
[0052] 3. Evr function: Determines whether there are months in the last p months where the variable value (i.e., the data value corresponding to each time node in the time series data) is greater than a threshold, reflecting whether the variable has been overactive.
[0053] Calculation method: The result of the Num function is binarized. If there is an active month, the result is 1; otherwise, it is 0.
[0054] 4. Nci function: Calculates the number of months in the last p months where the variable value increases compared to the previous month, reflecting the continuity of the growth trend.
[0055] Calculation method: Calculate the difference between the means of adjacent months and count the number of times the difference is positive.
[0056] 5. Ncd function: Calculates the number of months in the last p months where the variable value decreases compared to the previous month, reflecting the persistence of the downward trend.
[0057] Calculation method: Calculate the difference between the means of adjacent months and count the number of times the difference is negative.
[0058] 6. Ncn function: Calculates the number of months in the last p months where adjacent months have the same variable value, reflecting the stability of the data.
[0059] Calculation method: Calculate the difference between the means of adjacent months and count the number of times the difference is zero.
[0060] 7. Bup function: Determines whether the data for the most recent p months is strictly increasing and greater than a threshold, reflecting a continuous growth trend.
[0061] Calculation method: Check the size relationship of the data corresponding to each time point in adjacent months. If all the data meet the increasing condition, the result is 1; otherwise, the result is 0.
[0062] 8. Pdn function: Determines whether the data for the most recent p months is strictly decreasing and greater than a threshold, reflecting a continuous downward trend.
[0063] Calculation method: Check the size relationship of the data corresponding to each time point in adjacent months. If all data meet the decreasing condition, the result is 1; otherwise, the result is 0.
[0064] 9. Mai function: Calculates the maximum growth rate of a variable between two consecutive months in the most recent p months, reflecting the maximum growth potential.
[0065] Calculation method: Calculate the growth value between all adjacent months and take the maximum value.
[0066] 10. Mad function: Calculates the maximum decrease in a variable between two consecutive months in the most recent p months, reflecting the maximum risk of decline.
[0067] Calculation method: Calculate the decrease between all adjacent months and take the maximum value.
[0068] 11. Msg function: Calculates the number of months since the last time a variable value was greater than a threshold, reflecting the most recent active period.
[0069] Calculation method: Start from the most recent month and search backwards, recording the position of the first month that meets the condition.
[0070] 12. Msz function: Calculates the number of months since the last month in which the variable value was less than or equal to the threshold, reflecting the recent period of inactivity.
[0071] Calculation method: Start from the most recent month and search backwards, recording the position of the first month that does not meet the condition.
[0072] It should be noted that the aforementioned time window means that when performing multidimensional derived feature calculations on time series data corresponding to a preset time period, each feature calculation selects the data from the most recent p months from the time series data as the specific data corresponding to the feature function calculation; where the value of p is less than or equal to the preset time period. For example, if the preset time period is the most recent six months, then the value of p corresponding to the time window can be 3. The specific value of p corresponding to the time window can be flexibly adjusted according to the actual feature calculation needs.
[0073] Here, the aforementioned trend feature dimensions are mainly used to reflect the long-term change patterns of time series data within a preset time period. Specifically, for the aforementioned trend feature dimensions, the application device can calculate multiple trend-derived features of the time series data under the aforementioned trend feature dimensions based on multiple feature functions corresponding to the trend feature dimensions. The multiple trend-derived features are used to reflect the overall change trend of the time series data within the preset time period.
[0074] Specifically, as an optional embodiment, under the aforementioned trend feature dimension, 11 derived features of the time series data under the aforementioned trend feature dimension can be calculated based on the 11 feature functions shown below: 1. Avg function: Calculates the average of the variable values over the most recent p months, reflecting the central trend of the data.
[0075] Calculation method: Calculate the arithmetic mean of the data within a specified time window (i.e., the most recent p months).
[0076] 2. The Tot function: calculates the sum of the variable values over the most recent p months, reflecting the cumulative scale of the data.
[0077] Calculation method: sum the data within a specified time window (i.e., the most recent p months).
[0078] 3. Tot2T function: Calculates the sum of the variable values of the p months excluding the current month in the most recent (p+1) months, reflecting the cumulative scale of the earlier time period.
[0079] Calculation method: Summing the data within a specified earlier time window.
[0080] 4. Max function: Calculates the maximum value of a variable over the most recent p months, reflecting the peak level of the data.
[0081] Calculation method: Find the maximum value within a specified time window (i.e., the most recent p months).
[0082] 5. Min function: Calculates the minimum value of a variable over the most recent p months, reflecting the valley level of the data.
[0083] Calculation method: Find the minimum value within a specified time window (i.e., the most recent p months).
[0084] 6. TRM function: Calculates the trimmed mean of variable values over the most recent p months (the mean after removing the maximum and minimum values), reflecting the robust central tendency of the data.
[0085] Calculation method: After removing the maximum and minimum values within the time window (i.e., the most recent p months), calculate the average of the remaining data.
[0086] 7. Rpp function: Calculates the ratio of the average of the most recent p months to the average of the previous p months, reflecting the relative change between recent and previous levels.
[0087] Calculation method: Calculate the average value of two adjacent time windows, and then find the ratio.
[0088] 8. Dpp function: Calculates the difference between the average of the most recent p months and the average of the previous p months, reflecting the absolute change in recent and previous levels.
[0089] Calculation method: Calculate the average value of two adjacent time windows respectively, and then find the difference.
[0090] 9. Mpp function: Calculates the ratio of the maximum value in the most recent p months to the maximum value in the previous p months, reflecting the relative change between recent and previous peak values.
[0091] Calculation method: Calculate the maximum value of two adjacent time windows respectively, and then find the ratio.
[0092] 10. Npp function: Calculates the ratio of the minimum value in the most recent p months to the minimum value in the previous p months, reflecting the relative change between recent and previous trough values.
[0093] Calculation method: Calculate the minimum value of two adjacent time windows respectively, and then calculate the ratio.
[0094] 11. Msx function: Calculates the number of months from the current month when the maximum value occurred in the most recent p months, reflecting the time position of the peak.
[0095] Calculation method: Find the month where the maximum value is located, and calculate its distance from the current month.
[0096] It should be noted that the aforementioned time window means that when performing multidimensional derived feature calculations on time series data corresponding to a preset time period, each feature calculation selects the data from the most recent p months from the time series data as the specific data corresponding to the feature function calculation; where the value of p is less than or equal to the preset time period. For example, if the preset time period is the most recent six months, then the value of p corresponding to the time window can be 3. The specific value of p corresponding to the time window can be flexibly adjusted according to the actual feature calculation needs.
[0097] Here, the aforementioned fluctuation feature dimension is mainly used to reflect the stability (i.e., fluctuation amplitude) and dispersion of time series data within a preset time period; wherein, for the aforementioned fluctuation feature dimension, the application device can calculate multiple fluctuation-derived features of the time series data under the fluctuation feature dimension according to multiple feature functions corresponding to the fluctuation feature dimension; wherein, the multiple fluctuation-derived features are used to reflect the fluctuation amplitude and dispersion of the time series data within the preset time period.
[0098] Specifically, as an optional embodiment, taking the preset time period of the most recent p months as an example, under the above-mentioned fluctuation characteristic dimension, the 12 derived features of the time series data under the above-mentioned fluctuation characteristic dimension can be calculated according to the 12 feature functions shown below: 1. Std function: Calculates the standard deviation of variable values over the most recent p months, reflecting the dispersion of the data.
[0099] Calculation method: Calculate the standard deviation of data within a specified time window (i.e., the most recent p months).
[0100] 2. Cva function: Calculates the coefficient of variation (the ratio of standard deviation to mean) of variable values over the most recent p months, reflecting the relative dispersion of the data.
[0101] Calculation method: Calculate the ratio of standard deviation to mean to avoid the influence of dimensions.
[0102] 3. Ran function: Calculates the range of variable values over the most recent p months (the difference between the maximum and minimum values), reflecting the fluctuation range of the data.
[0103] Calculation method: Calculate the difference between the maximum and minimum values.
[0104] 4. Cav function: Calculates the ratio of the current month's variable value to the average value of the most recent p months, reflecting the relative relationship between the current level and the average level.
[0105] Calculation method: Divide the monthly value by the average value within the time window (i.e., the most recent p months).
[0106] 5. Cmn function: Calculates the ratio of the current month's variable value to the minimum value of the most recent p months, reflecting the relative relationship between the current level and the lowest level.
[0107] Calculation method: Divide the monthly value by the minimum value within the time window (i.e., the most recent p months).
[0108] 6. Cmx function: Calculates the ratio of the current month's variable value to the maximum value of the most recent p months, reflecting the relative relationship between the current level and the highest level.
[0109] Calculation method: Divide the current month's value by the maximum value within the time window (i.e., the most recent p months).
[0110] 7. Cmm function: Calculates the difference between the current month's variable value and the average value of the most recent p months, reflecting the absolute difference between the current level and the average level.
[0111] Calculation method: Subtract the average value within the time window (i.e., the most recent p months) from the current month's value.
[0112] 8. Cnm function: Calculates the difference between the current month's variable value and the minimum value of the most recent p months, reflecting the absolute difference between the current level and the lowest level.
[0113] Calculation method: Subtract the minimum value within the time window (i.e., the most recent p months) from the current month's value.
[0114] 9. Cxm function: Calculates the difference between the current month's variable value and the maximum value of the most recent p months, reflecting the absolute difference between the current level and the highest level.
[0115] Calculation method: Subtract the maximum value within the time window (i.e., the most recent p months) from the current month's value.
[0116] 10. Cmp function: Calculates the relative difference (difference divided by mean) between the current month's variable value and the average of the most recent p months, reflecting the degree to which the current level deviates from the average level.
[0117] Calculation method: (current month's value - average value) / average value.
[0118] 11. Cnp function: Calculates the relative difference between the current month's variable value and the minimum value of the most recent p months (difference divided by minimum value), reflecting the degree to which the current level deviates from the minimum level.
[0119] Calculation method: (current month's value - minimum value) / minimum value.
[0120] 12. Cxp function: Calculates the relative difference between the current month's variable value and the maximum value of the most recent p months (difference divided by the maximum value), reflecting the degree to which the current level deviates from the highest level.
[0121] Calculation method: (current month value - maximum value) / maximum value.
[0122] It should be noted that the aforementioned time window means that when performing multidimensional derived feature calculations on time series data corresponding to a preset time period, each feature calculation selects the data from the most recent p months from the time series data as the specific data corresponding to the feature function calculation; where the value of p is less than or equal to the preset time period. For example, if the preset time period is the most recent six months, then the value of p corresponding to the time window can be 3. The specific value of p corresponding to the time window can be flexibly adjusted according to the actual feature calculation needs.
[0123] S103, merge the multiple derived features corresponding to each feature dimension of the time series data to obtain the feature matrix corresponding to the user in the target business scenario.
[0124] Here, each element in the feature matrix corresponds to a feature function. For each feature function, simply fill the blank feature matrix with the calculation result (i.e., the derived feature) corresponding to that feature function to obtain the filled feature matrix, which serves as the feature matrix for the user in the target business scenario.
[0125] Specifically, for the target model in the target business scenario (i.e., the model that originally needed to use the user's historical behavior data in the target business scenario as the model input data), as an optional embodiment, the model input data of the target model can be replaced by the aforementioned historical behavior data with the aforementioned feature matrix, and the aforementioned feature matrix is input into the target model to obtain the target prediction result about the user output by the target model. In this case, compared with the aforementioned historical behavior data, since the feature matrix can more comprehensively reflect the characteristics of time series data under different feature dimensions, it is beneficial to provide richer feature input for the target model, thereby improving the accuracy of the prediction result output by the target model.
[0126] It should be noted that the above target model can be determined according to the target business scenario. For example, if the target business scenario is a financial risk control scenario, the above target model can be a risk prediction model used to predict the risk level of credit risk of users in the financial field. The specific model structure of the above target model is not limited in this application embodiment.
[0127] Specifically, when there are multiple users, the application device can process the historical behavior data of multiple users in the target business scenario in parallel according to the methods shown in steps b1-b3 below: Step b1: Input the historical behavior data of multiple users in the target business scenario into the main model, and use the main model to perform parallel prediction processing on the target evaluation results of the multiple users at each time node within the preset time period, and output the target time series data corresponding to the multiple users within the preset time period.
[0128] Here, when performing step b1, the specific processing method for each user's historical behavior data is the same as that in the aforementioned step S101, and the repetitive parts will not be repeated here.
[0129] Step b2: For the preset multiple feature dimensions, based on the multiple feature functions corresponding to each feature dimension, perform parallel calculations on the multiple derived features corresponding to each feature dimension of the multiple target time series data to obtain the multiple derived features corresponding to each feature dimension of each target time series data.
[0130] Here, when performing step b2, the specific processing method for the target time series data of each user is the same as the processing method in the aforementioned step S102, and the repetitions will not be repeated here.
[0131] Step b3: Merge the multiple derived features corresponding to each target time series data under each feature dimension to obtain the target feature matrix corresponding to the multiple users under the target business scenario.
[0132] Here, when performing step b3, the specific merging method for the multiple derived features corresponding to each user is the same as the merging method in the aforementioned step S103, and the repetitions will not be repeated here.
[0133] Specifically, in this application embodiment, as an optional embodiment, in order to improve the efficiency of feature calculation, in the underlying program design of the application device, multi-process parallel computing functionality can also be implemented based on Python's multiprocessing library during the execution of the above steps b1-b3. The core implementation method is as follows: 1. Process Count Configuration: The number of parallel processes can be dynamically determined based on the number of CPU cores in the application device (the default is the number of CPU cores minus 3), and developers can also manually specify the number of processes.
[0134] 2. Task allocation method: A process pool is created using the Pool class to perform parallel computation on 35 feature functions. Specifically, the task can be submitted asynchronously using the apply_async method, with each feature function's feature computation task being treated as an independent computation task.
[0135] 3. The specific calculation process for the underlying part is as follows: The outer layer can loop through all variable prefixes (feature_list), that is, loop through all data in the time series data.
[0136] The middle layer can loop through all time windows (p_list). For example, with the preset time period being the last 6 months, the time window p can be a dynamic sliding window of the last 3 months, the last 4 months, the last 5 months, and the last 6 months.
[0137] The inner layer can perform feature calculation tasks for 35 feature functions in parallel through multiple processes.
[0138] 4. The calculation results of all feature functions can be collected through the results list.
[0139] 5. Exception handling mechanism: Each feature function includes try-except exception handling logic to ensure that the failure of a single derived feature calculation will not affect the overall process. If the feature calculation task of the feature function fails, an empty value (np.nan) is returned.
[0140] 6. Result collection and merging: Use pool.join() to wait for all processes to complete the calculation, then obtain the calculation results of each feature function in turn, and merge them into the final feature matrix.
[0141] Based on the same inventive concept, this application also provides a multi-dimensional feature derivation device based on a master model, which corresponds to the above-mentioned multi-dimensional feature derivation method based on a master model. Since the principle of solving the problem by the multi-dimensional feature derivation device based on a master model in the embodiments of this application is similar to the above-mentioned multi-dimensional feature derivation method based on a master model in the embodiments of this application, the implementation of the multi-dimensional feature derivation device based on a master model can refer to the implementation of the above-mentioned multi-dimensional feature derivation method based on a master model, and the repeated parts will not be described again.
[0142] Reference Figure 2 As shown, Figure 2 The diagram illustrates a structural schematic of a multidimensional feature derivation device based on a master model provided in an embodiment of this application, wherein the multidimensional feature derivation device includes: The data input module 201 is used to input the user's historical behavior data in the target business scenario into the main model, and predict the target evaluation result of the user at each time node within a preset time period through the main model, and output the time series data of the user within the preset time period; wherein, the target evaluation result is determined according to the target business scenario; The feature derivation module 203 is used to calculate, based on multiple feature functions corresponding to each feature dimension, multiple derived features of the time series data under each feature dimension for multiple preset feature dimensions. The result output module 204 is used to merge multiple derived features corresponding to each feature dimension of the time series data to obtain the feature matrix corresponding to the user in the target business scenario.
[0143] In an optional embodiment, the multidimensional feature derivation device further includes: a data preprocessing module 202, wherein the data preprocessing module 202 is used for: In response to the presence of a target value in the time series data, based on the predicted values of the target evaluation results corresponding to other time nodes in the time series data, a target predicted value for the target evaluation result corresponding to the target time node in the time series data is calculated, and the target predicted value is used as the fill value corresponding to the target value to obtain the filled time series data; wherein, the target value is an outlier or missing value in the time series data, the target time node represents the time node corresponding to the target value within the preset time period, and the other time nodes are time nodes other than the target time node within the preset time period; In response to the mismatch between the dimension of the time series data and the target dimension corresponding to the target business scenario, the time series data is mapped to the value range corresponding to the target dimension to obtain time series data that matches the target dimension.
[0144] In an optional implementation, when calculating multiple derived features corresponding to each feature dimension of the time series data based on multiple feature functions corresponding to each feature dimension, the feature derivation module 203 is used to: Based on multiple feature functions corresponding to the momentum feature dimension, multiple momentum-derived features corresponding to the time series data under the momentum feature dimension are calculated respectively; wherein, the multiple momentum-derived features are used to reflect the direction and speed of change of the time series data within the preset time period.
[0145] In an optional implementation, when calculating multiple derived features corresponding to each feature dimension of the time series data based on multiple feature functions corresponding to each feature dimension, the feature derivation module 203 is used to: Based on multiple feature functions corresponding to the trend feature dimension, multiple trend-derived features corresponding to the time series data under the trend feature dimension are calculated respectively; wherein, the multiple trend-derived features are used to reflect the overall change trend of the time series data within the preset time period.
[0146] In an optional implementation, when calculating multiple derived features corresponding to each feature dimension of the time series data based on multiple feature functions corresponding to each feature dimension, the feature derivation module 203 is used to: Based on multiple feature functions corresponding to the fluctuation feature dimension, multiple fluctuation-derived features corresponding to the time series data under the fluctuation feature dimension are calculated respectively; wherein, the multiple fluctuation-derived features are used to reflect the fluctuation amplitude and dispersion of the time series data within the preset time period.
[0147] In an optional implementation, when the number of users of the user is multiple, the multidimensional feature derivation device further includes: a parallel processing module, wherein the parallel processing module is used for: The historical behavior data of multiple users in the target business scenario is input into the main model. The main model performs parallel prediction processing on the target evaluation results of the multiple users at each time node within the preset time period, and outputs the target time series data corresponding to the multiple users within the preset time period. For multiple preset feature dimensions, based on multiple feature functions corresponding to each feature dimension, multiple derived features corresponding to each of the target time series data are calculated in parallel for each feature dimension, so as to obtain multiple derived features corresponding to each of the target time series data for each feature dimension. For each target time series data, multiple derived features corresponding to each feature dimension are merged to obtain the target feature matrix corresponding to each user in the target business scenario.
[0148] In an optional implementation, the multidimensional feature derivation device further includes: a feature application module, wherein the feature application module is used for: The historical behavior data is replaced with the feature matrix in the input data of the target model, and the feature matrix is input into the target model to obtain the target prediction result of the target model about the user; wherein, the target model is determined according to the target business scenario.
[0149] like Figure 3 As shown, this application provides an electronic device 300 for executing the multidimensional feature derivation method based on the master model in this application. The device includes a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302. The memory 301 and the processor 302 are connected via a bus for communication. When the processor 302 executes the computer program, it implements the steps of the multidimensional feature derivation method based on the master model.
[0150] Specifically, the memory 301 and processor 302 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 302 runs the computer program stored in the memory 301, it can execute the above-mentioned multi-dimensional feature derivation method based on the master model.
[0151] Corresponding to the multidimensional feature derivation method based on the master model in this application, this application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the steps of the multidimensional feature derivation method based on the master model described above.
[0152] Specifically, the storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the storage medium is run, it can execute the aforementioned multi-dimensional feature derivation method based on the master model.
[0153] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0154] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0155] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0156] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0157] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0158] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A multidimensional feature derivation method based on a master model, characterized in that, The multidimensional feature derivation method includes: The user's historical behavior data in the target business scenario is input into the main model. The main model then predicts the target evaluation result of the user at each time node within a preset time period, and outputs the time series data of the user within the preset time period. The target evaluation result is determined based on the target business scenario. For multiple preset feature dimensions, based on multiple feature functions corresponding to each feature dimension, multiple derived features corresponding to the time series data under each feature dimension are calculated respectively. The multiple derived features corresponding to each feature dimension of the time series data are merged to obtain the feature matrix corresponding to the user in the target business scenario.
2. The multidimensional feature derivation method according to claim 1, characterized in that, After obtaining the time series data corresponding to the user within the preset time period from the output, the multidimensional feature derivation method further includes preprocessing the time series data using the following method: In response to the presence of a target value in the time series data, based on the predicted values of the target evaluation results corresponding to other time nodes in the time series data, a target predicted value for the target evaluation result corresponding to the target time node in the time series data is calculated, and the target predicted value is used as the fill value corresponding to the target value to obtain the filled time series data; wherein, the target value is an outlier or missing value in the time series data, the target time node represents the time node corresponding to the target value within the preset time period, and the other time nodes are time nodes other than the target time node within the preset time period; In response to the mismatch between the dimension of the time series data and the target dimension corresponding to the target business scenario, the time series data is mapped to the value range corresponding to the target dimension to obtain time series data that matches the target dimension.
3. The multidimensional feature derivation method according to claim 1, characterized in that, The step of calculating multiple derived features of the time series data under each feature dimension based on multiple feature functions corresponding to each feature dimension includes: Based on multiple feature functions corresponding to the momentum feature dimension, multiple momentum-derived features corresponding to the time series data under the momentum feature dimension are calculated respectively; wherein, the multiple momentum-derived features are used to reflect the direction and speed of change of the time series data within the preset time period.
4. The multidimensional feature derivation method according to claim 1, characterized in that, The step of calculating multiple derived features of the time series data under each feature dimension based on multiple feature functions corresponding to each feature dimension includes: Based on multiple feature functions corresponding to the trend feature dimension, multiple trend-derived features corresponding to the time series data under the trend feature dimension are calculated respectively; wherein, the multiple trend-derived features are used to reflect the overall change trend of the time series data within the preset time period.
5. The multidimensional feature derivation method according to claim 1, characterized in that, The step of calculating multiple derived features of the time series data under each feature dimension based on multiple feature functions corresponding to each feature dimension includes: Based on multiple feature functions corresponding to the fluctuation feature dimension, multiple fluctuation-derived features corresponding to the time series data under the fluctuation feature dimension are calculated respectively; wherein, the multiple fluctuation-derived features are used to reflect the fluctuation amplitude and dispersion of the time series data within the preset time period.
6. The multidimensional feature derivation method according to claim 1, characterized in that, When the number of users of the user is multiple, the multidimensional feature derivation method further includes: The historical behavior data of multiple users in the target business scenario is input into the main model. The main model performs parallel prediction processing on the target evaluation results of the multiple users at each time node within the preset time period, and outputs the target time series data corresponding to the multiple users within the preset time period. For multiple preset feature dimensions, based on multiple feature functions corresponding to each feature dimension, multiple derived features corresponding to each of the target time series data are calculated in parallel for each feature dimension, so as to obtain multiple derived features corresponding to each of the target time series data for each feature dimension. For each target time series data, multiple derived features corresponding to each feature dimension are merged to obtain the target feature matrix corresponding to each user in the target business scenario.
7. The multidimensional feature derivation method according to claim 1, characterized in that, The multidimensional feature derivation method also includes: The historical behavior data is replaced with the feature matrix in the input data of the target model, and the feature matrix is input into the target model to obtain the target prediction result of the target model about the user; wherein, the target model is determined according to the target business scenario.
8. A multi-dimensional feature derivation device based on a master model, characterized in that, The multidimensional feature derivation device includes: The data input module is used to input the user's historical behavior data in the target business scenario into the main model, and through the main model, predict the target evaluation result of the user at each time node within a preset time period, and output the time series data of the user within the preset time period; wherein, the target evaluation result is determined according to the target business scenario; The feature derivation module is used to calculate, based on multiple feature functions corresponding to each feature dimension, multiple derived features of the time series data under each feature dimension, for multiple preset feature dimensions. The result output module is used to merge multiple derived features corresponding to each feature dimension of the time series data to obtain the feature matrix corresponding to the user in the target business scenario.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the multi-dimensional feature derivation method based on the master model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the multidimensional feature derivation method based on the master model as described in any one of claims 1 to 7.