General field risk early warning method, system and equipment based on data processing and medium

Through multi-source data integration and model optimization, combined with time series and logistic regression models, the limitations of traditional risk warning systems have been overcome, accurate identification of risks and early warning have been achieved, and the accuracy and comprehensiveness of risk warnings have been improved.

CN120707269APending Publication Date: 2025-09-26INSPUR QILU SOFTWARE IND
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510749770.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional risk warning systems rely on a single data source and static models, which makes it difficult to adapt to the current situation of surging data volume and changing risk forms, resulting in incomplete capture of risk characteristics and inability to achieve efficient risk warning.

Method used

By integrating multi-source data, performing data cleaning and conversion, and combining time series models with logistic regression models, a risk warning model is constructed. Model parameters are optimized through cross-validation, and a visual risk analysis report is generated to achieve an organic combination of dynamic time series analysis and explainable classification.

Benefits of technology

It improves the accuracy and comprehensiveness of risk warnings, enables precise identification and early warning of risks, and ensures the reliability and accuracy of the model in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707269A_ABST
    Figure CN120707269A_ABST
Patent Text Reader

Abstract

The invention discloses a general field risk early warning method, system and equipment based on data processing and a medium, belongs to the technical field of data processing and risk prediction, and aims to solve the technical problem of how to fully mine an association relationship among general field data and improve the accuracy and comprehensiveness of risk early warning. According to the technical scheme, the method comprises the following steps: data collection: collecting multi-source data, and carrying out integration processing on the collected multi-source data to obtain integrated multi-source data; data preprocessing: carrying out preprocessing operation of data cleaning and data conversion on the integrated multi-source data to obtain preprocessed data; constructing a risk early warning model: constructing the risk early warning model by adopting hierarchical design of time sequence feature capture-classification decision based on the time sequence model and the logistic regression model; comprehensive evaluation and verification are carried out on the constructed risk early warning model; and risk early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing and risk prediction, and in particular to a general field risk early warning method, system, device and medium based on data processing. Background Art

[0002] In the digital economy, risk patterns are characterized by accelerated cross-domain transmission and strengthened nonlinear correlations. Traditional early warning systems based on threshold-based assessments have limitations. A single data source results in incomplete capture of risk characteristics, and static models struggle to adapt to a dynamically changing environment. Traditional risk warning methods rely on empirical judgment and simple statistical analysis, making them ill-suited to the current surge in data volumes and the variability of risk patterns.

[0003] Therefore, how to fully explore the correlation between general domain data and improve the accuracy and comprehensiveness of risk warnings is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The technical task of the present invention is to provide a general field risk warning method, system, equipment and medium based on data processing to solve the problem of how to fully explore the correlation between general field data and improve the accuracy and comprehensiveness of risk warning.

[0005] The technical task of the present invention is achieved in the following manner: a general field risk early warning method based on data processing, the method is as follows:

[0006] Data collection: Collect multi-source data, integrate and process the collected multi-source data, and obtain the integrated multi-source data;

[0007] Data preprocessing: Perform data cleaning and data conversion preprocessing operations on the integrated multi-source data to obtain preprocessed data;

[0008] Constructing a risk warning model: Based on a time series model and a logistic regression model, a hierarchical design combining time series feature capture and classification decision-making is employed to construct the risk warning model. This model retains the linear classification advantages of logistic regression while leveraging the time series model to capture the time series characteristics of data, achieving an organic combination of dynamic time series analysis and interpretable classification. The time series model (ARIMA / LSTM) is used to capture time series dynamics (trends, cycles, and abnormal fluctuations) to address the issue of data changing over time. The logistic regression model provides interpretable classification decisions through linear combination and probabilistic output, addressing the question of whether a risk has been triggered.

[0009] Comprehensively evaluate and validate the constructed risk warning model: By dividing the data set into training, validation, and test sets, using cross-validation techniques and combining accuracy, recall, and F1-value evaluation indicators, quantitatively evaluate the performance of the risk warning model from multiple dimensions: prediction accuracy, model generalization ability, and stability. Based on the evaluation results, adjust the risk warning model parameters or optimize the model structure to ensure the reliability of the model in practical applications.

[0010] Risk warning: Based on the risk results calculated by the risk warning model, graded warnings are issued according to the preset risk level thresholds, and a visual risk analysis report is generated to provide decision makers with intuitive risk information so that they can take timely response measures.

[0011] Preferably, the multi-source data includes internal data and external data;

[0012] Internal data comes from data within the internal systems of an enterprise or organization, including cash flow and revenue and profit data generated by the financial system, production progress and inventory turnover data recorded by the operation system, and customer information and transaction records stored in the customer relationship management system.

[0013] External data includes industry reports, social media, and market data, broadening the dimensions of risk analysis.

[0014] As a preference, data cleaning is specifically as follows:

[0015] Noise data processing: Set a reasonable threshold range, treat data outside the range as noise and correct or delete it;

[0016] Data deduplication: Identify and delete duplicate data records through data comparison algorithms;

[0017] Data conversion adjusts the data to the format and distribution characteristics that meet the input requirements of the risk warning model; the details are as follows:

[0018] Data normalization: Use the minimum-maximum normalization formula: x′=max(x)-min(x)x-min(x) to map the data to the [0,1] interval;

[0019] Standardization: Data is standardized using z = σx - μ, where μ is the data mean and σ is the standard deviation, so that the data follows a standard normal distribution with a mean of 0 and a standard deviation of 1. This eliminates dimensional differences between data and improves the training efficiency and prediction accuracy of the risk warning model.

[0020] Coding processing: For categorical variables, coding processing is performed to convert the categorical variables into numerical form to ensure that the risk warning model can be directly processed.

[0021] Preferably, the risk warning model is constructed as follows:

[0022] Time series data input layer: Convert the timestamp to date and time, and obtain a data frame containing the original timestamp, the converted date and time, and the extracted year, month, day, and time features;

[0023] Time series feature extraction layer: This layer generates sine and cosine periodic features from the extracted year-month-day-time features to mine potential information in the data. Specifically, for hourly features (value range: 0-23), the sine and cosine values ​​of the hour are calculated, expanding the one-dimensional hourly features into two-dimensional sine and cosine features. This more comprehensively describes the periodic changes in time and avoids information loss or misleading information caused by simple numerical representation.

[0024] Feature transfer and fusion layer: obtains static subject features and dynamic behavior features, cross-fuses static subject features and dynamic behavior features, and obtains behavior cross-features and spatiotemporal cross-features;

[0025] Logistic regression layer (Dense layer + Sigmoid): Maps input features to output variables through a linear model, and uses the sigmoid function to compress the output value to between [0,1] to obtain a probability value. The probability value is used to determine the probability of the sample belonging to any category.

[0026] More preferably, the behavioral cross-feature refers to the matching degree between the customer’s age and the transaction amount threshold;

[0027] The time-space intersection feature refers to the geographical distance and time difference between the transaction IP address and the customer's permanent residence (remote login + time difference exceeding 12 hours triggers an alert);

[0028] The feature transfer and fusion layer uses dynamic graph neural networks to perform feature fusion to obtain a financial transaction graph. The financial transaction graph consists of nodes and edges. Nodes represent users and accounts in financial transactions, while edges represent transaction relationships and transfer relationships. A message passing mechanism enables nodes to exchange information with each other, thereby updating their feature representations.

[0029] In a financial transaction graph, a user node generates a corresponding message based on its own transaction amount and transaction frequency characteristics, as well as the transaction relationship with other user nodes (such as transaction amount, transaction time interval, etc.); and passes the generated message to the neighboring nodes along the edge. After receiving the message, the neighboring node aggregates the message; the node updates its own feature representation based on the aggregated neighbor information and its own historical status.

[0030] More preferably, the probability value includes feature input, risk assessment, and result output;

[0031] Among them, feature input: the linear model receives the feature data obtained in the previous step;

[0032] Risk assessment: The linear model analyzes each feature and gives a risk probability value.

[0033] Result output: If the predicted financial risk probability value is higher than the set threshold, a risk warning will be triggered.

[0034] Preferably, cross-validation adopts K-fold cross-validation, which randomly divides the data set into K non-overlapping subsets, takes one of the subsets as the validation set each time, and uses the remaining K subsets as the training set. The training and validation process is repeated K times, and finally the results of the K validations are averaged to obtain the performance evaluation indicators of the risk warning model.

[0035] A general field risk early warning system based on data processing, the system comprising:

[0036] A data acquisition unit is used to collect multi-source data, integrate and process the collected multi-source data, and obtain integrated multi-source data;

[0037] A data preprocessing unit is used to perform preprocessing operations such as data cleaning and data conversion on the integrated multi-source data to obtain preprocessed data;

[0038] The risk warning model construction unit is used to build a risk warning model based on a time series model and a logistic regression model, using a hierarchical design that captures time series features and makes classification decisions. This model retains the linear classification advantages of logistic regression while leveraging the time series model to capture the time series features of data, achieving an organic combination of dynamic time series analysis and interpretable classification. The time series model (ARIMA / LSTM) is used to capture time series dynamics (trends, cycles, and abnormal fluctuations) to address the issue of data changing over time. The logistic regression model provides interpretable classification decisions through linear combination and probabilistic output, addressing the question of whether risks are triggered.

[0039] The evaluation and validation unit is used to quantitatively evaluate the performance of the risk warning model from multiple dimensions, including prediction accuracy, model generalization ability, and stability, by dividing the data set into training, validation, and test sets, using cross-validation techniques and combining accuracy, recall, and F1-value evaluation indicators. The unit then adjusts the risk warning model parameters or optimizes the model structure based on the evaluation results to ensure the reliability of the model in practical applications.

[0040] The risk warning unit is used to provide graded warnings based on the risk results calculated by the risk warning model according to the preset risk level thresholds, and at the same time generate a visual risk analysis report to provide decision makers with intuitive risk information so that they can take timely response measures.

[0041] An electronic device comprising: a memory and at least one processor;

[0042] Wherein, the memory stores a computer program;

[0043] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the general field risk early warning method based on data processing as described above.

[0044] A computer-readable storage medium stores a computer program, which can be executed by a processor to implement the general field risk warning method based on data processing as described above.

[0045] The general field risk early warning method, system, device, and medium based on data processing of the present invention have the following advantages:

[0046] (1) This invention achieves accurate risk identification and early warning through efficient processing of multi-source data. Through key technologies such as data collection, cleaning, and analysis, it integrates internal and external data, solves the problem of multi-source data integration, and constructs a risk warning model system that covers traditional and modern data-driven models. It clarifies the applicable scenarios and construction steps of each model; establishes an evaluation indicator system that includes accuracy, timeliness, and reliability, and proposes optimization strategies from the data, model, and process levels based on the evaluation results, providing scientific and universal risk warnings for risk management in various fields;

[0047] (2) The present invention integrates data from different sources, fully explores the correlation between data, improves the accuracy and comprehensiveness of risk warnings, constructs a risk warning model based on deep learning, automatically learns the characteristics and patterns in the data, and achieves accurate prediction of risks;

[0048] (3) The cross-validation of the present invention can make full use of the data set, reduce the random influence caused by sample division, and more accurately evaluate the generalization ability of the model. It can also be verified using an independent test set. The trained model is applied to the test set, and the various evaluation indicators of the model on the test set are calculated. If the indicators perform well, it means that the model has good generalization ability and can accurately warn of risks in practical applications; if the indicators are not ideal, the model needs to be adjusted and optimized, such as reselecting the model, adjusting model parameters, increasing the amount of data, or performing feature engineering, until the performance of the model on the test set meets the requirements; through strict model evaluation and verification, it can ensure that the risk warning model plays an effective warning role in practical applications and provide reliable support for risk management;

[0049] (4) The present invention sets reasonable warning thresholds based on historical data and business needs. In credit risk warning, different thresholds are set according to different credit rating standards and risk tolerance. When the warning probability predicted by the model exceeds the threshold, a risk warning of the corresponding level is issued. The constructed risk warning model is used to process and analyze real-time data. Once it is found that the risk indicator exceeds the warning threshold, warning information is promptly issued to relevant personnel.

[0050] (5) The present invention realizes the efficient integration of dynamic and static features of time series: the time series model is responsible for capturing the temporal dependency of data, and the logistic regression layer is responsible for linear classification decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The present invention will be further described below with reference to the accompanying drawings.

[0052] Attachment Figure 1 This is a flowchart of a general field risk warning method based on data processing. DETAILED DESCRIPTION

[0053] The general field risk warning method, system, device and medium based on data processing of the present invention are described in detail below with reference to the drawings and specific embodiments of the specification.

[0054] Example 1:

[0055] As attached Figure 1 As shown, this embodiment provides a general field risk warning method based on data processing, the method is as follows:

[0056] S1. Data collection: Collect multi-source data, integrate and process the collected multi-source data, and obtain the integrated multi-source data;

[0057] S2. Data preprocessing: Perform data cleaning and data conversion preprocessing operations on the integrated multi-source data to obtain preprocessed data;

[0058] S3. Constructing a risk warning model: Based on a time series model and a logistic regression model, a hierarchical design combining time series feature capture and classification decision-making is used to construct a risk warning model. This model retains the linear classification advantage of logistic regression while leveraging the time series model to capture the time series characteristics of data, achieving an organic combination of dynamic time series analysis and interpretable classification. The time series model (ARIMA / LSTM) is used to capture time series dynamics (trends, cycles, and abnormal fluctuations) to address the issue of data changing over time. The logistic regression model provides interpretable classification decisions through linear combination and probabilistic output, addressing the question of whether a risk is triggered.

[0059] S4. Conduct a comprehensive evaluation and validation of the constructed risk warning model: By dividing the data set into training, validation, and test sets, using cross-validation techniques and combining accuracy, recall, and F1-value evaluation indicators, quantitatively evaluate the performance of the risk warning model from multiple dimensions, including prediction accuracy, model generalization ability, and stability. Based on the evaluation results, adjust the risk warning model parameters or optimize the model structure to ensure the reliability of the model in practical applications.

[0060] S5. Risk Warning: Based on the risk results calculated by the risk warning model, graded warnings are issued according to the preset risk level thresholds, and a visual risk analysis report is generated to provide decision makers with intuitive risk information so that they can take timely response measures.

[0061] The multi-source data in this embodiment includes internal data and external data;

[0062] Internal data comes from data within the internal systems of an enterprise or organization, including cash flow and revenue and profit data generated by the financial system, production progress and inventory turnover data recorded by the operation system, and customer information and transaction records stored in the customer relationship management system.

[0063] External data includes industry reports, social media, and market data, broadening the dimensions of risk analysis.

[0064] This embodiment controls data quality as follows:

[0065] Completeness: The collected data should cover all relevant aspects and avoid data omissions. For example, when building an internet credit risk early warning model, if certain key customer information, such as income and credit history, is missing, it may lead to inaccurate credit risk assessment.

[0066] Accuracy: Data must be accurate; otherwise, analysis based on erroneous data will lead to erroneous conclusions. For financial data, even a single decimal point error can lead to serious misjudgments of risk. During the data collection process, data must be verified through various methods, such as comparison with authoritative data sources.

[0067] Consistency: Data from different sources should be consistent in terms of format and definition. For example, when integrating enterprise operational data from different departments, the definition and calculation method of the same indicator should be unified, otherwise it will cause difficulties in subsequent data processing.

[0068] The data cleaning in step S2 of this embodiment is specifically as follows:

[0069] ① Noise data processing: set a reasonable threshold range, treat the data outside the range as noise and correct or delete it;

[0070] ②Data deduplication: Identify and delete duplicate data records through data comparison algorithms.

[0071] The data conversion in step S2 of this embodiment adjusts the data to a format and distribution characteristics that meet the input requirements of the risk warning model; specifically, as follows:

[0072] ① Data normalization: Use the minimum-maximum normalization formula: x′=max(x)-min(x)x-min(x) to map the data to the [0,1] interval;

[0073] ② Standardization: Data is standardized through z = σx - μ, where μ is the data mean and σ is the standard deviation, so that the data follows a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby eliminating dimensional differences between data and improving the training efficiency and prediction accuracy of the risk warning model;

[0074] ③ Coding processing: For categorical variables, coding processing is performed to convert the categorical variables into numerical form to ensure that the risk warning model can be directly processed.

[0075] The specific steps of constructing the risk warning model in step S3 of this embodiment are as follows:

[0076] S301, Time Series Data Input Layer: Convert the timestamp to date and time, and obtain a data frame containing the original timestamp, the converted date and time, and the extracted year, month, day, and time features; the key code is as follows:

[0077] #Convert timestamp to datetime format

[0078] df['datetime']=pd.to datetime(df['timestamp'],unit='ms');

[0079] #Extract year, month, day and time features;

[0080] df['year']=df['datetime'].dt.year;

[0081] df['month']=df['datetime'].dt.month;

[0082] df['day']=df['datetime'].dt.day;

[0083] df['hour']=df['datetime'].dt.hour;

[0084] print(df[['timestamp',"datetime',','month','day','hour']]).

[0085] S302, Time Series Feature Extraction Layer: Generate sine and cosine periodic features from the extracted year-month-day-hour features to mine potential information in the data. Specifically, for the hour feature (value range is 0-23), calculate the sine and cosine values ​​of the hour, expand the one-dimensional hour feature into two-dimensional sine and cosine features, and more comprehensively describe the periodic changes of time, avoiding information loss or misleading caused by simple numerical representation. The key code is as follows:

[0086] import numpy as np;

[0087] #Assume that df is a data frame containing hourly features;

[0088] df=pd.DataFrame({'hour':[10,15]});

[0089] #Calculate sine and cosine periodic characteristics;

[0090] df['hour_sin']=np.sin(2*np.pi*df['hour'] / 23);

[0091] df['hour_cos']=np.cos(2*np.pi*df['hour'] / 23);

[0092] print(df[['hour','hour_sin','hour_cos']]);

[0093] S303, Feature Transfer and Fusion Layer: Obtain static subject features and dynamic behavior features, cross-fuse the static subject features and dynamic behavior features to obtain behavioral cross-features and spatiotemporal cross-features; wherein, static subject features include customer attributes, account attributes, and relationship networks; customer attributes include age, occupation, income, and credit score (reflecting long-term risk propensity); account attributes include account opening time, account type (savings card / credit card), and historical freeze records (reflecting account stability); relationship networks include emergency contact relevance and corporate equity structure; dynamic behavior features include transaction behavior, device features, and operation traces; transaction behavior includes transaction time (whether it is non-working hours), transaction frequency (high frequency transfers in a short period of time), and counterparty risk score; device features include device ID, IP address, and MAC address (to identify device anomalies); operation traces include login time, page jump path, and number of input errors (to determine whether it is an automated script operation);

[0094] S304, Logistic Regression Layer (Dense Layer + Sigmoid): Maps input features to output variables through a linear model, and uses the sigmoid function to compress the output value to between [0, 1] to obtain a probability value. The probability value is used to determine the probability of the sample belonging to any category.

[0095] The behavioral cross-feature in this embodiment refers to the matching degree between the customer age and the transaction amount threshold;

[0096] The time-space intersection feature refers to the geographical distance and time difference between the transaction IP address and the customer's permanent residence (remote login + time difference exceeding 12 hours triggers an alert);

[0097] The feature transfer and fusion layer uses dynamic graph neural networks to perform feature fusion to obtain a financial transaction graph. The financial transaction graph consists of nodes and edges. Nodes represent users and accounts in financial transactions, while edges represent transaction relationships and transfer relationships. A message passing mechanism enables nodes to exchange information with each other, thereby updating their feature representations.

[0098] In a financial transaction graph, a user node generates a corresponding message based on its own transaction amount and transaction frequency characteristics, as well as the transaction relationship with other user nodes (such as transaction amount, transaction time interval, etc.); and passes the generated message to the neighboring nodes along the edge. After receiving the message, the neighboring node aggregates the message; the node updates its own feature representation based on the aggregated neighbor information and its own historical status.

[0099] The probability value in this embodiment includes feature input, risk assessment and result output;

[0100] Among them, feature input: the linear model receives the feature data obtained in the previous step;

[0101] Risk assessment: The linear model analyzes each feature and gives a risk probability value.

[0102] Result output: If the predicted financial risk probability value is higher than the set threshold, a risk warning will be triggered.

[0103] The accuracy rate in this example measures the proportion of samples correctly predicted by the model to the total number of samples. In the internet credit risk early warning model, the accuracy rate reflects the accuracy of the model's credit risk assessment. For example, if the model correctly predicts the risk of 100 customers for 85 of them, the accuracy rate is 85%.

[0104] The recall rate in this embodiment refers to the proportion of actual positive examples that are correctly predicted as positive by the model. For risk warning models, a high recall rate means that they can detect as many potential risks as possible and avoid missing risks.

[0105] The F1 value in this embodiment comprehensively considers the precision and recall rates, and is the harmonic mean of the two, which can more comprehensively evaluate the model performance.

[0106] The cross-validation in this embodiment adopts K-fold cross-validation, which randomly divides the data set into K mutually non-overlapping subsets, takes one of the subsets as the validation set each time, and uses the remaining K subsets as the training set. The training and validation process is repeated K times, and finally the results of the K validations are averaged to obtain the performance evaluation index of the risk warning model.

[0107] Example 2:

[0108] This embodiment provides a general domain risk early warning system based on data processing, which includes:

[0109] A data acquisition unit is used to collect multi-source data, integrate and process the collected multi-source data, and obtain integrated multi-source data;

[0110] A data preprocessing unit is used to perform preprocessing operations such as data cleaning and data conversion on the integrated multi-source data to obtain preprocessed data;

[0111] The risk warning model construction unit is used to build a risk warning model based on a time series model and a logistic regression model, using a hierarchical design that captures time series features and makes classification decisions. This model retains the linear classification advantages of logistic regression while leveraging the time series model to capture the time series features of data, achieving an organic combination of dynamic time series analysis and interpretable classification. The time series model (ARIMA / LSTM) is used to capture time series dynamics (trends, cycles, and abnormal fluctuations) to address the issue of data changing over time. The logistic regression model provides interpretable classification decisions through linear combination and probabilistic output, addressing the question of whether risks are triggered.

[0112] The evaluation and validation unit is used to quantitatively evaluate the performance of the risk warning model from multiple dimensions, including prediction accuracy, model generalization ability, and stability, by dividing the data set into training, validation, and test sets, using cross-validation techniques and combining accuracy, recall, and F1-value evaluation indicators. The unit then adjusts the risk warning model parameters or optimizes the model structure based on the evaluation results to ensure the reliability of the model in practical applications.

[0113] The risk warning unit is used to provide graded warnings based on the risk results calculated by the risk warning model according to the preset risk level thresholds, and at the same time generate a visual risk analysis report to provide decision makers with intuitive risk information so that they can take timely response measures.

[0114] Example 3:

[0115] This embodiment also provides an electronic device, including: a memory and a processor;

[0116] wherein the memory stores computer-executable instructions;

[0117] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the general field risk early warning method based on data processing in any embodiment of the present invention.

[0118] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0119] The memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can also include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state memory devices.

[0120] Example 4:

[0121] This embodiment further provides a computer-readable storage medium storing a plurality of instructions, which are loaded by a processor to cause the processor to execute the general field risk warning method based on data processing according to any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, wherein the storage medium stores software program code that implements the functions of any of the above-described embodiments, and a computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.

[0122] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.

[0123] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RYMs, DVD-RWs, DVD+RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer via a communications network.

[0124] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.

[0125] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A general field risk early warning method based on data processing, characterized in that: The method is as follows: Data collection: Collect multi-source data, integrate and process the collected multi-source data, and obtain the integrated multi-source data; Data preprocessing: Perform data cleaning and data conversion preprocessing operations on the integrated multi-source data to obtain preprocessed data; Constructing a risk warning model: Based on a time series model and a logistic regression model, a hierarchical design combining time series feature capture and classification decision-making is used to construct the risk warning model. This model retains the linear classification advantage of logistic regression while leveraging the time series model to capture the temporal characteristics of data, achieving an organic combination of dynamic time series analysis and interpretable classification. The time series model is used to capture temporal dynamics, while the logistic regression model provides interpretable classification decisions through linear combination and probabilistic output, addressing the question of whether a risk has been triggered. Comprehensively evaluate and validate the constructed risk warning model: By dividing the data set into training, validation, and test sets, using cross-validation techniques and combining accuracy, recall, and F1-value evaluation indicators, quantitatively evaluate the performance of the risk warning model from multiple dimensions, including prediction accuracy, model generalization ability, and stability. Adjust the risk warning model parameters or optimize the model structure based on the evaluation results. Risk warning: Based on the risk results calculated by the risk warning model, graded warnings are issued according to the preset risk level thresholds, and a visual risk analysis report is generated to provide decision makers with intuitive risk information so that they can take timely response measures.

2. The general field risk early warning method based on data processing according to claim 1 is characterized in that: Multi-source data includes internal data and external data; Internal data comes from data within the internal systems of an enterprise or organization, including cash flow and revenue and profit data generated by the financial system, production progress and inventory turnover data recorded by the operation system, and customer information and transaction records stored in the customer relationship management system. External data includes industry reports, social media, and market data, broadening the dimensions of risk analysis.

3. The general field risk early warning method based on data processing according to claim 1 is characterized in that: The data cleaning process is as follows: Noise data processing: Set a reasonable threshold range, treat data outside the range as noise and correct or delete it; Data deduplication: Identify and delete duplicate data records through data comparison algorithms; Data conversion adjusts the data to the format and distribution characteristics that meet the input requirements of the risk warning model; the details are as follows: Data normalization: Use the minimum-maximum normalization formula: x′=max(x)-min(x)x-min(x) to map the data to the [0,1] interval; Standardization: Data is standardized using z = σx - μ, where μ is the data mean and σ is the standard deviation, so that the data follows a standard normal distribution with a mean of 0 and a standard deviation of 1. This eliminates dimensional differences between data and improves the training efficiency and prediction accuracy of the risk warning model. Coding processing: For categorical variables, coding processing is performed to convert the categorical variables into numerical form to ensure that the risk warning model can be directly processed.

4. The general field risk early warning method based on data processing according to any one of claims 1 to 3, characterized in that: The specific steps for building a risk warning model are as follows: Time series data input layer: Convert the timestamp to date and time, and obtain a data frame containing the original timestamp, the converted date and time, and the extracted year, month, day, and time features; Time series feature extraction layer: Generates sine and cosine periodic features from the extracted year-month-day-time features to mine potential information in the data. Specifically, for hourly features, the sine and cosine values ​​of the hour are calculated, expanding the one-dimensional hourly features into two-dimensional sine and cosine features to more comprehensively describe the periodic changes of time. Feature transfer and fusion layer: obtains static subject features and dynamic behavior features, cross-fuses static subject features and dynamic behavior features, and obtains behavior cross-features and spatiotemporal cross-features; Logistic regression layer: The input features are mapped to the output variables through a linear model, and the output values ​​are compressed to [0, 1] with the help of the sigmoid function to obtain a probability value. The probability value is used to determine the probability that the sample belongs to any category.

5. The general field risk early warning method based on data processing according to claim 4 is characterized in that: Behavioral cross-features refer to the matching degree between customer age and transaction amount threshold; The time-space intersection feature refers to the geographical distance and time difference between the transaction IP address and the customer's permanent residence; The feature transfer and fusion layer uses dynamic graph neural networks to perform feature fusion to obtain a financial transaction graph. The financial transaction graph consists of nodes and edges. Nodes represent users and accounts in financial transactions, while edges represent transaction relationships and transfer relationships. A message passing mechanism enables nodes to exchange information with each other, thereby updating their feature representations. In the financial transaction graph, a user node generates corresponding messages based on its own transaction amount and transaction frequency characteristics, as well as the transaction relationship with other user nodes; The generated message is passed along the edge to the neighboring node. After receiving the message, the neighboring node aggregates the message. The node updates its feature representation based on the aggregated neighbor information and its own historical status.

6. The general field risk early warning method based on data processing according to claim 4 is characterized in that: The probability value includes feature input, risk assessment and result output; Among them, feature input: the linear model receives the feature data obtained in the previous step; Risk assessment: The linear model analyzes each feature and gives a risk probability value. Result output: If the predicted financial risk probability value is higher than the set threshold, a risk warning will be triggered.

7. The general field risk early warning method based on data processing according to claim 1 is characterized in that: Cross-validation uses K-fold cross-validation, which randomly divides the data set into K mutually non-overlapping subsets. Each time, one of the subsets is taken as the validation set, and the remaining K subsets are used as the training set. The training and validation process is repeated K times, and finally the results of the K validations are averaged to obtain the performance evaluation indicators of the risk warning model.

8. A general field risk early warning system based on data processing, characterized in that: The system includes: A data acquisition unit is used to collect multi-source data, integrate and process the collected multi-source data, and obtain integrated multi-source data; A data preprocessing unit is used to perform preprocessing operations such as data cleaning and data conversion on the integrated multi-source data to obtain preprocessed data; The risk warning model construction unit is used to construct a risk warning model based on a time series model and a logistic regression model, using a hierarchical design of capturing time series features and making classification decisions. This model retains the linear classification advantages of logistic regression while utilizing the time series model to capture the time series features of data, achieving an organic combination of dynamic time series analysis and interpretable classification. The time series model is used to capture time series dynamics, while the logistic regression model provides interpretable classification decisions through linear combination and probabilistic output, addressing the question of whether a risk is triggered. The evaluation and validation unit is used to quantitatively evaluate the performance of the risk warning model from multiple dimensions, including prediction accuracy, model generalization ability, and stability, by dividing the data set into training, validation, and test sets, using cross-validation techniques and combining accuracy, recall, and F1-value evaluation indicators. The risk warning model parameters are adjusted or the model structure is optimized based on the evaluation results. The risk warning unit is used to provide graded warnings based on the risk results calculated by the risk warning model according to the preset risk level thresholds, and at the same time generate a visual risk analysis report to provide decision makers with intuitive risk information so that they can take timely response measures.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the general field risk early warning method based on data processing as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the general field risk early warning method based on data processing as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Risk prediction method based on mobile terminal equipment

    CN121052829A