Risk prediction method and device, equipment, storage medium and program product

Through the risk prediction model and behavior prediction model trained by federated learning, combined with static and dynamic data, the problem of low accuracy of risk prediction in the existing technology is solved, and accurate risk assessment of user behavior and data utilization is achieved.

CN120297982APending Publication Date: 2025-07-11INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510441013.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

Smart Images

  • Figure CN120297982A_ABST
    Figure CN120297982A_ABST
Patent Text Reader

Abstract

The invention provides a risk prediction method which can be applied to the field of artificial intelligence. The risk prediction method comprises the steps of obtaining static attribute data of a target object and dynamic behavior data generated when the target object performs business operation; inputting the static attribute data into a risk prediction model based on federated learning training of a business system and an anti-fraud system, and outputting a transaction risk probability; performing step-by-step prediction on a plurality of sub-behavior data in the dynamic behavior data by using a behavior prediction model obtained by training according to historical behavior data of the target object to generate predicted behavior data; and determining the target risk probability of the target object according to the deviation degree between the predicted behavior data and the dynamic behavior data and the transaction risk probability. The invention further provides a risk prediction device and equipment, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a risk prediction method, apparatus, device, medium, and program product. Background Art

[0002] In related technologies, a static defense mechanism based on preset rules or a supervised learning model is usually adopted to analyze the behavioral characteristics of users, and then evaluate the risk of users being defrauded.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in related technologies. When dealing with unstructured user behavior data, it is difficult to effectively capture potential risk behavior characteristics only relying on a static defense mechanism or a supervised learning model, resulting in low prediction accuracy. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides a risk prediction method, apparatus, device, medium, and program product.

[0005] According to a first aspect of the present disclosure, a risk prediction method is provided, including: obtaining static attribute data of a target object and dynamic behavior data generated when the target object performs a business operation; inputting the static attribute data into a risk prediction model trained by federated learning of a business system and an anti-fraud system, and outputting a transaction risk probability; using a behavior prediction model trained according to the historical behavior data of the target object to gradually predict multiple sub-behavior data in the dynamic behavior data, and generating predicted behavior data; determining a target risk probability of the target object according to the deviation degree between the predicted behavior data and the dynamic behavior data, and the transaction risk probability.

[0006] According to an embodiment of the present disclosure, the above using a behavior prediction model trained according to the historical behavior data of the above target object to gradually predict multiple sub-behavior data in the above dynamic behavior data, and generating predicted behavior data, includes: sorting the multiple sub-behavior data in the above dynamic behavior data according to a time stamp to obtain behavior sorted data; performing feature extraction and vector conversion on each sub-behavior data in the above behavior sorted data to generate a dynamic behavior sequence composed of multiple feature vectors; inputting the above dynamic behavior sequence into the above behavior prediction model, and outputting a predicted behavior sequence, where each prediction vector in the above predicted behavior sequence corresponds one-to-one to the feature vector in the above dynamic behavior sequence; converting the multiple prediction vectors in the above predicted behavior sequence into a preset format corresponding to the above dynamic behavior data to obtain the above predicted behavior data.

[0007] According to an embodiment of the present disclosure, the method further includes: aligning the feature vectors in the above dynamic behavior sequence with the prediction vectors in the above prediction behavior sequence in the time dimension; calculating the distance between each aligned feature vector and the corresponding prediction vector; taking the average of the distances between all the above feature vectors and the above prediction vectors to obtain the above deviation degree.

[0008] According to an embodiment of the present disclosure, determining the target risk probability of the above target object based on the deviation degree between the above prediction behavior data and the above dynamic behavior data, and the above transaction risk probability includes: performing normalization processing on the above deviation degree to obtain an intermediate value; performing probability conversion on the above intermediate value by using an activation function to obtain a deviation risk probability corresponding to the above deviation degree; based on a first weight corresponding to the above deviation risk probability and a second weight corresponding to the above transaction risk probability, performing weighted summation on the above deviation risk probability and the above transaction risk probability to obtain the above target risk probability.

[0009] According to an embodiment of the present disclosure, the above method further includes: obtaining the number of risk events occurring within a predetermined time period from the above anti-fraud system; in the case where the above number of events is greater than a preset number, performing normalization processing on the above number of events, and multiplying the normalized number of events by the above intermediate value to obtain a risk reference value; determining a first weight corresponding to the above risk reference value according to a preset mapping relationship between a preset weight and the risk reference value; determining a second weight corresponding to the above first weight according to a weight complementary relationship.

[0010] According to an embodiment of the present disclosure, the above behavior prediction model is trained in the following manner: converting the historical behavior data of the above target object into a time series format; inputting the converted historical behavior data into an initial prediction model, and calculating a prediction result through forward propagation, where the above initial prediction model includes a long short-term memory network layer, and the long short-term memory network layer is used to capture long-term dependencies in the above historical behavior data; comparing the above prediction result with the real behavior data in the above historical behavior data, and calculating a model loss value by using a loss function; based on the above model loss value, updating model parameters by using a backpropagation algorithm, and iteratively training until the model converges to obtain the above behavior prediction model.

[0011] According to an embodiment of the present disclosure, the conversion of the historical behavior data of the target object into a time series format includes: cleaning the historical behavior data to remove redundant information to obtain intermediate behavior data; constructing statistical features from multiple dimensions based on the intermediate behavior data, where the multiple dimensions include a time dimension, a numerical dimension, and a behavior pattern dimension; extracting target features with a correlation degree reaching a preset threshold with a preset variable from the intermediate behavior data through correlation analysis; using a random forest algorithm to extract key features from the intermediate behavior data through a feature importance evaluation mechanism; and converting the statistical features, the target features, and the key features into the time series format.

[0012] A second aspect of the present disclosure provides a risk prediction device, including: a data acquisition module for acquiring static attribute data of a target object and dynamic behavior data generated when the target object conducts business operations; a transaction prediction module for inputting the static attribute data into a risk prediction model trained by federated learning based on a business system and an anti-fraud system, and outputting a transaction risk probability; a behavior prediction module for using a behavior prediction model trained according to the historical behavior data of the target object to gradually predict multiple sub-behavior data in the dynamic behavior data to generate predicted behavior data; and a risk prediction module for determining a target risk probability of the target object according to the deviation between the predicted behavior data and the dynamic behavior data, and the transaction risk probability.

[0013] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, where the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0014] A fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instruction stored thereon, and the computer program or instruction, when executed by a processor, implements the steps of the above method.

[0015] A fifth aspect of the present disclosure further provides a computer program product including a computer program or instruction, and the computer program or instruction, when executed by a processor, implements the steps of the above method.

[0016] According to an embodiment of the present disclosure, by obtaining the static attribute data and dynamic behavior data of a target object, the target object is characterized from multiple perspectives. Using a risk prediction model trained by federated learning based on a business system and an anti-fraud system for risk prediction realizes the effective integration of cross-system data. Using a behavior prediction model to gradually predict the dynamic behavior data to generate predicted behavior data and calculating the deviation degree between it and the actual dynamic behavior data can effectively capture the abnormal behavior of the target object during the operation process. And since the anti-fraud system cannot obtain specific behavior data, an independent behavior prediction model is used to analyze it instead of incorporating it into the federated learning system. This separation processing method can effectively improve the data utilization rate. Determining the target risk probability of the target object according to the deviation degree and the transaction risk probability realizes the comprehensive consideration of static and dynamic risk factors. Even when dealing with unstructured data, the above method can significantly improve the accuracy of the target risk probability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0018] Figure 1 Schematically shows an application scenario diagram of a risk prediction method, device, equipment, medium, and program product according to an embodiment of the present disclosure;

[0019] Figure 2 Schematically shows a flowchart of a risk prediction method according to an embodiment of the present disclosure;

[0020] Figure 3 Schematically shows a schematic diagram of vector alignment in a risk prediction method according to an embodiment of the present disclosure;

[0021] Figure 4 Schematically shows a flowchart of a risk prediction method according to another embodiment of the present disclosure;

[0022] Figure 5 Schematically shows a structural block diagram of a risk prediction device according to an embodiment of the present disclosure; and

[0023] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing a risk prediction method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0025] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0027] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0028] In the technical solution of the present disclosure, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.

[0029] Embodiments of the present disclosure provide a risk prediction method, which includes: obtaining static attribute data of a target object and dynamic behavior data generated when the target object performs a business operation; inputting the static attribute data into a risk prediction model trained by federated learning based on a business system and an anti-fraud system, and outputting a transaction risk probability; using a behavior prediction model trained according to the historical behavior data of the target object to gradually predict multiple sub-behavior data in the dynamic behavior data to generate predicted behavior data; and determining the target risk probability of the target object according to the deviation between the predicted behavior data and the dynamic behavior data and the transaction risk probability.

[0030] Figure 1 Schematically shows an application scenario diagram of a risk prediction method, device, device, medium and program product according to an embodiment of the present disclosure.

[0031] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or fiber optic cables, etc.

[0032] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0034] The server 105 may be a server that provides various services, such as a background management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only for example). The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0035] It should be noted that the risk prediction method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the risk prediction device provided by the embodiments of the present disclosure can generally be set in the server 105. The risk prediction method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the risk prediction device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0036] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0037] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 6 scenario described below, the risk prediction method of the disclosure embodiments will be described in detail through

[0038] Figure 2 FIG. schematically shows a flowchart of the risk prediction method according to an embodiment of the present disclosure.

[0039] As Figure 2 shown, this embodiment includes operation S210 to operation S240.

[0040] In operation S210, static attribute data of the target object and dynamic behavior data generated when the target object performs a business operation are obtained.

[0041] In operation S220, the static attribute data is input into a risk prediction model trained by federated learning based on a business system and an anti-fraud system, and a transaction risk probability is output.

[0042] In operation S230, a behavior prediction model trained according to the historical behavior data of the target object is used to gradually predict multiple sub-behavior data in the dynamic behavior data to generate predicted behavior data.

[0043] In operation S240, according to the deviation degree between the predicted behavior data and the dynamic behavior data, and the transaction risk probability, the target risk probability of the target object is determined.

[0044] According to an embodiment of the present disclosure, when a user conducts business operations through mobile banking, the system will record their dynamic behavior data, such as click frequency, dwell time, and the order of accessed modules. By comprehensively analyzing the user's static attribute data and dynamic behavior data, it is possible to evaluate whether there are risks during the user's business operations.

[0045] According to an embodiment of the present disclosure, the static attribute data mainly includes the following types of information: User personal information, including financial data such as asset allocation and deposit size, which can be further broken down into specific proportions such as deposits, investments, and insurance. Device information, covering technical parameters such as the type of device used by the user and the operating system version. Time information, recording the user's login time and operation time, which can be further broken down into time periods such as weekdays / weekends, day / night, etc. And geographical location information, used to track the geographical location of the user's login and operation, which can identify abnormal geographical location changes. These data together form the basis of the user portrait, providing an important basis for risk identification and behavior analysis.

[0046] In an embodiment of the present disclosure, before obtaining the static attribute data and dynamic behavior data, the consent or authorization of the user can be obtained. For example, before operation S210, a request to obtain data can be sent to the target object. In the case where the target object consents or authorizes the acquisition of data, operation S210 is executed.

[0047] According to an embodiment of the present disclosure, based on the obtained static attribute data and the data provided by the anti-fraud system, both parties carry out federated learning cooperation. The specific implementation process is as follows: First, according to the actual needs of the anti-fraud business scenario, negotiate with the anti-fraud system on the types and scopes of user data required. Second, jointly complete the design and optimization of the federated learning model. Finally, cooperate with the deployment of the central server to ensure data security and the effectiveness of model training.

[0048] According to an embodiment of the present disclosure, the mobile banking constructs a first-party model of the user portrait based on the collected static attribute data according to the requirements of federated learning. The anti-fraud system, on the other hand, uses the data obtained from the user's device to establish a second-party model for device risk identification. After completing the model construction, both parties upload the key parameters of the first-party model and the second-party model to the federated learning central server, providing a data basis for subsequent joint modeling and secure computing.

[0049] According to an embodiment of the present disclosure, the central server integrates the parameters of the first-party model of the mobile banking and the second-party model of the anti-fraud system through a federated learning framework, adopts a distributed learning method, and performs weighted fusion on the data of both parties according to a predetermined protocol to construct a user behavior feature model under abnormal conditions. Through multiple rounds of iterative optimization, the system gradually converges to a stable global gradient and finally forms an accurate risk prediction model. By inputting the static attribute data of the user into this model, a quantified transaction risk probability can be output, providing a decision-making basis for risk prevention and control.

[0050] According to an embodiment of the present disclosure, the dynamically collected behavioral data is input into a behavior prediction model trained based on the historical behavioral data of the target object, and each sub-behavioral data is predicted one by one to generate predicted behavioral data. Subsequently, the system compares and analyzes the predicted behavioral data with the actual dynamic behavioral data to calculate the deviation degree. Finally, by fusing the deviation degree and the transaction risk probability, the target risk probability of the target object is comprehensively evaluated.

[0051] According to an embodiment of the present disclosure, by obtaining the static attribute data and dynamic behavioral data of the target object, the target object is characterized from multiple perspectives. Using a risk prediction model trained by federated learning based on the business system and the anti-fraud system realizes the effective integration of cross-system data. Using the behavior prediction model to gradually predict the dynamic behavioral data to generate predicted behavioral data and calculating the deviation degree between it and the actual dynamic behavioral data can effectively capture the abnormal behavior of the target object during the operation process. And since the anti-fraud system cannot obtain specific behavioral data, an independent behavior prediction model is used to analyze it instead of incorporating it into the federated learning system. This separation processing method can effectively improve the data utilization rate. Determining the target risk probability of the target object according to the deviation degree and the transaction risk probability realizes the comprehensive consideration of static and dynamic risk factors, making the target risk probability more accurate.

[0052] According to an embodiment of the present disclosure, using a behavior prediction model trained based on the historical behavioral data of the target object to gradually predict multiple sub-behavioral data in the dynamic behavioral data to generate predicted behavioral data, including: sorting the multiple sub-behavioral data in the dynamic behavioral data according to the time stamp to obtain behavior sorting data; performing feature extraction and vector conversion on each sub-behavioral data in the behavior sorting data to generate a dynamic behavior sequence composed of multiple feature vectors; inputting the dynamic behavior sequence into the behavior prediction model to output a predicted behavior sequence, where each predicted vector in the predicted behavior sequence corresponds one by one to the feature vector in the dynamic behavior sequence; converting the multiple predicted vectors in the predicted behavior sequence into a preset format corresponding to the dynamic behavioral data to obtain predicted behavioral data.

[0053] According to an embodiment of the present disclosure, the dynamic behavior data is a set composed of multiple sub-behavior data with independent characteristics, where each sub-behavior data contains a timestamp attribute that precisely records the moment when the behavior occurs. This timestamp serves as a key indicator for measuring the temporal relationship of behaviors and provides an important basis for data sorting. The system performs a temporal sorting operation on the dynamic behavior data based on the timestamp. By comparing the timestamp values of each sub-behavior data, they are arranged in ascending order, that is, the sub-behavior data with a smaller timestamp is arranged first, and the larger ones follow, finally generating behavior sorting data that strictly follows the chronological order. This sorting process ensures the temporal integrity of the behavior data, establishes a basic data structure that conforms to the actual behavior occurrence order for subsequent analysis and processing, and effectively guarantees the temporal consistency of the analysis process.

[0054] According to an embodiment of the present disclosure, after obtaining the behavior sorting data, the system will perform feature extraction processing on each sub-behavior data. Feature extraction is a process of screening and refining key feature information from sub-behavior data through specific algorithms, aiming to capture the feature dimensions that can effectively characterize the essence of the behavior. Specifically, the features extracted by the system mainly include: behavior types (such as click, slide, input, etc. operations), interface position coordinates (recording the specific position where the behavior occurs), duration (the behavior duration accurate to the millisecond level), and other core dimensions. Through predefined feature extraction algorithms, the system can accurately identify and extract these discriminative feature information from the original sub-behavior data, providing high-quality input for subsequent vectorization conversion and model analysis.

[0055] According to an embodiment of the present disclosure, vector conversion is a key step in converting the extracted feature information into a numerical vector representation. Since machine learning models have better processing capabilities for numerical data, this conversion process enables the behavior features to be effectively recognized and calculated by the model. Specifically, based on the type and number of dimensions of the features, the system maps the features of each sub-behavior data to a point in an n-dimensional vector space, where each dimension corresponds to a specific feature attribute, and its value represents the quantization result of the feature. Through this conversion process, the original behavior sorting data is converted into a dynamic behavior sequence composed of multiple feature vectors. This sequence not only completely retains the feature information of the original dynamic behavior data but also converts it into a standardized numerical representation form suitable for machine learning model processing, providing structured input data for subsequent model training and prediction.

[0056] According to an embodiment of the present disclosure, the generated dynamic behavior sequence is used as an input to a pre-trained behavior prediction model. The behavior prediction model is trained based on historical behavior data. By using deep learning algorithms to mine the latent patterns and evolution laws in the behavior sequence, a model architecture capable of accurately predicting future behavior sequences is constructed. After the dynamic behavior sequence is input into the behavior prediction model, the behavior prediction model performs layer-by-layer calculations and feature extractions on each feature vector in the input sequence according to the parameter weights and network structure obtained during its training, and finally outputs the corresponding prediction results.

[0057] According to an embodiment of the present disclosure, during the model inference process, the system predicts future behaviors for each feature vector in the dynamic behavior sequence based on the knowledge representation and behavior patterns obtained through training, and outputs the prediction results in vector form. These output prediction vectors form a complete predicted behavior sequence in order, where each prediction vector maintains a strict correspondence with the feature vector in the input dynamic behavior sequence. Through this prediction mechanism, the ability to predict the time series of dynamic behavior data is achieved, providing a reliable prediction benchmark for subsequent deviation analysis and risk decision-making.

[0058] According to an embodiment of the present disclosure, the predicted behavior sequence is composed of multiple prediction vectors, but its data format may be different from the original dynamic behavior data. To ensure the compatibility of the prediction results with the original data, the system needs to convert the prediction vectors into a preset format. The preset format is pre-designed based on business requirements and data logic, and is exactly the same as the original dynamic behavior data format, including complete information fields and a standardized data structure. Through this conversion process, the prediction results can be compared and analyzed with the original data under the same standard, ensuring the consistency and reliability of subsequent processing procedures.

[0059] According to an embodiment of the present disclosure, during the format conversion process, the system performs structured reconstruction and data mapping operations on the prediction vectors according to the preset format. Specifically, the system converts the numerical features in the prediction vectors into corresponding behavior semantic descriptions (such as operation types like click and slide), and maps the vector dimensions into a predefined field structure. After the conversion is completed, the system generates predicted behavior data that is completely compatible with the original dynamic behavior data format.

[0060] According to an embodiment of the present disclosure, by sorting the sub-behavior data according to timestamps, the temporal continuity and integrity of the behavior data are ensured, providing basic data that conforms to the actual order of behavior occurrences for subsequent analysis. By using the prediction model to generate a predicted behavior sequence corresponding to the input sequence, the evolution law of behaviors is accurately captured, and the future development trend of behaviors is predicted, providing reliable data support for risk prediction.

[0061] According to an embodiment of the present disclosure, the risk prediction method further includes: aligning the feature vectors in the dynamic behavior sequence with the prediction vectors in the prediction behavior sequence in the time dimension; calculating the distance between each aligned feature vector and the corresponding prediction vector; and taking the average of the distances between all feature vectors and prediction vectors to obtain the deviation degree.

[0062] According to an embodiment of the present disclosure, the feature vectors in the dynamic behavior sequence are quantitative representations of the user's actual behavior at specific time points, while the prediction vectors in the prediction behavior sequence are the prediction outputs of the model for future behavior at the corresponding time points. As a core preprocessing step, time dimension alignment ensures that each actual behavior feature vector is precisely synchronized with its corresponding prediction result on the time axis by establishing a strict time mapping relationship, providing a reliable time benchmark for subsequent behavior deviation analysis and risk assessment.

[0063] According to an embodiment of the present disclosure, due to factors such as data acquisition delay and transmission error, there may be a time deviation between the dynamic behavior sequence and the prediction behavior sequence. To ensure analysis accuracy, the system performs precise time dimension alignment based on timestamp information. First, a time indexing mechanism is established to precisely match the feature vectors and prediction vectors at the same time point, forming one-to-one corresponding vector pairs. Second, for the case where there is unilateral data at some time points (only feature vectors or prediction vectors), the system adopts strategies such as interpolation completion, data elimination, or retaining missing values according to business rules and data characteristics to handle the situation while maintaining data integrity and sequence consistency. Through this refined alignment process, the system provides accurate time-synchronized data pairs for subsequent distance calculation, effectively avoiding miscomparisons caused by time deviation, ensuring the reliability of the behavior deviation analysis results, and laying a solid foundation for the objective evaluation of the model's prediction performance.

[0064] Figure 3 Schematically shows a schematic diagram of vector alignment in the risk prediction method according to an embodiment of the present disclosure.

[0065] As Figure 3 shown, the dynamic behavior sequence 310 is [vector a, vector b, vector c...], and the prediction behavior sequence 320 is [vector 1, vector 2, vector 3...]. During the prediction process, it is necessary to know vector a first to predict the corresponding vector 1, so the arrangement of the dynamic behavior sequence and the prediction behavior sequence is in an interleaved pattern. After aligning them in the time dimension, each feature vector and prediction vector can form a vector pair 330, for example, vector a corresponds to vector 1.

[0066] According to an embodiment of the present disclosure, after completing the alignment in the time dimension, for each pair of corresponding feature vectors and prediction vectors, it is necessary to calculate the distance between them. The vector distance is a quantitative index to measure the degree of difference between two vectors. Different distance calculation methods can reflect the differences between vectors from different perspectives.

[0067] According to an embodiment of the present disclosure, common distance calculation methods include Euclidean distance, Manhattan distance, Chebyshev distance, etc. Taking Euclidean distance as an example, it calculates the straight-line distance between two vectors in a multi-dimensional space. This distance intuitively reflects the positional difference between two vectors in space. The smaller the distance, the closer the two vectors are, that is, the more similar the actual behavior and the predicted behavior are at this time point.

[0068] According to an embodiment of the present disclosure, in practical applications, it is necessary to select a suitable distance calculation method according to the specific business scenario and data characteristics. For the feature vectors and prediction vectors corresponding to each time point, use the selected distance calculation method to calculate, and obtain a set of distances reflecting the difference between the actual behavior and the predicted behavior at this time point. By calculating the vector distance, the difference between the actual behavior and the predicted behavior is quantified, providing specific data support for the subsequent calculation of the deviation degree. These distance values can intuitively reflect the deviation degree between the model prediction result and the actual situation at each time point, helping analysts deeply understand the prediction accuracy of the model at different time points.

[0069] According to an embodiment of the present disclosure, after obtaining the distance values between each pair of feature vectors and prediction vectors, it is necessary to summarize and analyze these distance values to obtain a comprehensive index to measure the overall prediction deviation degree, and this index is the deviation degree. The method of calculating the deviation degree is to add up all the distance values between the feature vectors and the prediction vectors, and then divide by the total number of distance values, that is, take the average value. The deviation degree is a comprehensive statistical index, which eliminates the influence of accidental fluctuations at individual time points and can reflect the difference degree between the predicted behavior sequence and the dynamic behavior sequence as a whole.

[0070] According to an embodiment of the present disclosure, as a comprehensive evaluation index, the deviation degree provides an intuitive and easy-to-understand standard for evaluating the performance of the behavior prediction model. A smaller deviation degree indicates that the prediction result of the model is closer to the actual behavior, and the model has higher accuracy and reliability. A larger deviation degree indicates that there is a large deviation in the model prediction, and it is necessary to further optimize and adjust the model, such as adjusting model parameters, increasing training data, improving feature engineering, etc., to improve the prediction performance of the model.

[0071] According to an embodiment of the present disclosure, determining the target risk probability of a target object based on the deviation degree between predicted behavior data and dynamic behavior data, and the trading risk probability includes: performing normalization processing on the deviation degree to obtain an intermediate value; using an activation function to perform probability conversion on the intermediate value to obtain a deviation risk probability corresponding to the deviation degree; based on a first weight corresponding to the deviation risk probability and a second weight corresponding to the trading risk probability, performing weighted summation on the deviation risk probability and the trading risk probability to obtain the target risk probability.

[0072] According to an embodiment of the present disclosure, the deviation degree is a comprehensive index for measuring the difference degree between a predicted behavior sequence and a dynamic behavior sequence. However, since its value range may vary due to different data characteristics and calculation methods, it is difficult to directly compare and analyze the deviation degrees in different data sets or different scenarios. Therefore, it is necessary to perform normalization processing on the deviation degree to map its value range to a fixed interval, usually the [0, 1] interval, so as to obtain an intermediate value. Specifically, the minimum value and the maximum value in the deviation degree data set can be found, and then normalization processing is performed on each deviation degree value.

[0073] According to an embodiment of the present disclosure, although the intermediate value obtained through normalization processing is already in a relatively unified interval, it is not a true probability value. In order to convert the intermediate value into a deviation risk probability corresponding to the deviation degree, it is necessary to use an activation function for processing. The activation function is a non-linear function that can map the input value to a probability distribution interval, usually the [0, 1] interval, so as to convert the intermediate value into a value with probability significance. Common activation functions include the sigmoid function, the softmax function, etc.

[0074] According to an embodiment of the present disclosure, taking the intermediate value as the input of the activation function, the corresponding output value is obtained through calculation, and this output value is the deviation risk probability. For each intermediate value, the activation function is used for calculation, so as to obtain a set of deviation risk probabilities corresponding to the deviation degree. After obtaining the deviation risk probability, it is also necessary to combine the trading risk probability to comprehensively evaluate the target risk. In order to comprehensively consider the influence degrees of the deviation risk and the trading risk on the target risk, corresponding weights, namely the first weight and the second weight, need to be assigned to the deviation risk probability and the trading risk probability respectively.

[0075] According to an embodiment of the present disclosure, the first weight and the second weight are preset according to specific business scenarios and risk assessment requirements. Their value ranges are usually between [0, 1], and the sum of the two is 1. For example, if in a certain business scenario, the deviation risk has a greater impact on the target risk, the first weight can be set higher. On the contrary, if the trading risk is more critical, the second weight can be set higher.

[0076] According to an embodiment of the present disclosure, the method for calculating the target risk probability by weighted summation can comprehensively consider the influence of deviation risk and transaction risk on the target risk, making the risk assessment more comprehensive and accurate.

[0077] According to an embodiment of the present disclosure, after calculating the target risk probability of the target object, the system executes a hierarchical early warning mechanism according to a preset risk level threshold. For the medium and low risk levels, the system pushes a pop-up reminder to the user to prompt them to pay attention to the current operation risk. For the high risk level, the system synchronously triggers a dual early warning mechanism. On the one hand, it blocks suspicious transactions in real time through the anti-fraud center client, and on the other hand, it automatically links with the regulatory agency for risk verification. This hierarchical response strategy not only ensures the timeliness of risk disposal but also realizes the accuracy of risk control, effectively reducing potential losses. At the same time, the system pushes the early warning information to the business personnel in real time to provide decision-making support for them, forming a complete risk prevention and control closed loop.

[0078] Figure 4 The flowchart of the risk prediction method according to another embodiment of the present disclosure is schematically shown.

[0079] As Figure 4 shown, another embodiment includes operations S410 to S460.

[0080] In operation S410, it is determined that the user logs in to the mobile banking. In operation S420, the static attribute data of the user and the dynamic behavior data generated when the user conducts business operations are obtained. In operation S430, the deviation degree is determined by using the behavior prediction model local to the mobile banking. In operation S440, the static attribute information is sent to the central server, and the transaction risk probability is determined through the risk prediction model. In operation S450, the deviation degree and the transaction risk probability are weighted and summed to obtain the target risk probability. In operation S460, a prompt is output according to the target risk probability to perform operation restrictions.

[0081] According to an embodiment of the present disclosure, a user portrait is constructed through user login confirmation and behavior data collection, and then the behavior deviation degree is calculated using a local model. At the same time, a comprehensive evaluation is carried out in combination with the risk prediction of the central server, and then risk control measures are implemented, effectively preventing potential risks while ensuring the user experience.

[0082] According to an embodiment of the present disclosure, the risk prediction method further includes: obtaining the number of risk events that occurred within a predetermined time period from an anti-fraud system; in the case where the number of events is greater than a preset number, performing a normalization process on the number of events, multiplying the normalized number of events by an intermediate value to obtain a risk reference value; determining a first weight corresponding to the risk reference value according to a preset mapping relationship between a preset weight and the risk reference value; and determining a second weight corresponding to the first weight according to a weight complementary relationship.

[0083] According to an embodiment of the present disclosure, as an information system dedicated to monitoring and recording risk events, the anti-fraud system stores a large amount of detailed data on risk events. To conduct subsequent risk assessments, it is necessary to extract the number of risk events that occurred within a predetermined time period from this system.

[0084] According to an embodiment of the present disclosure, the predetermined time period is a time range preset according to specific business requirements and analysis purposes. For example, it can be the past week, month, or quarter, etc. In actual operation, by sending a specific query request to the anti-fraud system, using the interfaces or query functions provided by the system, specifying the query time range, the system will screen out all risk events that occurred within this time period according to preset rules and count their numbers.

[0085] According to an embodiment of the present disclosure, after obtaining the number of events, it is necessary to judge it. The preset number is a threshold preset according to historical data, business experience, or industry standards, and is used to judge whether the number of risk events in the current time period is within the normal range.

[0086] According to an embodiment of the present disclosure, if the number of events is greater than the preset number, it indicates that risky activities are relatively frequent during this time period and further processing is required. To eliminate the differences in the value ranges of the number of events in different time periods and facilitate comparison and calculation with other data, it is necessary to perform a normalization process on the number of events.

[0087] According to an embodiment of the present disclosure, after obtaining the normalized number of events, multiply it by the previously obtained intermediate value (the value after deviation normalization processing) to obtain a risk reference value. The purpose of this step is to comprehensively consider the occurrence of risk events and the deviation information calculated previously, and provide a comprehensive reference index for subsequent weight determination.

[0088] According to an embodiment of the present disclosure, there is a preset mapping relationship between the preset weight and the risk reference value. This mapping relationship is established in advance according to a large amount of historical data, risk analysis models, and business experience. It stipulates the first weight corresponding to different risk reference values.

[0089] According to an embodiment of the present disclosure, after obtaining the risk reference value, according to a preset mapping relationship, the first weight corresponding to the risk reference value is found. For example, the mapping relationship can be a table or a function. According to the magnitude of the risk reference value, the corresponding first weight value is found in the table, or the first weight is calculated through the function.

[0090] According to an embodiment of the present disclosure, after determining the first weight, the second weight is determined according to the weight complementary relationship. The weight complementary relationship means that the sum of the first weight and the second weight is 1, that is, the two together constitute the weight distribution of each risk factor in the comprehensive risk assessment. Through the joint action of the first weight and the second weight, the relative importance of the deviation risk and the transaction risk in the comprehensive risk can be comprehensively and accurately reflected, providing a reasonable weight basis for finally calculating the target risk probability.

[0091] According to an embodiment of the present disclosure, the behavior prediction model is trained in the following manner: converting the historical behavior data of the target object into a time series format; inputting the converted historical behavior data into the initial prediction model, and calculating the prediction result through forward propagation, where the initial prediction model includes a long short-term memory network layer for capturing long-term dependencies in the historical behavior data; comparing the prediction result with the real behavior data in the historical behavior data to calculate the model loss value using a loss function; based on the model loss value, updating the model parameters using the backpropagation algorithm, and iteratively training until the model converges to obtain the behavior prediction model.

[0092] According to an embodiment of the present disclosure, the historical behavior data of the target object usually exists in various forms of original records, covering diverse behavior information shown by the target object at different time points, such as browsing, purchasing, and collecting operations of a user on an e-commerce platform, or state changes of a device during different periods of operation. To adapt to the subsequent model training requirements, special data processing algorithms and tools are needed to convert these scattered and unstructured historical behavior data into a regular time series format.

[0093] According to an embodiment of the present disclosure, the historical behavior data that has been converted into a time series format is used as input and fed into the initial prediction model. The core component of the initial prediction model includes a long short-term memory network layer (LSTM, Long Short-Term Memory), which has unique memory units and gating mechanisms and is specifically used to overcome the long-term dependence problem that is difficult to handle by traditional neural networks.

[0094] According to an embodiment of the present disclosure, when historical behavior data flows into the LSTM layer, the internal input gate, forget gate, and output gate operate in coordination. The input gate filters new information into the memory unit, the forget gate decides whether to retain or discard past memories, and the output gate controls the memory unit to output the information required at present. The LSTM layer processes the input data step by step, gradually extracting complex feature patterns hidden in historical behaviors, and with the assistance of subsequent structures such as fully connected layers, finally generates a prediction result for subsequent behaviors, realizing the inference from history to future trends.

[0095] According to an embodiment of the present disclosure, after obtaining the prediction result through forward propagation, the accuracy of the model prediction is measured. The prediction result is compared one by one with the corresponding real behavior data in the historical behavior data. The real behavior data, as a record of objective facts, carries the details of the actual behaviors occurring at each time node in history.

[0096] According to an embodiment of the present disclosure, a comparison calculation is performed using a pre-selected loss function. Common loss functions such as mean squared error, cross-entropy, etc. are adapted and selected according to different types of prediction tasks (regression or classification). Taking the mean squared error as an example, it calculates the mean of the squares of the deviations between the predicted value and the true value. The larger the deviation, the larger the function value, intuitively reflecting the degree of deviation of the prediction from the real situation, and thus accurately quantifying the loss value of this model prediction.

[0097] According to an embodiment of the present disclosure, based on the calculated model loss value, the backpropagation algorithm is started to drive model optimization. The backpropagation, according to the chain rule, derives the contribution of the weights of each layer to the loss layer by layer in the reverse direction from the loss value, and accurately locates the key parts that need to be adjusted.

[0098] According to an embodiment of the present disclosure, in the long short-term memory network layer and other components of the model, parameters such as neuron connection weights and biases are correspondingly adjusted to reduce the deviation between the prediction and the truth. Then, the processes of forward propagation, comparison calculation of loss, and backpropagation to update parameters are repeated, and continuous iterative training is carried out until the model loss value converges to the vicinity of a preset minimum value, indicating that the model performance reaches a relatively stable state. At this time, the training is completed to obtain the behavior prediction model. By means of the continuous and accurate self-optimization of the backpropagation algorithm, the model can continuously correct its own deficiencies and gradually improve the prediction accuracy.

[0099] According to an embodiment of the present disclosure, converting the historical behavior data of a target object into a time series format includes: performing data cleaning on the historical behavior data to remove redundant information, obtaining intermediate behavior data; constructing statistical features from multiple dimensions based on the intermediate behavior data, where the multiple dimensions include a time dimension, a numerical dimension, and a behavior pattern dimension; extracting target features with a correlation degree reaching a preset threshold with a preset variable from the intermediate behavior data through correlation analysis; using a random forest algorithm to extract key features from the intermediate behavior data through a feature importance evaluation mechanism; and converting the statistical features, target features, and key features into a time series format.

[0100] According to an embodiment of the present disclosure, historical behavior data often contains a large amount of redundant information, such as duplicate records, error data, incomplete data entries, etc. These information will interfere with the subsequent analysis and modeling processes. Therefore, it is necessary to perform data cleaning on the historical behavior data.

[0101] According to an embodiment of the present disclosure, identifying duplicate records in the data can specifically be done by comparing the key identification fields of the data, finding records that are exactly the same or have a very high similarity, and retaining one of them while deleting the rest of the duplicates. Checking for error data in the data, such as cases where the data type does not match, the value exceeds a reasonable range, etc. For error data, it can be corrected or deleted according to the specific situation. In addition, for incomplete data entries, if the missing part has little impact on the overall analysis, data filling can be performed, such as filling with statistical quantities such as the mean and median; if the missing part is severe, the data entry is directly deleted. After these processes, redundant information is removed, and intermediate behavior data is obtained.

[0102] According to an embodiment of the present disclosure, after obtaining the intermediate behavior data, statistical features are constructed from multiple dimensions to more comprehensively describe the features and patterns of the data. Constructing statistical features from the time dimension can analyze the variation pattern of the behavior over time. For example, the time interval between the occurrences of the behavior can be calculated to understand the periodicity of the behavior. Counting the occurrence frequency of the behavior in different time periods to find the peak and trough periods of the behavior. It is also possible to analyze the time distribution of the occurrence of the behavior, such as whether it shows seasonality, differences between weekdays and rest days, etc.

[0103] According to an embodiment of the present disclosure, in the numerical dimension, statistical analysis can be performed on the numerical variables in the intermediate behavior data. Calculating statistical quantities such as the mean, median, and standard deviation of the variables to understand the central tendency and dispersion degree of the data. Analyzing the maximum and minimum values of the variables to determine the value range of the data. It is also possible to perform binning processing to divide continuous numerical variables into different intervals to better observe the distribution characteristics of the data.

[0104] According to an embodiment of the present disclosure, constructing statistical features from the dimension of behavior patterns can uncover the associations and patterns between behaviors. For example, by analyzing the co-occurrence frequencies between different behaviors, combinations of behaviors that often occur simultaneously can be identified. Calculate the transition probabilities of behaviors to understand the likelihood of a behavior changing from one state to another. Sequence analysis can also be used to find the sequence and patterns of behaviors.

[0105] According to an embodiment of the present disclosure, in order to further screen out features that have a significant impact on a preset variable, correlation analysis needs to be performed. Select an appropriate correlation measurement method, such as the Pearson correlation coefficient, Spearman correlation coefficient, etc., and make the selection according to the type and distribution of the data. Calculate the correlation coefficient between each feature in the intermediate behavior data and the preset variable. Compare the correlation coefficient with a preset threshold. Only features whose correlation reaches the preset threshold are considered to have a relatively high degree of correlation with the preset variable, and these features are extracted as target features.

[0106] According to an embodiment of the present disclosure, the random forest algorithm is an ensemble learning method composed of multiple decision trees and can be used for feature importance evaluation. Use the intermediate behavior data as input to construct a random forest model. During the construction process, the random forest will train each decision tree. When each decision tree is trained, it will randomly select a part of the features and samples for learning. Then, according to indicators such as the number of splits and information gain of each feature in the decision tree in the random forest, evaluate the importance of each feature. Through the feature importance evaluation mechanism, find the features that have a greater impact on the model prediction result, and extract these features as key features.

[0107] According to an embodiment of the present disclosure, convert the previously constructed and extracted statistical features, target features, and key features into a time series format. The time series format requires the data to be arranged in chronological order, and each time point corresponds to a set of feature values. Therefore, it is necessary to sort and organize the features according to the timestamp information of the data. If there are missing time points in the feature data, interpolation processing can be performed according to the specific situation, such as using linear interpolation, spline interpolation, etc. Store the sorted feature data in chronological order to form data in a time series format for subsequent input into the behavior prediction model for training and analysis. By making full use of the time information of the data, improve the model's ability to capture the laws of behavior changes over time.

[0108] Based on the above risk prediction method, the present disclosure also provides a risk prediction device. The following will be combined with Figure 5 Describe this device in detail.

[0109] Figure 5 Schematically shows a structural block diagram of a risk prediction device according to an embodiment of the present disclosure.

[0110] As shown Figure 5 in FIG. Figure 5 , the risk prediction device 500 of this embodiment includes a data acquisition module 510, a transaction prediction module 520, a behavior prediction module 530, and a risk prediction module 540.

[0111] The data acquisition module 510 is configured to acquire the static attribute data of the target object and the dynamic behavior data generated when the target object performs business operations. In one embodiment, the data acquisition module 510 may be configured to execute the operation S210 described above, which will not be elaborated here.

[0112] The transaction prediction module 520 is configured to input the static attribute data into a risk prediction model trained by federated learning based on a business system and an anti-fraud system, and output a transaction risk probability. In one embodiment, the transaction prediction module 520 may be configured to execute the operation S220 described above, which will not be elaborated here.

[0113] The behavior prediction module 530 is configured to use a behavior prediction model trained according to the historical behavior data of the target object to gradually predict multiple sub-behavior data in the dynamic behavior data, and generate predicted behavior data. In one embodiment, the behavior prediction module 530 may be configured to execute the operation S230 described above, which will not be elaborated here.

[0114] The risk prediction module 540 is configured to determine the target risk probability of the target object according to the deviation degree between the predicted behavior data and the dynamic behavior data, and the transaction risk probability. In one embodiment, the risk prediction module 540 may be configured to execute the operation S240 described above, which will not be elaborated here.

[0115] According to an embodiment of the present disclosure, the behavior prediction module 530 includes a sorting sub-module, a vector conversion sub-module, a prediction sub-module, and a format conversion sub-module.

[0116] The sorting sub-module is configured to sort multiple sub-behavior data in the dynamic behavior data according to the time stamp to obtain behavior sorted data.

[0117] The vector conversion sub-module is configured to perform feature extraction and vector conversion on each sub-behavior data in the behavior sorted data to generate a dynamic behavior sequence composed of multiple feature vectors.

[0118] The prediction sub-module is configured to input the dynamic behavior sequence into the behavior prediction model and output a predicted behavior sequence, where each prediction vector in the predicted behavior sequence corresponds one-to-one with the feature vector in the dynamic behavior sequence.

[0119] The format conversion sub-module is configured to convert multiple prediction vectors in the predicted behavior sequence into a preset format corresponding to the dynamic behavior data to obtain predicted behavior data.

[0120] According to an embodiment of the present disclosure, the risk prediction device 500 further includes a sequence alignment module, a distance calculation module, and a deviation calculation module.

[0121] The sequence alignment module is configured to align the feature vectors in the dynamic behavior sequence with the prediction vectors in the prediction behavior sequence in the time dimension.

[0122] The distance calculation module is configured to calculate the distance between each aligned feature vector and the corresponding prediction vector.

[0123] The deviation calculation module is configured to take the average of the distances between all feature vectors and prediction vectors to obtain a deviation degree.

[0124] According to an embodiment of the present disclosure, the risk prediction module 540 includes a normalization sub-module, a probability conversion sub-module, and a risk determination sub-module.

[0125] The normalization sub-module is configured to normalize the deviation degree to obtain an intermediate value.

[0126] The probability conversion sub-module is configured to perform probability conversion on the intermediate value using an activation function to obtain a deviation risk probability corresponding to the deviation degree.

[0127] The risk determination sub-module is configured to perform weighted summation on the deviation risk probability and the transaction risk probability based on a first weight corresponding to the deviation risk probability and a second weight corresponding to the transaction risk probability to obtain a target risk probability.

[0128] According to an embodiment of the present disclosure, the risk prediction device 500 further includes a quantity determination module, a reference determination module, a first determination module, and a second determination module.

[0129] The quantity determination module is configured to obtain the number of risk events that occurred within a predetermined time period from the anti-fraud system.

[0130] The reference determination module is configured to, when the number of events is greater than a preset number, normalize the number of events, and multiply the normalized number of events by the intermediate value to obtain a risk reference value.

[0131] The first determination module is configured to determine a first weight corresponding to the risk reference value according to a preset mapping relationship between a preset weight and the risk reference value.

[0132] The second determination module is configured to determine a second weight corresponding to the first weight according to a weight complementary relationship.

[0133] According to an embodiment of the present disclosure, the risk prediction device 500 further includes a data conversion module, a data prediction module, a data comparison module, and an iterative training module.

[0134] A data conversion module for converting historical behavior data of a target object into a time series format.

[0135] A data prediction module for inputting the converted historical behavior data into an initial prediction model and calculating a prediction result through forward propagation. The initial prediction model includes a long short-term memory network layer for capturing long-term dependencies in the historical behavior data.

[0136] A data comparison module for comparing the prediction result with the true behavior data in the historical behavior data to calculate a model loss value using a loss function.

[0137] An iterative training module for updating model parameters based on the model loss value using the backpropagation algorithm and iteratively training until the model converges to obtain a behavior prediction model.

[0138] According to an embodiment of the present disclosure, the data conversion module includes a redundancy removal sub-module, a feature construction sub-module, a correlation analysis sub-module, a feature extraction sub-module, and a feature conversion sub-module.

[0139] The redundancy removal sub-module is used to clean the historical behavior data to remove redundant information and obtain intermediate behavior data.

[0140] The feature construction sub-module is used to construct statistical features from multiple dimensions based on the intermediate behavior data, where the multiple dimensions include a time dimension, a numerical dimension, and a behavior pattern dimension.

[0141] The correlation analysis sub-module is used to extract target features with a correlation degree reaching a preset threshold with a preset variable from the intermediate behavior data through correlation analysis.

[0142] The feature extraction sub-module is used to extract key features from the intermediate behavior data using a random forest algorithm through a feature importance evaluation mechanism.

[0143] The feature conversion sub-module is used to convert the statistical features, target features, and key features into a time series format.

[0144] According to embodiments of the present disclosure, any plurality of modules among the data acquisition module 510, the transaction prediction module 520, the behavior prediction module 530, and the risk prediction module 540 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of the data acquisition module 510, the transaction prediction module 520, the behavior prediction module 530, and the risk prediction module 540 may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-chip, a system-on-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the data acquisition module 510, the transaction prediction module 520, the behavior prediction module 530, and the risk prediction module 540 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0145] Figure 6 A block diagram of an electronic device suitable for implementing a risk prediction method according to an embodiment of the present disclosure is schematically shown.

[0146] As shown in FIG. 6, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0147] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in one or more memories.

[0148] According to an embodiment of the present disclosure, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read therefrom can be installed into the storage portion 608 as needed.

[0149] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0150] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 602 and / or RAM 603 and / or ROM 602 and RAM 603.

[0151] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the risk prediction method provided by the embodiment of the present disclosure.

[0152] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0153] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program code included in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0154] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0155] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0157] Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0158] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A risk prediction method, characterized in that, The method includes: Obtaining the static attribute data of the target object and the dynamic behavior data generated when the target object performs business operations; Inputting the static attribute data into a risk prediction model trained by federated learning based on a business system and an anti-fraud system, and outputting a transaction risk probability; Using a behavior prediction model trained according to the historical behavior data of the target object to gradually predict multiple sub-behavior data in the dynamic behavior data, and generating predicted behavior data; Determining the target risk probability of the target object according to the deviation degree between the predicted behavior data and the dynamic behavior data, and the transaction risk probability.

2. The method according to claim 1, characterized in that The step of using a behavior prediction model trained according to the historical behavior data of the target object to gradually predict multiple sub-behavior data in the dynamic behavior data and generate predicted behavior data includes: Sorting the multiple sub-behavior data in the dynamic behavior data according to timestamps to obtain behavior sorted data; Performing feature extraction and vector conversion on each sub-behavior data in the behavior sorted data to generate a dynamic behavior sequence composed of multiple feature vectors; Inputting the dynamic behavior sequence into the behavior prediction model, and outputting a predicted behavior sequence, where each predicted vector in the predicted behavior sequence corresponds one-to-one with the feature vector in the dynamic behavior sequence; Converting the multiple predicted vectors in the predicted behavior sequence into a preset format corresponding to the dynamic behavior data to obtain the predicted behavior data.

3. The method according to claim 2, wherein The method further includes: Aligning the feature vectors in the dynamic behavior sequence with the predicted vectors in the predicted behavior sequence in the time dimension; Calculating the distance between each aligned feature vector and the corresponding predicted vector; Taking the average value of the distances between all the feature vectors and the predicted vectors to obtain the deviation degree.

4. The method according to claim 1, wherein The step of determining the target risk probability of the target object according to the deviation degree between the predicted behavior data and the dynamic behavior data, and the transaction risk probability includes: Performing normalization processing on the deviation degree to obtain an intermediate value; Performing probability conversion on the intermediate value by using an activation function to obtain a deviation risk probability corresponding to the deviation degree; Based on a first weight corresponding to the deviation risk probability and a second weight corresponding to the transaction risk probability, performing weighted summation on the deviation risk probability and the transaction risk probability to obtain the target risk probability.

5. The method according to claim 4, wherein The method further includes: Obtaining the number of risk events that occurred within a predetermined time period from the anti-fraud system; In the case where the number of events is greater than a preset number, performing normalization processing on the number of events, and multiplying the normalized number of events by the intermediate value to obtain a risk reference value; Determining a first weight corresponding to the risk reference value according to a preset mapping relationship between a preset weight and the risk reference value; Determining a second weight corresponding to the first weight according to a weight complementary relationship.

6. The method according to claim 1, wherein The behavior prediction model is trained in the following manner: Converting the historical behavior data of the target object into a time series format; Input the converted historical behavior data into the initial prediction model, and calculate the prediction result through forward propagation. Among them, the initial prediction model includes a long short-term memory network layer, and the long short-term memory network layer is used to capture long-term dependencies in the historical behavior data; Compare the prediction result with the true behavior data in the historical behavior data to calculate the model loss value using a loss function; Based on the model loss value, use the backpropagation algorithm to update the model parameters, and iteratively train until the model converges to obtain the behavior prediction model.

7. The method according to claim 6, wherein The conversion of the historical behavior data of the target object into a time series format includes: Clean the historical behavior data to remove redundant information to obtain intermediate behavior data; Based on the intermediate behavior data, construct statistical features from multiple dimensions, where the multiple dimensions include a time dimension, a numerical dimension, and a behavior pattern dimension; Through correlation analysis, extract target features from the intermediate behavior data whose correlation with a preset variable reaches a preset threshold; Use the random forest algorithm to extract key features from the intermediate behavior data through a feature importance evaluation mechanism; Convert the statistical features, the target features, and the key features into the time series format.

8. A risk prediction device, characterized in that, The device includes: A data acquisition module for acquiring static attribute data of a target object and dynamic behavior data generated when the target object performs a business operation; A transaction prediction module for inputting the static attribute data into a risk prediction model trained by federated learning based on a business system and an anti-fraud system, and outputting a transaction risk probability; A behavior prediction module for using a behavior prediction model trained according to the historical behavior data of the target object to gradually predict multiple sub-behavior data in the dynamic behavior data and generate predicted behavior data; A risk prediction module for determining the target risk probability of the target object according to the deviation between the predicted behavior data and the dynamic behavior data, and the transaction risk probability.

9. An electronic device, comprising: One or more processors; A memory for storing one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction, when executed by the processor, implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instruction, when executed by the processor, implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Cross-border electronic customs declaration inspection early warning method and system based on dynamic risk learning

    CN120746425A

  • A Cross-Border Electronic Customs Declaration Inspection Early Warning Method and System Based on Dynamic Risk Learning

    CN120746425B

  • Network risk behavior identification method, electronic equipment, storage medium and program

    CN121125169A