Continuous Financial Fraud Detection Method and System Based on Multimodal Data Fusion
The method of reconstructing samples through multimodal data fusion and autoencoder generation and reconstruction of samples has been solved, and the problem that existing financial fraud detection methods are difficult to utilize multimodal data and adapt to market changes is achieved, achieving higher detection accuracy and model stability.
Patent Information
- Application Number
- CN202510460875.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-14
AI Technical Summary
Existing financial fraud detection methods are difficult to make full use of complementary information from multimodal data and cannot adapt to the rapid changes in the financial market, resulting in a decline in predictive performance of the detection model when new data is added and there is a risk of privacy leakage.
The continuous financial fraud detection method based on multimodal data fusion is adopted. By obtaining multimodal financial transaction samples, extracting and aligning different modal features, generating reconstructed samples using an autoencoder, and training the detection model with newly added financial transaction samples.
It improves the generalization ability of the model on new data, enhances the ability of multimodal data fusion processing, improves the accuracy of fraud detection, and avoids the problem of model forgetting historical features and privacy leakage.
Smart Images

Figure CN119991293B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial fraud detection, and particularly to a continuous financial fraud detection method and system based on multi-modal data fusion. Background Art
[0002] In the field of financial market fraud detection, existing methods face the dual challenges of diverse data sources and dynamic changes in the market environment, making it difficult to fully utilize the complementary information of different modal data and adapt to the rapid fluctuations in the financial market. Traditional prediction models usually perform static modeling on single-modal data, ignoring the potential correlations between multi-modal data, such as historical transaction records, news texts, market sentiment, and social media information. This not only leads to the one-sidedness of prediction results but also fails to capture the key driving forces in the dynamic changes of the market. In recent years, although deep learning methods have been applied to fraud detection to automatically learn high-dimensional features in the financial market and adapt to non-linear relationships, most rely on static data training and do not fully consider the highly dynamic and non-stationary characteristics of the financial market. As a result, the detection model is still difficult to adapt to the addition of new data, and the prediction performance of the model significantly decreases when dealing with new data. At the same time, there is also the problem of privacy leakage caused by reusing old data. Summary of the Invention
[0003] The purpose of the present invention is to overcome the problems of the prior art and provide a continuous financial fraud detection method and system based on multi-modal data fusion.
[0004] The purpose of the present invention is achieved through the following technical solutions: A continuous financial fraud detection method based on multi-modal data fusion, the method comprising the following steps:
[0005] S1: Obtain multi-modal financial transaction samples, including text data, time series data, and graph structure data;
[0006] S2: Extract the features of different modal financial transaction samples and perform alignment processing on the features of different modal financial transaction samples;
[0007] S3: Generate reconstructed samples of financial transaction samples based on an autoencoder;
[0008] S4: Train the detection model using the reconstructed samples and newly added financial transaction samples;
[0009] S5: Use the trained detection model to perform financial fraud detection and output the detection results.
[0010] In one example, after extracting the features of different modal financial transaction samples, the multi-modal feature space is optimized by sharing a contrastive learning objective to achieve alignment processing of the features of different modal financial transaction samples.
[0011] In one example, the objective loss function for contrastive learning The expression is:
[0012] ;
[0013] Wherein, represents the total number of samples; represents positive samples from different modalities; represents the similarity metric function between features; represents the temperature coefficient; represents negative sample pairs; 、 are both sequence number tags.
[0014] In one example, the autoencoder includes an encoder and a decoder. The encoder extracts the features of the input financial transaction samples, and the decoder generates a reconstructed sample of the financial transaction samples based on the extracted features of the financial transaction samples.
[0015] In one example, after the step of generating a reconstructed sample of the original financial transaction samples based on the autoencoder, it further includes:
[0016] Performing clustering processing on the reconstructed samples to obtain a clustering result presented as a spherical region in the feature space;
[0017] Using the centroid of the spherical region as the representative sample of all reconstructed samples within the current spherical region;
[0018] Training the detection model using the dataset composed of the representative sample and the newly added financial transaction samples.
[0019] In one example, the calculation expression of the centroid is:
[0020] ;
[0021] Wherein, represents the clustering centroid; represents the clustering result set; is the number of reconstructed samples within the clustering result; represents the th reconstructed sample in the clustering result.
[0022] In one example, when training the detection model using the dataset composed of the representative sample and the newly added financial transaction samples, it includes:
[0023] Introducing the loss function of the representative sample based on the clustering centroid Training the detection model, and the loss function The expression of is:
[0024] ;
[0025] Among them, represents the training loss of the newly added financial transaction samples; is the regularization coefficient; represents the number of clustering results; is the serial number label; represents a detection model, which is used to extract features and make predictions on the input data; represents the representative samples after clustering centroid processing; , respectively represent the current and historical model parameters.
[0026] It should be further noted that the technical features corresponding to the above examples can be combined with each other or replaced to form a new technical solution.
[0027] The present invention also includes a continuous financial fraud detection system based on multi-modal data fusion. The system includes interconnected terminals and servers, and the server includes interconnected storage servers and training servers;
[0028] The terminal inputs multi-modal financial transaction samples to the storage server by accessing the storage server;
[0029] The training server is used to execute steps S1-S4 in the method formed by any one of the above examples or a combination of multiple examples;
[0030] The storage server is used to save the trained detection model;
[0031] The terminal and / or the training server and / or the storage server use the trained detection model to perform financial fraud detection and output the detection results.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] 1. In one example, by extracting the features of different-modal financial transaction samples and aligning the features of different-modal financial transaction samples, it can better adapt to the changes of different-modal data and improve the generalization ability on new data; at the same time, by aligning the features of different-modal samples, multi-modal data fusion processing is realized, and the complementary information of different-modal data can be effectively obtained, enabling the model to more comprehensively understand the features of financial transaction samples, and thus more accurately identify abnormal transaction behaviors, thereby improving the accuracy of fraud detection.
[0034] Using newly added financial transaction samples (new data) to train the detection model can enable the model to absorb unknown data types in real time, thereby improving the prediction performance when processing new data; at the same time, the reconstructed samples can supplement the feature coverage of the current multimodal financial transaction samples. Using reconstructed samples and new data to train the model can not only avoid the model forgetting historical features due to focusing only on new data, but also prevent privacy leakage caused by reusing old data (historical financial transaction samples).
[0035] 2. In one example, contrastive learning can bring the feature representations of related samples (positive samples) in different modalities closer together, while at the same time distance the representations of unrelated samples. The resulting multimodal embedding space can not only effectively fuse heterogeneous data, but also be sensitive to the volatility characteristics of the financial market.
[0036] 3. In one example, through clustering centroid processing, without modifying the characteristics of the real sample (original multimodal financial transaction sample), the originally scattered and changeable single reconstructed sample is converted into a more stable and representative centroid point, so that the reconstructed sample as a whole can more easily form a clear boundary with the real sample in the feature space, thereby significantly improving the model's ability to distinguish between real and reconstructed samples, and maximally protecting the real sample characteristics from identification and leakage risks, thereby effectively protecting the privacy security of the real sample.
[0037] 4. In one example, a loss function based on representative samples of cluster centroids is introduced to train the model, and a regularization constraint is added to the loss function. The model's attention to past knowledge can be adjusted by adjusting the regularization coefficient, that is, increasing the regularization coefficient can make the model pay more attention to past knowledge, and vice versa, the model will pay more attention to the learning of new knowledge. Through this mechanism, the model can effectively retain the memory of historical data patterns while adapting to changes in the dynamic financial market, thereby improving the stability and accuracy of predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings. The accompanying drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The same reference numerals are used in these drawings to represent the same or similar parts. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0039] Figure 1 A method flow chart provided for an example of the present invention;
[0040] Figure 2 An autoencoder architecture diagram provided for an example of the present invention;
[0041] Figure 3The comparative learning framework diagram provided for an example of the present invention;
[0042] Figure 4 The system framework diagram provided for an example of the present invention.
[0043] In the figure: 101 - the first terminal; 102 - the second terminal; 103 - the server. Detailed implementation manners
[0044] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] In the description of the present invention, ordinal numbers (e.g., "first and second", etc.) are used to distinguish objects, and are not limited to this order, and should not be construed as indicating or implying relative importance. Unless otherwise clearly defined and limited, the term "connection" should be understood in a broad sense. For example, it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0046] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0047] In one example, as Figure 1 shown, a continuous financial fraud detection method based on multi-modal data fusion, the method includes the following steps:
[0048] S1: Obtain multi-modal financial transaction samples.
[0049] Among them, the multi-modal financial transaction samples include text data related to transactions, (transaction) time series data, graph structure data, such as consumption records, bank loan records, securities trading records, etc., which can reflect transaction time, transaction location, transaction frequency, transaction accounts, etc.
[0050] S2: Extract the features of different-modal financial transaction samples, and perform alignment processing on the features of different-modal financial transaction samples, that is, use inter-modal self-supervised distillation to jointly train the extracted embedding features to learn the deep feature representations across modalities.
[0051] To fully explore the potential correlations between features of different modalities, an inter-modal self-supervised distillation method is adopted to construct a multi-modal embedding space, thereby achieving effective data fusion. Specifically, in the present invention, feature extraction is separately performed on text data, time series data (numerical information), and graph-structured data. Among them, for text data, existing text analysis models such as BERT model and Transformer model are used to convert sentences or articles into a series of features that can reflect their meanings; for numerical information, a simple neural network such as a fully connected neural network can be used. After preprocessing the numerical data (such as adjusting the data range), through linear calculations and appropriate non-linear transformations, features for subsequent data fusion processing are extracted. For graph-structured data, an image recognition network is adopted to convert images into a set of digital features that can reflect the content of the pictures, such as convolutional neural network, graph convolutional neural network, etc. Through the above feature extraction methods, different types of data can be processed into similar expression formats, which is convenient for subsequent feature alignment processing. When mapping features of different modalities to a unified feature space, for example, through a neural network layer (such as a fully connected layer), the dimensions and scales of features of different modalities are matched. A loss function can also be further used to optimize the effect of feature alignment processing, realizing the effective fusion of multi-modal data, and then making full use of the complementary information of different modal data to improve the model's understanding ability of market behavior, thereby improving the accuracy of the prediction model.
[0052] Preferably, before extracting the features of financial transaction samples, data preprocessing can also be performed, such as data cleaning, data normalization processing, data augmentation processing, etc.
[0053] S3: Generate a reconstructed sample of the financial transaction sample based on the autoencoder.
[0054] As Figure 2 shown, an autoencoder is an unsupervised learning neural network model, including an encoder and a decoder. The encoder is used to compress the input data (financial transaction sample) into a low-dimensional latent space, and the decoder is used to reconstruct the original data to generate a reconstructed data (reconstructed sample) similar to the input data.
[0055] S4: Use the reconstructed sample and the newly added financial transaction sample to train the detection model of this round.
[0056] The method of the present invention is a continuous dynamic financial fraud detection method based on continuous learning. Specifically, continuous learning means training a model on a data stream of a sequence of tasks, and the training goal is that the trained model can achieve good performance on all learned tasks. Each task has its own separate training set, validation set, and test set. When the model is trained, it can only access the data of the current round of training task. As Figure 3 shown, the samples corresponding to different tasks , and labels , are different, so it is necessary to retrain the generator and solver for each task, so as to be able to generate reconstructed samples in the next task.
[0057] Preferably, while the model is being trained, the model parameters are recorded to guide subsequent calibration tasks. Further, on the basis of completing the model training, the trained model is calibrated, including:
[0058] While the model is making predictions, combined with the input of real-time dynamic data, calculate the change in the loss function of the model input and the predicted value, and adjust the input feature weights according to the real-time market dynamics to ensure the adaptability of the prediction to the dynamic environment, and use the latest data and the adjusted loss function to recalibrate the target model, and output the calibrated financial market fraud detection value. Then, use the latest data and the adjusted loss function to perform online calibration on the model.
[0059] In step S4, the detection model of this round is trained using the reconstructed samples and the newly added financial transaction samples, where the reconstructed samples can supplement the feature coverage of the current multi-modal financial transaction samples. At the same time, training the model with the reconstructed samples can also prevent the problem of privacy leakage caused by reusing old data (historical financial transaction samples). This is because the reconstructed samples are reconstructed based on the deep features extracted from the original financial transaction samples. These features are abstract representations after being processed by the model, rather than the specific content of the original data. Therefore, the reconstructed samples reflect more of the statistical characteristics of the data rather than the specific information of individuals, which greatly reduces the risk of privacy leakage. Further, the newly added financial transaction samples (new data) can provide the latest data features and patterns for the detection model, enabling the model to learn and adapt to unknown data types in real time. By training the model with these new data, the prediction performance of the model when processing new data is significantly improved. For example, in bank fraud detection, the new data may contain new types of fraud behavior patterns, and the trained model can better identify these newly emerging fraud behaviors. In addition, during the training process, using the reconstructed samples and the new data simultaneously can prevent the model from only focusing on the new data and causing catastrophic forgetting of historical features, and solve the problem of catastrophic forgetting in continuous learning.
[0060] S5: Use the detection model that has completed this round of training to perform financial fraud detection and output the detection result.
[0061] Preferably, after completing step S5, determine whether the model has reached the update cycle. If so, return to step S4 and enter the next round of model training to optimize the performance of the detection model to adapt to the dynamically changing financial fraud detection scenario.
[0062] In this example, cross-modal features are extracted through inter-modal self-supervised distillation technology, and an autoencoder is used to generate reconstructed samples of financial transaction samples to enhance the playback of the historical state of the market. During the dynamic market observation process, knowledge distillation technology is adopted to align the features of financial transaction samples in different modalities, effectively alleviating the catastrophic forgetting problem and ensuring the continuity and stability of the detection model during data update. This method can significantly improve the real-time performance and accuracy of financial market fraud detection, is applicable to dynamic financial transaction detection scenarios, and provides an efficient and intelligent solution for financial fraud detection.
[0063] In one example, after extracting the features of financial transaction samples in different modalities, a unified contrastive learning objective is shared to optimize the multi-modal feature space, thereby achieving the alignment of the features of financial transaction samples in different modalities. Among them, contrastive learning is an unsupervised learning method aimed at learning effective feature representations of data by comparing the similarities and differences between samples. The core idea of contrastive learning is to make similar samples closer in the feature space and dissimilar samples farther apart. In contrastive learning, the model is trained by constructing positive sample pairs (similar samples, such as different enhanced versions of the same image) and negative sample pairs (dissimilar samples, such as different images), thereby achieving high-quality feature extraction on unlabeled data and providing a good basic representation for downstream tasks (such as classification or clustering). Optionally, the objective loss function of contrastive learning is used to measure the consistency of representations between different modalities, using to represent, and is defined as follows:
[0064] ;
[0065] Among them, represents the total number of samples; represents positive samples from different modalities; represents the similarity measurement function between features; represents the temperature coefficient; represents negative sample pairs; 、 are both serial number labels.
[0066] In one example, in the dynamic data processing part, in order to cope with the dynamic changes in the data distribution in the financial market, a feature-level joint generation adjustment mechanism, hereinafter referred to as the adjustment mechanism, is introduced. The adjustment mechanism is mainly implemented based on an autoencoder. In the autoencoder, the encoder extracts the features of the input original financial transaction samples, and the decoder generates reconstructed samples of the original financial transaction samples based on the extracted features of the financial transaction samples. The adjustment mechanism can not only effectively extract the deep features in the financial transaction samples, but also generate reconstructed samples similar to the input financial transaction samples.
[0067] Furthermore, the training objective of the autoencoder is to minimize the reconstruction error, which is denoted by and the loss function is expressed as follows:
[0068] ;
[0069] where, is the input financial transaction sample, is the reconstructed sample (pseudo-sample) of the reconstructed output of the decoder.
[0070] Through the above adjustment mechanism based on the autoencoder, not only can the feature distribution of the current market be learned, but also in subsequent training, when new data needs to be updated, the autoencoder generates reconstructed samples based on the stored features, and these reconstructed samples are fed back into the model training to supplement the feature coverage of the current data. This not only avoids the model forgetting historical features due to only focusing on new data, but also prevents privacy leakage caused by reusing old data.
[0071] Optionally, a regularization method or a modular method can also be used to replace the sample replay method, so as to enable the model to remember the important sample information of the previous tasks and avoid the problem of catastrophic forgetting of the model in continuous learning.
[0072] In one example, in order to further improve the distinguishability between the reconstructed sample and the original sample, a fuzzy processing mechanism based on the clustering of the reconstructed samples is introduced. Specifically, after the step of generating the reconstructed samples of the original financial transaction samples based on the autoencoder, it further includes:
[0073] a. Clustering the reconstructed samples to obtain a clustering result that presents as a spherical region in the feature space.
[0074] Specifically, the generated reconstructed samples are regarded as a set of data points in the feature space, and these data points are clustered. Each cluster can be regarded as a spherical region in the feature space.
[0075] b. Using the centroid of the spherical region as the representative sample of all the reconstructed samples within the current spherical region.
[0076] Using the centroid (i.e., the center of the sphere) of the sphere as the representative point of the reconstructed samples means that: all the reconstructed samples under this clustering category are represented by the centroid, because the centroid can more stably reflect the central features of this clustering region, thereby reducing the impact of the feature fluctuations of individual reconstructed samples on model training.
[0077] Preferably, the calculation expression of the centroid is:
[0078] ;
[0079] where, represents the clustering centroid; represents the set of clustering results; is the number of reconstructed samples within the clustering result; represents the th reconstructed sample in the clustering result.
[0080] c. Training the detection model with a dataset composed of representative samples and newly added financial transaction samples.
[0081] In one example, to further enhance the effect of joint generation, a regularization constraint is added during the training of new data to ensure that the model retains the adaptability to old market fraud detection while learning new market trend features. Specifically, a loss function based on the representative samples of the clustering centroid is used to train the detection model, and the loss function has the following expression:
[0082] ;
[0083] where represents the training loss of the newly added financial transaction samples; represents the number of clustering results; represents the detection model, which is used to extract features and make predictions on the input data; represents the representative samples after being processed by the clustering centroid; , represent the current and historical model parameters respectively; is the regularization coefficient, and by adjusting the value, the attention degree of the model to past knowledge is adjusted, that is: increasing can make the model pay more attention to the knowledge learned in previous tasks, and vice versa, the model will pay more attention to the learning of new knowledge. Through this mechanism, while adapting to the changes in the dynamic financial market, the model can effectively retain the memory of the historical data pattern, thereby improving the stability and accuracy of prediction.
[0084] Combining the above examples, a preferred example of the present invention is obtained. At this time, the method includes the following steps:
[0085] S10: Obtaining multi-modal financial transaction samples;
[0086] S20: Extracting the features of different-modal financial transaction samples and performing alignment processing on the features of different-modal financial transaction samples;
[0087] S30: Generating reconstructed samples of financial transaction samples based on the autoencoder;
[0088] S40: Perform clustering on the reconstructed samples to obtain a clustering result that presents as a spherical region in the feature space; use the centroid of the spherical region as the representative sample of all the reconstructed samples within the current spherical region; use the dataset composed of the representative sample and the newly added financial transaction samples to train the detection model, and introduce the loss function of the representative sample based on the clustering centroid during the model training process;
[0089] S50: Use the trained detection model to perform financial fraud detection and output the detection result.
[0090] Through dynamic multi-modal learning, this method realizes the efficient fusion of data related to financial market fraud, and uses the feature-level joint generation adjustment mechanism to dynamically adapt to the changes in the market environment, thereby constructing a fraud detection model with market volatility adaptability. This method can accurately identify abnormal trading patterns in a multi-modal environment, improve the sensitivity and robustness of fraud detection, and provide effective support for financial risk control. After completing the two-stage training, the constructed model is the target fraud detection model of this embodiment. This model can detect fraud behaviors in the market based on the multi-modal financial data provided by users and return corresponding prediction results. When new data arrives, it will fuse multi-modal information in real time and dynamically adjust the model parameters to ensure that the detection results always conform to the latest market environment, thus maintaining continuous and efficient fraud detection capabilities.
[0091] The present invention also includes a continuous financial fraud detection system based on multi-modal data fusion, as Figure 4 shown. This system includes a terminal and a server 103. The terminal includes a first terminal 101 and a second terminal 102. The server 103 includes a storage server and a training server. The first terminal 101 and the second terminal 102 are both connected to the storage server and the training server.
[0092] In this example, the first terminal can be a personal user terminal, an enterprise institution terminal, or other third-party data source terminals. The first terminal can be a desktop computer, a smart phone, a tablet computer, a laptop computer, etc., but is not limited thereto. The first terminal uploads the modal financial transaction samples to the training server and / or the storage server through the network interface, providing a basis for subsequent data processing and model training.
[0093] The storage server and the training server can be independent physical servers or server clusters. Preferably, the storage server is used to receive and save the multi-modal financial transaction samples uploaded by the first terminal, preprocess and classify and store them, and at the same time provide efficient data reading support for the training server. The training server trains a dynamic financial market fraud detection model based on the multi-modal financial transaction samples (including newly added financial transaction samples) in the storage server to generate a high-precision market fraud detection model, that is, the training server is used to execute the method steps S1-S4 and steps S10-S40 of the present invention. Optionally, the training server can be equipped with a Linux operating system and high-performance GPU computing resources to meet the needs of large-scale data processing and complex model training.
[0094] Furthermore, the storage server saves the trained detection model for access by the first terminal and the second terminal. At this time, the terminal and / or the training server and / or the storage server use the detection model to perform financial fraud detection and output the detection results.
[0095] The second terminal includes a device for receiving and displaying the detection results. Users can access the detection model stored in the server through the second terminal and input financial transaction samples to the detection model to obtain the detection results. The second terminal can be a desktop computer, a smart phone, a tablet computer, a laptop computer, etc., but is not limited thereto.
[0096] Through the collaborative action of the devices in the above system, the multi-modal financial data can be processed efficiently and the financial market fraud detection results can be generated dynamically. It should be noted that this system supports multiple second-terminal users to access simultaneously and can meet the diverse needs of different users.
[0097] In summary, the system of the present invention first obtains different-modal financial transaction data and performs fusion processing on different-modal data. Specifically, through the inter-modal self-supervised distillation technology, it models the fusion of multi-modal data features to improve the comprehensive understanding ability of the market state; at the same time, it introduces a feature-level joint generation adjustment mechanism, enabling the model to not only absorb unknown data types in real time but also dynamically adjust the impact of outdated information on the prediction to ensure the stability and accuracy of the prediction results. While improving the prediction accuracy, this method reasonably optimizes the allocation of computing resources and provides an efficient and flexible solution for dealing with the complex and changeable financial market.
[0098] The above specific embodiments are detailed descriptions of the present invention. It cannot be determined that the specific embodiments of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for detecting persistent financial fraud based on multimodal data fusion, characterized in that: The following steps are involved: S1: Obtain multimodal financial transaction samples, including text data, time series data, and graph structure data; S2: Extract the features of financial transaction samples of different modalities and align the features of financial transaction samples of different modalities; S3: Reconstructed samples of financial transaction samples generated based on autoencoders; S4: Use reconstructed samples and newly added financial transaction samples to train the detection model; S5: Execute financial fraud detection using the trained detection model and output the detection results; After the step of reconstructing samples based on the autoencoder to generate the original financial transaction samples, the method further includes: Perform clustering processing on the reconstructed samples to obtain clustering results that appear as spherical regions in the feature space; The centroid of the spherical region is used as the representative sample of all reconstructed samples in the current spherical region; The detection model is trained using a dataset consisting of representative samples and newly added financial transaction samples.
2. The method for detecting persistent financial fraud based on multimodal data fusion according to claim 1, characterized in that: After extracting the features of the financial transaction samples of different modalities, the multimodal feature space is optimized by sharing a comparative learning objective to achieve alignment processing of the features of the financial transaction samples of different modalities.
3. The method for detecting persistent financial fraud based on multimodal data fusion according to claim 2 is characterized in that: Objective loss function for contrastive learning The expression is: ; in, represents the total number of samples; Represents positive samples from different modalities; Represents the similarity measurement function between features; represents the temperature coefficient; represents a negative sample pair; , All are serial number labels.
4. The method for detecting persistent financial fraud based on multimodal data fusion according to claim 1, characterized in that: The autoencoder includes an encoder and a decoder. The encoder extracts features of an input financial transaction sample, and the decoder generates a reconstructed sample of the financial transaction sample based on the extracted features of the financial transaction sample.
5. The method for detecting persistent financial fraud based on multimodal data fusion according to claim 1, characterized in that: The calculation expression of the centroid is: ; in, represents the cluster centroid; Represents the clustering result set; is the number of reconstructed samples in the clustering results; Indicates the clustering result A reconstruction sample.
6. The method for detecting persistent financial fraud based on multimodal data fusion according to claim 1, characterized in that: When the detection model is trained using a data set consisting of representative samples and newly added financial transaction samples, it includes: Introducing a loss function based on representative samples of cluster centroids Train the detection model, loss function The expression is: ; in, represents the training loss of newly added financial transaction samples; is the regularization coefficient; Indicates the number of clustering results; is a serial number label; Represents the detection model, which is used to extract features and predict input data; It represents the representative sample after the cluster centroid processing; , represent the current and historical model parameters respectively.
7. A continuous financial fraud detection system based on multimodal data fusion, characterized in that: The system includes interconnected terminals and servers, and the servers include interconnected storage servers and training servers; The terminal accesses the storage server and inputs the multimodal financial transaction sample into the storage server; The training server is used to perform steps S1-S4 in the method according to any one of claims 1-6; The storage server is used to save the trained detection model; The terminal and / or the training server and / or the storage server performs financial fraud detection using the trained detection model and outputs the detection result.
Citation Information
Patent Citations
Financial transaction data anomaly detection method and device and computer equipment
CN118279055A
Commercial credit evaluation and supervision method based on multi-modal coevolution algorithm
CN119250963A