Methods and devices for verifying transaction data anomalies, electronic devices and storage media

By using a pre-trained anomaly detection model to extract features and perform multi-dimensional detection on transaction data, the problem of low accuracy in transaction data verification is solved, and more efficient transaction data anomaly verification is achieved.

CN120687992BActive Publication Date: 2025-10-31深圳市深圳通有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511204230.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-31
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

In existing technologies, there is a problem with the low accuracy of anomaly verification when transaction data is received in the cloud, which makes it difficult to guarantee the integrity and consistency of transactions.

Method used

A pre-trained anomaly verification model is adopted, including a feature extraction sub-model, a first sub-anomaly verification model, and a second sub-anomaly verification model. Through feature extraction, anomaly classification detection, and anomaly clustering detection, combined with supervised and unsupervised learning, the accuracy of transaction data verification is improved.

Benefits of technology

By using multi-dimensional analysis, the risk of a single model failing to capture various transaction anomalies is reduced, the accuracy and efficiency of transaction data verification are improved, and the integrity and consistency of transaction data are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687992B_ABST
    Figure CN120687992B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, electronic device, and storage medium for transaction data anomaly verification, belonging to the field of public transportation technology. The method includes: invoking an anomaly verification model comprising a feature extraction sub-model, a first sub-anomaly verification model, and a second sub-anomaly verification model to perform anomaly verification; invoking the feature extraction sub-model to extract features from a first server record file and multiple second server record files to obtain transaction data features and corresponding transaction data derived features; invoking the first sub-anomaly verification model to perform anomaly classification detection on the transaction data features and transaction data derived features to obtain a first anomaly verification result; invoking the second sub-anomaly verification model to perform anomaly clustering detection on the transaction data features and transaction data derived features to obtain a second anomaly verification result; and determining a target verification result based on the first and second anomaly verification results. This application embodiment can improve the accuracy of anomaly verification for transaction data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of public transportation technology, and in particular to a method and apparatus for verifying abnormal transaction data, an electronic device, and a storage medium. Background Technology

[0002] With the rapid development of information technology, the transmission and interaction of data between different terminal devices and cloud servers are becoming increasingly frequent, and data synchronization has thus become a key element in the operation of various industries. For example, in the context of card transactions, the synchronization of transaction data between terminal devices (such as subway turnstiles and POS machines) and the cloud is a core link in ensuring the integrity and consistency of transactions.

[0003] In card-based transaction scenarios, when terminal devices synchronize transaction data to the cloud, factors such as network fluctuations, device malfunctions, or system delays may cause anomalies in the transaction data received by the cloud, such as missing or duplicate data. Therefore, to ensure the accuracy of the transaction data synchronized by the terminal devices to the cloud, it is necessary to perform anomaly verification on the transaction data received by the cloud, so as to promptly identify and handle errors that occur during the synchronization process based on the anomaly verification results.

[0004] However, the accuracy of anomaly verification for transaction data in related technologies is relatively low. Summary of the Invention

[0005] The main objective of this application is to provide a method, apparatus, electronic device, and storage medium for verifying transaction data anomalies, aiming to improve the accuracy of anomaly verification of transaction data.

[0006] To achieve the above objectives, a first aspect of this application proposes a method for verifying transaction data anomalies, the method comprising:

[0007] The system acquires multiple first server record files received by the cloud within a first time interval before the current time, and acquires multiple second server record files received by the cloud within a second time interval before the current time; wherein each first server record file and each second server record file is sent to the cloud by the terminal device, and each first server record file and each second server record file includes multiple transaction data; each second server record file has been pre-verified for anomalies.

[0008] For each of the first server record files, an anomaly verification model is invoked based on the first server record file and multiple second server record files to perform anomaly verification; wherein, the anomaly verification model includes a feature extraction sub-model, a first sub-anomaly verification model and a second sub-anomaly verification model;

[0009] For each first server record file, the feature extraction sub-model is invoked based on the first server record file and multiple second server record files to perform feature extraction, thereby obtaining transaction data features and transaction data derived features corresponding to the transaction data features;

[0010] For each of the first server record files, the first sub-anomaly verification model is invoked based on the transaction data features and the transaction data derived features to perform anomaly classification detection, thereby obtaining a first anomaly verification result. Then, the second sub-anomaly verification model is invoked based on the transaction data features and the transaction data derived features to perform anomaly clustering detection, thereby obtaining a second anomaly verification result.

[0011] For each of the first server record files, a target verification result is determined based on the first anomaly verification result and the second anomaly verification result; wherein, the target verification result is used to indicate the verification status of the first server record file, and the verification status includes one or more of the following: data normal, data duplicate, and data missing.

[0012] In some embodiments, the training method of the anomaly detection model includes:

[0013] The training sample data includes a first sample server record file received by the cloud within a first sample time interval prior to the historical time, a first classification label, and multiple second sample server record files received by the cloud within a second sample time interval prior to the sample time. The first classification label is used to guide the training of the first sub-anomaly verification model. Each first sample server record file and each second sample server record file are sent from the terminal device to the cloud, and each first sample server record file and each second sample server record file includes multiple transaction data. Each second sample server record file has undergone anomaly verification beforehand.

[0014] Based on the first sample server record file and multiple second sample server record files, the feature extraction sub-model is invoked to perform feature extraction, thereby obtaining sample transaction data features and sample transaction data derived features corresponding to the sample transaction data features;

[0015] Based on the characteristics of the sample transaction data and the derived characteristics of the sample transaction data, the first sub-anomaly verification model is invoked to perform anomaly classification detection to obtain the first sample anomaly verification result. Based on the characteristics of the sample transaction data and the derived characteristics of the sample transaction data, the second sub-anomaly verification model is invoked to perform anomaly clustering detection to obtain the second sample anomaly verification result.

[0016] The model parameters of the anomaly verification model are adjusted based on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, the first sample anomaly verification result, the second sample anomaly verification result, and the first classification label.

[0017] In some embodiments, adjusting the model parameters of the anomaly verification model based on the sample transaction data features, the sample transaction data derived features, the first sample anomaly verification result, the second sample anomaly verification result, and the classification label includes:

[0018] The anomaly detection result of the first sample and the first classification label are evaluated by classification indicators to obtain the first adjustment factor;

[0019] The clustering index is evaluated on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, and the anomaly verification results of the second sample to obtain the second adjustment factor;

[0020] A target adjustment factor is determined based on the preset weights, the first adjustment factor, and the second adjustment factor, and the model parameters of the anomaly verification model are adjusted according to the target adjustment factor.

[0021] In some embodiments, after adjusting the model parameters of the anomaly verification model based on the sample transaction data features, the sample transaction data derived features, the first sample anomaly verification result, the second sample anomaly verification result, and the first classification label, the method further includes:

[0022] The sample transaction data features and the sample transaction data derived features are perturbed to obtain the adversarial sample transaction data features and the corresponding adversarial sample transaction data derived features;

[0023] Based on the adversarial sample transaction data features and the adversarial sample transaction data derived features, the first sub-anomaly verification model is invoked to perform anomaly classification detection, resulting in the third sample anomaly verification result. Based on the adversarial sample transaction data features and the adversarial sample transaction data derived features, the second sub-anomaly verification model is invoked to perform anomaly clustering detection, resulting in the fourth sample anomaly verification result.

[0024] The model parameters of the anomaly verification model are adjusted based on the adversarial sample transaction data features, the adversarial sample transaction data derived features, the third sample anomaly verification result, the fourth sample anomaly verification result, and the first classification label.

[0025] In some embodiments, after determining the target verification result corresponding to the first server record file based on the first anomaly verification result and the second anomaly verification result, the method further includes:

[0026] The verification status distribution is determined based on the target verification results corresponding to multiple first server record files. The verification status distribution includes a first distribution corresponding to normal data, a second distribution corresponding to duplicate data, and a third distribution corresponding to missing data.

[0027] If the third distribution among the verification state distributions is the largest, the network parameters of the terminal device at the current time are obtained; if the network parameters are greater than a preset network threshold, the first server record file and the second classification label corresponding to the missing data are obtained as the target verification result, and missing training sample data are constructed based on the first server record file, multiple second server record files, and the second classification label.

[0028] The anomaly detection model is optimized and trained based on the missing training sample data.

[0029] If the second distribution quantity is the largest among the verification status distribution quantities, and the current time is within a preset peak period, the first server record file and the third classification label corresponding to the data duplication in the target verification result are obtained, and duplicate training sample data are constructed based on the first server record file, multiple second server record files, and the third classification label;

[0030] The anomaly detection model is optimized and trained based on the repeated training sample data.

[0031] In some embodiments, the step of calling the feature extraction sub-model based on the first server record file and multiple second server record files to perform feature extraction, obtaining transaction data features and transaction data derived features corresponding to the transaction data features, includes:

[0032] The feature extraction sub-model is invoked to extract the first device identifier corresponding to the terminal device that sent the first server record file, the first file identifier corresponding to the first server record file, and the transaction time, transaction amount, transaction identifier, and transaction type corresponding to the multiple transaction data included in the first server record file;

[0033] The transaction data features are defined as the first device identifier, the first file identifier, and the transaction time, the transaction amount, the transaction identifier, and the transaction type corresponding to the multiple transaction data included in the first server record file.

[0034] The feature extraction sub-model is invoked to extract the second device identifier corresponding to the terminal device that sent multiple second server record files and the transaction amount corresponding to the multiple transaction data included in each second server record file;

[0035] The first device identifier is matched with the second device identifier corresponding to multiple second server record files to obtain a device identifier matching value;

[0036] The transaction time difference sequence corresponding to the first server record file is determined based on the transaction time corresponding to the multiple transaction data included in the first server record file, and the transaction frequency corresponding to the first server record file is determined based on the multiple transaction data included in the first server record file and the first time interval.

[0037] A normalization factor is determined based on the transaction amount corresponding to the multiple transaction data included in the multiple second server record files, and the transaction amount corresponding to the multiple transaction data included in the first server record file is normalized based on the normalization factor to obtain the normalized transaction amount sequence corresponding to the first server record file.

[0038] The device identifier matching value, the transaction time difference sequence corresponding to the first server record file, the transaction frequency corresponding to the first server record file, and the normalized transaction amount sequence corresponding to the first server record file are determined as the transaction data derived features corresponding to the transaction data features.

[0039] In some embodiments, after determining the target verification result corresponding to the first server record file based on the first anomaly verification result and the second anomaly verification result, the method further includes:

[0040] If the target verification result indicates that the data is missing, obtain the third file identifier corresponding to the first server record file, generate a first data retransmission instruction based on the third file identifier, and send the first data retransmission instruction to the terminal device so that the terminal device retransmits the multiple transaction data corresponding to the third file identifier to the cloud.

[0041] If the target verification result indicates that the data is duplicated, a data deduplication instruction is generated and sent to the cloud so that the cloud can perform data deduplication on the first server record file to obtain a first deduplication server record file and store the first deduplication server record file.

[0042] If the target verification result is that the data is duplicated or the data is missing, the data in the first server record file is deduplicated to obtain a second deduplicated server record file, and the data in the second deduplicated server record file is matched with the first server record file to obtain the transaction data identifier corresponding to the duplicated transaction data in the first server record file.

[0043] Obtain the fourth file identifier corresponding to the first server record file, generate a second data retransmission instruction based on the fourth file identifier and the transaction data identifier, and send the second data retransmission instruction to the terminal device so that the terminal device performs data deduplication on multiple transaction data corresponding to the fourth file identifier based on the transaction data identifier, and resends the deduplicated multiple transaction data to the cloud.

[0044] If the target verification result indicates that the data is normal, a data storage instruction is generated and sent to the cloud so that the cloud stores the first server record file.

[0045] To achieve the above objectives, a second aspect of this application provides a transaction data anomaly verification device, the device comprising:

[0046] The file acquisition unit is used to acquire multiple first server record files received by the cloud within a first time interval before the current time, and to acquire multiple second server record files received by the cloud within a second time interval before the current time; wherein each first server record file and each second server record file is sent to the cloud by the terminal device, and each first server record file and each second server record file includes multiple transaction data; each second server record file has been pre-verified for anomalies.

[0047] The model invocation unit is used to invoke a pre-trained anomaly verification model for each of the first server record files, based on the first server record file and multiple second server record files, to perform anomaly verification; wherein, the anomaly verification model includes a feature extraction sub-model, a first sub-anomaly verification model, and a second sub-anomaly verification model.

[0048] The first anomaly verification unit is used to call the feature extraction sub-model to perform feature extraction for each first server record file based on the first server record file and multiple second server record files, so as to obtain transaction data features and transaction data derived features corresponding to the transaction data features.

[0049] The second anomaly verification unit is used to, for each of the first server record files, call the first sub-anomaly verification model to perform anomaly classification detection based on the transaction data features and the transaction data derived features to obtain a first anomaly verification result, and call the second sub-anomaly verification model to perform anomaly clustering detection based on the transaction data features and the transaction data derived features to obtain a second anomaly verification result.

[0050] The verification result unit is used to determine the target verification result corresponding to each of the first server record files based on the first abnormal verification result and the second abnormal verification result; wherein, the target verification result is used to indicate the verification status of the first server record file, and the verification status includes one or more of the following: data normal, data duplicate, and data missing.

[0051] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0052] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0053] The transaction data anomaly verification method, apparatus, electronic device, and storage medium proposed in this application acquire multiple first server record files received by the cloud within a first time interval before the current time, and multiple second server record files received by the cloud within a second time interval before the current time. Each first and second server record file is sent to the cloud by a terminal device and contains multiple transaction data sets; each second server record file undergoes pre-processing anomaly verification. For each first server record file, a pre-trained anomaly verification model is invoked based on the first server record file and the multiple second server record files. This anomaly verification model includes a feature extraction sub-model, a first sub-anomaly verification model, and a second sub-anomaly verification model. Next, for each first server record file, the feature extraction sub-model is invoked based on the first server record file and the multiple second server record files to extract features, obtaining transaction data features and their corresponding transaction data derived features. Then, for each first server record file, based on the extracted transaction data features and transaction data derived features, the first sub-anomaly verification model is invoked for anomaly classification detection to obtain a first anomaly verification result, and the second sub-anomaly verification model is invoked for anomaly clustering detection to obtain a second anomaly verification result. Finally, for each first server record file, the target verification result corresponding to the first server record file is determined based on the first anomaly verification result and the second anomaly verification result. The target verification result is used to indicate the verification status of the first server record file, and the verification status includes one or more of the following: data is normal, data is duplicated, and data is missing.

[0054] This application utilizes a pre-trained anomaly detection model's feature extraction sub-model to extract features from a first server record file and multiple second server record files, obtaining transaction data features and their derived features. Then, the first sub-anomaly detection model of the anomaly detection model is invoked to perform anomaly classification detection on the transaction data features and derived features, yielding a first anomaly detection result. Simultaneously, the second sub-anomaly detection model of the anomaly detection model is invoked to perform anomaly clustering detection on the transaction data features and derived features, yielding a second anomaly detection result. Finally, the verification status of the first server record file is determined based on the first and second anomaly detection results. This approach enables the analysis of multi-dimensional features of transaction data through two different techniques: classification detection and clustering detection. By combining the advantages of supervised and unsupervised learning, it reduces the risk that a single-type model cannot comprehensively capture various transaction anomalies and mitigates the potential for insufficient data utilization due to single-dimensional data. In other words, this application improves the accuracy of anomaly detection for transaction data. Attached Figure Description

[0055] Figure 1 This is a flowchart of the transaction data anomaly verification method provided in the embodiments of this application;

[0056] Figure 2 This is a flowchart of the training process of the anomaly verification model provided in the embodiments of this application;

[0057] Figure 3 This is a flowchart of the optimization training of the anomaly detection model provided in the embodiments of this application;

[0058] Figure 4 This is a schematic diagram of the transaction data anomaly verification device provided in the embodiments of this application;

[0059] Figure 5 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0061] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0063] First, let's analyze some of the terms used in this application:

[0064] Neural network models are computational models inspired by biological neural networks, used in machine learning and artificial intelligence. A neural network model consists of a hierarchical structure of multiple neurons, each connected to neurons in the next layer. These connections have weights, and through these weights and activation functions, the neural network can learn complex patterns and relationships in the input data.

[0065] Transaction data refers to the information records related to the transaction generated during the card transaction process. For example, transaction data may include information such as transaction time, transaction amount, transaction location (such as bus stop, subway station), transaction card number, and transaction type (such as entering the station, exiting the station, boarding the vehicle, alighting the vehicle, etc.).

[0066] The widespread application of transaction data anomaly verification methods provides technical support for the synchronization of transaction data in card transaction scenarios, ensuring the integrity and consistency of transactions. However, existing transaction data anomaly verification methods still have shortcomings in the accuracy of anomaly verification of transaction data received from the cloud, which greatly affects the reliability of transactions. For example, existing technologies typically use manual verification to check transaction data received from the cloud for anomalies. While this manual verification method can detect some abnormal transaction data, its accuracy and efficiency are both low. Furthermore, as the volume of transaction data continues to grow, this manual verification method can no longer meet the requirements for efficient and accurate anomaly verification. Alternatively, related technologies also use matching based on the transaction identifier or serial number corresponding to the transaction data for anomaly verification. However, this method can only determine whether the transaction data received from the cloud is duplicated to reduce the risk of duplicate transaction data, but it cannot dynamically assess the integrity of transaction data in real time, which may lead to misjudgment (such as normal data being intercepted due to mismatch) or missed judgment (such as duplicate data not being identified).

[0067] Furthermore, related technologies also employ redundant transmission or retry mechanisms to repeatedly send transaction data. This approach can only reduce the possibility of receiving incomplete transaction data in the cloud, thus mitigating the risk of missing transaction data. However, this approach may also lead to duplicate transaction data received by the cloud. For example, if a terminal does not receive a confirmation response from the server, it may send the same transaction record multiple times, and since the server lacks an efficient deduplication mechanism, this ultimately results in data redundancy. Based on this, embodiments of this application provide a transaction data anomaly verification method and apparatus, electronic device, and storage medium, aiming to improve the accuracy of anomaly verification for transaction data.

[0068] The transaction data anomaly verification method provided in this application relates to the field of public transportation technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the transaction data anomaly verification method, but is not limited to the above forms.

[0069] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0070] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0071] Figure 1 This is an optional flowchart of the transaction data anomaly verification method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105:

[0072] Step S101: Obtain multiple first server record files received by the cloud within a first time interval before the current time, and obtain multiple second server record files received by the cloud within a second time interval before the current time;

[0073] Step S102: For each first server record file, an anomaly verification model is called based on the first server record file and multiple second server record files to perform anomaly verification.

[0074] Step S103: For each first server record file, feature extraction is performed by calling the feature extraction sub-model based on the first server record file and multiple second server record files to obtain transaction data features and transaction data derived features corresponding to the transaction data features;

[0075] Step S104: For each first server record file, an anomaly classification detection is performed by calling the first sub-anomaly verification model based on transaction data features and transaction data derived features to obtain the first anomaly verification result, and anomaly clustering detection is performed by calling the second sub-anomaly verification model based on transaction data features and transaction data derived features to obtain the second anomaly verification result.

[0076] Step S105: For each first server record file, determine the target verification result corresponding to the first server record file based on the first anomaly verification result and the second anomaly verification result.

[0077] Steps S101 to S105 of this embodiment involve calling a pre-trained feature extraction sub-model of the anomaly verification model to extract features from the first server record file and multiple second server record files, obtaining transaction data features and their derived features. Then, the first sub-anomaly verification model of the anomaly verification model is called to perform anomaly classification detection on the transaction data features and derived features, obtaining a first anomaly verification result. Simultaneously, the second sub-anomaly verification model of the anomaly verification model is called to perform anomaly clustering detection on the transaction data features and derived features, obtaining a second anomaly verification result. Finally, the verification status of the first server record file is determined based on the first and second anomaly verification results. In this way, multi-dimensional features of transaction data can be analyzed using two different techniques: classification detection and clustering detection. This combines the advantages of supervised and unsupervised learning, reducing the risk that a single-type model cannot comprehensively capture all types of transaction anomalies, and mitigating the problem of insufficient data utilization that may result from single-dimensional data. Therefore, this application can improve the accuracy of anomaly verification for transaction data.

[0078] In step S101 of some embodiments, the current time may refer to the time point at which anomaly verification is performed on multiple first server record files received by the cloud within a first time interval. The first time interval may refer to a specific time range prior to the current time. For example, if the current time is 9:00, the first time interval may refer to one hour from 8:00 to 9:00; or, if the current time is 6:00, the first time interval may refer to two hours from 4:00 to 6:00. It is understood that the length of the first time interval can be adjusted according to actual needs. The first server record file may refer to a record file containing multiple transaction data received by the cloud within the first time interval.

[0079] It should be noted that each first server record file is sent to the cloud by the terminal device within the first time interval. The number of first server record files included in the first time interval and the amount of transaction data included in each first server record file can be adjusted according to actual needs.

[0080] The second time interval can refer to another specific time range preceding the current time. For example, if the current time is 9:00 AM on August 1st, and the first time interval is the one-hour period from 8:00 AM to 9:00 AM on August 1st, the second time interval could refer to the 24-hour period from 8:00 AM on July 31st to 8:00 AM on August 1st; or, if the current time is 6:00 AM on August 1st, and the first time interval is the two-hour period from 4:00 AM to 6:00 AM on August 1st, the second time interval could refer to the 12-hour period from 4:00 PM on July 31st to 4:00 AM on August 1st. It is understood that the length of the second time interval can be adjusted according to actual needs, but it must be ensured that the length of the second time interval is greater than the first time interval, and that the time period corresponding to the second time interval is after the first time interval. The second server record file can refer to a record file containing multiple transaction data received by the cloud within the second time interval.

[0081] It should be noted that each second server record file is sent to the cloud by the terminal device within the second time interval. The number of second server record files included in the second time interval and the amount of transaction data included in each second server record file can be adjusted according to actual needs. Each second server record file is pre-checked for anomalies.

[0082] In step S102 of some embodiments, the pre-trained anomaly detection model can refer to a trained neural network model with anomaly detection capabilities. Anomaly detection can refer to the process of calling the trained anomaly detection model to perform comprehensive analysis on the first server record file and multiple second server record files to detect abnormal transaction data in the first server record file.

[0083] It should be noted that the anomaly detection model includes a feature extraction sub-model, a first sub-anomaly detection model, and a second sub-anomaly detection model. The feature extraction sub-model can refer to the sub-model within the anomaly detection model that has feature extraction capabilities. For example, the feature extraction sub-model can be a Convolutional Neural Network (CNN) model or a Recurrent Neural Network (RNN) model, without specific limitations. The first sub-anomaly detection model can refer to the sub-model within the anomaly detection model that has anomaly classification and detection capabilities. For example, the first sub-anomaly detection model can be a Gradient Boosting Machine (GBM) model or a Light Gradient Boosting Machine (LightGBM) model, without specific limitations. The second sub-anomaly detection model can refer to the sub-model within the anomaly detection model that has anomaly clustering detection capabilities. For example, the second sub-anomaly detection model can be an Isolation Forest model or a Density-Based Spatial Clustering of Applications with Noise (DBSCAN) model, without specific limitations.

[0084] In step S103 of some embodiments, feature extraction can refer to the process of calling a feature extraction sub-model to comprehensively process the first server record file and multiple second server record files to identify and quantify the key attributes of the transaction data. Transaction data features can refer to the set of information describing the characteristics of the transaction data itself, obtained after calling the feature extraction sub-model to extract features from the first server record file and multiple second server record files. For example, transaction data features may include transaction time, transaction amount, transaction location, etc. It is understood that the information included in the transaction data features can be adjusted according to actual needs. Transaction data derived features corresponding to the transaction data features can refer to derived features obtained by further analysis or calculation based on the information included in the transaction data features. For example, if the transaction data features include transaction time, the transaction data derived features may refer to the difference in transaction time indicated by different transaction data; or, if the transaction data features include transaction amount, the transaction data derived features may refer to the transaction amount obtained after normalizing the transaction amounts indicated by different transaction data.

[0085] In step S104 of some embodiments, anomaly classification detection can refer to the process of calling a first sub-anomaly verification model to analyze and classify transaction data features and transaction data derived features to identify different anomaly types. The first anomaly verification result can refer to the anomaly classification status obtained after calling the first sub-anomaly verification model to perform anomaly classification detection on the transaction data features and transaction data derived features, used to indicate the anomaly classification status corresponding to the first server record file. For example, the first anomaly verification result can be one or more of data normal, data duplicate, and data missing. Anomaly clustering detection can refer to the process of calling a second sub-anomaly verification model to perform cluster analysis on the transaction data features and transaction data derived features to identify different anomaly types. The second anomaly verification result can refer to the anomaly clustering status obtained after calling the second sub-anomaly verification model to perform anomaly clustering detection on the transaction data features and transaction data derived features, used to indicate the anomaly clustering status corresponding to the first server record file. For example, the second anomaly verification result can be one or more of data normal, data duplicate, and data missing.

[0086] It should be noted that "data normal" can mean that multiple transaction data entries included in the first server's record file are complete. "Data duplicate" can mean that multiple transaction data entries included in the first server's record file are duplicated. "Data missing" can mean that multiple transaction data entries included in the first server's record file are missing.

[0087] In step S105 of some embodiments, the target verification result may refer to a result determined based on a combination of the first abnormal verification result and the second abnormal verification result, used to indicate the verification status of the first server record file. The verification status includes one or more of the following: data normal, data duplicate, and data missing. For example, if both the first abnormal verification result and the second abnormal verification result are data normal, the verification status of the first server record file can be determined to be data normal. Alternatively, if both the first abnormal verification result and the second abnormal verification result are data duplicate and data missing, the verification status of the first server record file can be determined to be data duplicate and data missing. Or, if both the first abnormal verification result and the second abnormal verification result are data duplicate and data missing, the verification status of the first server record file can be determined to be data duplicate and data missing.

[0088] It is understandable that the anomaly detection model can be pre-trained. The training process of the anomaly detection model is explained below. (Refer to...) Figure 2 , Figure 2 This is a flowchart of the training process of the anomaly verification model provided in the embodiments of this application. Figure 2 The steps may include, but are not limited to, steps S201 to S204:

[0089] Step S201: Obtain training sample data. The training sample data includes a first sample server record file received by the cloud within a first sample time interval before the historical time, a first classification label, and multiple second sample server record files received by the cloud within a second sample time interval before the sample time.

[0090] Step S202: Based on the first sample server record file and multiple second sample server record files, the feature extraction sub-model is called to extract features, thereby obtaining sample transaction data features and sample transaction data derived features corresponding to the sample transaction data features;

[0091] Step S203: Based on the characteristics of the sample transaction data and the derived characteristics of the sample transaction data, the first sub-anomaly verification model is called to perform anomaly classification detection to obtain the first sample anomaly verification result. Based on the characteristics of the sample transaction data and the derived characteristics of the sample transaction data, the second sub-anomaly verification model is called to perform anomaly clustering detection to obtain the second sample anomaly verification result.

[0092] Step S204: Adjust the model parameters of the anomaly verification model based on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, the anomaly verification result of the first sample, the anomaly verification result of the second sample, and the first classification label.

[0093] In step S201 of some embodiments, the training sample data may refer to the data set used to train the anomaly detection model. The training sample data includes a first sample server record file received by the cloud within a first sample time interval prior to the historical time, a first classification label, and multiple second sample server record files received by the cloud within a second sample time interval prior to the sample time. The first classification label may refer to a label used to identify the true anomaly type of the transaction data. The first classification label is used to guide the training of the first sub-anomaly detection model. The historical time may refer to a point in time prior to the current time in the cloud. The historical time is used to define the time range of the sample data. For example, if the current time is 9:00 AM on August 5th, the historical time could be three days before the current time, i.e., 9:00 AM on August 2nd; or, if the current time is 9:00 AM on August 5th, the historical time could be one month before the current time, i.e., 9:00 AM on July 5th, and so on.

[0094] The first sample time interval can refer to a time range prior to a historical date. For example, if the historical date is 9:00 AM on August 2nd, the first sample time interval could be one hour from 8:00 AM to 9:00 AM on August 2nd; or, if the historical date is 9:00 AM on July 5th, the first sample time interval could be two hours from 7:00 AM to 9:00 AM on July 5th. Understandably, the length of the first sample time interval can be adjusted according to actual needs. The first sample server record file can refer to a record file containing multiple transaction data received by the cloud within the first sample time interval.

[0095] It should be noted that each first sample server record file is sent to the cloud by the terminal device within the first sample time interval. The number of transaction data included in each first sample server record file can be adjusted according to actual needs, but it must be ensured that it is the same as the number of transaction data included in the first sample server record file.

[0096] The second sample time interval can refer to another specific time range preceding the historical time. For example, if the historical time is 9:00 AM on August 2nd, and the first sample time interval is the 1-hour period from 8:00 AM to 9:00 AM on August 2nd, the second sample time interval can refer to the 24-hour period from 8:00 AM on August 1st to 8:00 AM on August 2nd; or, if the historical time is 9:00 AM on July 5th, and the first sample time interval is the 2-hour period from 7:00 AM to 9:00 AM on July 5th, the second sample time interval can refer to the 12-hour period from 7:00 PM on July 4th to 7:00 AM on July 5th. It is understood that the length of the second sample time interval can be adjusted according to actual needs, but it must be ensured that the length of the second sample time interval is greater than the first sample time interval, and that the time period corresponding to the second sample time interval is after the first sample time interval. The second sample server record file can refer to a record file containing multiple transaction data received by the cloud within the second sample time interval.

[0097] It should be noted that each second sample server record file is sent to the cloud by the terminal device within the second sample time interval. The number of second sample server record files included within the second sample time interval and the amount of transaction data included in each second sample server record file can be adjusted according to actual needs, but it must be ensured that the amount of transaction data included in the second server record file is the same. Each second sample server record file undergoes pre-processing anomaly verification.

[0098] In step S202 of some embodiments, the sample transaction data features can refer to the set of information describing the characteristics of the transaction data itself, obtained after calling the feature extraction sub-model to extract features from the first sample server record file and multiple second sample server record files. For example, the sample transaction data features may include transaction time, transaction amount, transaction location, etc. It is understood that the information included in the sample transaction data features can be adjusted according to actual needs. The sample transaction data derived features corresponding to the sample transaction data features can refer to derived features derived from further analysis or calculation based on the information included in the sample transaction data features. For example, if the sample transaction data features include transaction time, the sample transaction data derived features may refer to the difference in transaction time indicated by different transaction data.

[0099] In steps S203 to S204 of some embodiments, the first sample anomaly verification result may refer to the anomaly classification status obtained after calling the first sub-anomaly verification model to perform anomaly classification detection on the sample transaction data features and sample transaction data derived features, used to indicate the anomaly classification status corresponding to the first sample server record file. For example, the first sample anomaly verification result may be one or more of data normal, data duplicate, and data missing. The second sample anomaly verification result may refer to the anomaly clustering status obtained after calling the second sub-anomaly verification model to perform anomaly clustering detection on the sample transaction data features and sample transaction data derived features, used to indicate the anomaly clustering status corresponding to the first sample server record file. For example, the second sample anomaly verification result may be one or more of data normal, data duplicate, and data missing.

[0100] It is understood that the model parameters of the anomaly detection model in this application embodiment can be adjusted by comprehensively considering the error in model prediction based on the first classification label, the anomaly detection result of the first sample, the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, and the anomaly detection result of the second sample, and updating the model parameters according to the error. Furthermore, this application embodiment can continuously iterate the parameters and train the anomaly detection model using multiple training sample data, thereby gradually improving the model's predictive performance.

[0101] In some embodiments, the model parameters of the anomaly detection model are adjusted based on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, the anomaly detection result of the first sample, the anomaly detection result of the second sample, and the classification label, including:

[0102] The anomaly detection result of the first sample and the first classification label are evaluated using classification indicators to obtain the first adjustment factor;

[0103] The clustering index is evaluated on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, and the anomaly verification results of the second sample to obtain the second adjustment factor;

[0104] The target adjustment factor is determined based on the preset weights, the first adjustment factor, and the second adjustment factor, and the model parameters of the anomaly verification model are adjusted according to the target adjustment factor.

[0105] In this embodiment, classification metric evaluation can refer to the process of quantifying the performance of the first sub-anomaly detection model by comparing the anomaly detection result of the first sample with the first classification label using a specific metric. The first adjustment factor can refer to the adjustment value used to optimize model parameters obtained after evaluating the classification metrics of the first sample anomaly detection result and the first classification label. For example, the first adjustment factor can be an adjustment value used to optimize model parameters obtained by calculating the precision or recall between the first sample anomaly detection result and the first classification label.

[0106] Clustering metric evaluation refers to the process of quantifying the performance of a second-sample anomaly detection model by comparing the characteristics of sample transaction data and derived features of the sample transaction data with the results of a second-sample anomaly detection using specific metrics. The second adjustment factor refers to the adjustment value used to optimize model parameters after evaluating the clustering metrics of the sample transaction data characteristics, derived features of the sample transaction data, and the results of the second-sample anomaly detection. For example, the second adjustment factor could be an adjustment value used to optimize model parameters by calculating the tightness or separation between the sample transaction data characteristics and the results of the second-sample anomaly detection.

[0107] The preset weight coefficient refers to a pre-defined weight value, including the weights corresponding to the first and second adjustment factors. For example, if the preset weight coefficient can be represented as 1, and the second adjustment factor is more important than the first adjustment factor, then the weight of the second adjustment factor can be set to 0.7, and the weight of the first adjustment factor can be set to 0.3; or, if the first adjustment factor is more important than the second adjustment factor, then the weight of the first adjustment factor can be set to 0.6, and the weight of the second adjustment factor can be set to 0.4. The target adjustment factor can refer to the comprehensive adjustment value obtained by weighting the first and second adjustment factors according to the preset weight coefficient. For example, if the first adjustment factor is 0.8, the second adjustment factor is 0.5, the weight of the first adjustment factor is set to 0.7, and the weight of the second adjustment factor is set to 0.3, then the target adjustment factor is 0.7 × 0.8 + 0.3 × 0.5 = 0.71.

[0108] It should be noted that adjusting the model parameters of the anomaly verification model can refer to determining the adjustment range of the model parameters based on the target adjustment factor, and then updating the model parameters of the anomaly verification model according to the adjustment range.

[0109] Understandably, this application determines the target adjustment factor of the anomaly detection model by combining a first adjustment factor obtained from evaluating the anomaly detection results of the first sample and the first classification label using classification indicators, a second adjustment factor obtained from evaluating the clustering indicators of the sample transaction data features, the derived features of the sample transaction data, and the anomaly detection results of the second sample, and preset weights. This allows for automatic adjustment of the model parameters based on the target adjustment factor, reducing the model's bias or omissions in identifying anomalies in transaction data, thereby improving the model's predictive performance.

[0110] In some embodiments, after adjusting the model parameters of the anomaly verification model based on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, the first sample anomaly verification result, the second sample anomaly verification result, and the first classification label, the method further includes:

[0111] By perturbing the features of the sample transaction data and the derived features of the sample transaction data, we obtain the adversarial sample transaction data features and the corresponding adversarial sample transaction data derived features.

[0112] Based on the characteristics of adversarial sample transaction data and the derived characteristics of adversarial sample transaction data, the first sub-anomaly verification model is called to perform anomaly classification detection, and the anomaly verification result of the third sample is obtained. Based on the characteristics of adversarial sample transaction data and the derived characteristics of adversarial sample transaction data, the second sub-anomaly verification model is called to perform anomaly clustering detection, and the anomaly verification result of the fourth sample is obtained.

[0113] The model parameters of the anomaly detection model are adjusted based on the characteristics of adversarial sample transaction data, the derived characteristics of adversarial sample transaction data, the anomaly detection results of the third sample, the anomaly detection results of the fourth sample, and the first classification label.

[0114] In this embodiment, feature perturbation refers to modifying or interfering with the features and derived features of sample transaction data through specific data preprocessing techniques to simulate network latency or clock skew that may occur during network transmission or system processing. Adversarial sample transaction data features refer to the set of information obtained after perturbing the features and derived features of sample transaction data, simulating the variation in transaction data caused by network latency or clock skew. For example, adversarial sample transaction data features can be obtained by simulating clock skew in the transaction time included in the sample transaction data features, thus obtaining a set of information describing the basic attributes of the transaction data. Adversarial sample transaction data derived features refer to derived features obtained by further analyzing or calculating the information included in the derived features of sample transaction data.

[0115] The third sample anomaly verification result can refer to the anomaly classification result obtained after calling the first sub-anomaly verification model to perform anomaly classification detection on the sample transaction data features and derived features, used to indicate the anomaly classification of the first server record file. The third sample anomaly verification result can be one or more of the following: normal data, duplicate data, and missing data. The fourth sample anomaly verification result can refer to the anomaly clustering result obtained after calling the second sub-anomaly verification model to perform anomaly clustering detection on the sample transaction data features and derived features, used to indicate the anomaly clustering of the first sample server record file. The fourth sample anomaly verification result can be one or more of the following: normal data, duplicate data, and missing data.

[0116] It should be noted that, in this embodiment of the application, the model parameters of the anomaly verification model can be adjusted by comprehensively considering the error of the model prediction based on the first classification label and the anomaly verification result of the third sample, the characteristics of the sample transaction data and the derived characteristics of the sample transaction data and the anomaly verification result of the fourth sample, and updating the model parameters based on the error.

[0117] Understandably, this application obtains adversarial sample transaction data features and adversarial sample transaction data derived features by perturbing the features of the sample transaction data features and features derived from the sample transaction data. These are then used to call the first sub-anomaly detection model for anomaly classification detection to obtain the third sample anomaly detection result, and the second sub-anomaly detection model for anomaly clustering detection to obtain the fourth sample anomaly detection result. Based on the obtained third and fourth sample anomaly detection results, the adversarial sample transaction data features, the adversarial sample transaction data derived features, and the first classification label, the model parameters of the anomaly detection model are adjusted. In this way, even when network latency or clock skew is lacking in the training sample data, simulating network latency or clock skew can reduce the model's sensitivity to network latency or clock skew, improve the model's generalization ability, and thus enhance the model's predictive performance.

[0118] In some embodiments, feature extraction is performed by calling a feature extraction sub-model based on a first server record file and multiple second server record files to obtain transaction data features and transaction data derived features corresponding to the transaction data features, including:

[0119] The feature extraction sub-model is invoked to extract the first device identifier corresponding to the terminal device that sent the first server record file, the first file identifier corresponding to the first server record file, and the transaction time, transaction amount, transaction identifier, and transaction type corresponding to the multiple transaction data included in the first server record file;

[0120] The first device identifier, the first file identifier, and the transaction time, transaction amount, transaction identifier, and transaction type corresponding to multiple transaction data included in the first server record file are determined as transaction data features;

[0121] The feature extraction sub-model is invoked to extract the second device identifier corresponding to the terminal device that sent multiple second server record files, as well as the transaction amount corresponding to the multiple transaction data included in each second server record file;

[0122] The first device identifier is matched with the second device identifier corresponding to multiple second server record files to obtain the device identifier matching value;

[0123] The transaction time difference sequence corresponding to the first server record file is determined based on the transaction time corresponding to the multiple transaction data included in the first server record file, and the transaction frequency corresponding to the first server record file is determined based on the multiple transaction data included in the first server record file and the first time interval.

[0124] A normalization factor is determined based on the transaction amounts corresponding to multiple transaction data included in multiple second server record files, and the transaction amounts corresponding to multiple transaction data included in the first server record file are normalized based on the normalization factor to obtain the normalized transaction amount sequence corresponding to the first server record file.

[0125] The device identifier matching value, the transaction time difference sequence corresponding to the first server record file, the transaction frequency corresponding to the first server record file, and the normalized transaction amount sequence corresponding to the first server record file are determined as the transaction data derived features corresponding to the transaction data features.

[0126] In this embodiment, the first device identifier can refer to the unique identification code of the terminal device that sends the first server record file. For example, the first device identifier can be the device's media access control address or device serial number, without specific limitations. The first file identifier can refer to the unique identification code of the first server record file. For example, the first file identifier can be the file's hash value or file number, without specific limitations. Transaction data characteristics can refer to a data set including the first device identifier, the first file identifier, and multiple transaction data items included in the first server record file, corresponding to transaction time, transaction amount, transaction identifier, and transaction type. The transaction time can refer to the point in time when the ride-hailing transaction occurs. The transaction amount can refer to the monetary value involved in the ride-hailing transaction. The transaction identifier can refer to a code or number used to identify the corresponding transaction data, such as a transaction card number. The transaction type can refer to the category of the ride-hailing transaction, such as entering the station, exiting the station, boarding the vehicle, or alighting from the vehicle.

[0127] The second device identifier can refer to a unique identification code of the terminal device that sends multiple second server record files. For example, the second device identifier can be the device's media access control address or device serial number; the specific method is not limited. Identifier matching can refer to the process of comparing and verifying the first device identifier against the second device identifiers corresponding to multiple second server record files to determine whether there is a correlation between the first and second device identifiers. The device identifier matching value can refer to a numerical value used to indicate the degree of matching between the first and second device identifiers. For example, the device identifier matching value can be 0 or 1, where 0 represents that the first and second device identifiers do not match, and 1 represents that the first and second device identifiers match. It is understood that the matching rules corresponding to the device identifier matching value can be adjusted according to actual needs.

[0128] The transaction time difference sequence is a sequence formed by calculating the time intervals between adjacent transactions based on the transaction times of multiple transaction data in the first server record file. For example, if the first server record file contains three transaction data points with transaction times of 9:00, 9:05, and 9:15, the transaction time difference sequence is 5 minutes and 10 minutes. Transaction frequency can refer to an indicator determined by the number of transactions within a first time interval based on multiple transaction data points in the first server record file. For example, if the first time interval is one hour and contains 10 transaction data points, the transaction frequency is 10 times / hour.

[0129] A normalization factor refers to a value determined based on transaction amounts from multiple second-server record files, used to standardize measurement. Normalization refers to scaling the transaction amounts in the first-server record file proportionally using a normalization factor to bring them within a specific range. A normalized transaction amount sequence refers to the sequence of values ​​obtained after normalizing the transaction amounts corresponding to multiple transaction data included in the first-server record file according to the amount normalization factor. For example, if the transaction amounts in the second-server record file range from 2 yuan to 20 yuan, the normalization factor could be 1 / 20. If the first-server record file includes three transaction data with transaction amounts of 5 yuan, 10 yuan, and 15 yuan respectively, the normalized transaction amount sequence would be 0.25, 0.5, and 0.75.

[0130] It is understood that this application defines transaction data features by identifying the first device identifier, the first file identifier, and the transaction time, transaction amount, transaction identifier, and transaction type corresponding to multiple transaction data included in the first server record file. It also defines the device identifier matching value, the transaction time difference sequence corresponding to the first server record file, the transaction frequency corresponding to the first server record file, and the normalized transaction amount sequence corresponding to the first server record file as transaction data derived features. In this way, it is possible to train the model by extracting multi-dimensional features from the transaction data, reducing the model's dependence on a single feature and thus improving the accuracy of the model in anomaly detection of transaction data.

[0131] In some embodiments, after determining the target verification result corresponding to the first server record file based on the first anomaly verification result and the second anomaly verification result, the method further includes:

[0132] The verification status distribution is determined based on the target verification results corresponding to multiple first server record files. The verification status distribution includes the first distribution corresponding to normal data, the second distribution corresponding to duplicate data, and the third distribution corresponding to missing data.

[0133] If the third distribution in the verification state distribution is the largest, obtain the network parameters of the terminal device at the current time; if the network parameters are greater than the preset network threshold, obtain the first server record file and the second classification label corresponding to the missing data in the target verification result, and construct the missing training sample data based on the first server record file, multiple second server record files, and the second classification label.

[0134] The anomaly detection model is optimized and trained based on the missing training sample data;

[0135] If the second distribution is the largest in the verification status distribution and the current time is within the preset peak period, obtain the first server record file and the third classification label corresponding to the data duplication in the target verification result, and construct duplicate training sample data based on the first server record file, multiple second server record files, and the third classification label;

[0136] The anomaly detection model is optimized and trained based on repeated training sample data.

[0137] In this embodiment, the verification status distribution quantity refers to a set of indicators reflecting the distribution of different types of verification statuses, obtained statistically from the target verification results corresponding to multiple first server record files. The verification status distribution quantity includes a first distribution quantity corresponding to normal data, a second distribution quantity corresponding to duplicate data, and a third distribution quantity corresponding to missing data. Specifically, the first distribution quantity can refer to the proportion or number of records in the multiple first server record files whose verification status is normal. The second distribution quantity can refer to the proportion or number of records in the multiple first server record files whose verification status is duplicate data. The third distribution quantity can refer to the proportion or number of records in the multiple first server record files whose verification status is missing data.

[0138] Please see Figure 3 , Figure 3 This is a flowchart illustrating the optimization training of the anomaly detection model provided in an embodiment of this application. For example... Figure 3 As shown, when the third distribution in the verification status distribution is the largest, firstly, the network parameters of the terminal device at the current time are obtained, and it is determined whether the network parameters are greater than a preset network threshold. Here, network parameters can refer to indicators used to measure the network status of the terminal device. For example, network parameters can be network latency, bandwidth utilization, packet loss rate, etc., without specific limitations. The preset network threshold can refer to a pre-set standard value used to judge network status. If the network parameters are greater than the preset network threshold, the first server record file and the second classification label corresponding to the missing data in the target verification result can be obtained. Here, the second classification label can refer to a label used to identify the actual missing transaction data. The second classification label is used to guide the training of the first sub-anomaly verification model. Then, missing training sample data can be constructed based on the first server record file, multiple second server record files, and the second classification label. Here, missing training sample data can refer to a data set containing data missing features used for optimizing the anomaly verification model training. If the network parameters are less than or equal to the preset network threshold, the anomaly verification model is not optimized.

[0139] It should be noted that the optimization training of the anomaly detection model can be achieved by iteratively training the model using pre-constructed training sample data, including a first server record file, multiple second server record files, and a second classification label. Furthermore, the embodiments of this application can continuously iterate the parameters and use multiple missing training sample data to perform iterative training on the anomaly detection model, thereby gradually improving the model's predictive performance.

[0140] like Figure 3As shown, when the second distribution value in the verification status distribution is the largest, the first step is to determine whether the current time is within a preset peak period. The preset peak period can refer to a time period with a large volume of transaction data or network traffic, set based on historical data and experience. For example, the preset peak period could be the morning or evening peak hours on weekdays, etc., without specific limitations. If the current time is within the preset peak period, the first server record file corresponding to the target verification result of data duplication and the third classification label can be obtained. The third classification label can refer to a label used to identify actual data duplication in transaction data. The third classification label is used to guide the training of the first sub-anomaly verification model. If the current time is not within the preset peak period, the anomaly verification model is not optimized during training.

[0141] It should be noted that the optimization training of the anomaly detection model can be achieved by iteratively training the model using pre-constructed training sample data, including a first server record file, multiple second server record files, and a third classification label. Furthermore, the embodiments of this application can continuously iterate the parameters and use multiple repeated training sample data to optimize the anomaly detection model, thereby gradually improving the model's predictive performance.

[0142] like Figure 3 As shown, when the first distribution quantity in the verification state distribution quantity is the largest, the abnormal verification model is not optimized and trained.

[0143] Understandably, this application determines the distribution of verification states by using the target verification results corresponding to multiple first server record files. When the distribution of target verification results indicating data duplication or data missing is the largest, the anomaly verification model is optimized and trained in conjunction with specific scenario conditions. For example, when the third distribution is the largest among the verification state distributions, if the network parameters exceed a preset network threshold, it indicates poor network conditions that may lead to data missing. In this case, optimizing the model by constructing missing training sample data can improve the model's ability to identify missing data under poor network conditions. When the second distribution is the largest among the verification state distributions and the current time is within a preset peak period, the transaction characteristics during peak periods may lead to data duplication. In this case, optimizing the model by constructing duplicate training sample data can enhance the model's ability to identify duplicate data during peak periods. This allows for targeted optimization of the model, avoiding excessive reliance on a single data type, further improving the model's generalization ability, and thus improving the accuracy of the model in anomaly verification of transaction data. Simultaneously, by optimizing the model training in conjunction with specific scenario conditions, the model becomes more aligned with actual application scenarios, improving its adaptability and predictive performance in different environments.

[0144] In some embodiments, after determining the target verification result corresponding to the first server record file based on the first anomaly verification result and the second anomaly verification result, the method further includes:

[0145] If the target verification result is that the data is missing, obtain the third file identifier corresponding to the first server record file, generate a first data retransmission instruction based on the third file identifier, and send the first data retransmission instruction to the terminal device so that the terminal device retransmits the multiple transaction data corresponding to the third file identifier to the cloud.

[0146] If the target verification result is that the data is duplicated, a data deduplication instruction is generated and sent to the cloud so that the cloud can perform data deduplication on the first server record file, obtain the first deduplication server record file, and store the first deduplication server record file;

[0147] If the target verification result is data duplication and data missing, the data in the first server record file is deduplicated to obtain the second deduplicated server record file. The data in the second deduplicated server record file is matched with the first server record file to obtain the transaction data identifier corresponding to the duplicate transaction data in the first server record file.

[0148] Obtain the fourth file identifier corresponding to the first server record file, generate a second data retransmission instruction based on the fourth file identifier and the transaction data identifier, and send the second data retransmission instruction to the terminal device so that the terminal device can perform data deduplication on multiple transaction data corresponding to the fourth file identifier based on the transaction data identifier, and resend the deduplicated multiple transaction data to the cloud.

[0149] If the target verification result is that the data is normal, a data storage instruction is generated and sent to the cloud so that the cloud stores the first server record file.

[0150] In this embodiment, when the target verification result corresponding to the first server record file indicates missing data, a third file identifier corresponding to the first server record file is obtained. The third file identifier may refer to an identification code associated with the first server record file corresponding to the missing data. Then, a first data retransmission instruction is generated based on the third file identifier and sent to the terminal device, causing the terminal device to resend multiple transaction data corresponding to the third file identifier to the cloud. The first data retransmission instruction may be a set of commands instructing the terminal device to resend multiple transaction data corresponding to the third file identifier to the cloud.

[0151] It should be noted that when the terminal device resends the multiple transaction data corresponding to the third file identifier to the cloud, it will perform data missing anomaly verification on the first server record file formed by the multiple transaction data received by the cloud again. After confirming that the anomaly verification result is that the data is normal, the cloud will store the first server record file to reduce storage errors caused by missing transaction data.

[0152] When the target verification result for the first server record file indicates duplicate data, a data deduplication instruction is generated and sent to the cloud. This instruction then causes the cloud to deduplicate the data in the first server record file, resulting in a first deduplicated server record file. The data deduplication instruction is a set of commands that instruct the cloud to perform data deduplication on the first server record file. Data deduplication refers to the process of identifying and removing duplicate transaction data from the first server record file. The first deduplicated server record file refers to the server record file containing non-duplicate transaction data obtained after deduplicating the first server record file whose target verification result indicates duplicate data. The cloud then stores the first deduplicated server record file.

[0153] When the target verification result for the first server record file is duplicate or missing data, the first server record file is first deduplicated to obtain a second deduplicated server record file. Then, data matching is performed between the second deduplicated server record file and the first server record file to obtain the transaction data identifiers corresponding to the duplicate transaction data in the first server record file. The second deduplicated server record file can refer to the server record file containing non-duplicate transaction data obtained after deduplicating the first server record file (where the target verification result is duplicate or missing data). Data matching refers to the process of identifying and locating duplicate transaction data by comparing the transaction data in the second deduplicated server record file and the first server record file. The transaction data identifier refers to the information used to identify duplicate transaction data obtained after matching the data in the second deduplicated server record file and the first server record file.

[0154] Then, the fourth file identifier corresponding to the first server record file is obtained, and a second data retransmission instruction is generated based on the fourth file identifier and the transaction data identifier. The fourth file identifier can refer to the identification code associated with the first server record file corresponding to missing or duplicate data. Finally, the second data retransmission instruction is sent to the terminal device, causing the terminal device to deduplicate the multiple transaction data corresponding to the fourth file identifier based on the transaction data identifier, and then resend the deduplicated transaction data to the cloud. The second data retransmission instruction can be a set of commands instructing the terminal device to deduplicate the multiple transaction data associated with the fourth file identifier based on the transaction data identifier, and then resend the deduplicated transaction data to the cloud.

[0155] It should be noted that when the terminal device resends the deduplicated transaction data to the cloud, it will perform data missing anomaly checks on the first server record file formed by the deduplicated transaction data received by the cloud. After confirming that the data is normal, the cloud will store the first server record file to avoid uploading duplicate transaction data again and reduce storage errors caused by missing transaction data and data duplication.

[0156] When the target verification result corresponding to the first server's record file is normal, a data storage instruction is generated and sent to the cloud so that the cloud can directly store the first server's record file. The data storage instruction can be a set of commands that instruct the cloud to directly store the first server's record file.

[0157] Understandably, this application can take corresponding processing measures based on the target verification result corresponding to the received first server record file. When the target verification result indicates missing data, the terminal device resends multiple transaction data corresponding to the third file identifier to the cloud; when the target verification result indicates duplicate data, the first server record file is deduplicated to obtain a first deduplicated server record file, which is then stored in the cloud; when the target verification result indicates both duplicate and missing data, first, data deduplication is performed to obtain a second deduplicated server record file, then the transaction data identifier of the duplicate transaction data is obtained through data matching, and finally, the terminal device resends multiple transaction data corresponding to the fourth file identifier after deduplication based on the transaction data identifier to the cloud. In this way, efficient data management and anomaly handling can be achieved based on the target verification result corresponding to the first server record file, reducing data redundancy and the possibility of storing missing transaction data, thereby improving the reliability of the model's anomaly verification of transaction data.

[0158] In some embodiments, it is considered that multiple first server record files received by the cloud within a first time interval may be missing. For example, the cloud should receive 10 first server record files within the first time interval, but may actually only receive 8. Based on this, embodiments of this application can also analyze the basic attributes corresponding to the multiple first server record files received within the first time interval. For example, the corresponding time interval sequence and the reception frequency of the first server record files can be calculated based on the time point at which each first server record file is received by the cloud. Furthermore, the file identifiers corresponding to the first server record files that were historically missing from the cloud can be comprehensively analyzed to determine the first server record files that were not uploaded. The missing first server record files are then re-uploaded for anomaly verification, and corresponding processing measures are taken based on the anomaly verification results. In this way, anomaly verification can be completed for all first server record files, avoiding inaccurate anomaly verification results due to incomplete data, thereby improving the reliability of anomaly verification and the accuracy of data processing.

[0159] Please see Figure 4 This application also provides a transaction data anomaly verification device, which can implement the above-mentioned transaction data anomaly verification method. The device includes:

[0160] The file acquisition unit 401 is used to acquire multiple first server record files received by the cloud within a first time interval before the current time, and to acquire multiple second server record files received by the cloud within a second time interval before the current time; wherein each first server record file and each second server record file are sent to the cloud by the terminal device, and each first server record file and each second server record file include multiple transaction data; each second server record file has been pre-verified for anomalies.

[0161] The model invocation unit 402 is used to invoke a pre-trained anomaly verification model for each first server record file, based on the first server record file and multiple second server record files, to perform anomaly verification; wherein, the anomaly verification model includes a feature extraction sub-model, a first sub-anomaly verification model, and a second sub-anomaly verification model;

[0162] The first anomaly verification unit 403 is used to perform feature extraction for each first server record file by calling the feature extraction sub-model based on the first server record file and multiple second server record files, so as to obtain transaction data features and transaction data derived features corresponding to the transaction data features.

[0163] The second anomaly verification unit 404 is used to, for each first server record file, call the first sub-anomaly verification model to perform anomaly classification detection based on transaction data features and transaction data derived features to obtain the first anomaly verification result, and call the second sub-anomaly verification model to perform anomaly clustering detection based on transaction data features and transaction data derived features to obtain the second anomaly verification result.

[0164] The verification result unit 405 is used to determine the target verification result corresponding to each first server record file based on the first abnormal verification result and the second abnormal verification result; wherein, the target verification result is used to indicate the verification status of the first server record file, and the verification status includes one or more of the following: data is normal, data is duplicated, and data is missing.

[0165] The specific implementation of the transaction data anomaly verification device is basically the same as the specific implementation of the transaction data anomaly verification method described above, and will not be repeated here.

[0166] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described transaction data anomaly verification method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0167] Please see Figure 5 , Figure 5 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0168] The processor 501 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0169] The memory 502 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 502 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 502 and is called and executed by the processor 501 using the transaction data anomaly verification method of the embodiments of this application.

[0170] The input / output interface 503 is used to implement information input and output;

[0171] The communication interface 504 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0172] Bus 505 transmits information between various components of the device (e.g., processor 501, memory 502, input / output interface 503, and communication interface 504);

[0173] The processor 501, memory 502, input / output interface 503, and communication interface 504 are connected to each other within the device via bus 505.

[0174] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described transaction data anomaly verification method.

[0175] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0176] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0177] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0179] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0180] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0181] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0182] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0183] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0184] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0185] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0186] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for verifying anomalies in transaction data, characterized in that, The method includes: The system acquires multiple first server record files received by the cloud within a first time interval before the current time, and acquires multiple second server record files received by the cloud within a second time interval before the current time; wherein each first server record file and each second server record file is sent to the cloud by the terminal device, and each first server record file and each second server record file includes multiple transaction data; each second server record file has been pre-verified for anomalies. For each of the first server record files, an anomaly verification model is invoked based on the first server record file and multiple second server record files to perform anomaly verification; wherein, the anomaly verification model includes a feature extraction sub-model, a first sub-anomaly verification model and a second sub-anomaly verification model; For each first server record file, the feature extraction sub-model is invoked based on the first server record file and multiple second server record files to perform feature extraction, thereby obtaining transaction data features and transaction data derived features corresponding to the transaction data features; For each of the first server record files, the first sub-anomaly verification model is invoked based on the transaction data features and the transaction data derived features to perform anomaly classification detection, thereby obtaining a first anomaly verification result. Then, the second sub-anomaly verification model is invoked based on the transaction data features and the transaction data derived features to perform anomaly clustering detection, thereby obtaining a second anomaly verification result. For each of the first server record files, a target verification result is determined based on the first anomaly verification result and the second anomaly verification result; wherein, the target verification result is used to indicate the verification status of the first server record file, and the verification status includes one or more of the following: data normal, data duplicate, and data missing.

2. The method according to claim 1, characterized in that, The training method for the anomaly detection model includes: The training sample data includes a first sample server record file received by the cloud within a first sample time interval prior to the historical time, a first classification label, and multiple second sample server record files received by the cloud within a second sample time interval prior to the sample time. The first classification label is used to guide the training of the first sub-anomaly verification model. Each first sample server record file and each second sample server record file are sent from the terminal device to the cloud, and each first sample server record file and each second sample server record file includes multiple transaction data. Each second sample server record file has undergone anomaly verification beforehand. Based on the first sample server record file and multiple second sample server record files, the feature extraction sub-model is invoked to perform feature extraction, thereby obtaining sample transaction data features and sample transaction data derived features corresponding to the sample transaction data features; Based on the characteristics of the sample transaction data and the derived characteristics of the sample transaction data, the first sub-anomaly verification model is invoked to perform anomaly classification detection to obtain the first sample anomaly verification result. Based on the characteristics of the sample transaction data and the derived characteristics of the sample transaction data, the second sub-anomaly verification model is invoked to perform anomaly clustering detection to obtain the second sample anomaly verification result. The model parameters of the anomaly verification model are adjusted based on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, the first sample anomaly verification result, the second sample anomaly verification result, and the first classification label.

3. The method according to claim 2, characterized in that, The step of adjusting the model parameters of the anomaly verification model based on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, the first sample anomaly verification result, the second sample anomaly verification result, and the first classification label includes: The anomaly detection result of the first sample and the first classification label are evaluated by classification indicators to obtain the first adjustment factor; The clustering index is evaluated on the characteristics of the sample transaction data, the derived characteristics of the sample transaction data, and the anomaly verification results of the second sample to obtain the second adjustment factor; A target adjustment factor is determined based on the preset weights, the first adjustment factor, and the second adjustment factor, and the model parameters of the anomaly verification model are adjusted according to the target adjustment factor.

4. The method according to claim 2 or 3, characterized in that, After adjusting the model parameters of the anomaly verification model based on the sample transaction data features, the sample transaction data derived features, the first sample anomaly verification result, the second sample anomaly verification result, and the first classification label, the method further includes: The sample transaction data features and the sample transaction data derived features are perturbed to obtain the adversarial sample transaction data features and the corresponding adversarial sample transaction data derived features; Based on the adversarial sample transaction data features and the adversarial sample transaction data derived features, the first sub-anomaly verification model is invoked to perform anomaly classification detection, resulting in the third sample anomaly verification result. Based on the adversarial sample transaction data features and the adversarial sample transaction data derived features, the second sub-anomaly verification model is invoked to perform anomaly clustering detection, resulting in the fourth sample anomaly verification result. The model parameters of the anomaly verification model are adjusted based on the adversarial sample transaction data features, the adversarial sample transaction data derived features, the third sample anomaly verification result, the fourth sample anomaly verification result, and the first classification label.

5. The method according to claim 1, characterized in that, After determining the target verification result corresponding to the first server record file based on the first anomaly verification result and the second anomaly verification result, the method further includes: The verification status distribution is determined based on the target verification results corresponding to multiple first server record files. The verification status distribution includes a first distribution corresponding to normal data, a second distribution corresponding to duplicate data, and a third distribution corresponding to missing data. If the third distribution among the verification state distributions is the largest, the network parameters of the terminal device at the current time are obtained; if the network parameters are greater than a preset network threshold, the first server record file and the second classification label corresponding to the missing data are obtained as the target verification result, and missing training sample data are constructed based on the first server record file, multiple second server record files, and the second classification label. The anomaly detection model is optimized and trained based on the missing training sample data. If the second distribution quantity is the largest among the verification status distribution quantities, and the current time is within a preset peak period, the first server record file and the third classification label corresponding to the data duplication in the target verification result are obtained, and duplicate training sample data are constructed based on the first server record file, multiple second server record files, and the third classification label; The anomaly detection model is optimized and trained based on the repeated training sample data.

6. The method according to claim 1, characterized in that, The step of calling the feature extraction sub-model based on the first server record file and multiple second server record files to perform feature extraction, obtaining transaction data features and corresponding transaction data derived features, includes: The feature extraction sub-model is invoked to extract the first device identifier corresponding to the terminal device that sent the first server record file, the first file identifier corresponding to the first server record file, and the transaction time, transaction amount, transaction identifier, and transaction type corresponding to the multiple transaction data included in the first server record file; The transaction data features are defined as the first device identifier, the first file identifier, and the transaction time, the transaction amount, the transaction identifier, and the transaction type corresponding to the multiple transaction data included in the first server record file. The feature extraction sub-model is invoked to extract the second device identifier corresponding to the terminal device that sent multiple second server record files and the transaction amount corresponding to the multiple transaction data included in each second server record file; The first device identifier is matched with the second device identifier corresponding to multiple second server record files to obtain a device identifier matching value; The transaction time difference sequence corresponding to the first server record file is determined based on the transaction time corresponding to the multiple transaction data included in the first server record file, and the transaction frequency corresponding to the first server record file is determined based on the multiple transaction data included in the first server record file and the first time interval. A normalization factor is determined based on the transaction amount corresponding to the multiple transaction data included in the multiple second server record files, and the transaction amount corresponding to the multiple transaction data included in the first server record file is normalized based on the normalization factor to obtain the normalized transaction amount sequence corresponding to the first server record file. The device identifier matching value, the transaction time difference sequence corresponding to the first server record file, the transaction frequency corresponding to the first server record file, and the normalized transaction amount sequence corresponding to the first server record file are determined as the transaction data derived features corresponding to the transaction data features.

7. The method according to claim 1, characterized in that, After determining the target verification result corresponding to the first server record file based on the first anomaly verification result and the second anomaly verification result, the method further includes: If the target verification result indicates that the data is missing, obtain the third file identifier corresponding to the first server record file, generate a first data retransmission instruction based on the third file identifier, and send the first data retransmission instruction to the terminal device so that the terminal device retransmits the multiple transaction data corresponding to the third file identifier to the cloud. If the target verification result indicates that the data is duplicated, a data deduplication instruction is generated and sent to the cloud so that the cloud can perform data deduplication on the first server record file to obtain a first deduplication server record file and store the first deduplication server record file. If the target verification result is that the data is duplicated or the data is missing, the data in the first server record file is deduplicated to obtain a second deduplicated server record file, and the data in the second deduplicated server record file is matched with the first server record file to obtain the transaction data identifier corresponding to the duplicated transaction data in the first server record file. Obtain the fourth file identifier corresponding to the first server record file, generate a second data retransmission instruction based on the fourth file identifier and the transaction data identifier, and send the second data retransmission instruction to the terminal device so that the terminal device performs data deduplication on multiple transaction data corresponding to the fourth file identifier based on the transaction data identifier, and resends the deduplicated multiple transaction data to the cloud. If the target verification result indicates that the data is normal, a data storage instruction is generated and sent to the cloud so that the cloud stores the first server record file.

8. A transaction data anomaly verification device, characterized in that, The device includes: The file acquisition unit is used to acquire multiple first server record files received by the cloud within a first time interval before the current time, and to acquire multiple second server record files received by the cloud within a second time interval before the current time; wherein each first server record file and each second server record file is sent to the cloud by the terminal device, and each first server record file and each second server record file includes multiple transaction data; each second server record file has been pre-verified for anomalies. The model invocation unit is used to invoke a pre-trained anomaly verification model for each of the first server record files, based on the first server record file and multiple second server record files, to perform anomaly verification; wherein, the anomaly verification model includes a feature extraction sub-model, a first sub-anomaly verification model, and a second sub-anomaly verification model. The first anomaly verification unit is used to call the feature extraction sub-model to perform feature extraction for each first server record file based on the first server record file and multiple second server record files, so as to obtain transaction data features and transaction data derived features corresponding to the transaction data features. The second anomaly verification unit is used to, for each of the first server record files, call the first sub-anomaly verification model to perform anomaly classification detection based on the transaction data features and the transaction data derived features to obtain a first anomaly verification result, and call the second sub-anomaly verification model to perform anomaly clustering detection based on the transaction data features and the transaction data derived features to obtain a second anomaly verification result. The verification result unit is used to determine the target verification result corresponding to each of the first server record files based on the first abnormal verification result and the second abnormal verification result; wherein, the target verification result is used to indicate the verification status of the first server record file, and the verification status includes one or more of the following: data normal, data duplicate, and data missing.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the transaction data anomaly verification method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the transaction data anomaly verification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Abnormal behavior detection method and device and verification system

    CN110177108A

  • Data anomaly detection method and related device

    CN114077872A