Data detection method, nonvolatile storage medium and electronic equipment

By receiving data detection requests in banking business data detection, determining potential anomalies and customizing anomaly detection model, the problem of low detection efficiency in traditional methods is solved, and more efficient and accurate abnormality detection is achieved.

CN120123940APending Publication Date: 2025-06-10INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510212697.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When facing complex and large amounts of banking data, traditional abnormality detection methods are inefficient and costly, and fail to effectively solve the problem of whether abnormalities are found in the detection data.

Method used

Provide a data detection method, by receiving data detection requests, determining potential abnormalities, and customizing an abnormality detection model based on these abnormalities, using the target model architecture and hyperparameters for training, and then performing abnormality detection.

Benefits of technology

By targeted adjustment of the model architecture and hyperparameters, the system can more accurately identify potential abnormalities, reduce the false alarm rate and missed alarm rate, improve detection efficiency, and shorten the abnormal detection cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123940A_ABST
    Figure CN120123940A_ABST
Patent Text Reader

Abstract

The invention discloses a data detection method, a nonvolatile storage medium and electronic equipment. The method comprises the following steps: receiving a data detection request; in response to the data detection request, determining a potential abnormal item corresponding to the to-be-detected data; an anomaly detection model corresponding to the potential anomaly item is determined, the anomaly detection model is obtained by training an initial detection model provided with a target model architecture by adopting a target hyper-parameter, the target model architecture is obtained according to the potential anomaly item, and the target hyper-parameter is obtained according to the potential anomaly item; and determining an anomaly detection result corresponding to the potential anomaly item according to the to-be-detected data and an anomaly detection model. According to the invention, the technical problem that the detection efficiency is low when the data is detected to be abnormal or not in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing, and in particular, to a data detection method, a non-volatile storage medium, and an electronic device. Background Art

[0002] With the continuous development of the industry, anomaly detection, as an important part of ensuring business compliance and risk control, has received increasing attention. Traditional anomaly detection methods often rely on rules and statistical techniques, depending on expert knowledge and anomaly patterns in historical data. However, with the increasing complexity of banking operations and the explosion of data volume, these traditional methods have gradually revealed limitations, such as low efficiency and high cost in traditional detection methods.

[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present invention provide a data detection method, a non-volatile storage medium, and an electronic device to at least solve the technical problem of low detection efficiency when detecting whether data is abnormal in related technologies.

[0005] According to one aspect of the embodiments of the present invention, a data detection method is provided, including: receiving a data detection request, where the data detection request carries data to be detected, and the data detection request is used to detect whether there is an anomaly in the data to be detected; in response to the data detection request, determining a potential anomaly item corresponding to the data to be detected, where the potential anomaly item is an anomaly item whose probability of occurrence of the corresponding anomaly in the data to be detected is greater than a predetermined threshold; determining an anomaly detection model corresponding to the potential anomaly item, where the anomaly detection model is obtained by training an initial detection model with a target model architecture using target hyperparameters, the target model architecture is obtained based on the potential anomaly item, and the target hyperparameters are obtained based on the potential anomaly item; and determining an anomaly detection result corresponding to the potential anomaly item according to the data to be detected and the anomaly detection model.

[0006] Optionally, the determining an anomaly detection result corresponding to the potential anomaly item according to the data to be detected and the anomaly detection model includes: determining a format processing method corresponding to the potential anomaly item; processing the data to be detected according to the format processing method to obtain updated detection data; and inputting the updated detection data into the anomaly detection model to obtain a detection result corresponding to the potential anomaly item.

[0007] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining the detection parameters corresponding to the potential anomaly item, where the detection parameters include a detection complexity index and a detection accuracy parameter; determining the target model architecture and the target hyperparameters according to the detection parameters.

[0008] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining a plurality of initial candidate models corresponding to the potential anomaly item, where different model architectures are set for the plurality of initial candidate models; determining the test parameters of the plurality of initial candidate models under a plurality of test metrics; determining the initial detection model from the plurality of initial candidate models according to the test parameters of the plurality of initial candidate models under the plurality of test metrics.

[0009] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining a predetermined hyperparameter range according to the potential anomaly item; selecting a plurality of first test hyperparameters from the predetermined hyperparameter range using a first span; determining the model quality indexes corresponding to a plurality of alternative detection models, where the plurality of alternative detection models correspond to the plurality of first test hyperparameters one by one; determining a plurality of restricted hyperparameters from the plurality of first test hyperparameters according to the plurality of model quality indexes; determining an updated hyperparameter range according to the plurality of restricted hyperparameters, and determining second test hyperparameters from the updated hyperparameter range using a second span in a loop until the target hyperparameters with the highest corresponding model quality index are determined.

[0010] Optionally, the determining of the potential anomaly item corresponding to the data to be detected in response to the data detection request includes: in response to the data detection request, determining the detection type corresponding to the data to be detected; determining a predetermined template corresponding to the detection type; performing pattern processing of the predetermined template on the data to be detected to obtain target template data; determining the potential anomaly item corresponding to the data to be detected according to the target template data.

[0011] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: when the data detection request carries a predetermined scenario, determining the sample data corresponding to the predetermined scenario; training the initial detection model according to the sample data to obtain the anomaly detection model.

[0012] According to one aspect of an embodiment of the present invention, a data detection device is provided, including: a receiving module, configured to receive a data detection request, wherein the data detection request carries data to be detected, and the data detection request is used to detect whether there is an abnormality in the data to be detected; a first determination module, configured to, in response to the data detection request, determine a potential abnormal item corresponding to the data to be detected, wherein the potential abnormal item is an abnormal item whose probability of occurrence of the corresponding abnormality in the data to be detected is greater than a predetermined threshold; a second determination module, configured to determine an anomaly detection model corresponding to the potential abnormal item, wherein the anomaly detection model is obtained by training an initial detection model with a target model architecture using target hyperparameters, the target model architecture is obtained based on the potential abnormal item, and the target hyperparameters are obtained based on the potential abnormal item; a third determination module, configured to determine an anomaly detection result corresponding to the potential abnormal item according to the data to be detected and the anomaly detection model.

[0013] According to one aspect of an embodiment of the present invention, a non-volatile storage medium is provided, the non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the data detection method described in any one of the above.

[0014] According to one aspect of an embodiment of the present invention, an electronic device is provided, including one or more processors and a memory, the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the data detection method described in any one of the above.

[0015] According to one aspect of an embodiment of the present invention, a computer program product is provided, including computer instructions, and the computer instructions are executed by a processor to implement the data detection method described in any one of the above.

[0016] In an embodiment of the present invention, a data detection request is received, where the data detection request carries the data to be detected, and the data detection request is used to detect whether there is an abnormality in the data to be detected. In response to the data detection request, potential abnormal items corresponding to the data to be detected are determined, where the potential abnormal items are abnormal items for which the probability of the data to be detected having the corresponding abnormality is greater than a predetermined threshold. An abnormal detection model corresponding to the potential abnormal items is determined, where the abnormal detection model is obtained by training an initial detection model with a target model architecture using target hyperparameters. The target model architecture is obtained based on the potential abnormal items, and the target hyperparameters are obtained based on the potential abnormal items. Based on the data to be detected and the abnormal detection model, an abnormal detection result corresponding to the potential abnormal items is determined. It can be seen that the abnormal detection is performed using the abnormal detection model corresponding to the abnormal items, and moreover, the abnormal detection model is customized for the model architecture and hyperparameters based on the potential abnormal items. By specifically adjusting the model architecture and hyperparameters, the system can more accurately identify potential abnormal items, reduce the false alarm rate and the missed alarm rate, improve the detection efficiency, and the customized model can process data more quickly, shortening the cycle of abnormal detection, thereby solving the technical problem of low detection efficiency when detecting whether there is an abnormality in the data in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0018] Figure 1 is a flowchart of a data detection method provided according to an embodiment of this application;

[0019] Figure 2 is a structural block diagram of a data detection method provided according to an optional embodiment of the present invention;

[0020] Figure 3 is a structural block diagram of a data detection device provided according to an embodiment of this application;

[0021] Figure 4 is a schematic diagram of an electronic device provided according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0023] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0024] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 is a flowchart of a data detection method provided according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0025] Step S101, receiving a data detection request, wherein the data detection request carries data to be detected, and the data detection request is used to detect whether there are anomalies in the data to be detected;

[0026] In step S101 provided in the present application, a data detection request is received.

[0027] Among them, a data detection request is involved. This can be a request initiated by a system or a user, and a specific algorithm or model can be used to analyze a certain data set to determine whether the data therein is abnormal.

[0028] Among them, the data to be detected is involved, which refers to the data submitted for analysis in the data detection request. These data can be original, unprocessed transaction records, or data sets that have been preliminarily cleaned and preprocessed. They can also be behavioral transaction records, account information, audit data, etc., to determine whether there are anomalies in terms of numerical values and formats of these data. In a banking scenario, the data to be detected may include account activities, loan records, fund flow situations, etc., and it is necessary to check whether there are anomalies or violations.

[0029] Among them, anomalies are involved. In the scenarios involved in the present application, anomalies can refer to data points or data patterns that deviate from the normal or expected range. Anomalies may be caused by data errors, system failures, improper activities, illegal operations, or special events. Anomalies may manifest as unusually large transactions, frequent account activities, loan applications that do not conform to historical behaviors, etc.

[0030] Step S102, in response to the data detection request, determining potential anomaly items corresponding to the data to be detected, wherein the potential anomaly items are anomaly items for which the probability of the data to be detected having corresponding anomalies is greater than a predetermined threshold;

[0031] In step S102 provided in the present application, in response to the data detection request, potential anomaly items corresponding to the data to be detected are determined.

[0032] Among them, potential abnormal items are involved. This refers to data points or patterns that may be abnormal identified by the system during the data detection process. For example, for a bank's bank card application form, potential abnormal items may include abnormal items regarding whether the format and filling requirements of the bank card form are correct, abnormal items regarding whether the filled content is true, and abnormal items regarding whether the filling user is an abnormal user applying for a bank card. Specific abnormal items need to be adaptively set according to the actual application and scenario.

[0033] Among them, a predetermined threshold is involved. This is a preset value used to determine whether the corresponding item should be regarded as a potential abnormal item. When the predetermined threshold is set to 0, it means that the possible abnormality that may occur in the data to be detected is the abnormal item. At this time, it is possible to determine to the greatest extent whether there are abnormalities in all aspects of the data to be detected. It should also be noted that the setting of the threshold can also consider the sensitivity and specificity of the detection, as well as the trade-off between false positives and false negatives, and be adaptively set.

[0034] Step S103, determine an anomaly detection model corresponding to the potential abnormal item. Among them, the anomaly detection model is obtained by training an initial detection model with a target model architecture using target hyperparameters. The target model architecture is obtained based on the potential abnormal item, and the target hyperparameters are obtained based on the potential abnormal item;

[0035] In step S103 provided in this application, an anomaly detection model corresponding to the potential abnormal item is determined.

[0036] Among them, an anomaly detection model is involved. The anomaly detection model is a statistical or machine learning model specifically used to identify anomalies or irregular patterns in data. For different abnormal items, there are different anomaly detection models. In the bank scenario, the anomaly detection model can detect corresponding anomalies according to the corresponding abnormal items. For example, it can be used to detect transaction behaviors, account activities, loan applications, etc. to discover possible risks or violations, as well as anomalies such as formats.

[0037] Among them, the target model architecture is involved. The target model architecture refers to the model structure designed according to specific anomaly detection requirements and data characteristics. The model architecture may include key parameters such as the type of the model (such as neural network, decision tree, support vector machine, etc.), the number of layers, the number of nodes, activation functions, etc. The selection of these parameters will affect the performance and learning ability of the model. A good architecture can better capture the patterns in the data. Different architectures may be required for different tasks (such as image recognition, natural language processing, sequence prediction, etc.). The setting of the architecture can include the depth and width of the network. The depth and width of the network affect the expressive ability and computational complexity of the model. Deeper and wider networks can usually learn more complex patterns, but may also face the problem of overfitting. The setting of the architecture can also include connection methods, such as fully connected layers, convolutional layers, skip connections, etc. Different connection methods are suitable for processing different types of data and tasks. In order to adapt to better detection of abnormal items.

[0038] Among them, the target hyperparameters are involved. Hyperparameters are parameters set before model training and have an important impact on the performance of the model. The target hyperparameters refer to the hyperparameter settings customized for the anomaly detection model according to the characteristics and types of potential abnormal items, such as learning rate, batch size, regularization coefficient, etc. The selection of these parameters should be able to maximize the ability of the model to identify potential abnormal items.

[0039] The strategy of customizing the model architecture and hyperparameters based on potential abnormal items can improve the accuracy of anomaly detection. By specifically adjusting the model architecture and hyperparameters, the system can more accurately identify potential abnormal items, reducing the false positive rate and false negative rate. And it can enhance the generalization ability of the model. The optimized model architecture and hyperparameters enable the model to not only perform well on the training data, but also make accurate judgments on unseen potential abnormal items, improving the generalization ability of the model. Moreover, it can improve the detection efficiency. The customized model can process data more quickly, shortening the cycle of anomaly detection.

[0040] Step S104: Determine the anomaly detection result corresponding to the potential abnormal item according to the data to be detected and the anomaly detection model.

[0041] In step S104 provided by this application, the anomaly detection result corresponding to the potential abnormal item is determined.

[0042] Among them, the anomaly detection result is involved. This refers to the output obtained after the anomaly detection model analyzes a specific potential abnormal item, usually including the judgment on whether the data to be detected is abnormal, and can also include the evaluation of the degree of abnormality. The result may be expressed as a probability score, anomaly level, anomaly type, etc.

[0043] Through the above steps, a data detection request is received, where the data detection request carries the data to be detected, and the data detection request is used to detect whether there are abnormalities in the data to be detected. In response to the data detection request, potential abnormal items corresponding to the data to be detected are determined, where the potential abnormal items are abnormal items for which the probability of the data to be detected having corresponding abnormalities is greater than a predetermined threshold. An abnormal detection model corresponding to the potential abnormal items is determined, where the abnormal detection model is obtained by training an initial detection model with a target model architecture using target hyperparameters. The target model architecture is obtained based on the potential abnormal items, and the target hyperparameters are obtained based on the potential abnormal items. Based on the data to be detected and the abnormal detection model, an abnormal detection result corresponding to the potential abnormal items is determined. It can be seen that an abnormal detection model corresponding to the abnormal items is used for abnormal detection, and moreover, the abnormal detection model customizes the model architecture and hyperparameters based on the potential abnormal items. By specifically adjusting the model architecture and hyperparameters, the system can more accurately identify potential abnormal items, reduce the false alarm rate and missed alarm rate, improve the detection efficiency, and the customized model can process data more quickly, shortening the abnormal detection cycle, thereby solving the technical problem of low detection efficiency when detecting whether there are abnormalities in the data in the related art.

[0044] As an optional embodiment, determining an abnormal detection result corresponding to the potential abnormal items based on the data to be detected and the abnormal detection model includes: determining a format processing method corresponding to the potential abnormal items; processing the data to be detected according to the format processing method to obtain updated detection data; and inputting the updated detection data into the abnormal detection model to obtain a detection result corresponding to the potential abnormal items.

[0045] In this embodiment, further operations for determining an abnormal detection result corresponding to the potential abnormal items based on the data to be detected and the abnormal detection model are described.

[0046] Among them, a format processing method is involved, and the format processing method is a processing strategy or method for data format abnormalities. This may include data conversion, format verification, template matching, standardization, etc., which can adjust the data to conform to the expected format for subsequent analysis.

[0047] That is, when there are multiple potential abnormal items in the data to be detected, such as potential abnormal items for detecting the format and potential abnormal items for detecting data values. Masking processing of numerical values or texts can be performed on the potential abnormal items for detecting the format to avoid interference from hybrid information and increase in computational complexity. De-templating, numericalization, normalization, etc. can be performed on the potential abnormal items for detecting data values to avoid interference from hybrid information and increase in computational complexity, and organize them into a language that the computer can understand, which can improve the processing efficiency.

[0048] Among them, the updated detection data, after being processed in terms of format, is more suitable for being input into the anomaly detection model for in-depth analysis.

[0049] In this step, determine the format processing method corresponding to the potential anomaly items, and process the data to be detected according to this method to ensure data quality and improve the effect of the anomaly detection model. When processing the potential anomaly items with detected formats, the system will identify potential format anomaly items through the data preprocessing stage. For example, in a bank card application form, there may be situations where the date format filled in by the user is incorrect, required fields are missing, or the number formats are inconsistent. For the identified format anomaly items, the most suitable processing method needs to be determined. This may include automatically correcting format errors (such as converting dates to a unified format through regular expressions), filling in missing values (such as filling with average values or predicted values), excluding data points that do not meet the format requirements, etc., to ensure data consistency and integrity. That is, when outputting the results, format anomalies will be output, but the corrected versions will be provided for determination. The updated detection data better meets the input requirements of the model and can improve the accuracy and reliability of the anomaly detection model.

[0050] As an optional embodiment, before determining the anomaly detection model corresponding to the potential anomaly items, it further includes: determining the detection parameters corresponding to the potential anomaly items, where the detection parameters include the detection complexity index and the detection accuracy parameter; determining the target model architecture and the target hyperparameters according to the detection parameters.

[0051] In this embodiment, the steps of determining the target model architecture and the target hyperparameters before determining the anomaly detection model corresponding to the potential anomaly items are described.

[0052] Among them, the detection parameters are involved. The detection parameters can be set or adjusted according to the characteristics of the potential anomaly items (such as data volume, time series characteristics, text patterns, etc.). For example, if the potential anomaly items involve complex time series analysis, a higher detection complexity index may need to be set; if the potential anomaly items are directly related to behavioral risks, a more stringent detection accuracy parameter may need to be set to reduce the false alarm rate.

[0053] Among them, the selection of the target model architecture is involved. According to the detection parameters, the system will select or design the most suitable architecture from a series of preset model architectures. For example, for detection tasks with a high complexity index, a neural network architecture containing recurrent layers or attention mechanisms may be selected; for requirements of high-precision parameters, a model architecture containing more detailed feature extraction may be designed.

[0054] This involves setting target hyperparameters. After the model architecture is determined, the system will set the best hyperparameter combination based on the detection parameters. For example, in order to improve the accuracy of detection, it may be necessary to adjust the learning rate and increase the regularization coefficient to prevent the model from overfitting; in order to adapt to a large amount of data to be detected, it may be necessary to set a larger batch size and an appropriate training cycle.

[0055] You can also perform model training and validation, and start training the anomaly detection model using the selected model architecture and hyperparameters. During the training process, you can adjust the model parameters and optimize the detection algorithm to ensure that the model can accurately and efficiently identify potential anomalies. After training, the model needs to be evaluated through methods such as cross-validation to verify its performance in detection accuracy and processing complexity.

[0056] In this embodiment, by setting a reasonable detection complexity index, the system can balance computing resources and processing time to ensure that anomaly detection achieves the best balance between efficiency and performance. The setting of detection accuracy parameters helps to control the false alarm rate and missed alarm rate, ensure the accuracy and rigor of anomaly detection, and reduce compliance risks. The determination of the target model architecture and target hyperparameters enables the anomaly detection model to better adapt to the characteristics of potential anomalies and improve the generalization and robustness of the model.

[0057] As an optional embodiment, before determining the anomaly detection model corresponding to the potential abnormal item, it also includes: determining multiple initial candidate models corresponding to the potential abnormal item, wherein the multiple initial candidate models are provided with different model architectures; determining test parameters of the multiple initial candidate models under multiple test indicators respectively; and determining an initial detection model from the multiple initial candidate models based on the test parameters of the multiple initial candidate models under the multiple test indicators respectively.

[0058] In this embodiment, a process of determining an initial detection model having a target model architecture before determining an anomaly detection model corresponding to a potential anomaly item is described.

[0059] Among them, multiple initial candidate models are involved. Before determining the anomaly detection model, a series of optional models will be prepared. Each model has a different architecture, such as decision tree, support vector machine, neural network, etc. The most suitable model for detecting potential anomalies can be found through comparison. Model architecture refers to the structure and composition of the model, including the type of model (such as linear model, nonlinear model), number of layers (for neural network), number of nodes, activation function, loss function, etc. Different model architectures are good at handling different types of data and problems. It can cover a variety of detection strategies and methods.

[0060] Among them, multiple test metrics are involved. During the testing and evaluation of the model, a series of metrics are used to measure the performance of the model, including accuracy, precision, recall, F1-score, etc. These metrics reflect the detection ability and efficiency of the model from different perspectives.

[0061] Among them, test parameters are also involved. Test parameters refer to the parameters used to adjust the model performance and reflect the evaluation results during the model testing process. The performance of each model can be tested under accuracy, precision, and recall to comprehensively evaluate the detection ability and efficiency of the model.

[0062] Among multiple initial candidate models, through testing and evaluation, the system will determine a model with the best performance as the preliminary model for anomaly detection. This model will be used for further training, optimization, and application to improve its anomaly detection ability in actual scenarios. Models with different architectures have different advantages and limitations when dealing with specific types of data. By testing and comparing multiple initial candidate models, the system can select a model with stronger generalization ability to adapt to diverse potential anomaly items and detection scenarios.

[0063] As an optional embodiment, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining a predetermined hyperparameter range according to the potential anomaly item; selecting multiple first test hyperparameters from the predetermined hyperparameter range using a first span; determining the model quality index corresponding to each of the multiple alternative detection models, where the multiple alternative detection models correspond one-to-one to the multiple first test hyperparameters; determining multiple restricted hyperparameters from the multiple first test hyperparameters according to the multiple model quality indices; determining an updated hyperparameter range according to the multiple restricted hyperparameters, and then using a second span to determine second test hyperparameters from the updated hyperparameter range, and repeating the process until the target hyperparameter with the highest corresponding model quality index is determined.

[0064] In this embodiment, the steps of determining the target hyperparameter before determining the anomaly detection model corresponding to the potential anomaly item are described.

[0065] Among them, a predetermined hyperparameter range is involved. Before selecting the anomaly detection model, according to the characteristics of the potential anomaly item (such as data type, data volume, complexity, etc.), the system will set an initial hyperparameter range. Hyperparameters are the parameters that need to be set before training a machine learning model, such as learning rate, regularization coefficient, number of network layers, etc. This range provides a basis for subsequent hyperparameter search and optimization.

[0066] Among them, the first span is involved. This is a strategy or method used by the system to select multiple first test hyperparameters within a predetermined hyperparameter range. The span can be a fixed step size. For example, in the search for the learning rate, from 0.01 to 0.1, with a step size of 0.01 for each span, multiple test learning rate values will be generated in this way.

[0067] Among them, the first test hyperparameters are involved. These are hyperparameter values selected from the predetermined hyperparameter range according to the first span and are used to conduct preliminary tests and evaluations on the alternative detection models. The purpose is to quickly screen out model and hyperparameter combinations that may perform well.

[0068] Among them, the model quality index is involved. The model quality index is a comprehensive index used to evaluate the performance of the model and can include accuracy, precision, recall, F1-score, etc. In the evaluation of multiple alternative detection models, the model quality index reflects the comprehensive performance of a model under a specific hyperparameter combination.

[0069] Among them, the alternative detection models are involved. The alternative detection models are models trained according to different hyperparameters and are used to evaluate the advantages and disadvantages of the hyperparameters.

[0070] Among them, the restricted hyperparameters are involved. After the first round of tests and evaluations, the system will, based on the feedback of the model quality index, screen out the hyperparameter values with better performance from the first test hyperparameters as the restricted hyperparameters, which are used to narrow the scope of subsequent hyperparameter search.

[0071] Among them, the updated hyperparameter range is involved. That is, based on the feedback of the restricted hyperparameters, the system will adjust the predetermined hyperparameter range and narrow the search interval to more precisely locate the optimal hyperparameter combination. For example, if it is found in the first round of tests that a smaller learning rate performs better, the updated hyperparameter range may focus on a smaller learning rate interval.

[0072] Among them, the second span is involved. This is a strategy or method used to select the second test hyperparameters within the updated hyperparameter range. Similar to the first span, the second span also defines the search interval or method within the updated hyperparameter range, but it is usually more detailed, that is, the second span is smaller than the first span, so as to more precisely explore the hyperparameter space.

[0073] Among them, the second test hyperparameters are involved. These are hyperparameter values selected from the updated hyperparameter range according to the second span and are used to conduct more in-depth tests and evaluations on the model. These hyperparameter values are based on the results of the first round of tests and the feedback of the restricted hyperparameters and can further optimize the model performance.

[0074] Among them, the target hyperparameters are involved. Through multiple rounds of testing and evaluation, the system finally determines the optimal combination of hyperparameters. This combination shows the highest performance in the model quality index and will be used for subsequent model training and optimization.

[0075] In this embodiment, through multiple rounds of hyperparameter optimization, the system can find the model and hyperparameter combination that is most suitable for processing potential abnormal items, significantly improving the detection accuracy and efficiency of the model. Moreover, by restricting the hyperparameters and updating the hyperparameter range strategy, the hyperparameter search space can be effectively reduced, unnecessary calculations can be reduced, and computing resources and time can be saved. The determination of the target hyperparameters helps to improve the stability of the model under different datasets and scenarios and reduce the performance fluctuations of the model caused by improper hyperparameter settings. According to the characteristics of potential abnormal items, the system can customize the selection and optimization of the model and hyperparameters to detect abnormalities more effectively.

[0076] As an alternative embodiment, in response to a data detection request, determining potential abnormal items corresponding to the data to be detected includes: in response to a data detection request, determining the detection type corresponding to the data to be detected; determining a predetermined template corresponding to the detection type; performing patterned processing of the predetermined template on the data to be detected to obtain target template data; and determining potential abnormal items corresponding to the data to be detected based on the target template data.

[0077] In this embodiment, the process of determining potential abnormal items is described.

[0078] Among them, the detection type is involved. The detection type can be the type of anomaly detection, such as transaction detection, report anomaly detection, abnormal behavior recognition, etc. Different detection types correspond to different detection targets and data characteristics.

[0079] Among them, the predetermined template is involved. The predetermined template is a structure or framework for data patterned processing, used to convert the data to be detected into a specific format that is conducive to the analysis of the anomaly detection model. The design of the template is based on specific detection types and data characteristics, which can highlight the key points of detection.

[0080] Among them, the target template data is involved. The target template data is the data format that the data to be detected is converted into after patterned processing and matches a specific anomaly detection model.

[0081] Among them, the potential abnormal items are involved. Potential abnormal items refer to behaviors or patterns that may be abnormal in the data to be detected, and these abnormal items need further analysis and confirmation.

[0082] In this step, when the system receives a data detection request, it first analyzes the information in the request to determine the specific detection type. According to the detection type, the system will automatically select a preset detection template that best matches this type, because different detection types often require attention to different data characteristics and detection patterns. After selecting the template, the system will perform a patterning process on the data to be detected, that is, perform preprocessing operations such as reorganizing, cleaning, and feature extraction on the data according to the structure of the template, so that the data format meets the input requirements and obtains the target template data.

[0083] By automatically selecting the template and performing the patterning process, the system can quickly prepare the data, provide accurate input for the anomaly detection model, greatly speed up the detection process, and save manpower and time costs. The patterning process ensures the consistency and format standardization of the data, enables the model to more accurately identify potential anomaly items, and reduces the possibility of false positives and false negatives. Moreover, the system can intelligently select the template according to the detection type, avoiding the waste of resources for unified processing of all data, and making the computing resources more efficiently serve the detection requirements.

[0084] As an alternative embodiment, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: when the data detection request carries a predetermined scenario, determining the sample data corresponding to the predetermined scenario; training the initial detection model based on the sample data to obtain the anomaly detection model.

[0085] In this embodiment, the process of obtaining the anomaly detection model before determining the anomaly detection model corresponding to the potential anomaly item is described.

[0086] Among them, a predetermined scenario is involved. When initiating a data detection request, when there is a clear scenario, such as it may involve a specific business process (such as a loan application), a specific time period (such as a holiday), etc., each scenario may correspond to different anomaly patterns and detection requirements. It can be set adaptively according to the targeted scenario.

[0087] Among them, sample data is involved, that is, according to the predetermined scenario, the most representative sample data will be selected from historical data or the database, including data of normal behavior and abnormal behavior. The selection of sample data should ensure the diversity and representativeness of the data, so that the model can learn all potential anomaly patterns in the scenario.

[0088] Among them, the process of training the initial detection model is involved, that is, the system uses the selected sample data to train the initial detection model. At this stage, the model gradually adapts to the anomaly detection task in a specific scenario by learning the features in the sample data.

[0089] After training, the initial detection model is transformed into an anomaly detection model capable of identifying abnormal behaviors in specific scenarios. This model can analyze new detection data based on the patterns of the training data and quickly and accurately discover potential abnormal items. Through training with sample data of specific scenarios, the anomaly detection model can more accurately identify abnormal behaviors in the scenarios, improving the pertinence and efficiency of detection.

[0090] Based on the above embodiments and optional embodiments, an optional implementation manner is provided, which is specifically described below.

[0091] An optional implementation manner of the present invention provides a data detection method based on a bank scenario. Figure 2 It is a structural block diagram of the data detection method provided according to the optional implementation manner of the present invention, as Figure 2 shown, and it will be introduced below.

[0092] In this scenario, the language large model technology can be utilized, combined with the actual requirements of bank data compliance detection, to develop a new type of bank data compliance detection system, so as to improve the detection efficiency, reduce costs, and ensure the accuracy and reliability of the data compliance detection results. The target users are data inspection personnel within the bank, which can help them perform detection work more quickly and accurately and improve the detection efficiency.

[0093] S1, Data collection and preparation:

[0094] Obtain relevant data for bank data compliance detection (the data to be detected carried in the above data detection request), including data systems, banking business data, data inspection reports, etc., and clean and preprocess the data.

[0095] The specific method is as follows:

[0096] Data cleaning and preprocessing: After data collection is completed, data cleaning and preprocessing work need to be carried out first. This includes removing irrelevant data, correcting incorrect data, filling in missing values, data desensitization, etc., to ensure the quality and security of the data. Data desensitization refers to replacing or hiding sensitive information to prevent the leakage of personal information and ensure compliance during data processing.

[0097] Data classification and grading: Classify and grade the data according to the content and use of the data. For example, classify the data into different levels such as public level, internal level, sensitive level, etc., so as to take corresponding security measures. At the same time, relevant data classification and processing standards need to be followed to ensure that data processing activities comply with the requirements of laws and regulations.

[0098] Build a knowledge base: Integrate the cleaned and classified data into the knowledge base. The knowledge base should contain various business rules, compliance detection provisions, historical cases (the same sample data corresponding to the predetermined scenarios above, i.e., historical sample data related to the bank scenario), etc., to support the training and application of the large model. During this process, it is necessary to ensure that the construction and maintenance of the knowledge base comply with data security and compliance requirements, such as by implementing security measures such as access control and data encryption.

[0099] S2, Model training and optimization:

[0100] In the model training and optimization stage, this invention adopts a series of innovative technical strategies to ensure the efficiency and accuracy of the model. The following is a detailed description of the key technical steps:

[0101] (1) Data preprocessing before model input:

[0102] Before model training, the following key data processing steps are implemented to ensure data quality and the effectiveness of model training:

[0103] (1) Data standardization: Unify data from different sources and formats, eliminate the influence of dimensions, and improve the generalization ability of the model.

[0104] (2) Outlier handling: Adopt advanced statistical methods to identify and handle outliers to ensure that model training is not misled.

[0105] 1) Data exploration: First, conduct data exploration to understand the distribution characteristics of the data, including central tendency (mean, median), dispersion degree (standard deviation, interquartile range), etc. This step helps to have an overall understanding of the data.

[0106] 2) Outlier identification: Use statistical methods to identify outliers in the data. Common methods include:

[0107] a. Standard deviation method: Calculate the mean and standard deviation of the data. Usually, it is considered that data points with a deviation from the mean exceeding 2 or 3 standard deviations are outliers.

[0108] b. Quartile method: Use the first quartile (Q1) and the third quartile (Q3) to determine the interquartile range (IQR) of the data. Outliers are usually defined as data points below Q1 - 1.5 * IQR or above Q3 + 1.5 * IQR.

[0109] c. Boxplot: A boxplot is a visualization tool that can intuitively display the distribution and outliers of the data.

[0110] d. Z - score method: Calculate the Z - score of each data point, that is, the difference between the data point and the mean divided by the standard deviation. An absolute value of the Z - score greater than 3 is usually considered an outlier.

[0111] 3) Outlier analysis: After identifying outliers, they need to be analyzed to determine whether they are real anomalies or data entry errors or measurement errors. This step may require collaboration with domain experts to understand the business context of the data.

[0112] 4) Ways to handle outliers:

[0113] a. Deletion: If outliers are due to obvious errors or irrelevant data points, these data can be selected for deletion.

[0114] b. Correction: If outliers are due to identifiable errors, such as data entry errors, these values can be corrected.

[0115] c. Transformation: Transform the data, such as logarithmic transformation or other transformations, to reduce the impact of outliers.

[0116] d. Retention: If outliers are real business situations, these data points may need to be retained and appropriately handled in the model.

[0117] 5) Model evaluation: After handling outliers, the performance of the model needs to be re-evaluated to ensure that the handling of outliers has not had a negative impact on the generalization ability of the model.

[0118] (3) Sequence analysis: For time-sensitive data, apply time series analysis techniques to capture and utilize the time dependence of the data. In the data, time series analysis can help understand trends, seasonal patterns, cyclic changes, etc. as the data changes over time. The following are the general steps of time series analysis in the data and the processing logic for each step:

[0119] 1) Data collection: Collect the data to be analyzed, such as sales, costs, profits, etc., and ensure that the data is arranged in chronological order.

[0120] 2) Data cleaning: Remove or fill in missing values and correct incorrect data.

[0121] 3) Data transformation: If necessary, convert the data into a format suitable for time series analysis, such as date-time format.

[0122] 4) Seasonal adjustment: If the data exhibits obvious seasonal patterns, the seasonal impact can be eliminated through seasonal adjustment methods (such as seasonal differencing).

[0123] 5) Data visualization: Intuitively observe the characteristics of the data, such as trends, seasonality, cyclicity, etc. by plotting a time series graph.

[0124] 6) Stationarity test: Time series analysis usually requires the data to be stationary, that is, the statistical characteristics of the data (such as mean and variance) are constant over time. Unit root tests (such as ADF test) can be used to test the stationarity of the data.

[0125] 7) Differencing operation: If the data is not stationary, it can be made stationary by differencing (i.e., calculating the changes in data over consecutive time periods). The order of differencing depends on the characteristics of the data.

[0126] 8) Model identification: Select an appropriate time series model according to the characteristics of the data. Common models include AR (Autoregressive Model), MA (Moving Average Model), ARMA (Autoregressive Moving Average Model), and ARIMA (Autoregressive Integrated Moving Average Model), etc.

[0127] 9) Parameter estimation: Use maximum likelihood estimation, Bayesian methods, or other statistical methods to estimate the model parameters.

[0128] 10) Model diagnosis: Check whether the residuals of the model conform to the assumption of white noise, that is, there is no autocorrelation between the residuals. ACF (Autocorrelation Function) and PACF (Partial Autocorrelation Function) can be used for diagnosis.

[0129] 11) Model validation: Evaluate the prediction ability of the model through methods such as cross-validation, AIC (Akaike Information Criterion), BIC (Bayesian Information Criterion), etc.

[0130] 12) Prediction: Use the established model to predict future data and give a prediction interval.

[0131] 13) Model update: As time goes by, new data is continuously generated, and the model needs to be updated regularly to maintain its accuracy.

[0132] 14) Anomaly detection: During the time series analysis process, it is also necessary to detect outliers or mutation points, which may be caused by special events and need special attention.

[0133] 15) Reporting and decision support: Finally, organize the analysis results into a report to provide decision support for management, such as inventory management, data planning, etc.

[0134] 16) In the time series analysis of data, each step requires careful consideration and professional statistical knowledge.

[0135] (2) Use the corresponding model (the same anomaly detection model corresponding to the potential anomaly items above) to conduct the interaction between model output and compliance assessment: Particularly emphasize the interpretability and interactivity of the model output to enhance the user experience and the practicality of the model:

[0136] (1) Output interpretability: The model provides detailed explanations of the evaluation results, enabling users to understand the logic and basis behind the evaluation.

[0137] (2) Interactive feedback mechanism: A user-friendly feedback system is built, allowing users to evaluate and provide feedback on the model output, enabling continuous improvement of the model.

[0138] (3) Dynamic compliance assessment: The model can dynamically adjust the evaluation criteria and processes based on user feedback and changes in compliance requirements, ensuring the timeliness and adaptability of the evaluation.

[0139] (III) Data augmentation and feature engineering: To improve the generalization ability and prediction accuracy of the model, the following techniques are adopted:

[0140] (1) Data augmentation: Techniques such as synthetic data generation are introduced to expand the training set and improve the model's adaptability to new situations.

[0141] (2) Feature selection: Advanced algorithms are used to identify key features, reduce feature redundancy, and enhance model performance.

[0142] (IV) Model architecture and hyperparameter optimization: During the training and feedback process of the model, the model architecture and hyperparameters can be carefully adjusted and optimized:

[0143] (1) Model architecture adjustment: According to the characteristics of banking business data, the model architecture is customized, such as introducing convolutional layers or recurrent layers, to better capture data patterns.

[0144] (2) Hyperparameter optimization: Automated machine learning techniques, such as Bayesian optimization, are used to automatically find the optimal combination of hyperparameters to improve model performance.

[0145] (V) Model fusion and interpretability: Through ensemble learning and model interpretability tools, the accuracy and transparency of the model are further enhanced:

[0146] (1) Model fusion: Ensemble learning techniques are adopted to combine the predictions of multiple models to improve the overall accuracy and robustness.

[0147] (2) Model interpretability: Model interpretability tools are integrated to provide detailed explanations of model decisions, enhancing users' trust in the model output.

[0148] (VI) Continuous learning mechanism: The model is designed to support online learning and incremental learning, ensuring that the model can adapt to changing data and compliance requirements:

[0149] (1) Continuous learning: The model can continuously learn from new data and adapt to new compliance scenarios without retraining.

[0150] Through these technical strategies, the model training and optimization process not only improves the prediction accuracy and generalization ability of the model, but also enhances the interactivity and adaptability of the model, ensuring the efficiency and reliability of the model in practical applications.

[0151] For functional design and software product development:

[0152] Design the functions of the application system to meet the actual needs of bank internal data inspection and detection personnel, including the establishment of a data compliance knowledge base, the setting of a dialogue agent robot (Agent), the analysis of institutional provisions, etc., and can be customized according to user needs. The following are the key functional modules:

[0153] 1. Upload and management of the knowledge base: Develop a user-friendly knowledge base management interface that allows data inspectors to upload and manage relevant materials such as data systems, banking business rules, inspection reports, etc. The interface should provide simple and intuitive operations, including functions such as uploading files, editing content, and classification management, so that users can quickly and effectively manage the audit knowledge base.

[0154] 2. Interface for users to chat with the large model: Develop an interactive chat interface that allows data inspectors to interact with the trained language large model. In this interface, set two Agents:

[0155] (1) System basis query Agent: Specifically responsible for answering users' questions about specific system bases or industry norms. Users can ask relevant questions in the conversation with this Agent, such as: "Please provide the detailed data system rules for entertainment expense reimbursement."

[0156] (2) Key inspection points and method suggestion Agent: Specifically responsible for providing users with key points of data compliance detection and method suggestions. Users can describe audit requirements and questions in the conversation with this Agent, such as: "What aspects do I need to focus on during the inspection?" or "Please give me some suggestions on how to effectively carry out off-site data detection?"

[0157] 3. Module for viewing and citing relevant regulatory documents: Directly view relevant regulatory documents, key points of data compliance detection, and method suggestions in the dialogue thread, so as to deeply understand the basis of data compliance detection and obtain guidance on data compliance detection. Through the development of the above software products, data inspectors can conveniently manage the data knowledge base and interact with the trained language large model to achieve a more efficient and intelligent data compliance detection work process.

[0158] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.

[0160] According to an embodiment of the present invention, there is also provided a device for implementing the above data detection method. Figure 3 It is a structural block diagram of a data detection device provided according to an embodiment of the present application, as Figure 3 shown. The device includes: a receiving module 301, a first determination module 302, a second determination module 303, and a third determination module 304. The device will be described in detail below.

[0161] The receiving module 301 is configured to receive a data detection request. Among them, the data detection request carries the data to be detected, and the data detection request is used to detect whether there is an abnormality in the data to be detected; the first determination module 302 is connected to the above receiving module 301, and is configured to, in response to the data detection request, determine a potential abnormal item corresponding to the data to be detected, where the potential abnormal item is an abnormal item whose probability of the data to be detected having the corresponding abnormality is greater than a predetermined threshold; the second determination module 303 is connected to the above first determination module 302, and is configured to determine an anomaly detection model corresponding to the potential abnormal item, where the anomaly detection model is an initial detection model provided with a target model architecture and trained using target hyperparameters, the target model architecture is obtained based on the potential abnormal item, and the target hyperparameters are obtained based on the potential abnormal item; the third determination module 304 is connected to the above second determination module 303, and is configured to determine an anomaly detection result corresponding to the potential abnormal item based on the data to be detected and the anomaly detection model.

[0162] It should be noted here that the above receiving module 301, first determining module 302, second determining module 303, and third determining module 304 correspond to steps S101 to S104 in the implementation of the data detection method. The instances and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.

[0163] The data detection device provided by the embodiment of the present application receives a data detection request, where the data detection request carries the data to be detected, and the data detection request is used to detect whether there is an abnormality in the data to be detected. In response to the data detection request, a potential abnormal item corresponding to the data to be detected is determined, where the potential abnormal item is an abnormal item whose probability of the data to be detected having the corresponding abnormality is greater than a predetermined threshold. An abnormal detection model corresponding to the potential abnormal item is determined, where the abnormal detection model is an initial detection model provided with a target model architecture and trained using target hyperparameters. The target model architecture is obtained based on the potential abnormal item, and the target hyperparameters are obtained based on the potential abnormal item. Based on the data to be detected and the abnormal detection model, an abnormal detection result corresponding to the potential abnormal item is determined. It can be seen that the abnormal detection model corresponding to the abnormal item is used for abnormal detection, and moreover, the abnormal detection model customizes the model architecture and hyperparameters based on the potential abnormal item. By specifically adjusting the model architecture and hyperparameters, the system can more accurately identify potential abnormal items, reduce the false alarm rate and missed alarm rate, improve the detection efficiency, and the customized model can process data more quickly, shortening the abnormal detection cycle, thereby solving the technical problem of low detection efficiency when detecting whether there is an abnormality in the relevant technology.

[0164] The data detection device includes a processor and a memory. The above multiple modules and the like are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory.

[0165] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the technical problem of low detection efficiency when detecting whether there is an abnormality in the relevant technology is solved.

[0166] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory includes at least one storage chip.

[0167] The embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the data detection method.

[0168] An embodiment of the present invention provides a processor for running a program, wherein when the program runs, it executes a data detection method.

[0169] An embodiment of the present invention provides a non-volatile storage medium, characterized in that the non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the above data detection method.

[0170] Figure 4 is a schematic diagram of an electronic device provided by an embodiment of the present invention, as Figure 4 shown, an embodiment of the present invention provides an electronic device, the device includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented: receiving a data detection request, wherein the data detection request carries the data to be detected, and the data detection request is used to detect whether there is an abnormality in the data to be detected; in response to the data detection request, determining a potential abnormality item corresponding to the data to be detected, wherein the potential abnormality item is an abnormality item whose probability of the data to be detected having the corresponding abnormality is greater than a predetermined threshold; determining an abnormality detection model corresponding to the potential abnormality item, wherein the abnormality detection model is an initial detection model with a target model architecture, and is obtained by training with target hyperparameters, the target model architecture is obtained based on the potential abnormality item, and the target hyperparameters are obtained based on the potential abnormality item; based on the data to be detected and the abnormality detection model, determining an abnormality detection result corresponding to the potential abnormality item.

[0171] Optionally, based on the data to be detected and the abnormality detection model, determining an abnormality detection result corresponding to the potential abnormality item includes: determining a format processing method corresponding to the potential abnormality item; processing the data to be detected according to the format processing method to obtain updated detection data; inputting the updated detection data into the abnormality detection model to obtain a detection result corresponding to the potential abnormality item.

[0172] Optionally, before determining the abnormality detection model corresponding to the potential abnormality item, it further includes: determining detection parameters corresponding to the potential abnormality item, wherein the detection parameters include a detection complexity index and a detection accuracy parameter; based on the detection parameters, determining the target model architecture and the target hyperparameters.

[0173] Optionally, before determining the abnormality detection model corresponding to the potential abnormality item, it further includes: determining a plurality of initial candidate models corresponding to the potential abnormality item, wherein the plurality of initial candidate models are provided with different model architectures; determining the test parameters of the plurality of initial candidate models under a plurality of test metrics; based on the test parameters of the plurality of initial candidate models under the plurality of test metrics, determining an initial detection model from the plurality of initial candidate models.

[0174] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining a predetermined hyperparameter range based on the potential anomaly item; selecting a plurality of first test hyperparameters from the predetermined hyperparameter range using a first span; determining the model quality indices corresponding to a plurality of alternative detection models, where the plurality of alternative detection models correspond to the plurality of first test hyperparameters one by one; determining a plurality of restricted hyperparameters from the plurality of first test hyperparameters based on the plurality of model quality indices; determining an updated hyperparameter range based on the plurality of restricted hyperparameters, and based on the updated hyperparameter range, determining a second test hyperparameter from the updated hyperparameter range using a second span, and looping until the target hyperparameter with the highest corresponding model quality index is determined.

[0175] Optionally, in response to a data detection request, determining a potential anomaly item corresponding to the data to be detected includes: in response to the data detection request, determining the detection type corresponding to the data to be detected; determining a predetermined template corresponding to the detection type; performing pattern processing of the predetermined template on the data to be detected to obtain target template data; and determining the potential anomaly item corresponding to the data to be detected based on the target template data.

[0176] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: when the data detection request carries a predetermined scenario, determining the sample data corresponding to the predetermined scenario; and training the initial detection model based on the sample data to obtain the anomaly detection model.

[0177] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0178] This application also provides a computer program product, which when executed on a data processing device is adapted to execute a program initialized with the following method steps: receiving a data detection request, where the data detection request carries the data to be detected, and the data detection request is used to detect whether there is an anomaly in the data to be detected; in response to the data detection request, determining a potential anomaly item corresponding to the data to be detected, where the potential anomaly item is an anomaly item whose probability of the data to be detected having the corresponding anomaly is greater than a predetermined threshold; determining an anomaly detection model corresponding to the potential anomaly item, where the anomaly detection model is obtained by training an initial detection model with a target model architecture using target hyperparameters, the target model architecture is obtained based on the potential anomaly item, and the target hyperparameters are obtained based on the potential anomaly item; and determining an anomaly detection result corresponding to the potential anomaly item based on the data to be detected and the anomaly detection model.

[0179] Optionally, based on the data to be detected and the anomaly detection model, determining the anomaly detection result corresponding to the potential anomaly item, including: determining the format processing method corresponding to the potential anomaly item; processing the data to be detected according to the format processing method to obtain updated detection data; inputting the updated detection data into the anomaly detection model to obtain the detection result corresponding to the potential anomaly item.

[0180] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining the detection parameters corresponding to the potential anomaly item, where the detection parameters include the detection complexity index and the detection accuracy parameter; determining the target model architecture and the target hyperparameters according to the detection parameters.

[0181] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining a plurality of initial candidate models corresponding to the potential anomaly item, where the plurality of initial candidate models are set with different model architectures; determining the test parameters of the plurality of initial candidate models under a plurality of test metrics; determining an initial detection model from the plurality of initial candidate models according to the test parameters of the plurality of initial candidate models under the plurality of test metrics.

[0182] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: determining a predetermined hyperparameter range according to the potential anomaly item; selecting a plurality of first test hyperparameters from the predetermined hyperparameter range using a first span; determining the model quality index corresponding to each of the plurality of alternative detection models, where the plurality of alternative detection models correspond one-to-one to the plurality of first test hyperparameters; determining a plurality of restricted hyperparameters from the plurality of first test hyperparameters according to the plurality of model quality indexes; determining an updated hyperparameter range according to the plurality of restricted hyperparameters, so as to determine second test hyperparameters from the updated hyperparameter range using a second span, and repeating until the target hyperparameters with the highest corresponding model quality index are determined.

[0183] Optionally, in response to a data detection request, determining the potential anomaly item corresponding to the data to be detected, including: in response to the data detection request, determining the detection type corresponding to the data to be detected; determining a predetermined template corresponding to the detection type; performing pattern processing of the predetermined template on the data to be detected to obtain target template data; determining the potential anomaly item corresponding to the data to be detected according to the target template data.

[0184] Optionally, before determining the anomaly detection model corresponding to the potential anomaly item, it further includes: when the data detection request carries a predetermined scenario, determining the sample data corresponding to the predetermined scenario; training the initial detection model according to the sample data to obtain the anomaly detection model.

[0185] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0186] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0187] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0189] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0190] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0191] A computer-readable medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0192] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0193] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0194] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A data detection method, characterized in that: include: Receiving a data detection request, wherein the data detection request carries data to be detected, and the data detection request is used to detect whether there is an abnormality in the data to be detected; In response to the data detection request, determining a potential abnormal item corresponding to the data to be detected, wherein the potential abnormal item is an abnormal item for which the probability of the corresponding abnormality occurring in the data to be detected is greater than a predetermined threshold; Determine an anomaly detection model corresponding to the potential anomaly item, wherein the anomaly detection model is obtained by training an initial detection model provided with a target model architecture using target hyperparameters, the target model architecture is obtained based on the potential anomaly item, and the target hyperparameters are obtained based on the potential anomaly item; An abnormality detection result corresponding to the potential abnormal item is determined based on the data to be detected and the abnormality detection model.

2. The method according to claim 1, characterized in that The determining, based on the data to be detected and the anomaly detection model, an anomaly detection result corresponding to the potential anomaly item includes: Determine a format processing method corresponding to the potential abnormal item; Processing the data to be detected according to the format processing method to obtain updated detection data; The updated detection data is input into the anomaly detection model to obtain a detection result corresponding to the potential anomaly item.

3. The method according to claim 1, characterized in that Before determining the anomaly detection model corresponding to the potential anomaly item, the method further includes: Determining detection parameters corresponding to the potential abnormal item, wherein the detection parameters include a detection complexity index and a detection accuracy parameter; The target model architecture and the target hyperparameters are determined based on the detection parameters.

4. The method according to claim 1, characterized in that: Before determining the anomaly detection model corresponding to the potential anomaly item, the method further includes: Determine a plurality of initial candidate models corresponding to the potential abnormal items, wherein the plurality of initial candidate models are provided with different model architectures; Determining test parameters of the multiple initial candidate models under multiple test indicators respectively; The initial detection model is determined from the multiple initial candidate models according to the test parameters of the multiple initial candidate models under multiple test indicators respectively.

5. The method according to claim 1, characterized in that Before determining the anomaly detection model corresponding to the potential anomaly item, the method further includes: Determining a predetermined hyperparameter range based on the potential abnormal item; Selecting a plurality of first test hyperparameters from the predetermined hyperparameter range using a first span; Determine model quality indexes corresponding to a plurality of candidate detection models, respectively, wherein the plurality of candidate detection models correspond one-to-one to the plurality of first test hyperparameters; Determining a plurality of limiting hyperparameters from the plurality of first testing hyperparameters according to a plurality of model quality indices; Based on the multiple restricted hyperparameters, an updated hyperparameter range is determined, and based on the updated hyperparameter range, a second test hyperparameter is determined from the updated hyperparameter range using a second span, and the cycle is repeated until a target hyperparameter with the highest corresponding model quality index is determined.

6. The method according to claim 1, characterized in that The step of determining, in response to the data detection request, a potential abnormal item corresponding to the data to be detected includes: In response to the data detection request, determining a detection type corresponding to the data to be detected; determining a predetermined template corresponding to the detection type; Performing pattern processing of a predetermined template on the data to be detected to obtain target template data; Based on the target template data, potential abnormal items corresponding to the data to be detected are determined.

7. The method according to any one of claims 1 to 6, characterized in that: Before determining the anomaly detection model corresponding to the potential anomaly item, the method further includes: In a case where the data detection request carries a predetermined scenario, determining sample data corresponding to the predetermined scenario; The initial detection model is trained according to the sample data to obtain the anomaly detection model.

8. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the data detection method according to any one of claims 1 to 7.

9. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the data detection method described in any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that: The computer instructions are executed by a processor to execute the data detection method according to any one of claims 1 to 7.