Abnormal service data determination method, apparatus and device

By using a pre-trained and optimized spatial clustering model combined with density reachability graphs, abnormal business data can be automatically identified, solving the problem of reliance on human experience in existing technologies and improving the efficiency and accuracy of business report download behavior analysis.

CN119316324BActive Publication Date: 2026-01-02INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410734265.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2026-01-02
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

In existing technologies, the identification of anomalies in business report download behavior relies on human experience, which is inefficient, time-consuming, and difficult to automate.

Method used

By employing a pre-trained and optimized spatial clustering model, combined with a density reachability graph, abnormal business data can be automatically identified by filtering out anomalous data points.

Benefits of technology

It enables automatic identification of abnormal business report download behavior from massive amounts of data, reducing the workload of manual judgment and improving efficiency and processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316324B_ABST
    Figure CN119316324B_ABST
Patent Text Reader

Abstract

The disclosure provides an abnormal business data determination method. It can be applied to the technical field of financial technology and the technical field of information security. The method comprises the following steps: acquiring a newly added business download data, inputting the a newly added business download data into a pre-trained and optimized spatial clustering model, outputting b non-outlier newly added business download data points as b first data points, wherein each first data point comprises the coordinates of a newly added business download data point. The coordinates of the b first data points are introduced into a pre-generated first density reachable graph to generate a second density reachable graph. Based on the second density reachable graph, d abnormal data points are screened out from the b first data points, and d abnormal business data are determined according to the d abnormal data points.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of financial technology, in particular to the technical field of information security, and more particularly to an abnormal business data determination method, device, equipment, medium and program product. BACKGROUND

[0002] In enterprises and organizations including banks, business reports are important basis for decision-making and management. However, business reports usually contain sensitive information such as financial data, customer information, etc. Therefore, the business report download system usually has strong access control and identity authentication mechanism to ensure that only authorized users can access and download the report, and to prevent unauthorized download behavior and various security risks such as data leakage. However, for business personnel who have obtained authorization, it is very difficult to define whether the business report they download is required by their work. Usually, the analysis of business report download behavior depends on professional personnel to capture these complex information through business rules and logic, and then judge whether the business report download behavior is an abnormal download behavior through personal experience. This method is extremely dependent on the business experience of professional personnel and consumes a lot of time of business personnel, and the work efficiency is very low. SUMMARY

[0003] In view of the above problems, the present disclosure provides an abnormal business data determination method, device, equipment, medium and program product.

[0004] According to a first aspect of the present disclosure, an abnormal business data determination method is provided, characterized in that the method comprises: obtaining a number of new business download data, inputting the a number of new business download data into a pre-trained and tuned spatial clustering model, outputting b number of non-outlier new business download data points as b number of first data points, wherein each first data point comprises the coordinates of the new business download data point, wherein a is greater than or equal to 2, b is less than or equal to a, and a and b are positive integers; introducing the coordinates of the b number of first data points into a pre-generated first density reachable graph to generate a second density reachable graph; and based on the second density reachable graph, screening d number of abnormal data points from the b number of first data points, and determining d number of abnormal business data according to the d number of abnormal data points, wherein d is less than or equal to b, and d is a positive integer.

[0005] According to the embodiment of the present disclosure, the first density reachable graph comprises: outlying point coordinates in the historical service download data points; and the filtering of the d abnormal data points from the b first data points based on the second density reachable graph comprises: obtaining a field radius of the outlying point coordinates in the historical service download data points from the second density reachable graph, to generate a plurality of first historical data points; obtaining a plurality of first historical data points within the field radius of each first data point from the second density reachable graph, to generate a plurality of second historical data points corresponding to each first data point; calculating the Euclidean distance between each first data point and the plurality of second historical data points, to generate a plurality of first distances of each first data point, wherein the field radius of each second historical data point generating the first distance with the first data point is not the same; obtaining the field radius of each second historical data point generating the first distance with the first data point, to generate a plurality of second distances, wherein the second distance and the first distance are mapped into a one-to-one corresponding relationship through the same historical data point; judging whether there is a second historical data point with a first distance smaller than the corresponding second distance in the plurality of second historical data points corresponding to each first data point, and if there is a second historical data point with a first distance smaller than the corresponding second distance, obtaining c first data points with a first distance smaller than the corresponding second distance, to generate c second data points, wherein c is less than or equal to b and c is a positive integer; obtaining the first distance of each second data point smaller than the corresponding second historical data point, to generate a related data point of each second data point; and filtering the d abnormal data points from the c second data points based on the related data point of each second data point, wherein d is less than or equal to c and d is a positive integer.

[0006] According to the embodiment of the present disclosure, the method further comprises: inputting the a new service download data into the pre-trained and optimized spatial clustering model, to output e outlying new service download data points, wherein e is a positive integer, and the sum of b and e is equal to a; and determining e abnormal service data according to the e outlying new service download data points.

[0007] According to the embodiment of the present disclosure, the method further comprises: obtaining f first data points in which there is no second historical data point whose first distance is less than the second distance corresponding thereto, and generating f third data points, wherein f is less than or equal to b and f is a positive integer; obtaining the number of second historical data points within the domain radius of each third data point, and if the number of second historical data points within the domain radius of the third data point is greater than or equal to a preset threshold, obtaining g third data points from the f third data points in which the number of second historical data points within the domain radius is greater than or equal to the preset threshold, and generating g fourth data points, wherein g is less than or equal to f and g is a positive integer; and determining g pieces of abnormal business data based on the g fourth data points.

[0008] According to the embodiment of the present disclosure, the first historical data point comprises a feature label, and the method further comprises: if the number of second historical data points within the domain radius of the third data point is less than the preset threshold, obtaining h third data points from the f third data points in which the number of second historical data points within the domain radius is less than the preset threshold, and generating h fifth data points, wherein the sum of h and g is equal to f, h is a positive integer; and obtaining the feature label of each second historical data point within the domain radius of each fifth data point and accumulating the feature label, selecting i abnormal data points from the h second data points, and determining i pieces of abnormal business data based on the i abnormal data points, wherein i is less than or equal to h and i is a positive integer.

[0009] According to the embodiment of the present disclosure, the method further comprises: based on the a pieces of newly added business download data, performing optimal parameter tuning on the pre-trained spatial clustering model by using a grid search method, and generating a tuned spatial clustering model, wherein the optimal parameters comprise a domain radius.

[0010] According to the embodiment of the present disclosure, the first data point comprises a domain radius of the batch of newly added business download data points, and the method further comprises: obtaining multiple batches of historical newly added business download data, wherein each batch of historical newly added business download data comprises coordinates of multiple business download data points, the coordinates of the multiple business download data points in the same batch of historical newly added business download data have the same domain radius, and the coordinates of the multiple business download data points in different batches of historical newly added business download data have different domain radii; and importing the multiple batches of historical newly added business download data into the same coordinate graph to generate the first density reachable graph.

[0011] According to an embodiment of the present disclosure, each of the first historical data points comprises a feature label, and the screening of the d abnormal data points from the c second data points based on the related data points of each second data point comprises: obtaining the feature labels of the related data points of each second data point, and generating a plurality of related feature labels of each second data point; accumulating the plurality of related feature labels of each second data point to generate a feature label of each second data point; and screening the d abnormal data points from the c second data points based on the feature label of each second data point.

[0012] According to a second aspect of the present disclosure, an abnormal service data determination apparatus is provided, and the apparatus comprises: a first obtaining module configured to obtain a plurality of pieces of newly added service download data, input the a pieces of newly added service download data into a pre-trained and tuned spatial clustering model, and output b pieces of non-outlier newly added service download data points as b first data points, wherein each first data point comprises a coordinate of a newly added service download data point, a is greater than or equal to 2, b is less than or equal to a, and a and b are positive integers; a first generating module configured to input the coordinates of the b first data points into a first density reachable graph generated in advance, and generate a second density reachable graph; and a first determining module configured to screen d abnormal data points from the b first data points based on the second density reachable graph, and determine d abnormal service data according to the d abnormal data points, wherein d is less than or equal to b, and d is a positive integer.

[0013] According to the embodiment of the present disclosure, the first density reachable graph comprises outlying point coordinates in the historical service download data points, the second determining module comprises: a fourth generating module configured to obtain the field radius of the outlying point coordinates in the historical service download data points from the second density reachable graph, and generate a plurality of first historical data points; a fifth generating module configured to obtain a plurality of first historical data points within the field radius of each first data point from the second density reachable graph, and generate a plurality of second historical data points corresponding to each first data point; a sixth generating module configured to calculate the Euclidean distance between each first data point and the plurality of second historical data points, and generate a plurality of first distances of each first data point, wherein the field radius of each second historical data point generating the first distance with the first data point is not the same; a seventh generating module configured to obtain the field radius of each second historical data point generating the first distance with the first data point, and generate a plurality of second distances, wherein the second distance and the first distance are mapped into a one-to-one corresponding relationship through the same historical data point; an eighth generating module configured to determine whether there is a second historical data point with a first distance smaller than the corresponding second distance in the plurality of second historical data points corresponding to each first data point, and if there is a second historical data point with a first distance smaller than the corresponding second distance, obtain c first data points with a first distance smaller than the corresponding second distance, and generate c second data points, wherein c is less than or equal to b, and c is a positive integer; a ninth generating module configured to obtain the first distance of each second data point smaller than the corresponding second historical data point, and generate the related data point of each second data point; and a first screening module configured to screen d abnormal data points from the c second data points based on the related data point of each second data point, wherein d is less than or equal to c, and d is a positive integer.

[0014] According to the embodiment of the present disclosure, the first obtaining module comprises: a second generating module configured to input the a new service download data into a pre-trained and optimized spatial clustering model, and output e outlying new service download data points, wherein e is a positive integer, and the sum of b and e is equal to a; and a second determining module configured to determine e abnormal service data according to the e outlying new service download data points.

[0015] According to the embodiment of the present disclosure, the second obtaining module comprises: a third obtaining module, configured to obtain f first data points without a second historical data point whose first distance is less than the corresponding second distance, and generate f third data points, wherein f is less than or equal to b and f is a positive integer; a tenth generating module, configured to obtain the number of second historical data points within the domain radius of each third data point, and if the number of second historical data points within the domain radius of the third data point is greater than or equal to a preset threshold, obtain g third data points from the f third data points, wherein the number of second historical data points within the domain radius of the third data points is greater than or equal to the preset threshold, and generate g fourth data points, wherein g is less than or equal to f and g is a positive integer; and a third determining module, configured to determine g abnormal business data based on the g fourth data points.

[0016] According to the embodiment of the present disclosure, the tenth generating module comprises: a tenth generating module, configured to obtain h third data points from the f third data points, wherein the number of second historical data points within the domain radius of the third data points is less than the preset threshold, and generate h fifth data points, wherein the sum of h and g is equal to f, h is a positive integer; and a fourth determining module, configured to obtain the feature label of each fifth data point within the domain radius of the second historical data point and accumulate, select i abnormal data points from the h second data points, and determine i abnormal business data according to the i abnormal data points, wherein i is less than or equal to h and i is a positive integer.

[0017] According to the embodiment of the present disclosure, the first obtaining module further comprises: an optimization module, configured to perform optimal parameter optimization on the pre-trained spatial clustering model based on the a new business download data by using the grid search method, and generate an optimized spatial clustering model, wherein the optimal parameters comprise a domain radius.

[0018] According to the embodiment of the present disclosure, the first generating module comprises: a second obtaining module, configured to obtain multiple batches of historical new business download data, wherein each batch of historical new business download data comprises the coordinates of multiple business download data points, the coordinates of the multiple business download data points in the same batch of historical new business download data have the same domain radius, and the coordinates of the multiple business download data points in different batches of historical new business download data have different domain radii; and a third generating module, configured to import the multiple batches of historical new business download data into the same coordinate graph, and generate the first density reachable graph.

[0019] According to the embodiment of the present disclosure, the first screening module comprises: an eleventh generation module configured to obtain feature labels of the related data points of each second data point, and generate a plurality of related feature labels of each second data point; a twelfth generation module configured to accumulate the plurality of related feature labels of each second data point, and generate a feature label of each second data point; and a fifth determination module configured to screen d abnormal data points from the c second data points based on the feature label of each second data point.

[0020] According to a third aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a storage device configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the above-mentioned abnormal business data determination method.

[0021] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, which stores executable instructions, and the instructions are executed by a processor to perform the above-mentioned abnormal business data determination method.

[0022] According to a fifth aspect of the present disclosure, a computer program product is also provided, comprising a computer program, which is executed by a processor to implement the above-mentioned abnormal business data determination method.

[0023] The embodiment of the present disclosure can realize the technical effect of excluding a large number of normal business report download behaviors from massive data, automatically identifying suspected abnormal business report download behaviors, greatly reducing the workload of personnel determination, and improving work efficiency and processing speed, by applying the optimized spatial clustering model to obtain the first data points, introducing the first data points into the first density reachable graph, and screening the first data points in combination with the first density reachable graph to screen out abnormal data points and determine the technical means of abnormal data. BRIEF DESCRIPTION OF DRAWINGS

[0024] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description of embodiments of the present disclosure, taken in conjunction with the accompanying drawings, in which:

[0025] Figure 1 An application scenario diagram of the abnormal business data determination method and device according to the embodiment of the present disclosure is schematically shown;

[0026] Figure 2 A flowchart of the abnormal business data determination method according to the embodiment of the present disclosure is schematically shown;

[0027] Figure 3 A DBSCAN clustering example diagram in the abnormal business data determination method according to the embodiment of the present disclosure is schematically shown;

[0028] Figure 4 An example diagram of grid search in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0029] Figure 5 A flowchart of determining abnormal business data by outputting outlier data points through a spatial clustering model in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0030] Figure 6 A schematic diagram of generating substitute data points in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure

[0031] Figure 7 A flowchart of generating a first density reachable graph in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0032] Figure 8 A flowchart of generating d abnormal data points in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0033] Figure 9 A flowchart of determining abnormal business data if there is no first data point of the first distance smaller than the second distance corresponding to the second historical data in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0034] Figure 10 A flowchart of determining abnormal business data if the second historical data points within the third data point field radius are smaller than a preset threshold in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0035] Figure 11 A flowchart of screening d abnormal data points in the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0036] Figure 12 A structural block diagram of the device for determining abnormal business data is schematically shown according to an embodiment of the present disclosure;

[0037] Figure 13 A block diagram of an electronic device suitable for implementing the method for determining abnormal business data is schematically shown according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. It should be understood, however, that the description is merely illustrative of the present disclosure, and does not limit the scope of the present disclosure. In the following detailed description of the embodiments, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one skilled in the art that the present disclosure can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present disclosure.

[0039] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used herein, the term "includes" and tautological expressions thereof, such as "including," "includes," "include," "contains," "containing," and so forth, shall not be taken to exclude

[0040] All terms used herein including technical and scientific terms have the same meanings as commonly understood by one of ordinary skill in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning that is consistent with the context of the specification, and not be interpreted in an idealized or overly formal way.

[0041] In the case where expressions similar to "at least one of A, B, and C, etc." are used, it is generally construed that the meaning of the expression is the same as that of "one or more of the members of a set including A, B, and C" (i.e., A or B or C or any combination of the members thereof). In the case where expressions similar to "one or more of A, B, and C, etc." are used, it is generally construed that the meaning of the expression is the same as that of "at least one of A, B, and C" (i.e., A or B or C or any combination of the members thereof).

[0042] Some of the blocks and / or flowcharts in the drawings represent computer program instructions or programs. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to create means for implementing the functions / acts specified in the block diagrams and / or flowchart block or blocks.

[0043] First, technical terms appearing in the present document are explained as follows:

[0044] The DBSCAN (Density-Based Spatial Clustering of Applications with Noise) model is a density-based spatial clustering model that is used to cluster data points in space. The core concept of DBSCAN is to define clusters around the density of data points. It can identify high-density regions separated by low-density regions and treat these regions as independent clusters.

[0045] Density reachable graph, the graph generated by DBSCAN is called density reachable graph.

[0046] Euclidean distance, the straight-line shortest distance between two points in multidimensional space.

[0047] Grid search, a method for selecting the optimal hyperparameters of the model, which finds the best parameter settings by exhaustively searching all possible hyperparameter combinations.

[0048] Cartesian product, a basic concept in set theory, used to describe the set of all possible ordered pairs between two sets.

[0049] The embodiment of the disclosure provides an abnormal business data determination method, which comprises: obtaining a new business download data, inputting the a new business download data into a pre-trained and optimized spatial clustering model, outputting b non-outlier new business download data points as b first data points, wherein each first data point comprises the coordinates of the new business download data point, wherein a is greater than or equal to 2, b is less than or equal to a, and a and b are positive integers. The coordinates of the b first data points are introduced into the first density reachable graph generated in advance to generate a second density reachable graph. And based on the second density reachable graph, d abnormal data points are selected from the b first data points, and d abnormal business data are determined according to the d abnormal data points, wherein d is less than or equal to b, and d is a positive integer.

[0050] The embodiment of the disclosure can realize the exclusion of a large number of normal business report download behaviors from massive data, automatic identification of suspected abnormal business report download behaviors, greatly reduce the workload of personnel determination, improve the work efficiency and processing speed, and achieve the technical effect of determining abnormal business data.

[0051] Figure 1 The application scenario of the abnormal business data determination method and device according to the embodiment of the disclosure is schematically shown. It should be noted that, Figure 1 The shown is only an example of the scene to which the embodiment of the disclosure can be applied, to help those skilled in the art understand the technical content of the disclosure, but does not mean that the embodiment of the disclosure cannot be used in other devices, systems, environments or scenes.

[0052] As Figure 1As shown, the application scenario 100 according to this embodiment can include a plurality of application terminals and an application server. For example, the plurality of application terminals include an application terminal 101, an application terminal 102, an application terminal 103, and the like. A network 104 is a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or fiber optic cables, and the like.

[0053] A user can use the application terminal devices 101, 102, 103 to interact with the application server 105 through the network 104 to receive or send messages, and the like. Various application programs can be installed on the application terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, and the like (only as examples).

[0054] The terminal devices 101, 102, 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, and the like.

[0055] The server 105 can be a server providing various services, such as a background management server providing support for a website browsed by a user using the terminal device 101, 102, 103 (only as an example). The background management server can analyze and process received user requests and the like, and feed back the processing results (such as a webpage, information, or data, and the like obtained or generated according to a user request) to the terminal device.

[0056] It should be noted that the abnormal business data determination method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the abnormal business data determination apparatus provided by the embodiments of the present disclosure can generally be arranged in the server 105. The abnormal business data determination method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the abnormal business data determination apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0057] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the application scenario 100 is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks, and servers.

[0058] The following will be described based on Figure 1 the scenario described in the foregoing embodiments, by Figures 2-11The abnormal service data determination method of the disclosed embodiments is described in detail. It should be noted that the above application scenarios are only shown for the purpose of facilitating understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0059] Figure 2 A flowchart of the abnormal service data determination method according to the embodiments of the present disclosure is schematically shown.

[0060] As shown in Figure 2 , the method 200 includes steps S201-S203.

[0061] In step S201, a number of newly added service download data are obtained, and the a number of newly added service download data are input into a pre-trained and optimized spatial clustering model, and b number of non-outlier newly added service download data points are output as b number of first data points, wherein each first data point includes a coordinate of a newly added service download data point and a field radius of the batch of newly added service download data points, wherein a is greater than or equal to 2, b is less than or equal to a, and a and b are positive integers.

[0062] For example, the newly added service download data includes newly added service download behavior data. A number of newly added service download behavior data can be obtained, and the a number of newly added service download behavior data are input into a pre-trained and optimized spatial clustering model, and b number of non-outlier newly added service download data points are output as b number of first data points.

[0063] In the DBSCAN clustering model, the maximum radius of the core point is defined as ε. If the distance between two data points is less than or equal to ε, they will be judged as the same class. In other words, ε is the distance threshold used by DBSCAN to determine whether two points are similar and belong to the same class, and ε is also referred to as the field radius of the DBSCAN clustering model. If the number of points contained in the ε neighborhood of a point is not less than MinPts, the point is regarded as a core point, and the initial point is also contained in MinPts, and MinPts is also referred to as the minimum number of points in the neighborhood of the DBSCAN clustering model.

[0064] Figure 3 A DBSCAN clustering example diagram in the abnormal service data determination method according to the embodiments of the present disclosure is schematically shown.

[0065] As shown in Figure 3As shown, there are 6 points in A, and if MinPts is less than or equal to 6, all the points in this part are judged as the same class. Specifically, DBSCAN divides all data points into three categories: core point: if the density of a point exceeds the threshold set by the algorithm, the point is a core point. That is, the number of points in the ε neighborhood of the point is not less than MinPts, where ε and MinPts are pre-set prior values. For example, Figure 1 The point A in FIG. 1 is a core point. If the density of a point is less than the threshold set by the algorithm, but the point is in the ε neighborhood of a core point, the point is a boundary point. The boundary point is a neighbor of the core point, but does not meet the condition to become a core point. For example, Figure 1 The point B in FIG. 1 is a boundary point. The remaining points that are neither core points nor boundary points are regarded as outliers or noise points. These points can be outliers in the data or points that do not belong to any cluster. For example, Figure 1 The point C in FIG. 1 is an outlier or noise point. By determining the core points, boundary points, and outliers, DBSCAN can discover clusters with different densities and shapes, and is more robust in handling noisy data than some other algorithms.

[0066] In the embodiments of the present disclosure, the historical business data can be obtained first, and the DBSCAN clustering model is trained by the historical business data. The DBSCAN clustering model is continuously improved by the DBSCAN clustering model training, so that the reliability of the DBSCAN clustering model is gradually improved. The specific method of training the DBSCAN clustering model includes: obtaining historical business data, performing data cleaning and feature selection on the historical business data according to the business scenario, ensuring the quality and suitability of the data, generating historical business data after data cleaning and feature selection, classifying the historical business data, and generating historical sample training data and historical verification training data. Then, the best parameter value of the DBSCAN clustering model is obtained, and the parameter of the DBSCAN clustering model is optimized. The parameter value here includes the radius domain and the minimum number of points in the neighborhood. Then, the historical sample training data is input into the DBSCAN clustering model after parameter optimization for model training, and then the historical verification training data is input into the trained DBSCAN clustering model for evaluation, and a training evaluation result is generated. Based on the training evaluation result, the parameter value of the DBSCAN clustering model is adjusted, and the above training and verification steps are repeated, and the DBSCAN clustering model parameters are continuously adjusted and the DBSCAN clustering model is optimized until the expected training evaluation result is reached, the training of the DBSCAN clustering model is completed, and the expected DBSCAN clustering model is generated.

[0067] The DBSCAN clustering model is very sensitive to the selection of parameters, and different parameters can lead to different clustering results. Therefore, selecting appropriate parameters for parameter tuning is an important step in applying the DBSCAN clustering model. In the embodiments of the present disclosure, the best parameter tuning of the pre-trained spatial clustering model can be performed based on the a pieces of new business download data by using the grid search method, to generate a tuned spatial clustering model. The best parameters include: a domain radius and a minimum number of points in a neighborhood MinPts. Through model parameter tuning, the accuracy of the DBSCAN clustering model output is realized, and the reliability of the DBSCAN clustering model is improved.

[0068] Figure 4 An example grid search diagram in the abnormal business data determination method according to the embodiments of the present disclosure is schematically shown.

[0069] As shown in Figure 4 , the grid search is an exhaustive search method that finds the optimal hyperparameters by traversing all possible combinations of hyperparameters. According to the application scenario, the maximum value of is set to ( , which represents the dimension of the data point feature), the minimum value is 0.001, and the step size is 0.001. The maximum value of MinPts is set to , the minimum value is set to 1, and the step size is set to 1. The setting idea of the maximum value here is that the data points at this time only have two features: normal business report download behavior and abnormal business report download behavior, so the maximum value of MinPts is set to , where represents the floor function. Next, the Cartesian product of ε and MinPts is generated to form a combination grid of hyperparameters. Then, the grid search performs DBSCAN clustering training and evaluation on each hyperparameter combination to find the hyperparameter combination with the best clustering accuracy. The grid search is shown in Figure 4 , where the black dots represent a hyperparameter combination.

[0070] In addition, inputting the a pieces of new business download data into the pre-trained and tuned spatial clustering model can also output e pieces of outlier new business download data points for determining abnormal business data points.

[0071] Figure 5 A flowchart of determining abnormal business data by outputting outlier data points by a spatial clustering model in the abnormal business data determination method according to the embodiments of the present disclosure is schematically shown.

[0072] As shown in Figure 5 , the method 500 includes steps S501-S502.

[0073] Step S501, input the a new service download data into the pre-trained and optimized spatial clustering model, and output e outlier new service download data points, wherein e is a positive integer, and the sum of b and e is equal to a.

[0074] Step S502, determine e abnormal service data according to the e outlier new service download data points.

[0075] For example, the accuracy of the model outputting e outlier new service download data points can be verified. If the verification is indeed a real outlier new service download data point, it is directly determined as an abnormal service data. If the verified outlier new service download data point is a false outlier new service download data point, the false outlier data point output in DBSCAN can be removed according to the real data label, and a replacement data point is generated and saved.

[0076] Figure 6 An illustrative diagram of generating a replacement data point in the abnormal service data determination method according to an embodiment of the present disclosure is shown.

[0077] As shown in Figure 6 (A, B, C, D) is a cluster of DBSCAN, C is a false outlier new service download data point, and C point needs to be removed from the cluster, and the label of C and the distance of the nearest point in the cluster are retained. Then, the center point E of the cluster is calculated according to (A, B, D) remaining after removing the C point. E is a replacement data point, which can be obtained by calculating the average value of the feature vectors of the remaining points, that is:

[0078]

[0079] wherein, represents the feature vector of the i-th point in the cluster, represents the number of remaining data points in the cluster after removing the error point.

[0080] The replacement data point E can then be determined as abnormal data according to the abnormal data determination step of the non-outlier data point.

[0081] By using the spatial clustering model to determine the outlier new service download data point and the abnormal service data, the dependence on manual intervention is reduced, the risk of human error is reduced, the automation degree and stability are improved, and the labor cost is reduced.

[0082] Referring back to Figure 2 ​In step S202, the coordinates of the b first data points are imported into a pre-generated first density reachability map to generate a second density reachability map. Each first data point includes the neighborhood radius of the newly added service download data points in this batch. The first density reachability map includes the coordinates of outliers in the historical service download data points.

[0083] Figure 7 The flowchart illustrating the generation of a first density reachability graph in the abnormal business data determination method according to an embodiment of the present disclosure is shown.

[0084] like Figure 7 As shown, the method 700 includes steps S701 to S702.

[0085] Step S701: Obtain multiple batches of historical new service download data. Each batch of historical new service download data includes the coordinates of multiple service download data points. The coordinates of multiple service download data points in the same batch of historical new service download data have the same domain radius, while the coordinates of multiple service download data points in different batches of historical new service download data have different domain radii.

[0086] For example, in the acquisition of multiple batches of historical new business download data, since parameter tuning is required for each acquisition of historical new business download data, the parameters of the clustering model are not the same for each batch. That is, the neighborhood radius and the minimum number of points in the neighborhood of the data points of each batch of historical new business download data are not necessarily the same.

[0087] Step S702: Import the multiple batches of historical new service download data into the same coordinate graph to generate a first density reachable graph.

[0088] For example, importing historical new business download data from multiple batches with different domain radii into the same coordinate graph generates a first density reachability graph. This first density reachability graph includes coordinates from multiple batches of historical new business download data with different domain radii.

[0089] Historical data is graphically displayed based on a pre-defined first density reachability map, which facilitates calculations in conjunction with newly added data points and improves the accuracy of identifying outliers.

[0090] After generating the first density reachability map, the coordinates of b first data points are imported into the pre-generated first density reachability map to generate a second density reachability map. The generated second density reachability map includes the coordinates of multiple batches of historical new business download data with different domain radii included in the original first density reachability map, as well as the coordinates of the newly imported b first data points. The coordinates of the b first data points have the same domain radius, but are not necessarily the same as the domain radius of the multiple batches of historical new business download data coordinates.

[0091] Referring back to Figure 2 In step S203, based on the second density reachable graph, d abnormal data points are screened out from the b first data points, and d abnormal service data are determined according to the d abnormal data points, where d is less than or equal to b, and d is a positive integer.

[0092] Figure 8 A flowchart of generating d abnormal data points in the abnormal service data determination method according to an embodiment of the present disclosure is schematically shown.

[0093] As Figure 8 shown, the method 800 includes steps S801-S807.

[0094] In step S801, the field radius of the outlier point coordinates in the historical service download data points is obtained from the second density reachable graph, and a plurality of first historical data points are generated.

[0095] For example, the outlier point coordinates in the historical service download data points herein include outlier point coordinates in multiple batches of historical data, where the field radius of the outlier point coordinates in the same batch of historical service download data points is the same, and the field radius of the outlier point coordinates in different batches of historical service download data points may be the same or different. The field radius mainly depends on the selection of model parameters when inputting historical data. Since the historical service download data points generally include a set of identification values and a feature label of whether to belong to abnormal service report download behavior, the set of identification values includes the number of customer information in the downloaded service report, whether the service report file contains sensitive data, download time, download IP address, employee permission, and the like. Therefore, the generated first historical data points can include the feature label.

[0096] In step S802, a plurality of first historical data points within the field radius of each first data point are obtained from the second density reachable graph, and a plurality of second historical data points corresponding to each first data point are generated.

[0097] For example, based on the second density reachable graph, a plurality of first historical data points within the field radius of each first data point are obtained, and a plurality of second historical data points corresponding to each first data point are generated. That is, a clustering cluster with each first data point as the core point and a plurality of first historical data points within the field radius of each first data point as the edge point is obtained, and the edge points of the clustering cluster formed by each first data point are extracted to generate a plurality of second historical data points corresponding to each first data point.

[0098] In step S803, the Euclidean distance between each first data point and the plurality of second historical data points is calculated, and a plurality of first distances of each first data point are generated, where the field radius of each second historical data point generating the first distance with the first data point is not the same.

[0099] For example, the Euclidean distance between each first data point and the plurality of second historical data points can be calculated to generate a plurality of first distances for each first data point, each first distance being formed by a first data point and a second historical data point, and each first data point corresponding to a plurality of first distances. Since the plurality of second historical data points are all within the domain radius of the first data point, all the first distances are less than the domain radius of the first data point.

[0100] Step S804, obtaining the domain radius of each second historical data point corresponding to the first data point to generate a plurality of second distances, wherein the second distance and the first distance are mapped into a one-to-one correspondence by the same historical data point.

[0101] For example, the domain radius of each second historical data point corresponding to each first data point can be obtained to generate a plurality of second distances, since each second distance is the domain radius of the second historical data point, each second distance has the same second historical data point as each first distance, and the second distance and the first distance can be mapped into a one-to-one correspondence by the same historical data point.

[0102] Step S805, determining whether there is a second historical data point whose first distance is less than the corresponding second distance in the plurality of second historical data points corresponding to each first data point, if there is a second historical data point whose first distance is less than the corresponding second distance, obtaining c first data points whose first distance is less than the corresponding second distance, and generating c second data points, wherein c is less than or equal to b, and c is a positive integer.

[0103] For example, the first distance and the second distance of each second historical data point corresponding to each first data point are compared, and the first data point whose first distance is less than the corresponding second distance is filtered out. Since each first data point can correspond to a plurality of second historical data points, as long as the first data point has a second historical data point whose first distance is less than the corresponding second distance, the first data point satisfying the condition is obtained, that is, c first data points satisfying the above condition are obtained from b first data points as c second data points.

[0104] In addition, if it is determined that there is no first data point whose first distance is less than the corresponding second distance, a preset threshold method can be used to determine the abnormal business data.

[0105] Figure 9Fig. 2 schematically shows a flowchart of a method for determining abnormal business data if there is no second historical data first data point whose first distance is less than a second distance corresponding thereto according to an embodiment of the present disclosure;

[0106] As shown in Fig. 2, the method 900 includes steps S901-S903. Figure 9

[0107] In step S901, f first data points, for which there is no second historical data point whose first distance is less than a second distance corresponding thereto, are obtained, and f third data points are generated, where f is less than or equal to b and f is a positive integer.

[0108] In step S902, the number of second historical data points within the field radius of each third data point is obtained, and if the number of second historical data points within the field radius of the third data point is greater than or equal to a preset threshold, g third data points, for which the number of second historical data points within the field radius is greater than or equal to the preset threshold, are obtained from the f third data points, and g fourth data points are generated, where g is less than or equal to f and g is a positive integer.

[0109] If the number of second historical data points within the field radius of the third data point is greater than or equal to the preset threshold, the abnormal business data can be determined in a feature label accumulation manner.

[0110] Figure 10 Fig. 3 schematically shows a flowchart of a method for determining abnormal business data if the number of second historical data points within the field radius of the third data point is less than the preset threshold according to an embodiment of the present disclosure.

[0111] As shown in Fig. 3, the method 1000 includes steps S1001-S1002. Figure 10

[0112] In step S1001, if the number of second historical data points within the field radius of the third data point is less than the preset threshold, h third data points, for which the number of second historical data points within the field radius is less than the preset threshold, are obtained from the f third data points, and h fifth data points are generated, where the sum of h and g is equal to f, and h is a positive integer.

[0113] In step S1002, the feature labels of the second historical data points within the field radius of each fifth data point are obtained and accumulated, i abnormal data points are selected from the h second data points, and i abnormal business data are determined according to the i abnormal data points, where i is less than or equal to h and i is a positive integer.

[0114] ​​For example, the feature label of the second historical data point within the radius of each fifth data point field can be acquired as the multiple related labels of each fifth data point, the multiple related labels are accumulated to generate the feature label of each fifth data point. According to the feature label of each fifth data point, i abnormal data points are screened from h second data points, and i abnormal service data are determined according to the i abnormal data points.

[0115] The third data point is screened through the accumulation of the feature label, and the accuracy of the abnormal data acquisition is further improved.

[0116] Referring back to Figure 9 In step S903, g abnormal service data are determined based on the g fourth data points.

[0117] The third data point is screened through the preset threshold, and the accuracy of the abnormal data acquisition is improved.

[0118] Referring back to Figure 8 In step S S806, the second historical data point with a first distance less than a second distance corresponding to each second data point is acquired from c second data points to generate the related data point of each second data point.

[0119] For example, the second historical data point with a first distance less than a second distance corresponding to each second data point is acquired from c second data points to generate the related data point of each second data point.

[0120] Step S807, based on the related data point of each second data point, d abnormal data points are screened from c second data points, wherein d is less than or equal to c, and d is a positive integer.

[0121] Figure 11 A flowchart of screening d abnormal data points in the abnormal service data determination method according to an embodiment of the present disclosure is schematically shown.

[0122] As Figure 11 shown, the method 1100 includes steps S1101-S1103.

[0123] Step S1101, the feature label of the related data point of each second data point is acquired to generate multiple related feature labels of each second data point.

[0124] For example, in the multiple related feature labels of one of the second data points, the feature label of 1 normal download behavior and the feature label of 2 normal download behaviors, at this time, the related feature label of this second data point can be expressed as [normal download behavior, abnormal download behavior] = [1, 2].

[0125] Step S1102, the plurality of related feature labels of each second data point are accumulated to generate a feature label of each second data point.

[0126] For example, the related feature label of the example second data point in step S1101 can be represented as [normal download behavior, abnormal download behavior] = [1, 2], and after accumulation, the feature label of this second data point is abnormal behavior. Accumulating the related feature labels can include comparing the normal download behavior and the abnormal download behavior in the related feature labels, if the normal download behavior is greater than the abnormal download behavior, the feature label of this second data point is normal behavior, and if the normal download behavior is less than or equal to the abnormal download behavior, the feature label of this second data point is abnormal behavior.

[0127] Step S1103, based on the feature label of each second data point, screening d abnormal data points from the c second data points, and determining d abnormal service data according to the d abnormal data points.

[0128] The feature label of each second data point is determined by accumulating the related data point labels of each second data point, which improves the reliability of determining abnormal data. And by obtaining c second data points that have a first distance less than a corresponding second distance, and combining the related data points of each second data point to screen d abnormal data points, the accuracy of determining d abnormal data can be improved, and the reliability of identifying suspected abnormal service report download behavior can be improved.

[0129] In addition, the abnormal service data determined by the above method can be further judged by error index calculation, residual analysis and other methods to generate a final result. And save the final result to the historical database, and update the historical data iteratively.

[0130] According to embodiments of this disclosure, the DBSCAN model effectively identifies clusters with different densities and shapes, demonstrating excellent performance in processing noisy data. By using a DBSCAN-based clustering method combined with outlier detection, business report download behavior is accurately classified. Grid search technology is employed to obtain optimal parameter values, and by correcting the labels and related information of false data points, the accuracy and generalization ability of the algorithm are improved, enhancing the precision and robustness of anomaly detection. Correcting false data points during the training phase effectively reduces the false alarm rate, minimizes interference with normal business operations, and enhances the reliability and stability of the algorithm. Calculating and comparing generated data points more accurately determines their classification, reducing the workload of manual judgment. Secondary judgment of anomaly results ensures accuracy and reliability, and all data point feature information and labels are fed back to the historical dataset for updates, further enhancing system reliability. By continuously updating the historical database and optimizing the model, the adaptability to new download behaviors is improved, forming a cyclical feedback mechanism that maintains the model's efficiency and accuracy. By automating feature learning and parameter tuning, this method reduces reliance on manual intervention, lowers the risk of human error, improves automation and stability, and reduces labor costs. It also automatically identifies suspected abnormal business report downloads by filtering out a large number of normal business report download behaviors from massive datasets, significantly reducing the workload of manual judgment and improving work efficiency and processing speed.

[0131] Figure 12 A schematic block diagram of an abnormal service data determination apparatus according to an embodiment of the present disclosure is shown.

[0132] like Figure 12 As shown, the device 1200 includes: a first acquisition module 1201, a first generation module 1202, and a first determination module 1203.

[0133] The first acquisition module 1201 is used to acquire *a* new service download data points, input the *a* new service download data points into a pre-trained and optimized spatial clustering model, and output *b* non-outlier new service download data points as *b* first data points. Each first data point includes the coordinates of the new service download data point and the neighborhood radius of the new service download data points in this batch. *a* is greater than or equal to 2, *b* is less than or equal to *a*, and both *a* and *b* are positive integers. In one embodiment, the first acquisition module 1201 can be used to execute step S201 described above, which will not be repeated here.

[0134] The first generation module 1202 is configured to introduce the coordinates of the b first data points into a pre-generated first density reachable graph to generate a second density reachable graph, where the first density reachable graph includes coordinates of outlier points in historical service download data. In an embodiment, the first generation module 1202 can be configured to perform the step S202 described above, and details are not repeated here.

[0135] The first determination module 1203 is configured to filter d abnormal data points from the b first data points based on the second density reachable graph, and determine d abnormal service data according to the d abnormal data points, where d is less than or equal to b, and d is a positive integer. In an embodiment, the first determination module 1203 can be configured to perform the step S203 described above, and details are not repeated here.

[0136] According to the embodiments of the present disclosure, the first acquisition module 1201 includes a seventh generation module and a second determination module.

[0137] The second generation module is configured to input the a new service download data into a pre-trained and optimized spatial clustering model to output e outlier new service download data points, where e is a positive integer, and the sum of b and e is equal to a. In an embodiment, the second generation module can be configured to perform the step S501 described above, and details are not repeated here.

[0138] The second determination module is configured to determine e abnormal service data according to the e outlier new service download data points. In an embodiment, the second determination module can be configured to perform the step S502 described above, and details are not repeated here.

[0139] According to the embodiments of the present disclosure, the first acquisition module 1201 further includes an optimization module configured to perform optimal parameter optimization on the pre-trained spatial clustering model based on the a new service download data by a grid search method to generate an optimized spatial clustering model, where the optimal parameters include a field radius.

[0140] According to the embodiments of the present disclosure, the first generation module 1202 includes a second acquisition module and a third generation module.

[0141] The second acquisition module is configured to acquire multiple batches of historical new service download data, where each batch of historical new service download data includes coordinates of multiple service download data points, the coordinates of the multiple service download data points in the same batch of historical new service download data have the same field radius, and the coordinates of the multiple service download data points in different batches of historical new service download data have different field radii. In an embodiment, the second acquisition module can be configured to perform the step S701 described above, and details are not repeated here.

[0142] The third generation module is configured to import the multiple batches of historical new service download data into a same coordinate graph to generate a first density reachable graph. In an embodiment, the third generation module can be configured to perform the step S702 described above, and details are not repeated here.

[0143] According to the embodiments of the present disclosure, the first determination module 1203 includes a fourth generation module, a fifth generation module, a sixth generation module, a seventh generation module, an eighth generation module, a ninth generation module, and a first screening module.

[0144] The fourth generation module is configured to obtain the domain radius of the outlier coordinates in the historical service download data points from the second density reachable graph to generate multiple first historical data points. In an embodiment, the fourth generation module can be configured to perform the step S801 described above, and details are not repeated here.

[0145] The fifth generation module is configured to obtain multiple first historical data points within the domain radius of each first data point from the second density reachable graph to generate multiple second historical data points corresponding to each first data point. In an embodiment, the fifth generation module can be configured to perform the step S802 described above, and details are not repeated here.

[0146] The sixth generation module is configured to calculate the Euclidean distance between each first data point and the multiple second historical data points to generate multiple first distances of each first data point, wherein the domain radius of each second historical data point generating a first distance with a first data point is not the same. In an embodiment, the sixth generation module can be configured to perform the step S803 described above, and details are not repeated here.

[0147] The seventh generation module is configured to obtain the domain radius of each second historical data point generating a first distance with a first data point to generate multiple second distances, wherein the second distances and the first distances are mapped into a one-to-one corresponding relationship by the same historical data points. In an embodiment, the seventh generation module can be configured to perform the step S804 described above, and details are not repeated here.

[0148] The eighth generation module is configured to determine whether there is a second historical data point with a first distance smaller than a corresponding second distance in the multiple second historical data points corresponding to each first data point, and if there is a second historical data point with a first distance smaller than a corresponding second distance, obtain c first data points with a first distance smaller than a corresponding second distance to generate c second data points, wherein c is smaller than or equal to b, and c is a positive integer. In an embodiment, the eighth generation module can be configured to perform the step S805 described above.

[0149] The ninth generation module is configured to acquire a second historical data point with a first distance less than a second distance corresponding to each second data point, and generate a related data point of each second data point. In an embodiment, the ninth generation module can be configured to perform the step S806 described above, and details are not repeated here.

[0150] The first screening module is configured to screen d abnormal data points from the c second data points based on the related data points of each second data point, and determine the d abnormal business data according to the d abnormal data points, where d is less than or equal to c, and d is a positive integer. In an embodiment, the first screening module can be configured to perform the step S807 described above.

[0151] According to the embodiments of the present disclosure, the eighth generation module comprises a third acquisition module, a fourth acquisition module and a third determination module.

[0152] The third acquisition module is configured to acquire f first data points without a second historical data point with a first distance less than a second distance corresponding to each first data point, and generate f third data points, where f is less than or equal to b, and f is a positive integer. In an embodiment, the third acquisition module can be configured to perform the step S901 described above, and details are not repeated here.

[0153] The fourth acquisition module is configured to acquire a number of second historical data points within a domain radius of each third data point, and if the number of second historical data points within the domain radius of the third data point is greater than or equal to a preset threshold, acquire g third data points with the number of second historical data points within the domain radius greater than or equal to the preset threshold from the f third data points, and generate g fourth data points, where g is less than or equal to f, and g is a positive integer. In an embodiment, the fourth acquisition module can be configured to perform the step S902 described above.

[0154] The third determination module is configured to determine g abnormal business data based on the g fourth data points. In an embodiment, the third determination module can be configured to perform the step S903 described above, and details are not repeated here.

[0155] According to the embodiments of the present disclosure, the fourth acquisition module comprises a tenth generation module and a fourth determination module.

[0156] The tenth generation module is configured to acquire h third data points with the number of second historical data points within the domain radius less than the preset threshold from the f third data points if the number of second historical data points within the domain radius of the third data point is less than the preset threshold, and generate h fifth data points, where the sum of h and g is equal to f, and h is a positive integer. In an embodiment, the tenth generation module can be configured to perform the step S1001 described above, and details are not repeated here.

[0157] The fourth determining module is configured to obtain and accumulate feature labels of the second historical data points in the radius of each fifth data point field, filter i abnormal data points from the h second data points, and determine i abnormal business data according to the i abnormal data points, where i is less than or equal to h and i is a positive integer. In an embodiment, the fourth determining module can be configured to perform the step S1002 described above, and details are not repeated here.

[0158] According to the embodiments of the present disclosure, the first screening module comprises an eleventh generating module, a twelfth generating module and a fifth determining module.

[0159] The eleventh generating module is configured to obtain feature labels of the related data points of each second data point, and generate a plurality of related feature labels of each second data point. In an embodiment, the twelfth generating module can be configured to perform the step S1101 described above, and details are not repeated here.

[0160] The twelfth generating module is configured to accumulate the plurality of related feature labels of each second data point, and generate a feature label of each second data point. In an embodiment, the twelfth generating module can be configured to perform the step S1102 described above, and details are not repeated here.

[0161] The fifth determining module is configured to filter d abnormal data points from the c second data points based on the feature label of each second data point, and determine d abnormal business data according to the d abnormal data points. In an embodiment, the fifth determining module can be configured to perform the step S1103 described above, and details are not repeated here.

[0162] According to an embodiment of the present disclosure, any of the first obtaining module 1201, the first generating module 1202 and the first determining module 1203 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present disclosure, at least one of the first obtaining module 1201, the first generating module 1202 and the first determining module 1203 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of hardware or firmware that can be integrated or packaged with a circuit, or any one of software, hardware and firmware or any appropriate combination of several of them. Alternatively, at least one of the first obtaining module 1201, the first generating module 1202 and the first determining module 1203 can be at least partially implemented as a computer program module that can perform corresponding functions when it is run.

[0163] Figure 13 A block diagram of an electronic device suitable for implementing the method for determining abnormal service data according to an embodiment of the present disclosure is schematically shown.

[0164] As shown in Figure 13 The electronic device 1300 according to an embodiment of the present disclosure includes a processor 1301 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1302 or loaded from a storage portion 1308 into a random access memory (RAM) 1303. The processor 1301 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and the like. The processor 1301 can also include an on-board memory for cache use. The processor 1301 can include a single processing unit or a plurality of processing units for performing different actions of the method processes according to embodiments of the present disclosure.

[0165] In the RAM 1303, various programs and data required for the operation of the electronic device 1300 are stored. The processor 1301, the ROM 1302, and the RAM 1303 are connected to each other via the bus 1304. The processor 1301 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 1302 and / or the RAM 1303. It should be noted that the programs can also be stored in one or more memories other than the ROM 1302 and the RAM 1303. The processor 1301 can also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.

[0166] According to an embodiment of the present disclosure, the electronic device 1300 can further include an input / output (I / O) interface 1305, which is also connected to the bus 1304. The electronic device 1300 can further include one or more of the following components connected to the I / O interface 1305: an input part 1306 including a keyboard, a mouse, etc.; an output part 1307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 1308 including a hard disk, etc.; and a communication part 1309 including a network interface card such as a LAN card, a modem, etc. The communication part 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as necessary. A removable medium 1311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 1310 as necessary, so that a computer program read out therefrom is installed in the storage part 1308 as necessary.

[0167] The present disclosure also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0168] According to an embodiment of the present disclosure, the computer readable storage medium can be a nonvolatile computer readable storage medium, for example, can include, but is not limited to, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer readable storage medium can include one or more memories, such as the ROM 1302 and / or the RAM 1303 described above, and / or one or more memories other than the ROM 1302 and the RAM 1303.

[0169] Embodiments of the present disclosure also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the abnormal service data determination method provided by the embodiments of the present disclosure.

[0170] The above-described functions defined in the system / apparatus / module / unit of the embodiments of the present disclosure are performed when the computer program is executed by the processor 1301. According to an embodiment of the present disclosure, the system, apparatus, module, unit, etc. described above can be implemented by computer program modules.

[0171] In one embodiment, the computer program can rely on a tangible storage medium, such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal on a network medium, and be downloaded and installed through the communication part 1309, and / or be installed from the detachable medium 1311. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any appropriate combination thereof.

[0172] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1309, and / or be installed from the detachable medium 1311. When the computer program is executed by the processor 1301, the above-described functions defined in the system of the embodiments of the present disclosure are performed. According to an embodiment of the present disclosure, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.

[0173] According to embodiments of the present disclosure, program code of the computer program for performing the methods provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages, and can be implemented in a computer program product. Specifically, the computer program can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. The programming language includes, but is not limited to, Java, C++, python, “C” language, or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's device and partly on a remote computing device, or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider (ISP).

[0174] The computer program code can also be loaded onto a computer (e.g., a server) to cause one or more processors in the computer to perform the functions of the program code. The program code can also be loaded onto a computer (e.g., a server) to cause one or more processors in the computer to perform the functions of the program code. The program code can also be loaded onto a computer (e.g., a server) to cause one or more processors in the computer to perform the functions of the program code.

[0175] Those skilled in the art will understand that features of the various embodiments and / or claims of the present disclosure can be combined or / and integrated with one another, even though such combinations or integrations are not expressly disclosed in the present disclosure. In particular, the features of the various embodiments and / or claims of the present disclosure can be combined or / and integrated with one another in any manner, without departing from the spirit and teachings of the present disclosure. All such combinations and / or integrations are within the scope of the present disclosure.

[0176] The above describes embodiments of the present disclosure. However, these embodiments are merely for illustrative purposes, and are not intended to limit the scope of the present disclosure. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present disclosure, and these substitutions and modifications should all fall within the scope of the present disclosure.

Claims

1. An abnormal traffic data determination method characterized by comprising: The method comprises: obtaining a plurality of newly added service download data, inputting the a plurality of newly added service download data into a pre-trained and optimized spatial clustering model, and outputting b non-outlier newly added service download data points as first data points, wherein each first data point comprises a coordinate of a newly added service download data point, a is greater than or equal to 2, b is less than or equal to a, and a and b are positive integers; introducing the coordinates of the b first data points into a pre-generated first density reachable graph to generate a second density reachable graph; and based on the second density reachable graph, screening d abnormal data points from the b first data points, and determining d abnormal service data according to the d abnormal data points, wherein d is less than or equal to b, and d is a positive integer, wherein each first data point comprises a domain radius of the batch of newly added service download data points, and the pre-generated first density reachable graph comprises: obtaining a plurality of batches of historical newly added service download data, wherein each batch of historical newly added service download data comprises coordinates of a plurality of service download data points, the coordinates of the plurality of service download data points in the same batch of historical newly added service download data have the same domain radius, and the coordinates of the plurality of service download data points in different batches of historical newly added service download data have different domain radii; and introducing the plurality of batches of historical newly added service download data into the same coordinate graph to generate a first density reachable graph.

2. The method of claim 1, wherein, The first density reachable graph comprises coordinates of outlier points in historical service download data points, and the screening of d abnormal data points from the b first data points based on the second density reachable graph comprises: obtaining the domain radius of the coordinates of the outlier points in the historical service download data points from the second density reachable graph to generate a plurality of first historical data points; obtaining a plurality of first historical data points within the domain radius of each first data point from the second density reachable graph to generate a plurality of second historical data points corresponding to each first data point; calculating the Euclidean distance between each first data point and the plurality of second historical data points to generate a plurality of first distances of each first data point, wherein the domain radius of each second historical data point generating a first distance with a first data point is different; obtaining the domain radius of each second historical data point generating a first distance with a first data point to generate a plurality of second distances, wherein the second distances and the first distances are mapped into a one-to-one correspondence relationship through the same historical data points; determining whether there is a second historical data point with a first distance less than a corresponding second distance in the plurality of second historical data points corresponding to each first data point, if there is a second historical data point with a first distance less than a corresponding second distance, obtaining c first data points of the second historical data point with a first distance less than a corresponding second distance to generate c second data points, wherein c is less than or equal to b, and c is a positive integer; obtaining the first distance of each second data point less than the second historical data point corresponding to the second data point to generate a related data point of each second data point; and Screening d abnormal data points from the c second data points based on the related data points of each second data point, wherein d is less than or equal to c, and d is a positive integer.

3. The method of claim 1, wherein, The method further comprises: inputting the a new service download data into the pre-trained and optimized spatial clustering model, and outputting e outlier new service download data points, wherein e is a positive integer, and the sum of b and e is equal to a; and determining e abnormal service data according to the e outlier new service download data points.

4. The method of claim 2, wherein, The method further comprises: obtaining f first data points that do not have a second historical data point with a first distance less than a corresponding second distance, and generating f third data points, wherein f is less than or equal to b, and f is a positive integer; obtaining the number of second historical data points within the domain radius of each third data point, if the number of second historical data points within the domain radius of the third data point is greater than or equal to a preset threshold, obtaining g third data points from the f third data points that have the number of second historical data points within the domain radius greater than or equal to the preset threshold, and generating g fourth data points, wherein g is less than or equal to f, and g is a positive integer; and determining g abnormal service data based on the g fourth data points.

5. The method of claim 4, wherein, The method further comprises: if the number of second historical data points within the domain radius of the third data point is less than the preset threshold, obtaining h third data points from the f third data points that have the number of second historical data points within the domain radius less than the preset threshold, and generating h fifth data points, wherein the sum of h and g is equal to f, and h is a positive integer; and obtaining the feature labels of the second historical data points within the domain radius of each fifth data point and accumulating them, screening i abnormal data points from the h second data points, and determining i abnormal service data according to the i abnormal data points, wherein i is less than or equal to h, and i is a positive integer.

6. The method of claim 1, wherein, Optimizing the pre-trained spatial clustering model comprises: based on the a new service download data, performing optimal parameter optimization on the pre-trained spatial clustering model by a grid search method, and generating an optimized spatial clustering model, wherein the optimal parameters include a domain radius.

7. The method of claim 2, wherein, The method further comprises: obtaining the feature labels of the related data points of each second data point, and generating multiple related feature labels of each second data point; accumulating the multiple related feature labels of each second data point, and generating a feature label of each second data point; and screening d abnormal data points from the c second data points based on the feature label of each second data point.

8. An abnormal service data determination apparatus characterized by comprising: The device comprises: The first obtaining module is configured to obtain a number of pieces of newly added service download data, input the a pieces of newly added service download data into a pre-trained and tuned spatial clustering model, and output b pieces of non-outlier newly added service download data points as b first data points, wherein each first data point includes a coordinate of a newly added service download data point, a is greater than or equal to 2, b is less than or equal to a, and a and b are positive integers; The first generating module is configured to input the coordinates of the b first data points into a pre-generated first density reachable graph, and generate a second density reachable graph; and The first determining module is configured to filter d pieces of abnormal data points from the b first data points based on the second density reachable graph, and determine d pieces of abnormal service data according to the d pieces of abnormal data points, wherein d is less than or equal to b, and d is a positive integer. Each first data point includes a domain radius of a newly added service download data point in the batch, and the pre-generated first density reachable graph includes: Obtaining a plurality of batches of historical newly added service download data, wherein each batch of historical newly added service download data includes coordinates of a plurality of service download data points, the coordinates of the plurality of service download data points in the same batch of historical newly added service download data have the same domain radius, and the coordinates of the plurality of service download data points in different batches of historical newly added service download data have different domain radii; and Inputting the plurality of batches of historical newly added service download data into the same coordinate graph to generate a first density reachable graph. 9.An electronic device comprising: one or more processors; a storage device for storing one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1-7.

11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1-7. The computer program is executed by the processor to implement the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Abnormity detection method and device, electronic equipment and storage medium

    CN111814910A

  • Application anomaly detection method and device, equipment and medium

    CN117131405A