A business data processing method, device, equipment and storage medium
By acquiring and analyzing time series type business data and determining the data status using preset models, the problem of real-time risk monitoring of server business data is solved, and effective guarantees for service safety are achieved.
Patent Information
- Application Number
- CN202111571221.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-21
AI Technical Summary
How to monitor the risks of a large amount of business data generated by the server in real time to prevent immeasurable losses to service security.
By obtaining business data of time series type, determining its characteristic data, and determining the data status based on preset models (including data prediction model and data clustering model), real-time monitoring of risks of business data is achieved.
It realizes the rapid and accurate status determination of business data, reasonably conducts real-time risk monitoring, and ensures service safety.
Smart Images

Figure CN114298194B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technologies, and in particular, to a method, apparatus, device, and storage medium for processing service data. Background Art
[0002] With the rapid development of Internet technologies, a large amount of service data is generated by servers. For example, in an e-commerce scenario, when the purchased goods cannot meet the user's needs, the user will choose to apply for a refund. To protect the reputation of the e-commerce platform and improve user satisfaction, the e-commerce platform will make an advance payment for the user's refund request. These payment data are part of the service data.
[0003] However, if the server does not monitor and limit a large amount of service data, it will cause inestimable losses to the service security of the server. Therefore, how to perform real-time risk monitoring on service data is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, and storage medium for processing service data, which can reasonably perform real-time risk monitoring on service data.
[0005] The technical solutions of the embodiments of the present disclosure are as follows:
[0006] According to a first aspect of the embodiments of the present disclosure, a method for processing service data is provided, and this method can be applied to an electronic device. The method for processing service data may include:
[0007] Obtain first service data corresponding to a time series type within a first time period;
[0008] Determine first feature data of the first service data;
[0009] Based on the first feature data and a preset model, determine the data status of the first service data; the preset model includes a data prediction model; the data prediction model is obtained by performing feature learning training on second service data within a second time period until convergence; the second time period is before the first time period.
[0010] Optionally, the preset model further includes a data clustering model; the data clustering model is obtained by clustering the second service data; the data clustering model includes at least one category feature; one category feature is used to represent one clustering result of the second service data;
[0011] Based on the first feature data and the preset model, determining the data status of the first service data includes:
[0012] Determine the Euclidean distance between the first feature data and each category feature in the data clustering model to obtain a plurality of distance values;
[0013] When the minimum distance value among multiple distance values is greater than a preset threshold, input the first feature data into the data prediction model to obtain the data status of the first service data.
[0014] Optionally, the data status includes a normal status and an abnormal status; determining the data status of the first service data based on the first feature data and a preset model includes:
[0015] When the minimum distance value is less than or equal to the preset threshold, determine that the data status of the first service data is the normal status.
[0016] Optionally, the generation method of the data prediction model includes:
[0017] Obtain the second service data corresponding to the time series type and the data status of the second service data;
[0018] Determine the second feature data; the second feature data includes the feature data of the second service data and the feature data of the data status of the second service data;
[0019] Input the second feature data into the support vector machine model, and perform feature learning training on the support vector machine model until convergence to obtain the data prediction model.
[0020] Optionally, obtaining the first service data corresponding to the time series type within the first time period includes:
[0021] Select, from the original service data corresponding to the time series type, the service data whose data generation time is within the first time period and determine it as the first service data.
[0022] Optionally, determining the first feature data of the first service data includes:
[0023] Perform noise reduction processing on the first service data, and perform feature engineering processing on the denoised first service data to obtain the first feature data.
[0024] Optionally, obtaining the second service data corresponding to the time series type includes:
[0025] Select, from the original service data corresponding to the time series type, the service data whose data generation time is within the second time period and determine it as the second service data.
[0026] Optionally, determining the second feature data includes:
[0027] Perform noise reduction processing on the second service data and the data status of the second service data, and perform feature engineering processing on the denoised second service data and the data status of the second service data to obtain the second feature data.
[0028] According to a second aspect of the embodiments of the present disclosure, a service data processing device is provided, which can be applied to an electronic device and includes: an acquisition unit and a processing unit;
[0029] The acquisition unit is configured to acquire first service data corresponding to a time series type within a first time period;
[0030] The processing unit is configured to determine first feature data of the first service data;
[0031] The processing unit is further configured to determine the data status of the first service data based on the first feature data and a preset model; the preset model includes a data prediction model; the data prediction model is obtained by performing feature learning training on second service data within a second time period until convergence; the second time period is before the first time period.
[0032] Optionally, the preset model further includes a data clustering model; the data clustering model is obtained by clustering the second service data; the data clustering model includes at least one category feature; one category feature is used to represent a clustering result of the second service data;
[0033] The processing unit is specifically configured to:
[0034] Determine the Euclidean distance between the first feature data and each category feature in the data clustering model to obtain a plurality of distance values;
[0035] When the minimum distance value among the plurality of distance values is greater than a preset threshold, input the first feature data into the data prediction model to obtain the data status of the first service data.
[0036] Optionally, the data status includes a normal status and an abnormal status; the processing unit is specifically configured to:
[0037] When the minimum distance value is less than or equal to the preset threshold, determine that the data status of the first service data is the normal status.
[0038] Optionally, the acquisition unit is further configured to acquire second service data corresponding to the time series type and the data status of the second service data;
[0039] The processing unit is further configured to determine second feature data; the second feature data includes the feature data of the second service data and the feature data of the data status of the second service data;
[0040] The processing unit is further configured to input the second feature data into a support vector machine model and perform feature learning training on the support vector machine model until convergence to obtain the data prediction model.
[0041] Optionally, the acquisition unit is specifically configured to:
[0042] From the original business data corresponding to the time series type, select the business data whose data generation time is within the first time period and determine it as the first business data.
[0043] Optionally, the processing unit is specifically configured to:
[0044] Perform noise reduction processing on the first business data, and perform feature engineering processing on the first business data after noise reduction to obtain first feature data.
[0045] Optionally, the obtaining unit is specifically configured to:
[0046] From the original business data corresponding to the time series type, select the business data whose data generation time is within the second time period and determine it as the second business data.
[0047] Optionally, the processing unit is specifically configured to:
[0048] Perform noise reduction processing on the second business data and the data status of the second business data, and perform feature engineering processing on the second business data and the data status of the second business data after noise reduction to obtain second feature data.
[0049] According to the third aspect of the embodiments of the present disclosure, an electronic device is provided, which may include: a processor and a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement any of the optional business data processing methods in the first aspect above.
[0050] According to the fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, and instructions are stored on the computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute any of the optional business data processing methods in the first aspect above.
[0051] According to the fifth aspect of the embodiments of the present disclosure, a computer program product is provided. The computer program product includes computer instructions. When the computer instructions run on an electronic device, the electronic device executes the business data processing method as described in any of the optional implementation manners in the first aspect.
[0052] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.
[0053] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0054] Based on any of the above aspects, in the present disclosure, after the electronic device obtains the first service data corresponding to the time series type within the first time period, it may determine the first feature data of the first service data, and based on the first feature data and the preset model, determine the data status of the first service data. First, since the data type of the first service data is the time series type, the first service data can describe the service data through the dimension of time. Second, since the preset model (including the one obtained by performing feature learning training on the second service data within the second time period until convergence) is used to determine the data status of the service data, the present application can quickly and accurately determine the data status of the first service data through the preset model and the first service data with the data type of time series type, and then reasonably perform real-time risk monitoring on the service data through the data status of the first service data. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.
[0056] Figure 1 The flowchart shows a service data processing method provided by an embodiment of the present disclosure;
[0057] Figure 2 The flowchart shows another service data processing method provided by an embodiment of the present disclosure;
[0058] Figure 3 The flowchart shows another service data processing method provided by an embodiment of the present disclosure;
[0059] Figure 4 The flowchart shows another service data processing method provided by an embodiment of the present disclosure;
[0060] Figure 5 The structural diagram shows a service data processing device provided by an embodiment of the present disclosure;
[0061] Figure 6 The structural diagram shows a terminal provided by an embodiment of the present disclosure;
[0062] Figure 7 The structural diagram shows a server provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0064] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0065] It should also be understood that the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements and / or components.
[0066] The data involved in the present disclosure can be data authorized by the user or fully authorized by all parties.
[0067] As described in the background art, if the e-commerce platform does not monitor and limit the large amount of business data generated by the server, it will cause inestimable losses to the service security of the e-commerce platform. Therefore, how to perform real-time risk monitoring on business data is a technical problem that needs to be solved urgently at present.
[0068] Based on this, the embodiments of the present disclosure provide a method for processing business data. After the electronic device obtains the first business data corresponding to the time series type within the first time period, it can determine the first feature data of the first business data, and based on the first feature data and a preset model, determine the data state of the first business data. First, since the data type of the first business data is the time series type, the first business data can describe the business data through the dimension of time. Second, since the preset model (including being obtained by performing feature learning training on the second business data within the second time period until convergence) is used to determine the data state of the business data, the present application can quickly and accurately determine the data state of the first business data through the preset model and the first business data with the data type of the time series type, and then reasonably perform real-time risk monitoring on the business data through the data state of the first business data.
[0069] The above-mentioned electronic device can be a server, a terminal, or other electronic devices used for real-time risk monitoring of business data. The present disclosure does not limit this.
[0070] When the electronic device is a server, the server can be a single server, or alternatively, it can be a server cluster composed of multiple servers. In some embodiments, the server cluster can also be a distributed cluster. The present disclosure does not limit the specific implementation manner of the server.
[0071] In some other embodiments, the server may further include a database or be connected to a database, and the page data of the web page can be stored in the database.
[0072] When the electronic device is a terminal, the terminal can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc., which can install and use a content community application (such as Kuaishou). The present disclosure does not impose special restrictions on the specific form of the terminal. It can perform human-computer interaction with the user through one or more of a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device.
[0073] As Figure 1 shown, when the business data processing method is applied to an electronic device, the business data processing method may include:
[0074] S101. The electronic device obtains first service data corresponding to a time series type within a first time period.
[0075] Optionally, the service data can be numerical transfer data. For example, when an e-commerce platform compensates a user, the numerical transfer data can be compensation data.
[0076] As described in the background art, when the goods purchased by the user do not meet the user's needs, the user will choose to apply for a refund. In this case, in order to protect the reputation of the e-commerce platform and improve user satisfaction, the e-commerce platform will make an advance compensation for the user's request for a refund. As a result, a compensation data, that is, the first service data, will be generated in the corresponding server of the e-commerce platform.
[0077] When determining the data status of the first service data, the electronic device can obtain the first service data corresponding to the time series type from the server.
[0078] Data of the time series type is defined as a series of tuples of values and times, and can be formally represented as: {(p1,t1),(p2,t2),…(pi,ti),…(pn,tn)}.
[0079] Among them, \(t_1 \lt t_2 \lt \cdots \lt t_n\), \(p_i\) is a data point in a \(d\) (\(d\geq1\))-dimensional data space, and \(t_i\) is the time when the data \(p_i\) is generated. If the sampling rates of two time series are the same, the timestamps can be omitted and regarded as a sequence of \(d\)-dimensional data points, and such a sequence is called the original representation of the time series.
[0080] Specifically, at different moments, the business data may be the same, but the status of the business data may be different. Therefore, compared with only obtaining the first business data (excluding time series features), the present disclosure can accurately determine the data status of the first business data by obtaining the first business data corresponding to the time series type.
[0081] Exemplarily, in the scenario where an e-commerce platform compensates users, the first piece of the first business data can be expressed as \((100, 10:15)\), that is, the first piece of the first business data with a compensation amount of 100 yuan is generated at 10:15. The second piece of the first business data can be expressed as \((100, 23:15)\), that is, the second piece of the first business data with a compensation amount of 100 yuan is generated at 23:15.
[0082] And from the second business data and the data status of the second business data at different moments within the second time period, it can be known that within the time period from 10:00 to 11:00, the first business data with a compensation amount of 100 yuan is normal, while within the time period from 23:00 to 24:00, the first business data with a compensation amount of 100 yuan is abnormal. Therefore, by obtaining the first business data corresponding to the time series type, the data status of the first business data can be accurately determined.
[0083] Optionally, the method for an electronic device to obtain the first business data corresponding to the time series type within the first time period specifically includes:
[0084] The electronic device selects the business data whose data generation moment is within the first time period from the original business data corresponding to the time series type and determines it as the first business data.
[0085] With the growth of the business, the server corresponding to the e-commerce platform can generate a large amount of original business data. Among them, the original business data corresponds to the time series type, that is, the original business data includes the data generation moment of each data. In this case, the electronic device can select the business data whose data generation moment is within the first time period from the original business data and determine it as the first business data.
[0086] Exemplarily, as the business grows, the server corresponding to the e-commerce platform generated 1,000 pieces of original business data from January 1, 2021 to February 1, 2021. Among them, each piece of business data includes a corresponding data generation time.
[0087] The preset first time period is from January 10, 2021 to January 11, 2021. In this case, the electronic device can obtain the business data whose data generation time is within the range from January 10, 2021 to January 11, 2021 from the 1,000 pieces of original business data generated by the server corresponding to the e-commerce platform, and determine it as the first business data.
[0088] S102. The electronic device determines the first feature data of the first business data.
[0089] Specifically, after obtaining the first business data within the first time period, in order to facilitate the subsequent determination of the data status of the first business data, the electronic device can determine the first feature data of the first business data.
[0090] Optionally, the method for the electronic device to determine the first feature data of the first business data specifically includes:
[0091] Perform noise reduction processing on the first business data, and perform feature engineering processing on the noise-reduced first business data to obtain the first feature data.
[0092] Optionally, the electronic device can perform noise reduction processing on the first business data based on the second-order wavelet algorithm, or can perform noise reduction processing on the first business data through other algorithms. The present disclosure does not limit this.
[0093] The second-order wavelet algorithm can smooth the tiny burr data in the first business data, and at the same time maximize the retention of the data characteristics of the first business data.
[0094] Optionally, when the electronic device performs feature engineering processing on the noise-reduced first business data, it can transform the noise-reduced first business data into the first feature data for input into the model for prediction. By performing feature engineering processing on the noise-reduced first business data, the electronic device can obtain better data characteristics, enabling the machine learning model to approximate the true value.
[0095] S103. The electronic device determines the data status of the first business data based on the first feature data and the preset model.
[0096] Specifically, after determining the first feature data of the first business data, the electronic device can determine the data status of the first business data based on the first feature data and the preset model.
[0097] Among them, the preset model includes a data prediction model; the data prediction model is obtained by performing feature learning training on the second service data in the second time period until convergence; the second time period is before the first time period.
[0098] Optionally, the preset model may further include a data clustering model; the data clustering model is obtained by clustering the second service data; the data clustering model includes at least one category feature; one category feature is used to represent a clustering result of the second service data.
[0099] The specific method for the electronic device to determine the data state of the first service data based on the first feature data and the preset model includes the following two methods:
[0100] The first method: The electronic device directly inputs the first feature data into the data prediction model to obtain the data state of the first service data.
[0101] The second method: The electronic device first determines the Euclidean distance between the first feature data and each category feature in the multiple category features according to the data clustering model to obtain multiple distance values. When the minimum distance value among the multiple distance values is greater than the preset threshold, the electronic device inputs the first feature data into the data prediction model to obtain the data state of the first service data. When the minimum distance value is less than or equal to the preset threshold, the electronic device directly determines that the data state of the first service data is the normal state.
[0102] In a realizable manner, when the data state of the first service data is an abnormal state, the electronic device may output an alarm message for prompting that the first service data is abnormal.
[0103] Optionally, the alarm message may be a prompt sound, or a prompt message, or other information for prompting that the first service data is abnormal, and the present disclosure does not limit this.
[0104] The technical solutions provided in the above embodiments at least bring the following beneficial effects: As can be seen from S101 - S103, after the electronic device obtains the first service data corresponding to the time series type in the first time period, it can determine the first feature data of the first service data, and based on the first feature data and the preset model, determine the data state of the first service data. First, since the data type of the first service data is the time series type, the first service data can describe the service data through the dimension of time. Second, since the preset model (including the one obtained by performing feature learning training on the second service data in the second time period until convergence) is used to determine the data state of the service data, the present application can quickly and accurately determine the data state of the first service data through the preset model and the first service data with the data type of time series type, and then reasonably perform real-time risk monitoring on the service data through the data state of the first service data.
[0105] In an implementable manner, in combination with Figure 1 , such as Figure 2 shown, the preset model further includes a data clustering model; the data clustering model is obtained by clustering the second service data; the data clustering model includes at least one category feature; one category feature is used to represent a clustering result of the second service data.
[0106] Specifically, the electronic device can obtain the second service data and determine the second feature data. Then, the electronic device can cluster the second feature data based on the hierarchical clustering algorithm to obtain the data clustering model.
[0107] The clustering algorithm can cluster data with the same features and divide data with different features. In this application, the electronic device can use the Agglomerative algorithm to cluster the second feature data to obtain the data clustering model.
[0108] Optionally, the clustering result can be a time period or other results, and the present disclosure does not limit this.
[0109] Exemplarily, in the scenario where an e-commerce platform compensates users, when the clustering result is a time period, the data clustering model includes: within the time period from 1:00 to 8:00, the compensation amount is 100-200 yuan; within the time period from 8:01 to 16:00, the compensation amount is 150-210 yuan; within the time period from 16:01 to 12:59, the compensation amount is 120-280 yuan.
[0110] In the above S103, the method for the electronic device to determine the data status of the first service data based on the first feature data and the preset model specifically includes:
[0111] S201. The electronic device determines the Euclidean distance between the first feature data and each category feature in the data clustering model to obtain a plurality of distance values.
[0112] Specifically, when the preset model includes a data prediction model and a data classification model, the electronic device can determine the Euclidean distance between the first feature data and each category feature among the plurality of category features to obtain a plurality of distance values.
[0113] Among them, one distance value is used to represent the similarity between the first feature data and a clustering result.
[0114] Optionally, the electronic device can also determine the similarity between the first feature data and each clustering result through other similarity algorithms, and the present disclosure does not limit this.
[0115] S202. When the minimum distance value among multiple distance values is greater than a preset threshold, the electronic device inputs the first feature data into the data prediction model to obtain the data status of the first service data.
[0116] Among them, the data status includes a normal status and an abnormal status.
[0117] Specifically, after determining the Euclidean distance between the first feature data and each category feature among multiple category features to obtain multiple distance values, the electronic device can determine the clustering result corresponding to the minimum distance value among the multiple distance values as the clustering result most similar to the first feature data. In this case, in order to ensure that the first feature data belongs to this clustering result, the electronic device can set a preset threshold.
[0118] When the minimum distance value is greater than the preset threshold, it indicates that although the distance between the first feature data and this category feature is the closest, it cannot meet the preset threshold. Therefore, the electronic device needs to input the first feature data into the data prediction model to further determine the data status of the first service data.
[0119] S203. When the minimum distance value is less than or equal to the preset threshold, the electronic device determines that the data status of the first service data is the normal status.
[0120] Specifically, after the electronic device sets the preset threshold, when the minimum distance value is less than or equal to the preset threshold, the electronic device determines that the data status of the first service data is the normal status.
[0121] The technical solutions provided in the above embodiments at least bring the following beneficial effects: As can be seen from S201 - S203, when the preset model includes a data prediction model and a data classification model, the electronic device can determine the Euclidean distance between the first feature data and each category feature among multiple category features to obtain multiple distance values. When the minimum distance value among the multiple distance values is greater than the preset threshold, the electronic device inputs the first feature data into the data prediction model to obtain the data status of the first service data. In this way, the electronic device can more accurately determine the data status of the first service data through the combination of the data prediction model and the data clustering model.
[0122] Correspondingly, when the minimum distance value is less than or equal to the preset threshold, the electronic device can directly determine, through the data clustering model, that the data status of the first service data is the normal status quickly and accurately.
[0123] In an implementable manner, as Figure 3 shown, the generation method of the data prediction model includes:
[0124] S301. The electronic device obtains the second service data corresponding to the time series type and the data status of the second service data.
[0125] Specifically, the electronic device can obtain the second service data corresponding to the time series type within the second time period, as well as the data status of the second service data, so that the electronic device can determine a data prediction model or a data classification model according to the second service data and the data status of the second service data.
[0126] It should be noted that at different times, the service data may be the same, but the status of the service data may be different. Therefore, compared with only obtaining the second service data (excluding time series features), the present disclosure can improve the accuracy of the model output of the data prediction model or the data classification model by obtaining the second service data corresponding to the time series type and the data status of the second service data.
[0127] Optionally, the second service data may be service data of the same data type as the first service data, or may include service data of the same data type as the first service data.
[0128] Exemplarily, if the data type of the preset first service data is type A, then the data type of the second service data may be service data of type A, or may be service data of type A and type B.
[0129] In practical applications, the second service data is usually service data of the same data type as the first service data. In this way, the data prediction model or the data classification model trained through the second service data and the data status of the second service data can accurately determine the data status of the first service data.
[0130] Exemplarily, in the scenario of compensating users on an e-commerce platform, the first service data may be service data of the compensation type generated within January 10, 2021 - January 11, 2021 (i.e., the first time period). The second service data may be service data of the compensation type generated within January 1, 2021 - January 7, 2021 (i.e., the second time period).
[0131] Optionally, the method for the electronic device to obtain the second service data corresponding to the time series type specifically includes:
[0132] Select the service data with the data generation time located within the second time period from the original service data corresponding to the time series type and determine it as the second service data.
[0133] Exemplarily, as the business grows, the server corresponding to the e-commerce platform generates 1000 pieces of original service data from January 1, 2021 to February 1, 2021. Each piece of service data includes a corresponding data generation time.
[0134] The preset second time period is from January 1, 2021 to January 7, 2021. In this case, the electronic device can obtain the service data whose generation time of the selected data is within the period from January 1, 2021 to January 7, 2021 from 1000 pieces of original service data generated by the server corresponding to the e-commerce platform and determine it as the second service data. S302. The electronic device determines the second feature data.
[0135] Specifically, after obtaining the second service data and the data status of the second service data, in order to facilitate subsequent training of the data prediction model, the electronic device can determine the second feature data.
[0136] Among them, the second feature data includes the feature data of the second service data and the feature data of the data status of the second service data.
[0137] Optionally, the method for the electronic device to determine the second feature data specifically includes:
[0138] Perform noise reduction processing on the second service data and the data status of the second service data, and perform feature engineering processing on the second service data and the data status of the second service data after noise reduction to obtain the second feature data.
[0139] Optionally, the electronic device can perform noise reduction processing on the second service data and the data status of the second service data based on the second-order wavelet algorithm, or can perform noise reduction processing on the second service data and the data status of the second service data through other algorithms, and the present disclosure does not limit this.
[0140] Optionally, when the electronic device performs feature engineering processing on the second service data and the data status of the second service data after noise reduction, it can transform the second service data and the data status of the second service data after noise reduction into the second feature data for input into the model for prediction.
[0141] The specific method for the electronic device to determine the second feature data can refer to the specific description of the electronic device to determine the first feature data, and will not be elaborated here.
[0142] S303. The electronic device inputs the second feature data into the support vector machine model and performs feature learning training on the support vector machine model until convergence to obtain the data prediction model.
[0143] Specifically, after determining the second feature data, the electronic device can input the second feature data into the support vector machine (Support Vector Machines, SVM) model and perform feature learning training on the support vector machine model until convergence to obtain the data prediction model.
[0144] The SVM model has a relatively high prediction accuracy and a lower time complexity compared to the neural network algorithm, which is of great benefit to the online prediction scenario. In this disclosure, the SVM model is used to train the second business data, and the model trained by SVM is stored for predictive analysis of the first business data.
[0145] When training the SVM model, it is first necessary to mark the abnormal data in the second business data. Then, the feature data of the second business data is determined, and the second feature data is input into the SVM model, and the SVM model is trained for feature learning until convergence to obtain a data prediction model. Since the data status includes abnormal status and normal status, the output categories of the data prediction model include two types: normal and abnormal.
[0146] When the SVM model performs data prediction, it first calculates the feature data of the data to be predicted. Then, the trained SVM model is used for prediction, and whether to alarm is determined according to the prediction result. When it is found that the prediction result is incorrect after manual intervention in the alarm, the result can be re-marked. The marked data can be added to the training set, and the SVM model can be retrained and updated.
[0147] The technical solutions provided in the above embodiments at least bring the following beneficial effects: As can be seen from S301 - S303, the electronic device can obtain the second business data within the second time period, as well as the data status of the second business data, and determine the second feature data. Subsequently, based on the second feature data and the genetic algorithm, the electronic device performs feature learning training on the support vector machine SVM model until convergence to obtain a data prediction model, and gives a specific implementation manner for training the data prediction model, so as to quickly and accurately determine the data status of the first business data according to the data prediction model subsequently.
[0148] In a feasible manner, in the scenario where an e-commerce platform compensates users, the business data can be compensation data. As Figure 4 shown, the data processing method includes:
[0149] S401. The electronic device obtains the second compensation data corresponding to the time series type and the data status of the second compensation data within the second time period.
[0150] S402. The electronic device performs noise reduction processing on the second compensation data and the data status of the second compensation data, and performs feature engineering processing on the denoised second compensation data and the data status of the second compensation data to obtain second feature data.
[0151] Among them, the second feature data includes the feature data of the second compensation data and the feature data of the data status of the second compensation data.
[0152] S403. The electronic device inputs the second feature data into the support vector machine model and performs feature learning training on the support vector machine model until convergence to obtain a data prediction model.
[0153] S404. The electronic device clusters the second feature data based on the hierarchical clustering algorithm to obtain a data clustering model.
[0154] Among them, the data clustering model includes at least one category feature; one category feature is used to represent a clustering result of the second claim data.
[0155] It should be noted that the electronic device may first execute S403 and then execute S404; or it may first execute S404 and then execute S403; or it may execute S403 and 404 simultaneously. The present disclosure does not make any limitation in this regard.
[0156] S405. The electronic device obtains the first claim data corresponding to the time series type within the first time period.
[0157] Among them, the second time period is before the first time period.
[0158] S406. The electronic device performs noise reduction processing on the first claim data and performs feature engineering processing on the noise-reduced first claim data to obtain the first feature data.
[0159] S407. The electronic device determines the Euclidean distance between the first feature data and each category feature in the data clustering model to obtain a plurality of distance values.
[0160] S408. When the minimum distance value is less than or equal to the preset threshold, the electronic device determines that the data state of the first claim data is the normal state.
[0161] S409. When the minimum distance value among the plurality of distance values is greater than the preset threshold, the electronic device inputs the first feature data into the data prediction model to obtain the data state of the first claim data.
[0162] S410. When the data status of the first compensation data output by the data prediction model is an abnormal status, the electronic device outputs an alarm message for prompting the abnormality of the first compensation data. It can be understood that in actual implementation, the electronic device described in the embodiments of the present disclosure may include one or more hardware structures and / or software modules for implementing the foregoing corresponding service data processing method, and these execution hardware structures and / or software modules may constitute an electronic device. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described function, but such implementation should not be considered to exceed the scope of the present disclosure.
[0163] Based on such an understanding, the embodiments of the present disclosure also correspondingly provide a service data processing device, which can be applied to an electronic device. Figure 5 The structural schematic diagram of the service data processing device provided by the embodiments of the present disclosure is shown. As Figure 5 shown, the service data processing device may include: an acquisition unit 501 and a processing unit 502;
[0164] The acquisition unit 501 is configured to acquire first service data corresponding to a time series type within a first time period;
[0165] The processing unit 502 is configured to determine first feature data of the first service data;
[0166] The processing unit 502 is further configured to determine the data status of the first service data based on the first feature data and a preset model; the preset model includes a data prediction model; the data prediction model is obtained by performing feature learning training on second service data within a second time period until convergence; the second time period is before the first time period.
[0167] Optionally, the preset model further includes a data clustering model; the data clustering model is obtained by clustering the second service data; the data clustering model includes at least one category feature; one category feature is used to represent a clustering result of the second service data;
[0168] The processing unit 502 is specifically configured to:
[0169] Determine the Euclidean distance between the first feature data and each category feature in the data clustering model to obtain a plurality of distance values;
[0170] When the minimum distance value among the plurality of distance values is greater than a preset threshold, input the first feature data into the data prediction model to obtain the data status of the first service data.
[0171] Optionally, the data status includes a normal status and an abnormal status; the processing unit 502 is specifically configured to:
[0172] When the minimum distance value is less than or equal to a preset threshold, determine that the data status of the first service data is the normal status.
[0173] Optionally, the obtaining unit 501 is further configured to obtain second service data corresponding to the time series type and the data status of the second service data;
[0174] The processing unit 502 is further configured to determine second feature data; the second feature data includes the feature data of the second service data and the feature data of the data status of the second service data;
[0175] The processing unit 502 is further configured to input the second feature data into a support vector machine model and perform feature learning training on the support vector machine model until convergence to obtain a data prediction model.
[0176] Optionally, the obtaining unit 501 is specifically configured to:
[0177] Select the service data whose data generation time is within the first time period from the original service data corresponding to the time series type and determine it as the first service data.
[0178] Optionally, the processing unit 502 is specifically configured to:
[0179] Perform noise reduction processing on the first service data and perform feature engineering processing on the noise-reduced first service data to obtain first feature data.
[0180] Optionally, the obtaining unit 501 is specifically configured to:
[0181] Select the service data whose data generation time is within the second time period from the original service data corresponding to the time series type and determine it as the second service data.
[0182] Optionally, the processing unit 502 is specifically configured to:
[0183] Perform noise reduction processing on the second service data and the data status of the second service data, and perform feature engineering processing on the noise-reduced second service data and the data status of the second service data to obtain second feature data.
[0184] As described above, the embodiments of the present disclosure may divide the functional modules of the electronic device according to the above method examples. Among them, the above integrated modules may be implemented in the form of hardware or in the form of software functional modules. In addition, it should be noted that the division of modules in the embodiments of the present disclosure is illustrative, only a logical function division, and there may be other division methods in actual implementation. For example, each functional module may be corresponding to each function, or two or more functions may be integrated in one processing module.
[0185] Regarding the service data processing device in the above embodiments, the specific manners in which each module performs operations and the beneficial effects thereof have been described in detail in the foregoing method embodiments, and will not be elaborated herein.
[0186] The embodiments of the present disclosure further provide a terminal, which may be a user terminal such as a mobile phone or a computer. Figure 6 The structural schematic diagram of the terminal provided by the embodiments of the present disclosure is shown. The terminal may be that the service data processing device may include at least one processor 61, a communication bus 62, a memory 63, and at least one communication interface 64.
[0187] The processor 61 may be a central processing unit (CPU), a microprocessing unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present disclosure solution. As an example, in combination with Figure 5 , the functions implemented by the acquisition unit 501 and the processing unit 502 in the electronic device are the same as the functions implemented by the processor 61 in Figure 6 .
[0188] The communication bus 62 may include a path for transmitting information between the above components.
[0189] The communication interface 64 uses any device such as a transceiver for communicating with other devices or communication networks, such as a server, an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0190] The memory 63 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processing unit through a bus. The memory can also be integrated with the processing unit.
[0191] Among them, the memory 63 is used to store the application program code for executing the solution of the present disclosure and is controlled by the processor 61 for execution. The processor 61 is used to execute the application program code stored in the memory 63, thereby implementing the functions in the method of the present disclosure.
[0192] In a specific implementation, as an embodiment, the processor 61 can include one or more CPUs, such as Figure 6 CPU0 and CPU1 in
[0193] In a specific implementation, as an embodiment, the terminal can include multiple processors, such as Figure 6 the processor 61 and the processor 65 in
[0194] Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). In a specific implementation, as an embodiment, the terminal can further include an input device 66 and an output device 67. The input device 66 and the output device 67 communicate and can receive user input in various ways. For example, the input device 66 can be a mouse, a keyboard, a touch screen device, or a sensing device, etc. The output device 67 communicates with the processor 61 and can display information in various ways. For example, the output device 61 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, etc.
[0195] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.
[0196] The embodiments of the present disclosure also provide a server. Figure 7 The schematic structural diagram of the server provided by the embodiments of the present disclosure is shown. The server may be a business data processing device. The server may vary greatly due to configuration or performance, and may include one or more processors 71 and one or more memories 72. Among them, at least one instruction is stored in the memory 72, and at least one instruction is loaded and executed by the processor 71 to implement the business data processing method provided by each of the above method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0197] The present disclosure also provides a computer-readable storage medium including instructions. When the instructions in the computer-readable storage medium are executed by the processor of the computer device, the computer is enabled to execute the business data processing method provided by the above-mentioned embodiments. For example, the computer-readable storage medium may be a memory 63 including instructions, and the above instructions may be executed by the processor 61 of the terminal to complete the above method. Another example is that the computer-readable storage medium may be a memory 72 including instructions, and the above instructions may be executed by the processor 71 of the server to complete the above method. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.
[0198] The present disclosure also provides a computer program product. The computer program product includes computer instructions. When the computer instructions run on the terminal, the terminal is enabled to execute the above-mentioned Figures 1-4 business data processing method shown in any of the accompanying drawings.
[0199] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0200] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for processing service data, characterized in that, it includes: obtaining first service data corresponding to a time series type within a first time period; the data corresponding to the time series type is a series of tuples of numerical values and time; determining first characteristic data of the first service data; determining the data status of the first service data based on the first characteristic data and a preset model; the preset model includes a data prediction model; the data prediction model is obtained by performing feature learning training on second service data within a second time period until convergence; the second time period is before the first time period; the generation method of the data prediction model includes: obtaining the second service data corresponding to the time series type and the data status of the second service data; determining second characteristic data; the second characteristic data includes the characteristic data of the second service data and the characteristic data of the data status of the second service data; inputting the second characteristic data into a support vector machine model, and performing feature learning training on the support vector machine model until convergence to obtain the data prediction model; the preset model further includes a data clustering model; the data clustering model is obtained by clustering the second service data; the data clustering model includes at least one category feature; one category feature is used to represent one clustering result of the second service data; the determining the data status of the first service data based on the first characteristic data and the preset model includes: determining the Euclidean distance between the first characteristic data and each category feature in the data clustering model to obtain a plurality of distance values; when the minimum distance value among the plurality of distance values is greater than a preset threshold, inputting the first characteristic data into the data prediction model to obtain the data status of the first service data.
2. The method for processing service data according to claim 1, characterized in that, the data status includes a normal status and an abnormal status; the determining the data status of the first service data based on the first characteristic data and the preset model includes: when the minimum distance value is less than or equal to the preset threshold, determining the data status of the first service data as the normal status.
3. The method for processing service data according to claim 1, characterized in that, the obtaining the first service data corresponding to the time series type within the first time period includes: selecting, from the original service data corresponding to the time series type, the service data whose data generation time is within the first time period and determining it as the first service data.
4. The method for processing service data according to claim 1, characterized in that, the determining the first characteristic data of the first service data includes: performing noise reduction processing on the first service data, and performing feature engineering processing on the first service data after noise reduction to obtain the first characteristic data.
5. The method for processing service data according to claim 1, characterized in that, the obtaining the second service data corresponding to the time series type includes: From the original business data corresponding to the time series type, select the business data whose data generation time is within the second time period and determine it as the second business data.
6. The business data processing method according to claim 1, wherein, the determining the second feature data includes: performing noise reduction processing on the second business data and the data status of the second business data, and performing feature engineering processing on the denoised second business data and the data status of the second business data to obtain the second feature data.
7. A business data processing device, wherein, it includes: an acquisition unit and a processing unit; the acquisition unit is used to acquire first business data corresponding to the time series type within the first time period; the data corresponding to the time series type is a series of tuples of numerical values and time; the processing unit is used to determine the first feature data of the first business data; the processing unit is further used to determine the data status of the first business data based on the first feature data and a preset model; the preset model includes a data prediction model; the data prediction model is obtained by performing feature learning training on the second business data within the second time period until convergence; the second time period is before the first time period; the acquisition unit is further used to acquire the second business data corresponding to the time series type and the data status of the second business data; the processing unit is further used to determine second feature data; the second feature data includes the feature data of the second business data and the feature data of the data status of the second business data; the processing unit is further used to input the second feature data into a support vector machine model and perform feature learning training on the support vector machine model until convergence to obtain the data prediction model; the preset model further includes a data clustering model; the data clustering model is obtained by clustering the second business data; the data clustering model includes at least one category feature; one category feature is used to represent a clustering result of the second business data; the processing unit is specifically used for: determining the Euclidean distance between the first feature data and each category feature in the data clustering model to obtain a plurality of distance values; when the minimum distance value among the plurality of distance values is greater than a preset threshold, inputting the first feature data into the data prediction model to obtain the data status of the first business data.
8. The business data processing device according to claim 7, wherein, the data status includes a normal status and an abnormal status; the processing unit is specifically used for: when the minimum distance value is less than or equal to the preset threshold, determining the data status of the first business data as the normal status.
9. The business data processing device according to claim 7, wherein, the acquisition unit is specifically used for: selecting the business data whose data generation time is within the first time period from the original business data corresponding to the time series type and determining it as the first business data.
10. The service data processing device according to claim 7, wherein, the processing unit is specifically configured to: perform noise reduction processing on the first service data, and perform feature engineering processing on the first service data after noise reduction to obtain the first feature data.
11. The service data processing device according to claim 1, wherein, the obtaining unit is specifically configured to: select the service data whose data generation time is within the second time period from the original service data corresponding to the time series type as the second service data.
12. The service data processing device according to claim 1, wherein, the processing unit is specifically configured to: perform noise reduction processing on the second service data and the data status of the second service data, and perform feature engineering processing on the second service data and the data status of the second service data after noise reduction to obtain the second feature data.
13. An electronic device, wherein, the electronic device includes: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the service data processing method according to any one of claims 1-6.
14. A computer-readable storage medium, on which instructions are stored, wherein, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the service data processing method according to any one of claims 1-6.
15. A computer program product, including instructions, wherein, when the instructions run on an electronic device, the electronic device is enabled to execute the service data processing method according to any one of claims 1-6.
Citation Information
Patent Citations
Service monitoring method, device, equipment and storage medium
CN111581508A
Abnormal data determination method and device, storage medium and electronic equipment
CN111831704A