Data Processing Method, Apparatus and Electronic Device

By generating feature data sets and sample labels from different periods to calculate the performance indicators of the model, the problem of difficulty in obtaining feedback data in a timely manner after model training and launch is solved, effective evaluation and timely repair of model performance is achieved, and the stability and accuracy of the model in financial risk control and marketing scenarios are improved.

CN115034322BActive Publication Date: 2025-07-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210698442.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-07-18
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

In the case where it is difficult to obtain feedback data in a timely manner during the model training stage and after it is launched, how to effectively evaluate the performance of the model to ensure that it meets the caller's requirements.

Method used

By acquiring the first data set corresponding to the pending model, using the feature data of the sample at different periods to generate multiple second data sets, and combining the sample labels to calculate the performance indicators of the model, such as AUC, KS and PSI, to generate data processing results to evaluate the performance of the model.

Benefits of technology

In the absence of real-time feedback data, the stability and classification effect of the model can be evaluated in a timely manner, potential problems can be discovered and repaired or updated, avoid model performance attenuation, and improve the model's performance in financial risk control and marketing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115034322B_ABST
    Figure CN115034322B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method, apparatus, electronic device, and storage medium, which relate to the field of deep learning technology in the field of artificial intelligence technology and can be used in scenarios such as financial risk control and marketing. The method is as follows: obtaining a first data set corresponding to a model to be processed, where the first data set includes samples and sample labels; obtaining feature data of features in different periods according to the features of the samples to generate multiple second data sets; obtaining the numerical value of the metric of the model according to the multiple second data sets and the sample labels; and generating a data processing result of the model to be processed according to the numerical value of the metric. The present disclosure obtains the feature data of the features in the known first data set in different periods, calculates the metrics related to the performance of the model to be processed according to the feature data in different periods and the known sample labels, and completes the data processing process of the model to be processed, and completes the inspection of the model performance in the case where it is difficult to obtain the feedback data of the model caller in a timely manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning in the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, and electronic device. Background Art

[0002] Currently, in order to ensure that the performance of the model meets the requirements of the caller, data processing processes such as scoring the model are required during the model training stage, before the model goes online, and after the model goes online. Usually, the model is processed based on an existing labeled dataset. For example, after the model goes online, the above-mentioned labeled dataset can be constructed according to the data feedback by the caller to complete the above data processing process. However, how to complete the inspection of the model performance by performing relevant data processing on the model in the case where it is difficult to obtain feedback data in a timely manner has become an urgent problem to be solved. Summary of the Invention

[0003] A data processing method, apparatus, and electronic device are provided.

[0004] According to a first aspect, a data processing method is provided, including: obtaining a first dataset corresponding to a model to be processed, where the first dataset includes samples and sample labels; obtaining feature data of the features at different times according to the features of the samples to generate a plurality of second datasets; obtaining a value of an index of the model to be processed according to the plurality of second datasets and the sample labels, where the index is used to characterize the performance of the model to be processed; and generating a data processing result of the model to be processed according to the value of the index.

[0005] According to a second aspect, a data processing apparatus is provided, including: a first obtaining module, configured to obtain a first dataset corresponding to a model to be processed, where the first dataset includes samples and sample labels; a second obtaining module, configured to obtain feature data of the features at different times according to the features of the samples to generate a plurality of second datasets; a third obtaining module, configured to obtain a value of an index of the model to be processed according to the plurality of second datasets and the sample labels, where the index is used to characterize the performance of the model to be processed; and a generating module, configured to generate a data processing result of the model to be processed according to the value of the index.

[0006] According to a third aspect, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data processing method according to the first aspect of the present disclosure.

[0007] According to a fourth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the data processing method according to the first aspect of the present disclosure.

[0008] According to a fifth aspect, there is provided a computer program product including a computer program which, when executed by a processor, implements the steps of the data processing method according to the first aspect of the present disclosure.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic flowchart of the data processing method according to the first embodiment of the present disclosure;

[0012] Figure 2 is a schematic flowchart of the data processing method according to the second embodiment of the present disclosure;

[0013] Figure 3 is a schematic flowchart of the data processing method according to the third embodiment of the present disclosure;

[0014] Figure 4 is a schematic flowchart of the data processing method according to the fourth embodiment of the present disclosure;

[0015] Figure 5 is a schematic block diagram of data processing on a model to be processed at different stages;

[0016] Figure 6 is a schematic diagram of data processing on a model to be processed at different stages according to the embodiments of the present disclosure;

[0017] Figure 7 is a block diagram of the data processing apparatus according to the first embodiment of the present disclosure;

[0018] Figure 8 is a block diagram of the data processing apparatus according to the second embodiment of the present disclosure;

[0019] Figure 9 is a block diagram of an electronic device for implementing the method according to the embodiments of the present disclosure. Detailed Embodiments

[0020] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0021] Artificial Intelligence (AI) is a technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Currently, AI technology has the advantages of high automation, high precision, and low cost, and has been widely applied.

[0022] Deep Learning (DL) is a new research direction in the field of Machine Learning (ML). It studies the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds. In terms of specific research content, it mainly includes neural network systems based on convolutional operations, namely convolutional neural networks; autoencoder neural networks based on multiple layers of neurons; and deep belief networks that are pre-trained in the form of multi-layer autoencoder neural networks and then further optimize the neural network weights by combining discriminative information. Deep learning has achieved many results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technology, and other related fields. Deep learning enables machines to imitate human activities such as audiovisual and thinking, solves many complex pattern recognition problems, and has made great progress in artificial intelligence-related technologies.

[0023] The data processing method, device, and electronic device of the embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0024] Figure 1 It is a schematic flowchart of the data processing method according to the first embodiment of the present disclosure.

[0025] As Figure 1 shown, the data processing method of the embodiments of the present disclosure may specifically include the following steps:

[0026] S101, obtain a first data set corresponding to the model to be processed, where the first data set includes samples and sample labels.

[0027] Specifically, the execution subject of the data processing method in the embodiments of the present disclosure may be the data processing device provided in the embodiments of the present disclosure. The data processing device may be a hardware device with data information processing capabilities and / or the necessary software for driving the hardware device to work. Optionally, the execution subject may include workstations, servers, computers, user terminals, and other devices. Among them, user terminals include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc.

[0028] In the embodiments of the present disclosure, for a model to be processed, a first data set corresponding to the model is obtained. The first data set includes samples and sample labels, and may also include feature data corresponding to the samples. Among them, the model to be processed may be a model that needs to be regularly inspected for performance after being put on the line. In practice, a labeled data set in the training stage may be used as a known first data set when performing data processing on the model to be processed. For example, in the financial risk control scenario, each sample in the first data set corresponds to a user, and the sample label may be whether the user defaults.

[0029] S102. According to the features of the samples, obtain the feature data of the features in different periods to generate multiple second data sets.

[0030] In the embodiments of the present disclosure, the feature data of each feature corresponding to the sample in different periods is obtained. The period described here is different from the period corresponding to the feature data in the first data set. Among them, the features of the samples can also be understood as the features of the first data set. For example, the features of a sample in a first data set corresponding to a risk control model may be the age of the user and relevant business indicators of the financial services handled by the user, etc.

[0031] Thus, the feature data corresponding to each feature of each sample in the first data set in different periods can be obtained, and the feature data in different periods can be put into different data sets to obtain multiple new second data sets. Among them, the above "period" may be the time length corresponding to a preset time window, such as a week or a month, etc.

[0032] S103. According to the multiple second data sets and the sample labels, obtain the numerical value of the index of the model to be processed, and the index is used to characterize the performance of the model to be processed.

[0033] In practice, if it is difficult to obtain the feedback data of the model invoker in a timely manner, it is difficult for us to determine the current corresponding label of the user, difficult to perform data processing processes such as scoring the model based on the current corresponding feature data of the user, and difficult to complete the inspection of the model performance. For example, in an e-commerce scenario, whether the user clicks on the products recommended by the platform can be used as feedback data and sent back to the server. This data can be used as the label corresponding to the current feature data of the user. For example, if the current feature data is the parameters corresponding to the user's current search query and other behaviors, the label is whether the user clicks on the products recommended by the platform. This label can be used to perform the above data processing processes on the recommendation model.

[0034] When it is difficult to obtain the feedback data in a timely manner, that is, the label corresponding to the current user behavior, the present disclosure uses the known label (i.e., the sample label in the first dataset) corresponding to the same sample (i.e., the same user) as the label corresponding to the feature data of the sample in the second dataset. In this way, a labeled dataset including the updated feature data and the known sample labels is formed.

[0035] In the embodiments of the present disclosure, according to the feature data and sample labels included in a certain period in the second dataset, the value of the index of the model to be processed is calculated. The index of the model to be processed can be used to characterize the performance of the model to be processed, and at least includes, but is not limited to, the area under the curve (AUC) corresponding to the Receiver Operating Characteristic Curve (ROC curve), the model discrimination index (Kolmogorov-Smirnov, KS for short), and the model stability index, that is, the Population Stability Index (PSI for short).

[0036] S104, generate the data processing result of the model to be processed according to the value of the index.

[0037] In the embodiments of the present disclosure, according to the obtained value of the index of the model to be processed, the stability, classification effect, positive and negative sample distribution of the model to be processed are judged, so as to obtain the data processing result of the model to be processed and complete the inspection of the model performance.

[0038] In summary, the data processing method according to the embodiments of the present disclosure obtains a first data set corresponding to a model to be processed, where the first data set includes samples and sample labels; obtains feature data of the features at different times according to the features of the samples to generate a plurality of second data sets; obtains the numerical values of the metrics of the model to be processed according to the plurality of second data sets and the sample labels, where the metrics can be used to characterize the performance of the model to be processed; and generates a data processing result of the model to be processed according to the numerical values of the metrics. The data processing method provided by the present disclosure obtains the feature data of the features in the first data set at different times according to the known first data set, calculates the metrics related to the performance of the model to be processed according to the feature data at different times and the known sample labels, and completes the data processing process of the model to be processed, and completes the inspection of the model performance in the case where it is difficult to obtain the feedback data of the model call party in a timely manner.

[0039] Figure 2 It is a schematic flowchart of the data processing method according to the second embodiment of the present disclosure.

[0040] As Figure 2 shown, on the basis of the embodiment shown in Figure 1 the data processing method according to the embodiments of the present disclosure may specifically include the following steps:

[0041] S201, obtain a first data set corresponding to the model to be processed, where the first data set includes samples and sample labels.

[0042] S202, obtain the feature data of the features at different times according to the features of the samples to generate a plurality of second data sets.

[0043] S203, determine two target data sets from the plurality of second data sets.

[0044] In the embodiments of the present disclosure, select the second data set corresponding to the preset time period from the plurality of second data sets as the target data set to obtain the feature data of different preset time periods.

[0045] S204, obtain two numerical values corresponding to the model discrimination index, two numerical values corresponding to the area under the curve, and the numerical value corresponding to the model stability index according to the two target data sets and the sample labels.

[0046] In the embodiments of the present disclosure, each target data set and the sample label are used as a labeled data set, and the AUC value and the KS value of the model are calculated based on the labeled data set, so as to obtain the scoring situation of the model to be processed based on the feature data at different times. In addition, the value of the model stability index PSI can also be calculated based on the above two target data sets.

[0047] In some embodiments, the numerical values of the metrics of the model to be processed can be obtained as needed, and it is not limited to obtaining the numerical values of the three metrics of AUC, KS, and PSI simultaneously.

[0048] S205. Generate a data processing result of the model to be processed according to the numerical values of the metrics.

[0049] Specifically, steps S201 - S202 are the same as the above steps S101 - S102, and step S205 is the same as the above step S104, which will not be elaborated here.

[0050] In some embodiments, the feature data of a certain period after the feature data is updated can also be used as a second data set, and a known labeled data set before the feature data is updated can be used as a first data set, so as to calculate the numerical values of the metrics of the model to be processed.

[0051] Based on the above embodiments, as Figure 3 shown, "Generate a data processing result of the model to be processed according to the numerical values of the metrics" in the above step S205 may include the following steps:

[0052] S301. Calculate a first difference between two numerical values corresponding to the model discrimination metric and a second difference between two numerical values corresponding to the area under the curve.

[0053] In the embodiments of the present disclosure, if the data processing process of the model to be processed is completed based on the calculation of the numerical values of the three metrics of AUC, KS, and PSI, it is necessary to calculate the difference between two KS values corresponding to the model to be processed, and use this difference as the first difference, calculate the difference between two AUC values corresponding to the model to be processed, and use this difference as the second difference, so as to determine whether the numerical values of the metrics of the model to be processed meet the corresponding conditions.

[0054] S302. In response to the numerical values of the metrics of the model to be processed meeting any of the following conditions: the first difference is greater than the first threshold, the second difference is greater than the second threshold, and the numerical value corresponding to the model stability metric is greater than the third threshold, determine that the data processing result of the model to be processed is model anomaly.

[0055] In the embodiments of the present disclosure, if the numerical values of the metrics of the model to be processed meet any of the following conditions, the data processing result of the model to be processed is determined to be model anomaly. The conditions corresponding to the numerical values of the metrics of the model to be processed are:

[0056] Condition 1: The difference between two KS values (i.e., the first difference) is greater than the first threshold;

[0057] Condition 2: The difference between two AUCs (i.e., the second difference) is greater than the second threshold;

[0058] Condition 3: The value corresponding to the model stability index PSI is greater than the third threshold.

[0059] S303. In response to the first difference being less than or equal to the first threshold, the second difference being less than or equal to the second threshold, and the value corresponding to the stability index being less than or equal to the third threshold, determine that the data processing result of the model to be processed is that the model is normal.

[0060] In the embodiments of the present disclosure, if the difference between two KS values (i.e., the first difference) is less than or equal to the first threshold, the difference between two AUC values (i.e., the second difference) is less than or equal to the second threshold, and the value corresponding to the stability index is less than or equal to the third threshold, it can be considered that the data processing result of the model to be processed is that the model is normal.

[0061] Among them, the first threshold, the second threshold, and the third threshold can be set in advance as needed. For example, the first threshold and the second threshold can be set to 0.03, and the third threshold can be set to 0.1. That is, when the AUC change corresponding to the model to be processed under the feature data in different periods does not exceed 0.03, the corresponding KS change does not exceed 0.03, and the PSI of the model to be processed does not exceed 0.1, it can be considered that the data processing result of the model to be processed is that the model is normal.

[0062] Based on the above embodiments, as Figure 4 shown, the data processing method of the embodiments of the present disclosure may further include a process of detecting or analyzing features, which may specifically include the following steps:

[0063] S401. In response to the data processing result of the model to be processed being that the model is abnormal, detect the distribution of features according to multiple second data sets.

[0064] In the embodiments of the present disclosure, if the data processing result of the model to be processed is that the model is abnormal, it is possible to check whether the model abnormality is caused by feature reasons by analyzing the distribution of features of the samples.

[0065] First, according to the updated feature data in multiple second data sets, the distribution of features can be viewed. For example, calculate the positive sample rate, coverage rate of features, and calculate the PSI value of features after binning the same features to judge the feature stability, so as to obtain the detection result of the distribution of features, such as low feature stability or large distribution of features.

[0066] S402. Analyze the reason why the data processing result of the model to be processed is that the model is abnormal according to the detection result of the distribution of features.

[0067] In the embodiments of the present disclosure, according to the detection result of the distribution of features, it is determined whether there is a feature problem. For example, if the distribution of features is large, it is manually checked whether it is caused by a feature value (i.e., feature data) problem. The feature value problem is generally caused by an error in the extraction program or a change in the underlying data without notifying the using end, and the feature value can be repaired. For the feature values that can be repaired, after the repair is completed, the above data processing is performed on the model to be processed again; for the feature values that cannot be repaired, the model can be re-iterated to update the model.

[0068] Thus, based on a known labeled first data set and updated feature data at different times, in the case of missing real-time feedback data, a data processing process related to the performance of the model to be processed can be performed to determine whether further inspection and update of the model to be processed are required, and problems such as attenuation of model features or the effect of the model itself can be detected early in financial risk control and marketing scenarios, avoiding losses to customers.

[0069] To illustrate the data processing method of the embodiments of the present disclosure in detail, now in combination with Figure 5 - Figure 6 it is described in detail, Figure 5 is a schematic block diagram of data processing on the model to be processed at different stages. As Figure 5 shown, during the entire life cycle of the model (before model training, before model going live, and after model going live), the above data processing needs to be performed on the model. The embodiments of the present disclosure can be applied to perform retrospective inspection and feature selection on the model through the above data processing process during model training; it can also be applied to verify the model scores and input features through the above data processing process before the model goes live; it can also be applied to regularly detect the model scores and input features through the above data processing process after the model goes live. For example, as Figure 6 shown, before the model goes live, the value of the index of the model to be processed is calculated according to the feature data in the latest time period, and data processing such as model scoring is performed on the model to be processed by analyzing whether the value of the index of the model to be processed exceeds the corresponding threshold. If the data processing result is that the model is abnormal, the features are checked, and the above data processing is performed on the model again. If the data processing result is that the model is normal, the model is put into production, and the data processing is performed on the model to be processed regularly according to the preset time to detect the changes of each index of the model to be processed. Whether the model is normal is checked by analyzing whether the value of the index of the model to be processed exceeds the corresponding threshold. If it is normal, the model is retained; if it is abnormal, the reason for the model abnormality is analyzed.

[0070] Figure 7 is a block diagram of the data processing device according to the first embodiment of the present disclosure.

[0071] As Figure 7As shown in the figure, the data processing device 700 according to an embodiment of the present disclosure includes: a first acquisition module 701, a second acquisition module 702, a third acquisition module 703, and a generation module 704.

[0072] The first acquisition module 701 is configured to acquire a first data set corresponding to a model to be processed, where the first data set includes samples and sample labels.

[0073] The second acquisition module 702 is configured to acquire feature data of a feature at different times according to the features of the samples, so as to generate a plurality of second data sets.

[0074] The third acquisition module 703 is configured to acquire the numerical value of an index of the model to be processed according to the plurality of second data sets and the sample labels, where the index is used to characterize the performance of the model to be processed.

[0075] The generation module 704 is configured to generate a data processing result of the model to be processed according to the numerical value of the index.

[0076] It should be noted that the above explanation of the data processing method embodiment is also applicable to the data processing device according to the embodiment of the present disclosure, and the specific process will not be elaborated here.

[0077] In summary, the data processing device according to the embodiment of the present disclosure acquires a first data set corresponding to a model to be processed, where the first data set includes samples and sample labels; acquires feature data of the features at different times according to the features of the samples, so as to generate a plurality of second data sets; acquires the numerical value of an index of the model to be processed according to the plurality of second data sets and the sample labels, where the index can be used to characterize the performance of the model to be processed; and generates a data processing result of the model to be processed according to the numerical value of the index. The data processing method provided by the present disclosure acquires the feature data of the features in the first data set at different times according to the known first data set, calculates the indexes related to the performance of the model to be processed according to the feature data at different times and the known sample labels, and completes the data processing process of the model to be processed, and completes the inspection of the model performance in the case where it is difficult to obtain feedback data in a timely manner.

[0078] Figure 8 It is a block diagram of a data processing device according to a second embodiment of the present disclosure.

[0079] As Figure 8 shown in the figure, the data processing device 800 according to an embodiment of the present disclosure includes: a first acquisition module 801, a second acquisition module 802, a third acquisition module 803, and a generation module 804.

[0080] Among them, the first acquisition module 801 has the same structure and function as the first acquisition module 701 in the previous embodiment, the second acquisition module 802 has the same structure and function as the second acquisition module 702 in the previous embodiment, the third acquisition module 803 has the same structure and function as the third acquisition module 703 in the previous embodiment, and the generation module 804 has the same structure and function as the generation module 704 in the previous embodiment.

[0081] Furthermore, the metrics of the model to be processed include at least one of the following: the area under the curve corresponding to the receiver operating characteristic curve, the model discrimination metric, and the model stability metric.

[0082] Furthermore, the third acquisition module 803 includes: a determination unit 8031 for determining two target data sets from multiple second data sets; and an acquisition unit 8032 for obtaining two values corresponding to the model discrimination metric, two values corresponding to the area under the curve, and the value corresponding to the model stability metric according to the two target data sets and the sample labels.

[0083] Furthermore, the generation module 804 includes: a calculation unit for calculating a first difference between two values corresponding to the model discrimination metric and a second difference between two values corresponding to the area under the curve; a first determination unit for determining that the data processing result of the model to be processed is model abnormal in response to the value of the metric of the model to be processed satisfying any one of the following conditions: the first difference is greater than a first threshold, the second difference is greater than a second threshold, and the value corresponding to the model stability metric is greater than a third threshold; and a second determination unit for determining that the data processing result of the model to be processed is model normal in response to the first difference being less than or equal to the first threshold, the second difference being less than or equal to the second threshold, and the value corresponding to the stability metric being less than or equal to the third threshold.

[0084] Furthermore, the data processing device 800 may further include: a detection module for detecting the distribution of features according to multiple second data sets in response to the data processing result of the model to be processed being model abnormal; and an analysis module for analyzing the reason why the data processing result of the model to be processed is model abnormal according to the detection result of the distribution of features.

[0085] In summary, the data processing device according to the embodiments of the present disclosure obtains a first data set corresponding to a model to be processed, where the first data set includes samples and sample labels; obtains feature data of the features at different times according to the features of the samples to generate a plurality of second data sets; obtains the numerical value of an index of the model to be processed according to the plurality of second data sets and the sample labels, and the index can be used to characterize the performance of the model to be processed; and generates a data processing result of the model to be processed according to the numerical value of the index. The data processing method provided by the present disclosure obtains the feature data of the features in the first data set at different times according to the known first data set, calculates the performance-related index of the model to be processed according to the feature data at different times and the known sample labels, and completes the data processing process of the model to be processed, so as to complete the inspection of the model performance in the case where it is difficult to obtain feedback data in a timely manner.

[0086] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0087] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0088] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0089] As Figure 9 shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 902 or the computer program loaded from the storage unit 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0090] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0091] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as Figure 1 to Figure 6 the data processing method shown. For example, in some embodiments, the data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the semantic parsing method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the data processing method by any other suitable means (e.g., by means of firmware).

[0092] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0093] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or a controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program codes may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0094] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0095] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0096] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.

[0097] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in a cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server can also be a server of a distributed system, or a server combined with blockchain.

[0098] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the steps of the data processing method shown in the above embodiments of the present disclosure.

[0099] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is imposed herein.

[0100] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A data processing method, comprising: Obtaining a first data set corresponding to a recommendation model, where the first data set includes samples and sample labels; Each sample corresponds to a user, and the sample label is whether to click on the product recommended by the platform; According to the features of the samples, obtaining the feature data of the features in different periods, and putting the feature data of different periods into different data sets to generate multiple second data sets; the features of the samples include parameters corresponding to user behaviors, and the behaviors include search query behaviors, and the periods corresponding to different periods are different from the periods corresponding to the feature data in the first data set; Determining two target data sets from the multiple second data sets; And According to the two target data sets and the sample labels, obtaining two values corresponding to the model discrimination index, two values corresponding to the area under the curve, and a value corresponding to the model stability index; and Generating a data processing result of the recommendation model according to the values of the indexes.

2. The method according to claim 1, wherein, The generating a data processing result of the recommendation model according to the values of the indexes includes: Calculating a first difference between the two values corresponding to the model discrimination index and a second difference between the two values corresponding to the area under the curve; In response to any one of the following conditions being satisfied by the values of the indexes of the recommendation model: the first difference is greater than a first threshold, the second difference is greater than a second threshold, and the value corresponding to the model stability index is greater than a third threshold, determining that the data processing result of the recommendation model is model abnormal; and In response to the first difference being less than or equal to the first threshold, the second difference being less than or equal to the second threshold, and the value corresponding to the stability index being less than or equal to the third threshold, determining that the data processing result of the recommendation model is model normal.

3. The method according to claim 2, further comprising: In response to the data processing result of the recommendation model being model abnormal, detecting the distribution of the features according to the multiple second data sets; And Analyzing the reason why the data processing result of the recommendation model is model abnormal according to the detection result of the distribution of the features.

4. A data processing apparatus, comprising: A first obtaining module, configured to obtain a first data set corresponding to a recommendation model, where the first data set includes samples and sample labels; Each sample corresponds to a user, and the sample label is whether to click on the product recommended by the platform; A second obtaining module, configured to obtain the feature data of the features in different periods according to the features of the samples, and put the feature data of different periods into different data sets to generate multiple second data sets; the features of the samples include parameters corresponding to user behaviors, and the behaviors include search query behaviors, and the periods corresponding to different periods are different from the periods corresponding to the feature data in the first data set; A third obtaining module, configured to determine two target data sets from the multiple second data sets; and obtaining two values corresponding to the model discrimination index, two values corresponding to the area under the curve, and a value corresponding to the model stability index according to the two target data sets and the sample labels; and a generation module, configured to generate a data processing result of the recommendation model according to the values of the indicators.

5. The device according to claim 4, wherein The generation module includes: a calculation unit, configured to calculate a first difference between the two values corresponding to the model discrimination index and a second difference between the two values corresponding to the area under the curve; a first determination unit, configured to determine that the data processing result of the recommendation model is model abnormal in response to any one of the following conditions being satisfied for the values of the indicators of the recommendation model: the first difference is greater than a first threshold, the second difference is greater than a second threshold, and the value corresponding to the model stability index is greater than a third threshold; and a second determination unit, configured to determine that the data processing result of the recommendation model is model normal in response to the first difference being less than or equal to the first threshold, the second difference being less than or equal to the second threshold, and the value corresponding to the stability index being less than or equal to the third threshold.

6. The apparatus according to claim 5, further comprising: a detection module, configured to detect the distribution of the features according to the multiple second data sets in response to the data processing result of the recommendation model being model abnormal; and an analysis module, configured to analyze the reason why the data processing result of the recommendation model is model abnormal according to the detection result of the distribution of the features.

7. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-3.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-3.

9. A computer program product, comprising a computer program, where the computer program implements the steps of the method according to any one of claims 1-3 when executed by a processor.

Citation Information

Patent Citations

  • Product recommendation method and device, medium and electronic equipment

    CN113837843A