Classification method for user data
By using slope change points to set thresholds in user data classification methods, the problem of low accuracy in the account distinction of shared network platform accounts in the prior art is solved, and higher classification model accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202110211262.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-02-25
AI Technical Summary
When identifying and distinguishing shared network platform accounts (such as family accounts), the prior art relies on experience to set thresholds, and the accuracy is unstable, resulting in low accuracy and poor generalization capabilities of the user classification model.
The threshold is set by extracting slope variable points and quantitatively distinguishing user data based on the set threshold, and two classification models are constructed and trained to improve the accuracy of the distinction.
It improves the accurate distinction stability between "family accounts" and "non-family accounts" and enhances the performance and generalization capabilities of the user classification model.
Smart Images

Figure CN114970650B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the fields of artificial intelligence and big data, and more particularly, to a method and apparatus for classifying user data. Background Art
[0002] In daily life, there is a situation where multiple people share a network platform account. Taking an e-commerce platform account as an example, when many people shop on the e-commerce platform, they will purchase goods for others (for example, other family members). This behavior of helping others purchase goods cannot reflect the personal behavior of the account, and these behavior data will also bring great errors to the classification model.
[0003] For example, when predicting the gender of a user using an e-commerce platform account, data features can be constructed by using user attribute information and the attribute information of the goods browsed, collected, and purchased by the user, and these data features can be used for training to obtain a gender classification model. If the user has the behavior of purchasing goods for others, then the features of the data corresponding to these behaviors may introduce noise to the gender classification model and confuse the training model, resulting in low accuracy and poor generalization ability of the model.
[0004] This behavior of sharing a network platform account is very common in practice. Generally, we refer to this kind of network platform account as a "family account".
[0005] In order to improve the accuracy of the user classification model, it is desirable to identify and extract these "family accounts" before training the classification model. However, the prior art usually sets thresholds in a qualitative manner based on experience. This mainly depends on the experience and judgment ability of developers, and its accuracy is unstable. Summary of the Invention
[0006] A brief overview of the present disclosure is given below in order to provide a basic understanding of some aspects of the present disclosure. However, it should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify the key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is only to present some concepts of the present disclosure in a simplified form as a prelude to the more detailed description given later.
[0007] The present disclosure aims to provide a method for classifying user data. The present disclosure mainly sets thresholds by extracting slope change points, and further quantitatively distinguishes user data (that is, "family accounts" and "non-family accounts") based on the set thresholds. Preferably, the present disclosure can further construct and train two classification models based on the distinguished user data, and the strong features of these two classification models are very different and the generalization ability is strong.
[0008] According to one aspect of the present disclosure, there is provided a method for classifying user data, including: using a first user dataset to construct and train a pre-classification model that meets preset requirements; inputting each sample in the first user dataset into the pre-classification model to obtain the probability that each sample in the first user dataset belongs to each category; obtaining the category discrimination degree of each sample in the first user dataset based on the probability that each sample in the first user dataset belongs to each category; arranging each sample in the first user dataset based on the magnitude of the category discrimination degree of each sample in the first user dataset, and fitting a category discrimination degree curve based on each arranged sample in the first user dataset; determining the slope change point of the category discrimination degree curve, and setting a threshold based on the determined slope change point; and determining the samples in the first user dataset with a category discrimination degree greater than the threshold as a first sample set, and determining the samples in the first user dataset with a category discrimination degree less than or equal to the threshold as a second sample set.
[0009] According to another aspect of the present disclosure, there is provided a classification device for user data, including: a memory on which instructions are stored; and a processor configured to execute the instructions stored on the memory to perform the methods according to the various aspects of the present disclosure.
[0010] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium, including computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the methods according to the various aspects of the present disclosure.
[0011] According to yet another aspect of the present disclosure, there is provided a computer program product, including computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the methods according to the various aspects of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0013] Referring to the drawings, the present disclosure can be more clearly understood from the following detailed description, wherein:
[0014] Figure 1 The flowchart shows an exemplary method for constructing and training a classification model in the prior art;
[0015] Figure 2 The flowchart shows the classification method of user data according to the embodiments of the present disclosure;
[0016] Figure 3Shows an example of a class discrimination curve fitted according to an embodiment of the present disclosure;
[0017] Figure 4a and Figure 4b respectively show the flowcharts of methods for optimizing a classification model according to an embodiment of the present disclosure;
[0018] Figure 5 Shows the flowchart of another method for classifying user data according to a preferred embodiment of the present disclosure; and
[0019] Figure 6 Shows an exemplary configuration of a computing device that can implement an embodiment of the present disclosure. Detailed implementation
[0020] The following detailed description is made with reference to the accompanying drawings, and the following detailed description is provided to help a comprehensive understanding of various exemplary embodiments of the present disclosure. The following description includes various details to help understanding, but these details are only considered as examples and are not intended to limit the present disclosure, which is defined by the appended claims and their equivalents. The words and phrases used in the following description are only for a clear and consistent understanding of the present disclosure. In addition, for clarity and conciseness, the description of well-known structures, functions, and configurations may be omitted. Those of ordinary skill in the art will recognize that various changes and modifications can be made to the examples described herein without departing from the spirit and scope of the present disclosure.
[0021] Figure 1 Shows the flowchart of an exemplary method for constructing and training a classification model in the prior art.
[0022] As Figure 1 shown, in step 102, preprocessing is performed on the training data to improve the efficiency of data mining and optimize the quality of the training data. Generally, the training data includes a large amount of diverse data with relatively reliable and accurate features. For example, a large amount of various user data with obvious features can be obtained from the database of a network platform (such as an e-commerce platform) as training data. By using such training data to construct and train a classification model, the performance of the classification model can be made better.
[0023] Preferably, the training data includes validation data for verifying the attributes of the classification model to determine whether the classification model meets certain preset requirements. The validation data is data similar to the training data, which can be used to assist in the construction of the model and can be reused.
[0024] The preprocessing of the training data can be performed in any manner known in the prior art. For example, operations such as outlier removal, missing value filling, normalization, and standardization can be performed on the training data based on data analysis.
[0025] In step 104, perform feature engineering on the preprocessed sample data to obtain the features of the sample data. Feature engineering can be performed in any manner known in the prior art. According to some embodiments, feature engineering may include encoding conversion, timestamp processing, extending features based on statistical analysis methods, extraction of cross features, and feature selection, etc.
[0026] In step 106, construct and train a model based on the processing results of feature engineering. This step can be completed by any suitable means known in the art.
[0027] In step 108, evaluate the model training results, and in step 110, determine whether the evaluation results meet the preset requirements. According to some embodiments, validation data can be input into the model to determine whether the model meets the preset requirements. Specifically, it can be determined whether one or more attributes of the model meet the preset requirements. The attributes of the model can include, for example, precision, recall rate, accuracy, etc.
[0028] As shown in the figure, if the judgment result is "no", then return to step 106 to continue training the model until the attributes of the model meet the preset requirements; if the judgment result is "yes", then proceed to step 112 and output this model that meets the preset requirements as a classification model.
[0029] Figure 2 A flowchart of a classification method for user data according to an embodiment of the present disclosure is shown.
[0030] As Figure 2 shown, in step 202, use the first user data set to construct and train a pre-classification model that meets the preset requirements. The first user data set can be a set of training data. The specific implementation here can be similar to the steps in Figure 1 and will not be elaborated here.
[0031] Then, in step 204, input each sample in the first user data set into the pre-classification model to obtain the probability that each sample in the first user data set belongs to each category.
[0032] According to some embodiments, each category refers to a group of categories based on the same classification criterion. For example, taking gender as the classification criterion, the output of the pre-classification model can be the probability that a sample belongs to male and female respectively. Taking age group as the classification criterion, the output of the pre-classification model can be the probability that a sample belongs to each age group such as 0 - 20 years old, 21 - 40 years old, 41 - 60 years old, and over 60 years old respectively. The classification criterion is not limited to gender and age group, and can also be various other classification criteria such as occupation, etc.
[0033] However, due to the relatively large amount of noise in the first user dataset used to construct and train the pre-classification model (the "family account" and "non-family account" are not distinguished), the accuracy and confidence of the classification results output by the pre-classification model are relatively low.
[0034] In step 206, the class discrimination degree of each sample in the first user dataset is obtained based on the probability that each sample in the first user dataset belongs to each class.
[0035] When the pre-classification model is configured to determine the probabilities that each sample in the first user dataset belongs to two classes, the absolute value of the difference between the two probabilities that each sample belongs to the two classes is determined as the class discrimination degree of each sample. Taking the classification standard of gender as an example, if the pre-classification model outputs the probability P 1 that a sample belongs to male as 0.7 and the probability P 2 that it belongs to female as 0.3, then the class discrimination degree of this sample is 0.4 (i.e., |P 1 -P 2 | = 0.4).
[0036] When the pre-classification model is configured to determine the probabilities that each sample in the first user dataset belongs to at least three classes, the maximum value among the absolute values of the differences between every two of the at least three probabilities that each sample belongs to the at least three classes is determined as the class discrimination degree of each sample. Taking the classification standard of age group as an example, if the pre-classification model outputs the probability P 1 that a sample belongs to the age group of 0 - 20 years old as 0.2, the probability P 2 that it belongs to the age group of 21 - 40 years old as 0.5, the probability P 3 that it belongs to the age group of 41 - 60 years old as 0.3, and the probability P 4 that it belongs to the age group over 60 years old as 0, then the class discrimination degree of this sample is 0.5. Specifically, the absolute values of the differences between every two probabilities are 0.3 (i.e., |P 1 -P 2 |), 0.1 (i.e., |P 1 -P 3 |), 0.2 (i.e., |P 1 -P 4 |), 0.2 (i.e., |P 2 -P 3 |), 0.5 (i.e., |P 2 -P 4 |), 0.3 (i.e., |P 3 -P 4 |). In this case, the maximum value 0.5 (i.e., |P 2 -P 4 |) of the absolute values of the differences is taken as the class discrimination degree of this sample.
[0037] Then, in step 208, the samples in the first user dataset are arranged based on the magnitude of the class discrimination of each sample in the first user dataset, and a class discrimination curve is fitted based on the arranged samples in the first user dataset.
[0038] The horizontal and vertical coordinates of the class discrimination curve can be set in any suitable manner. Preferably, the horizontal coordinate of the class discrimination curve is set to the equally spaced integer identification numbers assigned to each sample, and the vertical coordinate of the class discrimination curve is set to the class discrimination, so that the position of the slope change point can be determined more conveniently and accurately in subsequent steps. Taking 5 samples as an example, for instance, after arranging the 5 samples in ascending order of their class discrimination, the identification numbers corresponding to these 5 samples can be, for example, 1, 2, 3, 4, 5, and the identification number corresponding to the sample with the smallest class discrimination is 1, while the identification number corresponding to the sample with the largest class discrimination is 5.
[0039] In step 210, the slope change point of the class discrimination curve is determined, and a threshold is set based on the determined slope change point.
[0040] The slope change point of the class discrimination curve can be extracted in any suitable way. An embodiment of extracting the slope change point of the class discrimination curve will be described in detail later in the description of Figure 3 Then, the class discrimination corresponding to the slope change point can be set as the threshold. Still taking the 5 samples arranged in ascending order as an example, if the position of the slope change point is determined to be at sample 3, then the class discrimination of sample 3 can be set as the threshold. In the case of only distinguishing between "family accounts" and "non-family accounts" in the first user dataset, optionally, the integer identification number corresponding to the slope change point can also be set as the threshold, in which case the identification number corresponding to the sample needs to be compared with this threshold in subsequent steps.
[0041] In step 212, the samples in the first user dataset with a class discrimination greater than the threshold are determined as the first sample set, and the samples in the first user dataset with a class discrimination less than or equal to the threshold are determined as the second sample set.
[0042] Still taking the 5 samples arranged in ascending order as an example for the following description.
[0043]
[0044] If it is determined that the position of the slope change point is at sample 3, and the class discrimination degree of sample 3 (assumed to be "0.2") is set as the threshold, then samples with a class discrimination degree greater than "0.2" (i.e., sample 4 and sample 5) can be determined as the first sample set, while samples with a class discrimination degree less than or equal to "0.2" (i.e., sample 1, sample 2, and sample 3) can be determined as the second sample set.
[0045] Optionally, if it is determined that the position of the slope change point is at sample 3, and the identification number "3" of sample 3 is set as the threshold, then samples with an identification number greater than "3" (i.e., sample 4 and sample 5) can be determined as the first sample set, while samples with an identification number less than or equal to "3" (i.e., sample 1, sample 2, and sample 3) can be determined as the second sample set.
[0046] Regarding "family accounts" and "non-family accounts", since "family accounts" are used by multiple people, the class discrimination degree of samples as "family accounts" is usually greater than that of samples as "non-family accounts". That is to say, samples with a class discrimination degree greater than the threshold can be determined as "family accounts", while samples with a class discrimination degree less than or equal to the threshold can be determined as "non-family accounts".
[0047] By using the slope change point instead of the developer's experience to set the threshold, the determination of the threshold can be achieved in a quantitative rather than qualitative manner, thereby distinguishing "family accounts" and "non-family accounts". This greatly improves the stability of accurately distinguishing "family accounts" and "non-family accounts".
[0048] Figure 3 An example of the class discrimination degree curve fitted according to an embodiment of the present disclosure is shown.
[0049] Next, an embodiment of extracting the slope change point will be specifically described by taking Figure 3 as an example.
[0050] As Figure 3 The shown class discrimination degree curve is a class discrimination degree curve fitted based on 20 samples with the class discrimination degrees arranged in ascending order. The ordinate of the class discrimination degree curve is the class discrimination degree, and the abscissa of the class discrimination degree curve is the integer identification number assigned to each sample (i.e., 1, 2,..., 20).
[0051] The average change rate of the class discrimination degree in the interval [1, i] can be expressed by the following expression (1):
[0052]
[0053] The average change rate of the class discrimination degree in the interval [i, n] It can be represented by the slope of the line segment between two points, that is, it can be specifically expressed as the following expression (2):
[0054]
[0055] where {1,..., i,..., n} are equally spaced integer identification numbers assigned to each sample, and {S 1 ,..., S i ,..., S n} are the class discrimination degrees of each sample.
[0056] The slope change point can be defined as the point where the ratio of to
[0057] is the maximum value. Figure 3 Specifically, taking the class discrimination degree curve fitted based on 20 samples as shown in as an example for illustration. The ratio of to at each sample can be calculated. Taking sample 2 as an example, the ratio of to
[0058] is approximately 4.449. By calculation, it can be obtained that the slope change point is at sample 17, and its ratio is approximately 6.667.
[0059] Then the threshold can be set to the class discrimination degree of this slope change point. That is, the threshold can be set to "6.667", and the samples with class discrimination degrees greater than this threshold "6.667" are determined as the first sample set (or "family account"), and the samples with class discrimination degrees less than or equal to this threshold "6.667" are determined as the second sample set (or "non-family account"). Thus, it can be seen that the "family account" includes samples 18 and 19, and the "non-family account" includes samples 1 - 17.
[0060] Figure 4a and Figure 4b show a flowchart of a method for optimizing a classification model according to an embodiment of the present disclosure. After performing all the steps of the method as shown in Figure 2 , the steps as shown in Figure 4a and Figure 4b are further performed.
[0061] In steps 402a and 402b, preprocessing is performed on the first sample set and the second sample set, respectively. In steps 404a and 404b, feature engineering is performed on the preprocessed first sample set and the second sample set, respectively. In steps 406a and 406b, models are constructed and trained based on the processing results of feature engineering, respectively. In steps 408a and 408b, the model training results are evaluated, respectively, and in steps 410a and 410b, it is determined whether the evaluation results meet the preset requirements, respectively.
[0062] As shown in the figure, if the judgment result is "no", return to step 406a or 406b to continue training the model until the attributes of the model meet the preset requirements; if the judgment result is "yes", proceed to step 412a or 412b and output the model that meets the preset requirements as the first classification model or the second classification model.
[0063] Since the specific implementation of steps 402a / 402b to 412a / 412b is substantially the same as steps 102 to 112, they are not described in detail herein.
[0064] By training different classification models for the first sample set (i.e., "family accounts") and the second sample set (i.e., "non-family accounts") respectively, the performance and generalization ability of the user classification model can be greatly improved.
[0065] Therefore, the classification result of the first classification model or the second classification model for a sample has higher accuracy and confidence than the classification result of the pre-classification model for the sample. For example, for a sample, the pre-classification model may determine that the probability of the sample belonging to "male" is 0.6 based on the gender classification standard, and the probability of it belonging to "female" is 0.4. If the sample is input into the first classification model or the second classification model, it may be obtained that the probability of it belonging to "male" is 0.9, and the probability of it belonging to "female" is 0.1.
[0066] Figure 5 FIG. 4 is a flowchart showing a method for classifying user data according to a preferred embodiment of the present disclosure. After executing all steps of the method shown in FIG. 4 , further executing steps as shown in FIG. Figure 5 The specific implementation of the steps similar to the previous steps will not be repeated here.
[0067] like Figure 5 As shown, in step 502, each sample in the second user data set is input into the pre-classification model to obtain a first probability that each sample in the second user data set belongs to each category. According to some embodiments, the second user data set is a set of test data different from the first user data set.
[0068] In step 504, similarly to step 206, the class discrimination degrees of the samples in the second user dataset are obtained based on the first probabilities that the respective samples in the second user dataset belong to the respective classes.
[0069] In step 506, similarly based on the threshold set in step 210, the samples in the second user dataset with class discrimination degrees greater than the threshold are determined as the third sample set, and the samples in the second user dataset with class discrimination degrees less than or equal to the threshold are determined as the fourth sample set.
[0070] Then, in steps 508a and 508b, the respective samples in the third sample set are input into the first classification model to obtain the second probabilities that the respective samples in the third sample set belong to the respective classes, and the respective samples in the fourth sample set are input into the second classification model to obtain the second probabilities that the respective samples in the fourth sample set belong to the respective classes.
[0071] Figure 6 An exemplary configuration of a computing device 600 capable of implementing embodiments according to the present disclosure is shown.
[0072] The computing device 600 is an example of a hardware device capable of applying the above aspects of the present disclosure. The computing device 600 can be any machine configured to perform processing and / or computing. The computing device 600 can be, but is not limited to, a workstation, a server, a desktop computer, a laptop computer, a tablet computer, a personal digital assistant (PDA), a smart phone, an in-vehicle computer, or a combination thereof.
[0073] As Figure 6As shown, the computing device 600 may include one or more components that may be connected or communicate with a bus 602 via one or more interfaces. The bus 602 may include, but is not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, etc. The computing device 600 may include, for example, one or more processors 604, one or more input devices 606, and one or more output devices 608. The one or more processors 604 may be any type of processor and may include, but is not limited to, one or more general-purpose processors or special-purpose processors (such as special processing chips). The processor 604 may be configured, for example, to implement a method for classifying user data according to various embodiments of the present disclosure. The input device 606 may be any type of input device capable of inputting information to the computing device and may include, but is not limited to, a mouse, a keyboard, a touch screen, a microphone, and / or a remote controller. The output device 608 may be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer.
[0074] The computing device 600 may also include or be connected to a non-transitory storage device 614, which may be any non-transitory storage device capable of implementing data storage and may include, but is not limited to, a disk drive, an optical storage device, a solid-state memory, a floppy disk, a flexible disk, a hard disk, a magnetic tape, or any other magnetic medium, a compact disk, or any other optical medium, a cache memory, and / or any other storage chip or module, and / or any other medium from which a computer can read data, instructions, and / or code. The computing device 600 may also include a Random Access Memory (RAM) 610 and a Read-Only Memory (ROM) 612. The ROM 612 may store programs, utilities, or processes to be executed in a non-volatile manner. The RAM 610 may provide volatile data storage and store instructions related to the operation of the computing device 600. The computing device 600 may also include a network / bus interface 616 coupled to a data link 618. The network / bus interface 616 may be any type of device or system capable of enabling communication with external devices and / or networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication device, and / or a chipset (such as Bluetooth TM devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication facilities, etc.).
[0075] The present disclosure also provides a computer-readable storage medium, which includes computer-executable instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to the above aspects of the present disclosure.
[0076] The present disclosure also provides a computer program product, which includes computer-executable instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to the above aspects of the present disclosure.
[0077] The present disclosure can be implemented as any combination of a device, a system, an integrated circuit, and a computer program on a non-transitory computer-readable medium. One or more processors can be implemented as an integrated circuit (IC), an application-specific integrated circuit (ASIC), or a large-scale integrated circuit (LSI), a system LSI, a super LSI, or an ultra LSI component that executes some or all of the functions described in the present disclosure.
[0078] The present disclosure includes the use of software, an application program, a computer program, or an algorithm. The software, the application program, the computer program, or the algorithm can be stored on a non-transitory computer-readable medium to enable a computer such as one or more processors to execute the steps described above and in the drawings. For example, one or more memories store software or an algorithm in executable instructions, and one or more processors can be associated with a set of instructions for executing the software or the algorithm to provide various functions according to the embodiments described in the present disclosure.
[0079] Software and a computer program (which can also be referred to as a program, a software application, an application, a component, or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural language, an object-oriented programming language, a functional programming language, a logic programming language, or an assembly language or a machine language. The term "computer-readable medium" refers to any computer program product, device, or apparatus for providing machine instructions or data to a programmable data processor, such as a magnetic disk, an optical disk, a solid-state storage device, a memory, and a programmable logic device (PLD), including a computer-readable medium that receives the machine instructions as a computer-readable signal.
[0080] The subject matter of the present disclosure is provided as an example of a device, a system, a method, and a program for performing the features described in the present disclosure. However, other features or variations can be expected in addition to the above features. It is expected that the implementation of the components and functions of the present disclosure can be completed with any emerging technology that may replace any of the above implementations.
[0081] In addition, the above description provides examples and does not limit the scope, applicability, or configuration set forth in the claims. Without departing from the spirit and scope of the present disclosure, changes may be made to the functions and arrangements of the elements discussed. Various embodiments may appropriately omit, substitute, or add various processes or components. For example, the features described with respect to certain embodiments may be combined in other embodiments.
[0082] In addition, in the description of the present disclosure, the terms "first", "second", etc. are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.
[0083] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for classifying user data, comprising: using a first user dataset to construct and train a pre-classification model that meets preset requirements; inputting each sample in the first user dataset into the pre-classification model to obtain the probability that each sample in the first user dataset belongs to each category; obtaining the category discrimination degree of each sample in the first user dataset based on the probability that each sample in the first user dataset belongs to each category; arranging each sample in the first user dataset based on the magnitude of the category discrimination degree of each sample in the first user dataset, and fitting a category discrimination degree curve based on each sample in the arranged first user dataset; determining the slope change point of the category discrimination degree curve, and setting a threshold based on the determined slope change point; determining the samples in the first user dataset whose category discrimination degree is greater than the threshold as a first sample set, and determining the samples in the first user dataset whose category discrimination degree is less than or equal to the threshold as a second sample set; and using the first sample set to construct and train a first classification model that meets preset requirements, and using the second sample set to construct and train a second classification model that meets preset requirements; wherein, when the pre-classification model is configured to determine the probability that each sample in the first user dataset belongs to two categories, the absolute value of the difference between the two probabilities that each sample belongs to the two categories is determined as the category discrimination degree of each sample; when the pre-classification model is configured to determine the probability that each sample in the first user dataset belongs to at least three categories, the maximum value among the absolute values of the differences between every two of the at least three probabilities that each sample belongs to the at least three categories is determined as the category discrimination degree of each sample.
2. The method according to claim 1, wherein, by performing preprocessing and feature engineering on the first user dataset to obtain the features of each sample in the first user dataset, and using the features of each sample to construct and train a pre-classification model that meets preset requirements.
3. The method according to claim 2, wherein, the preprocessing includes removing outliers, filling missing values, normalizing, and standardizing.
4. The method according to claim 2, wherein, the feature engineering includes encoding conversion, timestamp processing, extension based on statistical analysis methods, extraction of cross features, and feature selection.
5. The method according to claim 1, wherein, using a validation dataset in the first user dataset to determine whether the pre-classification model meets preset requirements.
6. The method according to claim 1, wherein, the abscissa of the category discrimination degree curve is set to an equally spaced integer identification number assigned to each sample, and the ordinate of the category discrimination degree curve is set to the category discrimination degree.
7. The method according to claim 1, wherein, setting the threshold based on the determined slope change point includes setting the category discrimination degree at the slope change point as the threshold.
8. The method according to claim 1, wherein, using a validation dataset in the first user dataset to determine whether the first classification model and the second classification model meet preset requirements.
9. The method according to claim 1, further Including: Inputting each sample in the second user dataset into a pre-classification model to obtain the first probability that each sample in the second user dataset belongs to each category; Obtaining the category discrimination degree of each sample in the second user dataset based on the first probability that each sample in the second user dataset belongs to each category; Determining the samples in the second user dataset with a category discrimination degree greater than the threshold as the third sample set, and determining the samples in the second user dataset with a category discrimination degree less than or equal to the threshold as the fourth sample set; And Inputting each sample in the third sample set into a first classification model to obtain the second probability that each sample in the third sample set belongs to each category, and inputting each sample in the fourth sample set into a second classification model to obtain the second probability that each sample in the fourth sample set belongs to each category.
10. An apparatus for classifying user data, Including: A memory storing instructions thereon; And A processor configured to execute the instructions stored on the memory to perform the method according to any one of claims 1 to 9.
11. A computer-readable storage medium including computer-executable instructions, which when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 9.
12. A computer program product including computer-executable instructions, which when executed by one or more processors, implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Cross-domain adaptive image classification method for improving local category discrimination
CN110020674A
Garbage classification motivation method and device based on block chains
CN110844394A