A feature encoding method, device, equipment and medium

By calculating the weight of objects in different categories and binning, the feature distinction and contribution of sparse samples are improved, and the problem of low feature distinction and contribution of sparse samples in the existing technology is solved, and the recognition ability and training effect of the model are improved.

CN114881163BActive Publication Date: 2025-06-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210564917.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-06-06
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

In the prior art, the characteristic distinction and contribution of sparse samples are low, resulting in the model being unable to accurately identify objects belonging to this category, especially in scenarios such as credit risk control.

Method used

By calculating the first weight of the object in different categories, after binning, the second weight of the binning in different categories is calculated, and it is used as feature values ​​to train the model.

Benefits of technology

The coverage and feature distinction of sparse samples are improved, the model's ability to identify the categories of sparse samples belonging to, and the model's training effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114881163B_ABST
    Figure CN114881163B_ABST
Patent Text Reader

Abstract

The present disclosure provides a feature encoding method, device, equipment, medium and program product, which relates to the field of machine learning technology, especially to smart finance, artificial intelligence and deep learning technology. The specific implementation scheme is: according to the number of samples of multiple objects and the number of samples of multiple objects under at least two categories, the first weights of multiple objects in at least two categories are calculated, wherein the goal of the model training is to enable the model to classify the input objects in at least two categories; bin multiple objects according to the first weight to obtain multiple object bins; according to the number of samples of multiple object bins and the number of samples of multiple object bins under at least two categories, the second weights of multiple object bins in at least two categories are calculated, and the second weights of multiple object bins are used as the feature values ​​of multiple object bins. The present disclosure can improve the coverage, monotonicity and discrimination of sparse features, thereby enhancing the model training effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of machine learning technology, in particular to smart finance, artificial intelligence and deep learning technology, and specifically to a feature encoding method, device, equipment, medium and program product. Background Art

[0002] In the field of machine learning, it is necessary to perform feature encoding on sample data before training the model. The quality of feature encoding directly affects the training effect of the model.

[0003] For the training of linear models for classification tasks, existing feature encoding methods usually have high feature dimensions. Moreover, for categories with low sample coverage, when applied to modeling, the feature discrimination and contribution of these sparse samples are extremely low, making it impossible for the model to accurately identify the category. Summary of the invention

[0004] The present disclosure provides a feature encoding method, apparatus, device, medium and program product.

[0005] According to one aspect of the present disclosure, there is provided a feature encoding method, comprising:

[0006] Calculating first weights of the multiple objects in the at least two categories according to the sample numbers of the multiple objects and the sample numbers of the multiple objects in the at least two categories, wherein the goal of the model training is to enable the model to classify the input objects in the at least two categories;

[0007] Binning the multiple objects according to the first weight to obtain multiple object bins;

[0008] According to the number of samples in the multiple object bins and the number of samples in the multiple object bins under the at least two categories, second weights of the multiple object bins in the at least two categories are calculated, and the second weights of the multiple object bins are used as feature values ​​of the multiple object bins.

[0009] According to another aspect of the present disclosure, there is provided a feature encoding device, comprising:

[0010] A first weight calculation module, configured to calculate first weights of the multiple objects in at least two categories according to the number of samples of the multiple objects and the number of samples of the multiple objects in at least two categories, wherein the goal of the model training is to enable the model to classify the input objects in the at least two categories;

[0011] A binning module, configured to bin the plurality of objects according to the first weight to obtain a plurality of object bins;

[0012] A second weight calculation module is used to calculate second weights of the multiple object bins in the at least two categories based on the number of samples in the multiple object bins and the number of samples in the multiple object bins under the at least two categories, and use the second weights of the multiple object bins as feature values ​​of the multiple object bins.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the feature encoding method described in any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the feature encoding method described in any embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the feature encoding method described in any embodiment of the present disclosure is implemented.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0021] Figure 1 is a schematic diagram of a feature encoding method according to an embodiment of the present disclosure;

[0022] Figure 2 is a schematic diagram of a feature encoding method according to an embodiment of the present disclosure;

[0023] Figure 3 is a structural schematic diagram of a feature encoding device according to an embodiment of the present disclosure;

[0024] Figure 4 It is a block diagram of an electronic device used to implement the feature encoding method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] Figure 1 This is a flow chart of a feature encoding method according to an embodiment of the present disclosure. This embodiment can be applied to feature encoding of training samples to train a model using the sample features obtained after encoding, especially for feature encoding of sparse samples, and relates to the field of machine learning technology, especially to smart finance, artificial intelligence and deep learning technology. The method can be executed by a feature encoding device, which is implemented in software and / or hardware, and is preferably configured in an electronic device, such as a computer device or a server. Figure 1 As shown, the method specifically includes the following:

[0027] S101. Calculate first weights of the multiple objects in at least two categories based on the number of samples of the multiple objects and the number of samples of the multiple objects in at least two categories, wherein a goal of model training is to enable the model to classify input objects in at least two categories.

[0028] For classification models, the purpose of model training is to enable the model to have the ability to classify inputs, including binary classification models and multi-classification models. For binary classification models, the inputs are classified into two categories, while for multi-classification models, the inputs are classified into more than two categories.

[0029] If the samples belonging to a certain category in the training samples used to train the classification model are sparse samples, that is, the number of samples accounts for a very low proportion of the total samples, then when using the samples to train the model, if the existing multi-hot or one-hot encoding methods are adopted, and the full amount of object samples are used as the feature dimension, the obtained feature encoding will not have feature discrimination in the category, and the trained model will not be able to identify the objects belonging to the category and cannot accurately classify them. For example, in the credit risk control scenario, the model is required to judge whether a certain application APP has a high risk of fraud after being installed, that is, to achieve a binary classification task of high risk and low risk. However, since the number of samples of such high-risk APPs installed and fraudulent behavior is originally small, they are sparse samples and the positive sample coverage rate is very low. The existing multi-hot or one-hot encoding methods all use the full amount of APP as the feature dimension, and the samples of the installed APP are marked as 1, and the samples of the uninstalled APP are marked as 0. Not only is the feature dimension high, but also the sample features of the sparse samples have extremely low feature discrimination and contribution when used for model training, which affects the model training effect and cannot accurately identify objects with high risks.

[0030] In the disclosed embodiment, the first weights of the multiple objects in at least two categories are calculated based on the number of samples of the multiple objects and the number of samples of the multiple objects under at least two categories. The first weight is used to represent the importance of each object in each category among the multiple objects, and can be represented by parameters such as TF-IDF (term frequency-inverse document frequency), lift value (a measure to evaluate whether the prediction model is effective), woe (Weight of Evidence) or category proportion. Among them, the total number of samples can be determined according to the number of samples of each object in the multiple objects, and then the first weight of each object in each category is calculated by calculating TF-IDF, lift value, woe or category proportion in combination with the number of samples of each object under each category. Moreover, in the scenario where the sample of the object belongs to a sparse sample, but the sparse sample belongs to a positive sample, even if the number of sparse samples is small, the first weight of these objects in the positive sample is higher than that of other samples, which is equivalent to extracting important sparse values ​​through the first weight, thereby improving the contribution of the sparse sample. For example, in a credit risk control scenario, the two categories to be classified are high-risk positive samples and low-risk negative samples. Then, although the number of samples of apps with high credit risks installed is smaller than that of other samples, their first weight in the positive samples is higher than that of other samples.

[0031] S102: bin multiple objects according to a first weight to obtain multiple object bins.

[0032] By using the binning method, the first weight of each object in the multiple objects can be divided, and each of the multiple object bins obtained contains multiple objects, and the first weights of these objects are similar. The specific binning process can be implemented according to the existing technology and will not be repeated here.

[0033] S103. Calculate second weights of the multiple object bins in at least two categories based on the number of samples in the multiple object bins and the number of samples in the multiple object bins under at least two categories, and use the second weights of the multiple object bins as feature values ​​of the multiple object bins.

[0034] The second weight is used to indicate the importance of each object bin in each category, and the calculation method can be the same as or different from the calculation method of the first weight. The calculated second weight is the feature value of each object bin, which is used for training the classification model.

[0035] It should be noted that in the disclosed embodiment, before binning, samples are based on each object, and the number of samples is very large. After binning according to the first weight, each object bin is used as a unit, so that the number of samples in the object bin is reduced exponentially. Therefore, the coverage of the object bins corresponding to the sparse samples in the overall object bin samples is improved. By training the model with the feature values ​​of the object bins, the discrimination and contribution of the sparse sample features can be improved, so that the model has the ability to classify objects in the category to which the sparse samples belong. For example, in the credit risk control scenario, it is necessary to predict whether there is a high risk of fraud in the installation of the APP, or to perform classification and prediction in binary classification scenarios such as anti-fraud identification and marketing response.

[0036] The technical solution of the embodiment of the present disclosure first calculates the first weight of each object in each category, and then bins each object according to the first weight to obtain multiple object bins. Next, the second weight of each object bin in each category is calculated based on each object bin, and the second weight is used as the feature value of each object bin to train the model. Through binning, the original large number of sample values ​​is reduced exponentially, which naturally increases the coverage of a small number of sparse samples in the total sample, improves the problem of low discrimination and contribution of sparse sample features in the prior art, and improves the training effect of the model.

[0037] In one embodiment, the model can be a binary classification model, and correspondingly, the categories include the category to which the positive sample belongs and the category to which the negative sample belongs, and the positive sample is a sparse sample. Then, by calculating the first weight of each object in the category to which the positive sample belongs according to the method of the embodiment of the present disclosure, the importance of the sparse sample in the sample population can be expressed, and data preparation can be done for further binning operations.

[0038] Figure 2 is a flow chart of a feature encoding method according to an embodiment of the present disclosure. Based on the above embodiment, this embodiment makes further optimization for the case where the model is a binary classification model and the weight is calculated using word frequency-inverse document frequency. Figure 2 As shown, the method specifically includes the following:

[0039] S201. Calculate reverse file frequencies of multiple objects according to the sample numbers of multiple objects and the sum of the sample numbers of multiple objects.

[0040] Among them, the inverse document frequency is IDF (Inverse Document Frequency). Generally, the IDF of a specific word can be obtained by dividing the total number of documents by the number of documents containing the word, and then taking the logarithm of the quotient. In the embodiment of the present disclosure, the "document" here is the object, so the calculation formula of the inverse document frequency of each object is: Idf = log [total number of samples / (number of samples of each object + 1)], where the total number of samples is the sum of the number of samples of each object.

[0041] In one embodiment, taking the credit risk control scenario as an example, the model is used to identify whether the act of installing a certain APP has a high risk of fraud. Accordingly, the object refers to the application APP, the sample number of the object refers to the number of APPs installed, and the categories include two categories: the APP is installed and there is a high risk of financial fraud and the APP has a low risk of financial fraud. The positive sample in the sample of the object refers to the APP being installed and there is a high risk of financial fraud. Therefore, in the credit risk control scenario, the calculation formula for the reverse file frequency of each APP can be adjusted to: Idf = log [total number of samples / (number of samples of installing the APP + 1)], where the total number of samples is the sum of the number of samples of installing each APP.

[0042] S202: Calculate the proportion of positive samples of the multiple objects according to the total number of positive samples in the samples of the multiple objects and the sum of the number of samples of the multiple objects, wherein the positive samples are sparse samples.

[0043] Specifically, the total number of positive samples in the sample of a certain object is divided by the sum of the number of samples of each object to get the proportion of positive samples of the object, which is similar to the proportion of words in the text TF (word frequency). In the credit risk control scenario, it is the total number of positive samples of the APP installation divided by the sum of the number of samples of each APP installation. In the credit risk control scenario, the proportion of positive samples of installing such high-risk APPs and committing fraudulent behaviors in the total sample is very small, so these positive samples are sparse samples.

[0044] S203: Multiply the reverse file frequencies of the multiple objects by the proportion of positive samples, and use the result as the first weight of the multiple objects in the category to which the positive samples belong.

[0045] The term frequency-inverse document frequency (TF-IDF) technology is a commonly used weighting technology for information retrieval and text mining, which can be used to evaluate the importance of a word to a document set or a document in a corpus. Its meaning is that the importance of a word increases in direct proportion to the number of times it appears in the file, but at the same time it decreases in inverse proportion to the frequency of its appearance in the corpus. In the disclosed embodiment, the term frequency-inverse document frequency technology is applied to the first weight of the calculation object in the category to which the positive sample belongs. Since the positive sample is a sparse sample, the term frequency-inverse document frequency technology can be used to find important sparse values, and then the relatively similar values ​​are aggregated together through binning, thereby improving the discrimination and contribution of sparse features. Moreover, in the calculation results, the first weight of the APP with high risk true value is higher than that of other APPs, and among these APPs with high risk true value, the higher the risk, the greater the value of the first weight of the corresponding APP.

[0046] S204: Reversely sort the first weights, and bin the multiple objects based on the reversely sorted first weights to obtain multiple object bins.

[0047] Specifically, it can be implemented by equal-frequency binning or custom binning. Among them, if equal-frequency binning is adopted, the number of object samples included in each object bin obtained is the same. As for custom binning, the principle is to ensure that the object sample coverage rate corresponding to each object bin must be higher than a certain preset threshold, and some bins cannot be particularly large and some bins cannot be particularly small. Accordingly, the reverse sorting method is adopted in the embodiment of the present disclosure, and its purpose is to facilitate manual determination of the binning threshold when custom binning, that is, how many object samples are included in the bin, and how many bins there are in total.

[0048] S205: Calculate the reverse file frequencies of the multiple object bins according to the number of samples of the multiple object bins and the sum of the number of samples of the multiple object bins.

[0049] S206. Calculate the proportion of positive samples in the multiple object bins according to the total number of positive samples in the samples of the multiple object bins and the sum of the number of samples in the multiple object bins, wherein the positive samples are sparse samples.

[0050] S207: taking the result of multiplying the reverse file frequency of the multiple object bins by the positive sample proportion as the second weights of the multiple object bins in the category to which the positive samples belong, and taking the second weights of the multiple object bins as the feature values ​​of the multiple object bins.

[0051] The method of calculating the second weight is the same as the first weight, both of which are calculated using the word frequency-inverse document frequency technology, but in the process of calculating the second weight, each object bin is treated as an independent unit. Specifically, the calculation formula for the inverse document frequency of each object bin is: Idf = log [total number of samples / (number of samples in each object bin + 1)], where the total number of samples is the sum of the number of samples in each object bin. Then, the total number of positive samples in the samples of a certain object bin is divided by the sum of the number of samples in each object bin to obtain the proportion of positive samples in the object bin. Finally, the inverse document frequency of each object bin is multiplied by the proportion of positive samples to obtain the second weight.

[0052] Correspondingly, in the credit risk control scenario, the calculation formula for the reverse file frequency of each sub-box APP can be adjusted to: Idf = log [total number of samples / (number of samples of the sub-box APP set installed + 1)], where the total number of samples is the sum of the number of samples of each sub-box APP installed; the total number of positive samples of a sub-box APP set installed divided by the sum of the number of samples of each sub-box APP installed can be calculated to get the proportion of positive samples of the sub-box APP; finally, multiply the reverse file frequency of the sub-box APP by the positive sample proportion to get the second weight of each sub-box APP.

[0053] The technical solution of the embodiment of the present disclosure first uses the word frequency-inverse file frequency technology to calculate the first weight of each object in the category to which the positive sample as the sparse sample belongs, and then divides each object into bins according to the first weight after reverse sorting to obtain multiple object bins. Next, taking each object bin as a unit, the second weight of each object bin in the category to which the positive sample belongs is calculated, and the second weight is used as the feature value of each object bin to train the model. In this way, on the one hand, the word frequency-inverse file frequency technology is used to find out important sparse values ​​as a measure of the importance of sparse samples. On the other hand, through binning, the original large number of sample values ​​is reduced exponentially. Naturally, the coverage rate of a small number of sparse samples in the total sample is increased, improving the problem of low discrimination and contribution of sparse sample features in the prior art, and improving the training effect of the model.

[0054] Furthermore, in the credit risk control scenario, since the samples of installing risky APPs and used for fraudulent behavior are few in themselves, they are sparse samples. Therefore, when performing feature encoding, if the sparse samples with a small number are feature encoded according to the prior art and applied to modeling, the contribution is extremely low. However, by adopting the technical solution of the embodiment of the present disclosure, the corresponding weight of each APP in the high-risk positive sample is first calculated using TF-IDF, and the important sparse values ​​are found to increase the importance of the sparse samples. Then, the bins are divided according to the weights, so that the original large number of sample values ​​are reduced exponentially after binning, and naturally, the coverage rate of a small number of sparse samples in the sample population is increased. Finally, the weight of each bin is calculated again with each bin APP as the unit, and the feature value of the bin can be obtained. Moreover, the feature value obtained by the above feature encoding method is monotonic, that is, the more high-risk APPs are installed, the higher the feature value of the bin where the high-risk APP is located, and vice versa. In this way, better results and interpretability can be achieved without further feature engineering in the linear model.

[0055] Figure 3 This is a schematic diagram of the structure of a feature encoding device according to an embodiment of the present disclosure. This embodiment can be applied to the case where feature encoding of training samples is performed to train a model using the sample features obtained after encoding, especially for the case of feature encoding of sparse samples. It relates to the field of machine learning technology, especially to smart finance, artificial intelligence and deep learning technology. The device can implement the feature encoding method described in any embodiment of the present disclosure. Figure 3 As shown, the device 300 specifically includes:

[0056] A first weight calculation module 301 is used to calculate first weights of the multiple objects in at least two categories according to the number of samples of the multiple objects and the number of samples of the multiple objects in at least two categories, wherein the goal of the model training is to enable the model to classify the input objects in the at least two categories;

[0057] A binning module 302, configured to bin the plurality of objects according to the first weights to obtain a plurality of object bins;

[0058] The second weight calculation module 303 is used to calculate the second weights of the multiple object bins in the at least two categories based on the number of samples in the multiple object bins and the number of samples in the multiple object bins under the at least two categories, and use the second weights of the multiple object bins as feature values ​​of the multiple object bins.

[0059] Optionally, the model is a binary classification model, the categories include the category to which the positive sample belongs and the category to which the negative sample belongs, and the positive sample is a sparse sample.

[0060] Optionally, the first weight calculation module 301 includes:

[0061] A first reverse file frequency calculation unit, configured to calculate the reverse file frequencies of the plurality of objects according to the number of samples of the plurality of objects and the sum of the number of samples of the plurality of objects;

[0062] A first positive sample ratio calculation unit, configured to calculate the positive sample ratio of the multiple objects according to the total number of positive samples in the samples of the multiple objects and the sum of the number of samples of the multiple objects;

[0063] The first weight calculation unit is used to calculate the result of multiplying the reverse file frequencies of the multiple objects by the positive sample proportion as the first weights of the multiple objects in the category to which the positive sample belongs.

[0064] Optionally, the second weight calculation module 303 includes:

[0065] a second reverse file frequency calculation unit, configured to calculate the reverse file frequencies of the plurality of object bins according to the number of samples of the plurality of object bins and the sum of the number of samples of the plurality of object bins;

[0066] A second positive sample ratio calculation unit, configured to calculate the positive sample ratios of the multiple object bins according to the total number of positive samples in the samples of the multiple object bins and the sum of the number of samples of the multiple object bins;

[0067] A second weight calculation unit is used to use the result of multiplying the reverse file frequency of the multiple object bins by the proportion of the positive samples as the second weights of the multiple object bins in the category to which the positive samples belong, and use the second weights of the multiple object bins as feature values ​​of the multiple object bins.

[0068] Optionally, the binning module 302 is specifically used for:

[0069] The first weights are reversely sorted, and the plurality of objects are binned based on the reversely sorted first weights to obtain a plurality of object bins.

[0070] Optionally, the binning module 302 is specifically used for:

[0071] According to the first weight and using an equal frequency binning method, the multiple objects are binned to obtain multiple object bins.

[0072] Optionally, the model is used for credit risk control;

[0073] Correspondingly, the object refers to an application, the sample number of the object refers to the number of times the application is installed, the category includes two categories: the application is installed and there is a high risk of financial fraud and the application has a low risk of financial fraud, and the positive sample in the sample of the object refers to the application being installed and there is a high risk of financial fraud.

[0074] The above-mentioned product can execute the method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0075] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0076] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0077] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0078] like Figure 4 As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0079] A number of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0080] The computing unit 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as feature encoding methods. For example, in some embodiments, the feature encoding method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the feature encoding method described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the feature encoding method in any other appropriate manner (e.g., by means of firmware).

[0081] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0082] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0083] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0084] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0085] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0086] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services. The server may also be a server of a distributed system, or a server combined with a blockchain.

[0087] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, knowledge graph technology, and other major directions.

[0088] Cloud computing refers to a technology system that uses network access to elastically scalable shared physical or virtual resource pools. Resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for technical applications such as artificial intelligence and blockchain, as well as model training.

[0089] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions provided by this disclosure can be achieved, and this document does not limit this.

[0090] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A feature encoding method, include: According to the number of samples of multiple objects and the number of samples of the multiple objects in at least two categories, first weights of the multiple objects in the at least two categories are calculated, wherein the goal of model training is to enable the model to classify input objects in the at least two categories; the model is a binary classification model, the categories include the category to which the positive sample belongs and the category to which the negative sample belongs, and the positive sample is a sparse sample; the sparse sample means that the number of samples accounts for a low proportion in the total samples; the object refers to an application; Binning the multiple objects according to the first weight to obtain multiple object bins; Calculating second weights of the multiple object bins in the at least two categories according to the number of samples in the multiple object bins and the number of samples in the multiple object bins under the at least two categories, and using the second weights of the multiple object bins as feature values ​​of the multiple object bins; The step of calculating first weights of the plurality of objects in the at least two categories includes: Calculating reverse file frequencies of the multiple objects according to the sample numbers of the multiple objects and the sum of the sample numbers of the multiple objects; Calculate the proportion of positive samples of the multiple objects according to the total number of positive samples in the samples of the multiple objects and the sum of the number of samples of the multiple objects; A result of multiplying the reverse file frequencies of the multiple objects by the positive sample ratio is used as a first weight of the multiple objects in the category to which the positive sample belongs; the first weight is used to represent the importance of each object in each category.

2. The method according to claim 1, in, The calculating second weights of the plurality of object bins in the at least two categories comprises: Calculating the reverse file frequencies of the plurality of object bins according to the number of samples of the plurality of object bins and the sum of the number of samples of the plurality of object bins; Calculate the proportion of positive samples in the multiple object bins according to the total number of positive samples in the samples of the multiple object bins and the sum of the number of samples in the multiple object bins; A result of multiplying the inverse file frequencies of the multiple object bins by the positive sample ratio is used as a second weight of the multiple object bins in the category to which the positive sample belongs.

3. The method according to claim 1, in, The binning the multiple objects according to the first weight to obtain multiple object bins includes: The first weights are reversely sorted, and the plurality of objects are binned based on the reversely sorted first weights to obtain a plurality of object bins.

4. The method according to claim 1, in, The binning the multiple objects according to the first weight to obtain multiple object bins includes: According to the first weight and using an equal frequency binning method, the multiple objects are binned to obtain multiple object bins.

5. The method according to claim 1, in, The model is used for credit risk control; The sample number of the object refers to the number of times the application is installed, the categories include two categories: the application is installed and there is a high risk of financial fraud and the application has a low risk of financial fraud, and the positive samples in the samples of the object refer to the samples in which the application is installed and there is a high risk of financial fraud.

6. A feature encoding device, include: A first weight calculation module is used to calculate the first weights of the multiple objects in the at least two categories according to the number of samples of the multiple objects and the number of samples of the multiple objects in the at least two categories, wherein the goal of model training is to enable the model to classify the input objects in the at least two categories; the model is a binary classification model, the categories include the category to which the positive sample belongs and the category to which the negative sample belongs, and the positive sample is a sparse sample; the sparse sample means that the number of samples accounts for a low proportion of the total samples; the object refers to an application; A binning module, configured to bin the plurality of objects according to the first weight to obtain a plurality of object bins; a second weight calculation module, configured to calculate second weights of the plurality of object bins in the at least two categories according to the number of samples in the plurality of object bins and the number of samples in the plurality of object bins under the at least two categories, and use the second weights of the plurality of object bins as feature values ​​of the plurality of object bins; Wherein, the first weight calculation module includes: A first reverse file frequency calculation unit, configured to calculate the reverse file frequencies of the plurality of objects according to the number of samples of the plurality of objects and the sum of the number of samples of the plurality of objects; A first positive sample ratio calculation unit, configured to calculate the positive sample ratio of the multiple objects according to the total number of positive samples in the samples of the multiple objects and the sum of the number of samples of the multiple objects; A first weight calculation unit is used to calculate the result of multiplying the reverse file frequency of the multiple objects by the positive sample ratio as the first weight of the multiple objects in the category to which the positive sample belongs; the first weight is used to represent the importance of each object in each category.

7. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the feature encoding method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to enable a computer to execute the feature encoding method according to any one of claims 1-5.

9. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the feature encoding method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Credit object classification method and device, terminal and storage medium

    CN111815436A

  • Credit risk control model generation method and device, score card generation method, machine readable medium and equipment

    CN111898675A