Device Characterization Method and Apparatus, and Device Identification Reset Fraud Detection Method and Apparatus
By constructing the frequency distribution of the basic information elements of the device and the automatic encoder output vector, combined with classifier detection, the problem of device identification reset fraud detection is solved, and effective identification and distinction of fraud devices is achieved.
Patent Information
- Application Number
- CN202010542814.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-06-15
AI Technical Summary
The prior art cannot effectively detect device identification reset fraud, and fraudsters simulate new users in fraudulent activities by resetting device identification, resulting in no real value in the growth of advertisers or applications.
By constructing the frequency distribution of basic information elements of the device, using an automatic encoder to output vectors to characterize the device, and using a classifier to detect the device identification and reset fraud, collecting basic information elements that do not endanger user privacy for distinction.
It realizes the effective distinction between normal equipment and fraudulent equipment, provides effective detection of device identification reset fraud, and avoids insufficient detection relying on IP and BSSID relationships.
Smart Images

Figure CN113869080B_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, mobile communication technologies, and particularly to a device characterization method and apparatus, and a device identifier reset fraud detection method and apparatus. Background Art
[0002] Mobile applications (APPs) use device identifiers (device IDs) to track multiple access records of a device. However, the use of device IDs has also brought new fraudsters, who have taken the opportunity to massively reset their device IDs. Among them, device ID reset fraud is a periodic fraud activity. Fraudsters click on advertisements, install applications, generate in-app interactions, then reset the device ID, and then uninstall and reinstall the application to simulate a new installation from a "new user". Although the installation and participation of the application seem to be genuine, these so-called "new users" have no real value to advertisers or the growth of the application and are fraudsters.
[0003] The related art does not provide an effective method for detecting device ID reset fraud. Summary of the Invention
[0004] This application provides a device characterization method and apparatus, and a device identifier reset fraud detection method and apparatus, which can provide guarantee for effective detection of device ID reset fraud.
[0005] An embodiment of the present invention provides a device characterization method, including:
[0006] Constructing a frequency distribution according to basic information elements of a device;
[0007] Representing an instance of a basic information element of the device according to the constructed frequency distribution;
[0008] Inputting the instance of the basic information element of the device represented by the frequency distribution into an autoencoder, and using the vector output by the autoencoder to characterize the device.
[0009] In an exemplary instance, the constructing a frequency distribution according to basic information elements of the device includes:
[0010] Selecting one or more pairs of elements from the basic information elements of the device; wherein, for the first element in the pair of elements, most of the data subsets obtained by splitting the first element instance corresponding to the first element have sufficient samples; and for any first element instance, the second element instances corresponding to the second element are diversified;
[0011] Based on the law of large numbers of Bernoulli, constructing a frequency distribution for any basic information element instance i in the basic information elements in the pair of elements, to obtain a distribution set including one or more frequency distributions.
[0012] In an exemplary instance, an element with an information entropy lower than a first threshold is used as the first element in the element pair;
[0013] For a given first element, an element with a difference higher than a set second threshold is used as the second element in the element pair, where the difference is: the difference between the information gain rate of the information gain of the second element with respect to the first element and the information gain rate of the information gain of the first element with respect to the second element.
[0014] In an exemplary instance, the frequency distribution constructed according to the basic information elements of the device represents an instance of the basic information elements of the device, including:
[0015] Replacing the basic information elements of the device with the distribution set.
[0016] In an exemplary instance, the vector output by the autoencoder is the root mean square error.
[0017] The embodiments of the present application further provide a computer-readable storage medium storing computer-executable instructions for executing the device characterization method described in any one of the above.
[0018] The embodiments of the present application further provide a device for implementing device characterization, including a memory and a processor, where the memory stores the following instructions executable by the processor: steps for executing the device characterization method described in any one of the above.
[0019] The embodiments of the present application further provide a device for implementing device characterization, including: a construction module, a replacement module, and a characterization module; where
[0020] The construction module is configured to construct a frequency distribution according to the basic information elements of the device;
[0021] The replacement module is configured to represent an instance of the basic information elements of the device according to the constructed frequency distribution;
[0022] The characterization module is configured to input the instance of the basic information elements of the device represented by the frequency distribution into an autoencoder, and use the vector output by the autoencoder to characterize the device.
[0023] The embodiments of the present application further provide a method for detecting device identity reset fraud, including:
[0024] Inputting the device information characterized by the frequency distribution into a classifier;
[0025] The output of the classifier indicates whether the device is a device with device identity reset fraud;
[0026] where the classifier is trained according to the device information samples characterized by the frequency distribution;
[0027] Among them, device information of the frequency distribution representation is obtained by the device characterization method according to any one of the above.
[0028] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions for executing the device identification reset fraud detection method according to any one of the above.
[0029] An embodiment of the present application further provides a device for implementing device identification reset fraud detection, including a memory and a processor. Among them, the following instructions executable by the processor are stored in the memory: steps for the device identification reset fraud detection method according to any one of the above.
[0030] An embodiment of the present application only needs to collect basic information elements of devices that do not endanger user privacy to well characterize the devices, which has a strong discrimination for distinguishing normal devices and fraud devices, provides an effective guarantee for device identification reset fraud detection, and enables effective device identification reset fraud detection.
[0031] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the specification, claims, and drawings. Description of the Drawings
[0032] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application and do not constitute a limitation to the technical solutions of the present application.
[0033] Figure 1 It is a schematic diagram of the IP frequency distribution under the dimension of the basic information elements of the device in the present application;
[0034] Figure 2 It is a flowchart of the device characterization method in the present application;
[0035] Figure 3 It is a process example diagram of an embodiment for implementing device characterization in the present application;
[0036] Figure 4 It is a schematic diagram of the composition structure of the device characterization device in the present application. Detailed Embodiments
[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be described in detail below with reference to the drawings. It should be noted that, without conflict, the embodiments and features in the embodiments of the present application can be combined with each other arbitrarily.
[0038] In a typical configuration of the present application, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0039] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0040] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0041] The steps illustrated in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions. And, although a logical order is illustrated in the flowchart, in some cases, the steps shown or described may be executed in a different order than herein.
[0042] Suppose μ is the number of times event A occurs in n independent trials, and the probability of event A occurring in each trial is P. Then, for any positive number ε, there is the formula shown in (1):
[0043]
[0044] What formula (1) shows is the Bernoulli's law of large numbers. The Bernoulli's law of large numbers is a special case of the Chebyshev's law of large numbers. Its meaning is that when n is large enough, the frequency of the occurrence of event A will be almost close to its probability of occurrence, that is, the stability of frequency.
[0045] According to the law of large numbers of Bernoulli, if there are enough samples, then the frequency of an event will be almost close to its probability. Through analysis, the inventors of this application have obtained that the frequency distributions under the basic information elements of normal devices are usually similar and smooth, while the frequency distributions under the basic information elements of abnormal devices usually contain some spikes caused by synchronous behaviors. A set of comparable frequency distributions includes two elements. The first element is used to determine from which dimension to construct the distribution set, and the second element is used to determine the type of frequency distribution to be established. Here is an example. As Figure 1 shown, in this embodiment, taking the IP frequency distribution with the first element being the Brand dimension and the second element being the IP address (i.e., the frequency distribution of the IP address) as an example, "Brand A", "Brand C", and "Brand D" are Brand instances with almost no device ID forgery. The IP frequency distributions under these element instances are similar and smooth; "Brand B" is a Brand instance with a relatively large number of device ID forgeries. The IP frequency distribution under this element instance is not similar to other distributions and has many protrusions. Based on the analysis results, the inventors of this application propose that the basic information elements of the device can be represented as frequency distributions, and then an algorithm can be used to extract the main components of the distribution. The degree of deviation from the main components is used as the outlier of the distribution, and the devices corresponding to these distributions with outliers can be determined as devices with device ID reset fraud, that is, fraudulent devices. Since a device can collect multiple basic information elements of the device, a device can be represented as an ordered set of outliers.
[0046] Among them, the basic information elements of the device include but are not limited to: device ID (DeviceID), processor name (ProcName), brand of the device (Brand), model of the device (Model), hardware information of the device (Hardware), manufacturer information (Manufacturer), motherboard information of the device (Board), device CPU, device name (DeviceName), Host, operating system (OS), resolution (Resolution), Display, wifi (Bssid) connected when the device ID is generated, IP address (Internet Protocol Address, abbreviated as IP in this article), etc. Correspondingly, an instance of the basic information element of the device is a specific value corresponding to the basic information element of the device. For example, if the Brand of the device is the basic information element, then "Brand A" is an instance of the brand of the device. That is to say, "Brand A" is an instance of the basic information element of the device.
[0047] In this article, it is assumed that the ID (Di) of device i, i.e., DeviceID = Di, and the basic information element of device i is represented as DeviceInf(Di). Then, DeviceInf(Di) can be represented by a multi-tuple [DeviceID, ProcName, Brand, Model, Hardware, Manufacturer, Board, CPU, DeviceName, Host, OS, resolution, display, Bssid, IP, fingerprint,...].
[0048] This application provides a device characterization method, as Figure 2 shown, including:
[0049] Step 200: Construct a frequency distribution based on the basic information elements of the device.
[0050] From the analysis of the inventors of this application, it can be seen that a set of comparable frequency distributions includes two elements, called an element pair. The first element in the element pair is used to determine from which dimension to construct the distribution set, and the second element in the element pair is used to determine the type of frequency distribution to be established.
[0051] In an exemplary instance, the element pairs formed by the basic information elements of the device can include one or more. For example, by selecting b pairs of element pairs, a distribution set including b frequency distributions can be constructed.
[0052] It should be particularly noted here that the element pair composed of the first element and the second element is a meaningful element pair. That is to say, according to the law of large numbers of Bernoulli, when b is large enough, the frequency distribution is stable. The inventors of this application have analyzed and obtained that the first condition for constructing a meaningful frequency distribution is that most of the data subsets sliced by the first element instance corresponding to the first element have sufficient samples. The second condition is that under any first element instance, the second element instance corresponding to the second element needs to be sufficiently diversified. Therefore, the second element instance should provide sufficient information to the first element instance, while the first element instance should provide little information to the second element instance. For example: the first element is Brand and the second element is IP address is a meaningful element pair; another example: the first element is Brand and the second element is Model is a meaningful element pair; another example: the first element is Model and the second element is Brand is a meaningless element pair, and so on, and no more examples are given here.
[0053] Assume that a set of devices Di ∈ D and the basic information elements DeviceInf(Di) of this set of devices are given, and an element pair {Pair(tk, th)} is selected to construct a frequency distribution, where tk represents the first element in the element pair and th represents the second element in the element pair.
[0054] Let \(K\) denote the set of instances of the first element \(t_k\). The first - element instances in the dataset contain \(r\) data samples with \(o\) different values. Then the information entropy of the first element \(t_k\) is shown in Equation (2):
[0055]
[0056] In Equation (2), \(p\) i is the probability that any sample belongs to a specific value \(c\) i . Let \(r\) i be the number of samples with value \(c\) i . Then \(p\) i = \(r\) i / \(r\). Taking \(t_k\) as the brand, assuming there are a total of 100 brands, then \(o = 100\).
[0057] For example, taking \(K\) as the data set representing brands, such as \(K=\{\text{Brand A},\text{Brand B},\text{Brand C},\cdots\}\), then \(c_1=\text{‘Brand A’}\). In the entire data set of \(K\) above, if there are 10,000 devices, then \(r = 10,000\). If the number of devices of Brand A is 2,000, then \(r_1 = 2,000\) and \(p_1=2,000 / 10,000 = 0.2\). If the number of devices of Brand B is 3,000, then \(r_2 = 3,000\) and \(p_2 = 3,000 / 10,000 = 0.3\), \(\cdots\). For example, thus, \(I(K)=-(0.2\times\log0.2 + 0.3\times\log0.3+\cdots)\).
[0058] Let \(H\) denote the set of instances of the second element \(t_h\). Let \(r\) ij be the number of samples with value \(c\) j in the subset \(H\) i . The information entropy of the second element \(t_h\) under the condition of the first element \(t_k\) is shown in Equation (3):
[0059]
[0060] For example, assume that H represents a set of models and K represents a set of brands. In this embodiment, assume that there are a total of 10,000 devices, and K = {Brand A, Brand B, Brand C, … p}, where the number of devices of Brand A is 2,000, i.e., r1 = 2,000, the number of devices of Brand B is 3,000, i.e., r2 = 3,000, …; In this embodiment, assume that H = {Model 1, Model 2, Model 3, Model 4, Model 5, Model 6, … i}, the number of devices with model being Model 1 is 1,000, among which the number of devices with brand = ‘Brand A’ and model = ‘Model 1’ is 400, the number of devices with brand = ‘Brand B’ and model = ‘Model 1’ is 600, the number of devices with brand = ‘Brand A’ and model = ‘Model 2’ is 1,100, the number of devices with brand = ‘Brand A’ and model = ‘Model 3’ is 500. Then, r 11 = 400, r 12 = 1,100, r 13 = 500, r 21 = 6,000. According to formula (2) and formula (3), in this embodiment, I(H|K =, Brand A’) = -((400 / 2,000)*log(400 / 2,000) + (1,100 / 2,000)*log(1,100 / 2,000) + (500 / 2,000)*log(500 / 2,000)).
[0061] Therefore, the information gain Gain(H|K) brought by the second element to the first element is as shown in formula (4):
[0062] Gain(H|K) = I(K) - E(H|K) (4)
[0063] To standardize the information gain, the normalization formula is defined as shown in (5):
[0064]
[0065] Therefore, the information gain ratio Gain_Ratio(H|K) of the information gain brought by H to K is as shown in formula (6):
[0066] Gain_Ratio(H|K) = Gain(H|K) / SplitInfroH(K) (6)
[0067] In this way, the gain ratio can be defined for Pair_gain_ratio(H|K) as shown in formula (7):
[0068] Pair_gain_ratio(H|K) = Gain_Ratio(H|K) - Gain_Ratio(K|H) (7)
[0069] Formula (7) represents the difference between the information gain ratio of the second element with respect to the information gain of the first element and the information gain ratio of the first element with respect to the information gain of the second element for a given first element.
[0070] According to the analysis of the inventors of the present application, the first condition for constructing the frequency distribution of the present application is that most data subsets sliced by the first element instances have sufficient samples. Therefore, in an embodiment of the present application, an element with I(K) lower than the set first threshold can be used as the first element tk. In an exemplary instance, the empirical value of this low threshold can be, for example, lower than 0.2;
[0071] The second condition for constructing the frequency distribution of the present application is that the second element instances need to be fully diversified under any first element instance. Therefore, the second element instances should provide sufficient information to the first element instances, while the first element instances should provide little information to the second element instances. Therefore, in an embodiment of the present application, for a given first element, an element with Pair_gain_ratio(H|K) higher than the set second threshold can be selected as the second element th. In an exemplary instance, the empirical value of this high threshold can be, for example, greater than 0.5.
[0072] Assume that a basic information element instance k is given i , where k i ∈K, and a selected element pair Pair(tk, th), then constructing a unique frequency distribution includes:
[0073] Represent the frequency distribution as DTk i (th), where DTk i (th) = [d1, d2, d3,..., d m . Let H(k i ) be the set of second elements when the first element instance is k i , and l can be used to represent how many second elements there are. Let fre(h j ) be the number of occurrences of h j extracted from this data set. Then, the value of d i can be as shown in formula (8):
[0074]
[0075] For example, for a certain first element instance k i = 'Brand A', assuming that the overall data set has 10,000 devices, k i = 'There are 2,000 devices of Brand A, and H(k i) For the set of all models in these 2000 devices, H(k i ) = {Model 1, Model 2, Model 3, Model 4}. At this time, l = 4. In this embodiment, it is assumed that Model 1 appears 100 times, Model 2 appears 200 times, Model 3 appears 300 times, and Model 4 appears 1400 times. Then, fre('Model 1') = 100, fre('Model 2') = 200, fre('Model 3') = 300, fre('Model 4') = 1400. At this time,
[0076] d i = |{h j ∈ H(k i )|ln(fre(h j )) = (i - 1) / 10 - |, which represents the number of h j such that ln(fre(h i )) = (i - 1) / 10. For example, if there are 0 second elements that satisfy ln(fre(h j )) = 0, then, d1 = 0.
[0077] In an exemplary instance, step 200 may include:
[0078] Select one or more element pairs {Pair(tk, th)} from the basic information elements DeviceInf(Di) of the device;
[0079] An element that can divide the data set into multiple parts and each part has a sufficient amount of data, that is, an element with a low information entropy (i.e., low diversity) of the first element below the first threshold is used as the first element tk in the element pair {Pair(tk, th)}, and an element with a stable frequency distribution in each sub - data set after the first element is divided, that is, an element with a difference higher than the set second threshold relative to the first element is used as the second element th in the element pair {Pair(tk, th)}, where the difference is: the difference between the information gain rate of the information gain of the second element with respect to the first element and the information gain rate of the information gain of the first element with respect to the second element;
[0080] Based on the law of large numbers of Bernoulli, construct a frequency distribution for any basic information element instance ki in the basic information element DeviceInf(Di) in the element pair.
[0081] Step 201: Represent the basic information element instance of the device according to the constructed frequency distribution.
[0082] A device has multiple basic information elements. Therefore, a device is represented as a combination of multiple distributions.
[0083] In an exemplary instance, replacing the DeviceInf (Di) with n selected frequency distributions x is as follows: x = [x1, x2,... x n , where x represents an ordered set of n frequency distributions, and x i = [d i1 , d i2 ,... d im . That is, the basic information elements of the device are replaced with a distribution set.
[0084] In this way, the device is represented by constructing n frequency distribution curves based on the information gain rate.
[0085] Step 202: Input the basic information element instance of the device represented by the frequency distribution into the autoencoder, and use the vector output by the autoencoder to characterize the device.
[0086] An autoencoder (AE, AutoEncoder) is a powerful unsupervised deep model for functional learning. AE has been widely used in various machine learning applications. The characteristic of an autoencoder is that the encoder creates a hidden layer (or multiple hidden layers) that contains a low-dimensional vector representing the meaning of the input data. AE can assist in data classification, visualization, and storage. AE is an unsupervised learning mode that only requires input data. An autoencoder usually includes three layers: an input layer, a hidden layer, and an output layer. The definition of the autoencoder is shown in formula (9):
[0087]
[0088] In formula (9), d i ∈ R m , represents the i-th input data, h i ∈ R m′ , represents the hidden representation of the autoencoder, represents the reconstructed data from the auto-decoder; Θ = {W(1), W(2), b(1), b(2)} represents the model parameters of the autoencoder; δ(.) represents the non-linear activation function. In an exemplary instance, the model parameters can be learned by minimizing the root mean square reconstruction error.
[0089] The output result of the autoencoder is represented by the root mean square error (RMSE, root-mean-square error), as shown in formula (10):
[0090]
[0091] Therefore, for the DeviceInf(Di) represented by n selected frequency distributions x in the embodiments of the present application, where x represents an ordered set of n frequency distributions, x = [x1, x2,... x n , where x i = [d i1 , d i2 ,... d im .
[0092] In an exemplary instance, assume that the input of the i-th autoencoder is: x i = [d i1 , d i2 ,... d im , then, according to the working principle of the autoencoder, the output is: RMSE i .
[0093] For the instance characterized by the frequency distribution x of the embodiments of the present application, the encoding result of this instance after passing through the autoencoder can be expressed as: x' = [RMSE1, RMSE2,... RMSE n .
[0094] In this step, the constructed frequency distribution is used as the input, and the RMSE output after quantization processing by a group of autoencoders is used to characterize the device. That is to say, the device is represented as an ordered set of RMSEs output by these autoencoders.
[0095] The device characterization method of the present application only needs to collect the basic information elements of the device that do not endanger user privacy to well characterize the device, has a strong discrimination ability for distinguishing normal devices and fraudulent devices, provides an effective guarantee for device identity reset fraud detection, and enables effective device identity reset fraud detection. The device characterization method of the present application can achieve better performance in supervised and unsupervised tasks.
[0096] Figure 3 is a process example diagram of the embodiments of the present application for realizing device characterization. As Figure 3 shown, first, element pairs are selected, such as Pair(Brand, IP) in Figure 3 . In this embodiment, Brand includes, for example: brand1, brand2, brand3... brandN; then, the first element instances in the device are respectively represented as the frequency distributions of their corresponding second elements, such as the device information element characterization in Figure 3 ; then, the autoencoder is used to perform feature encoding on the frequency distribution; finally, the ordered set of RMSEs output by the autoencoder is used as the final characterization of the device.
[0097] The present application also provides a computer-readable storage medium storing computer-executable instructions for executing the device characterization method of any one of the above.
[0098] The present application further provides a device for implementing device characterization, including a memory and a processor. Among them, the memory stores the following instructions executable by the processor: steps for executing the device characterization method described in any one of the above.
[0099] The present application also provides a device for implementing device characterization, as Figure 4 shown, including: a construction module, a replacement module, and a characterization module; among them,
[0100] The construction module is used to construct a frequency distribution according to the basic information elements of the device;
[0101] The replacement module is used to represent the instance of the basic information element of the device according to the constructed frequency distribution;
[0102] The characterization module is used to input the instance of the basic information element of the device represented by the frequency distribution into an autoencoder, and use the vector output by the autoencoder to characterize the device.
[0103] In an exemplary instance, the construction module is specifically used for:
[0104] Select one or more element pairs {Pair(tk,th)} from the basic information elements DeviceInf(Di) of the device;
[0105] Take the element with a low Gain_Ratio(K|D) below the set low threshold as the first element tk in the element pair {Pair(tk,th)}, and take the element with a high Pair_gain_ratio(H|K) above the set high threshold as the second element th in the element pair {Pair(tk,th)};
[0106] Based on the law of large numbers of Bernoulli, construct a frequency distribution for any basic information element instance ki in the basic information element DeviceInf(Di) in the element pair.
[0107] In an exemplary instance, the characterization module is an autoencoder.
[0108] The present application also provides a method for detecting device identity reset fraud, including:
[0109] Input the device information characterized by the frequency distribution into a classifier;
[0110] The output of the classifier indicates whether the device is a device with device identity reset fraud;
[0111] Among them, the classifier is trained according to the device information samples characterized by the frequency distribution.
[0112] Among them, the device information characterized by the frequency distribution is obtained according to the device characterization method described in any one of the embodiments of the present application.
[0113] In an exemplary example, the classifier can be obtained by giving some black-and-white samples and using the device information characterized by the frequency distribution of the black-and-white samples as input for classification training. The specific implementation is not used to limit the protection scope of the present application and will not be elaborated here.
[0114] The device identification reset fraud detection method of the present application does not rely on the relationship established by the IP and BSSID to detect devices, and realizes the effective detection of device ID reset fraud that cannot be detected due to frequent replacement of the IP and BSSID.
[0115] The present application also provides a computer-readable storage medium storing computer-executable instructions for executing the device identification reset fraud detection method described in any one of the above.
[0116] The present application further provides a device for implementing device identification reset fraud detection, including a memory and a processor. Among them, the memory stores the following instructions executable by the processor: steps for executing the device identification reset fraud detection method described in any one of the above.
[0117] Although the disclosed embodiments of the present application are as above, the content described is only an embodiment adopted for the convenience of understanding the present application and is not used to limit the present application. Any person skilled in the art within the scope of the present application can make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed by the present application. However, the patent protection scope of the present application shall still be subject to the scope defined by the appended claims.
Claims
1. A device characterization method, comprising: Constructing a frequency distribution based on basic information elements of a device, including: selecting one or more element pairs from the basic information elements of the device; wherein, most data subsets segmented by the first element instance corresponding to the first element in the element pair have sufficient samples; under any first element instance, the second element instances corresponding to the second element are diverse; based on the law of large numbers of Bernoulli, constructing a frequency distribution for any basic information element instance i in the basic information elements of the element pair, obtaining a distribution set including one or more frequency distributions; Representing the basic information element instances of the device according to the constructed frequency distribution; Inputting the basic information element instances of the device represented by the frequency distribution into an autoencoder, and using the vector output by the autoencoder to characterize the device.
2. The device characterization method according to claim 1, wherein Taking the element with the information entropy of the first element lower than the first threshold as the first element in the element pair; For a given first element, taking the element with the difference higher than the set second threshold as the second element in the element pair, wherein the difference is: the difference between the information gain rate of the information gain of the second element with respect to the first element and the information gain rate of the information gain of the first element with respect to the second element.
3. The device characterization method according to claim 1, wherein The representing the basic information element instances of the device according to the constructed frequency distribution includes: Replacing the basic information elements of the device with the distribution set.
4. The device characterization method according to claim 1, wherein, The vector output by the autoencoder is the root mean square error.
5. A computer-readable storage medium storing computer-executable instructions for executing the device characterization method according to any one of claims 1 to 4.
6. A device for implementing device characterization, comprising a memory and a processor, wherein, Instructions executable by a processor are stored in a memory: for performing the steps of the device characterization method according to any one of claims 1 to 4.
7. An apparatus for implementing device characterization, comprising: A construction module, a replacement module, and a characterization module; wherein The construction module is configured to construct a frequency distribution based on basic information elements of a device, including: selecting one or more element pairs from the basic information elements of the device; wherein, most data subsets segmented by the first element instance corresponding to the first element in the element pair have sufficient samples; under any first element instance, the second element instances corresponding to the second element are diverse; based on the law of large numbers of Bernoulli, constructing a frequency distribution for any basic information element instance i in the basic information elements of the element pair, obtaining a distribution set including one or more frequency distributions; The replacement module is configured to represent the basic information element instances of the device according to the constructed frequency distribution; The characterization module is configured to input the basic information element instances of the device represented by the frequency distribution into an autoencoder, and use the vector output by the autoencoder to characterize the device.
8. A device identity reset fraud detection method, comprising: Inputting device information characterized by a frequency distribution into a classifier; The output of the classifier indicates whether the device is a device identity reset fraud device; wherein the classifier is trained according to device information samples characterized by a frequency distribution; Among them, device information of the frequency distribution representation is obtained by the device characterization method according to any one of claims 1 to 4.
9. A computer-readable storage medium storing computer-executable instructions for performing the device identification reset fraud detection method according to claim 8.
10. A device for implementing fraud detection of device identifier reset, including a memory and a processor, wherein, Instructions executable by a processor are stored in a memory: steps for performing the device identification reset fraud detection method according to claim 8.
Citation Information
Patent Citations
Group membership data visualization method and system
CN108280644A
Method and system for product recognition
US20180081877A1