Account detection method and device, computer device, and storage medium
By acquiring data on the relationship between account registration location and time, the target size is adaptively determined, overcoming the limitations of existing account detection technologies, enabling security detection even without data, and improving the accuracy and efficiency of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-03-02
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, account detection methods cannot effectively perform security checks when no transaction data or chat information has been generated, which has limitations.
By acquiring data on the relationship between the target account and other accounts in terms of registration location, neighborhood size, and probability density, the target size is adaptively determined, and the security of the account is predicted based on the clustering characteristics of registration location and time.
It enables account security detection even without transaction data or chat information, avoiding detection limitations and improving detection accuracy and efficiency.
Smart Images

Figure CN116738385B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an account detection method, apparatus, computer device, and storage medium. Background Technology
[0002] With the development of computer technology, more and more people are registering accounts to transfer resources or exchange messages online. To ensure account security, security checks are usually performed to determine whether an account is secure.
[0003] In related technologies, account security needs to be assessed based on transaction data and chat information generated by the account after registration. However, without transaction data and chat information generated by the account, account security cannot be assessed, which limits the effectiveness of this method. Summary of the Invention
[0004] This application provides an account detection method, apparatus, computer device, and storage medium, avoiding the limitations of account detection. The technical solution is as follows:
[0005] On the one hand, an account detection method is provided, the method comprising:
[0006] Obtain first relational data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account. The neighborhood size is used to determine a range centered on the registration location of the target account and with the neighborhood size as the radius. The first probability density of the target account represents the probability that the target account belongs to a non-safe type, predicted based on the type of each account registered within the range.
[0007] Each time a neighborhood size is determined, based on the first relationship data, the coordinates of the registration locations of the multiple first accounts, and the currently determined neighborhood size, a first probability density is determined for each first account, and a first error is determined based on the first probability density of each first account. The multiple first accounts are accounts belonging to the non-safe type, and the first error represents the accuracy of predicting the type to which the multiple first accounts belong.
[0008] From the determined neighborhood sizes, the neighborhood size with the smallest first error is determined as the target size;
[0009] Based on the first relationship data, the coordinates of the registration locations of the plurality of first accounts, and the target size, second relationship data is determined, wherein the second relationship data represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account;
[0010] Based on the coordinates of the registration location of the second account to be predicted and the second relationship data, the first probability density of the second account is determined.
[0011] On the other hand, an account detection device is provided, the device comprising:
[0012] The relationship acquisition module is used to acquire first relationship data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account. The neighborhood size is used to determine the range centered on the registration location of the target account and with the neighborhood size as the radius. The first probability density of the target account represents the probability that the target account belongs to a non-safe type based on the type of each account registered within the range.
[0013] The first error determination module is used to determine a neighborhood size each time, and based on the first relationship data, the coordinates of the registration locations of the multiple first accounts and the currently determined neighborhood size, determine the first probability density of each first account, and determine the first error based on the first probability density of each first account, wherein the multiple first accounts are accounts belonging to the non-security type, and the first error represents the accuracy of predicting the type to which the multiple first accounts belong.
[0014] The target size determination module is used to determine the neighborhood size with the smallest first error from among the determined neighborhood sizes as the target size;
[0015] The relationship determination module is used to determine second relationship data based on the first relationship data, the coordinates of the registration locations of the plurality of first accounts, and the target size. The second relationship data represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account.
[0016] The account detection module is used to determine the first probability density of the second account based on the coordinates of the registration location of the second account to be predicted and the second relationship data.
[0017] In one possible implementation, the first error determination module is configured to:
[0018] Determine the first neighborhood size, and based on the first relationship data, the coordinates of the registration locations of the multiple first accounts, and the first neighborhood size, determine the first probability density of each first account, and calculate the root mean square error based on the first probability density of the multiple first accounts to obtain the first error corresponding to the first neighborhood size;
[0019] A second neighborhood size is determined. Based on the first relationship data, the coordinates of the registration locations of the multiple first accounts, and the second neighborhood size, a first probability density is determined for each first account. The root mean square error is calculated based on the first probability density of the multiple first accounts to obtain the first error corresponding to the second neighborhood size. This process continues until a minimum first error is found among the first errors corresponding to the multiple neighborhood sizes.
[0020] In another possible implementation, the first error determination module is used to adjust the first neighborhood size based on the adjustment coefficient to obtain the second neighborhood size.
[0021] In another possible implementation, the registration location information for any account includes primary location information and secondary location information;
[0022] The primary location information includes primary longitude coordinates and primary latitude coordinates, and the primary location information refers to the lowest-level administrative region to which the registered location belongs; the secondary location information includes secondary longitude coordinates and secondary latitude coordinates, and the secondary location information refers to the next higher-level administrative region of the administrative region indicated by the primary location information; the device further includes:
[0023] A coordinate determination module is used to determine the primary longitude coordinates and the secondary latitude coordinates as the coordinates of the registered location; or,
[0024] The coordinate determination module is further configured to determine the primary latitude coordinates and the secondary longitude coordinates as the coordinates of the registered location.
[0025] In another possible implementation, the device further includes:
[0026] The account partitioning module is used to divide the registration data of multiple candidate accounts into multiple datasets. Each dataset includes the registration data of at least one candidate account, and the registration data includes the registration location. The number of registration data in each dataset is no greater than a first reference number, or the size of the range formed by the registration locations in each dataset is no greater than a reference size.
[0027] The account segmentation module is also used to determine the multiple candidate accounts corresponding to multiple registration data in any dataset as the multiple first accounts.
[0028] In another possible implementation, the account partitioning module is used for:
[0029] The registration data of the multiple candidate accounts is divided into a target number of first datasets;
[0030] For each first dataset, if the number of registered data in the first dataset is greater than the first reference number and the size of the range formed by the registered positions in the first dataset is greater than the reference size, the first dataset is divided into the target number of second datasets until the number of registered data in each of the currently divided datasets is no greater than the first reference number, or the size of the range formed by the registered positions in each dataset is no greater than the reference size.
[0031] In another possible implementation, the registration data further includes the registration time, and the device further includes:
[0032] The relationship acquisition module is used to acquire third relationship data, which represents the relationship between the registration time and duration of the target account and at least one other account and the second probability density of the target account. The duration is used to determine a time period centered on the registration time of the target account and containing twice the duration. The second probability density of the target account represents the probability that the target account belongs to the non-safe type, predicted based on the type of each account registered within the time period.
[0033] The second error determination module is used to determine a duration each time, and based on the third relationship data, the registration time of the multiple first accounts and the currently determined duration, determine the second probability density of each first account, and determine the second error based on the second probability density of each first account. The second error represents the accuracy of predicting the type of the multiple first accounts.
[0034] The target duration determination module is used to determine the duration with the smallest second error from among the determined multiple durations as the target duration;
[0035] The relationship determination module is used to determine fourth relationship data based on the third relationship data, the registration time of the multiple first accounts, and the target duration, wherein the fourth relationship data represents the relationship between the registration time of the target account and the second probability density of the target account;
[0036] The device further includes:
[0037] The relationship determination module is used to determine fifth relationship data based on the second relationship data and the fourth relationship data. The fifth relationship data represents the relationship between the coordinates of the registration location of the target account, the registration time, and the probability density of the target account. The probability density of the target account is the product of the first probability density and the second probability density of the target account. The probability density of the target account represents the probability predicted based on the type of each account registered within the range and time period that the target account belongs to the non-safe type.
[0038] The account detection module is also used to determine the probability density of the second account based on the coordinates of the registration location of the second account, the registration time, and the fifth relationship data.
[0039] In another possible implementation, the second error determination module is used for:
[0040] A first duration is determined. Based on the third relationship data, the coordinates of the registration locations of the multiple first accounts, and the first duration, a second probability density of each first account is determined. The root mean square error is calculated based on the second probability density of the multiple first accounts to obtain the second error corresponding to the first duration.
[0041] A second duration is determined. Based on the third relationship data, the coordinates of the registration locations of the multiple first accounts, and the second duration, a second probability density is determined for each first account. The root mean square error is calculated based on the second probability density of the multiple first accounts to obtain the second error corresponding to the second duration. This process continues until a minimum second error is found among the multiple durations.
[0042] In another possible implementation, the device further includes:
[0043] The verification module is used to determine the density probability of the third account based on the coordinates of the registration location, the registration time, and the fifth relationship data, wherein the type of the third account is either the safe type or the non-safe type.
[0044] The verification module is further configured to update the target size and the target duration based on the coordinates of the registration location, registration time, first relationship data, and third relationship data of the plurality of first accounts when the type of the probability density representation of the third account is inconsistent with the type to which the third account belongs.
[0045] In another possible implementation, the device further includes:
[0046] The verification module is used to determine the area of a first region and a first number, wherein the area of the first region is the area of the region where the coordinates of the registration locations of the plurality of first accounts are located, and the first number is the number of the plurality of first accounts.
[0047] The inspection module is also used to determine the area of the second region and the number of the second region, where the area of the second region is the area of the reference region, and the number of the second region is the number of the multiple fourth accounts, and the area where the coordinates of the registration locations of the multiple fourth accounts are located is the reference region.
[0048] The inspection module is further configured to determine a first ratio between the first number and the second number, and a second ratio between the area of the first region and the area of the second region;
[0049] The verification module is further configured to determine a prediction accuracy parameter based on the ratio between the first ratio and the second ratio, wherein the prediction accuracy parameter represents the accuracy of the probability density predicted based on the second relationship data.
[0050] In another possible implementation, the device further includes:
[0051] The account selection module is used to select the plurality of first accounts that meet the target conditions from a plurality of candidate accounts. The target conditions include that the candidate accounts have a corresponding registration location and that the registration location meets the target location format.
[0052] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to perform the operations performed by the account detection method as described above.
[0053] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored therein, the at least one computer program being loaded and executed by a processor to perform the operations performed by the account detection method as described above.
[0054] On the other hand, a computer program product is provided, including a computer program that, when executed by a processor, performs the operations performed by the account detection method described above.
[0055] The technical solution provided in this application embodiment addresses the issue that accounts belonging to the non-secure type tend to cluster in their registration locations. Therefore, by utilizing the registration location of a first account already identified as belonging to the non-secure type, and based on first relationship data, a target size is determined. This target size describes the degree of clustering in the registration location distribution of multiple first accounts. Then, second relationship data is determined based on the target size and the first relationship data. By processing the registration location of a second account based on this second relationship data, the likelihood that the second account belongs to the non-secure type can be determined. This account detection method can perform detection based on the account's registration location without needing to obtain data generated by the account, thus avoiding the limitations of account detection. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a flowchart of an account detection method provided in an embodiment of this application;
[0058] Figure 2 This is a flowchart of another account detection method provided in the embodiments of this application;
[0059] Figure 3 This is a flowchart illustrating a process for partitioning a dataset, as provided in an embodiment of this application.
[0060] Figure 4 This is a schematic diagram illustrating a process of partitioning a dataset according to an embodiment of this application;
[0061] Figure 5 This is a flowchart of another account detection method provided in the embodiments of this application;
[0062] Figure 6 This is a schematic diagram illustrating a fifth relational data provided in an embodiment of this application;
[0063] Figure 7 This is a schematic diagram illustrating a kernel density estimation principle provided in an embodiment of this application;
[0064] Figure 8 This is a flowchart of another account detection method provided in the embodiments of this application;
[0065] Figure 9 This is a schematic diagram of a prediction region provided in an embodiment of this application;
[0066] Figure 10This is a schematic diagram of a longitude and latitude distribution provided in an embodiment of this application;
[0067] Figure 11 This is a flowchart of another account detection method provided in the embodiments of this application;
[0068] Figure 12 This is a schematic diagram of the structure of an account detection device provided in an embodiment of this application;
[0069] Figure 13 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0070] Figure 14 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0071] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0072] It is understood that the terms "first," "second," etc., used in this application may be used to describe various concepts herein, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, the first arrangement order may be referred to as the second arrangement order, and the second arrangement order may be referred to as the first arrangement order.
[0073] As used in this application, the terms "at least one," "multiple," "each," and "any" have the following meanings: at least one includes one, two, or more; multiple includes two or more; each refers to each of the corresponding multiple; and any refers to any one of the multiple. For example, multiple accounts include three accounts, where each account refers to every one of the three accounts, and any refers to any one of the three accounts, which could be the first, the second, or the third.
[0074] To facilitate understanding of the embodiments of this application, the keywords involved in the embodiments of this application will be explained first:
[0075] Kernel Density Estimation (KDE) is a nonparametric testing method suitable for estimating unknown probability density functions. It features good applicability, high stability, and strong continuous optimization capabilities. Kernel density estimation has wide applications in fields such as power, finance, healthcare, crime prediction, and urban planning.
[0076] Octree Algorithm: The octree algorithm is a classic spatial partitioning and indexing technique that is widely used in image decomposition and spatial indexing. This algorithm recursively decomposes the data space into octet-like blocks of varying densities. In denser areas, there are more and smaller blocks; conversely, in sparser areas, there are fewer and larger blocks.
[0077] The execution subject of this application embodiment is a computer device, which is a terminal or a server. Optionally, the terminal is a computer, mobile phone, tablet computer, vehicle terminal, or other terminal. The server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0078] Figure 1 This is a flowchart illustrating an account detection method provided in an embodiment of this application. The execution subject of this embodiment is a computer device. See also... Figure 1 The method includes the following steps:
[0079] 101. Obtain first relational data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account. The neighborhood size is used to determine the range centered on the registration location of the target account and with the neighborhood size as the radius. The first probability density of the target account represents the probability that the target account belongs to the non-safe type based on the type of each account registered within this range.
[0080] In this context, the target account and at least one other account belong to the non-secure type; the neighborhood size is used to describe the degree of clustering of the distribution of multiple registration locations; the first probability density of the target account varies with the coordinates of the registration locations of the target account and at least one other account, as well as the neighborhood size.
[0081] 102. Each time a neighborhood size is determined, based on the first relation data, the coordinates of the registration locations of multiple first accounts, and the currently determined neighborhood size, the first probability density of each first account is determined, and based on the first probability density of each first account, the first error is determined. Multiple first accounts are accounts belonging to the non-safe type, and the first error represents the accuracy of predicting the type to which multiple first accounts belong.
[0082] 103. From the determined multiple neighborhood sizes, the neighborhood size with the smallest first error is determined as the target size.
[0083] Since the size of the neighborhood significantly impacts the accuracy of subsequent account detection, a suitable neighborhood size needs to be determined to ensure accurate account detection. In this embodiment, given a determined neighborhood size, a first probability density for each first account is obtained based on first relationship data and the coordinates of the registration locations of multiple first accounts. Then, a first error is calculated from the first probability densities of multiple first accounts. This process of calculating the first error is repeated multiple times. The smaller the first error, the more accurate the corresponding neighborhood size. The neighborhood size with the smallest first error is determined as the target size, thereby achieving an adaptive process for determining the target size.
[0084] 104. Based on the first relational data, the coordinates of the registration locations of multiple first accounts, and the target size, determine the second relational data, which represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account.
[0085] 105. Based on the coordinates of the registration location of the second account to be predicted and the second relationship data, determine the first probability density of the second account.
[0086] Since the target size is determined based on the coordinates of the registration locations of multiple first accounts, and this target size is used to describe the degree of clustering of the registration location distribution of multiple first accounts, that is, the target size is closely related to the registration locations of multiple first accounts, it is necessary to determine the second relationship data based on the first relationship data, the coordinates of the registration locations of multiple first accounts, and the target size, so that the second relationship data can predict whether the second account belongs to the non-safe type according to the degree of clustering of the registration location distribution of multiple first accounts and the coordinates of the registration location of the second account to be predicted.
[0087] The method provided in this application embodiment, since accounts belonging to the non-secure type tend to cluster in their registration locations, utilizes the registration location of the first account already determined to be of the non-secure type. Based on the first relationship data, a target size is determined, which describes the degree of clustering of the registration location distribution of multiple first accounts. Then, based on the target size and the first relationship data, second relationship data can be determined. By processing the registration location of the second account based on the second relationship data, the probability that the second account belongs to the non-secure type can be determined. This account detection method can perform detection based on the registration location of the account without obtaining the data generated by the account, thus avoiding the limitations of account detection.
[0088] In the account detection method provided in this application embodiment, the first method detects the account based on the registration location; the second method detects the account based on both the registration location and registration time. The following describes... Figure 2 The first embodiment will be described in detail below.
[0089] Figure 2 This is a flowchart illustrating an account detection method provided in an embodiment of this application. The execution subject of this embodiment is a computer device. See also... Figure 2 The method includes the following steps:
[0090] 201. Obtain registration data for multiple primary accounts, which are non-secure accounts. The registration data includes the registration location.
[0091] The account type is divided into non-secure and secure types. Non-secure type means that the account has risks, that is, based on the account having engaged in risky behavior. Secure type means that the account has no risks.
[0092] Due to certain regional factors, accounts registered in these regions are more likely to be of the insecure type, resulting in a spatial clustering of the registration locations of insecure accounts. That is, the registration locations of multiple insecure accounts are relatively close to each other. Therefore, in this embodiment of the application, by obtaining multiple first accounts of the insecure type, the account to be predicted is predicted based on the distribution characteristics of the registration locations of these multiple first accounts and the registration location of the account to be predicted.
[0093] Optionally, the account registration location can be in the form of coordinates, including longitude and latitude coordinates; or the registration location can be in the form of an administrative region of at least one level, including country, province, city, county or district, and a detailed address, where the detailed address refers to the address down to a specific neighborhood or village. The detailed address represents the lowest level of administrative region. When the registration location is in the form of an administrative region of at least one level, the coordinates corresponding to that registration location need to be determined based on that administrative region. For example, if each administrative region has corresponding coordinates, then the coordinates corresponding to the detailed address will be used as the coordinates of the registration location.
[0094] In one possible implementation, since the account's registration location is needed in subsequent processing, but the corresponding registration location may not have been filled in during account registration, even if the account is a non-secure account, subsequent processing based on the registration location is impossible. Therefore, to ensure that multiple primary accounts have corresponding registration locations, multiple candidate accounts (non-secure accounts) need to be obtained first. From these candidate accounts, multiple primary accounts that meet the target conditions are selected. These target conditions include that the candidate accounts have corresponding registration locations, and that the registration locations meet the target location format. The target location format refers to a format containing multiple levels of administrative regions, such as "Country-Province-City-County or District-Detailed Address," or "County or District-Detailed Address," or other formats that can identify the detailed address in the location information.
[0095] In one possible implementation, the account has corresponding registration location information, which includes primary and secondary location information. The primary location information includes primary longitude and primary latitude coordinates, and the secondary location information includes secondary longitude and secondary latitude coordinates. The primary location information refers to the lowest-level administrative region to which the registration location belongs, and the secondary location information refers to the next higher-level administrative region. For example, if an account's registration location information is Province A, City B, County C, then the primary location information is County C, and the secondary location information is City B.
[0096] Considering that the coordinates of the lowest-level administrative region may be inaccurate, when determining the coordinates of the registration location, not only the coordinates corresponding to the lowest-level administrative region are considered, but also the coordinates corresponding to the administrative region at the next higher level. That is, the primary longitude coordinates and the secondary latitude coordinates are determined as the coordinates of the registration location, or the primary latitude coordinates and the secondary longitude coordinates are determined as the coordinates of the registration location.
[0097] In another possible implementation, the registration data of multiple candidate accounts is divided into multiple datasets. Each dataset includes the registration data of at least one candidate account. The number of registration data in each dataset is no greater than a first reference number, or the size of the range formed by the registration positions in each dataset is no greater than a reference size. Multiple candidate accounts corresponding to multiple registration data in any dataset are then determined as multiple first accounts. The first reference number and reference size are pre-set. Since the distribution characteristics of the registration positions of multiple accounts are considered, the registration positions of multiple accounts will affect each other during subsequent processing. If the number of accounts is too large, it will lead to low processing efficiency. Therefore, in this embodiment, by dividing the registration data of multiple candidate accounts into multiple datasets and determining multiple candidate accounts corresponding to multiple registration data in one dataset as multiple first accounts, the processing efficiency can be improved by subsequently processing the registration data of multiple first accounts.
[0098] Optionally, the registration data of multiple candidate accounts is divided into a target number of first datasets. For each first dataset, if the number of registered data in the first dataset is greater than a first reference number and the size of the range formed by the registration positions in the first dataset is greater than the reference size, the first dataset is divided into a target number of second datasets until the number of registered data in each of the currently divided datasets is no greater than the first reference number, or the size of the range formed by the registration positions in each dataset is no greater than the reference size. For any first dataset, if the number of registered data in the first dataset is no greater than the first reference number, or if the number of registered data in the first dataset is greater than the first reference number, but the size of the range formed by the registration positions in the first dataset is no greater than the reference size, the first dataset is taken as the divided dataset and no further division is performed.
[0099] For example, the octree algorithm can be used to partition the registration data of multiple candidate accounts, see [link to relevant documentation]. Figure 3The flowchart shown has a target number of 8. First, the registration data of multiple candidate accounts is divided into 8 first datasets. Then, the number of registrations and the size of the range formed by the registration locations in each first dataset are determined. For each first dataset, it is determined whether the number of registrations in the first dataset is greater than a first reference number. If the number of registrations in the first dataset is not greater than the first reference number, the first dataset is considered as the second dataset and no further division is performed. If the number of registrations in the first dataset is greater than the first reference number, it is determined whether the size of the range formed by the registration locations in the first dataset is greater than the reference size. If the size of the range formed by the registration locations in the first dataset is not greater than the reference size, the first dataset is considered as the second dataset and no further division is performed. If the size of the range formed by the registration locations in the first dataset is greater than the reference size, the first dataset is further divided into 8 second datasets. Then, for each second dataset, the process is similar to that of the first dataset, until the number of registrations in each second dataset is not greater than the first reference number, or the size of the range formed by the registration locations in each second dataset is not greater than the reference size. Optionally, it is also necessary to determine whether each of the partitioned datasets has performed the steps of determining whether the number of registered data in the dataset is greater than the first reference number, and whether the steps of determining whether the size of the range formed by the registered positions in the dataset is greater than the reference size. If there are still datasets that have not performed at least one of the above steps, the dataset is judged based on the above steps. If at least one of the above steps has been performed in each dataset, the partitioning process ends.
[0100] For example, see Figure 4 The diagram illustrates the partitioning method of the octree algorithm using a cube, treating the registration data of multiple candidate accounts as a whole. Figure 4 The cube 401 in the middle, and then the registration data of multiple candidate accounts are divided into 8 first datasets, that is... Figure 4 The cube in the image is divided into 8 parts, each part being a smaller cube 402. Then, the registration data from the two first datasets are further divided into 8 second datasets, i.e. Figure 4 The cube 402 in the dataset is divided into eight parts, and then the registration data in one of the second datasets is further divided into eight third datasets, that is... Figure 4 The cube 403 is divided into eight parts. Figure 4 The diagram represents dividing a circle into 8 circles.
[0101] It should be noted that the embodiments of this application are only used as an example of registration data including registration location. In another embodiment, the registration data may also include registration time, such as the registration time being a certain year, month, day and hour. The embodiments of this application do not limit other data included in the registration data.
[0102] 202. Obtain first relation data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account.
[0103] In this context, the target account is the account to be predicted, and at least one other account is an account belonging to the non-safe type. The neighborhood size is used to determine the range centered on the registration location of the target account and with the neighborhood size as the radius. The first probability density of the target account represents the probability that the target account belongs to the non-safe type based on the type of each account registered within this range.
[0104] In one possible implementation, the coordinates of the registration location include longitude and latitude coordinates, and the neighborhood size includes longitude and latitude dimensions. Then, the first relational data represents the relationship between the longitude and latitude coordinates, longitude dimensions, latitude dimensions, and the first probability density of the registration locations of the target account and at least one other account.
[0105] 203. Determine the first neighborhood size. Based on the first relation data, the coordinates of the registration locations of multiple first accounts, and the first neighborhood size, determine the first probability density of each first account. Calculate the root mean square error based on the first probability density of multiple first accounts to obtain the first error corresponding to the first neighborhood size.
[0106] For each of the multiple first accounts, the coordinates of its registration location are used as the coordinates of the target account's registration location in the first relation data. The coordinates of the registration locations of other first accounts besides the target account are used as the coordinates of the registration locations of at least one other account in the first relation data. The first neighborhood size is used as the neighborhood size in the first relation data. Based on this first relation data, the first probability density of the first account is obtained, thus obtaining the first probability density of each first account. The first neighborhood size is a preset size.
[0107] In one possible implementation, the first neighborhood size includes a first longitude size and a first latitude size, and the first relational data is represented by the following formula. For any first account, the first probability density of the first account is determined using the following formula:
[0108]
[0109] in, This represents the first probability density of the first account. This indicates the longitude coordinates of the registered location of the first account. This represents the latitude and longitude coordinates of the registration location of the first account. This refers to the first account among multiple first accounts. The first account. Indicates the first The longitude coordinates of the first account Indicates the first The latitude and longitude coordinates of the first account Indicates longitude dimension. Indicates longitude dimension. Indicates the adjustment factor. This indicates the number of multiple primary accounts. Represents the kernel function.
[0110] in, This can be any one of a polynomial kernel, a quadratic kernel, or a Gaussian kernel. A polynomial kernel is suitable for registration locations accurate to the city / county level; a quadratic kernel is suitable for registration locations that are coordinates; and a Gaussian kernel is suitable for registration locations with significant noise. The polynomial, quadratic, and Gaussian kernels are as follows:
[0111] Polynomial kernel function: ;
[0112] Quadratic kernel function: ;
[0113] Gaussian kernel function: ;
[0114] In the first relational data mentioned above, z,m in the polynomial kernel function is p is the first reference parameter, and m in the quadratic kernel function and Gaussian kernel function is... .
[0115] Based on the aforementioned first relational data, the first probability density of each first account is obtained, and then the first error is determined using the following formula:
[0116]
[0117] in, Indicates the first error. This refers to the first account among multiple first accounts. The first probability density of the first account This indicates the number of multiple primary accounts.
[0118] 204. Determine the second neighborhood size. Based on the first relation data, the coordinates of the registration locations of multiple first accounts, and the second neighborhood size, determine the first probability density of each first account. Calculate the root mean square error based on the first probability density of multiple first accounts to obtain the first error corresponding to the second neighborhood size. Continue this process until the first error corresponding to the multiple neighborhood sizes is found to be the smallest.
[0119] The implementation method of step 204 is the same as that of step 203 described above. The difference is that the second neighborhood size is determined based on the first neighborhood size.
[0120] In one possible implementation, the size of the first neighborhood is adjusted based on an adjustment coefficient to obtain the size of the second neighborhood. For example, the second neighborhood size is obtained by multiplying the adjustment coefficient by the first neighborhood size. Optionally, the adjustment coefficient is a preset value, or the adjustment coefficient also changes with the first probability density, for example, by determining the adjustment coefficient using the following formula:
[0121]
[0122] in, Indicates the adjustment factor. This represents the first probability density of the first account. This is the second reference parameter.
[0123] It should be noted that the embodiments of this application are only illustrated by taking the above two processes of obtaining the first error as an example. In another embodiment, the above steps 203 or 204 are repeated multiple times to obtain the first error corresponding to multiple neighborhood sizes.
[0124] 205. From the determined neighborhood sizes, the neighborhood size with the smallest first error is determined as the target size.
[0125] Through steps 203-204 above, the first error corresponding to various neighborhood sizes can be obtained. The smaller the first error, the more accurate the corresponding neighborhood size is. That is, the corresponding neighborhood size can accurately describe the degree of clustering of the registration locations of multiple first accounts. The neighborhood size with the smallest first error is determined as the target size, thereby obtaining a suitable target size. This facilitates the prediction of whether any account belongs to the non-safe type based on the degree of clustering of the registration locations of multiple first accounts.
[0126] The larger the target size, the more concentrated the registration locations of multiple first accounts are; the smaller the target size, the more dispersed the registration locations of multiple first accounts are.
[0127] 206. Based on the first relational data, the coordinates of the registration locations of multiple first accounts, and the target size, determine the second relational data, which represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account.
[0128] In this embodiment of the application, after the target size is determined through the above implementation method, the target size can be used as the neighborhood size in the first relation data, and the coordinates of the registration positions of multiple first accounts can be used as the coordinates of the registration positions of at least one account to obtain the second relation data. The independent variable in the second relation data is the coordinates of the registration position of the target account, so that the second relation data can represent the relationship between the coordinates of the registration position of the target account and the first probability density of the target account.
[0129] 207. Based on the coordinates of the registration location of the second account to be predicted and the second relationship data, determine the first probability density of the second account.
[0130] In this context, the second account is any account different from the multiple first accounts. Using the coordinates of the second account's registration location as the coordinates of the target account's registration location in the second relational data, a first probability density for the second account can be obtained. This first probability density represents the probability that the target account, predicted based on the type of each first account registered within the target range, belongs to a non-safe type. The target range is defined by the second account's registration location as the center and the target size as the radius. A higher first probability density indicates a greater probability that the second account belongs to a non-safe type, and a lower first probability density indicates a lower probability that the second account belongs to a non-safe type.
[0131] The method provided in this application embodiment, since accounts belonging to the non-secure type tend to cluster in their registration locations, utilizes the registration location of the first account already determined to be of the non-secure type. Based on the first relationship data, a target size is determined, which describes the degree of clustering of the registration location distribution of multiple first accounts. Then, based on the target size and the first relationship data, second relationship data can be determined. By processing the registration location of the second account based on the second relationship data, the probability that the second account belongs to the non-secure type can be determined. This account detection method can perform detection based on the registration location of the account without obtaining the data generated by the account, thus avoiding the limitations of account detection.
[0132] The following is through Figure 5 The second method will be described in detail first in the following embodiments.
[0133] Figure 5 This is a flowchart illustrating an account detection method provided in an embodiment of this application. The execution subject of this embodiment is a computer device. See also... Figure 5 The method includes the following steps:
[0134] 501. Obtain registration data for multiple primary accounts. These primary accounts are non-secure accounts. The registration data includes the registration location and registration time.
[0135] 502. Obtain first relation data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account.
[0136] 503. Each time a neighborhood size is determined, based on the first relationship data, the coordinates of the registration locations of the multiple first accounts, and the currently determined neighborhood size, a first probability density is determined for each first account, and a first error is determined based on the first probability density of each first account, wherein the multiple first accounts are accounts belonging to the non-security type, and the first error represents the accuracy of predicting the type to which the multiple first accounts belong.
[0137] 504. From the determined multiple neighborhood sizes, the neighborhood size with the smallest first error is determined as the target size.
[0138] 505. Based on the first relational data, the coordinates of the registration locations of multiple first accounts, and the target size, determine the second relational data, which represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account.
[0139] Steps 501-505 are the same as steps 201-206 above, and will not be repeated here.
[0140] 506. Obtain third relation data, which represents the relationship between the registration time and duration of the target account and at least one other account and the second probability density of the target account.
[0141] The duration is used to determine a time period centered on the registration time of the target account and including twice the duration. The second probability density of the target account represents the probability that the target account belongs to the non-safe type based on the type of each account registered within the time period.
[0142] 507. Each time a duration is determined, based on the third relation data, the registration time of multiple first accounts, and the currently determined duration, the second probability density of each first account is determined, and based on the second probability density of each first account, the second error is determined. The second error represents the accuracy of predicting the type of multiple first accounts.
[0143] Since the duration of the event significantly impacts the accuracy of subsequent account detection, a suitable duration needs to be determined to ensure accuracy. In this embodiment, given a fixed duration, a second probability density for each first account is obtained based on third-party relationship data and the registration times of multiple first accounts. Then, a second error is calculated from the second probability densities of multiple first accounts. This process of calculating the second error is repeated multiple times. The smaller the second error, the more accurate the corresponding duration. The duration with the smallest second error is determined as the target duration, thus achieving an adaptive process for determining the target duration.
[0144] In one possible implementation, a first duration is determined. Based on the third relation data, the coordinates of the registration locations of multiple first accounts, and the first duration, a second probability density for each first account is determined. The root mean square error is calculated based on the second probability densities of the multiple first accounts to obtain the second error corresponding to the first duration. Here, the first duration is a preset duration. For each of the multiple first accounts, the registration time of that first account is used as the registration time of the target account in the third relation data. The registration times of other first accounts besides that first account are used as the registration times of at least one other account in the third relation data. The first duration is used as the duration in the third relation data. Based on the third relation data, the second probability density of that first account is obtained, thus obtaining the second probability density for each first account.
[0145] Then, a second duration is determined. This second duration is either another preset duration or a duration determined based on the first duration. Based on the third relationship data, the coordinates of the registration locations of multiple first accounts, and the second duration, the second probability density of each first account is determined. The root mean square error is calculated based on the second probability density of multiple first accounts to obtain the second error corresponding to the second duration. This process continues until the second error corresponding to the multiple durations is found to be the smallest.
[0146] In this embodiment, the method for determining the target duration is similar to the method for determining the target size in the above embodiments, and will not be repeated here.
[0147] 508. From the various determined durations, the duration with the smallest second error is determined as the target duration.
[0148] Through step 507 above, the second error corresponding to various durations can be obtained. The smaller the second error, the more accurate the corresponding duration is. That is, the corresponding duration can accurately describe the degree of clustering of the registration time of multiple first accounts. The duration with the smallest corresponding second error is determined as the target duration, so as to obtain a suitable target duration. This is convenient for predicting whether any account belongs to the non-safe type based on the degree of clustering of the registration time of multiple first accounts.
[0149] The longer the target duration, the more concentrated the registration time of multiple first accounts is; the shorter the target duration, the more dispersed the registration time of multiple first accounts is.
[0150] 509. Based on the third relation data, the registration time and target duration of multiple first accounts, determine the fourth relation data, which represents the relationship between the registration time of the target account and the second probability density of the target account.
[0151] In this embodiment of the application, after the target duration is determined through the above implementation method, the target duration can be used as the duration in the third relation data, and the registration time of multiple third accounts can be used as the registration time of at least one account to obtain the fourth relation data. The independent variable in the fourth relation data is the registration time of the target account, so that the fourth relation data can represent the relationship between the registration time of the target account and the second probability density of the target account.
[0152] It should be noted that the embodiments of this application are only illustrated by taking the execution of steps 502-505 first and then steps 506-509 as an example. In another embodiment, steps 506-509 can be executed first and then steps 502-505 can be executed, or steps 502-505 and steps 506-509 can be executed simultaneously. The embodiments of this application do not restrict the order of execution of the steps.
[0153] 510. Based on the second and fourth relation data, determine the fifth relation data, which represents the relationship between the coordinates of the target account's registration location, the registration time, and the probability density of the target account.
[0154] The probability density of the target account represents the probability that the target account belongs to an unsafe type, based on the type of each account registered within the range and time period.
[0155] In one possible implementation, the second, fourth, and fifth relation data are all relation data determined using a kernel density estimation algorithm, and the fifth relation data is represented by the following formula:
[0156]
[0157] in, This represents the probability density of the target account. This represents the longitude coordinates of the target account's registration location. This represents the latitude and longitude coordinates of the target account's registration location. Indicates the registration time of the target account. This refers to the first account among multiple first accounts. The first account. Indicates the first The longitude coordinates of the first account Indicates the first The latitude and longitude coordinates of the first account Indicates the first The registration time of the first account Indicates the target size. Indicates the target duration. This indicates the number of multiple primary accounts. and Represents the kernel function.
[0158] In another possible implementation, with the longitude and latitude dimensions determined separately, the fifth relation data is represented by the following formula:
[0159]
[0160] in, This represents the probability density of the target account. This represents the longitude coordinates of the target account's registration location. This represents the latitude and longitude coordinates of the target account's registration location. Indicates the registration time of the target account. This refers to the first account among multiple first accounts. The first account. Indicates the first The longitude coordinates of the first account Indicates the first The latitude and longitude coordinates of the first account Indicates the first The registration time of the first account Indicates longitude dimension. Indicates longitude dimension. Indicates the target duration. This indicates the number of multiple primary accounts. Represents the kernel function.
[0161] For example, see Figure 6 The aforementioned fifth relation data can be represented as Figure 6 The diagram shown is as follows. Figure 6 The x-axis represents longitude, the y-axis represents latitude, and the t-axis represents time. Figure 6 Multiple location points in the text represent multiple primary accounts. Figure 6 Curve 1 corresponding to the x-axis represents The curve 2 corresponding to the y-axis represents The curve 3 corresponding to the t-axis represents .
[0162] 511. Based on the coordinates of the registration location, registration time, and fifth relationship data of the second account, determine the probability density of the second account.
[0163] By using the coordinates of the second account's registration location as the coordinates of the target account's registration location in the fifth relation data, and the second account's registration time as the target account's registration time in the fifth relation data, the probability density of the second account can be obtained based on the fifth relation data. A higher probability density indicates a greater likelihood that the second account belongs to an insecure type, while a lower probability density indicates a lower likelihood. Optionally, if the probability density of the second account is greater than a reference probability density, the second account is determined to be insecure; if the probability density of the second account is not greater than the reference probability density, the second account is determined to be secure.
[0164] In one possible implementation, after obtaining the fifth relationship data based on the above-described embodiments, it is necessary to verify the fifth relationship data. If the verification of the fifth relationship data is successful, that is, if it is determined that the type of the account can be accurately predicted based on the fifth relationship data, then the coordinates of the registration location and the registration time of the second account are processed based on the fifth relationship data. If the verification of the fifth relationship data fails, it is necessary to sample the above-described embodiments again to update the target size and target duration.
[0165] In one possible implementation, the density probability of the third account is determined based on the coordinates of its registration location, registration time, and fifth relationship data. The type of the third account is either safe or unsafe. If the type represented by the probability density of the third account is inconsistent with the type to which the third account belongs, the target size and target duration are updated based on the coordinates of the registration locations of multiple first accounts, registration time, first relationship data, and third relationship data.
[0166] Optionally, when examining the fifth relationship data based on multiple third accounts, a chi-square test can be performed on the fifth relationship data. The chi-square test result indicates that the fifth relationship data can reflect randomness characteristics. The chi-square test formula is as follows:
[0167]
[0168] in, This represents the output value of the chi-square test. This represents the frequency of occurrence of the probability density of multiple third accounts within the range of the i-th probability density. Let k represent the frequency of the probability density corresponding to the type to which multiple third accounts actually belong in the i-th probability density range, k represent the number of multiple probability density ranges, and n represent the number of multiple third accounts.
[0169] In another possible implementation, for multiple fourth accounts, the fifth relationship data is validated based on the hit rate metric and the PAI (Prediction Accuracy Index). The hit rate represents the ratio between the number of accounts predicted to be non-safe from the multiple fourth accounts and the actual number of accounts that are non-safe from the multiple fourth accounts; the closer this ratio is to 1, the more accurate the fifth relationship data. However, since the hit rate metric is effective for a defined reference area, its result becomes meaningless when the reference area is too large. Therefore, after determining the hit rate, it is also necessary to determine the PAI metric, namely, to determine the area and number of the first region, where the area is the area of the coordinates of the registration locations of multiple first accounts, and the number is the number of multiple first accounts; to determine the area and number of the second region, where the area is the area of the reference region, and the number is the number of multiple fourth accounts, and the area of the coordinates of the registration locations of multiple fourth accounts is the reference region; to determine the first ratio between the first number and the second number, and the second ratio between the area of the first region and the area of the second region; and based on the ratio between the first and second ratios, to determine the prediction accuracy parameter, which represents the accuracy of the probability density predicted based on the second relational data.
[0170] For example, the formula for calculating the PAI indicator is as follows:
[0171]
[0172] Wherein, PAI represents the prediction accuracy parameter. l Indicates the first number. L Indicates the area of the first region. a Indicates the second number. A This indicates the area of the second region.
[0173] If both the hit rate metric and the PAI metric reach their respective thresholds, the fifth relationship data is confirmed to be accurate and can be used for subsequent account detection.
[0174] It should be noted that the above embodiment detects the second account based on the fifth relationship data. In another embodiment, multiple registered accounts can be processed based on the fifth relationship data to predict the area and time period where accounts that may be registered as non-secure can be identified. This allows for subsequent processing of newly registered accounts within that area and time period, such as further security testing of the newly registered accounts.
[0175] Furthermore, the algorithm used in the embodiments of this application is a kernel density estimation algorithm, the principle of which can be found in [link to relevant documentation]. Figure 7 , Figure 7 The points in the diagram represent the registration data of the first account. The kernel function is the determined fifth relation data. The space covered by the kernel function is the range in which accounts of non-secure types may be registered. The time period covered by the kernel function is the time period in which accounts of non-secure types may be registered.
[0176] The method provided in this application embodiment, since accounts belonging to the non-secure type tend to cluster in terms of registration location and registration time, utilizes the registration location and registration time of the first account already determined to be of the non-secure type. Based on the first relationship data and the third relationship data, a target size and a target duration are determined. The target size describes the degree of clustering of the registration location distribution of multiple first accounts, and the target duration describes the degree of clustering of the registration time distribution of multiple first accounts. Then, based on the target size, target duration, first relationship data, and third relationship data, a fifth relationship data can be determined. By processing the registration location and registration time of the second account based on the fifth relationship data, the probability that the second account belongs to the non-secure type can be determined. This account detection method can detect based on the registration location and registration time of the account without obtaining the data generated by the account, thus avoiding the limitations of account detection.
[0177] Furthermore, in this embodiment, the fifth relationship data is tested using chi-square test, hit rate index, and PAI index to determine its accuracy. If the fifth relationship data is inaccurate, the target size and target duration can be updated again to obtain accurate fifth relationship data, thereby ensuring that the detection results of accounts subsequently detected based on the fifth relationship data are accurate and improving the accuracy of account detection.
[0178] Furthermore, the account detection method provided in this application embodiment can automatically detect whether an account is at risk, without relying on human experience to make judgments, thus reducing manual labor, lowering inspection costs, and improving inspection efficiency.
[0179] In one possible implementation, see [link to relevant documentation]. Figure 8 The flowchart shown illustrates the account detection process of this application embodiment from another perspective, with the execution entity being a computer device.
[0180] 1. Obtain registration data for multiple alternative accounts.
[0181] 2. Preprocess the registration data of multiple candidate accounts to obtain multiple accounts that meet the target conditions.
[0182] Preprocessing refers to selecting accounts that meet the target criteria from multiple candidate accounts.
[0183] 3. Divide the multiple accounts that meet the target conditions into training dataset and test dataset.
[0184] 4. Using the octree algorithm, the registration data of multiple accounts in the training dataset is divided into multiple datasets.
[0185] 5. Using root mean square error, determine the target size and target duration based on the registration data of multiple accounts in any dataset.
[0186] 6. Build a prediction model based on the test dataset, target size, and target duration, and perform a chi-square test on the prediction model, as well as calculate the hit rate and PAI metrics.
[0187] The prediction model is the fifth relation data in the above embodiments.
[0188] 7. Determine whether both the hit rate metric and the PAI metric have reached their corresponding thresholds. If both the hit rate metric and the PAI metric have reached their corresponding thresholds, output the prediction results.
[0189] The prediction results indicate the geographical range and time period in which the prediction model predicts the possible registration of accounts belonging to non-secure types.
[0190] 8. Based on the prediction results, process the prediction results to obtain a spatiotemporal prediction map.
[0191] For example, see Figure 9 The spatiotemporal prediction map shown indicates that the areas marked with circles are the regions where accounts belonging to non-secure types may be registered.
[0192] In one possible implementation, the spatiotemporal prediction map can also be represented as a schematic diagram of longitude and latitude distribution, for example, see [reference needed]. Figure 10 .
[0193] 9. Based on the prediction results, subsequent processing is carried out through the prediction risk processing module.
[0194] In one possible implementation, after obtaining the prediction results, a verification program can be added to the computer device to further verify accounts registered in the region and during the time period using risk control strategies.
[0195] In another possible implementation, see Figure 11 The flowchart shown indicates that the executing entity is a computer device, through... Figure 11 The flowchart shown illustrates the account verification process:
[0196] 1. Obtain the candidate account dataset, which includes the registration data of multiple candidate accounts.
[0197] 2. Preprocess the registration data of multiple candidate accounts to obtain multiple accounts that meet the target conditions.
[0198] Preprocessing refers to selecting accounts that meet the target criteria from multiple candidate accounts.
[0199] 3. Determine the kernel density estimation algorithm to be used, that is, determine the first relation data and the third relation data.
[0200] 4. Select any kernel function from the polynomial kernel function, quadratic kernel function, and Gaussian kernel function as the kernel function used in the kernel density estimation algorithm.
[0201] 5. Improve the kernel density estimation algorithm by using the coordinates of the lowest-level administrative region and the next higher-level administrative region in the registration address information to determine the coordinates of the registration location.
[0202] 6. Use the octree algorithm to divide the registration data of multiple accounts.
[0203] 7. Use root mean square error to determine the target size and target duration.
[0204] 8. Based on the kernel density estimation algorithm determined in step 3, determine the optimized kernel density estimation algorithm based on steps 4-7.
[0205] 9. The optimized kernel density estimation algorithm is validated based on validation metrics.
[0206] 10. Construct a prediction model and apply it, that is, determine the fifth relationship data.
[0207] 11. Determine the prediction results for the account and process the account based on the prediction results.
[0208] It should be noted that, Figure 8 and Figure 11 The implementation method of the account detection process shown is the same as that described above. Figure 5 The implementation methods shown are similar and will not be described in detail here.
[0209] Figure 12 This is a schematic diagram of the structure of an account detection device provided in an embodiment of this application. See also... Figure 12 The device includes:
[0210] The relationship acquisition module 1201 is used to acquire first relationship data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account. The neighborhood size is used to determine the range centered on the registration location of the target account and with the neighborhood size as the radius. The first probability density of the target account represents the probability that the target account belongs to the non-safe type based on the type of each account registered within the range.
[0211] The first error determination module 1202 is used to determine a neighborhood size each time, and based on the first relationship data, the coordinates of the registration locations of multiple first accounts and the currently determined neighborhood size, determine the first probability density of each first account, and determine the first error based on the first probability density of each first account. The multiple first accounts are accounts belonging to the non-safe type, and the first error represents the accuracy of predicting the type of the multiple first accounts.
[0212] The target size determination module 1203 is used to determine the target size from the multiple determined neighborhood sizes, with the corresponding first error being the smallest.
[0213] The relationship determination module 1204 is used to determine the second relationship data based on the first relationship data, the coordinates of the registration locations of multiple first accounts, and the target size. The second relationship data represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account.
[0214] The account detection module 1205 is used to determine the first probability density of the second account based on the coordinates of the registration location of the second account to be predicted and the second relationship data.
[0215] In one possible implementation, the first error determination module 1202 is used for:
[0216] Determine the size of the first neighborhood. Based on the first relation data, the coordinates of the registration locations of multiple first accounts, and the size of the first neighborhood, determine the first probability density of each first account. Calculate the root mean square error based on the first probability density of multiple first accounts to obtain the first error corresponding to the first neighborhood size.
[0217] Determine the second neighborhood size. Based on the first relation data, the coordinates of the registration locations of multiple first accounts, and the second neighborhood size, determine the first probability density of each first account. Calculate the root mean square error based on the first probability densities of multiple first accounts to obtain the first error corresponding to the second neighborhood size. Continue this process until the smallest first error exists among the first errors corresponding to the multiple neighborhood sizes.
[0218] In another possible implementation, the first error determination module 1202 is used to adjust the first neighborhood size based on the adjustment coefficient to obtain the second neighborhood size.
[0219] In another possible implementation, the registration location information for any account includes primary location information and secondary location information;
[0220] The primary location information includes primary longitude and primary latitude coordinates, which indicate the lowest level of administrative region to which the registered location belongs; the secondary location information includes secondary longitude and secondary latitude coordinates, which indicate the next higher level of administrative region to which the primary location information indicates; the device also includes:
[0221] The coordinate determination module is used to determine the coordinates of the registered location using the primary longitude coordinates and the secondary latitude coordinates; or,
[0222] The coordinate determination module is also used to determine the primary latitude coordinates and secondary longitude coordinates as the coordinates of the registered location.
[0223] In another possible implementation, the device also includes:
[0224] The account partitioning module is used to divide the registration data of multiple candidate accounts into multiple datasets. Each dataset includes the registration data of at least one candidate account, and the registration data includes the registration location. The number of registration data in each dataset is no greater than the first reference number, or the size of the range formed by the registration locations in each dataset is no greater than the reference size.
[0225] The account segmentation module is also used to identify multiple candidate accounts corresponding to multiple registration data in any dataset as multiple primary accounts.
[0226] In another possible implementation, the account partitioning module is used for:
[0227] The registration data of multiple candidate accounts is divided into a target number of primary datasets;
[0228] For each first dataset, if the number of registered data in the first dataset is greater than the first reference number and the size of the range formed by the registered positions in the first dataset is greater than the reference size, the first dataset is divided into a target number of second datasets until the number of registered data in each of the currently divided datasets is no greater than the first reference number, or the size of the range formed by the registered positions in each dataset is no greater than the reference size.
[0229] In another possible implementation, the registration data also includes the registration time, and the device further includes:
[0230] The relationship acquisition module 1201 is also used to acquire third relationship data, which represents the relationship between the registration time and duration of the target account and at least one other account and the second probability density of the target account. The duration is used to determine a time period centered on the registration time of the target account and including twice the duration. The second probability density of the target account represents the probability that the target account belongs to a non-safe type based on the type of each account registered within the time period.
[0231] The second error determination module is used to determine a duration each time, and based on the third relationship data, the registration time of multiple first accounts and the currently determined duration, determine the second probability density of each first account, and determine the second error based on the second probability density of each first account. The second error represents the accuracy of predicting the type of multiple first accounts.
[0232] The target duration determination module is used to determine the duration with the smallest second error from among the various determined durations as the target duration;
[0233] The relationship determination module 1204 is also used to determine the fourth relationship data based on the third relationship data, the registration time and target duration of multiple first accounts, and the fourth relationship data represents the relationship between the registration time of the target account and the second probability density of the target account.
[0234] The device also includes:
[0235] The relationship determination module 1204 is also used to determine the fifth relationship data based on the second relationship data and the fourth relationship data. The fifth relationship data represents the relationship between the coordinates of the registration location of the target account, the registration time and the probability density of the target account. The probability density of the target account is the product of the first probability density and the second probability density of the target account. The probability density of the target account represents the probability that the target account belongs to the non-safe type based on the type of each account registered within the range and within the time period.
[0236] The account detection module 1205 is also used to determine the probability density of the second account based on the coordinates of the registration location, the registration time, and the fifth relationship data of the second account.
[0237] In another possible implementation, the second error determination module is used for:
[0238] The first duration is determined. Based on the third relationship data, the coordinates of the registration locations of multiple first accounts, and the first duration, the second probability density of each first account is determined. The root mean square error is calculated based on the second probability density of multiple first accounts to obtain the second error corresponding to the first duration.
[0239] The second duration is determined. Based on the third relationship data, the coordinates of the registration locations of multiple first accounts, and the second duration, the second probability density of each first account is determined. The root mean square error is calculated based on the second probability density of multiple first accounts to obtain the second error corresponding to the second duration. This process continues until the second error corresponding to the multiple durations is found to be the smallest.
[0240] In another possible implementation, the device also includes:
[0241] The verification module is used to determine the density probability of a third account based on the coordinates of its registration location, registration time, and fifth relationship data, and to determine whether the third account belongs to a safe or non-safe type.
[0242] The verification module is also used to update the target size and target duration based on the coordinates of the registration location, registration time, first relationship data, and third relationship data of multiple first accounts when the type of probability density representation of the third account is inconsistent with the type to which the third account belongs.
[0243] In another possible implementation, the device also includes:
[0244] The verification module is used to determine the area of the first region and the first number. The area of the first region is the area of the region where the coordinates of the registration locations of the multiple first accounts are located, and the first number is the number of the multiple first accounts.
[0245] The verification module is also used to determine the area of the second region and the second number. The area of the second region is the area of the reference region, and the second number is the number of multiple fourth accounts. The area where the coordinates of the registration locations of the multiple fourth accounts are located is the reference region.
[0246] The inspection module is also used to determine a first ratio between the first number and the second number, and a second ratio between the area of the first region and the area of the second region;
[0247] The verification module is also used to determine the prediction accuracy parameter based on the ratio between the first ratio and the second ratio. The prediction accuracy parameter represents the accuracy of the probability density predicted based on the second relationship data.
[0248] In another possible implementation, the device also includes:
[0249] The account selection module is used to select multiple primary accounts that meet the target conditions from multiple candidate accounts. The target conditions include that the candidate accounts have corresponding registration locations and that the registration locations meet the target location format.
[0250] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.
[0251] It should be noted that the account detection device provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the account detection device and the account detection method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0252] This application also provides a computer device, which includes a processor and a memory. The memory stores at least one computer program, which is loaded and executed by the processor to perform the operations of the account detection method described above.
[0253] Optionally, the computer device is provided as a terminal. Figure 13 This is a schematic diagram of the structure of a terminal 1300 provided in an embodiment of this application. The terminal 1300 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1300 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.
[0254] Terminal 1300 includes a processor 1301 and a memory 1302.
[0255] Processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1301 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1301 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1301 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0256] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 are used to store at least one computer program, which is executed by the processor 1301 to implement the account detection method provided in the method embodiments of this application.
[0257] In some embodiments, the terminal 1300 may also optionally include a peripheral device interface 1303 and at least one peripheral device. The processor 1301, memory 1302, and peripheral device interface 1303 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1303 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, and a power supply 1308.
[0258] Peripheral device interface 1303 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1301 and memory 1302. In some embodiments, processor 1301, memory 1302 and peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1301, memory 1302 and peripheral device interface 1303 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0259] The radio frequency (RF) circuit 1304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1304 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1304 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1304 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1304 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0260] Display screen 1305 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1305 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1301 for processing. In this case, display screen 1305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 1305 may be a single screen, disposed on the front panel of terminal 1300; in other embodiments, display screen 1305 may be at least two screens, disposed on different surfaces of terminal 1300 or in a folded design; in still other embodiments, display screen 1305 may be a flexible display screen, disposed on a curved or folded surface of terminal 1300. Furthermore, display screen 1305 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1305 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0261] The camera assembly 1306 is used to acquire images or videos. Optionally, the camera assembly 1306 includes a front-facing camera and a rear-facing camera. The front-facing camera is disposed on the front panel of the terminal, and the rear-facing camera is disposed on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1306 may also include a flash. The flash may be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.
[0262] The audio circuit 1307 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1301 for processing, or input to the radio frequency circuit 1304 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 1300. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1301 or the radio frequency circuit 1304 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1307 may also include a headphone jack.
[0263] Power supply 1308 is used to power the various components in terminal 1300. Power supply 1308 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1308 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0264] Those skilled in the art will understand that Figure 13 The structure shown does not constitute a limitation on terminal 1300 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0265] Optionally, the computer device is provided as a server. Figure 14 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1400 can vary considerably due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1401 and one or more memories 1402. The memory 1402 stores at least one computer program, which is loaded and executed by the processor 1401 to implement the methods provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.
[0266] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to implement the operations performed by the account detection method of the above embodiments.
[0267] This application also provides a computer program product, which includes a computer program that, when executed by a processor, performs the operations of the account detection method described above.
[0268] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.
[0269] It is understood that in the specific implementation of this application, data related to account, registration location, registration time, etc. are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0270] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0271] The above are merely optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present application should be included within the protection scope of the present application.
Claims
1. An account detection method, characterized in that, The method includes: Obtain first relational data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account. The neighborhood size is used to determine a range centered on the registration location of the target account and with the neighborhood size as the radius. The first probability density of the target account represents the probability that the target account belongs to a non-safe type, predicted based on the type of each account registered within the range. Each time a neighborhood size is determined, based on the first relationship data, the coordinates of the registration locations of the multiple first accounts, and the currently determined neighborhood size, a first probability density is determined for each first account, and a first error is determined based on the first probability density of each first account. The multiple first accounts are accounts belonging to the non-safe type, and the first error represents the accuracy of predicting the type to which the multiple first accounts belong. From the determined neighborhood sizes, the neighborhood size with the smallest first error is determined as the target size; Based on the first relationship data, the coordinates of the registration locations of the plurality of first accounts, and the target size, second relationship data is determined, wherein the second relationship data represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account; Based on the coordinates of the registration location of the second account to be predicted and the second relationship data, the first probability density of the second account is determined.
2. The method according to claim 1, characterized in that, Each time a neighborhood size is determined, based on the first relationship data, the coordinates of the registration locations of multiple first accounts, and the currently determined neighborhood size, a first probability density is determined for each first account, and a first error is determined based on the first probability density of each first account, including: Determine the first neighborhood size, and based on the first relationship data, the coordinates of the registration locations of the multiple first accounts, and the first neighborhood size, determine the first probability density of each first account, and calculate the root mean square error based on the first probability density of the multiple first accounts to obtain the first error corresponding to the first neighborhood size; A second neighborhood size is determined. Based on the first relationship data, the coordinates of the registration locations of the multiple first accounts, and the second neighborhood size, a first probability density is determined for each first account. The root mean square error is calculated based on the first probability density of the multiple first accounts to obtain the first error corresponding to the second neighborhood size. This process continues until a minimum first error is found among the first errors corresponding to the multiple neighborhood sizes.
3. The method according to claim 2, characterized in that, Determining the second neighborhood size includes: The second neighborhood size is obtained by adjusting the first neighborhood size based on the adjustment factor.
4. The method according to claim 1, characterized in that, The registration location information for any account includes primary location information and secondary location information; The primary location information includes primary longitude coordinates and primary latitude coordinates, and the primary location information refers to the lowest level of administrative region to which the registered location belongs; The secondary location information includes secondary longitude coordinates and secondary latitude coordinates. The secondary location information refers to the administrative region above the administrative region indicated by the primary location information. The method further includes: The primary longitude coordinates and the secondary latitude coordinates are determined as the coordinates of the registered location; or, The primary latitude coordinates and the secondary longitude coordinates are determined as the coordinates of the registered location.
5. The method according to claim 1, characterized in that, The method further includes: The registration data of multiple candidate accounts are divided into multiple datasets. Each dataset includes the registration data of at least one candidate account, and the registration data includes the registration location. The number of registration data in each dataset is no greater than the first reference number, or the size of the range formed by the registration locations in each dataset is no greater than the reference size. Multiple candidate accounts corresponding to multiple registration data in any dataset are determined as the multiple first accounts.
6. The method according to claim 5, characterized in that, The process of dividing the registration data of multiple candidate accounts into multiple datasets includes: The registration data of the multiple candidate accounts is divided into a target number of first datasets; For each first dataset, if the number of registered data in the first dataset is greater than the first reference number and the size of the range formed by the registered positions in the first dataset is greater than the reference size, the first dataset is divided into the target number of second datasets until the number of registered data in each of the currently divided datasets is no greater than the first reference number, or the size of the range formed by the registered positions in each dataset is no greater than the reference size.
7. The method according to claim 1, characterized in that, The registration data also includes the registration time, and the method further includes: Obtain third relation data, which represents the relationship between the registration time and duration of the target account and at least one other account and the second probability density of the target account. The duration is used to determine a time period centered on the registration time of the target account and containing twice the duration. The second probability density of the target account represents the probability that the target account belongs to the non-safe type, predicted based on the type of each account registered within the time period. Each time a duration is determined, based on the third relationship data, the registration time of the multiple first accounts, and the currently determined duration, a second probability density of each first account is determined, and based on the second probability density of each first account, a second error is determined, the second error representing the accuracy of predicting the type of the multiple first accounts; From the various determined durations, the duration with the smallest corresponding second error is determined as the target duration; Based on the third relationship data, the registration time of the multiple first accounts, and the target duration, fourth relationship data is determined, which represents the relationship between the registration time of the target account and the second probability density of the target account. After determining the second relationship data based on the first relationship data, the coordinates of the registration locations of the plurality of first accounts, and the target size, the method further includes: Based on the second relationship data and the fourth relationship data, a fifth relationship data is determined. The fifth relationship data represents the relationship between the coordinates of the registration location of the target account, the registration time, and the probability density of the target account. The probability density of the target account is the product of the first probability density and the second probability density of the target account. The probability density of the target account represents the probability that the target account belongs to the non-safe type, predicted based on the type of each account registered within the range and time period. Based on the coordinates of the registration location of the second account, the registration time, and the fifth relationship data, the probability density of the second account is determined.
8. The method according to claim 7, characterized in that, Each time a duration is determined, based on the third relationship data, the registration time of the multiple first accounts, and the currently determined duration, a second probability density is determined for each first account, and a second error is determined based on the second probability density of each first account, including: A first duration is determined. Based on the third relationship data, the coordinates of the registration locations of the multiple first accounts, and the first duration, a second probability density of each first account is determined. The root mean square error is calculated based on the second probability density of the multiple first accounts to obtain the second error corresponding to the first duration. A second duration is determined. Based on the third relationship data, the coordinates of the registration locations of the multiple first accounts, and the second duration, a second probability density is determined for each first account. The root mean square error is calculated based on the second probability density of the multiple first accounts to obtain the second error corresponding to the second duration. This process continues until a minimum second error is found among the multiple durations.
9. The method according to claim 7, characterized in that, The method further includes: Based on the coordinates of the registration location of the third account, the registration time, and the fifth relationship data, the density probability of the third account is determined, and the type to which the third account belongs is either the safe type or the non-safe type; If the type of probability density representation of the third account is inconsistent with the type to which the third account belongs, the target size and the target duration are updated based on the coordinates of the registration location of the plurality of first accounts, the registration time, the first relationship data and the third relationship data.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Determine the area of the first region and the number of the first accounts. The area of the first region is the area of the region where the coordinates of the registration locations of the multiple first accounts are located, and the number of the multiple first accounts is the number of the multiple first accounts. Determine the area of the second region and the number of the second region, where the area of the second region is the area of the reference region, and the number of the second region is the number of the multiple fourth accounts, and the area where the coordinates of the registration locations of the multiple fourth accounts are located is the reference region; Determine a first ratio between the first number and the second number, and a second ratio between the area of the first region and the area of the second region; Based on the ratio between the first ratio and the second ratio, a prediction accuracy parameter is determined, wherein the prediction accuracy parameter represents the accuracy of the probability density predicted based on the second relationship data.
11. The method according to any one of claims 1-9, characterized in that, The method further includes: Select the first accounts that meet the target conditions from a plurality of candidate accounts. The target conditions include that the candidate accounts have a corresponding registration location and that the registration location meets the target location format.
12. An account detection device, characterized in that, The device includes: The relationship acquisition module is used to acquire first relationship data, which represents the relationship between the coordinates of the registration location of the target account and at least one other account, the neighborhood size, and the first probability density of the target account. The neighborhood size is used to determine the range centered on the registration location of the target account and with the neighborhood size as the radius. The first probability density of the target account represents the probability that the target account belongs to a non-safe type based on the type of each account registered within the range. The first error determination module is used to determine a neighborhood size each time, and based on the first relationship data, the coordinates of the registration locations of the multiple first accounts and the currently determined neighborhood size, determine the first probability density of each first account, and determine the first error based on the first probability density of each first account, wherein the multiple first accounts are accounts belonging to the non-security type, and the first error represents the accuracy of predicting the type to which the multiple first accounts belong. The target size determination module is used to determine the neighborhood size with the smallest first error from among the determined neighborhood sizes as the target size; The relationship determination module is used to determine second relationship data based on the first relationship data, the coordinates of the registration locations of the plurality of first accounts, and the target size. The second relationship data represents the relationship between the coordinates of the registration location of the target account and the first probability density of the target account. The account detection module is used to determine the first probability density of the second account based on the coordinates of the registration location of the second account to be predicted and the second relationship data.
13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to perform the operations of the account detection method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to perform the operations of the account detection method as described in any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it performs the operations of the account detection method according to any one of claims 1 to 11.