A clustering method and device for evaluation objects
By constructing the sample points of the evaluation object and judging based on the number of sample points in the neighborhood and preset thresholds, the problem of the inability to accurately identify real and abnormal driving records in the prior art is solved, and the rapid and accurate classification of the evaluation object is achieved.
Patent Information
- Application Number
- CN201911055761.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2039-10-31
AI Technical Summary
The prior art cannot accurately identify real driving records and abnormal driving records, especially when the characteristic mean method cannot distinguish real records from abnormal records.
By constructing the sample points of the evaluation object, determining the number of sample points in the neighborhood, and determining the cluster cluster to which the sample points belong based on the relationship between the number of sample points in the neighborhood and the preset threshold, and then determining that the evaluation object is normal or abnormal data.
It realizes rapid and accurate classification of the evaluation objects, improves the recognition efficiency of normal samples and abnormal samples, especially when the majority of the evaluation objects are normal samples, providing a clear classification basis.
Smart Images

Figure CN110866549B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of financial technology (Fintech), and in particular, to a clustering method and device for evaluation objects. Background Art
[0002] With the development of computer technology, more and more technologies (such as blockchain, cloud computing, or big data) are applied in the financial field. The traditional financial industry is gradually transforming into financial technology, and big data technology is no exception. However, due to the security and real-time requirements of the financial and payment industries, higher requirements are also imposed on big data technology.
[0003] For example, the financial field can use big data technology to review the trade background of customers. For loan requests submitted by customers in the transportation and logistics industries, banks need to conduct trade background reviews on them. Customers who pass the trade background review can obtain loans from the bank. Generally speaking, trade background reviews can involve the analysis of transportation route data. Banks can require customers to provide all driving record information about a fixed transportation route. However, most of the data in all the driving record information about this fixed route fed back by customers to the bank is real, but there is still a small amount of data that is false driving records, that is, abnormal driving records. Banks need to analyze all the driving records provided by customers about a fixed route.
[0004] When the prior art solves the above problems, it adopts the method of feature mean. By calculating the feature mean of all driving records, and then comparing each driving record with the feature mean: if the difference between each driving record and the feature mean is not large, it is considered that this driving record is a real driving record; if the difference between each driving record and the feature mean is very large, it is considered that this driving record is an abnormal driving record. However, for the situation where the difference between a real driving record and the feature mean is very large and the difference between an abnormal driving record and the feature mean is not large, this method of feature mean cannot accurately classify real driving records and abnormal driving records. Summary of the Invention
[0005] The present invention provides a clustering method and device for evaluation objects to solve the problem that the prior art cannot accurately identify normal data and abnormal data.
[0006] In a first aspect, an embodiment of the present invention provides a clustering method for evaluation objects, the method comprising: constructing sample points corresponding to the respective evaluation objects according to the attribute information of the respective evaluation objects; determining the clustering clusters to which the respective sample points belong; wherein, for any sample point, the clustering cluster to which it belongs is determined by the following method: determining the number of sample points within the neighborhood of the sample point; if the number of sample points within the neighborhood meets the clustering point requirement, determining the clustering cluster to which the sample point belongs as the clustering cluster to which the sample points within the neighborhood belong; the neighborhood is a set area range based on the sample point; the clustering point requirement is that the number of sample points within the neighborhood is not less than a preset threshold, or the number of sample points within the neighborhood is greater than the number of sample points in the clustering cluster to which the sample point belongs; determining, according to the number of sample points in each clustering cluster, each clustering cluster as the cluster where normal sample points are located or the cluster where abnormal sample points are located.
[0007] Based on this solution, by constructing sample points corresponding to the respective evaluation objects, the clustering clusters to which the respective sample points belong are determined through the clustering point requirement; at the same time, the clustering point requirement includes judging the relationship between the number of sample points within the neighborhood and the preset threshold, and the relationship between the number of sample points within the neighborhood and the number of sample points in the clustering cluster to which the sample point belongs. Determining the clustering cluster to which the sample point belongs from multiple judgment bases helps to determine the attribution of the respective evaluation objects; finally, it is determined whether each evaluation object is normal data or abnormal data according to the number of sample points in each clustering cluster.
[0008] In a possible implementation method, if the sample point currently has no belonging clustering cluster, taking the sample points within the neighborhood as a clustering cluster; or if the number of sample points within the neighborhood is less than the preset threshold and the number of sample points within the neighborhood is not greater than the number of sample points in the clustering cluster to which the sample point belongs, selecting the next sample point in the clustering cluster to which the sample point belongs for judgment on whether it meets the clustering point requirement until whether any sample point in the clustering cluster to which the sample point belongs has completed the judgment on whether it is a clustering point.
[0009] The above implementation method further classifies the sample points that currently have no belonging clustering cluster, and at the same time, realizes a cyclic judgment on the sample points in the clustering cluster, making the entire clustering process faster.
[0010] In a possible implementation method, the respective evaluation objects are N travel record information; constructing sample points corresponding to the respective evaluation objects according to the attribute information of the respective evaluation objects includes: for each travel record information, constructing a sample point corresponding to the travel record information according to the starting position and the ending position in the travel record information; determining the number of sample points within the neighborhood of the sample point includes: determining the distance between any two sample points; determining the sample points whose distance from the sample point is within the neighborhood as the sample points within the neighborhood.
[0011] In the above implementation method, taking travel records as the evaluation object, it further refines how to form sample points and how to determine the sample points within the neighborhood, thereby realizing the evaluation of travel records.
[0012] In a possible implementation method, according to the number of sample points in each clustering cluster, it is determined whether each clustering cluster is a cluster where normal sample points are located or a cluster where abnormal sample points are located, including: determining the clustering cluster with the largest number of sample points as the cluster where normal sample points are located; determining the clustering clusters other than the clustering cluster with the largest number of sample points as the clusters where abnormal sample points are located.
[0013] In the above implementation method, in the case where the vast majority of the evaluation objects are normal samples, it gives the basis for dividing normal samples and abnormal samples. At the same time, each abnormal sample will also have its own affiliated clustering cluster, which is convenient for classifying and analyzing abnormal samples.
[0014] In a second aspect, an embodiment of the present invention provides a clustering device for an evaluation object, and the device includes: a construction unit, configured to construct each sample point corresponding to each evaluation object according to the attribute information of each evaluation object; a first determination unit, configured to determine the clustering cluster to which each sample point belongs; wherein, for any sample point, the clustering cluster to which it belongs is determined through the following method: determining the number of sample points within the neighborhood of the sample point; if the number of sample points within the neighborhood meets the clustering point requirement, determining the clustering cluster to which the sample point belongs as the clustering cluster to which the sample points within the neighborhood belong; the neighborhood is a set regional range based on the sample point; the clustering point requirement is that the number of sample points within the neighborhood is not less than a preset threshold, or the number of sample points within the neighborhood is greater than the number of sample points in the clustering cluster to which the sample point belongs; a second determination unit, configured to determine whether each clustering cluster is a cluster where normal sample points are located or a cluster where abnormal sample points are located according to the number of sample points in each clustering cluster.
[0015] Based on this solution, by constructing each sample point corresponding to each evaluation object, it is determined to which clustering cluster each sample point belongs through the clustering point requirement; at the same time, the clustering point requirement includes judging the relationship between the number of sample points within the neighborhood and the preset threshold, and the relationship between the number of sample points within the neighborhood and the number of sample points in the clustering cluster to which the sample point belongs. Among the multiple judgment bases for determining the clustering cluster to which the sample point belongs, it helps to determine the attribution of each evaluation object; finally, it is determined whether each evaluation object is normal data or abnormal data according to the number of sample points in each clustering cluster.
[0016] In a possible implementation method, the first determination unit is configured to: if the sample point currently has no affiliated clustering cluster, use the sample points in the neighborhood as a clustering cluster; or if the number of sample points in the neighborhood is less than the preset threshold and the number of sample points in the neighborhood is not greater than the number of sample points in the clustering cluster to which the sample point belongs, select the next sample point in the clustering cluster to which the sample point belongs to determine whether it meets the requirements of a clustering point until any sample point in the clustering cluster to which the sample point belongs has completed the determination of whether it is a clustering point.
[0017] The above implementation method further classifies the sample points that currently have no affiliated clustering cluster. At the same time, it realizes a cyclic judgment of the sample points in the clustering cluster, making the entire clustering process faster.
[0018] In a possible implementation method, the evaluation objects are N trip record information; the construction unit is specifically configured to: for each trip record information, construct a sample point corresponding to the trip record information according to the starting position and the ending position in the trip record information; the first determination unit is specifically configured to: determine the distance between any two sample points; determine the sample points whose distance from the sample point is within the neighborhood as the sample points in the neighborhood.
[0019] The above implementation method takes the trip record as the evaluation object and further refines how to form sample points and how to determine the sample points in the neighborhood, thereby realizing the evaluation of the trip record.
[0020] In a possible implementation method, the second determination unit is specifically configured to: determine the clustering cluster with the largest number of sample points as the cluster where the normal sample points are located; determine the clustering clusters other than the clustering cluster with the largest number of sample points as the clusters where the abnormal sample points are located.
[0021] The above implementation method gives the basis for dividing normal samples and abnormal samples in the case that the vast majority of the evaluation objects are normal samples. At the same time, each abnormal sample will also have its own affiliated clustering cluster, which is convenient for classifying and analyzing abnormal samples.
[0022] In a third aspect, an embodiment of the present invention provides a computing device, including:
[0023] A memory for storing program instructions;
[0024] A processor for calling the program instructions stored in the memory and executing the method according to any one of the first aspects as obtained by the program.
[0025] Fourthly, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the method according to any one of the first aspect. Description of the Drawings
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 A clustering method for evaluation objects provided by an embodiment of the present invention;
[0028] Figure 2 A clustering schematic diagram provided by an embodiment of the present invention;
[0029] Figure 3 A clustering device for evaluation objects provided by an embodiment of the present invention. Detailed Embodiments
[0030] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0031] As Figure 1 shown, a clustering method for evaluation objects provided by an embodiment of the present invention includes the following steps:
[0032] Step 101: Construct respective sample points corresponding to the respective evaluation objects according to the attribute information of the respective evaluation objects.
[0033] Step 102: Determine the clustering clusters to which the respective sample points belong; wherein, for any sample point, the clustering cluster to which it belongs is determined by the following method: determine the number of sample points within the neighborhood of the sample point; if the number of sample points within the neighborhood meets the clustering point requirement, then determine the clustering cluster to which the sample point belongs as the clustering cluster to which the sample points within the neighborhood belong; the neighborhood is a set regional range based on the sample point; the clustering point requirement is that the number of sample points within the neighborhood is not less than a preset threshold, or the number of sample points within the neighborhood is greater than the number of sample points in the clustering cluster to which the sample point belongs.
[0034] Step 103: Determine whether each cluster is a cluster where normal sample points are located or a cluster where abnormal sample points are located according to the number of sample points in each cluster.
[0035] Based on this solution, by constructing sample points corresponding to each evaluation object, determine the cluster to which each sample point belongs by meeting the requirements for clustering points. At the same time, the requirements for clustering points include judging the relationship between the number of sample points in the neighborhood and a preset threshold, and the relationship between the number of sample points in the neighborhood and the number of sample points in the cluster to which the sample point belongs. Determining the cluster to which the sample point belongs from multiple judgment bases helps to determine the attribution of each evaluation object. Finally, determine whether each evaluation object is normal data or abnormal data according to the number of sample points in each cluster.
[0036] In the above step 101, the evaluation object can be the multiple trip record information of multiple transportation routes. For example, there can be 3 transportation routes, namely the transportation route between place A and place B, the transportation route between place C and place D, and the transportation route between place E and place F. Among them, the number of occurrences of the transportation route between place A and place B can be 50 times, the number of occurrences of the transportation route between place C and place D can be 8 times, and the number of occurrences of the transportation route between place E and place F can be 2 times. Each of the above 60 transportation routes, that is, each transportation route in these 60 trips, is an evaluation object.
[0037] The attribute information of the evaluation object can include many contents. Due to the differences in evaluation objects, the specific attribute information is also different. For example, for each transportation route, its attribute information can be the time period formed between the time when the transportation route occurs and the time when the transportation route arrives, or it can be the spatial information formed between the specific geographical location where the transportation route occurs and the specific geographical location where the transportation route arrives.
[0038] After obtaining the attribute information of each evaluation object, construct sample points corresponding to each evaluation object. For example, for the above 60 transportation routes in total, the time period formed between the time when each transportation route occurs and the time when the transportation route arrives can be used as a sample point, so that sample points corresponding to 60 transportation routes can be formed; or the spatial information formed between the specific geographical location where each transportation route occurs and the specific geographical location where the transportation route arrives can be used as a sample point, so that sample points corresponding to 60 transportation routes can be formed.
[0039] It should be noted that due to the differences among the evaluation objects, for the same attribute, the attribute information shown by each evaluation object is not the same. Therefore, when constructing the sample points corresponding to each evaluation object according to the attribute information of each evaluation object, it is for the same attribute information of each evaluation object, rather than different attribute information of each evaluation object. For example, among the sample points corresponding to each evaluation object constructed, it can be the time period formed between the moment when each transportation route occurs and the moment when the transportation route arrives as a sample point, or it can be the spatial information formed between the specific geographical location where each transportation route occurs and the specific geographical location where the transportation route arrives as a sample point; it cannot be that the time periods formed between the moments when some transportation routes occur and the moments when the transportation routes arrive are used as some sample points, and at the same time, the spatial information formed between the specific geographical locations where some other transportation routes occur and the specific geographical locations where the transportation routes arrive is used as some other sample points.
[0040] In the above step 102, taking the 60 transportation routes mentioned above as an example, assume that the sample points formed by these 60 transportation routes are based on the spatial information formed between the specific geographical location where each transportation route occurs and the specific geographical location where the transportation route arrives.
[0041] Determine the cluster to which each sample point belongs, that is, it is necessary to determine which one of the 3 transportation routes, namely the transportation route between place A and place B, the transportation route between place C and place D, and the transportation route between place E and place F, each transportation route in the 60 transportation routes mentioned above specifically corresponds to.
[0042] For any sample point, determine the cluster to which it belongs in the following way: determine the number of sample points within the neighborhood of the sample point; if the number of sample points within the neighborhood meets the clustering point requirement, then determine the cluster to which the sample point belongs as the cluster to which the sample points within the neighborhood belong; the neighborhood is a set area range based on the sample point; the clustering point requirement is that the number of sample points within the neighborhood is not less than a preset threshold, or the number of sample points within the neighborhood is greater than the number of sample points in the cluster to which the sample point belongs.
[0043] Among them, the neighborhood is a set area range based on the sample point. For example, it can be a circle with the sample point as the center and a preset distance as the radius. So the set area range based on the sample point is a circle, that is, the neighborhood of the sample point is a circle. Of course, the neighborhood can also be defined in other ways. For example, it can be a preset arbitrary shape as the set area range based on the sample point. So the set area range based on the sample point is a preset arbitrary shape, that is, the neighborhood of the sample point is a preset arbitrary shape. In this regard, the present invention does not make a limitation.
[0044] For the number of sample points in the neighborhood to be not less than a preset threshold, and the preset threshold can be a value set by R & D personnel according to experience.
[0045] Specifically, when it is necessary to determine the cluster to which each sample point formed by the transportation route in each of the foregoing 60 transportation routes belongs, it can be carried out in the following manner: Take any one of the foregoing 60 transportation routes as an inspection sample point, and determine the number of sample points in the neighborhood of the inspection sample point; If the number of sample points in the neighborhood of the inspection sample point meets the requirements of the clustering point, then the cluster to which the inspection sample point belongs can be determined as the cluster to which the sample points in the neighborhood of the inspection sample point belong.
[0046] The requirements for the above-mentioned clustering points are that the number of sample points in the neighborhood of the inspection sample point is not less than the preset threshold, or the number of sample points in the neighborhood of the inspection sample point is greater than the number of sample points in the cluster with the majority of the inspection sample point.
[0047] Through the clustering of the above process, the corresponding clusters can be obtained. Taking the foregoing 60 transportation routes as an example, 3 corresponding clusters can be obtained, which are respectively the clusters corresponding to the transportation routes between place A and place B, the transportation routes between place C and place D, and the transportation routes between place E and place F.
[0048] In the above step 103, taking the 3 clusters obtained above as an example, that is, taking the clusters corresponding to the transportation routes between place A and place B, the transportation routes between place C and place D, and the transportation routes between place E and place F as an example, determine the number of sample points in these 3 clusters. For example, the number of sample points in the cluster corresponding to the transportation route between place A and place B is 50, the number of sample points in the cluster corresponding to the transportation route between place C and place D is 8, and the number of sample points in the cluster corresponding to the transportation route between place E and place F is 2.
[0049] When determining whether each of the above three clusters belongs to the cluster where normal sample points are located or the cluster where abnormal sample points are located, a threshold can be set. For example, the set threshold is 40. If the number of sample points in a cluster is not less than 40, then it is considered that this cluster is the cluster where normal sample points are located; if the number of sample points in a cluster is less than 40, then it is considered that this cluster is the cluster where abnormal sample points are located. As in the previous example, since the number of sample points in the cluster corresponding to the transportation route between Place A and Place B is 50, which is greater than the threshold 40, it is considered that the cluster corresponding to the transportation route between Place A and Place B is the cluster where normal sample points are located; since the number of sample points in the cluster corresponding to the transportation route between Place C and Place D is 8, which is less than the threshold 40, it is considered that the cluster corresponding to the transportation route between Place C and Place D is the cluster where abnormal sample points are located; for the same reason, it can also be considered that the cluster corresponding to the transportation route between Place E and Place F is the cluster where abnormal sample points are located.
[0050] When determining whether each of the above three clusters belongs to the cluster where normal sample points are located or the cluster where abnormal sample points are located, it can also be to arrange the number of sample points in these three clusters in descending order, and consider several clusters whose sorting positions meet the preset sorting as the clusters where normal sample points are located, and consider several clusters whose sorting positions do not meet the preset sorting as the clusters where abnormal sample points are located. For example, consider the cluster ranked second and the clusters before the second (i.e., the cluster ranked first) as the clusters where normal sample points are located, and consider the cluster ranked third and the clusters after the third as the clusters where abnormal sample points are located. As in the previous example, since the number of sample points in the cluster corresponding to the transportation route between Place A and Place B is 50, the number of sample points in the cluster corresponding to the transportation route between Place C and Place D is 8, and the number of sample points in the cluster corresponding to the transportation route between Place E and Place F is 2, arranging the number of sample points in these three clusters in descending order, we can get the cluster corresponding to the transportation route between Place A and Place B ranked first, the cluster corresponding to the transportation route between Place C and Place D ranked second, and the cluster corresponding to the transportation route between Place E and Place F ranked third. Then it can be considered that these two clusters, namely the cluster corresponding to the transportation route between Place A and Place B and the cluster corresponding to the transportation route between Place C and Place D, are the clusters where normal sample points are located, and the cluster corresponding to the transportation route between Place E and Place F is the cluster where abnormal sample points are located.
[0051] As a possible implementation, if the sample point does not currently belong to a cluster, the sample points in the neighborhood are taken as a cluster; or if the number of sample points in the neighborhood is less than the preset threshold and the number of sample points in the neighborhood is not greater than the number of sample points in the cluster to which the sample point belongs, then the next sample point in the cluster to which the sample point belongs is selected to determine whether it meets the requirements of a clustering point until whether any sample point in the cluster to which the sample point belongs has been determined to be a clustering point.
[0052] As Figure 2 shown, it is a clustering schematic diagram provided by an embodiment of the present invention. Figure 2 Each circled number in it represents a sample point. The circled number "1" represents the No. 1 sample point, the circled number "2" represents the No. 2 sample point, and the meanings of the remaining circled numbers are not elaborated here. Let the preset threshold be 4.
[0053] Taking the No. 14 sample point as an example, the cluster to which it belongs is determined in the following way: Since the starting point of clustering is the No. 14 sample point, it means that the No. 14 sample point does not currently belong to a cluster. Therefore, the sample points in the neighborhood of the No. 14 sample point can be taken as a cluster. In addition to itself, the sample points in the neighborhood of the No. 14 sample point also include the No. 15 sample point. Therefore, the two sample points, namely the No. 14 sample point and the No. 15 sample point, can be taken as a cluster.
[0054] After the inspection of the No. 14 sample point is completed and the cluster to which the No. 14 sample point belongs is obtained, any sample point in the neighborhood of the No. 14 sample point needs to be selected for the inspection of clustering points. Since there is only the No. 15 sample point in addition to itself in the cluster to which the No. 14 sample point belongs, the No. 15 sample point is selected for the inspection of clustering points. It can be found that in addition to itself, the sample points in the neighborhood of the No. 15 sample point also include the No. 14 sample point, that is, the number of sample points in the neighborhood of the No. 15 sample point is 2. When judging whether the No. 15 sample point meets the requirements of a clustering point, it can be found that the No. 15 sample point does not meet the requirements of a clustering point, as shown below: the number of sample points in the neighborhood of the No. 15 sample point is 2, 2 is less than the preset threshold 4; and 2 is equal to the number of sample points in the cluster to which the No. 15 sample point belongs, which is 2 (here the cluster to which the No. 15 sample point belongs refers to the cluster to which the No. 14 sample point and the No. 15 sample point belong).
[0055] Therefore, the No. 14 sample point is selected as the starting sample point, and the finally obtained cluster is the cluster containing the No. 14 sample point and the No. 15 sample point.
[0056] After the inspection of the No. 14 sample point and the No. 15 sample point is completed, it is necessary to determine the clusters to which the other sample points except the No. 14 sample point and the No. 15 sample point belong.
[0057] For example, taking sample point No. 12 as an example next, determine the cluster it belongs to through the following method: Since the starting point of clustering is sample point No. 12, it means that sample point No. 12 currently does not belong to any cluster. Therefore, the sample points within the neighborhood of sample point No. 12 can be regarded as a cluster. The sample points within the neighborhood of sample point No. 12, in addition to itself, also include sample point No. 13 and sample point No. 11. Therefore, the three sample points, namely sample point No. 12, sample point No. 13, and sample point No. 11, can be regarded as a cluster.
[0058] After the inspection of sample point No. 12 is completed and the cluster to which sample point No. 12 belongs is obtained, it is necessary to continue to select any sample point within the neighborhood of sample point No. 12 for the inspection of clustering points. Since, in addition to itself, the cluster to which sample point No. 12 belongs also includes sample point No. 13 and sample point No. 11, at this time, any one of the two sample points, sample point No. 13 and sample point No. 11, can be arbitrarily selected for the inspection of clustering points. For example, sample point No. 13 within the neighborhood of sample point No. 12 can be selected first for the inspection of clustering points, and then sample point No. 11 within the neighborhood of sample point No. 12 can be selected for the inspection of clustering points; or sample point No. 11 within the neighborhood of sample point No. 12 can be selected first for the inspection of clustering points, and then sample point No. 13 within the neighborhood of sample point No. 12 can be selected for the inspection of clustering points.
[0059] It should be noted that the present invention does not limit the order of inspection of any sample point within the neighborhood for clustering points, but it is necessary to inspect any sample point within the neighborhood for clustering points, that is, it is necessary to inspect all sample points within the neighborhood for clustering points. For example, in the embodiment of the present invention, sample point No. 13 within the neighborhood of sample point No. 12 is selected first for the inspection of clustering points, and then sample point No. 11 within the neighborhood of sample point No. 12 is selected for the inspection of clustering points.
[0060] When sample point No. 13 within the neighborhood of sample point No. 12 is selected for the inspection of clustering points, it can be found that the sample points within the neighborhood of sample point No. 13, in addition to itself, also include sample point No. 12, that is, the number of sample points within the neighborhood of sample point No. 13 is 2. When judging whether sample point No. 13 meets the requirements of a clustering point, it can be found that sample point No. 13 does not meet the requirements of a clustering point, as shown below: The number of sample points within the neighborhood of sample point No. 13 is 2, and 2 is less than the preset threshold of 4; and 2 is less than the number of sample points 3 in the cluster to which sample point No. 13 belongs (here, the cluster to which sample point No. 13 belongs refers to the cluster where sample point No. 12, sample point No. 13, and sample point No. 11 are located).
[0061] After determining that the sample point No. 13 is not a clustering point, it is necessary to continue to select the sample point No. 11 within the neighborhood of the sample point No. 12 for the investigation of clustering points. It can be found that the sample points within the neighborhood of the sample point No. 11, in addition to itself, also include the sample point No. 12 and the sample point No. 9, that is, the number of sample points within the neighborhood of the sample point No. 11 is 3. When judging whether the sample point No. 11 meets the requirements of a clustering point, it can be found that the sample point No. 11 also does not meet the requirements of a clustering point, as shown below: the number of sample points within the neighborhood of the sample point No. 11 is 3, 3 is less than the preset threshold 4; and 3 is equal to the number of sample points in the clustering cluster to which the sample point No. 11 belongs (here, the clustering cluster to which the sample point No. 11 belongs refers to the clustering cluster where the sample points No. 12, No. 13, and No. 11 are located).
[0062] Therefore, the sample point No. 12 is selected as the starting sample point, and the final obtained clustering cluster is the clustering cluster that includes the sample points No. 12, No. 13, and No. 11.
[0063] Furthermore, after the investigation of the sample points No. 12, No. 13, and No. 11 is completed, it is necessary to determine the clustering clusters to which the other sample points belong except for the sample points No. 12, No. 13, and No. 11.
[0064] For example, next, taking the sample point No. 1 as an example, the clustering cluster to which it belongs is determined in the following way: Since the starting point of clustering is the sample point No. 1, which means that the sample point No. 1 currently does not belong to any clustering cluster, therefore, the sample points within the neighborhood of the sample point No. 1 can be regarded as a clustering cluster. The sample points within the neighborhood of the sample point No. 1, in addition to itself, also include the sample points No. 2, No. 3, and No. 4. Therefore, the 4 sample points, namely the sample points No. 1, No. 2, No. 3, and No. 4, can be regarded as a clustering cluster. At this time, it can also be found that the sample point No. 1 meets the requirements of a clustering point, as shown below: the number of sample points within the neighborhood of the sample point No. 1 is 4, 4 is equal to the preset threshold 4; and 4 is greater than the number of sample points in the clustering cluster to which the sample point No. 1 belongs (since the sample point No. 1 currently does not belong to any clustering cluster, it can be considered that the number of sample points in the clustering cluster to which the sample point No. 1 currently belongs is 0). Therefore, the sample point No. 1 is a clustering point.
[0065] After the investigation of the sample point No. 1 is completed and the clustering cluster to which the sample point No. 1 belongs is obtained, it is necessary to continue to select any sample point within the neighborhood of the sample point No. 1 for the investigation of clustering points.
[0066] When selecting the 2nd sample point within the neighborhood of the 1st sample point for the examination of clustering points, it can be found that the sample points within the neighborhood of the 2nd sample point, besides itself, also include the 1st sample point, the 3rd sample point, and the 5th sample point. That is, the number of sample points within the neighborhood of the 2nd sample point is 4. When judging whether the 2nd sample point meets the requirements of a clustering point, it can be found that the 2nd sample point meets the requirements, as shown below: the number of sample points within the neighborhood of the 2nd sample point is 4, and 4 is equal to the preset threshold of 4. Although the number of sample points within the neighborhood of the 2nd sample point, which is 4, is equal to the number of sample points within the clustering cluster to which the 2nd sample point belongs (here, the clustering cluster to which the 2nd sample point belongs refers to the clustering cluster where the 1st sample point, the 2nd sample point, the 3rd sample point, and the 4th sample point are located), the judgment rule for a clustering point only needs to meet one of them. For example, the reason why the 2nd sample point can become a clustering point is that it meets the condition that the number of sample points within the neighborhood of the 2nd sample point, which is 4, is equal to the preset threshold of 4.
[0067] Therefore, after determining that the 2nd sample point is a clustering point, the sample points within the neighborhood of the 2nd sample point can be added to the clustering cluster to which the 2nd sample point currently belongs. Specifically, the sample points within the neighborhood of the 2nd sample point include the 1st sample point, the 3rd sample point, the 5th sample point, and itself. The sample points within the clustering cluster to which the 2nd sample point currently belongs include the 1st sample point, the 2nd sample point, the 3rd sample point, and the 4th sample point. Therefore, after determining that the 2nd sample point is a clustering point, the current clustering cluster can be updated to a clustering cluster that includes the 1st sample point, the 2nd sample point, the 3rd sample point, the 4th sample point, and the 5th sample point.
[0068] After examining the 2nd sample point within the neighborhood of the 1st sample point for clustering points, continue to select any sample point within the neighborhood of the 1st sample point other than the 2nd sample point for the examination of clustering points. For example, the 3rd sample point can be selected for the examination of clustering points.
[0069] When selecting the 3rd sample point within the neighborhood of the 1st sample point for the examination of clustering points, it can be found that the sample points within the neighborhood of the 3rd sample point, besides itself, also include the 1st sample point, the 2nd sample point, the 5th sample point, the 6th sample point, the 7th sample point, and the 4th sample point. That is, the number of sample points within the neighborhood of the 3rd sample point is 7. When judging whether the 3rd sample point meets the requirements of a clustering point, it can be found that the 3rd sample point meets the requirements, as shown below: the number of sample points within the neighborhood of the 3rd sample point is 7, 7 is greater than the preset threshold of 4; and 7 is greater than the number of sample points within the clustering cluster to which the 3rd sample point belongs (here, the clustering cluster to which the 3rd sample point belongs refers to the clustering cluster where the 1st sample point, the 2nd sample point, the 3rd sample point, the 4th sample point, and the 5th sample point are located).
[0070] Therefore, after determining that the sample point No. 3 is a clustering point, the sample points within the neighborhood of the sample point No. 3 can be added to the clustering cluster to which the sample point No. 3 currently belongs. Specifically, the sample points within the neighborhood of the sample point No. 3 include the sample point No. 1, the sample point No. 2, the sample point No. 5, the sample point No. 6, the sample point No. 7, the sample point No. 4, and itself. The sample points in the clustering cluster to which the sample point No. 3 currently belongs include the sample point No. 1, the sample point No. 2, the sample point No. 3, the sample point No. 4, and the sample point No. 5. Therefore, after determining that the sample point No. 3 is a clustering point, the current clustering cluster can be updated to a clustering cluster that includes the sample point No. 1, the sample point No. 2, the sample point No. 3, the sample point No. 4, the sample point No. 5, the sample point No. 6, and the sample point No. 7.
[0071] The process of determining whether the remaining other sample points are clustering points will not be elaborated here.
[0072] In the above manner, by selecting the sample point No. 1 as the starting sample point, the finally obtained clustering cluster is a clustering cluster that includes the sample point No. 1, the sample point No. 2, the sample point No. 3, the sample point No. 4, the sample point No. 5, the sample point No. 6, the sample point No. 7, the sample point No. 8, and the sample point No. 9.
[0073] As a possible implementation manner, the evaluation objects are N travel record information; according to the attribute information of each evaluation object, each sample point corresponding to each evaluation object is constructed, including: for each travel record information, according to the starting position and the ending position in the travel record information, the sample point corresponding to the travel record information is constructed; determining the number of samples within the neighborhood of the sample point, including: determining the distance between any two sample points; determining the sample points whose distance from the sample point is within the neighborhood as the sample points within the neighborhood.
[0074] When a bank conducts a trade background review on customers in the transportation and logistics industries, the content of the review can be the tax information and invoice information of the customers, or the transportation route information. For example, when a customer in the transportation and logistics industries submits a loan request to the bank, the bank can require the customer to provide information about the transportation route. For example, if the customer claims that the main transportation route of its company is from place P in Beijing to place Q in Shanghai, the bank can require the customer to provide all travel record information between place P in Beijing and place Q in Shanghai within a preset time (such as within the most recent six months). However, in fact, most of the travel record information submitted by the customer to the bank is true travel record information from place P in Beijing to place Q in Shanghai, while a small part is not true travel record information from place P in Beijing to place Q in Shanghai. Therefore, the bank needs to perform data analysis on all the travel record information submitted by the customer to determine the normal travel record information and the abnormal travel record information.
[0075] For each trip record information submitted by the customer, the bank can obtain the starting location and ending location from the trip record information. For example, each trip record information includes the longitude and latitude (x1, y1) of the starting location and the longitude and latitude (x2, y2) of the ending location of the trip. By taking the starting location and ending location (x1, y1, x2, y2) of each trip record information as a sample point, all the sample points corresponding to the trip record information can be obtained. The longitude and latitude of the starting location and the longitude and latitude of the ending location can be obtained from in-vehicle GPS (Global Positioning System) data.
[0076] When clustering the sample points corresponding to all the trip record information, the number of sample points within the neighborhood of the sample point can be determined in the following way:
[0077] Determine the distance between any two sample points. The distance between any two sample points can be determined using the Euclidean distance formula.
[0078] Determine that the sample points whose distances from the sample point are within the neighborhood are the sample points within the neighborhood. Taking the sample point corresponding to a certain trip record information as the reference point, if there are another 3 sample points and the distances from these 3 sample points to the reference point are within the neighborhood, then these 3 sample points and the reference point can be regarded as the number of sample points within the neighborhood of the reference point. At this time, the number of sample points within the neighborhood of the sample point is 4.
[0079] As a possible implementation, according to the number of sample points in each clustering cluster, determine whether each clustering cluster is the cluster where normal sample points are located or the cluster where abnormal sample points are located, including: determining the clustering cluster with the largest number of sample points as the cluster where normal sample points are located; determining the clustering clusters other than the clustering cluster with the largest number of sample points as the clusters where abnormal sample points are located.
[0080] For example, the total number of trip record information of the transportation route from place P in Beijing to place Q in Shanghai within the preset time submitted by the customer to the bank is 112. Through the above clustering method, the number of clustering clusters finally obtained is 3, namely the 1st clustering cluster including 100 sample points, the 2nd clustering cluster including 5 sample points, and the 3rd clustering cluster including 7 sample points.
[0081] For the above clustering results, the following method can be used to determine whether each clustering cluster is the cluster where normal sample points are located or the cluster where abnormal sample points are located: Since the number of sample points in the 1st clustering cluster is the largest, the 1st clustering cluster can be determined as the cluster where normal sample points are located. That is, the 100 sample points in the 1st clustering cluster represent 100 repeated occurrences of the trip record from place P in Beijing to place Q in Shanghai. Compared with the 1st clustering cluster, the number of sample points in the 2nd clustering cluster and the 3rd clustering cluster is not the largest. Therefore, both the 2nd clustering cluster and the 3rd clustering cluster can be determined as the clusters where abnormal sample points are located. For example, the 5 sample points in the 2nd clustering cluster can represent 5 repeated occurrences of the trip record from place M in Suzhou to place N in Nanjing, and the 7 sample points in the 3rd clustering cluster can represent 7 repeated occurrences of the trip record from place O in Guangzhou to place K in Tianjin. The trip record information from place M in Suzhou to place N in Nanjing and from place O in Guangzhou to place K in Tianjin can be considered abnormal compared with the trip record from place P in Beijing to place Q in Shanghai.
[0082] Based on the same concept, an embodiment of the present invention further provides a clustering device for an evaluation object, as Figure 3 shown. The device includes:
[0083] A construction unit 301, configured to construct each sample point corresponding to each evaluation object according to the attribute information of each evaluation object;
[0084] A first determination unit 302, configured to determine the clustering cluster to which each sample point belongs. Among them, for any sample point, the clustering cluster to which it belongs is determined by the following method: Determine the number of sample points in the neighborhood of the sample point; if the number of sample points in the neighborhood meets the clustering point requirement, then determine the clustering cluster to which the sample point belongs as the clustering cluster to which the sample points in the neighborhood belong; the neighborhood is a set area range based on the sample point; the clustering point requirement is that the number of sample points in the neighborhood is not less than a preset threshold, or the number of sample points in the neighborhood is greater than the number of sample points in the clustering cluster to which the sample point belongs;
[0085] A second determination unit 303, configured to determine whether each clustering cluster is the cluster where normal sample points are located or the cluster where abnormal sample points are located according to the number of sample points in each clustering cluster.
[0086] Further, for the device, the first determination unit 302 is specifically configured to: if the sample point currently has no affiliated clustering cluster, use the sample points in the neighborhood as a clustering cluster; or if the number of sample points in the neighborhood is less than the preset threshold and the number of sample points in the neighborhood is not greater than the number of sample points in the clustering cluster to which the sample point belongs, select the next sample point in the clustering cluster to which the sample point belongs to determine whether it meets the requirements of a clustering point until it is determined whether any sample point in the clustering cluster to which the sample point belongs is a clustering point.
[0087] Further, for the device, each evaluation object is N trip record information; the construction unit 301 is specifically configured to: for each trip record information, construct a sample point corresponding to the trip record information according to the starting position and the ending position in the trip record information; the first determination unit 302 is specifically configured to: determine the distance between any two sample points; determine the sample points whose distances from the sample point are within the neighborhood as the sample points in the neighborhood.
[0088] Further, for the device, the second determination unit 303 is specifically configured to: determine the clustering cluster with the largest number of sample points as the cluster where the normal sample points are located; determine the clustering clusters other than the clustering cluster with the largest number of sample points as the clusters where the abnormal sample points are located.
[0089] An embodiment of the present invention provides a computing device, which may specifically be a desktop computer, a portable computer, a smart phone, a tablet computer, a personal digital assistant (PDA), etc. The computing device may include a central processing unit (CPU), a memory, an input / output device, etc. The input device may include a keyboard, a mouse, a touch screen, etc., and the output device may include a display device, such as a liquid crystal display (LCD), a cathode ray tube (CRT), etc.
[0090] The memory may include a read-only memory (ROM) and a random access memory (RAM), and provide program instructions and data stored in the memory to the processor. In the embodiment of the present invention, the memory may be used for the program instructions of the clustering method for the evaluation object;
[0091] The processor is configured to call the program instructions stored in the memory and execute the clustering method for the evaluation object according to the obtained program.
[0092] An embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a clustering method for an evaluation object.
[0093] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0094] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0095] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0097] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0098] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A clustering method for evaluation objects, characterized in that, Including: Constructing each sample point corresponding to each evaluation object according to the attribute information of each evaluation object; Determining the clustering cluster to which each sample point belongs; wherein, for any sample point, the clustering cluster to which it belongs is determined in the following manner: determining the number of sample points within the neighborhood of the sample point; if the number of sample points within the neighborhood meets the clustering point requirement, then determining the clustering cluster to which the sample point belongs as the clustering cluster to which the sample points within the neighborhood belong; the neighborhood is a set regional range based on the sample point; the clustering point requirement is that the number of sample points within the neighborhood is not less than a preset threshold, or the number of sample points within the neighborhood is greater than the number of sample points in the clustering cluster to which the sample point belongs; Determining, according to the number of sample points in each clustering cluster, whether each clustering cluster is a cluster where normal sample points are located or a cluster where abnormal sample points are located; If the sample point currently has no belonging clustering cluster, then taking the sample points within the neighborhood as a clustering cluster; or If the number of sample points within the neighborhood is less than the preset threshold, and the number of sample points within the neighborhood is not greater than the number of sample points in the clustering cluster to which the sample point belongs, then selecting the next sample point in the clustering cluster to which the sample point belongs for determining whether it meets the clustering point requirement until it is determined whether any sample point in the clustering cluster to which the sample point belongs is a clustering point; Each of the evaluation objects is N travel record information; Constructing each sample point corresponding to each evaluation object according to the attribute information of each evaluation object, including: For each travel record information, constructing the sample point corresponding to the travel record information according to the starting position and the ending position in the travel record information; Determining the number of sample points within the neighborhood of the sample point, including: Determining the distance between any two sample points; Determining the sample points whose distances from the sample point are within the neighborhood as the sample points within the neighborhood.
2. The method according to claim 1, wherein Including: Determining, according to the number of sample points in each clustering cluster, whether each clustering cluster is a cluster where normal sample points are located or a cluster where abnormal sample points are located, including: Determining the clustering cluster with the largest number of sample points as the cluster where normal sample points are located; Determining the clustering clusters other than the clustering cluster with the largest number of sample points as the clusters where abnormal sample points are located.
3. A clustering device for an evaluation object, characterized in that, Including: A construction unit, configured to construct each sample point corresponding to each evaluation object according to the attribute information of each evaluation object; A first determination unit, configured to determine the clustering cluster to which each sample point belongs; wherein, for any sample point, the clustering cluster to which it belongs is determined in the following manner: determining the number of sample points within the neighborhood of the sample point; if the number of sample points within the neighborhood meets the clustering point requirement, then determining the clustering cluster to which the sample point belongs as the clustering cluster to which the sample points within the neighborhood belong; the neighborhood is a set regional range based on the sample point; the clustering point requirement is that the number of sample points within the neighborhood is not less than a preset threshold, or the number of sample points within the neighborhood is greater than the number of sample points in the clustering cluster to which the sample point belongs; A second determination unit, configured to determine, according to the number of sample points in each clustering cluster, whether each clustering cluster is a cluster where normal sample points are located or a cluster where abnormal sample points are located; The first determination unit is further configured to: If the sample point does not currently belong to any clustering cluster, the sample points within the neighborhood are used as a clustering cluster; or If the number of sample points within the neighborhood is less than the preset threshold and the number of sample points within the neighborhood is not greater than the number of sample points in the clustering cluster to which the sample point belongs, the next sample point in the clustering cluster to which the sample point belongs is selected to determine whether it meets the requirements of a clustering point until it is determined whether any sample point in the clustering cluster to which the sample point belongs is a clustering point; Each of the evaluation objects is N trip record information; The construction unit is specifically configured to: for each trip record information, construct a sample point corresponding to the trip record information according to the starting position and the ending position in the trip record information; The first determination unit is specifically configured to: determine the distance between any two sample points; determine the sample points whose distances from the sample point are within the neighborhood as the sample points within the neighborhood.
4. The device according to claim 3, characterized in that, The second determination unit is specifically configured to: determine the clustering cluster with the largest number of sample points as the cluster where the normal sample points are located; determine the clustering clusters other than the clustering cluster with the largest number of sample points as the clusters where the abnormal sample points are located.
5. A computing device, characterized in that, Including: A memory for storing program instructions; A processor for calling the program instructions stored in the memory and executing the method according to claims 1 or 2 according to the obtained program.
6. A computer-readable storage medium, characterized in that, The storage medium stores computer-executable instructions for causing a computer to execute the method according to claims 1 or 2.
Citation Information
Patent Citations
Clustering method and device
CN110163280A