Method and apparatus for efficiently determining the closest neighbors of data points included in a dataset

By sorting and grouping data points by multiple dimensions and using matrix operations on smaller datasets, the method effectively addresses the complexity challenges in determining closest neighbors, achieving efficient and accurate results.

WO2025122165A1PCT designated stage expired Publication Date: 2025-06-12TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Patent Information

Application Number
PCT/US2023/083225
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing methods for determining the closest neighbors of data points in a dataset face challenges with high time and space complexities, particularly when dealing with large datasets.

Method used

The method involves sorting data points by multiple dimensions, grouping them, expanding group limits, and calculating distances using matrix operations on smaller datasets, thereby reducing complexity.

Benefits of technology

This approach allows for efficient determination of closest neighbors with lower time and space complexities compared to traditional methods, achieving a balance between computational efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2023083225_12062025_PF_FP_ABST
    Figure US2023083225_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A method to determine closest neighbors of data points is disclosed. The method includes sorting the data points with respect to a first dimension to generate a first sorted dataset and grouping the data points based on their order in the first sorted dataset to generate a first plurality of groups. The method further includes, for each group, generating an expanded group based on expanding limits of the group along the first dimension and determining a first set of distances between data points included in the group and the expanded group. The method further includes performing similar operations for a second dimension to determine a second set of distances for each group of a second plurality of groups. The method further includes determining a set of closest neighbors of a data point based on distances in the first set of distances and the second set of distances involving the data point.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR EFFICIENTLY DETERMINING THE CLOSESTNEIGHBORS OF DATA POINTS INCLUDED IN A DATASETTECHNICAL FIELD

[0001] Embodiments of the invention relate to the field of computer-implemented search algorithms, and more specifically, to efficiently determining the closest neighbors of data points included in a dataset.BACKGROUND

[0002] Data analytics refers to the process of analyzing datasets to extract useful information from them. A dataset may include multiple data points. Finding the n closest neighbors to each data point in a dataset is a common problem that needs to be solved during data analytics. The closest neighbors to a given data point are a set of points that have the shortest distances to the given data point. For example, the data points in a dataset may represent geographical coordinates in terms of longitude and latitude and it may be desirable to find the n closest neighbors to each data point in terms of geographical distance.

[0003] Time and space complexities are a major challenge when finding the n closest neighbors. Time complexity refers to the amount of computing time it takes to execute an algorithm. Time complexity may be measured in terms of the number of elementary operations needed to execute an algorithm. Space complexity refers to the amount of memory space needed to execute an algorithm. One approach to finding the n closest neighbors is to iterate through each data point included in the dataset and determine the distances between that data point and every other data point included in the dataset. However, while this “brute force” approach has a relatively small space complexity, it has a relatively large time complexity that increases exponentially with respect to the size of the dataset. Another approach to finding the n closest neighbors is to apply matrix operations to the entire dataset. However, while the use of matrix operations can significantly reduce the time complexity compared to the brute force approach, it significantly increases the space complexity.SUMMARY

[0004] A method is disclosed for determining closest neighbors of data points included in a dataset. Each data point included in the dataset may have at least two dimensions including a first dimension and a second dimension. The method includes sorting the data points included in the dataset with respect to the first dimension to generate a first sorted dataset and grouping the data points included in the dataset based on their order in the first sorted dataset to generate afirst plurality of groups. The method further includes, for each group in the first plurality of groups, generating an expanded group based on expanding limits of the group along the first dimension and determining a first set of distances between data points included in the group and data points included in the expanded group. The method further includes sorting the data points included in the dataset with respect to the second dimension to generate a second sorted dataset and grouping the data points included in the dataset based on their order in the second sorted dataset to generate a second plurality of groups. The method further includes, for each group in the second plurality of groups, generating an expanded group based on expanding limits of the group along the second dimension and determining a second set of distances between data points included in the group and data points included in the expanded group. The method further includes determining a set of closest neighbors of each of one or more data points included in the dataset based on distances in the first set of distances involving the data point and the second set of distances involving the data point.

[0005] A non-transitory machine-readable storage medium is disclosed that provides instructions that, if executed by a processor of a computing device, will cause the computing device to carry out operations for determining closest neighbors of data points included in a dataset. Each data point included in the dataset may have at least two dimensions including a first dimension and a second dimension. The operations include sorting the data points included in the dataset with respect to the first dimension to generate a first sorted dataset and grouping the data points included in the dataset based on their order in the first sorted dataset to generate a first plurality of groups. The operations further include, for each group in the first plurality of groups, generating an expanded group based on expanding limits of the group along the first dimension and determining a first set of distances between data points included in the group and data points included in the expanded group. The operations further include sorting the data points included in the dataset with respect to the second dimension to generate a second sorted dataset and grouping the data points included in the dataset based on their order in the second sorted dataset to generate a second plurality of groups. The operations further include, for each group in the second plurality of groups, generating an expanded group based on expanding limits of the group along the second dimension and determining a second set of distances between data points included in the group and data points included in the expanded group. The operations further include determining a set of closest neighbors of each of one or more data points included in the dataset based on distances in the first set of distances involving the data point and the second set of distances involving the data point.

[0006] A computing device to determine the closest neighbors of data points included in a dataset is disclosed. Each data point included in the dataset may have at least two dimensionsincluding a first dimension and a second dimension. The computing device includes a set of one or more processors and a non-transitory machine-readable storage medium that provides instructions that, if executed by the set of one or more processors, will cause the computing device to carry out operations for determining closest neighbors of data points included in a dataset. The operations include sorting the data points included in the dataset with respect to the first dimension to generate a first sorted dataset and grouping the data points included in the dataset based on their order in the first sorted dataset to generate a first plurality of groups. The operations further include, for each group in the first plurality of groups, generating an expanded group based on expanding limits of the group along the first dimension and determining a first set of distances between data points included in the group and data points included in the expanded group. The operations further include sorting the data points included in the dataset with respect to the second dimension to generate a second sorted dataset and grouping the data points included in the dataset based on their order in the second sorted dataset to generate a second plurality of groups. The operations further include, for each group in the second plurality of groups, generating an expanded group based on expanding limits of the group along the second dimension and determining a second set of distances between data points included in the group and data points included in the expanded group. The operations further include determining a set of closest neighbors of each of one or more data points included in the dataset based on distances in the first set of distances involving the data point and the second set of distances involving the data point.

[0007] Advantageously, embodiments are able to determine the closest neighbors of data points (or an approximation thereof) with lower time and / or space complexity. In particular, embodiments achieve this by splitting a sorted dataset into multiple smaller datasets, performing matrix operations on the smaller datasets, and merging the results of the matrix operations performed on the smaller datasets, thereby avoiding having to calculate distances between every pair of data points. Accordingly, using the disclosed techniques, the n closest neighbors can be identified with lower time complexity compared to a brute force approach and with lower space complexity compared to an approach that applies matrix operations to the entire dataset. For example, the brute force approach has a time complexity of O( / / 2), where n is the number of data points, whereas embodiments have a time complexity in the realm of O(k*l), where k is the size of a group of data points (a smaller dataset) and / is the size of the expanded group, where k and I are expected to be less than n. Embodiments may lower the space complexity compared by a similar order of magnitude compared to approaches that apply matrix operations on the entire dataset.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The invention may best be understood by referring to the following description and accompanying drawings that are used to illustrate embodiments of the invention. In the drawings:

[0009] Figure l is a flow diagram of a method for determining the closest neighbors of data points included in a dataset representing geographical coordinates, according to some embodiments.

[0010] Figure 2 is a diagram showing groupings of data points based on their order in a sorted dataset, according to some embodiments.

[0011] Figure 3 is a diagram showing the latitude limits of a group and an expanded group, according to some embodiments.

[0012] Figure 4 is a diagram showing the longitude limits of a group and an expanded group, according to some embodiments.

[0013] Figure 5 is a diagram showing an outlier identification method, according to some embodiments.

[0014] Figure 6 is a diagram showing a matrix of outlier data points and all data points, according to some embodiments.

[0015] Figure 7 is a flow diagram of a method for determining the closest neighbors of data points included in a dataset, according to some embodiments.

[0016] Figure 8 is a diagram showing connectivity between network devices (NDs) within an example network, as well as three example implementations of the NDs, according to some embodiments.DETAILED DESCRIPTION

[0017] The following description describes methods and apparatus for efficiently determining the closest neighbors of data points included in a dataset (or an acceptable approximation thereof). In the following description, numerous specific details such as logic implementations, opcodes, means to specify operands, resource partitioning / sharing / duplication implementations, types and interrelationships of system components, and logic partitioning / integration choices are set forth in order to provide a more thorough understanding of the present invention. It will be appreciated, however, by one skilled in the art that the invention may be practiced without such specific details. In other instances, control structures, gate level circuits and full software instruction sequences have not been shown in detail in order not to obscure the invention. Those of ordinary skill in the art, with the included descriptions, will be able to implement appropriate functionality without undue experimentation.

[0018] References in the specification to “one embodiment,” “an embodiment,” “an example embodiment,” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0019] Bracketed text and blocks with dashed borders (e.g., large dashes, small dashes, dotdash, and dots) may be used herein to illustrate optional operations that add additional features to embodiments of the invention. However, such notation should not be taken to mean that these are the only options or optional operations, and / or that blocks with solid borders are not optional in certain embodiments of the invention.

[0020] In the following description and claims, the terms “coupled” and “connected,” along with their derivatives, may be used. It should be understood that these terms are not intended as synonyms for each other. “Coupled” is used to indicate that two or more elements, which may or may not be in direct physical or electrical contact with each other, co-operate or interact with each other. “Connected” is used to indicate the establishment of communication between two or more elements that are coupled with each other.

[0021] An electronic device stores and transmits (internally and / or with other electronic devices over a network) code (which is composed of software instructions and which is sometimes referred to as computer program code or a computer program) and / or data using machine-readable media (also called computer-readable media), such as machine-readable storage media (e.g., magnetic disks, optical disks, solid state drives, read only memory (ROM), flash memory devices, phase change memory) and machine-readable transmission media (also called a carrier) (e.g., electrical, optical, radio, acoustical or other form of propagated signals - such as carrier waves, infrared signals). Thus, an electronic device (e.g., a computer) includes hardware and software, such as a set of one or more processors (e.g., wherein a processor is a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application specific integrated circuit, field programmable gate array, other electronic circuitry, a combination of one or more of the preceding) coupled to one or more machine-readable storage media to store code for execution on the set of processors and / or to store data. For instance, an electronic device may include non-volatile memory containing the code since the non-volatile memory can persist code / data even when the electronic device is turned off (when power is removed), and while the electronic device is turned on that part of the code that is to beexecuted by the processor(s) of that electronic device is typically copied from the slower nonvolatile memory into volatile memory (e.g., dynamic random access memory (DRAM), static random access memory (SRAM)) of that electronic device. Typical electronic devices also include a set of one or more physical network interface(s) (NI(s)) to establish network connections (to transmit and / or receive code and / or data using propagating signals) with other electronic devices. For example, the set of physical NIs (or the set of physical NI(s) in combination with the set of processors executing code) may perform any formatting, coding, or translating to allow the electronic device to send and receive data whether over a wired and / or a wireless connection. In some embodiments, a physical NI may comprise radio circuitry capable of receiving data from other electronic devices over a wireless connection and / or sending data out to other devices via a wireless connection. This radio circuitry may include transmitter(s), receiver(s), and / or transceiver s) suitable for radiofrequency communication. The radio circuitry may convert digital data into a radio signal having the appropriate parameters (e.g., frequency, timing, channel, bandwidth, etc.). The radio signal may then be transmitted via antennas to the appropriate recipient(s). In some embodiments, the set of physical NI(s) may comprise network interface controller(s) (NICs), also known as a network interface card, network adapter, or local area network (LAN) adapter. The NIC(s) may facilitate in connecting the electronic device to other electronic devices allowing them to communicate via wire through plugging in a cable to a physical port connected to a NIC. One or more parts of an embodiment of the invention may be implemented using different combinations of software, firmware, and / or hardware.

[0022] A network device (ND) is an electronic device that communicatively interconnects other electronic devices on the network (e.g., other network devices, end-user devices). Some network devices are “multiple services network devices” that provide support for multiple networking functions (e.g., routing, bridging, switching, Layer 2 aggregation, session border control, Quality of Service, and / or subscriber management), and / or provide support for multiple application services (e.g., data, voice, and video).

[0023] As mentioned above, finding the n closest neighbors of data points included in a dataset is a common problem that needs to be solved during data analytics. One approach to finding the n closest neighbors is to iterate through each data point included in the dataset and determine the distances between that data point and every other data point included in the dataset. However, while this “brute force” approach has a relatively small space complexity, it has a relatively large time complexity that increases exponentially with respect to the size of the dataset.Another approach to finding the n closest neighbors is to apply matrix operations to the entire dataset. However, while this use of matrix operations can significantly reduce the timecomplexity compared to the brute-force approach, it significantly increases the space complexity.

[0024] Embodiments are described herein that are able to determine the n closest neighbors of data points included in a dataset (or an acceptable approximation thereof) with lower time complexity compared to the brute force approach and with lower space complexity compared to approaches that apply matrix operations to the entire dataset. Embodiments are thus able to achieve a balance between time complexity and space complexity.

[0025] An embodiment is a method performed by one or more computing devices to determine the closest neighbors of data points included in a dataset. Each data point included in the dataset may have at least a first dimension and a second dimension. For example, the data points may represent geographical coordinates, where the first dimension represents latitude and the second dimension represents longitude. The method includes sorting the data points included in the dataset with respect to the first dimension to generate a first sorted dataset and grouping the data points included in the dataset based on their order in the first sorted dataset to generate a first plurality of groups. The method further includes, for each group in the first plurality of groups, generating an expanded group based on expanding limits of the group along the first dimension and determining a first set of distances between data points included in the group and data points included in the expanded group. The method further includes sorting the data points included in the dataset with respect to the second dimension to generate a second sorted dataset and grouping the data points included in the dataset based on their order in the second sorted dataset to generate a second plurality of groups. The method further includes, for each group in the second plurality of groups, generating an expanded group based on expanding limits of the group along the second dimension and determining a second set of distances between data points included in the group and data points included in the expanded group. The method further includes determining closest neighbors of each of one or more data points included in the dataset based on distances in the first set of distances involving the data point and the second set of distances involving the data point. In an embodiment, the method further includes identifying one or more data points included in the dataset that are outliers, determining a third set of distances between each of the one or more data points identified as being outliers and all other data points included in the dataset, and determining closest neighbors for each of the one or more data points identified as being outliers based on the third set of distances.

[0026] Embodiments are able to determine the closest neighbors (or an approximation thereof that is acceptable for most use cases) with lower time complexity and / or space complexity compared to existing approaches. For example, the brute force approach to finding the n closest neighbors mentioned above has a time complexity of O( / / 2) (that is, the time to run the algorithmincreases exponentially as the input size grows). Embodiments have a lower time complexity. For example, the time complexity of processing each group may be ()(&* / ), where k is the size of the group and I is the size of the expanded group (k and I are less than ri). The time complexity of processing the outliers may be O(n*m), where n is the number of data points and m is the number of outliers (m is expected to be much less than ri). The time complexity is expressed herein in Big-0 notation. As is known in the art, Big-0 notation is a tool used to describe the time complexity of algorithms in terms of the input size. Also, embodiments improve the space complexity compared to an approach that applies matrix operations to the entire dataset by the same / similar order of magnitude. Embodiments are able to achieve a balance between time complexity and space complexity by sorting data points included in a dataset along multiple dimensions, splitting the dataset into multiple smaller datasets, performing matrix operations on the smaller datasets, and merging the results of the matrix operations. Also, embodiments may perform an outlier analysis to identify outlier data points to further improve accuracy. Embodiments are further described herein with reference to the accompanying figures.

[0027] Figure l is a flow diagram of a method for determining the closest neighbors of data points included in a dataset representing geographical coordinates, according to some embodiments. In an embodiment, the method is performed by one or more computing devices.

[0028] As shown in the diagram, at operation 105, the one or more computing devices may load a dataset of data points representing geographical coordinates on the Earth. Each data point may have a latitude and longitude.

[0029] At operation 110A, the computing device may sort the data points included in the dataset based on latitude to generate a latitude-sorted dataset. In an embodiment, the computing device uses a merge sort algorithm to sort the data points included in the dataset based on latitude, although it will be appreciated that other types of sorting algorithms can be used for this purpose.

[0030] At operation 115A, the one or more computing devices may group the data points included in the dataset based on their order in the latitude-sorted dataset. In an embodiment, a maximum size is defined for the grouping (the maximum number of data points that can be included in a single group). In an embodiment, the maximum size of each group, k, is determined based on the amount of available system memory (volatile memory). The value of k can be configured to achieve a balance between time complexity and space complexity. In general, increasing k increases the space requirement (more volatile memory is needed) but decreases the computation time, while decreasing k decreases the space requirement (less volatile memory is needed) but increases the computation time. In an embodiment, k is configured such that A2* 10 data points can be stored within the allocated / available volatilememory of the one or more computing devices (the computing device(s) executing the method) at the same time. This has been found to achieve a good balance between time complexity and space complexity. As an example, assuming there are m total data points, the one or more computing devices may group the first k data points included in the latitude- sorted dataset into a first group, group the next k data points included in the latitude- sorted dataset into a second group, and so on. The last group may have a size of m-(n- \ )k data points, where m is the total number of data points, n is the number of groups, and k is the maximum group size.

[0031] The one or more computing devices may perform operations 120 A and 125 A for each group of data points generated as part of operation 115A.

[0032] At operation 120 A, the one or more computing devices may generate an expanded group based on expanding the limits of the group north and south. Each group may have a southern limit and a northern limit. The southern limit may correspond to the latitude of the point having the lowest (southernmost) latitude in the group and the northern limit may correspond to the latitude of the point having the highest latitude in the group. The one or more computing devices may generate the expanded group based on expanding the northern limit of the group north by a predefined distance and expanding the southern limit of the group south by the predefined distance. For example, the expanded group may be generated based on expanding the northern limit of the group north by 20 miles and expanding the southern limit of the group south by 20 miles. Any points having a latitude within those limits may be included in the expanded group. The predefined distance may correspond to a certain amount of latitude (e.g., 20 miles corresponds to approximately 17 minutes in latitude). Thus, the one or more computing devices may determine the latitudes of the expanded limits based on converting the predefined distance to an amount of latitude, adding the this amount of latitude to the northern limit of the group, and subtracting this amount of latitude from the southern limit of the group.

[0033] At operation 125 A, the one or more computing devices may determine a first set of distances between data points included in the group and data points included in the expanded group. The one or more computing devices may determine the second set of distances by generating a matrix of data points and performing matrix operations on the matrix to determine the distances between the data points. In an embodiment, the matrix operations involve [vectorized calculations. Vectorization refers to the process of converting an algorithm from operating on a single value at a time to operating on a set of values (vector) at one time. The rows of the matrix may represent data points included in the group and the columns of the matrix may represent data points included in the expanded group. The matrix may thus have a size of k *j, where k is the number of data points included in the (original / non-expanded) group and j is the number of data points included in the expanded group. The distances between a pairof data points may be defined using any suitable distance function for determining the distance between two geographical coordinates. For example, the distance function may be a Haversine distance function or similar function. The Haversine distance function calculates the great-circle distance between two points on a sphere (e.g., the Earth). It assumes a perfect sphere, which provides a good approximation for short distances. As another example, the distance function may be a Euclidian distance function.

[0034] Performing operations 120A and 125A for each group can be visualized as a diagonalization of a larger matrix but with some overlaps. The diagonalization effectively expands the size of the larger matrix to N for the row and AT for the column, albeit with some overlap, where N is the total number of data points and AT is the number of data points included in the expanded scopes / limits.

[0035] Similar operations as operations 110A-125A may be performed for longitude, as will be further described herein below. At operation HOB, the one or more computing devices may sort the data points included in the dataset based on longitude to generate a longitude-sorted dataset. In an embodiment, the computing device uses a merge sort algorithm to sort the data point included in the dataset based on longitude, although it will be appreciated that other types of sorting algorithms can be used for this purpose.

[0036] At operation 115B, the one or more computing devices may group the data points included in the dataset based on their order in the longitude- sorted dataset. In an embodiment, a maximum size is defined for the grouping (the maximum number of data points that can be included in a single group). In an embodiment, the maximum size of each group, k, is determined based on the amount of available system memory (volatile memory). For example, as mentioned above, k may be configured such that A2* 10 data points can be stored within the available / allocated volatile memory of the one or more computing devices (the computing device(s) executing the method) at the same time. As an example, assuming there are m total data points, the one or more computing devices may group the first k data points included in the longitude- sorted dataset into a first group, group the next k data points included in the longitude- sorted dataset into a second group, and so on. The last group may have a size of m-(n- \ )k data points, where m is the total number of data points, n is the number of groups, and k is the maximum group size.

[0037] The one or more computing devices may perform operations 120B and 125B for each group of data points generated as part of operation 115B.

[0038] At operation 120B, the one or more computing devices may generate an expanded group based on expanding the limits of the group west and east. Each group may have a western limit and an eastern limit. The western limit may correspond to the longitude of the pointhaving the lowest (westmost) latitude in the group and the eastern limit may correspond to the longitude of the point having the highest longitude in the group. The one or more computing devices may generate the expanded group based on expanding the western limit of the group west by an amount of longitude and expanding the eastern limit of the group east by the amount of longitude. For example, the expanded group may be generated based on expanding the western limit of the group west by 2 degrees in longitude and expanding the eastern limit of the group east by 2 degrees in longitude. Any points having a longitude within those limits may be included in the expanded group.

[0039] The groups that were formed based on the latitude-sorted dataset were expanded north and south by a predefined distance to generate the expanded groups. However, the groups that were formed based on the longitude- sorted datasets may not be expanded west and east in the same way. Due to the curvature of the Earth, the longitude that is a predefined distance (e.g., 20 miles) east or west from a given longitude at one latitude is different from the longitude that is a predefined distance east or west from the given longitude at a different latitude. Longitude lines are furthest apart near the equator and converge near the poles so the distance covered by a given amount of longitude is different depending on the latitude. For example, in the northern hemisphere of the Earth, a longitude that is 20 miles east of a longitudinal line at latitude 40 degrees north is higher than a longitude that is 20 miles east of the same longitudinal line at latitude 35 degrees north. Thus, to take this into consideration, in an embodiment, the one or more computing devices determine the amount of longitude to use for expanding a group (the amount by which the western limit and eastern limit of the group is expanded) based on identifying the data point included in the group that has the greatest absolute value in terms of latitude within the group (i.e., the data point closest to the poles) and determining the amount of longitude that corresponds to a predefined distance (e.g., 20 miles) at the latitude of the identified data point.

[0040] At operation 125B, the one or more computing devices may determine a second set of distances between data points included in the group and data points included in the expanded group. The one or more computing devices may determine the second set of distances by generating a matrix of data points and performing matrix operations on the matrix. In an embodiment, the matrix operations involve vectorized calculations. The rows of the matrix may represent data points included in the group and the columns of the matrix may represent data points included in the expanded group. The matrix may thus have a size of k *j, where k is the number of data points included in the (original / non-expanded) group and j is the number of data points included in the expanded group. The distances between a pair of data points may bedefined using any suitable distance function for determining the distance between two geographical coordinates.

[0041] Performing operations 120B and 125B for each group can be visualized as a diagonalization of a larger matrix but with some overlaps. The diagonalization effectively expands the size of the larger matrix to N for the row and M for the column, albeit with some overlap, where N is the total number of data points and M is the number of data points included in the expanded scopes.

[0042] A given data point may belong to two groups: (1) a first group, which is one of the groups formed based on the latitude- sorted dataset; and (2) a second group, which is one of the groups formed based on the longitude-sorted dataset. Thus, the first set of distances determined for the first group (determined at operation 125A) may include distances between the given data point and every other data point included in the first group. Also, the second set of distances determined for the second group (determined at operation 125B) may include distances between the given data point and every other data point included in the second group.

[0043] At operation 130, the one or more computing devices may identify data points that are outliers. In an embodiment, a data point is identified as being an outlier if the shortest distance in the first set of distances and the second set of distances involving the data point is longer than a predefined outlier threshold distance. The predefined outlier threshold distance may be configurable. In an embodiment, the predefined outlier threshold distance is set to be the same as the predefined distance by which the groups are expanded north and south (in operation 120A) and / or expanded west and east (in operation 120B).

[0044] At operation 135, the one or more computing devices may determine a third set of distances between each outlier data point and all data points included in the dataset. The one or more computing devices may determine the third set of distances by generating a matrix of the outlier data points and all data points and performing matrix operations on the matrix.

[0045] The one or more computing devices may perform operation 140 for each data point of interest (the data points for which the closest neighbors are desired). At operation 140, the one or more computing devices may determine whether the data point is an outlier (e.g., based on the result of operation 130). If the data point is not determined to be an outlier, at operation 145, the one or more computing devices may determine the closest neighbors of the data point based on the distances in the first set of distances and the second set of distances involving the data point. For example, the one or more computing devices may determine that the n shortest distances in the first set of distances and the second set of distances involving a given data point and designate the other data points involved in the determined distances as being the closestneighbors of the given data point. If more than one data point has the same distance from the given data point, one of them may be chosen at random.

[0046] Returning to operation 140, if the data point is determined to be an outlier (e.g., based on the result of operation 130), at operation 150, the one or more computing devices may determine the closest neighbors of the data point based on the third set of distances.

[0047] Thus, embodiments may only use a brute force approach for outlier data points, which helps improve the accuracy of the result while conserving resources (e.g., computational and storage resources) compared to a full brute force approach. The impact on system resources may vary depending on the number of outliers, but the number of outliers is expected to be small relative to the total number of data points in the dataset (and the number of outliers can be controlled by adjusting the predefined outlier threshold distance).

[0048] In an embodiment, the one or more computing devices apply special processing to data points near the poles. For example, the one or more computing devices may identify one or more data points that have a latitude that is greater than a predefined north pole limit latitude as being north pole data points and identify one or more data points that have a latitude that is less than a predefined south pole limit latitude as being south pole data points. In an embodiment, the predefined north pole limit latitude and the predefined south pole limit latitude are set based on the predefined distance that is used for expanding the limits of groups. For example, if the predefined distance that is used for expanding the limits of groups is 20 miles, then the predefined north pole limit latitude may be approximately 89 degrees 40 minutes North (which is ~20 miles away from the north pole) and the predefined north pole limit latitude may be approximately 89 degrees 40 minutes South (which is ~20 miles away from the south pole). The one or more computing devices may then determine a fourth set of distances between each of the north pole data points and determine a fifth set of distances between each of the south pole data points. The one or more computing devices may then determine the closest neighbors of each of one or more of the north pole data points based on the fourth set of distances and determine the closest neighbors of each of one or more of the south pole data points based on the fifth set of distances (e.g., based on performing matrix operations on a matrix of the data points). If the north pole data points and the south pole data points receive special processing, they may be excluded from the operations shown in Figure 1.

[0049] Embodiments are able to achieve a balance between time complexity and space complexity by working with smaller datasets and leveraging matrix operations. By considering both latitude and longitude, as well as separately addressing outliers, embodiments provide a comprehensive approach to solving the problem of determining the closest neighbors of data points included in a dataset (or an acceptable approximation thereof).

[0050] Figure 2 is a diagram showing groupings of data points based on their order in a sorted dataset, according to some embodiments. In the example shown in the diagram, it is assumed that there is a total of m data points that are to be split into groups. The data points are sorted based on latitude or longitude. The first k data points (in the sorted order) are added to the first group, the next k data points are added to the second group, and so on. The last group (group n in this example) will have m - (n-V)k data points. The data points included in the dataset may be grouped based on their order in the latitude- sorted dataset to generate a first plurality of groups and also may be grouped based on their order in the longitude- sorted dataset to generate a second plurality of groups. As a result, each data point may belong to two groups: (1) one group in the first plurality of groups; and (2) one group in the second plurality of groups.

[0051] Figure 3 is a diagram showing the latitude limits of a group and an expanded group, according to some embodiments. As shown in the diagram, the data points included in a given group in the first plurality of groups (the groups that were generated based on the order of data points in the latitude- sorted dataset) may include k data points that lie on a latitude strip 330 on the Earth 310. It is assumed that the data point included in the group that has the lowest latitude (southernmost latitude) has a latitude of Pi and that the point included in the group that has the highest latitude (northernmost latitude) has a latitude of Pk. Thus, the original southern limit of the group is Pi and the original northern limit of the group is Pk. The southern limit and the northern limit may be referred to as latitude limits. An expanded group may be generated based on expanding the latitude limits of the group north and south by a predefined distance. For example, the expanded group may include all points that have a latitude between Pi - y and Pk + y. Thus, the southern limit of the expanded group is Pi -y and the northern limit of the expanded group is Pk + y. The value of y may be an amount of latitude that corresponds to the predefined distance (e.g., 20 miles). The expanded group may include j data points, where j is greater than or equal to k.

[0052] A matrix 340 may be generated, where the rows of the matrix 340 represent the data points included in the group and the columns of the matrix 340 represent the data points included in the expanded group (or vice versa). That is, a k x j matrix of data points 340 may be generated. Matrix operations may be performed on the matrix 340 to determine the set of distances between data points included in the group and the data points included in the expanded group.

[0053] For sake of simplicity, the diagram shows the strip 330 associated with a single group. It should be appreciated that there can be additional non-overlapping strips on the Earth 310 associated with additional groups. An expanded group may be generated for each group by expanding the limits of the group north and south, in a similar manner as described above. Also,a matrix may be generated for each group and matrix operations may be performed on the matrix to determine the set of distances between the data points included in the group and the data points included in the expanded group, in a similar manner as described above.

[0054] In an embodiment, the data points near the poles of the Earth 310 receive special processing. For example, the data points that lie within a north pole strip 315 of the Earth 310 (referred to as north pole data points) may form their own group and the data points that lie within a south pole strip 320 of the Earth 310 (referred to as south pole data points) may form their own group. In an embodiment, any data points having a latitude that is greater than a predefined north pole limit latitude is designated as being a north pole data point. Also, any datapoints having a latitude that is less than a predefined south pole limit latitude is designated as being a south pole data point. In an embodiment, the predefined north pole limit latitude and the predefined south pole limit latitude are set based on the predefined distance that is used for expanding the limits of groups. For example, if the predefined distance that is used for expanding the limits of groups is 20 miles, then the predefined north pole limit latitude may be approximately 89 degrees 40 minutes North (which is ~20 miles away from the north pole) and the predefined north pole limit latitude may be approximately 89 degrees 40 minutes South (which is ~20 miles away from the south pole). The distances between the north pole data points may be determined and the distances between the south pole data points may be determined without performing group expansion.

[0055] Figure 4 is a diagram showing the longitude limits of a group and an expanded group, according to some embodiments. As shown in the diagram, the data points included in a given group in the second plurality of groups (the groups that were generated based on the order of data points in the longitude-sorted dataset) may include k data points that lie on a longitude strip 430 on the Earth 310. It is assumed that the data point included in the group that has the lowest longitude (westernmost longitude) has a latitude of Qi and that the point included in the group that has the highest longitude (easternmost longitude) has a latitude of Qk. Thus, the original western limit of the group is Qi and the original eastern limit of the group is Qk. The western limit and the eastern limit may be referred to as longitude limits. An expanded group may be generated based on expanding the longitude limits of the group west and east by a predefined amount of longitude. For example, the expanded group may include all points that have a latitude between Qi - x and Qk + x. Thus, the western limit of the expanded group is Qi - x and the eastern limit of the expanded group is Qk + x. The value of x may be an amount of latitude that corresponds to the predefined distance (e.g., 20 miles) at the latitude of the data point having the highest latitude (in terms of absolute value) in the group. Thus, as shown in the upper-right side of the diagram, the distance between the longitude limits of the (original / non-expanded) group and the longitude limits of the expanded group may be the predefined distance (20 miles) at a higher latitude but the distance between longitude limits of the (original / non-expanded) group and the longitude limits of the expanded group may be greater than the predefined distance (> 20 miles) at a lower latitude (assuming that the lower latitude is near the equator). The expanded group may include j data points, where j is greater than or equal to k.

[0056] A matrix 440 may be generated, where the rows of the matrix 440 represent the data points included in the group and the columns of the matrix 440 represent the data points included in the expanded group (or vice versa). That is, a k x j matrix of data points 440 may be generated. Matrix operations may be performed on the matrix 440 to determine the set of distances between data points included in the group and the data points included in the expanded group.

[0057] For sake of simplicity, the diagram shows the strip 430 associated with a single group. It should be appreciated that there can be additional non-overlapping strips on the Earth 310 associated with additional groups. An expanded group may be generated for each group by expanding the limits of the group west and east, in a similar manner as described above. Also, a matrix may be generated for each group and matrix operations may be performed on the matrix to determine the set of distances between the data points included in the group and the data points included in the expanded group, in a similar manner as described above.

[0058] In an embodiment, the data points near the poles of the Earth 310 may be excluded from the groups that are formed based on the longitude- sorted dataset. For example, the data points that lie within a north pole strip 315 of the Earth 310 (the north pole data points) and the data points that lie within a south pole strip 320 of the Earth 310 (the south pole data points) may be excluded from the groups that are formed based on the longitude- sorted dataset. These data points may receive special processing, as described above (by being grouped into their own respective groups (group of north pole data points and group of south pole data points)).

[0059] Figure 5 is a diagram showing an outlier identification method, according to some embodiments. The outlier identification method may be performed by one or more computing devices to determine whether a given data point is considered to be an outlier. The method is initiated by the one or more computing devices obtaining the distances involving the data point 510. The distances involving the data point 510 may include a first set of distances involving the data point from the latitude-based groupings 520 and a second set of distances involving the data point from the longitude-based groupings 530. At operation 540, the one or more computing devices may determine the shortest distance in the distances involving the data point. At operation 550, the one or more computing devices may determine whether the shortestdistance is longer than a predefined outlier threshold distance. If the shortest distance is not longer than the predefined outlier threshold distance, then at operation 560, the one or more computing devices determine that the data point is not an outlier data point. Otherwise, if the shortest distance is longer than the outlier threshold distance, then at operation 570, the one or more computing devices determine that the data point is an outlier.

[0060] The diagram shows one example way to identify outliers. It should be appreciated that outliers can be identified in other ways. For example, a data point may be identified as being an outlier if the r shortest distances involving the data point are all longer than the predefined outlier threshold distance, where r is a predefined threshold number.

[0061] Figure 6 is a diagram showing a matrix of outlier data points and all data points, according to some embodiments.

[0062] As shown in the diagram, a matrix 610 can be generated, where the rows of the matrix 610 represent outlier data points and the columns of the matrix represent all data points. The matrix 610 may be a / ? x m matrix, where p is the number of outlier data points and m is the number of total data points. Matrix operations may be performed on the matrix 610 to determine the set of distances between the outlier data points and all other data points.

[0063] Embodiments have been described in a context where the data points represent geographical coordinates having two dimensions (latitude and longitude). It should be appreciated that embodiments can be used / applied in contexts where the data points have more than two dimensions (e.g., a third dimension representing elevation can be added to latitude and longitude). Also, embodiments can be applied in contexts where the data points represent other types of data (e.g., any type of data that can be represented using values along multiple dimensions).

[0064] Embodiments can be used for various use cases. For example, embodiments can be used in Geographic Information Systems (GIS) applications. GIS applications often deal with large datasets of geographical information. GIS applications may use embodiments to efficiently identify the closest locations / objects to various locations / objects on a map.

[0065] As another example, embodiments can be used in transportation / logistics use cases. For example, embodiments can be used to perform tasks such as finding the nearest service locations, tracking assets, and / or optimizing routing for vehicles.

[0066] As another example, embodiments can be used in social media platforms. For example, social media platforms, and online communities in general, can use embodiments to find the closest connections or recommendations for users, help with personalized content delivery, provide friend recommendations, and more.

[0067] As another example, embodiments can be used in Internet of Things (loT) networks. In loT networks, loT devices need to discover and communicate with nearby sensors or actuators. loT networks may use embodiments to efficiently identify the closest network elements that can help loT devices interact with their surroundings.

[0068] As another example, embodiments can be used for autonomous vehicle use cases. Autonomous vehicles rely on being able to identify the closest objects, intersections, and / or other vehicles to make real-time navigation and safety decisions. Embodiments disclosed herein can be used to do this efficiently.

[0069] As another example, embodiments can be used in a telecommunications network. In large-scale telecommunication networks, identifying the closest cell tower or access point can be used for improving network performance, optimizing handovers, and enhancing quality of service (QoS). A telecommunications network may use embodiments disclosed herein to do this efficiently.

[0070] Figure 7 is a flow diagram of a method for determining the closest neighbors of data points included in a dataset (or an acceptable approximation thereof), according to some embodiments. In an embodiment, the method is performed by one or more computing devices. Each data point included in the dataset may have at least two dimensions including a first dimension and a second dimension.

[0071] The operations in the flow diagram will be described with reference to the examples embodiments of the other figures. However, it should be understood that the operations of the flow diagram can be performed by embodiments other than those discussed with reference to the other figures, and the embodiments discussed with reference to these other figures can perform operations different than those discussed with reference to the flow diagram.

[0072] Also, while the flow diagrams in the figures show a particular order of operations performed by certain embodiments, it should be understood that such order is provided by way of example and not intended to be limiting (e.g., alternative embodiments may perform the operations in a different order, combine certain operations, overlap certain operations, etc.).

[0073] At operation 705, the one or more computing devices sort the data points included in the dataset with respect to the first dimension to generate a first sorted dataset.

[0074] At operation 710, the one or more computing devices group the data points included in the dataset based on their order in the first sorted dataset to generate a first plurality of groups.

[0075] At operation 715, for each group in the first plurality of groups, the one or more computing devices generate an expanded group based on expanding limits of the group along the first dimension and determine a first set of distances between data points included in the group and data points included in the expanded group.

[0076] At operation 720, the one or more computing devices sort the data points included in the dataset with respect to the second dimension to generate a second sorted dataset.

[0077] At operation 725, the one or more computing devices group the data points included in the dataset based on their order in the second sorted dataset to generate a second plurality of groups. In an embodiment, each group in the first plurality of groups and the second plurality of groups has a size that is equal to or less than k, wherein k is configured such that A2* 10 data points can be stored within volatile memory of the one or more computing devices at a same time.

[0078] At operation 730, for each group in the second plurality of groups, the one or more computing devices generate an expanded group based on expanding limits of the group along the second dimension and determine a second set of distances between data points included in the group and data points included in the expanded group.

[0079] At operation 735, the one or more computing devices determine a set of closest neighbors of each of one or more data points included in the dataset based on distances in the first set of distances involving the data point and the second set of distances involving the data point. The set of closest neighbors of a given data point is a set of data points that are determined to have the shortest distances to the given data point.

[0080] In an embodiment, at operation 740, the one or more computing devices identify one or more data points included in the dataset that are outliers. In an embodiment, a data point is identified as being an outlier if a shortest distance in the first set of distances and the second set of distances involving the data point is longer than a predefined threshold distance. At operation 745, the one or more computing devices determine a third set of distances between each of the one or more data points identified as being outliers and all other data points included in the dataset. At operation 750, the one or more computing devices determines a set of closest neighbors of each of the one or more data points identified as being outliers based on the third set of distances.

[0081] In an embodiment, the data points included in the dataset represent geographical coordinates, wherein the first dimension represents latitude and the second dimension represents longitude. In an embodiment, the expanded group for a group in the first plurality of groups is generated based on expanding a northern limit of the group north by a predefined distance and expanding a southern limit of the group south by the predefined distance. In an embodiment, the expanded group for a group in the second plurality of groups is generated based on expanding a western limit of the group east by an amount of longitude and expanding an eastern limit of the group west by the amount of longitude. In an embodiment, the amount of longitude is determined based on identifying a data point included in the group that has a greatest absolutevalue in terms of latitude within the group and determining an amount of longitude corresponding to a predefined distance at the latitude of the identified data point.

[0082] In an embodiment, the one or more computing devices identify one or more data points that have a latitude that is greater than a predefined north pole limit latitude as being north pole data points, determines a fourth set of distances between each of the north pole data points, determines a set of closest neighbors of each of one or more of the north pole data points based on the fourth set of distances, identifies one or more data points that have a latitude that is less than a predefined south pole limit latitude as being south pole data points, determines a fifth set of distances between each of the south pole data points, and determines a set of closest neighbors of each of one or more of the south pole data points based on the fifth set of distances.

[0083] Figure 8 illustrates connectivity between network devices (NDs) within an example network, as well as three example implementations of the NDs, according to some embodiments of the invention. Figure 8 shows NDs 800A-H, and their connectivity by way of lines between 800A-800B, 800B-800C, 800C-800D, 800D-800E, 800E-800F, 800F-800G, and 800A- 800G, as well as between 800H and each of 800A, 800C, 800D, and 800G. These NDs are physical devices, and the connectivity between these NDs can be wireless or wired (often referred to as a link). An additional line extending from NDs 800A, 800E, and 800F illustrates that these NDs act as ingress and egress points for the network (and thus, these NDs are sometimes referred to as edge NDs; while the other NDs may be called core NDs).

[0084] Two of the example ND implementations in Figure 8 are: 1) a special-purpose network device 802 that uses custom application-specific integrated-circuits (ASICs) and a specialpurpose operating system (OS); and 2) a general purpose network device 804 that uses common off-the-shelf (COTS) processors and a standard OS.

[0085] The special -purpose network device 802 includes networking hardware 810 comprising a set of one or more processor(s) 812, forwarding resource(s) 814 (which typically include one or more ASICs and / or network processors), and physical network interfaces (NIs) 816 (through which network connections are made, such as those shown by the connectivity between NDs 800A-H), as well as non-transitory machine readable storage media 818 having stored therein networking software 820. During operation, the networking software 820 may be executed by the networking hardware 810 to instantiate a set of one or more networking software instance(s) 822. Each of the networking software instance(s) 822, and that part of the networking hardware 810 that executes that network software instance (be it hardware dedicated to that networking software instance and / or time slices of hardware temporally shared by that networking software instance with others of the networking software instance(s) 822), form a separate virtual network element 830A-R. Each of the virtual network element(s) (VNEs) 830 A-R includes a control communication and configuration module 832A-R (sometimes referred to as a local control module or control communication module) and forwarding table(s) 834A-R, such that a given virtual network element (e.g., 830 A) includes the control communication and configuration module (e.g., 832A), a set of one or more forwarding table(s) (e.g., 834A), and that portion of the networking hardware 810 that executes the virtual network element (e.g., 830A).

[0086] In an embodiment, software 820 includes code such as closest neighbor component 823, which when executed by networking hardware 810, causes the special-purpose network device 802 to perform operations of one or more embodiments disclosed herein as part of networking software instances 822 (e.g., operations to determine the closest neighbors of data points included in a dataset).

[0087] The special-purpose network device 802 is often physically and / or logically considered to include: 1) a ND control plane 824 (sometimes referred to as a control plane) comprising the processor(s) 812 that execute the control communication and configuration module(s) 832A-R; and 2) a ND forwarding plane 826 (sometimes referred to as a forwarding plane, a data plane, or a media plane) comprising the forwarding resource(s) 814 that utilize the forwarding table(s) 834A-R and the physical NIs 816. By way of example, where the ND is a router (or is implementing routing functionality), the ND control plane 824 (the processor(s) 812 executing the control communication and configuration module(s) 832A-R) is typically responsible for participating in controlling how data (e.g., packets) is to be routed (e.g., the next hop for the data and the outgoing physical NI for that data) and storing that routing information in the forwarding table(s) 834A-R, and the ND forwarding plane 826 is responsible for receiving that data on the physical NIs 816 and forwarding that data out the appropriate ones of the physical NIs 816 based on the forwarding table(s) 834A-R.

[0088] As shown in Figure 8, the general purpose network device 804 includes hardware 840 comprising a set of one or more processor(s) 842 (which are often COTS processors) and physical NIs 846, as well as non-transitory machine readable storage media 848 having stored therein software 850. During operation, the processor(s) 842 execute the software 850 to instantiate one or more sets of one or more applications 864A-R. While one embodiment does not implement virtualization, alternative embodiments may use different forms of virtualization. For example, in one such alternative embodiment the virtualization layer 854 represents the kernel of an operating system (or a shim executing on a base operating system) that allows for the creation of multiple instances 862A-R called software containers that may each be used to execute one (or more) of the sets of applications 864A-R; where the multiple software containers (also called virtualization engines, virtual private servers, or jails) are user spaces(typically a virtual memory space) that are separate from each other and separate from the kernel space in which the operating system is run; and where the set of applications running in a given user space, unless explicitly allowed, cannot access the memory of the other processes. In another such alternative embodiment the virtualization layer 854 represents a hypervisor (sometimes referred to as a virtual machine monitor (VMM)) or a hypervisor executing on top of a host operating system, and each of the sets of applications 864A-R is run on top of a guest operating system within an instance 862A-R called a virtual machine (which may in some cases be considered a tightly isolated form of software container) that is run on top of the hypervisor - the guest operating system and application may not know they are running on a virtual machine as opposed to running on a “bare metal” host electronic device, or through para-virtualization the operating system and / or application may be aware of the presence of virtualization for optimization purposes. In yet other alternative embodiments, one, some or all of the applications are implemented as unikernel(s), which can be generated by compiling directly with an application only a limited set of libraries (e.g., from a library operating system (LibOS) including drivers / libraries of OS services) that provide the particular OS services needed by the application. As a unikernel can be implemented to run directly on hardware 840, directly on a hypervisor (in which case the unikernel is sometimes described as running within a LibOS virtual machine), or in a software container, embodiments can be implemented fully with unikernels running directly on a hypervisor represented by virtualization layer 854, unikernels running within software containers represented by instances 862A-R, or as a combination of unikernels and the above-described techniques (e.g., unikernels and virtual machines both run directly on a hypervisor, unikernels and sets of applications that are run in different software containers).

[0089] The instantiation of the one or more sets of one or more applications 864A-R, as well as virtualization if implemented, are collectively referred to as software instance(s) 852. Each set of applications 864A-R, corresponding virtualization construct (e.g., instance 862A-R) if implemented, and that part of the hardware 840 that executes them (be it hardware dedicated to that execution and / or time slices of hardware temporally shared), forms a separate virtual network element(s) 860A-R.

[0090] The virtual network element(s) 860A-R perform similar functionality to the virtual network element(s) 830A-R - e.g., similar to the control communication and configuration module(s) 832A and forwarding table(s) 834A (this virtualization of the hardware 840 is sometimes referred to as network function virtualization (NFV)). Thus, NFV may be used to consolidate many network equipment types onto industry standard high volume server hardware, physical switches, and physical storage, which could be located in Data centers, NDs, andcustomer premise equipment (CPE). While embodiments of the invention are illustrated with each instance 862A-R corresponding to one VNE 860A-R, alternative embodiments may implement this correspondence at a finer level granularity (e.g., line card virtual machines virtualize line cards, control card virtual machine virtualize control cards, etc.); it should be understood that the techniques described herein with reference to a correspondence of instances 862A-R to VNEs also apply to embodiments where such a finer level of granularity and / or unikernels are used.

[0091] In certain embodiments, the virtualization layer 854 includes a virtual switch that provides similar forwarding services as a physical Ethernet switch. Specifically, this virtual switch forwards traffic between instances 862A-R and the physical NI(s) 846, as well as optionally between the instances 862A-R; in addition, this virtual switch may enforce network isolation between the VNEs 860A-R that by policy are not permitted to communicate with each other (e.g., by honoring virtual local area networks (VLANs)).

[0092] In an embodiment, software 850 includes code such as closest neighbor component 853, which when executed by processor(s) 842, causes the general purpose network device 804 to perform operations of one or more embodiments described herein as part of software instances 862 A-R (e.g., operations to determine the closest neighbors of data points included in a dataset).

[0093] The third example ND implementation in Figure 8 is a hybrid network device 806, which includes both custom ASICs / special-purpose OS and COTS processors / standard OS in a single ND or a single card within an ND. In certain embodiments of such a hybrid network device, a platform VM (i.e., a VM that that implements the functionality of the special-purpose network device 802) could provide for para-virtualization to the networking hardware present in the hybrid network device 806.

[0094] Some portions of the preceding detailed descriptions have been presented in terms of algorithms and symbolic representations of transactions on data bits within a computer memory. These algorithmic descriptions and representations are the ways used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consi stent sequence of transactions leading to a desired result. The transactions are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0095] It should be bome in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise as apparent from the above discussion, it is appreciated that throughout the description, discussions utilizing terms such as "processing" or "computing" or "calculating" or "determining" or "displaying" or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.

[0096] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various general purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method transactions. The required structure for a variety of these systems will appear from the description above. In addition, embodiments are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of embodiments as described herein.

[0097] An embodiment may be an article of manufacture in which a non-transitory machine- readable storage medium (such as microelectronic memory) has stored thereon instructions (e.g., computer code) which program one or more data processing components (generically referred to here as a “processor”) to perform the operations described above. In other embodiments, some of these operations might be performed by specific hardware components that contain hardwired logic (e.g., dedicated digital filter blocks and state machines). Those operations might alternatively be performed by any combination of programmed data processing components and fixed hardwired circuit components.

[0098] Throughout the description, embodiments have been presented through flow diagrams. It will be appreciated that the order of transactions and transactions described in these flow diagrams are only intended for illustrative purposes and not intended to be limiting. One having ordinary skill in the art would recognize that variations can be made to the flow diagrams.

[0099] In the foregoing specification, embodiments have been described with reference to specific example embodiments thereof. It will be evident that various modifications may be made thereto without departing from the broader spirit and scope of the disclosure provided herein. The specification and drawings are, accordingly, to be regarded in an illustrative sense rather than a restrictive sense.

Claims

CLAIMSWhat is claimed is:

1. A method performed by one or more computing devices to determine closest neighbors of data points included in a dataset, each data point included in the dataset having at least two dimensions including a first dimension and a second dimension, the method comprising: sorting (705) the data points included in the dataset with respect to the first dimension to generate a first sorted dataset; grouping (710) the data points included in the dataset based on their order in the first sorted dataset to generate a first plurality of groups; for each group in the first plurality of groups, generating (715) an expanded group based on expanding limits of the group along the first dimension and determining a first set of distances between data points included in the group and data points included in the expanded group; sorting (720) the data points included in the dataset with respect to the second dimension to generate a second sorted dataset; grouping (725) the data points included in the dataset based on their order in the second sorted dataset to generate a second plurality of groups; for each group in the second plurality of groups, generating (730) an expanded group based on expanding limits of the group along the second dimension and determining a second set of distances between data points included in the group and data points included in the expanded group; and determining (735) a set of closest neighbors of each of one or more data points included in the dataset based on distances in the first set of distances involving the data point and the second set of distances involving the data point.

2. The method of claim 1, further comprising: identifying (740) one or more data points included in the dataset that are outliers; determining (745) a third set of distances between each of the one or more data points identified as being outliers and all other data points included in the dataset; and determining (750) a set of closest neighbors of each of the one or more data points identified as being outliers based on the third set of distances.

3. The method of claim 2, wherein a data point is identified as being an outlier if a shortest distance in the first set of distances and the second set of distances involving the data point is longer than a predefined threshold distance.

4. The method of claim 1, wherein each group in the first plurality of groups and the second plurality of groups has a size that is equal to or less than k, wherein k is configured such that *10 data points can be stored within volatile memory of the one or more computing devices at a same time.

5. The method of claim 1, wherein the data points included in the dataset represent geographical coordinates, wherein the first dimension represents latitude and the second dimension represents longitude.

6. The method of claim 5, wherein the expanded group for a group in the first plurality of groups is generated based on expanding a northern limit of the group north by a predefined distance and expanding a southern limit of the group south by the predefined distance.

7. The method of claim 5, wherein the expanded group for a group in the second plurality of groups is generated based on expanding a western limit of the group east by an amount of longitude and expanding an eastern limit of the group west by the amount of longitude.

8. The method of claim 7, wherein the amount of longitude is determined based on identifying a data point included in the group that has a greatest absolute value in terms of latitude within the group and determining an amount of longitude corresponding to a predefined distance at the latitude of the identified data point.

9. The method of claim 5, further comprising: identifying one or more data points that have a latitude that is greater than a predefined north pole limit latitude as being north pole data points; determining a fourth set of distances between each of the north pole data points; determining a set of closest neighbors of each of one or more of the north pole data points based on the fourth set of distances; identifying one or more data points that have a latitude that is less than a predefined south pole limit latitude as being south pole data points; determining a fifth set of distances between each of the south pole data points; and determining a set of closest neighbors of each of one or more of the south pole data points based on the fifth set of distances.

10. A machine-readable medium comprising computer program code which when executed by a computer carries out the method steps of any of claims 1-9.

11. A computing device (804) to determine the closest neighbors of data points included in a dataset, each data point included in the dataset having at least two dimensions including a first dimension and a second dimension, the computing device comprising: a set of one or more processors (842); and a non-transitory machine-readable storage medium (848) that provides instructions (853) that, if executed by the set of one or more processors, will cause the computing device to carry out the method steps of any one of claims 1-9.

Citation Information

Patent Citations

  • Method, apparatus and product for efficient solution of nearest object problems

    US10746562B2

  • Mutual neighbors

    US20200265270A1

  • GEO-visual search

    US20220391437A1

Cited By

  • Intelligent design method and system for mountain tunnel lining

    CN120745064A