An Improved Random Forest-based WiFi Fingerprint Location Method and Device
By calculating the similarity of the decision tree and merging the decision tree of the same category, the random forest model is improved, and the problem of lower model accuracy in WiFi indoor positioning is solved, achieving higher positioning accuracy and computing efficiency.
Patent Information
- Application Number
- CN202310330109.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-03-30
AI Technical Summary
There is a problem that some decision trees in the existing random forest model have lowered model accuracy in WiFi indoor positioning, especially in terms of calculation and positioning accuracy.
By calculating the similarity of the decision tree, merging decision trees of the same category, deleting other decision trees, updating the random forest model, improving the accuracy of the decision tree and reducing the amount of calculation.
While reducing the calculation amount, it improves the accuracy and real-timeness of WiFi indoor positioning, solving the problem of lower accuracy in the trained random forest model.
Smart Images

Figure CN116390031B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of WiFi positioning, and in particular, to a WiFi fingerprint positioning method and device for improving a random forest. Background Art
[0002] With the popularization of smart phones, applications based on Location Based Services (LBS) have attracted much attention in daily life, and high-precision indoor positioning technology has become a research hotspot at the present stage. Currently, domestic and foreign researchers have proposed indoor positioning technologies such as Bluetooth, WiFi, RFID, and Ultra Wide Band (UWB). However, different positioning technologies have different application scenarios due to limitations in aspects such as positioning range, positioning accuracy, and device deployment.
[0003] The WiFi-based positioning technology has advantages such as low device deployment cost, low positioning cost, large positioning signal transceiver range, and strong applicability. There are many existing WiFi-based indoor positioning methods. Representative research results have emerged in WiFi indoor positioning technology, such as indoor positioning systems like the RADAR system, Nibble system, and Weyes system. Currently, the commonly used WiFi fingerprint positioning method often constructs a WiFi fingerprint database through features such as received signal strength and basic service set identifier, and uses machine learning models such as random forest, SVM, and decision tree to perform indoor location perception;
[0004] When using the random forest model in the prior art for WiFi indoor positioning, a large number of decision trees usually need to be established in the random forest model to ensure the model accuracy. However, there is often diversity among many decision trees. Therefore, there are some decision trees that cause the model accuracy to decrease. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a WiFi fingerprint positioning method for improving a random forest to eliminate or improve one or more defects existing in the prior art.
[0006] One aspect of the present invention provides a WiFi fingerprint positioning method for improving a random forest, and the steps of the method include:
[0007] Obtain a pre-trained random forest model, and obtain the decision tree parameters of the decision trees in the random forest model;
[0008] Calculate the similarity between every two nodes in the decision tree based on the decision tree parameters of the two nodes, and determine whether the two decision trees are decision trees of the same category based on the similarity;
[0009] Calculate the quantitative calculation parameters of each decision tree in the same category based on the decision tree parameters, determine the standard decision tree for each category based on the quantitative calculation parameters, delete the other decision trees except the standard decision tree in the random forest model, and update the random forest model;
[0010] Receive the WiFi fingerprint data, and input the WiFi fingerprint data into the updated random forest model to obtain a positioning result.
[0011] Adopting the above solution, this solution determines the similarity between two decision trees based on the decision tree parameters of each node in the decision tree, determines the decision trees in the same category based on the similarity, merges the decision trees in the same category. The merged random forest model improves the decision tree accuracy while reducing the number of decision trees, that is, improves the positioning accuracy while reducing the calculation amount, and solves the problem that the accuracy of the trained random forest model becomes low due to some decision trees.
[0012] In some embodiments of the present invention, the step of calculating the similarity between two decision trees based on the decision tree parameters of every two nodes in the decision tree includes:
[0013] Based on the decision tree parameters of two corresponding nodes in the two decision trees, determine the similar nodes in the two decision trees;
[0014] Perform qualitative calculation based on the similar nodes in the two decision trees to obtain the qualitative similarity parameter of the two decision trees;
[0015] Obtain the decision tree parameters of the similar nodes in the two decision trees, and perform quantitative calculation based on the decision tree parameters of the similar nodes in the two decision trees to obtain the quantitative similarity parameter of the two decision trees;
[0016] Calculate the similarity between the two decision trees based on the qualitative similarity parameter and the quantitative similarity parameter of the two decision trees.
[0017] In some embodiments of the present invention, the WiFi scenario includes multiple WiFi signal points for positioning, and each node of the decision tree is used to determine a WiFi signal point. In the step of determining the similar nodes in the two decision trees based on the decision tree parameters of two corresponding nodes in the two decision trees, start comparing from the root nodes of the two decision trees. If the WiFi signal points determined by the two corresponding nodes are the same signal point, then the two nodes are similar nodes, and continue to compare the downstream nodes at the corresponding positions; if the WiFi signal points determined by the two corresponding nodes are different signal points, then the two nodes are not similar nodes, and stop comparing the downstream nodes at the corresponding positions.
[0018] In some embodiments of the present invention, in the step of performing qualitative calculation based on similar nodes in two decision trees to obtain the qualitative similarity parameter of the two decision trees, the number of similar nodes in the two decision trees is obtained, and the qualitative similarity parameter between the two decision trees is calculated based on the total number of nodes in each of the two decision trees and the number of similar nodes in the two decision trees.
[0019] In some embodiments of the present invention, in the step of calculating the qualitative similarity parameter of the two decision trees based on the total number of nodes in each of the two decision trees and the number of similar nodes in the two decision trees, the qualitative similarity parameter of the two decision trees is calculated based on the following formula:
[0020]
[0021] where q1 represents the qualitative similarity parameter between the two decision trees, N L represents the number of similar nodes in the two decision trees, N1 represents the total number of nodes in decision tree 1, and N2 represents the total number of nodes in decision tree 2.
[0022] In some embodiments of the present invention, the decision tree parameter includes the mean square error parameter of each node during the training process. In the step of performing quantitative calculation based on the decision tree parameters of similar nodes in the two decision trees to obtain the quantitative similarity parameter of the two decision trees, the quantitative similarity parameter between the two decision trees is calculated according to the following formula:
[0023]
[0024] where q2 represents the quantitative similarity parameter between the two decision trees, N L represents the number of similar nodes in the two decision trees, sq1 represents the mean square error parameter of the root node in decision tree 1, represents the mean square error parameter of the i-th node in decision tree 1 that is similar to decision tree 2, represents the mean square error parameter of the i-th node in decision tree 2 that is similar to decision tree 1.
[0025] In some embodiments of the present invention, the coordinate parameters of all input data of each node of the decision tree during the training process are obtained, the average coordinate of the coordinate parameters of all input data of each node during the training process is calculated, and the mean square error of the coordinate parameters of all input data of each node during the training process is calculated based on the average value.
[0026] In some embodiments of the present invention, in the step of calculating the quantitative calculation parameter of each decision tree of the same category based on the decision tree parameter and determining the standard decision tree of each category based on the quantitative calculation parameter, the quantitative calculation parameter of each decision tree is calculated according to the following formula:
[0027]
[0028] Among them, D represents the quantitative calculation parameter of the decision tree, and N δ represents the total number of nodes of the decision tree, and m j represents the mean square error parameter of the j-th node of the decision tree, and sq represents the mean square error parameter of the root node of the decision tree.
[0029] In some embodiments of the present invention, in the step of calculating the similarity of two decision trees based on the qualitative similarity parameter and the quantitative similarity parameter of the two decision trees, the similarity of the two decision trees is calculated according to the following formula:
[0030] Q = αq1 + βq2;
[0031] Among them, Q represents the similarity of the decision tree, q1 represents the qualitative similarity parameter between the two decision trees, q2 represents the quantitative similarity parameter between the two decision trees, α is the weight parameter of the preset qualitative similarity parameter, and β is the weight parameter of the preset quantitative similarity parameter.
[0032] The second aspect of the present invention also provides a WiFi fingerprint positioning device for improving a random forest. The device includes a computer device, the computer device includes a processor and a memory, a computer instruction is stored in the memory, and the processor is used to execute the computer instruction stored in the memory. When the computer instruction is executed by the processor, the device realizes the steps implemented by the method described above.
[0033] The third aspect of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps implemented by the aforementioned WiFi fingerprint positioning method for improving a random forest are realized.
[0034] The additional advantages, objects, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objects and other advantages of the present invention can be pointed out and obtained specifically in the description and the drawings.
[0035] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not limit the present invention.
[0037] Figure 1 Schematic diagram of an implementation manner of the WiFi fingerprint positioning method for improving the random forest according to the present invention;
[0038] Figure 2 Schematic diagram of another implementation manner of the WiFi fingerprint positioning method for improving the random forest according to the present invention;
[0039] Figure 3 Schematic diagram of the architecture of this solution;
[0040] Figure 4 Schematic diagram for comparing the positioning effects of the improved random forest and the original algorithm;
[0041] Figure 5 Schematic diagram for comparing the effects of the original random forest model and the random forest model improved by this solution as the number of trees changes;
[0042] Figure 6 Schematic diagram for comparing the positioning accuracies of the KNN algorithm and the algorithm of this solution;
[0043] Figure 7 Schematic diagram of the structure of the decision tree;
[0044] Figure 8 Schematic diagram of a partial structure of the decision tree. Specific implementation manner
[0045] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the implementation manners and the accompanying drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0046] Herein, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0047] To solve the above problems, as Figure 1 shown, the present invention proposes a WiFi fingerprint positioning method for improving the random forest, and the steps of the method include:
[0048] Step S100, obtaining a pre-trained random forest model, and obtaining the decision tree parameters of the decision trees in the random forest model;
[0049] In some implementation manners of the present invention, in machine learning, a random forest model is a classifier including multiple decision trees, and the output category thereof is determined by the mode of the categories output by individual trees
[0050] In recent years, due to the significant improvement in the computing power of computers and mobile devices, machine learning technology has received extensive attention and in-depth development. More and more people have introduced machine learning methods into indoor positioning technology. WiFi-based fingerprint positioning technology is a location-based service that uses the unique signal characteristics of WiFi access points to determine the location of mobile devices. This technology relies on the fact that WiFi access points have unique identifiers and signal strengths, which can be used to determine the location of a device relative to the access point. In recent years, many researchers have conducted research on the online matching accuracy problem based on WiFi signals. The currently used matching algorithms include KNN, decision trees, SVM, random forests, etc. The process of the positioning method based on machine learning is as follows: First, in the offline phase, an RSSI fingerprint is used to train a classifier, which can be a random forest model. Then, in the online phase, the location of the user is estimated based on the trained classifier.
[0051] In some embodiments of the present invention, the parameters of the decision tree are represented using a data table.
[0052] Step S200, calculate the similarity between two decision trees based on the decision tree parameters of every two nodes in the decision tree, and determine whether the two decision trees are decision trees of the same category based on the similarity.
[0053] In some embodiments of the present invention, a WiFi scenario includes multiple WiFi signal points for positioning. Each node of the decision tree is used to determine a WiFi signal point. The decision tree parameters include the WiFi signal point determined by each node in the decision tree and the mean square error parameter of this node.
[0054] In some embodiments of the present invention, the similarities of the decision trees in the random forest model are represented in the same table, and the table format is as shown in the following table:
[0055]
[0056] Step S300, calculate the quantitative calculation parameters of each decision tree of the same category based on the decision tree parameters, determine the standard decision tree of each category based on the quantitative calculation parameters, delete other decision trees except the standard decision tree in the random forest model, and update the random forest model.
[0057] In some embodiments of the present invention, for the quantitative calculation parameters of each decision tree in a category, the higher the quantitative calculation parameter of the decision tree, the higher the calculation accuracy of the decision tree. Then, this decision tree is used as the standard decision tree of this category. Then, only the standard decision trees corresponding to the number of categories exist in the updated random forest model. This method performs regularization processing on the decision tree to improve the calculation accuracy of the decision tree.
[0058] In the specific implementation process, decision tree regularization is a technique used to prevent decision trees from overfitting the training data. Overfitting occurs when the decision tree model fits the training data too closely, resulting in poor generalization to new, unseen data. Regularization techniques can help reduce overfitting in decision trees by adding constraints to the tree construction process. The following are some common regularization techniques for decision trees:
[0059] Tree pruning: This technique involves removing nodes or branches from the tree that do not improve the model's performance on the validation data;
[0060] Maximum depth: This technique limits the depth of the decision tree. Shallow trees are less likely to overfit the training data;
[0061] Minimum samples to split: This technique sets a threshold for the number of training samples required to split a node. Nodes with fewer samples than the minimum will not be split;
[0062] Minimum samples per leaf: This technique sets a threshold for the number of training samples required in a leaf node. Nodes with fewer samples than the minimum will be merged with their parent node;
[0063] Feature selection: This technique limits the number of features used to split nodes, considering only a subset of features at each split, thus reducing the complexity of the model.
[0064] These techniques can be used individually or in combination to regularize the decision tree and improve its generalization performance. The principle of regularization is to minimize the overfitting phenomenon as much as possible without significantly affecting the performance of the random forest regression.
[0065] Step S400: Receive the WiFi fingerprint data, input the WiFi fingerprint data into the updated random forest model, and obtain the positioning result.
[0066] Using the above solution, as Figure 3 shown, when calculating two identical decision trees, their structures must be the same, and quantitative analysis can well reflect the performance of this decision tree. After obtaining the similarity matrix, first perform clustering on the decision trees in the random forest. If a decision tree metric is less than threshold1, it is considered to have poor performance and is discarded; otherwise, find decision trees with a similarity greater than threshold2 in the similarity matrix and group them into one category. Then select the decision tree with the best performance in each cluster to form a new random forest. Based on the decision tree set established in the previous step, use each decision tree for regression prediction, and then take the average of the prediction results of all decision trees to obtain the final prediction result.
[0067] In some embodiments of the present invention, the WiFi fingerprint data is a vector composed of RSSI (Received Signal Strength Indication) parameters received from multiple WiFi signal points for positioning in a WiFi scenario.
[0068] In the specific implementation process, the positioning result is a coordinate parameter.
[0069] The existing KNN-based WiFi fingerprint online location matching algorithm requires a large amount of computing resources, especially when dealing with large datasets, because it has to calculate the distance between each data point in the training set and the query point; it performs poorly in high-dimensional spaces. As the number of features increases, the distance between any two points tends to become larger, reducing the efficiency of the algorithm. This problem is called the curse of dimensionality. When performing regression operations based on the traditional random forest algorithm, randomness is introduced during the construction of decision trees, and some decision trees with high similarity and poor regression effects will be constructed during the training phase, affecting the accuracy of the model.
[0070] Furthermore, as a combined classifier, the performance of a random forest depends on the performance of each base classifier and the correlation between the base classifiers in the combined classifier, and is positively correlated with the former and negatively correlated with the latter. In a random forest, the performance of each decision tree is relatively accurate in the corresponding dataset, but due to randomly sampling sample values for training, its learning effect is one-sided. At the same time, it is easy to cause low diversity in the decision tree cluster, making the effect of the random forest worse.
[0071] With the above solution, this solution determines the similarity between two decision trees through the decision tree parameters of each node in the decision tree, determines the decision trees of the same category based on the similarity, merges the decision trees of the same category, and the merged random forest model improves the accuracy of the decision tree while reducing the number of decision trees, that is, improves the positioning accuracy while reducing the computational amount. This solution can further improve the WiFi fingerprint positioning accuracy and real-time performance in a complex indoor environment, and solves the problem that some decision trees in the trained random forest model cause the model accuracy to become low.
[0072] As Figure 2 shown, in some embodiments of the present invention, the steps of calculating the similarity between two decision trees based on the decision tree parameters of every two nodes in the decision tree include:
[0073] Step S210, based on the decision tree parameters of two corresponding nodes in the two decision trees, determine the similar nodes in the two decision trees;
[0074] As Figure 7 and 8As shown, in the specific implementation process, the decision tree adopts a binary tree architecture and starts from the root node.
[0075] In the specific implementation process, two nodes at corresponding positions in the two decision trees start from the root node. First, the root nodes of the two decision trees correspond to each other. The left child node of the root node of one decision tree corresponds to the left child node of the root node of the other decision tree, and the right child node of the root node of one decision tree corresponds to the right child node of the root node of the other decision tree, and then continue to judge the next-level child nodes of the two child nodes.
[0076] In the specific implementation process, the constructed decision tree model is converted into interpretable data. Each decision tree after training is a model and cannot directly participate in calculations. It is necessary to convert the model into data for interpretation and calculation. This algorithm uses an array storage method similar to a binary tree to represent the decision tree. For each node in the regression model of the decision tree, the mean squared error, the feature and feature value of node splitting, and the output result under the current node are used to describe the numerical information of the node; the nodes are numbered to describe the structural information between the nodes. According to the constructed data model, all decision trees in the random forest are represented by this data model, as Figure 7 shown, Figure 7 in which RSSI1 is the RSSI value of the first WiFi signal point determined by this node, RSSI2 is the RSSI value of the second WiFi signal point determined by this node, and similarly represents other WiFi signal points; Figure 7 in which squareerror represents the mean squared error parameter of this node.
[0077] Step S220, perform qualitative calculation based on similar nodes in the two decision trees to obtain the qualitative similarity parameter of the two decision trees;
[0078] In the specific implementation process, starting from the root nodes of the two decision trees, first compare the root nodes of the two decision trees. If the two root nodes are not used to determine the same WiFi signal point, then the two root nodes are not similar nodes, terminate the comparison, and the number of similar nodes of the two decision trees is 0; if the two root nodes are used to determine the same WiFi signal point, then the two root nodes are similar nodes, and at this time the number of similar nodes of the two decision trees is 1; if the two root nodes are similar nodes, then continue to judge the two child nodes of the two root nodes. If the two nodes under the corresponding root nodes of the two decision trees are both used to determine the same WiFi signal point, then the two nodes under the corresponding root nodes of the two decision trees are similar nodes, and at this time the number of similar nodes of the two decision trees is 3, and continue to judge the downstream nodes; if one of the two nodes under the corresponding root nodes of the two decision trees is a similar node, at this time the number of similar nodes of the two decision trees is 2, and continue to judge the nodes under the similar nodes.
[0079] Step S230: Obtain the decision tree parameters of the similar nodes in the two decision trees, perform quantitative calculation based on the decision tree parameters of the similar nodes in the two decision trees, and obtain the quantitative similarity parameter of the two decision trees.
[0080] In some embodiments of the present invention, in the step of performing quantitative calculation based on the decision tree parameters of the similar nodes in the two decision trees, the decision tree parameters of the similar nodes include the mean squared error parameter of the node.
[0081] Step S240: Calculate the similarity of the two decision trees based on the qualitative similarity parameter and the quantitative similarity parameter of the two decision trees.
[0082] Adopting the above solution, this solution first determines the similar nodes of the two decision trees based on the WiFi signal points determined by the nodes, and determines the qualitative similarity parameter of the two decision trees. Further, quantitative calculation is performed through the decision tree parameters of the similar nodes in the two decision trees. That is, this solution takes into account both qualitative and quantitative indicators, and comprehensively determines the similarity of the two decision trees based on the two indicators, improving the determination accuracy of the decision tree similarity.
[0083] In the specific implementation process, step S250: Determine whether the two decision trees are decision trees of the same category based on the similarity.
[0084] In some embodiments of the present invention, the WiFi scenario includes multiple WiFi signal points for positioning, and each node of the decision tree is used to determine a WiFi signal point. In the step of determining the similar nodes in the two decision trees based on the decision tree parameters of the two nodes at the corresponding positions in the two decision trees, start comparing from the root nodes of the two decision trees. If the WiFi signal points determined by the two nodes at the corresponding positions are the same signal point, the two nodes are similar nodes, and continue to compare the downstream nodes at the corresponding positions; if the WiFi signal points determined by the two nodes at the corresponding positions are different signal points, the two nodes are not similar nodes, and stop comparing the downstream nodes at the corresponding positions.
[0085] In some embodiments of the present invention, in the step of performing qualitative calculation based on the similar nodes in the two decision trees to obtain the qualitative similarity parameter of the two decision trees, obtain the number of similar nodes in the two decision trees, and calculate the qualitative similarity parameter between the two decision trees based on the total number of nodes in each of the two decision trees and the number of similar nodes in the two decision trees.
[0086] In some embodiments of the present invention, in the step of calculating the qualitative similarity parameter of the two decision trees based on the total number of nodes in each of the two decision trees and the number of similar nodes in the two decision trees, calculate the qualitative similarity parameter of the two decision trees based on the following formula:
[0087]
[0088] Among them, q1 represents the qualitative similarity parameter between two decision trees, N L represents the number of similar nodes in the two decision trees, N1 is the total number of nodes in decision tree 1, and N2 is the total number of nodes in decision tree 2.
[0089] In some embodiments of the present invention, the decision tree parameters include the mean square error parameter of each node during the training process. In the step of quantitatively calculating the quantitative similarity parameter of the two decision trees based on the decision tree parameters of the similar nodes in the two decision trees, the quantitative similarity parameter between the two decision trees is calculated according to the following formula:
[0090]
[0091] Among them, q2 represents the quantitative similarity parameter between two decision trees, N L represents the number of similar nodes in the two decision trees, sq1 represents the mean square error parameter of the root node in decision tree 1, represents the mean square error parameter of the i-th node in the nodes of decision tree 1 that are similar to decision tree 2, represents the mean square error parameter of the i-th node in the nodes of decision tree 2 that are similar to decision tree 1.
[0092] In some embodiments of the present invention, the coordinate parameters of all input data of each node of the decision tree during the training process are obtained, the average coordinate of the coordinate parameters of all input data of each node during the training process is calculated, and the mean square error of the coordinate parameters of all input data of each node during the training process is calculated based on the average value.
[0093] In the specific implementation process, each piece of training data includes the RSSI value received by each WiFi signal point and the coordinate parameters of the emission point. The step of calculating the average coordinate of the coordinate parameters of all input data of each node during the training process is to calculate the average coordinate composed of the average abscissa and the average ordinate of the coordinate parameters input to each node.
[0094] In the specific implementation process, in the step of calculating the mean square error of the coordinate parameters of all input data of each node during the training process based on the average value, the mean square error is calculated according to the following formula:
[0095]
[0096] The average coordinate is calculated according to the following formula:
[0097]
[0098] Among them, H represents the mean square error of the node, and N m represents the total number of data input to this node, and y i represents the coordinates of the i-th data input to this node, and c m represents the average coordinate of this node;
[0099] Measurement of the similarity between decision trees and the performance of decision trees themselves. The similarity between decision trees requires qualitative analysis of the decision tree structure and quantitative analysis of the decision tree performance. The former is calculated using the in-order traversal method. If the splitting criteria of a certain node in two decision trees are the same, then calculate the similarity of the left and right subtrees of this node, and so on recursively. The ratio of the number of similar nodes to the total number of nodes in the two trees is used as an indicator to measure the structural similarity of the two decision trees. The latter is carried out during the traversal of the decision tree. If only one node in a certain subtree of two decision trees is similar, then record the MSE of the two nodes, respectively record them in a list, and the minimum standard MSE determines the position of the splitting node in the future.
[0100] In some embodiments of the present invention, in the step of calculating the quantitative calculation parameters of each decision tree of the same category based on the decision tree parameters and determining the standard decision tree of each category based on the quantitative calculation parameters, calculate the quantitative calculation parameters of each decision tree according to the following formula:
[0101]
[0102] Among them, D represents the quantitative calculation parameter of the decision tree, and N δ represents the total number of nodes of the decision tree, and m j represents the mean square error parameter of the j-th node of the decision tree, and sq represents the mean square error parameter of the root node of the decision tree.
[0103] In some embodiments of the present invention, in the step of calculating the similarity of two decision trees based on the qualitative similarity parameter and the quantitative similarity parameter of the two decision trees, calculate the similarity of the two decision trees according to the following formula:
[0104] Q = αq1 + βq2;
[0105] Among them, Q represents the similarity of the decision tree, q1 represents the qualitative similarity parameter between the two decision trees, q2 represents the quantitative similarity parameter between the two decision trees, α is the weight parameter of the preset qualitative similarity parameter, and β is the weight parameter of the preset quantitative similarity parameter.
[0106] Experimental Example 1:
[0107] Deploy eight WiFi AP nodes in a small indoor environment. Then use a mobile phone App to collect the signal strength of all AP nodes during the walking process of people, store it in the mobile phone memory, and then transmit it to a PC to verify the effectiveness of the algorithm.
[0108] First, preprocess the collected data, including filtering out noise, removing irrelevant data, and normalizing the signal strength values. Among them, use box plots to remove irrelevant data. First, calculate the quartiles of each location fingerprint information, and then obtain the upper whisker and lower whisker of the box plot according to the quartiles. After removing outliers, calculate the mean value of the RSSI value of each AP node at the reference point to construct the fingerprint database.
[0109] In the offline stage, use an improved random forest algorithm to learn the fingerprint database, and use the L-shaped route of people walking in the area where the fingerprint database is located for positioning verification at this stage. First, perform pruning operations on the decision trees. The purpose of this operation is to reduce the overfitting phenomenon of the model by setting parameters, and then reduce the computational complexity of subsequent operations.
[0110] As Figure 4 shown, through experiments, it is obtained that when the number of decision trees in the random forest is about 80, the positioning accuracy is basically stable at about 2.2m. So set the number of decision trees to 80 for the next experiment. Figure 4 For the comparison of the positioning effects between the improved random forest and the original algorithm.
[0111] Figure 5 Set the number of random forest trees to 80, and then run 100 times to obtain the positioning effects of the original random forest algorithm and the improved algorithm on the validation set. It can be seen from the two figures that the algorithm proposed in this paper can improve the positioning results obtained by the original random forest regression operation. The construction process of the random forest itself brings randomness in the construction process of the decision tree, so the degree of improvement in positioning accuracy is also random. But on average, this algorithm can improve the positioning accuracy by 10.8608%. Figure 5 For the schematic diagram of the comparison of the effects of the two algorithms with the number of trees.
[0112] Figure 5 In, the upper two lines represent the effects of the two algorithms on the training set, and the lower two lines represent the effects of the two algorithms on the validation set. Among them, the straight line is the effect of the random forest algorithm, and the dotted line is the effect of the improved random forest algorithm. From Figure 5 it can be seen that the improved random forest algorithm can not only improve the accuracy of the random forest algorithm, but also effectively suppress the overfitting effect.
[0113] KNN, as a baseline algorithm commonly used for comparison in machine learning, can be used for both classification and regression. We are in Figure 6The comparison of the positioning accuracy between the KNN algorithm and the algorithm proposed in this paper is shown in Figure 6 As can be seen from
[0114] An embodiment of the present invention further provides a WiFi fingerprint positioning device that improves the random forest. The device includes a computer device, and the computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps implemented by the method described above.
[0115] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps implemented by the aforementioned WiFi fingerprint positioning method that improves the random forest. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0116] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to execute in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to execute the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0117] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0118] In the present invention, features described and / or illustrated for one embodiment may be used in the same manner or in a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0119] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An improved WiFi fingerprint positioning method based on random forest, characterized in that The steps of the method include: Obtain a pre-trained random forest model, and obtain the decision tree parameters of the decision trees in the random forest model; Calculate the similarity between two decision trees based on the decision tree parameters of every two nodes in the decision trees. Based on the decision tree parameters of two corresponding nodes in the two decision trees, determine the similar nodes in the two decision trees. Perform qualitative calculation based on the similar nodes in the two decision trees to obtain the qualitative similarity parameters of the two decision trees. The decision tree parameters include the mean squared error parameters of each node during the training process. Calculate the quantitative similarity parameters between the two decision trees according to the following formula: Among them, q2 represents the quantitative similarity parameter between two decision trees, and N L represents the number of similar nodes in the two decision trees, sq1 represents the mean square error parameter of the root node in decision tree 1, represents the mean square error parameter of the i-th node among the nodes similar to decision tree 2 in decision tree 1, represents the mean square error parameter of the i-th node among the nodes similar to decision tree 1 in decision tree 2; Obtain the decision tree parameters of the similar nodes in the two decision trees, perform quantitative calculation based on the decision tree parameters of the similar nodes in the two decision trees to obtain the quantitative similarity parameters of the two decision trees. Calculate the similarity between the two decision trees based on the qualitative similarity parameters and quantitative similarity parameters of the two decision trees, and determine whether the two decision trees are decision trees of the same category based on the similarity; Calculate the quantitative calculation parameters of each decision tree of the same category based on the decision tree parameters, determine the standard decision tree of each category based on the quantitative calculation parameters, delete the other decision trees except the standard decision tree in the random forest model, and update the random forest model; Receive WiFi fingerprint data, and input the WiFi fingerprint data into the updated random forest model to obtain a positioning result.
2. The improved WiFi fingerprint positioning method based on the random forest according to claim 1, wherein The WiFi scenario includes multiple WiFi signal points for positioning. Each node of the decision tree is used to determine a WiFi signal point. In the step of determining the similar nodes in the two decision trees based on the decision tree parameters of two corresponding nodes in the two decision trees, start comparing from the root nodes of the two decision trees. If the WiFi signal points determined by the two corresponding nodes are the same signal point, then the two nodes are similar nodes, and continue to compare the downstream nodes at the corresponding positions; if the WiFi signal points determined by the two corresponding nodes are different signal points, then the two nodes are not similar nodes, and stop comparing the downstream nodes at the corresponding positions.
3. The improved WiFi fingerprint positioning method based on the random forest according to claim 1, characterized in that, In the step of performing qualitative calculation based on the similar nodes in the two decision trees to obtain the qualitative similarity parameters of the two decision trees, obtain the number of similar nodes in the two decision trees, and calculate the qualitative similarity parameters of the two decision trees based on the total number of nodes in each of the two decision trees and the number of similar nodes in the two decision trees.
4. The improved WiFi fingerprint positioning method based on the random forest according to claim 3, characterized in that In the step of calculating the qualitative similarity parameters of the two decision trees based on the total number of nodes in each of the two decision trees and the number of similar nodes in the two decision trees, calculate the qualitative similarity parameters of the two decision trees according to the following formula: Among them, q1 represents the qualitative similarity parameter of two decision trees, N L represents the number of similar nodes in two decision trees, N1 is the total number of nodes in decision tree 1, and N2 is the total number of nodes in decision tree 2.
5. The improved WiFi fingerprint positioning method based on the random forest according to claim 1, characterized in that, Obtain the coordinate parameters of all the input data of each node of the decision tree during the training process, calculate the average coordinate of the coordinate parameters of all the input data of each node during the training process, and calculate the mean squared error of the coordinate parameters of all the input data of each node during the training process based on the average value.
6. The improved WiFi fingerprint positioning method based on the random forest according to claim 1, characterized in that, In the step of calculating the quantitative calculation parameters of each decision tree of the same category based on the decision tree parameters and determining the standard decision tree of each category based on the quantitative calculation parameters, the quantitative calculation parameters of each decision tree are calculated according to the following formula: Among them, D represents the quantitative calculation parameter of the decision tree, N δ represents the total number of nodes of the decision tree, m j represents the mean square error parameter of the j-th node of the decision tree, and sq represents the mean square error parameter of the root node of the decision tree.
7. The improved WiFi fingerprint positioning method based on the random forest according to claim 1, characterized in that In the step of calculating the similarity of two decision trees based on the qualitative similarity parameter and the quantitative similarity parameter of the two decision trees, the similarity of the two decision trees is calculated according to the following formula: Q = αq1 + βq2; where Q represents the similarity of the decision tree, q1 represents the qualitative similarity parameter between the two decision trees, q2 represents the quantitative similarity parameter between the two decision trees, α is the weight parameter of the preset qualitative similarity parameter, and β is the weight parameter of the preset quantitative similarity parameter.
8. An improved random forest-based WiFi fingerprint positioning device, characterized in that, The device includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps implemented by the method according to any one of claims 1 to 7.