A Wireless Data Processing Method and System Based on Text and Spatiotemporal Characteristics

The method classifies wireless data into unique and random MAC sets, using time and location filters to accurately remove duplicates, improving processing efficiency and reliability for datasets with random MAC addresses.

CN117131025BActive Publication Date: 2025-07-15Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310846958.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-11
Publication Date
2025-07-15
Estimated Expiration
2043-07-11

AI Technical Summary

Technical Problem

Existing wireless data processing algorithms cannot efficiently and accurately remove redundant data caused by random MAC addresses in large-scale wireless data sets, affecting the availability of data.

Method used

Through a method based on text and spatiotemporal characteristics, the wireless data set is divided into MAC address unique data set and MAC address random data set, and deduplication is performed based on their respective characteristics. The uniqueness and randomness characteristics of the MAC address are used to filter and retain valid data in combination with time and location information.

Benefits of technology

It realizes efficient deduplication of large-scale wireless data, improves data availability, reduces the impact of random MAC address policies, and ensures the accuracy of data processing and the effective utilization of storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117131025B_ABST
    Figure CN117131025B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data processing, and specifically relates to a wireless data processing method and system based on text and spatio-temporal characteristics. The method divides a wireless data set into a unique MAC address data set and a random MAC address data set, and respectively performs corresponding data deduplication processing on the above two data sets based on the time information, attribute information, and spatial information of the wireless data. Different wireless data deduplication processing is performed on the two data sets. Specifically, for the random MAC address data set, not only wireless data with the same time characteristics, the same wireless network name, and the same MAC address is deduplicated, but also among the wireless data with the same wireless network name, by only retaining the wireless data whose data volume at the same position information is greater than a set value, the random MAC address data set is deduplicated, thereby ensuring that the wireless data in the random MAC address data set can also be effectively and accurately deduplicated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a wireless data processing method and system based on text and spatio-temporal characteristics. Background Art

[0002] With the development of the Internet and the development of location awareness capabilities, a large amount of spatio-temporal location data recording human production and life has been obtained based on various location sensors. According to the "Statistical Report on the Development of China's Internet Network" released by the China Internet Network Information Center (CNNIC), as of December 2022, the total number of terminal connections to the mobile network in China has reached 3.528 billion households, and WiFi has become the preferred way for Internet users to access the Internet in fixed locations. Wireless (network) data refers to data containing spatio-temporal information obtained by collecting the information of various wireless network devices and portable intelligent devices accessing the wireless network within the monitoring range relying on the WIFI probe technology. Wireless (network) data is a semi-structured data, and the main format parameters are as follows: time: timestamp, the time when the MAC is collected; MAC: the MAC address of the collected wireless device; SSID: network name ID; lat: latitude; lon: longitude. Wireless data has accurate spatial and temporal attributes, and at the same time contains rich text information.

[0003] Affected by wireless network devices and the collection environment, there are a large number of duplicate data in the collected wireless data. Usually, the MAC address is used as the physical unique identifier of the wireless router device, and by default, one MAC address corresponds to one device. However, affected by the random MAC policy, in order to protect the personal privacy information of users, when the intelligent device actively searches for new wireless networks around the device, it will hide the MAC address of the intelligent device and send a random MAC address, resulting in a large amount of redundant data in the collected data, which is not conducive to the relevant data analysis based on wireless data and reduces the usability of wireless data. Traditional wireless data processing algorithms are only applicable to small-scale data, mainly for deduplication based on data characteristics. For large-scale wireless data sets with random MAC addresses, traditional methods cannot perform data deduplication efficiently and accurately, affecting the value of wireless (network) data in related research. Summary of the Invention

[0004] The purpose of the present invention is to provide a wireless data processing method and system based on text and spatio-temporal characteristics, so as to solve the problem that the existing method of deduplication based on data characteristics cannot accurately perform data deduplication on wireless data sets with random MAC addresses.

[0005] To solve the above technical problems, the present invention provides a wireless data processing method based on text and spatio-temporal characteristics, including the following steps:

[0006] 1) Classify the obtained wireless data according to the set rules to form multiple wireless data sets;

[0007] 2) Re-divide the wireless data in each wireless data set according to the uniqueness of the MAC address to obtain a MAC address unique data set and a MAC address random data set;

[0008] 3) In the MAC address unique data set, filter the wireless data with the same time characteristics and the same wireless network name, and retain one of the wireless data. Filter the wireless data with the same wireless network name, the same MAC address, and the same location information, and retain one of the wireless data. Then form a new MAC address unique data set with the retained wireless data and the remaining unfiltered wireless data. In the MAC address random data set, filter the wireless data with the same time characteristics, the same wireless network name, and the same MAC address, and retain one of the wireless data. Filter the wireless data with the same wireless network name, and retain the wireless data with the amount of data with the same location information greater than the set value. Then form a new MAC address random data set with the retained wireless data and the remaining unfiltered wireless data;

[0009] 4) Store the wireless data in the new MAC address unique data set and the new MAC address random data set.

[0010] The beneficial effects are as follows: The method of the present invention divides the wireless data in the wireless data set into a MAC address unique data set and a MAC address random data set, and performs different duplicate removal processes on the two data sets. Specifically, for the MAC address random data set, not only the wireless data with the same time characteristics, the same wireless network name, and the same MAC address is subjected to duplicate removal, but also by retaining only the wireless data with the amount of data with the same location information greater than the set value among the wireless data with the same wireless network name, the MAC address random data set is subjected to duplicate removal, thereby ensuring that the wireless data in the MAC address random data set can also be effectively and accurately de-duplicated. And the method of the present invention is a process of removing duplicates of wireless data by combining the time and space characteristics of the data for both the MAC address unique data set and the MAC address random data set, and can accurately remove duplicate values in the wireless data.

[0011] Further, in step 3), in the MAC address unique data set, filter the wireless data with the same MAC address and the same wireless network name, and retain the wireless data with the largest amount of data with the same location information.

[0012] In the method of the present invention, considering the influence of the acquisition accuracy of the device, devices with the same MAC address and wireless network name may have inconsistent geographical coordinates (i.e., location information), and it is necessary to discriminate through coordinate information, and retain the data row with the most repeated latitude and longitude coordinates, which further ensures that the method of the present invention accurately removes duplicate values in the wireless data.

[0013] Further, in step 3), retaining one of the wireless data means retaining the wireless data that appears first in the screening process.

[0014] In the screening process of the method of the present invention, taking the first wireless data as a reference, comparing the remaining wireless data with the first wireless data. When the comparison result indicates that duplicate removal is required in the present invention, by directly removing the wireless data that duplicates the first wireless data, it ensures that the reference remains unchanged. Then, after the comparison based on this reference is completed, the reference is changed to conduct the comparison and duplicate removal process again.

[0015] Further, in step 2), according to the wireless network name, the wireless data in each wireless data set is first divided into the wireless data of fixed wireless router devices and the wireless data of custom wireless router devices. Then, the wireless data with random MAC addresses is screened out from the wireless data of custom wireless router devices to form a MAC address random data set, and the other data in the wireless data of custom wireless router devices and the wireless data of fixed wireless router devices form a MAC address unique data set.

[0016] The method of the present invention considers that wireless router devices include fixed router devices and mobile router devices. The MAC address of fixed router devices is unique, while the MAC address of mobile router devices may be random. Therefore, by first dividing the devices into fixed router devices and custom wireless router devices (mobile router devices are included therein), the wireless data of the fixed router devices is used as the MAC address unique data set, and then the corresponding wireless data is integrated into the corresponding data set through further screening in the custom wireless router devices.

[0017] Further, in step 2), the method of screening out the wireless data with random MAC addresses from the wireless data of custom wireless router devices to form a MAC address random data set is: screening the wireless data whose second character in the MAC address field is 2, 6, A, or E as the wireless data with random MAC addresses.

[0018] The method of the present invention makes full use of the characteristics of the MAC address, that is, the second character of the MAC address field can identify whether the MAC address is a random address. Therefore, the method of the present invention screens the wireless data whose second character in the MAC address field is 2, 6, A, or E as the wireless data with random MAC addresses.

[0019] Further, in step 2), the method for dividing the wireless data in each wireless data set according to the wireless network name is as follows: perform fuzzy matching between the wireless network names of the wireless data in each wireless data set and the device names in the preset wireless fixed routing device name rule library, divide the wireless data corresponding to the wireless network names with successful fuzzy matching into the wireless data of the fixed wireless routing device, and divide the remaining wireless data into the wireless data of the custom wireless routing device.

[0020] Further, the method for establishing the preset wireless fixed routing device name rule library is as follows: convert the wireless network names in the original wireless data set into text data, extract keywords from the text data, and match the keywords with the corresponding existing wireless routing device brand names to obtain the preset wireless fixed routing device name rule library.

[0021] In the early stage of establishing this rule library of the present invention, the original wireless data set is directly subjected to text conversion. After extracting keywords using a word cloud, they are matched with common fixed router brand names to obtain a rule library. The common fixed router brands can be collected through relevant platforms such as Baidu, shopping platforms, and router product manufacturers.

[0022] Further, in step 1), the obtained wireless data is wireless data with null values and outliers removed.

[0023] Since null values are unknown values, null values do not carry any valid information and are not required by the device in subsequent device operation data. Instead, they will affect the efficiency of subsequent screening of data to be stored. Therefore, in the present invention, the process of removing these null values reduces the data volume of the data set while ensuring that the valid information in the data set is not lost. Furthermore, when screening the data to be stored subsequently, the screening efficiency can be improved, so there is no situation where the storage space is wasted by null value data. And in the present invention, considering that there are also outliers with data errors in the data set, the process of removing these outliers ensures that the data stored in the storage space is all accurate data that can be used by other devices, ensuring the accuracy of the data stored in the storage space, and this process of removing outliers can also improve the efficiency of subsequent screening of stored data.

[0024] Further, in step 4), the wireless data in the new MAC address unique data set and the new MAC address random data set is stored by batch importing the wireless data into the database according to the preset data table design.

[0025] To solve the above technical problems, the present invention also provides a wireless data processing system based on text and spatio-temporal characteristics, including a memory and a processor. The processor is used to execute instructions to implement the steps of the wireless data processing method introduced above, and achieve the same beneficial effects as the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of the wireless data processing method based on text and spatio-temporal characteristics of the present invention;

[0027] Figure 2 is a flowchart of the preprocessing of raw data in the wireless data processing method based on text and spatio-temporal characteristics of the present invention;

[0028] Figure 3 is a flowchart of the initial classification of wireless data in the wireless data processing method based on text and spatio-temporal characteristics of the present invention;

[0029] Figure 4 is a flowchart of the construction of the wireless router device name rule library in the wireless data processing method based on text and spatio-temporal characteristics of the present invention;

[0030] Figure 5 is a flowchart of the deduplication process of wireless data by combining spatio-temporal characteristics in the wireless data processing method based on text and spatio-temporal characteristics of the present invention;

[0031] Figure 6 is a flowchart of the data warehousing in the wireless data processing method based on text and spatio-temporal characteristics of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0033] Embodiment of the wireless data processing method based on text and spatio-temporal characteristics:

[0034] In order to improve the efficiency of data cleaning of wireless data, reduce the impact of the random MAC policy on data processing, and improve the practicality of wireless data in the wireless data processing method of this embodiment, the present invention performs deduplication processing on wireless data by combining the text features and spatio-temporal features of wireless data. That is, the method of this embodiment classifies and deduplicates data by fully considering the text features and spatio-temporal features of wireless data, can accurately remove duplicate values in wireless data, improves the efficiency of wireless data processing, and at the same time effectively reduces the impact of the random MAC address policy on the wireless data processing process, improving the reliability of wireless data.

[0035] Such as Figure 1As shown in the figure, the wireless data processing method based on text and spatio-temporal characteristics in this embodiment includes the following steps:

[0036] Step 1: Preprocessing of original data.

[0037] As Figure 2 shown, the original data and processing process in this embodiment are to remove null values and outliers from the original data set obtained by the acquisition device. Specifically, by structuring the original wireless data set obtained by the acquisition device, slicing and splitting the data, and then performing preliminary cleaning to remove null values and outliers, mainly including attribute null values and spatial coordinate outliers.

[0038] Since null values are unknown values, null values do not carry any valid information and are not required by the device in subsequent device operation data. Instead, they will affect the efficiency of subsequent screening of data to be stored. Moreover, when this null value data is not removed and stored, the storage space for storing data will increase, resulting in a waste of storage space. And when the subsequent device accesses this storage space to obtain data, due to the existence of such invalid data control, the screening efficiency will also be affected. Therefore, in this embodiment, the process of removing this null value reduces the data volume of the data set while ensuring that the valid information of the data set is not lost. Furthermore, when screening the data to be stored subsequently, the screening efficiency of the data can be improved, and it is ensured that the data stored in the storage space are all data that need to be used. Therefore, there is no situation where the storage space is wasted by null value data. And in this embodiment, considering that there are also outliers of data errors in the data set, if this outlier is stored and applied, it will lead to an incorrect result obtained by using this abnormal data. Therefore, by removing this outlier, it is ensured that the data stored in the storage space are all accurate data that can be used by other devices, ensuring the accuracy of the data stored in the storage space. And this process of removing outliers can also improve the efficiency of subsequent screening of stored data.

[0039] Step 2: Initial classification of wireless data.

[0040] As Figure 3As shown in the figure, the initial classification of wireless data in this embodiment is a process of classifying data according to name and time. Since the wireless router device has two types of network cards, 2.4G and 5G, and 2.4G is the default network card, and the routing device can be distinguished as 2.4G or 5G through the wireless network name ID. Therefore, in this embodiment, classifying data by name is to divide the data set into a 2.4G wireless data set and a 5G wireless data set. Among them, the 2.4G wireless data set is the data obtained from the wireless router device with a 2.4G network card, and the 5G wireless data set is the data obtained from the wireless router device with a 5G network card. And time classification is to classify data according to a certain time period. For example, data in the same month is classified into one category, data in the same week is classified into one category, or multiple time ranges are set to classify data in the same time range into one category. In this embodiment, monthly statistics are adopted, that is, data in the same month is classified into one category.

[0041] 2.1 Utilize the text features of wireless data and the physical characteristics of wireless network devices to classify wireless network data by name:

[0042] Perform fuzzy matching on the network name ID field through text features to obtain a 2.4G wireless data set and a 5G wireless data set.

[0043] 2.2 Classify wireless network data by time based on the timestamps of wireless network data:

[0044] Use the time conversion function to convert the timestamps of the two types of wireless data sets after preliminary classification into the standard time format, and then generate a monthly wireless data set for the whole year with months as the statistical unit.

[0045] Step 3: Construct a wireless router device name rule library.

[0046] As Figure 4 shown, construct a wireless router device name rule library according to the text features of wireless data. This step uses a word cloud diagram to construct the name rule library. A word cloud diagram visually highlights the "keywords" with higher frequencies in the text to form a keyword cloud layer.

[0047] 3.1 Convert the network name ID field in the data set into text data;

[0048] 3.2 Perform jieba word segmentation and word frequency statistics on the text data to obtain the network name ID keywords that appear most frequently in the data, and generate a word cloud diagram;

[0049] 3.3 Compare the network name ID keywords that appear in the word cloud diagram with the existing wireless router device brand names to form a wireless fixed router device name rule library.

[0050] Step 4: Combine spatio-temporal characteristics to deduplicate wireless data.

[0051] 4.1 Fixed wireless router device data processing;

[0052] Such as Figure 5 , use the wireless network name ID and the wireless router device name rule library for fuzzy matching to obtain the original dataset of fixed router devices, and then use the uniqueness of the MAC address and geographical coordinate constraints for data deduplication.

[0053] ① Use the network name ID field and the router rule library to perform fuzzy matching on the monthly dataset to obtain the wireless fixed router device dataset;

[0054] ② Process the data obtained in ①, extract the data with unique network name ID and output it to the wireless dataset with unique network name;

[0055] ③ Combine the data time characteristics and text characteristics to deduplicate the duplicate values in the remaining data. Screen the data with the same time characteristics and the same network name ID, and retain the data that appears for the first time to obtain the wireless dataset with unique time;

[0056] ④ Use the uniqueness of the MAC address and geographical coordinate constraints for data deduplication. Screen the data with different times but the same network name ID, compare the MAC address and geographical coordinates of the data, if the MAC address is the same and the longitude and latitude coordinates are consistent, then retain the first piece of data;

[0057] ⑤ Affected by the device collection accuracy, the geographical coordinates of devices with the same MAC address and network name ID are inconsistent, and it is necessary to make a judgment through the coordinate information and retain the data row with the most repeated longitude and latitude coordinates;

[0058] ⑥ Summarize and merge the processed data.

[0059] 4.2 Custom wireless router device data processing;

[0060] Such as Figure 5 , a custom device refers to a user's custom setting of the wireless network name ID, mainly including fixed router devices and mobile router devices whose network name ID has been changed.

[0061] ① Use the naming rule of the MAC address to view the second character of the data MAC address field. If it is 2, 6, A or E, it is a random address, and it can be considered that the data is mobile router device data. Extract the qualified data to generate the original wireless dataset of mobile router devices, and then perform subsequent operations; the remaining data is processed according to step 4.1, and the processing results are merged into the fixed wireless router device dataset and stored in the database;

[0062] ② Filter the data with the same time, network name ID and MAC address, retain the data that appears for the first time, and output it to the mobile wireless routing device data set;

[0063] ③ Filter the data rows with the same network name ID from the remaining data, count the number of occurrences of data with the same longitude and latitude, set the filtering threshold (the number of occurrences of data with the same longitude and latitude > 2), retain the data that meets the threshold, and delete the rest;

[0064] ④Summarize and merge the processed data.

[0065] In the present embodiment, during the deduplication process, when duplicate data needs to be removed, the deduplication process is implemented by retaining only the first occurrence of the data. The reason for retaining this item is the influence of the original data. Affected by the data collection cycle, multiple data are collected at the same time and the data format and content are exactly the same. Therefore, only the first occurrence of the duplicate data is retained to implement the deduplication process.

[0066] Step 5: Data storage.

[0067] like Figure 6 As shown, after completing the deduplication processing of step 4, the obtained fixed wireless routing device data set and the mobile wireless routing device data set are stored. In this embodiment, after the wireless data in the data set is coordinate-converted, the coordinate-converted wireless data is imported and stored according to the data table design.

[0068] Specifically, in this embodiment, the coordinates of the data aggregated and merged in the above steps are converted from the Baidu coordinate system to the WGS-1984 coordinate system. Finally, a data table is designed in the PostgreSQL database according to the data structure characteristics, and the data that has completed the coordinate conversion is imported into the database in batches.

[0069] This embodiment performs a wireless data processing process based on text and spatiotemporal characteristics. The process implements the wireless data processing process through the steps of preprocessing the original data, initially classifying the wireless data, building a wireless routing device name rule base, deduplicating the wireless data in combination with spatiotemporal characteristics, and storing the data. This method is aimed at large-scale wireless (network) data sets with pseudo-MAC addresses. It classifies and deduplicates the data in combination with the text, time and space characteristics of the data. It can more accurately remove duplicate values in the wireless data, improve the efficiency of wireless data processing, and effectively reduce the impact of the random MAC address strategy on the wireless data processing process, thereby improving the availability of wireless (network) data.

[0070] Embodiment of wireless data processing system based on text and spatiotemporal characteristics:

[0071] The system of this embodiment includes a memory and a processor. The processor is used to execute instructions to implement the steps of the wireless data processing method. Specifically, the steps of the wireless data processing method have been introduced in detail in the embodiment of the wireless data processing method steps and will not be elaborated here.

[0072] As described above, only the preferred embodiment of the present invention is provided and is not intended to limit the present invention. The patent protection scope of the present invention is subject to the claims. All equivalent structural changes made by using the content of the specification and drawings of the present invention should, by the same token, be included in the protection scope of the present invention.

Claims

1. A wireless data processing method based on text and spatio-temporal characteristics, characterized in that, It includes the following steps: 1) Using the text features of the acquired wireless data and the physical characteristics of the wireless network devices, initially divide the wireless network data into a 2.4G wireless data set and a 5G wireless data set; then, according to the set time range, further divide the wireless data in the two initially classified wireless data sets into one category for the wireless data within the same time range, forming multiple wireless data sets; 2) First, divide the wireless data in each wireless data set into the wireless data of fixed wireless router devices and the wireless data of custom wireless router devices according to the wireless network name. Then, in the wireless data of custom wireless router devices, screen out the wireless data whose second character in the MAC address field is 2, 6, A, or E as the wireless data with random MAC addresses to form a random MAC address data set. The other data in the wireless data of custom wireless router devices and the wireless data of fixed wireless router devices form a unique MAC address data set; 3) In the unique MAC address data set, screen out the wireless data with the same time characteristics and the same wireless network name, and retain one of the wireless data. Screen out the wireless data with the same wireless network name, the same MAC address, and the same location information, and retain one of the wireless data. Then, form a new unique MAC address data set with the retained wireless data and the remaining un-screened wireless data. In the random MAC address data set, screen out the wireless data with the same time characteristics, the same wireless network name, and the same MAC address, and retain one of the wireless data. Screen out the wireless data with the same wireless network name, and retain the wireless data with the data volume of the same location information greater than the set value. Then, form a new random MAC address data set with the retained wireless data and the remaining un-screened wireless data; 4) Store the wireless data in the new unique MAC address data set and the new random MAC address data set. In step 3), in the unique MAC address data set, also screen out the wireless data with the same MAC address and the same wireless network name, and retain the wireless data with the largest data volume of the same location information.

2. The wireless data processing method based on text and spatio-temporal characteristics according to claim 1, characterized in that, In step 3), the one of the wireless data to be retained is the wireless data that appears first in the screening process.

3. The wireless data processing method based on text and spatio-temporal characteristics according to claim 1 or 2, characterized in that, In step 2), the method for initially dividing the wireless data in each wireless data set according to the wireless network name is as follows: perform fuzzy matching between the wireless network name of the wireless data in each wireless data set and the device names in the preset wireless fixed router device name rule library. Divide the wireless data corresponding to the successfully fuzzy-matched wireless network name into the wireless data of fixed wireless router devices, and divide the remaining wireless data into the wireless data of custom wireless router devices.

4. The wireless data processing method based on text and spatio-temporal characteristics according to claim 1, wherein The method for establishing the preset wireless fixed router device name rule library is as follows: convert the wireless network name in the original wireless data set into text data, extract keywords from this text data, and match these keywords with the corresponding existing wireless router device brand names to obtain the preset wireless fixed router device name rule library.

5. The wireless data processing method based on text and spatio-temporal characteristics according to claim 4, wherein ​ 6. The wireless data processing method based on text and spatio-temporal characteristics according to claim 1, characterized in that In step 1), the obtained wireless data is wireless data from which null values and outliers have been removed.

7. The wireless data processing method based on text and spatio-temporal characteristics according to claim 1, characterized in that, In step 4), the wireless data of the new MAC address unique data set and the new MAC address random data set is stored by batch importing the wireless data into the database according to the preset data table design.

8. A wireless data processing system based on text and spatio-temporal characteristics, characterized in that, It includes a memory and a processor, and the processor is used to execute instructions to implement the steps of the wireless data processing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Wireless multi-step attack mode excavation method for WLAN

    CN103944919A

  • System and method for distinguishing random MAC address of smart mobile phone

    CN110493363A