Internet of Things-based Personal Information Secure Storage Method and System
The IoT-based personal information storage method addresses the inefficiencies and vulnerabilities of traditional data storage by clustering and filtering data, improving storage efficiency and security.
Patent Information
- Application Number
- CN202411197215.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-29
AI Technical Summary
Traditional data storage methods lead to waste of resources and vulnerability to attacks, and the data stored in centralized data is insufficient.
The Internet of Things-based personal information security storage method is adopted to carefully classify data points and eliminate useless data through clustering analysis and feature similarity judgment to achieve reasonable data storage and security protection.
Effectively reduce storage space consumption, reduce the risk of data loss, and improve the stability and reliability of information systems.
Smart Images

Figure CN119089499B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for secure storage of personal information based on the Internet of Things. Background Art
[0002] In the Internet, data privacy and security are of utmost importance. It not only relates to the protection of personal information but also involves the security of the entire network. With the continuous development of the information industry, information security technology has become one of the key forces driving the progress of this industry. A computer network is a distributed system that ensures the security of data interaction through communication and encryption operations between nodes. However, traditional data storage methods are data-centric and concentrate high-density information on specific nodes. Although this approach simplifies data management, it also brings some problems. Firstly, due to the over-concentration of data, the consumption rate of storage space may be too fast, resulting in resource waste. Secondly, centrally stored data is more likely to become a target of attack. Once these nodes are attacked or malfunction, it may lead to data loss or invalidation. Summary of the Invention
[0003] The present invention provides a method and system for secure storage of personal information based on the Internet of Things to solve existing problems.
[0004] The method and system for secure storage of personal information based on the Internet of Things of the present invention adopt the following technical solutions:
[0005] An embodiment of the present invention provides a method for secure storage of personal information based on the Internet of Things, which includes the following steps:
[0006] Obtain the sample space of all information data points to be stored, and obtain several clustering clusters of the data points in the sample space;
[0007] According to the distance between data points in each clustering cluster and the distance from the data points to the center point of the clustering cluster, obtain the classification consistency of any two data points in each clustering cluster belonging to the same small classification; cluster all the data points in each clustering cluster according to the classification consistency of any two data points in each clustering cluster belonging to the same small classification, and obtain several small classifications and data scatter points in each clustering cluster, where each small classification contains several data points;
[0008] According to the distance between each data scatter point and the center point of each small classification in each clustering cluster, and the classification consistency of each data scatter point and the center point of each small classification belonging to the same small classification, obtain the feature similarity between each data scatter point and each small classification in each clustering cluster and the suspected target scatter points of each small classification;
[0009] According to the changes of each small classification before and after the suspected target scattered points of each small classification in each cluster are added, and the feature similarity between each data scattered point and each small classification, the similarity between each small classification in each cluster and each suspected target scattered point in each small classification is obtained; according to the similarity between each small classification in each cluster and each suspected target scattered point in each small classification, it is determined whether the suspected target scattered points of each small classification in each cluster are added;
[0010] Obtain useless data classification based on the number of small categories in each cluster, the distance between the updated center points of each small category, the number of data points in each small category and the number of suspected target scattered points finally added; eliminate the data in the useless data classification and complete the data storage.
[0011] Preferably, the step of obtaining the classification consistency of any two data points in each cluster belonging to the same sub-classification according to the distance between the data points in each cluster and the distance from the data point to the center point of the cluster includes the following specific steps:
[0012]
[0013] In the formula, Q v,k,y is the classification consistency that the kth data point and the yth data point in the vth cluster belong to the same sub-category, l v,k,y is the Euclidean distance between the kth data point and the yth data point in the vth cluster in the sample space, L v,k is the distance between the kth data point in the vth cluster and the center point of the vth cluster, L v,y is the distance between the yth data point in the vth cluster and the center point of the vth cluster, || is the absolute value, and Sigmoid() is the normalization function.
[0014] Preferably, the method of obtaining the feature similarity between each data scatter point in each cluster and each small classification and the suspected target scatter points of each small classification according to the distance between each data scatter point in each cluster and the center point of each small classification, and the classification consistency that each data scatter point and the center point of each small classification belong to the same small classification, includes the following specific steps:
[0015]
[0016] Where W v,i,j is the feature similarity between the i-th data point and the j-th small classification in the v-th cluster, l v,i,j is the Euclidean distance between the ith data point in the vth cluster and the jth small classification center point, Q v,i,j is the classification consistency that the i-th data scatter point in the v-th cluster and the j-th small classification center point belong to the same small classification, max(Q v,i,j) is the maximum value among the classification consistencies that the $i$-th data scatter point in the $v$-th cluster and all data points in the $j$-th sub-classification belong to the same sub-classification;
[0017] According to the feature similarity between each data scatter point and each sub-classification, obtain the suspected target scatter points of each sub-classification.
[0018] Preferably, the step of obtaining the suspected target scatter points of each sub-classification according to the feature similarity between each data scatter point and each sub-classification specifically includes the following steps:
[0019] Among the feature similarities between the $i$-th data scatter point in the $v$-th cluster and all sub-classifications, select a sub-classification corresponding to the maximum feature similarity, and use the $i$-th data scatter point as the suspected target scatter point of the sub-classification corresponding to the maximum feature similarity.
[0020] Preferably, the step of obtaining the similarity of each sub-classification in each cluster to each suspected target scatter point in each sub-classification according to the changes of each sub-classification before and after adding the suspected target scatter points of each sub-classification in each cluster, and the feature similarity between each data scatter point and each sub-classification specifically includes the following steps:
[0021] Use the principal component analysis method to obtain the projection difference on the principal components before and after adding the $f$-th suspected target scatter point of the $j$-th sub-classification in the $v$-th cluster to the $j$-th sub-classification to quantify the change value of the direction, and record the change value as the change value of the overall direction of the $j$-th sub-classification in the $v$-th cluster after adding the $f$-th suspected target scatter point to the $j$-th sub-classification;
[0022] Record the center point of the $j$-th sub-classification after adding the $f$-th suspected target scatter point of the $j$-th sub-classification as the updated center point of the $j$-th sub-classification;
[0023] Calculate the absolute value of the difference between the distance from the $f$-th suspected target scatter point of the $j$-th sub-classification to the center point of the $j$-th sub-classification and the distance from the $f$-th suspected target scatter point of the $j$-th sub-classification to the updated center point of the $j$-th sub-classification, and record the absolute value of the distance difference as the change value of the spatial size of the $j$-th sub-classification in the $v$-th cluster after adding the $f$-th suspected target scatter point to the $j$-th sub-classification;
[0024] According to the change value of the overall direction and the change value of the spatial size of each sub-classification in each cluster after adding each suspected target scatter point, and the feature similarity between each data scatter point and each sub-classification, obtain the similarity of each sub-classification in each cluster to each suspected target scatter point in each sub-classification.
[0025] Preferably, after each suspected target scatter point is added, the similarity between each sub-classification in each cluster and each suspected target scatter point in each sub-classification is obtained based on the change value of the overall direction and the change value of the spatial size of each sub-classification in each cluster, and the feature similarity between each data scatter point and each sub-classification. The specific steps are as follows:
[0026]
[0027] In the formula, R v,f,j is the similarity between the j-th sub-classification in the v-th cluster and the f-th suspected target scatter point in the j-th sub-classification, W v,f,j is the feature similarity between the j-th sub-classification in the v-th cluster and the f-th suspected target scatter point in the j-th sub-classification, S v,f,j is the change value of the overall direction of the j-th sub-classification in the v-th cluster after the f-th suspected target scatter point is added to the j-th sub-classification, θ v,f,j is the change value of the spatial size of the j-th sub-classification in the v-th cluster after the f-th suspected target scatter point is added to the j-th sub-classification, N v,j,f is the number of suspected target scatter points other than the f-th suspected target scatter point in the j-th sub-classification in the v-th cluster, S j,m,v is the change value of the spatial size of the j-th sub-classification in the v-th cluster after the m-th suspected target scatter point is added to the j-th sub-classification, θ j,m,v is the change value of the overall direction of the j-th sub-classification in the v-th cluster after the m-th suspected target scatter point is added to the j-th sub-classification, || represents taking the absolute value.
[0028] Preferably, based on the similarity between each sub-classification in each cluster and each suspected target scatter point in each sub-classification, it is determined whether the suspected target scatter points in each sub-classification of each cluster are added. The specific steps are as follows:
[0029] When the similarity between the j-th sub-classification in the v-th cluster and the f-th suspected target scatter point in the j-th sub-classification is greater than the preset threshold, the f-th suspected target scatter point in the j-th sub-classification is added to the j-th sub-classification in the v-th cluster to obtain the updated classification of each sub-classification in the v-th cluster.
[0030] Preferably, based on the number of sub-classifications in each cluster, the distance between the updated center points of each sub-classification, the number of data points and the finally added suspected target scatter points in each sub-classification, the useless data classification is obtained. The specific steps are as follows:
[0031]
[0032] In the formula, U v,e is the uselessness of the e-th updated classification data in the v-th cluster, nv,e is the number of all data points in the e-th updated classification in the v-th cluster, n v is the number of updated classifications in the v-th cluster, l v,e,h is the distance between the central points of the e-th updated classification and the h-th updated classification in the v-th cluster, N v,e is the number of data scatter points in the e-th updated classification in the v-th cluster, R v,e,b is the similarity between the sub-classification corresponding to the e-th updated classification in the v-th cluster and the b-th data scatter point in the e-th updated classification; Sigmoid() is a normalization function;
[0033] Obtain the useless data classification according to the uselessness of each updated classification data in each cluster.
[0034] Preferably, the step of obtaining the useless data classification according to the uselessness of each updated classification data in each cluster specifically includes the following steps:
[0035] When the uselessness of the e-th updated classification in the v-th cluster is greater than a preset threshold, it is determined that the e-th updated classification belongs to the useless data classification.
[0036] The present invention also proposes an Internet of Things-based personal information security storage system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned Internet of Things-based personal information security storage method.
[0037] The beneficial effects of the technical solution of the present invention are: By classifying the data more carefully and deleting abnormal classification data, the present invention can effectively reduce the total amount of data and the data transmission volume per unit time. This not only reduces the consumption speed of storage space, but also avoids data loss and invalidation. The detailed classification ensures that similar data is reasonably stored, while identifying and eliminating abnormal data, further optimizing the storage utilization efficiency. This method protects data privacy and security while also enhancing the stability and reliability of the entire information system. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 is the flowchart of the steps of the Internet of Things-based personal information security storage method of the present invention. Detailed Implementation Manner
[0040] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manner, structure, features, and effects of the method and system for secure storage of personal information based on the Internet of Things proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0042] The following specifically describes the specific solutions of the method and system for secure storage of personal information based on the Internet of Things provided by the present invention with reference to the accompanying drawings.
[0043] Please refer to Figure 1 , which shows a flowchart of the steps of the method for secure storage of personal information based on the Internet of Things provided by an embodiment of the present invention. The method includes the following steps:
[0044] Step S001: Obtain the sample space of all information data points to be stored, and obtain several clustering clusters of the data points in the sample space.
[0045] Obtain all information data points to be stored, classify a large amount of data to be stored, and regard each piece of data as a data point. Map each data point into a multi-dimensional sample space according to data characteristics, where each dimension is regarded as an axis, representing an attribute or feature of the data.
[0046] Use the K-means clustering algorithm for clustering to obtain multiple clustering clusters. In this embodiment, K = 30, and the K-means clustering algorithm is a well-known technology, taking this as an example.
[0047] It should be noted that the information data in this embodiment takes device data, user behavior data, geographical location information, personal privacy information, etc. as examples.
[0048] Step S002: According to the distance between data points in each clustering cluster and the distance from the data points to the center point of the clustering cluster, obtain the classification consistency that any two data points in each clustering cluster belong to the same small classification; cluster all the data points in each clustering cluster according to the classification consistency that any two data points in each clustering cluster belong to the same small classification, and obtain several small classifications and data scatter points in each clustering cluster, where each small classification contains several data points.
[0049] When classifying data using the clustering method, especially when the data content is the main storage classification condition, there may be a situation where the amount of data in some classifications is excessive. This will not only lead to a rapid reduction in storage space but also increase the risk of data loss. For example, when processing the data of a certain nursing home, the amount of data in the medical examination category may be particularly large, and this category can be further subdivided into multiple small classifications such as heart examination, brain examination, blood examination, etc. To optimize the storage process and ensure data security, the large classification can be further subdivided into multiple small classifications for storage, which can effectively slow down the reduction speed of storage space and reduce the risk of data loss.
[0050] The center point of each cluster is obtained through the centroid formula. It should be noted that the clustering center point of the cluster may not be a data point in the cluster.
[0051] In any one cluster, obtain the classification consistency that any two data points belong to the same small classification.
[0052] This implementation gives a method for calculating the classification consistency that the k-th data point and the y-th data point in the same cluster belong to the same small classification:
[0053]
[0054] In the formula, Q v,k,y is the classification consistency that the k-th data point and the y-th data point in the v-th cluster belong to the same small classification, l v,k,y is the Euclidean distance between the k-th data point and the y-th data point in the v-th cluster in the sample space, L v,k is the distance between the k-th data point in the v-th cluster and the center point of the v-th cluster, L v,y is the distance between the y-th data point in the v-th cluster and the center point of the v-th cluster, || is to take the absolute value, and Sigmoid( ) is the normalization function.
[0055] In the above formula, |L v,k -L v,y | is the absolute value of the difference in the distances of the two data points from the center point of the v-th cluster. The smaller the value, the more likely the two data points belong to the same small classification.
[0056] According to the reciprocal of the classification consistency that all any two data points in any one cluster belong to the same small classification, use the DBSCAN clustering algorithm to cluster all the data points in the any one cluster to obtain several small classifications and data scatter points. In this embodiment, the preset parameter radius value is 20 and the preset number of points in the MinPts neighborhood is 5 during the clustering process. Among them, the DBSCAN clustering algorithm is a well-known technology, and this is used as an example.
[0057] It should be noted that the core idea of the DBSCAN clustering algorithm is to define clusters by density, and a cluster consists of a set of density-connected points. If a point lies within the density region of a core point, it will be assigned to a cluster; otherwise, it may be marked as a noise point, that is, a data scatter point.
[0058] Step S003: According to the distance between each data scatter point and each small classification center point in each clustering cluster, and the classification consistency that each data scatter point and each small classification center point belong to the same small classification, obtain the feature similarity between each data scatter point and each small classification in each clustering cluster and the suspected target scatter points of each small classification.
[0059] Finally, the data points in each clustering cluster can be divided into multiple small classifications and multiple undefined scatter points. These scatter points may contain the characteristics of multiple categories. For example, blood pressure detection may simultaneously involve the characteristics of blood detection and heart detection. In order to store these data more orderly and ensure that each data point is accurately classified, it is necessary to further classify these data scatter points with unclear boundaries.
[0060] Use a method based on distance and classification consistency to classify these data scatter points and judge the feature similarity between all undefined scatter points and small classifications. Specifically, if the distance between a data scatter point and the center point of a certain small classification is relatively close and the classification consistency with this small classification is relatively high, it can be considered that this data scatter point is more likely to belong to this small classification.
[0061] This embodiment gives a method for calculating the feature similarity between the i-th data scatter point and the j-th small classification in the v-th clustering cluster:
[0062] For each small classification in the v-th clustering cluster, calculate the average position of all data points in the sample space of each small classification, and record the average value as the center point of each small classification.
[0063] According to the method of calculating the classification consistency that the k-th data point and the y-th data point in the same clustering cluster belong to the same small classification, obtain the classification consistency that the i-th data scatter point and the center point of the j-th small classification in the same clustering cluster belong to the same small classification.
[0064]
[0065] In the formula, W v,i,j is the feature similarity between the i-th data scatter point and the j-th small classification in the v-th clustering cluster, l v,i,j is the Euclidean distance between the i-th data scatter point and the center point of the j-th small classification in the v-th clustering cluster, Q v,i,j is the classification consistency that the i-th data scatter point and the center point of the j-th small classification in the v-th clustering cluster belong to the same small classification, max(Q v,i,j) is the maximum value of the classification consistency that the i-th data scatter point in the v-th cluster and all data points in the j-th sub-classification belong to the same sub-classification.
[0066] Among the feature similarities between the i-th data scatter point in the v-th cluster and all sub-classifications, select a sub-classification corresponding to the maximum feature similarity, and regard the i-th data scatter point as a suspected target scatter point of the sub-classification corresponding to the maximum feature similarity. It should be noted that when there are multiple sub-classifications corresponding to the maximum feature similarity in this embodiment, any one of the sub-classifications can be selected for analysis.
[0067] Step S004: According to the changes of each sub-classification before and after the addition of suspected target scatter points in each sub-classification of each cluster, and the feature similarities between each data scatter point and each sub-classification, obtain the similarity of each sub-classification in each cluster and each suspected target scatter point in each sub-classification; according to the similarity of each sub-classification in each cluster and each suspected target scatter point in each sub-classification, determine whether the suspected target scatter points in each sub-classification of each cluster should be added.
[0068] In cluster analysis, it is crucial to maintain the consistency of information description. When new information causes deviations, screening is required to maintain the accuracy of clustering. In the sample space, the distribution of similar data points shows regularity, and the degree of conformity of a single data scatter point with a sub-classification determines its classification. If the changes of the sub-classification after the addition of the data scatter point are similar to those before, and the overall direction and area changes are small, then the data scatter point is very likely to belong to this sub-classification. Such precise classification helps to improve the accuracy of data analysis and the consistency of clustering.
[0069] The method for calculating the similarity between the suspected target scatter point of the j-th sub-classification in the v-th cluster and the j-th sub-classification in this embodiment is as follows:
[0070] Obtain the projection difference on the principal components before and after the addition of the f-th suspected target scatter point of the j-th sub-classification in the v-th cluster by the principal component analysis method to quantify the change value of the direction, and record the change value as the change value of the overall direction of the j-th sub-classification in the v-th cluster after adding the f-th suspected target scatter point to the j-th sub-classification. Among them, the principal component analysis method is a well-known technology and will not be elaborated here.
[0071] According to the acquisition method of the center point of each sub-classification in each cluster, obtain the center point of each sub-classification after adding the suspected target scatter point, and record the center point after adding the f-th suspected target scatter point of the j-th sub-classification to the j-th sub-classification as the updated center point of the j-th sub-classification.
[0072] Calculate the absolute value of the difference between the distance from the f-th suspected target scatter point in the j-th sub-category to the center point of the j-th sub-category and the distance from the f-th suspected target scatter point to the updated center point of the j-th sub-category, and denote the absolute value of the difference in distance as the change value of the spatial size of the j-th sub-category in the v-th clustering cluster after adding the f-th suspected target scatter point to the j-th sub-category.
[0073]
[0074] In the formula, R v,f,j is the similarity of the j-th sub-category in the v-th clustering cluster to the f-th suspected target scatter point in the j-th sub-category, and W v,f,j is the feature similarity of the j-th sub-category in the v-th clustering cluster to the f-th suspected target scatter point in the j-th sub-category, and S v,f,j is the change value of the overall direction of the j-th sub-category in the v-th clustering cluster after adding the f-th suspected target scatter point to the j-th sub-category, and θ v,f,j is the change value of the spatial size of the j-th sub-category in the v-th clustering cluster after adding the f-th suspected target scatter point to the j-th sub-category, and N v,j,f is the number of suspected target scatter points other than the f-th suspected target scatter point in the j-th sub-category in the v-th clustering cluster, and S j,m,v is the change value of the spatial size of the j-th sub-category in the v-th clustering cluster after adding the m-th suspected target scatter point to the j-th sub-category, and θ j,m,v is the change value of the overall direction of the j-th sub-category in the v-th clustering cluster after adding the m-th suspected target scatter point to the j-th sub-category, and || represents taking the absolute value.
[0075] In the above formula, S v,f,j ×θ v,f,j is the distribution difference of the sub-category range before and after adding the suspected target scatter point. is the difference sum of the distribution differences of the sub-category range before and after adding the f-th suspected target scatter point and the suspected target scatter points other than the f-th suspected target scatter point in the j-th sub-category. The smaller the value, the more similar the features of the f-th suspected target scatter point are to the other added suspected target scatter points, and the more likely it belongs to this sub-category.
[0076] When the similarity of the j-th sub-category in the v-th clustering cluster to the f-th suspected target scatter point in the j-th sub-category is greater than the preset threshold, add the f-th suspected target scatter point in the j-th sub-category to the j-th sub-category in the v-th clustering cluster. Thus, the updated classification of each sub-category in the v-th clustering cluster is obtained.
[0077] Step S005: Obtain useless data classification according to the number of small categories in each cluster, the distance between the updated center points of each small category, the number of data points in each small category and the number of suspected target scattered points finally added; eliminate the data in the useless data classification and complete data storage.
[0078] The useless data classification has a different distribution in the sample interval compared to the useful data classification. Because there are fewer useless data than useful data, and because the difference between useless data and useful data is large, the useless data will be distributed closer to the edge in the cluster and farther away from other useful data. The characteristics of different data points contained in the useless data classification are generally very different, and most of them become data scatter points during the initial classification. Therefore, the data points in the useless data classification are more likely to be added later as data scatter points. Therefore, the sub-classification can be screened for useless data classification based on the amount of data in the sub-classification, the distance from other sub-classifications, and the similarity of the data scatter points in the sub-classification and the sub-classification.
[0079] This implementation provides a method for calculating the uselessness of the e-th updated classification data in the v-th cluster as follows:
[0080]
[0081] Where U v,e The uselessness of updating the classification data for the eth in the vth cluster, n v,e Update the number of all data points in the e-th updated classification in the v-th cluster, n v is the number of updated categories in the vth cluster, l v,e,h is the distance between the center points of the e-th updated classification and the h-th updated classification in the v-th cluster, N v,e is the number of scattered data points in the e-th updated classification in the v-th cluster, R v,e,b is the homogeneity between the small category corresponding to the e-th updated category in the v-th cluster and the b-th data scatter point in the e-th updated category, and Sigmoid() is the normalization function.
[0082] When the uselessness of the e-th update classification in the v-th cluster is greater than a preset threshold, it is determined that the e-th update classification belongs to the useless data classification. In this embodiment, the preset threshold is 0.5, which is used as an example for description.
[0083] Eliminate useless data categories, use hash storage method to store the remaining useful data, and use a single updated category as the amount of data stored at one time to complete data storage.
[0084] It should be noted that in this embodiment, through density clustering analysis, the data points in different clustering clusters are reclassified, and the useless data classifications are identified. After removing these useless data, the total data volume and the data transmission volume of a single time length can be effectively reduced, the security of the stored data can be improved, and data loss can be avoided. This storage scheme is not only economical and efficient, but also easy to expand and suitable for processing large-scale data sets.
[0085] It should be noted that: in this embodiment, when the denominator in the formula is 0, the denominator is set to 1 to ensure the formula holds, and this is used as an example for description.
[0086] The present invention also provides a personal information security storage system based on the Internet of Things, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program stored in the memory to implement the steps of the aforementioned personal information security storage method based on the Internet of Things.
[0087] Thus far, the present invention is completed.
[0088] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for securely storing personal information based on the Internet of Things, characterized in that The method comprises the following steps: Obtain a sample space of all information data points to be stored, and obtain several clusters of data points in the sample space; According to the distance between the data points in each cluster and the distance from the data point to the center point of the cluster, the classification consistency of any two data points in each cluster belonging to the same small category is obtained; according to the classification consistency of any two data points in each cluster belonging to the same small category, all the data points in each cluster are clustered to obtain several small categories and data scatter points in each cluster, and the small category contains several data points; According to the distance between each data scatter point in each cluster and the center point of each small classification, and the classification consistency that each data scatter point and the center point of each small classification belong to the same small classification, the feature similarity between each data scatter point in each cluster and each small classification and the suspected target scatter points of each small classification are obtained; According to the changes of each small classification before and after the suspected target scattered points of each small classification in each cluster are added, and the feature similarity between each data scattered point and each small classification, the similarity between each small classification in each cluster and each suspected target scattered point in each small classification is obtained; according to the similarity between each small classification in each cluster and each suspected target scattered point in each small classification, it is determined whether the suspected target scattered points of each small classification in each cluster are added; Obtain useless data classification according to the number of small classifications in each cluster, the distance between the updated center points of each small classification, the number of data points in each small classification and the number of suspected target scattered points finally added; remove the data in the useless data classification, use the hash storage method to store the remaining useful data, and use a single updated classification as the amount of data stored at one time to complete the data storage; The specific steps of judging whether to add the suspected target scattered points of each small category in each cluster cluster according to the similarity of each small category in each cluster cluster and each suspected target scattered point in each small category are as follows: When the similarity between the th sub-category in the th cluster and the th sub-category in the th suspected target scatter point is greater than the preset threshold, add the th suspected target scatter point in the th sub-category to the th sub-category in the th cluster to obtain the updated classification of each sub-category in the th cluster; The specific steps of obtaining useless data classification according to the number of small categories in each cluster, the distance between the updated center points of each small category, the number of data points in each small category and the number of suspected target scattered points finally added are as follows: In the formula, is the uselessness of the th updated classification data in the th clustering cluster, is the number of all data points in the th updated classification in the th clustering cluster, is the number of updated classifications in the th clustering cluster, is the distance between the th and th updated classifications and the center point of the th updated classification in the th clustering cluster, is the number of data scatter points in the th updated classification in the th clustering cluster, is the similarity between the sub - classification corresponding to the th updated classification and the th data scatter point in the th updated classification in the Obtain useless data classification based on the uselessness of each updated classification data in each cluster.
2. The method for securely storing personal information based on the Internet of Things according to claim 1, wherein, The method of obtaining the classification consistency of any two data points in each cluster belonging to the same sub-classification according to the distance between the data points in each cluster and the distance from the data point to the center point of the cluster includes the following specific steps: In the formula, is the classification consistency that the th data point and the th data point in the th clustering cluster belong to the same sub - classification, is the Euclidean distance in the sample space between the th data point and the th data point in the th clustering cluster, is the distance between the th data point and the center point of the th clustering cluster in the th clustering cluster, is the distance between the th data point and the center point of the th clustering cluster in the th clustering cluster, means taking the absolute value, is the normalization function.
3. The method for securely storing personal information based on the Internet of Things according to claim 1, characterized in that According to the distance between each data scatter point in each cluster and the center point of each small classification, and the classification consistency that each data scatter point and the center point of each small classification belong to the same small classification, the feature similarity between each data scatter point in each cluster and each small classification and the suspected target scatter point of each small classification are obtained, and the specific steps include the following: In the formula, is the feature similarity between the -th data scatter point in the -th cluster and the -th sub-category; is the Euclidean distance between the -th data scatter point in the -th cluster and the center point of the -th sub-category; is the classification consistency that the -th data scatter point in the -th cluster and the center point of the -th sub-category belong to the same sub-category; is the maximum value among the classification consistencies that all data points of the -th data scatter point in the -th cluster and the -th sub-category belong to the same sub-category; According to the feature similarity between each data scatter point and each small category, the suspected target scatter points of each small category are obtained.
4. The method for securely storing personal information based on the Internet of Things according to claim 3, characterized in that, The specific steps of obtaining the suspected target scattered points of each small category according to the feature similarity between each data scattered point and each small category are as follows: Among the feature similarities between the th data scatter point in the th clustering cluster and all small classifications, select one small classification corresponding to the maximum feature similarity, and use the th data scatter point as the suspected target scatter point of the small classification corresponding to the maximum feature similarity.
5. The personal information security storage method based on the Internet of Things according to claim 1, wherein Obtain the similarity between each sub - classification in each cluster and each suspected target scatter point in each sub - classification according to the changes of each sub - classification before and after adding the suspected target scatter points in each sub - classification, and the feature similarity between each data scatter point and each sub - classification. The specific steps are as follows: Obtain the projection difference of the th small classification in the th clustering cluster on the principal components before and after adding the th suspected target scatter point to quantify the change value of the direction, and record the change value as the change value of the overall direction of the th small classification after adding the th suspected target scatter point in the th clustering cluster and the th small classification in the th clustering cluster; Add the th suspected target scatter point of the th small classification to the center point after the th small classification, denoted as the updated center point of the th small classification; Calculate the distance from the th suspected target scatter point of the th small classification to the center point of the th small classification, and the absolute value of the difference between the distance to the updated center point of the th small classification. Denote the absolute value of the distance difference as the change value of the th small classification in the th clustering cluster after adding the th suspected target scatter point to the spatial size of the th small classification; According to the change value of the overall direction and the change value of the spatial size of each sub - classification in each cluster after adding each suspected target scatter point, and the feature similarity between each data scatter point and each sub - classification, obtain the similarity between each sub - classification in each cluster and each suspected target scatter point in each sub - classification.
6. The method for securely storing personal information based on the Internet of Things according to claim 5, characterized in that The method of obtaining the similarity between each sub - classification in each cluster and each suspected target scatter point in each sub - classification according to the change value of the overall direction and the change value of the spatial size of each sub - classification in each cluster after adding each suspected target scatter point, and the feature similarity between each data scatter point and each sub - classification, includes the following specific steps: Wherein, is the similarity of the th sub-classification in the th clustering cluster and the th sub-classification and the th suspected target scatter point, is the feature similarity of the th sub-classification in the th clustering cluster and the th sub-classification and the th suspected target scatter point, is the change value of the overall direction of the th sub-classification in the th clustering cluster after adding the th suspected target scatter point, is the change value of the overall direction of the is the change value of the spatial size of the th sub-classification in the th clustering cluster after adding the th suspected target scatter point, is the change value of the spatial size of the is the number of suspected target scatter points except the th sub-classification in the th clustering cluster except the th suspected target scatter point, is the change value of the spatial size of the th sub-classification in the th clustering cluster after adding the th suspected target scatter point, is the change value of the overall direction of the is the change value of the overall direction of the th sub-classification in the th clustering cluster after adding the th suspected target scatter point, is the change value of the overall direction of the is to take the absolute value.
7. The method for securely storing personal information based on the Internet of Things according to claim 6, characterized in that Obtain the useless data classification according to the uselessness of each updated classification data in each cluster. The specific steps are as follows: When the uselessness of the th updated classification in the th clustering cluster is greater than the preset threshold, it is determined that the th updated classification belongs to the useless data classification.
8. An Internet of Things-based personal information security storage system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for secure personal information storage based on the Internet of Things according to any one of claims 1 - 7.
Citation Information
Patent Citations
Method and device for warning about abnormal behavior
CN109509327A
Network abnormal flow analysis method and system based on Spark and clustering
CN112511547A