High-dimensional multi-target Internet of Things data preprocessing method with privacy protection function

Through the federated learning model and multi-objective preprocessing optimization model, the feature selection problem of high-dimensional data in the Internet of Things is solved, privacy protection and efficient data preprocessing are achieved, and analysis accuracy and efficiency are improved.

CN120336793AInactive Publication Date: 2025-07-18NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510813967.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot effectively select feature when processing data in the Internet of Things with high-dimensional, high redundant and privacy protection requirements, resulting in inaccurate data analysis results and high computing resources consumption.

Method used

Using the federated learning model, the local subset of each participant is extracted by confidence participants, forming a target global feature set, and a multi-objective preprocessing optimization model is built to solve it to obtain the target feature subset, and realize privacy protection and efficient data preprocessing.

Benefits of technology

On the premise of protecting data privacy, the accuracy and reliability of IoT data analysis are improved, the computational complexity is reduced, and the preprocessing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336793A_ABST
    Figure CN120336793A_ABST
Patent Text Reader

Abstract

The invention discloses a high-dimensional multi-target Internet of Things data preprocessing method with a privacy protection function, and the method comprises the steps: extracting a local feature subset of participation sub-data of participants in the Internet of Things based on a federated learning mode, the confidence participant in the Internet of Things determines a target global feature set of the target data in the Internet of Things according to the local feature subset of the participant; according to the target global feature set of the target data, constructing a multi-target preprocessing optimization model of the target data; and solving the multi-target preprocessing optimization model to obtain a target feature subset of the target data. According to the method provided by the invention, the high-dimensional data which is owned by different participants and cannot be directly shared with each other in the Internet of Things can be processed with high quality on the premise of protecting data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data preprocessing, and in particular to a high-dimensional multi-objective Internet of Things (IoT) data preprocessing method with privacy protection function. Background Art

[0002] With the rapid development of IoT technology, the amount of data generated by participants in the IoT, such as sensor networks and IoT devices, has increased explosively, and the dimension of the data has also increased rapidly, making the data in the existing IoT exhibit the characteristics of large quantity and high dimension, which contains a large amount of redundant or irrelevant data. Moreover, in many cases, the same type of data in the IoT is owned by different participants and cannot be directly shared with each other due to privacy protection functions. If the data analysis methods of data science and machine learning are directly used to analyze IoT data with such characteristics, it will consume a large amount of computing resources and affect the subsequent analysis results.

[0003] In related technologies, feature selection methods are used to preprocess IoT data, but most of them take the number of features and classification accuracy as a single optimization goal. In the feature selection process, not only the privacy protection function of participants in the IoT is not considered, but also other factors in the feature selection process are ignored, such as feature selection cost, classification recall rate, and classification error. This makes the finally selected feature subset unable to truly reflect the characteristics of the data, difficult to improve the quality of IoT data, and thus affect the analysis results of IoT data. Summary of the Invention

[0004] In view of this, this application provides a high-dimensional multi-objective IoT data preprocessing method with privacy protection function, which can process a class of high-dimensional data owned by different participants in the IoT and cannot be directly shared with each other with high quality under the premise of protecting data privacy.

[0005] According to one aspect of this application, a high-dimensional multi-objective IoT data preprocessing method with privacy protection function is provided, including: Based on the federated learning mode, extract the local feature subsets of the participating sub-data of the participants in the IoT, so that the trusted participants in the IoT can determine the target global feature set of the target data in the IoT according to the local feature subsets of the participants. The target data is the data in the IoT belonging to the target type, the participating sub-data is the data owned by the participant in the target data, the trusted participant is an independent trusted device, and the trusted participant can communicate with the participant; Construct a multi-objective preprocessing optimization model for the target data according to the target global feature set of the target data; Solve the multi-objective preprocessing optimization model to obtain the target feature subset of the target data.

[0006] According to another aspect of the present application, there is provided a high-dimensional multi-objective Internet of Things data preprocessing device with a privacy protection function, including: A federation module, configured to extract a local feature subset of the participating sub-data of the participants in the Internet of Things based on the federated learning mode, so that the confident participants in the Internet of Things determine the target global feature set of the target data in the Internet of Things according to the local feature subset of the participants. The target data is the data belonging to the target type in the Internet of Things, the participating sub-data is the data owned by the participants in the target data, the confident participants are independent trusted devices, and the confident participants can communicate with the participants; A construction module, configured to construct a multi-objective preprocessing optimization model for the target data according to the target global feature set of the target data; A solving module, configured to solve the multi-objective preprocessing optimization model to obtain the target feature subset of the target data.

[0007] According to still another aspect of the present application, there is provided a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the above-mentioned high-dimensional multi-objective Internet of Things data preprocessing method with a privacy protection function are implemented.

[0008] According to yet another aspect of the present application, there is provided a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the steps of the above-mentioned high-dimensional multi-objective Internet of Things data preprocessing method with a privacy protection function are implemented.

[0009] By means of the above technical solutions, the present application provides a high-dimensional multi-objective Internet of Things data preprocessing method with a privacy protection function. For a type of target data owned by different participants in the Internet of Things, based on the federated learning mode, each participant extracts its own local feature subset from the participating sub-data it locally owns. The confident participants collect the local feature subsets of each participant on their respective participating sub-data, and form the target global feature set of this type of data in the Internet of Things according to the collected local feature subsets. There is no direct information interaction between the participants, realizing the privacy protection function. Further, according to the target global feature set of this type of data in the Internet of Things, a multi-objective preprocessing optimization model for this type of data is constructed, and the multi-objective preprocessing optimization model is solved to efficiently search for the solution of the multi-objective preprocessing optimization model, reduce the computational complexity, improve the preprocessing efficiency, so as to obtain diverse target feature subsets of the same type of high-dimensional data in the Internet of Things, enabling the target feature subsets to comprehensively and truly reflect the characteristics of the same type of data in the Internet of Things, and improving the accuracy and reliability of subsequent analysis.

[0010] The above description is only an overview of the technical solution of the present application. In order to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Brief Description of the Drawings

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 A flowchart showing the high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function provided by an embodiment of the present application; Figure 2 A structural block diagram showing the high-dimensional multi-objective Internet of Things data preprocessing device with privacy protection function provided by an embodiment of the present application. Detailed Embodiments

[0012] The present application will be described in detail below with reference to the drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0013] The embodiments of the present application are described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0014] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "joined" to another element, it can be directly connected or joined to other elements, or there may also be intermediate elements. In addition, the "connection" or "joining" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0015] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. It should be understood that these embodiments are provided so that the disclosure of the present application is thorough and complete, and the concept of these exemplary embodiments is fully conveyed to those of ordinary skill in the art.

[0016] In this embodiment, a high-dimensional multi-objective Internet of Things data preprocessing method with a privacy protection function is provided. As Figure 1 shown, the method includes: Step 101, based on the federated learning mode, extract the local feature subsets of the participating sub-data of the participants in the Internet of Things, so that the trusted participants in the Internet of Things can determine the target global feature set of the target data in the Internet of Things according to the local feature subsets of the participants.

[0017] Among them, the target data is the data belonging to the target type in the Internet of Things, the participating sub-data is the data owned by the participants in the target data, and the trusted participants are independent trusted devices that can communicate with the participants.

[0018] It should be noted that in the Internet of Things environment, a participant refers to an entity that owns or generates data in the Internet of Things system, such as sensors, smart devices, institutions, or organizations, etc.

[0019] In this embodiment, for a type of data (i.e., target data) in the Internet of Things that has a large amount of data, high dimensions, is owned by different participants, and cannot be directly shared due to privacy protection functions, first, let the participants extract their own local feature subsets on the participating sub-data they own in the target data. Then, a trusted participant trusted by the participants is introduced. The trusted participant is an independent trusted device that can communicate with the participants, so that the trusted participant can learn and integrate the local feature subsets of different participants to form the target global feature set of this type of data in the Internet of Things, realizing the preliminary feature selection of the same type of data in the Internet of Things.

[0020] Specifically, for example, the target type is determined according to the data type of the same type of data in the Internet of Things in the actual application scenario. For example, in intelligent healthcare, the patient diagnosis data of a certain disease owned by different hospitals. Due to the protection of patient privacy, hospitals cannot directly share the relevant data and can only provide some features of the diagnosis data. Or in an intelligent factory, the device status data collected by the production equipment sensors of multiple factories. Or in environmental monitoring, the air quality data of the same area monitored by different meteorological stations. Or in a smart home, the energy consumption data collected by the smart devices of multiple families.

[0021] In this embodiment, there is no need for participants to share the original data. Instead, only the features need to be shared by the participants, reducing the risk of privacy leakage. The trusted participant acts as an intermediate coordinator, responsible for feature integration and screening, avoiding direct contact between participants with sensitive information, and reducing the trust cost in collaboration.

[0022] Step 102: Construct a multi-objective preprocessing optimization model for the target data according to the target global feature set of the target data.

[0023] Step 103: Solve the multi-objective preprocessing optimization model to obtain the target feature subset of the target data.

[0024] In this embodiment, according to the target global feature set of a certain type of data in the Internet of Things, a multi-objective preprocessing optimization model for this type of data is constructed. By solving the multi-objective preprocessing optimization model, a target feature subset that can truly reflect the characteristics of this type of data is obtained, improving the accuracy and reliability of subsequent analysis.

[0025] In this application, for the situation where a certain type of data in the Internet of Things is owned by different participants, based on the federated learning mode, each participant extracts its own local feature subset from the local participating sub-data it owns. The trusted participant collects the local feature subsets of each participant on their respective participating sub-data, and forms the target global feature set of this type of data in the Internet of Things according to the collected local feature subsets. There is no direct information interaction between the participants, realizing the privacy protection function. Further, according to the target global feature set of this type of data in the Internet of Things, a multi-objective preprocessing optimization model for this type of data is constructed, and the multi-objective preprocessing optimization model is solved to efficiently search for the solution of the multi-objective preprocessing optimization model, reducing the computational complexity, improving the preprocessing efficiency, and thus obtaining a diverse target feature subset of the same type of high-dimensional data in the Internet of Things, enabling the target feature subset to comprehensively and truly reflect the characteristics of the same type of data in the Internet of Things, and improving the accuracy and reliability of subsequent analysis.

[0026] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, a local feature subset of the participation sub-data of the participants in the Internet of Things is extracted, which specifically includes: The participant initializes the position of the particle according to the initial feature subset of the participant's participation sub-data and the initial global feature set of the target data to form a particle population, where the particle is a candidate feature subset of the participant's participation sub-data, and the position of the particle is used to represent the selection probability of the known features of the target data by the participant. The initial feature subset of the participant and the initial global feature set of the target data are determined according to the known features; The participant evaluates the particle using a preset classifier according to the participant's participation sub-data to determine the first evaluation result of the particle; The participant determines the optimal position of the particle and the population optimal position of the particle population according to the first evaluation result of the particle; The participant updates the position of the particle according to the optimal position of the particle and the population optimal position to update the particle population until the particle population meets the preset termination condition; The participant determines the population optimal position of the particle population that meets the preset termination condition as the local feature subset of the participant's participation sub-data.

[0027] In this embodiment, a feature integrator is constructed according to the participants and the confidence participants in the Internet of Things. Each participant in the feature integrator performs a search operation of particle swarm optimization with fast search advantages on their respective local participation sub-data to extract their respective feature subsets. The confidence participant in the feature integrator collects the feature subsets of each participant and determines the global feature set of this type of data (i.e., the target data) in the Internet of Things according to the feature subsets of each participant.

[0028] Specifically, this embodiment runs the feature integrator twice. When the feature integrator runs for the first time, the participant randomly initializes the position of the particle with a real number between 0 and 1, and randomly selects a classifier from the classification pool including random forest, support vector machine, k-nearest neighbors (KNN), logistic regression, and decision tree as the preset classifier, and uses the preset classifier to evaluate the fitness of each particle to enhance the generality and adaptive ability, so that the participant can perform the search operation of particle swarm optimization on their own participation sub-data to obtain their own initial feature subset. The confidence participant determines the initial global feature set of this type of data in the Internet of Things from the initial feature subsets of each participant.

[0029] It should be noted that the target data in the Internet of Things has a large number of known features, and these known features have a fixed order. Each real number in the position of the particle corresponds to the selection probability of a known feature, so the particle is used as a candidate feature subset of the participant. The participant's participation sub-data includes multiple samples, and each sample includes the specific data corresponding to each known feature and the label corresponding to the sample.

[0030] Specifically, in intelligent healthcare, different hospitals (i.e., participants) jointly own the diagnostic data of diabetic patients (i.e., target data). The diagnostic data of each hospital (i.e., participating sub-data) includes the specific diagnostic data of each patient (i.e., sample) (the specific data of the known features corresponding to diabetes) and the disease type (i.e., label). Suppose there are 4 known features with a fixed order for diabetes, namely age, blood pressure, body temperature, and blood glucose level. A hospital randomly initializes particles in the form of real numbers between 0 and 1. The initial position of the particle may be [0.8, 0.4, 0, 0.7]. Here, 0.8, 0.4, 0, and 0.7 represent that the selection probabilities of the hospital for the four known features of age, blood pressure, body temperature, and blood glucose level are 80%, 40%, 0, and 70% respectively. Then, the hospital will also decode the initialized particle into a binary form to update the position of the initialized particle. Specifically, if the real number in the position of the particle is greater than 0.5, the real number is decoded into 1; if the real number in the position of the particle is less than or equal to 0.5, the real number is decoded into 0. Here, 1 represents selection and 0 represents non-selection. For example, the particle [0.8, 0.4, 0, 0.7] will be decoded into the binary form 1001, which means that the hospital selects age and blood glucose level among the known features as its own candidate feature subset.

[0031] It can be understood that after the participants obtain their own initial feature subsets using the particle swarm optimization algorithm, the confident participant determines the initial global feature set from the initial feature subsets of each participant. For example, the initial global feature set is 110011, which means that the confident participant selects the 1st, 2nd, 5th, and 6th known features as the initial global feature set among all 6 known features with a fixed order of the target data in the Internet of Things, while the 3rd and 4th known features are not selected.

[0032] Furthermore, in this embodiment, the feature integrator is run again. During the process of running the feature integrator again, the participant uses the initial global feature set to guide the initialization of the particles, so as to optimize the search direction of the particle swarm optimization, reduce the interference of redundant features, and improve the search efficiency. Then, the participant uses the preset classifier to evaluate the particles according to their own participation sub-data, and determines the first evaluation result of the particles. Then, according to the first evaluation result of the particles, the participant updates the optimal position of the particles and the population optimal position of the particle population formed by the particles, so as to guide the particles to move in the direction of higher evaluation results through the feedback of local and global optimal solutions. Based on the optimal position of the particles and the population optimal position, the participant updates the positions of the particles to continue searching for better solutions and gradually approach the solution that can most truly reflect the participant's participation sub-data. Moreover, the participant performs a mutation operation on the particles to randomly perturb the positions of some particles, such as randomly flipping the selection status of some features, introducing diversity, avoiding the particle swarm optimization algorithm from falling into local optimality, and enhancing the search ability of the particle swarm optimization algorithm on the participant's participation sub-data. Then, the participant determines whether the entire particle population meets the preset termination conditions, such as reaching the maximum number of iterations or the first evaluation result of the particles converges. If not, the participant re-evaluates the particles using the preset classifier according to their own participation sub-data to update the first evaluation result of the particles until the particle population meets the preset termination conditions. If so, the participant takes the population optimal position in the particle population that meets the preset termination conditions as its own local feature subset.

[0033] Here, the participant takes the classification accuracy of the particles on their participation sub-data as the first evaluation result of the particles. Specifically, the participant divides their participation sub-data into a training set and a validation set. The particles are used as the candidate feature subsets of the participant. The participant finds the data corresponding to the particles and the labels corresponding to this part of the data in the training set of their participation sub-data to form the first training sample, and uses the first training sample to train the preset classifier so that the preset classifier learns the features selected by the particles. Then, the participant uses the preset classifier trained with the particles to classify the validation set of their participation sub-data, and calculates the classification accuracy of the preset classifier under the training of the particles on the validation set of their participation sub-data to obtain the first evaluation result of the particles on the participant's participation sub-data.

[0034] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, the positions of the particles are initialized according to the initial feature subset of the participant's participation sub-data and the initial global feature set of the target data, specifically including: comparing the initial feature subset of the participant with the initial global feature set of the target data; if the comparison result is consistent, initializing the positions of the particles according to the initial global feature set of the target data; if the comparison result is inconsistent, initializing the positions of the particles according to the initial feature subset of the participant and the initial global feature set of the target data.

[0035] Exemplarily, first, the participant compares the initial feature subset of himself with the initial global feature set of the target data in the Internet of Things . If the initial feature subset of the participant himself is the initial global feature set (that is, the comparison result is consistent), during the process of running the feature integrator again, the participant uses the initial global feature set to initialize the positions of his particles according to the following formula: ; where, is the selection probability of the k -th known feature in the position of the p -th particle of the participant, is the selection probability of the -th known feature in the initial global feature set p , p = 1, 2, 3,..., n , n is the number of known features, is a random number between 0 and 1.

[0036] Here, for a participant whose initial feature subset is the initial global feature set , for any known feature p of the target data in the Internet of Things, if the known feature is selected in the initial global feature subset p , that is, , this means that the known feature p is important to some extent. When initializing the initial particles during the process of running the feature integrator again, through the above formula, the known feature p can have a higher probability of being selected into the next global feature set. When the known feature is not selected in the initial global feature set p , that is, , this means that this known feature pIt is likely to be a redundant feature. When initializing particles during the process of running the feature integrator again, through the above formula, the known features p can have a lower probability of being selected into the next global feature set.

[0037] If the initial feature subset of the participant itself is not the initial global feature set , during the process of running the feature integrator again, the participant initializes the position of its particles based on its own initial feature subset and the initial global feature set according to the following formula: ; where is the selection probability of l known features in the initial feature subset of the participant p for

[0038] Here, for other participants whose initial feature subsets are not the initial global feature set l , for any known feature p of the target data in the Internet of Things, if the known feature is not selected in the initial global feature set p , while the known feature l is selected in the initial feature subset of the participant p , that is , the quality of the known feature p is difficult to determine. For this situation, when initializing particles during the process of running the feature integrator again, randomly initialize the selection probability of the known feature p . For other situations, for example, if the known feature is selected in the initial global feature set p , that is , this means that the known feature p is important to some extent. When initializing particles during the process of running the feature integrator again, through , the known feature p can have a higher probability of being selected into the next global feature set. When the known feature is not selected in both the initial feature subset and the initial global feature set p , that is , this indicates that the known feature p is very likely to be a redundant feature. Through , the known feature p will only have a very small chance of being selected into the next global feature set.

[0039] In this embodiment, according to the prior value provided by the initial global feature set, the search efficiency of participants on their own participating sub-data is improved, redundant calculations are reduced, and the convergence speed of the particle swarm optimization algorithm is accelerated. For example, when a hospital searches for features, it directly starts from the globally important features without randomly exploring from the entire feature space.

[0040] Further, as a refinement and extension of the specific implementation manner of the above embodiment, to fully illustrate the specific implementation process of this embodiment, the trusted participant determines the target global feature set of the target data in the Internet of Things according to the local feature subset of the participant, specifically including: the participant sends the local feature subset of the participant to the trusted participant; the trusted participant sends the local feature subsets of all received participants to the participant; the participant uses a preset classifier to evaluate the received local feature subsets of the participant on the participant's participating sub-data, and determines the second evaluation result of the local feature subset of the participant; the participant sends the second evaluation result of the local feature subset of the participant to the trusted participant; the trusted participant determines the average evaluation result of the local feature subset of the participant according to the second evaluation result of the received local feature subset of the participant; the trusted participant takes the local feature subset with the largest average evaluation result as the target global feature set.

[0041] In this embodiment, based on the federated learning mode, the trusted participant collects the local feature subsets of each participant on their respective participating sub-data, and sends the collected all local feature subsets to each participant respectively. When a participant receives the local feature subsets of other participants, each participant evaluates all the local feature subsets on their respective participating sub-data, and feeds back the second evaluation result of each local feature subset to the trusted participant. After the trusted participant receives all the feedbacks from each participant, it calculates the average evaluation result of each local feature subset, and determines the target global feature set from the local feature subsets according to the average evaluation result.

[0042] Exemplarily, after extracting the local feature subset on its own participating sub-data, the participant sends its local feature subset to the trusted participant. After the trusted participant receives the local feature subsets sent by each participant, it sends the received all local feature subsets to each participant respectively. The participant uses a preset classifier to evaluate each received local feature subset on its participating sub-data, determines the second evaluation result of each received local feature subset on its participating sub-data, and sends the second evaluation result of each local feature subset to the trusted participant. The trusted participant receives the second evaluation result of the local feature subset, and determines the target global feature set according to the following formula: ; where is the target global feature set, Mis the total number of participants, is a participant i of the local feature subset in the participant j of the participation sub - data on the second evaluation result.

[0043] Here, for each local feature subset, the confident participant calculates the average value of its second evaluation results on the participation sub - data of all participants to obtain the average evaluation result of the local feature subset. The confident participant selects the local feature subset with the largest average evaluation result as the target global feature set.

[0044] It should be noted that the process of a participant using a preset classifier to evaluate the received local feature subset on its participation sub - data is the same as the process of a participant using a preset classifier to evaluate particles according to its participation sub - data. Specifically, the participant divides its participation sub - data into a training set and a validation set. The participant finds the data corresponding to the received local feature subset and the labels corresponding to this part of the data in the training set of its participation sub - data to form a second training sample, and uses the second training sample to train the preset classifier so that the preset classifier learns the features selected by the received local feature subset. Then, the participant uses the preset classifier trained with the received local feature subset to classify the validation set of its participation sub - data, and calculates the classification accuracy of the preset classifier trained with the received local feature subset on the validation set of its participation sub - data to obtain the second evaluation result of the received local feature subset on its participation sub - data.

[0045] In this embodiment, the learning and integration process of the local feature subsets of different participants only occurs in the confident participant, and the information interaction only occurs between the confident participant and each participant, and there is no direct information interaction between the participants. In addition, the information interaction between the confident participant and each participant is only limited to the feature subset and the classification accuracy, so as to achieve the privacy protection of the data.

[0046] Further, as a refinement and extension of the specific implementation manner of the above embodiment, to fully illustrate the specific implementation process of this embodiment, step 102, that is, according to the target global feature set of the target data, a multi-objective preprocessing optimization model of the target data is constructed, which specifically includes: determining an optimized feature subset of the target global feature set according to the target global feature set of the target data; constructing a first objective function with the goal of minimizing the number of features in the optimized feature subset; evaluating the optimized feature subset using a preset classifier according to the target sub-data, and determining a third evaluation result of the optimized feature subset, where the target sub-data is determined according to the participating sub-data corresponding to the target global feature set; respectively constructing a second objective function, a third objective function, and a fourth objective function with the goal of minimizing the third evaluation result; constructing a fifth objective function with the goal of minimizing the feature selection cost of the optimized feature subset; and constructing a multi-objective preprocessing optimization model according to the first objective function, the second objective function, the third objective function, the fourth objective function, and the fifth objective function.

[0047] Among them, the optimized feature subset is a candidate solution of the multi-objective preprocessing optimization model, and is used to represent the selection probability of the selected features in the target global feature set. The selected features are the known features of the target data with a selection probability of a preset value in the target global feature set.

[0048] In this embodiment, different from the traditional preprocessing that often focuses on a single or a few objectives, this embodiment makes the preprocessing process more conform to the complex and multi-dimensional characteristics of Internet of Things data through multi-objective optimization. Moreover, the data sources of the Internet of Things are extensive, with diverse and high-dimensional features. The multi-objective preprocessing optimization model can preprocess a more representative and low-redundancy target feature subset through multi-objective constraints. At the same time, the multi-objective preprocessing optimization model directly evaluates the classification effect of the optimized feature subset through a preset classifier, ensuring that the selected features can retain the key information of the original data and avoiding the decline of the model performance caused by simply reducing the number of features, so as to provide high-quality data for subsequent machine learning, decision analysis, etc., and improve the overall efficiency of the Internet of Things system. For example, real-time decision-making in intelligent transportation and precise analysis in environmental monitoring.

[0049] Further, as a refinement and extension of the specific implementation manner of the above embodiment, to fully illustrate the specific implementation process of this embodiment, evaluating the optimized feature subset using a preset classifier according to the target sub-data specifically includes: training the preset classifier according to the data corresponding to the optimized feature subset in the target sub-data; classifying the target sub-data using the trained preset classifier, and determining the classification accuracy, classification recall rate, and classification error of the preset classifier under the training of the optimized feature subset, and taking the negative of the classification accuracy, the negative of the classification recall rate, and the classification error of the preset classifier under the training of the optimized feature subset as the third evaluation result.

[0050] Among them, the negative value of the classification accuracy is used to construct the second objective function, the negative value of the classification recall rate is used to construct the third objective function, and the classification error is used to construct the fourth objective function.

[0051] In this embodiment, by optimizing the feature subset to train the preset classifier and calculating the classification accuracy, classification recall rate, and classification error, the classification effect of the optimized feature subset on the target sub-data is quantitatively evaluated from different perspectives, and the overall improvement of the classification performance is achieved through multi-objective optimization.

[0052] Specifically, first, among the known features selected from the target global feature set (i.e., the selected features, which are the known features with a selection probability of 1 (preset value) in the target global feature set), features are randomly selected, and the selection probabilities of the randomly selected features are randomly initialized with real numbers between 0 and 1. At the same time, binary decoding is performed to generate an optimized feature subset.

[0053] It can be understood that the optimized feature subset is a candidate solution for the subsequent multi-objective preprocessing optimization model to be established.

[0054] Then, the preset classifier is used to analyze the optimized feature subset to construct a multi-objective preprocessing optimization model for the target data.

[0055] Specifically, taking the minimization of the number of known features with a selection probability of 1 in the optimized feature subset (i.e., the selected known features in the optimized feature subset) as the goal, the first objective function is constructed. The participation sub-data of the participant corresponding to the local feature subset as the target global feature set is used as the target sub-data. According to the target sub-data, the preset classifier is used to evaluate the optimized feature subset to determine the third evaluation result of the optimized feature subset. This process is the same as the process in which the participant uses the preset classifier to evaluate the particle according to its own participation sub-data.

[0056] Specifically, for example, the target sub-data is divided into a target training set and a target validation set. The data corresponding to the optimized feature subset and the labels corresponding to this part of the data are found in the target training set to form a third training sample, and the third training sample is used to train the preset classifier so that the preset classifier learns the features selected by the optimized feature subset. Then, the preset classifier trained based on the optimized feature subset is used to classify and predict the labels of the validation samples in the target validation set. The classification accuracy, classification recall rate, and classification error of the preset classifier under the training of the optimized feature subset are calculated on the target validation set, and the negative value of the classification accuracy, the negative value of the classification recall rate, and the classification error of the preset classifier under the training of the optimized feature subset are used as the third evaluation result.

[0057] Next, taking the minimization of the negative value of the classification accuracy of the preset classifier under the training of the optimized feature subset as the objective, a second objective function is constructed; taking the minimization of the negative value of the classification recall rate of the preset classifier under the training of the optimized feature subset as the objective, a third objective function is constructed; taking the minimization of the classification error of the preset classifier under the training of the optimized feature subset as the objective, a fourth objective function is constructed; taking the minimization of the feature selection cost of the optimized feature subset as the objective, a fifth objective function is constructed. Thus, according to the first objective function, the second objective function, the third objective function, the fourth objective function and the fifth objective function, a multi-objective preprocessing optimization model is constructed.

[0058] Exemplarily, the multi-objective preprocessing optimization model F is expressed as: ; wherein, ; ; ; ; .

[0059] Here, is the first objective function, is the number of known features with a selection probability of 1 in the optimized feature subset X . is the second objective function, is the classification accuracy of the preset classifier under the training of the optimized feature subset X . is the third objective function, is the classification recall rate of the preset classifier under the training of the optimized feature subset X . is the fourth objective function, is the classification error of the preset classifier under the training of the optimized feature subset X , is used to calculate the binary cross-entropy loss of the preset classifier under the training of the optimized feature subset X , is used to calculate the multi-class cross-entropy loss of the preset classifier under the training of the optimized feature subset X , Y is the number of validation samples in the target validation set, is the label of the validation sample h in the binary classification, is the probability of correctly predicting the validation sample X in the binary classification under the training of the optimized feature subset by the preset classifier, h , Kis the number of classifications of the preset classifier, is the validation sample for multi-classification h label, is the probability that the preset classifier correctly predicts the validation sample during multi-classification under the training of the optimized feature subset X Specifically, in intelligent healthcare, binary classification is applicable to predicting whether a patient is sick (labeled 0 / 1), and multi-classification is applicable to predicting the classification of diseases, such as diabetes types 1, 2, and 3. h is the fifth objective function, which refers to summing the costs of known features with a selection probability of 1 in the optimized feature subset to obtain the feature selection cost of the optimized feature subset X The cost of the optimized feature subset X selecting known features can be randomly initialized or specifically set according to the actual application scenario, and this embodiment does not make specific restrictions. X selecting known features p The cost can be randomly initialized or specifically set according to the actual application scenario, and this embodiment does not make specific restrictions.

[0060] and in and are calculated by the following two formulas: , , represents the number of negative samples whose labels are predicted as positive classes by the preset classifier under the training of the optimized feature subset X , represents the number of positive samples whose labels are predicted as negative classes by the preset classifier under the training of the optimized feature subset X , represents the number of positive samples whose labels are predicted as positive classes by the preset classifier under the training of the optimized feature subset X .

[0061] In this embodiment, the first objective function is used to reduce the data dimension, lower the computational complexity, improve the processing efficiency, and avoid the curse of dimensionality. The second objective function and the third objective function are equivalent to maximizing the classification accuracy and the classification recall rate, ensuring the prediction accuracy of the preset classifier on the optimized feature subset and the recognition ability for positive samples. For example, in intelligent healthcare, high accuracy reduces misdiagnosis, and high recall rate avoids missed diagnosis, improving the reliability of diagnosis. The fourth objective function minimizes the classification error of the preset classifier under the training of the optimized feature subset, enhances the generalization ability of the multi-objective preprocessing optimization model, makes the target feature subset obtained after preprocessing the target data applicable to unknown data, and improves the stability of Internet of Things applications. The fifth objective function comprehensively considers the costs such as feature acquisition and storage, and avoids excessive investment in resources in pursuit of accuracy. For example, in industrial Internet of Things, low-cost and efficient features are preferentially selected to balance benefits and costs.

[0062] Further, as a refinement and extension of the specific implementation manner of the above embodiment, to fully illustrate the specific implementation process of this embodiment, step 103, that is, solving the multi-objective preprocessing optimization model to obtain the target feature subset of the target data, specifically includes: determining a preset number of optimized feature subsets of the target global feature set according to the target global feature set of the target data; using the optimized feature subsets as individuals to form a parent population; calculating the objective function values of the individuals through the objective functions in the multi-objective preprocessing optimization model, where the objective functions are constructed based on the optimized feature subsets, and the objective functions include the first objective function, the second objective function, the third objective function, the fourth objective function, and the fifth objective function; performing genetic operations on the parent population to generate an offspring population, and merging the parent population and the offspring population to obtain a target population; screening out the next-generation parent population from the target population based on the objective function values of the individuals; using the individuals in the parent population with the evolution generation equal to the preset threshold as the target feature subset of the target data.

[0063] Among them, the optimized feature subset is a candidate solution of the multi-objective preprocessing optimization model, used to represent the selection probability of the selected features in the target global feature set, and the selected features are the known features of the target data in the target global feature set with the selection probability being the preset value.

[0064] In this embodiment, based on the feedforward control and feedback control ideas in the field of automatic control, an evolutionary algorithm is used to solve the multi-objective preprocessing optimization model, reduce the computational complexity, improve the preprocessing efficiency, and obtain high-quality preprocessing results for the same type of high-dimensional data in the Internet of Things.

[0065] Specifically, among the selected features in the target global feature set, features are randomly selected, and the selection probabilities of the randomly selected features are randomly initialized with real numbers between 0 and 1. At the same time, binary decoding is performed to generate a preset number of optimized feature subsets. It is worth mentioning that here, the optimized feature subsets are regarded as individuals, thus forming an initialized parental population with a preset number. It can be understood that the individuals are also candidate solutions for the multi-objective preprocessing optimization model.

[0066] Exemplarily, each individual x can be expressed as: ; where represents the selection probability of individual x for the known feature p . It can be understood that for the known features not selected in the target global feature set, the selection probability of the known feature in individual x is also 0.

[0067] Next, according to the following formula, individual x is decoded into binary form to update individual x : ; where the individual x after binary decoding is expressed as: . is the after binary decoding.

[0068] Then, through the objective functions in the multi-objective preprocessing optimization model, the objective function values of the individuals are calculated to evaluate the advantages and disadvantages of the individuals through the objective function values, providing a basis for subsequent optimization. Among them, the objective functions include the first objective function, the second objective function, the third objective function, the fourth objective function, and the fifth objective function. The objective function values correspond to the objective functions and include the first objective function value, the second objective function value, the third objective function value, the fourth objective function value, and the fifth objective function value. Next, the objective function values of the individuals are normalized to overcome the influence of different objective scales.

[0069] Exemplarily, according to the following formula, the objective function values of individual x are normalized: ; ; ; where m is the number of objective functions, that is, 5. is the x -th o objective function value of individual For an individual x The o normalized objective function value is the parent population P where the individual x of the o minimum value among the objective function values is the parent population P where the individual x of the o maximum value among the objective function values

[0070] Furthermore, perform random matching selection on the parent population, simulate binary crossover and polynomial mutation operations to generate an offspring population, increase population diversity, and explore new solution spaces. Also, merge the parent population and the offspring population to obtain the target population, so as to retain the excellent individuals in the parent population, and at the same time combine the new individuals of the offspring to expand the selection range and provide more individuals.

[0071] Exemplarily, the merge operation is expressed as: , where is the offspring population is the target population

[0072] Thus, based on the objective function values of the individuals, evaluate the distribution and quality of the individuals in the target population, and select the next-generation parent population from the target population until the evolutionary generation of the parent population reaches the maximum evolutionary generation (i.e., the preset threshold), to obtain the final feature selection scheme, that is, take the individuals in the parent population with the evolutionary generation equal to the preset threshold as the target feature subset of the target data, and further obtain a high-quality preprocessing result for the same type of data in the Internet of Things, so that the solution process achieves a balance between the optimal solution and population diversity, ensuring that the solution process converges to the optimal region and does not lose the ability to explore new solution spaces.

[0073] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, based on the objective function values of individuals, the next-generation parent population is selected from the target population, which specifically includes: based on the objective function values of individuals, non-dominated sorting is performed on the target population to divide the individuals in the target population into non-dominated layers of different levels; according to the levels of the non-dominated layers, superposition processing is performed on the non-dominated layers until the number of individuals in the superposed non-dominated layer is greater than or equal to the preset number, obtaining a target set; if the number of individuals in the target set is equal to the preset number, the target set is used as the next-generation parent population, and the individuals in the target set are the individuals in the target set; if the number of individuals in the target set is greater than the preset number, according to the objective function values of the individuals in the target set, the convergence evaluation values of the individuals in the target set and the cosine values between the individuals in the target set and the coordinate axes in the coordinate system are respectively determined, the coordinate system is established according to the objective function, and the coordinate axes correspond to the objective function; according to the convergence evaluation values and cosine values of the individuals in the target set, the next-generation parent population is selected from the target set.

[0074] In this embodiment, clustering environmental selection based on information feedback is performed. By non-dominated sorting, the individuals in the target population are divided into different levels, and the high-level individuals are preferentially selected to form a target set to ensure the quality of the basic population. If the number of individuals in the target set exceeds the preset number, a convergence evaluation value is introduced to measure the degree to which an individual approaches the optimal solution, and a coordinate system is constructed according to the objective function to form a target space, and the cosine value between the individual and the coordinate axis corresponding to the objective function is used to reflect the distribution position of the individual in the target space, providing multi-dimensional indicators for further screening, balancing convergence and distribution, and avoiding the quality decline caused by simple quantity superposition.

[0075] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, the next-generation parent population is screened from the target set according to the convergence evaluation value and cosine value of the individuals in the target set, which specifically includes: clustering the individuals in the target set according to the objective function values of the individuals to obtain a preset number of clusters; screening out the convergent individuals and boundary individuals from the target set according to the convergence evaluation value and cosine value respectively; adding the convergent individuals and boundary individuals to the next-generation parent population, and deleting the clusters to which the convergent individuals and boundary individuals belong from the target set to update the target set; if the number of selected individuals in the next-generation parent population is less than the preset number, determining the fitness evaluation value of the individuals in the target set according to the convergence evaluation value of the individuals in the target set; screening out the adaptable individuals from the target set according to the fitness evaluation value, and adding the adaptable individuals to the next-generation parent population; deleting the clusters to which the adaptable individuals belong from the target set to update the target set, and re-determining the fitness evaluation value of the individuals in the target set according to the convergence evaluation value of the individuals in the target set until the number of selected individuals in the next-generation parent population is equal to the preset number.

[0076] In this embodiment, after clustering the individuals in the target set, the convergent individuals with high convergence evaluation values and closer to the optimal solution, as well as the boundary individuals at the edge positions of the target space, are first screened out, added to the next-generation parent population, and the clusters to which they belong are deleted to avoid repeated selection and maintain diversity. If the number of individuals in the next-generation parent population is still insufficient, further screening is performed through the fitness evaluation value to ensure that the parent population contains both high-quality solutions and has a wide spatial distribution, improving the global search ability of the solution.

[0077] Among them, the selected individuals in the next-generation parent population are the individuals that have entered the next-generation parent population.

[0078] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, the fitness evaluation value of the individuals in the target set is determined according to the convergence evaluation value of the individuals in the target set, which specifically includes: re-representing the individuals in the target set in the coordinate system according to the objective function values of the individuals in the target set; re-representing the selected individuals in the next-generation parent population in the coordinate system according to the objective function values of the selected individuals in the next-generation parent population; determining the local diversity value of the individuals in the target set according to the re-represented individuals in the target set; determining the global diversity value of the individuals in the target set according to the re-represented individuals in the target set and the selected individuals in the next-generation parent population; determining the comprehensive diversity evaluation value of the individuals in the target set according to the local diversity value and global diversity value of the individuals in the target set; determining the fitness evaluation value of the individuals in the target set according to the convergence evaluation value and comprehensive diversity evaluation value of the individuals in the target set.

[0079] In this embodiment, according to the objective function values of individuals, the individuals in the target set and the selected individuals in the next-generation parent population are placed in the coordinate system established according to the objective function. The difference between an individual in the target set and its neighboring individuals is represented by the local diversity value, and the difference between an individual in the target set and the selected individuals in the next-generation parent population is represented by the global diversity value, thereby forming a comprehensive diversity evaluation value. Combining with the convergence evaluation value, the fitness evaluation value of the individuals in the target set is obtained, so that the fitness evaluation value can comprehensively reflect the "quality" and "uniqueness" of the individuals in the target set, ensuring that the selected fit individuals can not only drive the solution process to converge to the optimal solution, but also avoid population homogenization, and enhance the robustness and optimization efficiency of the solution process in high-dimensional complex data.

[0080] Specifically, non-dominated sorting is performed on the target population to divide the individuals in the target population into non-dominated levels at different levels. The individuals at higher levels are superior in the objective function, which helps to select better individuals to enter the next generation subsequently.

[0081] Then, starting from the largest to the smallest level of the non-dominated levels, the number of individuals in each non-dominated level is superimposed until the sum exceeds or equals the population size (i.e., the preset number), a critical level is determined, and the individuals in these non-dominated levels are placed in the target set, so as to retain as many excellent individuals as possible while ensuring that the number of finally selected individuals can meet the population size.

[0082] Exemplarily, the critical level L is expressed as: , where is the l th non-dominated level, N is the preset number, is the number of individuals in the l th non-dominated level.

[0083] Then, compare the number of individuals in the target set with the population size: If the number of individuals in the target set is exactly equal to the population size, directly use the target set as the next-generation parent population to ensure the quality of the individuals in the parent population; If the number of individuals in the target set is greater than the population size, then the target set needs to be further screened.

[0084] Specifically, first, through the formula in the above steps Normalize the objective function values of individuals in the target set to update the objective function values of individuals in the target set, thereby eliminating the dimensional differences between different objectives, making subsequent operations fairer and not being affected by the overly large ranges of certain objectives. Then, cluster the individuals in the target set according to their objective function values to obtain a preset number of clusters, in order to maintain the diversity of the population and avoid the solution set being too concentrated in a certain area. Through clustering, it can be ensured that a certain number of individuals are selected in each cluster, thus maintaining the distribution of solutions.

[0085] Furthermore, according to the objective function values of individuals, calculate the convergence evaluation values of individuals in the target set. The larger the convergence evaluation value of an individual, the better the performance of the individual on each objective function, so the convergence is better, which means that the individual is closer to the optimal solution.

[0086] Exemplarily, calculate the individuals in the target set according to the following formula x of the convergence evaluation value : .

[0087] Thus, select the individual with the best convergence in the target set as the convergent individual, and directly select the convergent individual into the next-generation parental population to retain the optimal solution, ensure the convergence of the solution process, and embody the feedforward control idea in the field of control.

[0088] Next, establish a coordinate system according to the objective function. Each coordinate axis in the coordinate system corresponds to an objective function, thus forming an objective space. In the objective space, represent the individuals in the target set again according to the objective function values of the individuals in the target set, that is, connect the individuals in the target set with the origin in the coordinate system to obtain a first vector. Each dimension in the first vector corresponds to an objective function value of an individual in the target set, thus placing the individuals in the target set into the objective space.

[0089] Then, find the individuals closest to each coordinate axis as the boundary individuals on each coordinate axis. Specifically, for any coordinate axis, use the cosine value between the first vector of the individuals in the target set and the coordinate axis to measure the closeness of the individuals in the target set to the coordinate axis. The individual corresponding to the largest cosine value is the boundary individual on that coordinate axis. The boundary individuals help maintain the expansibility of the population, ensure that the solution process can explore all corners of the objective space, and avoid the contraction of the front. Then, directly select the screened boundary individuals in the target set into the next-generation parental population.

[0090] Exemplarily, screen out the boundary individuals in the target set according to the following formula: ; where represents the closest to theo The boundary individuals of the coordinate axes represent the individuals in the target set x and the cosine value of the angle between the first vector of the individual o and the th coordinate axis. o represents the unit vector of the

[0091] It is worth mentioning that the convergence individual with the best convergence in the target set and the boundary individuals that have a greater impact on the solution process are directly selected into the next-generation parent population to actively guide the evolution direction of the population, rather than relying solely on the natural selection process, which reflects the feedforward control idea in the control field.

[0092] Furthermore, the convergence individuals and boundary individuals are removed from the target set, and the clusters where the convergence individuals and boundary individuals are located are deleted from the target set to update the target set, avoid repeated selection, and enhance the population diversity.

[0093] If the number of selected individuals in the next-generation parent population is still less than the population size at this time, further calculate the local diversity and global diversity of the individuals in the target set.

[0094] Exemplarily, calculate the local diversity S of the individuals in the target set x and the global diversity according to the following formula: : ; ; where represents the neighboring individual closest to the individual S in the target set x , is the selected individual in the next-generation parent population P . is the cosine value of the angle between the first vector of the individual S in the target set x and the second vector of the selected individual P in the next-generation parent population . is the cosine value of the angle between the first vector of the individual S in the target set x and the first vector of its neighboring individual .

[0095] Here, in the same way as the process of determining the boundary individuals, in the target space, according to the objective function values of the selected individuals in the next-generation parent population, the selected individuals in the next-generation parent population are re-represented, that is, by connecting the selected individuals in the next-generation parent population with the origin of the coordinate system to obtain a second vector. Each dimension in the second vector corresponds to an objective function value of a selected individual in the next-generation parent population, so as to place the selected individuals in the next-generation parent population into the target space. Moreover, the angle between the first vectors is calculated. The smaller the angle, the closer the individuals corresponding to these two first vectors in the target set are, and thus the neighboring individuals of any individual in the target set can be determined.

[0096] In this embodiment, the local diversity value is used to measure the distribution sparsity of individuals in the local area of the target set, avoiding overcrowding of individuals within a cluster and preventing redundancy of solutions in the local area. If the local diversity value of an individual in the target set is large, it indicates that the individual is unique in its nearby area and should be preferentially retained to maintain local diversity. The global diversity value is used to measure the distribution breadth of individuals in the entire target space, ensuring that the individuals in the target set are significantly different from the selected individuals in the next-generation parent population globally and covering different regions of the target space. If the global diversity value of an individual in the target set is large, it indicates that the individual makes a high contribution to global diversity and helps to explore the uncovered frontier regions.

[0097] Furthermore, according to the local diversity value and the global diversity value of the individuals in the target set, the comprehensive diversity evaluation value of the individuals in the target set is determined.

[0098] Exemplarily, the target set is determined according to the following formula S for the individuals x in the : ; where is a weight factor, and is the number of selected individuals in the next-generation parent population P .

[0099] Here, when an individual in the target set has a large comprehensive diversity evaluation value, it indicates that the individual has good diversity. Using the selected individuals in the next-generation parent population as a reference to evaluate the diversity of individuals in the target set fully reflects the feedback control idea in the control field.

[0100] Next, according to the convergence evaluation value and the comprehensive diversity evaluation value of the individuals in the target set, the fitness evaluation value of the individuals in the target set is determined.

[0101] Exemplarily, the target set is determined according to the following formula SIndividual in the middle x Fitness evaluation value : 。

[0102] Here, when an individual in the target set has a large fitness evaluation value, it indicates that the individual has good convergence and diversity.

[0103] Take the individual with the largest fitness evaluation value in the target set as the adaptable individual, add the adaptable individual to the next-generation parent population, and delete the cluster to which the adaptable individual belongs from the target set to improve the population diversity, so as to update the target set.

[0104] Then, re-determine the fitness evaluation value of the individuals in the target set to loop through the above steps until the number of selected individuals in the next-generation parent population reaches the population size, that is 。

[0105] It should be noted that the magnitudes of the serial numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0106] Furthermore, as Figure 2 shown, as a specific implementation of the above multi-objective Internet of Things data preprocessing method with privacy protection function, the embodiments of the present application provide a multi-objective Internet of Things data preprocessing device 200 with privacy protection function. The multi-objective Internet of Things data preprocessing device 200 with privacy protection function includes: a federation module 201, a construction module 202, and a solution module 203.

[0107] Among them, the federation module 201 is used to extract the local feature subset of the participating sub-data of the participants in the Internet of Things based on the federated learning mode, so that the trusted participants in the Internet of Things can determine the target global feature set of the target data in the Internet of Things according to the local feature subset of the participants. The target data is the data in the Internet of Things belonging to the target type, the participating sub-data is the data owned by the participants in the target data, and the trusted participants are independent trusted devices. The trusted participants can communicate with the participants; The construction module 202 is used to construct a multi-objective preprocessing optimization model of the target data according to the target global feature set of the target data; The solution module 203 is used to solve the multi-objective preprocessing optimization model to obtain the target feature subset of the target data.

[0108] In one embodiment, the federated module 201 is specifically configured for the participant to initialize the positions of the particles according to the initial feature subset of the participant's participation sub-data and the initial global feature set of the target data to form a particle population, where the particle is a candidate feature subset of the participant's participation sub-data, the position of the particle is used to represent the selection probability of the known features of the target data by the participant, and the initial feature subset of the participant and the initial global feature set of the target data are determined according to the known features; the participant evaluates the particles according to the participant's participation sub-data using a preset classifier to determine the first evaluation result of the particles; the participant determines the optimal position of the particles and the population optimal position of the particle population according to the first evaluation result of the particles; the participant updates the positions of the particles according to the optimal position of the particles and the population optimal position to update the particle population until the particle population meets the preset termination condition; the participant determines the population optimal position of the particle population that meets the preset termination condition as the local feature subset of the participant's participation sub-data.

[0109] In one embodiment, the federated module 201 is specifically configured for comparing the initial feature subset of the participant with the initial global feature set of the target data; if the comparison result is consistent, initializing the positions of the particles according to the initial global feature set of the target data; if the comparison result is inconsistent, initializing the positions of the particles according to the initial feature subset of the participant and the initial global feature set of the target data.

[0110] In one embodiment, the federated module 201 is specifically configured for the participant to send the local feature subset of the participant to the confident participant; the confident participant sends the received local feature subsets of all participants to the participant; the participant evaluates the received local feature subsets of the participants according to the participant's participation sub-data using a preset classifier to determine the second evaluation result of the local feature subset of the participant; the participant sends the second evaluation result of the local feature subset of the participant to the confident participant; the confident participant determines the average evaluation result of the local feature subset of the participant according to the received second evaluation result of the local feature subset of the participant; the confident participant takes the local feature subset with the largest average evaluation result as the target global feature set.

[0111] In one embodiment, the construction module 202 is specifically configured to determine an optimized feature subset of the target global feature set according to the target global feature set of the target data, where the optimized feature subset is a candidate solution of the multi-objective preprocessing optimization model and is used to represent the selection probability of the selected features in the target global feature set, and the selected features are the known features of the target data with a selection probability of a preset value in the target global feature set; taking the minimization of the number of features in the optimized feature subset as the objective, construct the first objective function; according to the target sub-data, use a preset classifier to evaluate the optimized feature subset to determine the third evaluation result of the optimized feature subset, and the target sub-data is determined according to the participation sub-data of the participant corresponding to the local feature subset that is the target global feature set; taking the minimization of the third evaluation result as the objective, construct the second objective function, the third objective function, and the fourth objective function respectively; taking the minimization of the feature selection cost of the optimized feature subset as the objective, construct the fifth objective function; according to the first objective function, the second objective function, the third objective function, the fourth objective function, and the fifth objective function, construct a multi-objective preprocessing optimization model.

[0112] In one embodiment, the construction module 202 is specifically configured to train a preset classifier according to the data corresponding to the optimized feature subset in the target sub-data; use the trained preset classifier to classify the target sub-data to determine the classification accuracy, classification recall rate, and classification error of the preset classifier under the training of the optimized feature subset, and use the negative value of the classification accuracy, the negative value of the classification recall rate, and the classification error of the preset classifier under the training of the optimized feature subset as the third evaluation result; among them, the negative value of the classification accuracy is used to construct the second objective function, the negative value of the classification recall rate is used to construct the third objective function, and the classification error is used to construct the fourth objective function.

[0113] In one embodiment, the solution module 203 is specifically configured to determine a preset number of optimized feature subsets of the target global feature set according to the target global feature set of the target data, where the optimized feature subset is a candidate solution of the multi-objective preprocessing optimization model and is used to represent the selection probability of the selected features in the target global feature set, and the selected features are the known features of the target data with a selection probability of a preset value in the target global feature set; use the optimized feature subset as an individual so that the individuals form a parent population; calculate the objective function value of the individual through the objective functions in the multi-objective preprocessing optimization model, and the objective functions are constructed according to the optimized feature subset, and the objective functions include the first objective function, the second objective function, the third objective function, the fourth objective function, and the fifth objective function; perform genetic operations on the parent population to generate an offspring population, and merge the parent population and the offspring population to obtain a target population; based on the objective function values of the individuals, select the next-generation parent population from the target population; use the individuals in the parent population with the evolution generation equal to the preset threshold as the target feature subset of the target data.

[0114] In one embodiment, the solving module 203 is specifically configured to perform non-dominated sorting on the target population based on the objective function values of the individuals, so as to divide the individuals in the target population into non-dominated layers at different levels; perform superposition processing on the non-dominated layers according to the levels of the non-dominated layers until the number of individuals in the superposed non-dominated layers is greater than or equal to a preset number, to obtain a target set; if the number of individuals in the target set is equal to the preset number, use the target set as the next-generation parent population; if the number of individuals in the target set is greater than the preset number, respectively determine the convergence evaluation values of the individuals in the target set and the cosine values between the individuals in the target set and the coordinate axes in the coordinate system, where the coordinate system is established according to the objective function and the coordinate axes correspond to the objective function; select the next-generation parent population from the target set according to the convergence evaluation values and cosine values of the individuals in the target set.

[0115] In one embodiment, the solving module 203 is specifically configured to perform clustering processing on the individuals in the target set according to the objective function values of the individuals, to obtain a preset number of clusters; respectively select convergent individuals and boundary individuals from the target set according to the convergence evaluation values and cosine values; add the convergent individuals and boundary individuals to the next-generation parent population, and delete the clusters to which the convergent individuals and boundary individuals belong from the target set to update the target set; if the number of selected individuals in the next-generation parent population is less than the preset number, determine the fitness evaluation values of the individuals in the target set according to the convergence evaluation values of the individuals in the target set; select adaptable individuals from the target set according to the fitness evaluation values and add the adaptable individuals to the next-generation parent population; delete the clusters to which the adaptable individuals belong from the target set to update the target set, and re-determine the fitness evaluation values of the individuals in the target set according to the convergence evaluation values of the individuals in the target set until the number of selected individuals in the next-generation parent population is equal to the preset number.

[0116] In one embodiment, the solving module 203 is specifically configured to re-represent the individuals in the target set in the coordinate system according to the objective function values of the individuals in the target set; re-represent the selected individuals in the next-generation parent population in the coordinate system according to the objective function values of the selected individuals in the next-generation parent population; determine the local diversity values of the individuals in the target set according to the re-represented individuals in the target set; determine the global diversity values of the individuals in the target set according to the re-represented individuals in the target set and the re-represented selected individuals in the next-generation parent population; determine the comprehensive diversity evaluation values of the individuals in the target set according to the local diversity values and global diversity values of the individuals in the target set; determine the fitness evaluation values of the individuals in the target set according to the convergence evaluation values and comprehensive diversity evaluation values of the individuals in the target set.

[0117] For the specific limitations of the high-dimensional multi-objective Internet of Things data preprocessing device with privacy protection function, reference can be made to the limitations of the high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function in the above text, which will not be elaborated here. Each module in the above high-dimensional multi-objective Internet of Things data preprocessing device with privacy protection function can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0118] Based on the above as Figure 1 shown in the method, correspondingly, an embodiment of the present application further provides a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above as Figure 1 shown in the high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function.

[0119] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0120] Based on the above as Figure 1 shown in the method, and Figure 2 shown in the virtual device embodiment, in order to achieve the above purpose, an embodiment of the present application further provides a computer device, which can specifically be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the above as Figure 1 shown in the high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function.

[0121] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.

[0122] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not limit the computer device, and it may include more or fewer components, or combine certain components, or have different component arrangements.

[0123] The storage medium may also include an operating system and a network communication module. The operating system is a program for managing and storing the hardware and software resources of a computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components within the storage medium, as well as communication with other hardware and software in the entity device.

[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or the embodiments of this application can be implemented through hardware.

[0125] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment scenario, and the modules or processes in the drawings are not necessarily essential for implementing this application. Those skilled in the art can understand that the modules in the device in the embodiment scenario can be distributed in the device in the embodiment scenario according to the description of the embodiment scenario, or can be correspondingly changed and located in one or more devices different from this embodiment scenario. The modules in the above embodiment scenario can be combined into one module, or further split into multiple sub-modules.

[0126] The above serial numbers of this application are only for description and do not represent the advantages or disadvantages of the embodiment scenario. The above-disclosed are only several specific embodiment scenarios of this application. However, this application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of this application.

Claims

1. A high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function, characterized in that, The method includes: Based on the federated learning mode, extract a local feature subset of the participating sub-data of the participants in the Internet of Things, so that the trusted participants in the Internet of Things determine the target global feature set of the target data in the Internet of Things according to the local feature subset of the participants. The target data is the data belonging to the target type in the Internet of Things, the participating sub-data is the data owned by the participants in the target data, the trusted participants are independent credit-granting devices, and the trusted participants can communicate with the participants; Construct a multi-objective preprocessing optimization model for the target data according to the target global feature set of the target data; Solve the multi-objective preprocessing optimization model to obtain the target feature subset of the target data.

2. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 1, wherein The extraction of the local feature subset of the participating sub-data of the participants in the Internet of Things specifically includes: The participant initializes the position of the particle according to the initial feature subset of the participating sub-data of the participant and the initial global feature set of the target data to form a particle population. Wherein, the particle is a candidate feature subset of the participating sub-data of the participant, and the position of the particle is used to represent the selection probability of the participant for the known features of the target data. The initial feature subset of the participant and the initial global feature set of the target data are determined according to the known features; The participant evaluates the particle according to the participating sub-data of the participant by using a preset classifier to determine the first evaluation result of the particle; The participant determines the optimal position of the particle and the population optimal position of the particle population according to the first evaluation result of the particle; The participant updates the position of the particle according to the optimal position of the particle and the population optimal position to update the particle population until the particle population meets the preset termination condition; The participant determines the population optimal position of the particle population that meets the preset termination condition as the local feature subset of the participating sub-data of the participant.

3. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 2, characterized in that, The initialization of the position of the particle according to the initial feature subset of the participating sub-data of the participant and the initial global feature set of the target data specifically includes: Compare the initial feature subset of the participant with the initial global feature set of the target data; If the comparison result is consistent, initialize the position of the particle according to the initial global feature set of the target data; If the comparison result is inconsistent, initialize the position of the particle according to the initial feature subset of the participant and the initial global feature set of the target data.

4. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 1, characterized in that, The trusted participant determines the target global feature set of the target data in the Internet of Things according to the local feature subset of the participant, specifically including: The participant sends the local feature subset of the participant to the trusted participant; The trusted participant sends all the received local feature subsets of the participants to the participant; The participant evaluates the received local feature subset of the participant by using a preset classifier according to the participating sub-data of the participant to determine the second evaluation result of the local feature subset of the participant; The participant sends the second evaluation result of the local feature subset of the participant to the confidence participant; The confidence participant determines the average evaluation result of the local feature subset of the participant according to the received second evaluation result of the local feature subset of the participant; The confidence participant uses the local feature subset with the maximum average evaluation result as the target global feature set.

5. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 1, characterized in that, Constructing the multi-objective preprocessing optimization model for the target data according to the target global feature set of the target data specifically includes: Determining an optimized feature subset of the target global feature set according to the target global feature set of the target data, where the optimized feature subset is a candidate solution of the multi-objective preprocessing optimization model and is used to represent the selection probability of the selected features in the target global feature set, and the selected features are the known features of the target data with a selection probability of a preset value in the target global feature set; Constructing a first objective function with the goal of minimizing the number of features in the optimized feature subset; Evaluating the optimized feature subset according to the target sub-data by using a preset classifier, where the target sub-data is determined according to the participating sub-data corresponding to the target global feature set, and determining the third evaluation result of the optimized feature subset; Constructing a second objective function, a third objective function, and a fourth objective function respectively with the goal of minimizing the third evaluation result; Constructing a fifth objective function with the goal of minimizing the feature selection cost of the optimized feature subset; Constructing the multi-objective preprocessing optimization model according to the first objective function, the second objective function, the third objective function, the fourth objective function, and the fifth objective function.

6. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 5, characterized in that Evaluating the optimized feature subset according to the target sub-data by using a preset classifier specifically includes: Training the preset classifier according to the data corresponding to the optimized feature subset in the target sub-data; Classifying the target sub-data by using the trained preset classifier, determining the classification accuracy, classification recall rate, and classification error of the preset classifier under the training of the optimized feature subset, and using the negative value of the classification accuracy of the preset classifier under the training of the optimized feature subset, the negative value of the classification recall rate, and the classification error as the third evaluation result; Among them, the negative value of the classification accuracy is used to construct the second objective function, the negative value of the classification recall rate is used to construct the third objective function, and the classification error is used to construct the fourth objective function.

7. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 1, characterized in that, Solving the multi-objective preprocessing optimization model to obtain the target feature subset of the target data specifically includes: Determining a preset number of optimized feature subsets of the target global feature set according to the target global feature set of the target data, where the optimized feature subset is a candidate solution of the multi-objective preprocessing optimization model and is used to represent the selection probability of the selected features in the target global feature set, and the selected features are the known features of the target data with a selection probability of a preset value in the target global feature set; Use the optimized feature subset as an individual to form a parental population with these individuals; Calculate the objective function values of the individuals by means of the objective functions in the multi-objective preprocessing optimization model. The objective functions are constructed based on the optimized feature subset and include a first objective function, a second objective function, a third objective function, a fourth objective function, and a fifth objective function; Perform genetic operations on the parental population to generate an offspring population, and merge the parental population and the offspring population to obtain a target population; Based on the objective function values of the individuals, select the next-generation parental population from the target population; Take the individuals in the parental population whose evolution generation is equal to the preset threshold as the target feature subset of the target data.

8. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 7, characterized in that, The step of selecting the next-generation parental population from the target population based on the objective function values of the individuals specifically includes: Based on the objective function values of the individuals, perform non-dominated sorting on the target population to divide the individuals in the target population into non-dominated levels at different levels; Perform a superposition process on the non-dominated levels according to the levels of the non-dominated levels until the number of individuals in the superposed non-dominated level is greater than or equal to the preset number to obtain a target set; If the number of individuals in the target set is equal to the preset number, use the target set as the next-generation parental population; If the number of individuals in the target set is greater than the preset number, respectively determine the convergence evaluation values of the individuals in the target set and the cosine values between the individuals in the target set and the coordinate axes in the coordinate system according to the objective function values of the individuals in the target set. The coordinate system is established based on the objective functions, and the coordinate axes correspond to the objective functions; Select the next-generation parental population from the target set according to the convergence evaluation values and cosine values of the individuals in the target set.

9. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 8, wherein, The step of selecting the next-generation parental population from the target set according to the convergence evaluation values and cosine values of the individuals in the target set specifically includes: Perform clustering processing on the individuals in the target set according to the objective function values of the individuals to obtain the preset number of clusters; Respectively select the convergent individuals and the boundary individuals from the target set according to the convergence evaluation values and the cosine values; Add the convergent individuals and the boundary individuals to the next-generation parental population, and delete the clusters to which the convergent individuals and the boundary individuals belong from the target set to update the target set; If the number of selected individuals in the next-generation parental population is less than the preset number, determine the fitness evaluation values of the individuals in the target set according to the convergence evaluation values of the individuals in the target set; Select the fit individuals from the target set according to the fitness evaluation values and add the fit individuals to the next-generation parental population; Delete the cluster to which the adapted individual belongs from the target set to update the target set, and re-determine the fitness evaluation value of the individuals in the target set according to the convergence evaluation value of the individuals in the target set until the number of selected individuals in the next-generation parent population is equal to the preset number.

10. The high-dimensional multi-objective Internet of Things data preprocessing method with privacy protection function according to claim 9, characterized in that, The determination of the fitness evaluation value of the individuals in the target set according to the convergence evaluation value of the individuals in the target set specifically includes: Redraw the individuals in the target set in the coordinate system according to the objective function values of the individuals in the target set; Redraw the selected individuals in the next-generation parent population in the coordinate system according to the objective function values of the selected individuals in the next-generation parent population; Determine the local diversity value of the individuals in the target set according to the individuals in the target set after redrawing; Determine the global diversity value of the individuals in the target set according to the individuals in the target set after redrawing and the selected individuals in the next-generation parent population after redrawing; Determine the comprehensive diversity evaluation value of the individuals in the target set according to the local diversity value and the global diversity value of the individuals in the target set; Determine the fitness evaluation value of the individuals in the target set according to the convergence evaluation value and the comprehensive diversity evaluation value of the individuals in the target set.