A data feature extraction method based on artificial intelligence
Through an artificial intelligence-based data feature extraction method, the problem of not considering the importance of data in the prior art is solved, and data features are efficiently extracted while protecting private data, reducing processing time and computing costs.
Patent Information
- Application Number
- CN202411975543.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The prior art does not consider the importance of data in the data feature extraction process, resulting in useless data information being analyzed, increasing time consumption and calculation costs. At the same time, in order to protect private data, all data needs to be encrypted, further increasing complexity and cost.
A data feature extraction method based on artificial intelligence is proposed. By acquiring heterogeneous data, feature extraction and model aggregation, important features are determined, sensitive information is identified, and data preprocessing and feature extraction are carried out according to importance and sensitivity, so as to achieve efficient extraction of target data.
While protecting private data, reasonable data feature extraction is carried out according to the importance of the data, which significantly reduces the processing time of data feature extraction and reduces the complexity and calculation cost of data processing.
Smart Images

Figure CN119397222B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of feature extraction technology, and in particular to a data feature extraction method based on artificial intelligence. Background Art
[0002] In modern data processing systems, data feature extraction is a crucial part of artificial intelligence (AI) technology. It involves extracting representative features from a large amount of raw data to facilitate subsequent model training and analysis.
[0003] However, with the continuous development of data processing technology, the issue of privacy data protection has become increasingly important. Privacy data refers to information involving personal privacy, such as name, ID number, address, bank card number, etc. If these data are leaked without authorization, it may pose a serious threat to personal privacy.
[0004] In the existing technology, the importance of data is not taken into consideration during the feature extraction process, which results in some useless data information being analyzed together with the data. Large amounts of data analysis result in data features taking a lot of time, which reduces efficiency. In addition, in order to protect privacy data, all data need to be encrypted, which also increases the complexity and computational cost of data processing. Summary of the invention
[0005] In view of the above problems, the present application is proposed to provide an artificial intelligence-based data feature extraction method and system thereof to overcome the above problems or at least partially solve the above problems, including:
[0006] A data feature extraction method based on artificial intelligence, the method comprising:
[0007] Get heterogeneous data from several clients;
[0008] Extracting features based on the heterogeneous data to generate a standardized feature set;
[0009] Performing model aggregation processing based on the standardized feature set to determine important features;
[0010] Determining target data in the heterogeneous data according to the important features;
[0011] Performing data preprocessing according to the target data to generate standardized data;
[0012] Determining sensitive information in the standardized data based on the standardized data and a preset deep learning model;
[0013] A target feature is determined based on the sensitive information and the standardized data.
[0014] Furthermore, the step of extracting features from the heterogeneous data to generate a standardized feature set includes:
[0015] Performing principal component analysis on the heterogeneous data to generate dimension-reduced data;
[0016] Determining a feature set based on the dimension reduction data;
[0017] Determining initial features according to the feature set and a preset algorithm, wherein the preset algorithm is a LASSO algorithm;
[0018] The initial features are standardized to generate a standardized feature set.
[0019] Furthermore, the step of performing standardization processing on the initial features to generate a standardized feature set includes:
[0020] Determining a mean and a variance corresponding to each of the initial features;
[0021] Determining a sample value according to the mean and the variance;
[0022] The standardized feature set is generated according to the mean, the variance and the sample value.
[0023] Furthermore, the step of performing model aggregation processing according to the standardized feature set to determine important features includes:
[0024] Determining model parameters corresponding to a plurality of said clients according to said standardized feature set;
[0025] Performing weighted averaging processing on the model parameters to obtain initial parameters;
[0026] Performing regularization processing according to the target parameter to obtain the target parameter;
[0027] generating an important feature model according to the target parameter and the standardized feature set;
[0028] Important features are determined based on the important feature model.
[0029] Furthermore, the step of performing data preprocessing based on the target data to generate standardized data includes:
[0030] Convert the target data into format data according to a preset format;
[0031] Determining local data belonging to the participant within the formatted data;
[0032] Performing encryption processing on the local data to generate encrypted data;
[0033] Performing federated learning processing on the encrypted data to generate a standardized model;
[0034] The standardized data is generated according to the standardized model.
[0035] Furthermore, the step of determining the sensitive information in the standardized data based on the standardized data and a preset deep learning model includes:
[0036] Performing feature extraction on the standardized data by using a convolutional neural network and a recurrent neural network to generate text features;
[0037] Performing named entity recognition on the text features to determine key information;
[0038] The key information is semantically analyzed to determine the sensitive information.
[0039] Furthermore, the step of determining target features based on the sensitive information and the standardized data includes:
[0040] Extracting features from the standardized data based on the sensitive information to generate an initial feature extraction model;
[0041] Determining a feature type corresponding to each extracted feature according to the initial feature extraction model and the sensitive information, wherein the feature type includes filtering and normal;
[0042] Generate a target feature extraction model for the normal extraction feature, the initial feature extraction model and a preset evaluation index according to the feature type, wherein the preset evaluation index includes one or more of accuracy, recall rate and F1 value;
[0043] Generate target features according to the target feature extraction model.
[0044] An embodiment of the present application further discloses a data feature extraction system based on artificial intelligence, the system comprising:
[0045] The acquisition module is used to obtain heterogeneous data from several clients;
[0046] A first generating module, used for performing feature extraction based on the heterogeneous data to generate a standardized feature set;
[0047] A first determination module, configured to perform model aggregation processing according to the standardized feature set to determine important features;
[0048] A second determination module, configured to determine target data in the heterogeneous data according to the important features;
[0049] A third determination module is used to perform data preprocessing according to the target data to generate standardized data;
[0050] A fourth determination module, configured to determine sensitive information in the standardized data based on the standardized data and a preset deep learning model;
[0051] The fifth determination module is used to determine the target feature based on the sensitive information and the standardized data.
[0052] An embodiment of the present application also discloses a computer device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the data feature extraction method based on artificial intelligence are implemented as described above.
[0053] An embodiment of the present application further discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data feature extraction method based on artificial intelligence are implemented as described above.
[0054] This application has the following advantages:
[0055] In the embodiments of the present application, the prior art does not consider the importance of data during feature extraction, resulting in some useless data information being analyzed during data feature extraction, and a large amount of data analysis results in a large amount of time required for data features, thereby reducing efficiency; and because data encryption technology is required for all data to protect privacy data, the complexity and computational cost of data processing are also increased. The present application provides a solution that can perform data feature extraction according to the importance of data while protecting privacy data, specifically: obtaining heterogeneous data from several clients; performing feature extraction based on the heterogeneous data to generate a standardized feature set; performing model aggregation processing based on the standardized feature set to determine important features; determining target data in the heterogeneous data based on the important features; performing data preprocessing based on the target data to generate standardized data; determining sensitive information in the standardized data based on the standardized data and a preset deep learning model; and determining target features based on the sensitive information and the standardized data. By "determining the target data in the heterogeneous data based on the important features; performing data preprocessing to generate standardized data based on the target data; determining the sensitive information in the standardized data based on the standardized data and a preset deep learning model; and determining the target features based on the sensitive information and the standardized data", the problem of "not considering the importance of the data during feature extraction, resulting in some useless data information being analyzed when extracting features from the data, and a large amount of data analysis results in a lot of time spent on data features, thereby reducing efficiency; and because data encryption technology needs to be performed on all data in order to protect privacy data, the complexity and computational cost of data processing are also increased" is achieved. The effect of "being able to perform reasonable data feature extraction according to the importance of the data while protecting privacy data, thereby greatly reducing the processing time of data feature extraction, reducing the complexity and computational cost of data processing" is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0057] Figure 1 This is a flowchart of a method for extracting data features based on artificial intelligence provided by an embodiment of the present application;
[0058] Figure 2It is a structural block diagram of a data feature extraction system based on artificial intelligence provided by an embodiment of the present application;
[0059] Figure 3 It is a structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0060] In order to make the objects, features and advantages of the present application more obvious and understandable, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0061] Reference Figure 1 , showing a flowchart of a method for extracting data features based on artificial intelligence provided by an embodiment of the present application;
[0062] A data feature extraction method based on artificial intelligence, the method comprising:
[0063] S110, obtaining heterogeneous data of several clients;
[0064] S120, extracting features based on the heterogeneous data to generate a standardized feature set;
[0065] S130, performing model aggregation processing according to the standardized feature set to determine important features;
[0066] S140, determining target data in the heterogeneous data according to the important features;
[0067] S150, performing data preprocessing according to the target data to generate standardized data;
[0068] S160, determining sensitive information in the standardized data according to the standardized data and a preset deep learning model;
[0069] S170. Determine target features based on the sensitive information and the standardized data.
[0070] In the embodiments of the present application, the prior art does not consider the importance of data during feature extraction, resulting in some useless data information being analyzed during data feature extraction, and a large amount of data analysis results in a large amount of time required for data features, thereby reducing efficiency; and because data encryption technology is required for all data to protect privacy data, the complexity and computational cost of data processing are also increased. The present application provides a solution that can perform data feature extraction according to the importance of data while protecting privacy data, specifically: obtaining heterogeneous data from several clients; performing feature extraction based on the heterogeneous data to generate a standardized feature set; performing model aggregation processing based on the standardized feature set to determine important features; determining target data in the heterogeneous data based on the important features; performing data preprocessing based on the target data to generate standardized data; determining sensitive information in the standardized data based on the standardized data and a preset deep learning model; and determining target features based on the sensitive information and the standardized data. By "determining the target data in the heterogeneous data based on the important features; performing data preprocessing to generate standardized data based on the target data; determining the sensitive information in the standardized data based on the standardized data and a preset deep learning model; and determining the target features based on the sensitive information and the standardized data", the problem of "not considering the importance of the data during feature extraction, resulting in some useless data information being analyzed when extracting features from the data, and a large amount of data analysis results in a lot of time spent on data features, thereby reducing efficiency; and because data encryption technology needs to be performed on all data in order to protect privacy data, the complexity and computational cost of data processing are also increased" is achieved. The effect of "being able to perform reasonable data feature extraction according to the importance of the data while protecting privacy data, thereby greatly reducing the processing time of data feature extraction, reducing the complexity and computational cost of data processing" is achieved.
[0071] Next, a data feature extraction method based on artificial intelligence in this exemplary embodiment will be further described.
[0072] As described in step S110, heterogeneous data of several clients are obtained.
[0073] It should be noted that heterogeneous data is collected from multiple clients, and these data may come from different fields or application scenarios, and therefore have different characteristics and distributions. Since data from different clients may have different characteristics and distributions, it is difficult to find a unified method to process all heterogeneous data; and due to the heterogeneity of the data, the model parameters of different clients may not converge, resulting in reduced model performance; in a specific embodiment of the present application, heterogeneous data may be data from the financial field, which may include transaction records, credit scores, etc. In another embodiment, it can also be used for data from the medical field, which may include patient medical history, physical examination results, etc.
[0074] As described in step S120, feature extraction is performed based on the heterogeneous data to generate a standardized feature set.
[0075] In an embodiment of the present invention, the specific process of "generating a standardized feature set by extracting features based on the heterogeneous data" in step S120 can be further explained in combination with the following description.
[0076] As described in the following steps,
[0077] S210, performing principal component analysis on the heterogeneous data to generate dimension-reduced data;
[0078] S220, determining a feature set according to the dimension reduction data;
[0079] S230, determining initial features according to the feature set and a preset algorithm, wherein the preset algorithm is a LASSO algorithm;
[0080] S240: Standardize the initial features to generate a standardized feature set.
[0081] It should be noted that Principal Component Analysis (PCA) is a commonly used dimensionality reduction technique that is used to project high-dimensional data into a low-dimensional space while retaining the main features of the data as much as possible. The main purpose of PCA is to reduce the dimensionality of the data while minimizing information loss.
[0082] As an example, the heterogeneous data is standardized to make its mean 0 and variance 1. This is to eliminate the influence of different feature dimensions; the covariance matrix reflects the linear relationship between the features of the data. For data with n features, the covariance matrix is an n*n matrix; the covariance matrix is decomposed into eigenvalues to obtain n eigenvalues and corresponding eigenvectors, where the eigenvalue represents the variance of the data in the direction of the eigenvector, and the eigenvector represents the direction of the data projection; the eigenvectors corresponding to the first k eigenvalues are selected in descending order, where k is the dimension to be reduced, and these k eigenvectors constitute a new coordinate system, and the projection of the reduced-dimensional data in the new coordinate system is the reduced-dimensional data; the heterogeneous data is projected onto the selected eigenvector to obtain the reduced-dimensional data, i.e., the reduced-dimensional data; representative features can be extracted based on the reduced-dimensional data, and a feature set is formed based on the representative features. After obtaining the feature set, the LASSO (least absolute shrinkage and selection operator) algorithm is used for feature selection to obtain more meaningful features after further optimization, i.e., the initial features, and the standardized feature set is generated after the initial features are standardized.
[0083] As described in step S240, the initial features are standardized to generate a standardized feature set.
[0084] In an embodiment of the present invention, the specific process of "standardizing the initial features to generate a standardized feature set" in step S240 can be further explained in combination with the following description.
[0085] As described in the following steps,
[0086] S310, determining a mean and a variance corresponding to each of the initial features;
[0087] S320, determining a sample value according to the mean and the variance;
[0088] S330, generating the standardized feature set according to the mean, the variance and the sample value.
[0089] It should be noted that the means of all initial features are adjusted to 0 and the variances are adjusted to 1, thereby generating the standardized feature set; through standardization, data of different features can be compared on the same scale.
[0090] As an example, by calculating the mean (μ) and variance (σ), for each feature \( X_i \), calculate its mean \( \mu_i \) and variance \( \sigma_i \).
[0091] \[ \mu_i = \frac{1}{n} \sum_{j=1}^{n} x_{ij} \] \[ \sigma_i^2 = \frac{1}{n} \sum_{j=1}^{n} (x_{ij} - \mu_i)^2 \]
[0092] where \( x_{ij} \) is the \( j \)th sample value of feature \( X_i \) and \( n \) is the total number of samples;
[0093] Using the calculated mean and variance, each sample value is normalized.
[0094] \[ z_{ij} = \frac{x_{ij} - \mu_i}{\sigma_i} \]
[0095] So that the standardized data \( z_{ij} \) will have mean 0 and variance 1; replace the original data \( x_{ij} \) with the standardized value \( z_{ij} \).
[0096] As described in step S130, model aggregation processing is performed based on the standardized feature set to determine important features.
[0097] In an embodiment of the present invention, the specific process of "performing model aggregation processing according to the standardized feature set to determine important features" in step S130 can be further explained in combination with the following description.
[0098] As described in the following steps,
[0099] S410, determining model parameters corresponding to a plurality of the clients according to the standardized feature set;
[0100] S420, performing weighted average processing on the model parameters to obtain initial parameters;
[0101] S430, performing regularization processing according to the target parameter to obtain the target parameter;
[0102] S440, generating an important feature model according to the target parameter and the standardized feature set;
[0103] S450: Determine important features according to the important feature model.
[0104] It should be noted that model aggregation is performed on the extracted and standardized features, that is, the standardized feature set.
[0105] As an example, the model parameters of different clients are first weighted averaged to obtain initial parameters, and then the weighted averaged initial parameters are regularized to obtain target parameters, so that the target parameters can converge better, and then the important feature model is generated from the target parameters and the standardized feature set, thereby obtaining important features.
[0106] In a specific implementation, the regularization process is L2 regularization, which can prevent overfitting.
[0107] As described in step S150, data preprocessing is performed based on the target data to generate standardized data.
[0108] In an embodiment of the present invention, the specific process of "preprocessing the target data to generate standardized data" in step S150 can be further explained in combination with the following description.
[0109] As described in the following steps,
[0110] S510, converting the target data into format data according to a preset format;
[0111] S520, determining the local data belonging to the participant in the format data;
[0112] S530, performing encryption processing according to the local data to generate encrypted data;
[0113] S540, performing federated learning processing on the encrypted data to generate a standardized model;
[0114] S550: Generate the standardized data according to the standardized model.
[0115] It should be noted that formatted data refers to the unified conversion of different types of data into a data format that can be used for secure multi-party computing; by encrypting the data for transmission, that is, encrypting the data, and using distributed computing to train and optimize the model, knowledge sharing can be achieved while protecting data privacy.
[0116] As an example, the target data is first uniformly converted into a data format that can be used for secure multi-party computing. The encrypted data is generated by encrypting and perturbing the local data belonging to the participants in the format data to protect data privacy. The secure multi-party computing technology is used to encrypt the model parameters and protect the security of the collaborative training process. For example, the Paillier encryption algorithm or the ElGamal encryption algorithm can be used to encrypt the model parameters so that only the participants can decrypt and access the model parameters; a standardized model is generated by performing federated learning processing on the encrypted data, in which a more effective adversarial attack defense mechanism can be introduced in the federated learning process according to actual conditions; finally, standardized data is obtained from the standardized model.
[0117] In a specific implementation, text data is converted into numerical data, image data is converted into pixel value data, etc.; local data is denoised using the Laplace mechanism or the Gaussian mechanism; model parameters are encrypted using the Paillier encryption algorithm or the ElGamal encryption algorithm so that only the participating parties can decrypt and access the model parameters.
[0118] As described in step S160, sensitive information in the standardized data is determined based on the standardized data and a preset deep learning model.
[0119] In one embodiment of the present invention, the specific process of "determining the sensitive information in the standardized data based on the standardized data and the preset deep learning model" in step S160 can be further explained in combination with the following description.
[0120] As described in the following steps,
[0121] S610, extracting features from the standardized data by using a convolutional neural network and a recurrent neural network to generate text features;
[0122] S620, performing named entity recognition processing on the text features to determine key information;
[0123] S630: Perform semantic analysis on the key information to determine the sensitive information.
[0124] It should be noted that sensitive information refers to information that, once leaked, may cause significant losses to individuals, companies or the country, such as personal identity information (such as ID number, bank card number), business secrets (such as customer lists, financial statements), etc.
[0125] As an example, by using machine learning algorithms and natural language processing technology, sensitive information is identified in the target data; specifically, deep learning models, convolutional neural networks (CNN) and recurrent neural networks (RNN) are used to train the data to identify sensitive privacy information in the data.
[0126] In a specific implementation, natural language processing technology uses word embedding technology to convert the text in the target data into a vector representation to facilitate processing by the machine learning model; key information in the extracted text is obtained through named entity recognition (NER) and partial syntactic parsing to improve the accuracy of sensitive information identification; and the extracted key information is comprehensively analyzed and judged through semantic analysis and context understanding to achieve comprehensive identification and processing of various types of sensitive information.
[0127] As described in step S170, target features are determined based on the sensitive information and the standardized data.
[0128] In one embodiment of the present invention, the specific process of "determining the target feature based on the sensitive information and the standardized data" in step S170 can be further explained in combination with the following description.
[0129] As described in the following steps,
[0130] S710, extracting features from the standardized data according to the sensitive information to generate an initial feature extraction model;
[0131] S720, determining a feature type corresponding to each extracted feature according to the initial feature extraction model and the sensitive information, wherein the feature type includes filtering and normal;
[0132] S730, generating a target feature extraction model according to the feature type for the normal extraction feature, the initial feature extraction model and a preset evaluation index, wherein the preset evaluation index includes one or more of accuracy, recall rate and F1 value;
[0133] S740: Generate target features according to the target feature extraction model.
[0134] It should be noted that, during the feature extraction process, the results of sensitive information identification are used to filter the extracted features and shield those parts containing sensitive privacy information.
[0135] As an example, for each feature, determine whether it contains sensitive privacy information. If it does, the feature type of the extracted feature is considered to be filtering, and the current extracted feature needs to be shielded and does not participate in subsequent feature extraction and application; the feature extraction model is evaluated, and the target feature extraction model is obtained through the accuracy, recall rate and F1 value indicators in the preset indicators. The target feature extraction model can ensure its performance and stability, and ensure the accuracy and reliability of the extracted features; the target feature is obtained through the target feature extraction model.
[0136] In a specific implementation, if the model evaluation result is not ideal, it is necessary to adjust the model parameters or algorithm to improve the performance of the model.
[0137] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0138] Reference Figure 2 , showing a structural block diagram of a data feature extraction system based on artificial intelligence provided by an embodiment of the present application;
[0139] A data feature extraction system based on artificial intelligence, the system comprising:
[0140] An acquisition module 210 is used to acquire heterogeneous data of a plurality of clients;
[0141] A first generating module 220, configured to extract features from the heterogeneous data and generate a standardized feature set;
[0142] A first determination module 230, configured to perform model aggregation processing according to the standardized feature set to determine important features;
[0143] A second determination module 240, configured to determine target data in the heterogeneous data according to the important features;
[0144] A third determination module 250 is used to perform data preprocessing according to the target data to generate standardized data;
[0145] A fourth determination module 260, configured to determine sensitive information in the standardized data based on the standardized data and a preset deep learning model;
[0146] The fifth determination module 270 is used to determine target features based on the sensitive information and the standardized data.
[0147] In one embodiment of the present invention, the first generating module 220 includes:
[0148] A first generating submodule is used to perform principal component analysis on the heterogeneous data to generate dimension-reduced data;
[0149] A first determination submodule, used to determine a feature set based on the dimension reduction data;
[0150] A second determination submodule is used to determine the initial features according to the feature set and a preset algorithm, wherein the preset algorithm is a LASSO algorithm;
[0151] The second generating submodule is used to perform standardization processing on the initial features to generate a standardized feature set.
[0152] In one embodiment of the present invention, the second generating submodule includes:
[0153] A first determining unit, configured to determine a mean and a variance corresponding to each of the initial features;
[0154] A second determining unit, configured to determine a sample value according to the mean and the variance;
[0155] The first generating unit is used to generate the standardized feature set according to the mean, the variance and the sample value.
[0156] In one embodiment of the present invention, the first determining module 230 includes:
[0157] A third determination submodule, configured to determine model parameters corresponding to a plurality of said clients according to said standardized feature set;
[0158] A first processing submodule, used for performing weighted average processing on the model parameters to obtain initial parameters;
[0159] A second processing submodule, used for performing regularization processing according to the target parameter to obtain the target parameter;
[0160] A third generating submodule, used for generating an important feature model according to the target parameter and the standardized feature set;
[0161] The fourth determining submodule is used to determine important features according to the important feature model.
[0162] In one embodiment of the present invention, the third determining module 250 includes:
[0163] A fourth generating submodule, used for converting the target data into format data according to a preset format;
[0164] A fourth determination submodule, used to determine the local data belonging to the participant in the format data;
[0165] A fifth generating submodule, used for performing encryption processing on the local data to generate encrypted data;
[0166] a sixth generation submodule, configured to perform federated learning processing on the encrypted data to generate a standardized model;
[0167] The seventh generating submodule is used to generate the standardized data according to the standardized model.
[0168] In one embodiment of the present invention, the fourth determining module 260 includes:
[0169] An eighth generation submodule, used for performing feature extraction on the standardized data through a convolutional neural network and a recurrent neural network to generate text features;
[0170] A fifth determination submodule is used to perform named entity recognition processing on the text features to determine key information;
[0171] The sixth determination submodule is used to perform semantic analysis on the key information to determine the sensitive information.
[0172] In one embodiment of the present invention, the fifth determining module 270 includes:
[0173] A ninth generating submodule, configured to extract features from the standardized data according to the sensitive information to generate an initial feature extraction model;
[0174] a seventh determination submodule, configured to determine a feature type corresponding to each extracted feature according to the initial feature extraction model and the sensitive information, wherein the feature type includes filtering and normal;
[0175] a tenth generating submodule, configured to generate a target feature extraction model for the normal extraction feature, the initial feature extraction model and a preset evaluation index according to the feature type, wherein the preset evaluation index includes one or more of accuracy, recall rate and F1 value;
[0176] The eleventh generating submodule is used to generate target features according to the target feature extraction model.
[0177] Reference Figure 3 , shows a computer device of a data feature extraction method based on artificial intelligence of the present invention, which may specifically include the following:
[0178] The computer device 12 is in the form of a general-purpose computing device, and the components of the computer device 12 may include but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0179] The bus 18 represents one or more of several types of bus 18 structures, including a memory bus 18 or memory controller, a peripheral bus 18, an accelerated graphics port, a processor, or a local bus 18 using any of a variety of bus 18 architectures. These architectures include, by way of example, but are not limited to, an Industry Standard Architecture (ISA) bus 18, a Micro Channel Architecture (MAC) bus 18, an Enhanced ISA bus 18, an Audio Video Electronics Standards Association (VESA) local bus 18, and a Peripheral Component Interconnect (PCI) bus 18.
[0180] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0181] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write to non-removable, non-volatile magnetic media (commonly referred to as a "hard drive"). Although Figure 3 Not shown, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory may include at least one program product having a set (e.g., at least one) of program modules 42, which are configured to perform the functions of various embodiments of the present invention.
[0182] A program / utility 40 having a set (at least one) of program modules 42 may be stored in, for example, a memory, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules 42, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0183] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, cameras, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., local area networks (LANs)), wide area networks (WANs), and / or public networks (e.g., the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the computer device 12 via a bus 18. It should be understood that although Figure 3 Not shown, other hardware and / or software modules may be used in conjunction with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units 16, external disk drive arrays, RAID systems, tape drives, and data backup storage systems 34, etc.
[0184] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing a data feature extraction method based on artificial intelligence provided by an embodiment of the present invention.
[0185] That is, when the processing unit 16 executes the program, it achieves: obtaining heterogeneous data from several clients; performing feature extraction based on the heterogeneous data to generate a standardized feature set; performing model aggregation processing based on the standardized feature set to determine important features; determining target data within the heterogeneous data based on the important features; performing data preprocessing based on the target data to generate standardized data; determining sensitive information within the standardized data based on the standardized data and a preset deep learning model; and determining target features based on the sensitive information and the standardized data.
[0186] In an embodiment of the present invention, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a data feature extraction method based on artificial intelligence as provided in all embodiments of the present application:
[0187] That is, when the program is executed by the processor, it is implemented as follows: obtaining heterogeneous data from several clients; performing feature extraction based on the heterogeneous data to generate a standardized feature set; performing model aggregation processing based on the standardized feature set to determine important features; determining target data within the heterogeneous data based on the important features; performing data preprocessing based on the target data to generate standardized data; determining sensitive information within the standardized data based on the standardized data and a preset deep learning model; and determining target features based on the sensitive information and the standardized data.
[0188] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device or device.
[0189] Computer-readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0190] The computer program code for performing the operation of the present invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet). The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other.
[0191] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0192] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0193] The above is a detailed introduction to the artificial intelligence-based data feature extraction method and system provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A data feature extraction method based on artificial intelligence, characterized in that: The method comprises: Obtain heterogeneous data from several clients; Extract features from the heterogeneous data to generate a standardized feature set; perform principal component analysis on the heterogeneous data to generate dimension reduction data; determine a feature set based on the dimension reduction data; determine initial features based on the feature set and a preset algorithm, wherein the preset algorithm is a LASSO algorithm; perform standardization on the initial features to generate a standardized feature set; determine a mean and a variance corresponding to each of the initial features; determine a sample value based on the mean and the variance; generate the standardized feature set based on the mean, the variance and the sample value; Performing model aggregation processing based on the standardized feature set to determine important features; Determining target data in the heterogeneous data according to the important features; Performing data preprocessing according to the target data to generate standardized data; Determining sensitive information in the standardized data based on the standardized data and a preset deep learning model; Determine target features based on the sensitive information and the standardized data; extract features from the standardized data based on the sensitive information to generate an initial feature extraction model; determine the feature type corresponding to each extracted feature based on the initial feature extraction model and the sensitive information, wherein the feature types include filtered and normal; generate a target feature extraction model based on the feature type for the normal extracted features, the initial feature extraction model and preset evaluation indicators, wherein the preset evaluation indicators include one or more of accuracy, recall rate and F1 value; generate target features based on the target feature extraction model.
2. The method according to claim 1, characterized in that: The step of performing model aggregation processing according to the standardized feature set to determine important features includes: Determining model parameters corresponding to a plurality of said clients according to said standardized feature set; Performing weighted averaging processing on the model parameters to obtain initial parameters; Performing regularization processing according to the initial parameters to obtain target parameters; generating an important feature model according to the target parameter and the standardized feature set; Important features are determined based on the important feature model.
3. The method according to claim 1, characterized in that The step of performing data preprocessing according to the target data to generate standardized data comprises: Convert the target data into format data according to a preset format; Determining local data belonging to the participant in the formatted data; Performing encryption processing on the local data to generate encrypted data; Performing federated learning processing on the encrypted data to generate a standardized model; The standardized data is generated according to the standardized model.
4. The method according to claim 1, characterized in that: The step of determining the sensitive information in the standardized data based on the standardized data and a preset deep learning model includes: Performing feature extraction on the standardized data by using a convolutional neural network and a recurrent neural network to generate text features; Performing named entity recognition on the text features to determine key information; The key information is semantically analyzed to determine the sensitive information.
5. A data feature extraction system based on artificial intelligence, characterized in that: The system comprises: The acquisition module is used to obtain heterogeneous data from several clients; The first generation module is used to extract features from the heterogeneous data to generate a standardized feature set; perform principal component analysis on the heterogeneous data to generate dimension reduction data; determine a feature set based on the dimension reduction data; determine initial features based on the feature set and a preset algorithm, wherein the preset algorithm is a LASSO algorithm; perform standardization on the initial features to generate a standardized feature set; determine a mean and a variance corresponding to each of the initial features; determine a sample value based on the mean and the variance; and generate the standardized feature set based on the mean, the variance, and the sample value; A first determination module, configured to perform model aggregation processing according to the standardized feature set to determine important features; A second determination module, configured to determine target data in the heterogeneous data according to the important features; A third determination module is used to perform data preprocessing according to the target data to generate standardized data; A fourth determination module, configured to determine sensitive information in the standardized data based on the standardized data and a preset deep learning model; The fifth determination module is used to determine the target feature based on the sensitive information and the standardized data; perform feature extraction on the standardized data based on the sensitive information to generate an initial feature extraction model; determine the feature type corresponding to each extracted feature based on the initial feature extraction model and the sensitive information, wherein the feature types include filtered and normal; generate a target feature extraction model based on the feature type for the normal extracted feature, the initial feature extraction model and preset evaluation indicators, wherein the preset evaluation indicators include one or more of accuracy, recall rate and F1 value; generate the target feature based on the target feature extraction model.
6. A computer device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the method according to any one of claims 1 to 4 when executed by the processor.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Time sequence industrial big data feature extraction and anomaly detection method and system
CN118114034A
Data feature extraction method and system and computer equipment
CN118171076A