Sample Data Volume Joint Expansion Method, Device, Equipment, System and Storage Medium
By generating and encrypting feature vectors, and using the server to perform horizontal joint expansion of sample data, the problem of insufficient sample data is solved, and the generalization ability and recognition accuracy of machine learning models are improved.
Patent Information
- Application Number
- CN202111265684.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-10-28
AI Technical Summary
It is difficult to obtain sufficient amount of sample data in the prior art to improve the generalization ability and recognition accuracy of machine learning models, especially when the sample data volume is small.
Time series data is generated by obtaining the original data, feature vectors are generated and encrypted, and uploaded to the server for processing. The server filters out similar data from the sample data sets of other terminals and adds them to the original data sets, thereby amplifying the sample data volume.
It is realized to obtain sufficient number of sample data similar to local data through non-leakage horizontal joints, improving the generalization ability and recognition accuracy of the machine learning model.
Smart Images

Figure CN113902135B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of machine learning technologies, and in particular, to a method, apparatus, device, system, and storage medium for jointly expanding the amount of sample data. Background Art
[0002] Generally, training machine learning applications or algorithm models requires a large amount of sample data. And how to obtain a sufficient number of sample data to provide for the machine to learn, so as to obtain an application or algorithm model that can solve a specific problem, is a very challenging task.
[0003] For example, when the initiator wants to train an algorithm model that can solve a specific problem (for example, to build an algorithm model for predicting traffic flow to solve the problem of road congestion), due to cold start or other reasons, the amount of sample data it has is small. And only using the sample data owned by the initiator itself to train the machine learning algorithm, the generalization ability and recognition accuracy of the obtained algorithm model are often poor, so it cannot be put into practical use.
[0004] Therefore, how to obtain a sufficient number of sample data to improve the generalization ability and recognition accuracy of the model obtained by machine learning is one of the hot issues that need to be urgently solved in current machine learning. Summary of the Invention
[0005] In view of this, the embodiments of the present disclosure provide a method, apparatus, device, system, and storage medium for jointly expanding the amount of sample data, so as to obtain a sufficient number of sample data for machine learning to use, thereby improving the generalization ability and recognition accuracy of the model obtained by machine learning.
[0006] In the first aspect of the embodiments of the present disclosure, a method for jointly expanding the amount of sample data is provided, which is applied to a first terminal and includes:
[0007] Obtain original data, and generate first time series data according to the original data;
[0008] Generate a first feature vector according to all or part of the data in the first time series data;
[0009] Encrypt the first feature vector to obtain first encrypted data, and upload a first sample data set containing the first encrypted data to the server, so that the server screens out second encrypted data similar to the first encrypted data from second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data amount of the first sample data set.
[0010] In the second aspect of the embodiments of the present disclosure, another method for jointly expanding the amount of sample data is provided, which is applied to a server and includes:
[0011] Receive the first sample data set uploaded by the first terminal and the second sample data sets uploaded by at least one second terminal, where the first sample data set includes at least one first encrypted data, and the second sample data sets include multiple second encrypted data;
[0012] Screen out the second encrypted data similar to the first encrypted data from the second sample data sets, and add the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0013] In a third aspect of the embodiments of the present disclosure, a sample data volume joint expansion device is provided, including:
[0014] A data acquisition module, configured to acquire original data and generate first time series data according to the original data;
[0015] A feature vector generation module, configured to generate a first feature vector according to all or part of the data in the first time series data;
[0016] A data volume expansion module, configured to encrypt the first feature vector to obtain first encrypted data, and upload the first sample data set including the first encrypted data to a server, so that the server screens out the second encrypted data similar to the first encrypted data from the second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0017] In a fourth aspect of the embodiments of the present disclosure, another sample data volume joint expansion device is provided, including:
[0018] A data reception module, configured to receive the first sample data set uploaded by the first terminal and the second sample data sets uploaded by at least one second terminal, where the first sample data set includes at least one first encrypted data, and the second sample data sets include multiple second encrypted data;
[0019] A data screening module, configured to screen out the second encrypted data similar to the first encrypted data from the second sample data sets, and add the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0020] In a fifth aspect of the embodiments of the present disclosure, a sample data volume joint expansion system is provided, including:
[0021] A server, where the server includes the above-mentioned (first) sample data volume joint expansion device;
[0022] A first terminal communicatively connected to the server, where the first terminal includes the above-mentioned (second) sample data volume joint expansion device; and
[0023] At least one second terminal communicatively connected to the server.
[0024] In a sixth aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0025] In a seventh aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0026] The beneficial effects of the embodiments of the present disclosure compared with the prior art at least include: When the initiator (the first terminal) wants to train an algorithm model that can solve a specific problem and currently has a small amount of sample data, in order to improve the generalization ability and recognition accuracy of its algorithm model, it is possible to obtain original data, generate first time series data according to the original data; generate first feature vectors according to all or part of the data in the first time series data; encrypt the first feature vectors to obtain first encrypted data, and upload a first sample data set containing the first encrypted data to the server, so that the server screens out second encrypted data similar to the first encrypted data from the second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data volume of the first sample data set. Through the above method, the first terminal (the initiator lacking sample data) can non-leakage horizontally combine the sample data it has locally with the sample data of the second terminal (other participating parties) via the server, so as to obtain a sufficient number of sample data similar to its local data for machine learning use, thereby improving the generalization ability and recognition accuracy of the algorithm model it wants to build. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0028] Figure 1 is a schematic diagram of the application scenario of the embodiments of the present disclosure;
[0029] Figure 2 is a schematic flowchart of a method for jointly expanding the sample data volume provided by the embodiments of the present disclosure;
[0030] Figure 3It is a schematic flowchart of another method for jointly expanding the sample data volume provided by an embodiment of the present disclosure;
[0031] Figure 4 It is a schematic structural diagram of a device for jointly expanding the sample data volume provided by an embodiment of the present disclosure;
[0032] Figure 5 It is a schematic structural diagram of another device for jointly expanding the sample data volume provided by an embodiment of the present disclosure;
[0033] Figure 6 It is a schematic structural diagram of a system for jointly expanding the sample data volume provided by an embodiment of the present disclosure;
[0034] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0035] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0036] A method and a device for jointly expanding the sample data volume according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0037] Figure 1 It is a schematic architecture diagram of a federated learning according to an embodiment of the present disclosure. As Figure 1 shown, the architecture of federated learning may include a server 101, at least one first terminal 102, and at least one second terminal 103.
[0038] Specifically, federated learning (also known as joint learning) is a distributed machine learning framework with privacy protection and security encryption technologies, which allows decentralized participants to collaborate in training a machine learning model on the premise of not disclosing private data to other participants.
[0039] During the federated learning process, the server 101 establishes a basic model and sends the basic structure and model parameters of the model to at least one first terminal 102 and at least one second terminal 103 that have established a communication connection with it. The first terminal 102 and the second terminal 103 construct a model according to the downloaded basic structure and model parameters, use local data for model training, obtain updated model parameters, and encrypt and upload the updated model parameters to the server 101. The server 101 aggregates the model parameters sent by the first terminal 102 and the second terminal 103 to obtain global model parameters, and sends the global model parameters back to the first terminal 102 and the second terminal 103. The first terminal 102 and the second terminal 103 update their respective models according to the received global model parameters, thereby realizing the training of the model. Since the data uploaded by the first terminal 102 and the second terminal 103 during the federated learning process are model parameters, the local data will not be uploaded to the server 101, and all participating parties can share the final model parameters. Therefore, joint modeling can be achieved on the basis of ensuring data privacy.
[0040] It should be noted that the number of the first terminal 102 and the second terminal 103 is not limited to one or two as described above, but can be set as needed, and the embodiments of the present disclosure do not limit this.
[0041] The sample data volume joint expansion method provided by the embodiments of the present disclosure can be applied to the above-mentioned federated learning architecture. Specifically, the first terminal 102 is usually the party lacking sample data. For example, it is the a power company in City A. When the a power company constructs a power transmission prediction algorithm model required for a smart city project and finds that its local sample data volume is insufficient for training the power transmission prediction algorithm model, the first terminal 102 can first convert its local raw data into first time series data, then generate a first feature vector according to all or part of the data in the first time series data, encrypt the first feature vector to obtain first encrypted data, and then upload a first sample data set containing the first encrypted data to the server 101, so as to perform non-leaking horizontal combination with the data in the second sample data set provided by at least one second terminal 103 via the server 101 (a third party), so as to obtain a sufficient number of sample data similar to its local data for machine federated learning, thereby improving the generalization ability and recognition accuracy of the algorithm model to be constructed.
[0042] Although some initiators can increase the amount of sample data to a certain extent by jointly adopting the sample data owned by other participants (for example, peers in other cities or regions), when jointly adopting the sample data provided by the participants, the internal relationship of the data of each party is not deeply considered, and the generalization ability of the model learned by the machine using these sample data is still poor. For the technical solution provided by the present disclosure, when the first terminal 102 laterally combines with the data in the second sample data set provided by at least one second terminal 103 via the server 101 (a third party), it fully considers the similarity between its own local sample data and the sample data provided by the other party (the second terminal 103), and selects the sample data similar to its own local sample data to jointly expand the amount of its local sample data, and provides these jointly combined sample data sets to the machine for federated learning to construct a target algorithm model. The obtained target algorithm model has good generalization ability and high recognition accuracy.
[0043] Figure 2 It is a schematic flowchart of a method for jointly expanding the amount of sample data provided by an embodiment of the present disclosure. Figure 2 The method for jointly expanding the amount of sample data can be performed by Figure 1 the first terminal 102. As Figure 2 shown, the method for jointly expanding the amount of sample data includes:
[0044] Step S201, obtain the original data, and generate the first time series data according to the original data.
[0045] Among them, the original data usually refers to the data stored locally by the first terminal 102 (the initiator, such as a gas company, an electric power company, a weather forecast company, etc. in a certain city). For example, the weather data (including temperature, humidity, light intensity, etc.) collected in real time by the data collectors deployed in some weather stations by a weather forecast company and stored locally by it. And the data stored locally by it is the original data.
[0046] As an example, the data collected in real time by the data collectors of a weather forecast company through weather stations is usually time series data, that is, time series data, which refers to the data column recorded in the order of time for the same unified index. Among them, the time series data can be period data or time point data. For example, the time series data can be the weather data of a certain month or a certain year, or the weather data at a certain time point of a certain day.
[0047] Step S202, generate the first feature vector according to all or part of the data in the first time series data.
[0048] As an example, in combination with the foregoing, assume that the original data stored locally by a certain weather forecasting company is the weather data collected in 12 months from January to December 20XX. The weather data for each month includes the weather data collected at each time point of 24 hours per day. Sort the weather data for these 12 months in the order from January to December, and the first time series data can be obtained.
[0049] Exemplarily, a first feature vector is generated based on all or part of the data in the first time series data. Specifically, the first feature vector can be generated according to all the weather data for 12 months in the above first time series data, or can be generated according to the weather data for one or more months in the above first time series data.
[0050] Among them, the first feature vector refers to extracting features from the first time series data to generate new features of the first time series data, and arranging these new features in sequence, that is, generating the first feature vector. For example, the first time series data is the weather data for December. Extract features from the data collected at each time point (in hours) of each day in December, that is, a total of 31*24 = 744 time point data (such as calculating the sum of squares, mean, variance, etc. of these 744 data), to obtain new features (i.e., the sum of squares, mean, variance), and then arrange the sum of squares, mean, variance in sequence, that is, generate the first feature vector.
[0051] Step S203: Encrypt the first feature vector to obtain first encrypted data, and upload the first sample data set containing the first encrypted data to the server, so that the server filters out second encrypted data similar to the first encrypted data from the second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0052] Among them, the second terminal 103 is usually the party that may have more sample data required by the first terminal 102 when constructing or improving its target algorithm model. Or, it is a cooperation party of the first terminal 102 in jointly developing a certain project, etc. For example, the first terminal 102 is the initiator of the project, and the second terminal 103 is the participant of the project.
[0053] As an example, the Locality Sensitive Hashing (LSH) algorithm can be used to encrypt the first feature vector to obtain the first encrypted data. Similarly, before uploading the second sample data set to the server 101, the second terminal 103 (participant) can also process its original data in the same way as the first terminal 102 processes its original data and encrypts it to obtain the first encrypted data, so as to obtain the second encrypted data, and then package the second encrypted data into the second sample data set and upload it to the server 101.
[0054] The first terminal 102 packages the first encrypted data obtained through the above processing into the first sample data set and uploads it to the server 101. After receiving the first sample data set uploaded by the first terminal 102, the server 101 can screen out the second encrypted data similar to the first encrypted data from the second sample data sets uploaded by at least one second terminal, and add the second encrypted data to the first sample data set, so as to expand the data volume of the first sample data set, provide sufficient sample data for subsequent machine learning, and then improve the generalization ability and recognition accuracy of the model.
[0055] The technical solution provided by the embodiments of the present disclosure includes: obtaining original data, generating first time series data according to the original data; generating a first feature vector according to all or part of the data in the first time series data; encrypting the first feature vector to obtain the first encrypted data, and uploading the first sample data set including the first encrypted data to the server, so that the server screens out the second encrypted data similar to the first encrypted data from the second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data volume of the first sample data set. Through the above method, the first terminal (the initiating party lacking sample data) can non-leakage horizontally combine the sample data it owns locally with the sample data of the second terminal (other participants) via the server, so as to obtain a sufficient number of sample data similar to its local data for machine learning, and then improve the generalization ability and recognition accuracy of the algorithm model to be constructed.
[0056] In some embodiments, the above step S202 includes:
[0057] Select the latest data ranked last from the first time series data, and generate a first feature vector according to the latest data.
[0058] As an example, as described above, the latest data ranked last can be selected from the first time series data (including weather data for 12 months from January to December), that is, the weather data for December, and a first feature vector can be generated according to the weather data for December.
[0059] Generally, data closer to the current time point or period can better reflect recent data changes. By selecting the latest data ranked last in the first time series data, a first feature vector is generated, and the first feature vector is encrypted to obtain first encrypted data. Then, the server 101 filters out second encrypted data similar to the first encrypted data from the second sample data set provided by the second terminal 103, and adds the second encrypted data to the first sample data set. Subsequently, using the sample data of the first sample data set to train the model, the obtained model can better predict the changes at the next time point or period.
[0060] In some embodiments, the above step S203 includes:
[0061] Select the recent data ranked last M bits from the first time series data, generate a third feature vector according to the recent data, and encrypt the third feature vector to obtain third encrypted data, where M is a positive integer greater than or equal to 2;
[0062] Upload the first sample data set containing the first encrypted data and the third encrypted data to the server, so that the server filters out second encrypted data similar to the first encrypted data and / or the third encrypted data from the second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set.
[0063] As an example, in combination with the foregoing example, the weather data ranked last 2 bits (i.e., November and December) can be selected from the first time series data as the recent data, and a third feature vector is generated according to the recent data. Specifically, after summarizing the weather data of November and December, feature extraction is performed to obtain at least 2 new features, and these new features are arranged in order, that is, a third feature vector is generated.
[0064] Similarly, the local sensitive hashing algorithm can be used to encrypt the above third feature vector to obtain third encrypted data.
[0065] The first terminal 102 uploads the first sample data set containing the above first encrypted data and third encrypted data to the server 101. After receiving the first sample data set, the server 101 can filter out second encrypted data similar to the first encrypted data and / or the third encrypted data from the second sample data sets uploaded by the second terminal, and add the second encrypted data to the first sample data set.
[0066] For example, the first sample dataset contains A (the first encrypted data) and B (the third encrypted data), and the second sample dataset includes six second encrypted data, namely a, b, c, d, e, and f. The server 101 compares A with a, b, c, d, e, and f respectively, and finds that the data similar to A are a and b; it compares B with a, b, c, d, e, and f respectively, and finds that the data similar to B is f. Then, a, b, and f can be added to the first sample dataset to obtain a new first sample dataset containing A, B, a, b, and f, that is, the data volume of the original first sample dataset is jointly expanded from 2 to 5.
[0067] The technical solution provided by the embodiments of the present disclosure generates a third feature vector by selecting the last M pieces of recent data in the first time series data, encrypts the third feature vector to obtain third encrypted data, and uploads the first sample dataset containing the first encrypted data and the third encrypted data to the server 101. The server 101 filters out the second encrypted data similar to the first encrypted data and / or the third encrypted data (the latest and / or recent data) from the second sample dataset provided by the second terminal 103, so as to quickly expand the data volume of its first sample dataset and reduce the cost of data collection.
[0068] In some embodiments, in the above step S203, encrypting the first feature vector to obtain the first encrypted data may specifically be:
[0069] Initialize and randomly generate a two-dimensional matrix, where the number of rows of the two-dimensional matrix is the same as the dimension of the first feature vector, and the number of columns is a random number;
[0070] Multiply the first feature vector by the randomly generated two-dimensional matrix to obtain the first encrypted data, and the first encrypted data is a hash code.
[0071] As an example, assume that the first feature vector is a 1*3-dimensional vector [0.1, 0.3, 0.5], and the randomly generated two-dimensional matrix initialized is a 3*4 matrix Multiply the first feature vector by this two-dimensional matrix, that is The 1*4 matrix [1.4, 1.9, 3.8, 5.2] is the first encrypted data (hash code).
[0072] The technical solution provided by the embodiments of the present disclosure encrypts the first feature vector by using the local hashing sensitive hashing algorithm to obtain the first encrypted data. Specifically, the first feature vector is encrypted by using a two-dimensional feature vector randomly generated with the same dimension as the first feature vector to obtain the first encrypted data, and then the first encrypted data is packaged into the first sample data set and uploaded to the server 101. Similarly, when the second terminal 103 uploads the second sample data set, the above method can also be used to encrypt its local data and then upload it. By encrypting their local data and then uploading it, and screening out the sample data required by the first terminal 102 via a third party (the server 101), the first terminal 102 and the second terminal 103 can jointly expand the first sample data set in a non-disclosure manner, that is, while protecting the privacy of the local data of both the first terminal 102 and the second terminal 103, the problem that the first terminal 102 lacks sample data for training the model can be solved at the same time.
[0073] Figure 3 It is a schematic flowchart of another method for jointly expanding the sample data volume provided by the embodiments of the present disclosure. Figure 3 The method for jointly expanding the sample data volume can be executed by Figure 1 the server 101. As Figure 2 shown, the method for jointly expanding the sample data volume includes:
[0074] Step S301, receiving the first sample data set uploaded by the first terminal and the second sample data sets uploaded by at least one second terminal, where the first sample data set includes at least one first encrypted data, and the second sample data sets include multiple second encrypted data.
[0075] Step S302, screening out the second encrypted data similar to the first encrypted data from the second sample data sets, and adding the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0076] Specifically, when the initiator (the first terminal 102) wants to train an algorithm model that can solve a specific problem and currently has a small amount of sample data, it can send a request to the server 101 to horizontally combine the data of other participating parties. This request contains the first sample data set, and the first sample data set contains the first encrypted data. After the server 101 receives the first sample data set uploaded by the first terminal 102 and the second sample data sets uploaded by at least one second terminal that contain multiple second encrypted data, it can screen out the second encrypted data similar to the first encrypted data by comparing the similarity between the first encrypted data and each second encrypted data, and add the second encrypted data to the first sample data set, thereby expanding the data volume of the first sample data set.
[0077] The technical solution provided by the embodiments of the present disclosure, through the above method, the server 101, as a third party, can help the first terminal 102 (the initiator lacking sample data) and the second terminal (other participating parties) to perform non-disclosure horizontal combination of sample data, so as to obtain a sufficient number of sample data similar to its local data for machine learning, thereby improving the generalization ability and recognition accuracy of the algorithm model to be constructed.
[0078] In some embodiments, the above step S302 includes:
[0079] Calculate the similarity between each second encrypted data and the first encrypted data respectively, and add the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set.
[0080] Among them, the preset threshold range can be set according to the actual situation. For example, it can be set to be greater than or equal to 85%, or greater than or equal to 90%, etc.
[0081] As an example, assume that the preset threshold range is greater than or equal to 85%, the first sample data set contains the first encrypted data A, and the second sample data set contains three second encrypted data a, b, and c. Then calculate the similarities between A and a, A and b, and A and c respectively to obtain three similarities. For example, the similarities between A and a, A and b, and A and c are 90%, 75%, and 60% respectively. Then add a whose similarity with A ≥ 85% to the first sample data set.
[0082] In some embodiments, the above-mentioned calculating the similarity between each second encrypted data and the first encrypted data respectively, and adding the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set can be specifically:
[0083] Perform an exclusive OR operation on the string of each second encrypted data and the string of the first encrypted data respectively to obtain the Hamming distance between each second encrypted data and the first encrypted data;
[0084] Add the second encrypted data whose Hamming distance from the first encrypted data meets the preset threshold range to the first sample data set.
[0085] Among them, exclusive OR is a mathematical operator. It is applied to logical operations. The mathematical symbol of exclusive OR is The computer symbol is "xor". Its operation rule is: If the two values of a and b are different, the exclusive OR result is 1. If the two values of a and b are the same, the exclusive OR result is 0. Exclusive OR is also called half adder operation, and its operation rule is equivalent to binary addition without carry.
[0086] The Hamming distance is used in error control coding for data transmission. The Hamming distance is a concept that represents the number of different corresponding bits between two words of the same length. It is denoted as d(x,y) for the Hamming distance between two words x and y. Perform an exclusive OR operation on two strings and count the number of 1s in the result, and that number is the Hamming distance.
[0087] As an example, when the string of the first encrypted data A obtained after the above encryption process is [1, 0.8, 0.7, 0.5], the string of the second encrypted data a is [1, 0.3, 0.7, 0.5], the string of b is [1.2, 0.4, 0.8, 0.5], and the string of c is [0.2, 2.4, 1, 0.7]. Perform an exclusive OR operation on the first encrypted data and the second encrypted data respectively. Specifically, the Hamming distance obtained after the exclusive OR operation of the string [1, 0.8, 0.7, 0.5] of the first encrypted data A and the string [1, 0.3, 0.7, 0.5] of the second encrypted data a is 1; the Hamming distance obtained after the exclusive OR operation of the string [1, 0.8, 0.7, 0.5] of the first encrypted data A and the string [1.2, 0.4, 0.8, 0.5] of the second encrypted data b is 3; the Hamming distance obtained after the exclusive OR operation of the string [1, 0.8, 0.7, 0.5] of the first encrypted data A and the string [0.2, 2.4, 1, 0.7] of the second encrypted data c is 4.
[0088] The preset threshold range here can also be flexibly set according to the actual situation. For example, it can be that the Hamming distance is less than or equal to 3, or it can be that the Hamming distance is less than or equal to 2, etc.
[0089] As an example, combining the foregoing example, assuming that the preset threshold range is that the Hamming distance is less than or equal to 3, then the Hamming distances between the second encrypted data a and b and the first encrypted data A meet this preset threshold range. At this time, the second encrypted data a and b can be added to the first sample data set.
[0090] In some other embodiments, the above-mentioned calculating the similarity between each second encrypted data and the first encrypted data respectively, and adding the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set can also be specifically:
[0091] Sort the similarities between each second encrypted data and the first encrypted data from high to low to obtain a sorting result;
[0092] According to the sorting result, add the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set one by one until the current data volume in the first sample data set reaches the preset data volume required for training the model.
[0093] As an example, in combination with the foregoing example, the smaller the Hamming distance, the more similar the two data are. Sort the similarities between the above second encrypted data a, b, and c and the first encrypted data A in descending order, and the sorting result obtained is a > b > c.
[0094] Suppose the preset data volume required for training the model is 5, and the current data volume in the first sample data set is 3, that is, 2 data are still missing. The preset threshold range is that the Hamming distance is less than or equal to 3. Then, the second encrypted data a can be added to the first sample data set first, and then the second encrypted data b can be added to the first sample data set until the data volume of the first sample data set reaches 5, that is, the joint expansion of the first sample data set is completed.
[0095] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here.
[0096] The following is an embodiment of the apparatus of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For details not disclosed in the embodiment of the apparatus of the present disclosure, please refer to the method embodiment of the present disclosure.
[0097] Figure 4 It is a schematic diagram of a device for jointly expanding the sample data volume provided by an embodiment of the present disclosure. As Figure 4 shown, the device for jointly expanding the sample data volume includes:
[0098] A data acquisition module 401, configured to acquire original data and generate first time series data according to the original data;
[0099] A feature vector generation module 402, configured to generate a first feature vector according to all or part of the data in the first time series data;
[0100] A data volume expansion module 403, configured to encrypt the first feature vector to obtain first encrypted data, and upload a first sample data set including the first encrypted data to a server, so that the server screens out second encrypted data similar to the first encrypted data from second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0101] The technical solution provided by the embodiments of the present disclosure obtains original data through the data acquisition module 401, and generates first time series data according to the original data; the feature vector generation module 402 generates a first feature vector according to all or part of the data in the first time series data; the data volume expansion module 403 encrypts the first feature vector to obtain first encrypted data, and uploads a first sample data set including the first encrypted data to the server, so that the server filters out second encrypted data similar to the first encrypted data from second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data volume of the first sample data set. Through the above device, the first terminal (the initiator lacking sample data) can non-leakably horizontally combine the sample data it owns locally with the sample data of the second terminal (other participating parties) via the server, so as to obtain a sufficient number of sample data similar to its local data for machine learning, thereby improving the generalization ability and recognition accuracy of the algorithm model to be constructed.
[0102] In some embodiments, the above-mentioned feature vector generation module 402 includes:
[0103] The first vector generation unit is configured to select the latest data ranked last from the first time series data, and generate a first feature vector according to the latest data.
[0104] In some embodiments, the above-mentioned data volume expansion module 403 includes:
[0105] The first encryption unit is configured to select the recent data ranked last M bits from the first time series data, generate a third feature vector according to the recent data, and encrypt the third feature vector to obtain third encrypted data, where M is a positive integer greater than or equal to 2;
[0106] The data upload unit is configured to upload a first sample data set including the first encrypted data and the third encrypted data to the server, so that the server filters out second encrypted data similar to the first encrypted data and / or the third encrypted data from second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set.
[0107] In some embodiments, the above-mentioned data volume expansion module 403 further includes:
[0108] The matrix generation unit is configured to initialize and randomly generate a two-dimensional matrix, where the number of rows of the two-dimensional matrix is the same as the dimension of the first feature vector, and the number of columns is a random number;
[0109] The second encryption unit is configured to multiply the first feature vector by the randomly generated two-dimensional matrix to obtain first encrypted data, and the first encrypted data is a hash code.
[0110] Figure 5 It is a schematic diagram of another sample data volume joint expansion device provided by an embodiment of the present disclosure. As Figure 5 shown, the sample data volume joint expansion device includes:
[0111] A data receiving module 501, configured to receive a first sample data set uploaded by a first terminal, and second sample data sets uploaded by at least one second terminal, wherein the first sample data set includes at least one first encrypted data, and the second sample data sets include a plurality of second encrypted data;
[0112] A data screening module 502, configured to screen out second encrypted data similar to the first encrypted data from the second sample data sets, and add the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0113] The technical solution provided by the embodiment of the present disclosure, through the above device, can help the first terminal 102 (the initiating party lacking sample data) and the second terminal (other participating parties) to perform non-disclosed horizontal association of sample data, so as to obtain a sufficient number of sample data similar to its local data for machine learning, and further improve the generalization ability and recognition accuracy of the algorithm model to be constructed.
[0114] In some embodiments, the above data screening module 502 includes:
[0115] A first data adding unit, configured to calculate the similarity between each second encrypted data and the first encrypted data respectively, and add the second encrypted data whose similarity with the first encrypted data meets a preset threshold range to the first sample data set.
[0116] In some embodiments, the above data screening module 502 further includes:
[0117] A Hamming distance calculation unit, configured to perform an exclusive OR operation on the string of each second encrypted data and the string of the first encrypted data respectively to obtain the Hamming distance between each second encrypted data and the first encrypted data;
[0118] A second data adding unit, configured to add the second encrypted data whose Hamming distance from the first encrypted data meets a preset threshold range to the first sample data set.
[0119] In some embodiments, the above data screening module 502 further includes:
[0120] A sorting unit, configured to sort the similarity between each second encrypted data and the first encrypted data from high to low to obtain a sorting result;
[0121] A third data increment unit, configured to incrementally add second encrypted data that meets a preset threshold range of similarity to the first encrypted data to the first sample data set one by one according to the sorting result until the current data volume in the first sample data set reaches the preset data volume required for training the model.
[0122] Figure 6 It is a schematic structural diagram of a sample data volume joint expansion system provided by an embodiment of the present disclosure.
[0123] As Figure 6 shown, the sample data volume joint expansion system includes:
[0124] A server 101, the server 101 includes a sample data volume joint expansion device as Figure 4 shown; a first terminal 102 communicatively connected to the server 101, the first terminal 102 includes a sample data volume joint expansion device as Figure 3 shown; and at least one second terminal 103 communicatively connected to the server 101.
[0125] Specifically, the first terminal 102 (initiator) and the server 101 can communicate through networks, Bluetooth, etc. The first terminal 102 can process its local original data through the above encryption method to obtain first encrypted data, and pack the first encrypted data into a first sample data set and upload it to the server 101. At least one second terminal 103 (participant) and the server 101 can communicate through networks, Bluetooth, etc. The second terminal 103 can refer to the encryption processing method of the first terminal 102 for its local original data, encrypt its local original data to obtain second encrypted data, and pack the second encrypted data into a second sample data set and upload it to the server 101. After the server 101 receives the first sample data set uploaded by the first terminal 102 and the second sample data set uploaded by the second terminal 103, it can compare the similarity between the first encrypted data in the first sample data set and each second encrypted data in the second sample data set respectively, then screen out the second encrypted data similar to the first encrypted data, and add the second encrypted data to the first sample data set to expand the data volume of the first sample data set.
[0126] The technical solution provided by the embodiment of the present disclosure can use the server 101 as a third party to help the first terminal 102 (the initiator lacking sample data) and the second terminal 103 (other participants) perform non-leaking horizontal joint of sample data, so as to obtain a sufficient number of sample data similar to its local data for machine learning, thereby improving the generalization ability and recognition accuracy of the algorithm model to be constructed.
[0127] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0128] Figure 7 FIG. 4 is a schematic diagram of an electronic device 700 provided by an embodiment of the present disclosure. As Figure 7 shown, the electronic device 700 of this embodiment includes: a processor 701, a memory 702, and a computer program 703 stored in the memory 702 and executable on the processor 701. When the processor 701 executes the computer program 703, the steps in the above method embodiments are implemented. Alternatively, when the processor 701 executes the computer program 703, the functions of each module / unit in the above device embodiments are implemented.
[0129] Exemplarily, the computer program 703 may be divided into one or more modules / units. The one or more modules / units are stored in the memory 702 and executed by the processor 701 to complete the present disclosure. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 703 in the electronic device 7.
[0130] The electronic device 700 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 700 may include, but is not limited to, the processor 701 and the memory 702. Those skilled in the art can understand that Figure 7 merely an example of the electronic device 700, and does not constitute a limitation to the electronic device 700. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0131] The processor 701 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0132] The memory 702 can be an internal storage unit of the electronic device 700, for example, the hard disk or memory of the electronic device 700. The memory 702 can also be an external storage device of the electronic device 700, for example, a plug-in hard disk equipped on the electronic device 700, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 702 can also include both the internal storage unit of the electronic device 700 and an external storage device. The memory 702 is used to store computer programs and other programs and data required by the electronic device. The memory 702 can also be used to temporarily store data that has been output or will be output.
[0133] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.
[0134] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0135] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0136] In the embodiments provided in the present disclosure, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0137] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0139] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present disclosure, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0140] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit it; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the present disclosure's various embodiments, and should all be included within the protection scope of the present disclosure.
Claims
1. A method for jointly expanding the sample data volume, characterized in that, applied to the first terminal, including: Obtain the original data and generate first time-series data according to the original data; Generate a first feature vector according to all or part of the data in the first time-series data; Encrypt the first feature vector to obtain first encrypted data, and upload a first sample data set containing the first encrypted data to the server, so that the server screens out second encrypted data similar to the first encrypted data from second sample data sets uploaded by at least one second terminal, and add the second encrypted data to the first sample data set to expand the data volume of the first sample data set; Wherein, each of the second terminals encrypts its local original data according to the encryption processing method of the local original data of the first terminal to obtain second encrypted data, and packs the second encrypted data into a second sample data set and uploads it to the server.
2. The method for jointly expanding the sample data volume according to claim 1, characterized in that, The generating a first feature vector according to all or part of the data in the first time-series data includes: Select the latest data ranked last in the first time-series data, and generate a first feature vector according to the latest data.
3. The method for jointly expanding the sample data volume according to claim 2, characterized in that, The encrypting the first feature vector to obtain first encrypted data, and uploading a first sample data set containing the first encrypted data to the server, so that the server screens out second encrypted data similar to the first encrypted data from second sample data sets uploaded by at least one second terminal, and add the second encrypted data to the first sample data set includes: Select the recent data ranked last M bits in the first time-series data, generate a third feature vector according to the recent data, and encrypt the third feature vector to obtain third encrypted data, where M is a positive integer greater than or equal to 2; Upload a first sample data set containing the first encrypted data and the third encrypted data to the server, so that the server screens out second encrypted data similar to the first encrypted data and / or the third encrypted data from second sample data sets uploaded by at least one second terminal, and add the second encrypted data to the first sample data set.
4. The method for jointly expanding the sample data volume according to any one of claims 1 to 3, characterized in that, The encrypting the first feature vector to obtain first encrypted data includes: Initialize and randomly generate a two-dimensional matrix, where the number of rows of the two-dimensional matrix is the same as the dimension of the first feature vector, and the number of columns is a random number; Multiply the first feature vector by the randomly generated two-dimensional matrix to obtain first encrypted data, and the first encrypted data is a hash code.
5. A method for jointly expanding the sample data volume, characterized in that, applied to the server, including: Receive the first sample data set uploaded by the first terminal and the second sample data sets uploaded by at least one second terminal, where the first sample data set includes at least one first encrypted data, and the second sample data sets include multiple second encrypted data; Screen out the second encrypted data similar to the first encrypted data from the second sample data sets, and add the second encrypted data to the first sample data set to expand the data volume of the first sample data set; Among them, the first terminal obtains the original data, generates first time series data according to the original data; generates a first feature vector according to all or part of the data in the first time series data; encrypts the first feature vector to obtain the first encrypted data, and uploads the first sample data set containing the first encrypted data to the server; Each of the second terminals encrypts its local original data with reference to the encryption processing method of the local original data of the first terminal to obtain second encrypted data, and packs the second encrypted data into a second sample data set and uploads it to the server.
6. The method for jointly expanding the sample data volume according to claim 5, characterized in that The screening out the second encrypted data similar to the first encrypted data from the second sample data sets and adding the second encrypted data to the first sample data set includes: Calculate the similarity between each second encrypted data and the first encrypted data respectively, and add the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set.
7. The method for jointly expanding the sample data volume according to claim 6, characterized in that The calculating the similarity between each second encrypted data and the first encrypted data respectively, and adding the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set includes: Perform an exclusive OR operation on the string of each second encrypted data and the string of the first encrypted data respectively to obtain the Hamming distance between each second encrypted data and the first encrypted data; Add the second encrypted data whose Hamming distance from the first encrypted data meets the preset threshold range to the first sample data set.
8. The method for jointly expanding the sample data volume according to claim 6, characterized in that The calculating the similarity between each second encrypted data and the first encrypted data respectively, and adding the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set includes: Sort the similarities between each second encrypted data and the first encrypted data from high to low to obtain a sorting result; According to the sorting result, add the second encrypted data whose similarity with the first encrypted data meets the preset threshold range to the first sample data set one by one until the current data volume in the first sample data set reaches the preset data volume required for training the model.
9. A device for jointly expanding the sample data volume, characterized in that including: A data acquisition module, configured to acquire original data and generate first time-series data according to the original data; A feature vector generation module, configured to generate a first feature vector according to all or part of the data in the first time-series data; A data volume expansion module, configured to encrypt the first feature vector to obtain first encrypted data, and upload a first sample data set containing the first encrypted data to a server, so that the server screens out second encrypted data similar to the first encrypted data from second sample data sets uploaded by at least one second terminal, and adds the second encrypted data to the first sample data set to expand the data volume of the first sample data set; Wherein, each of the second terminals encrypts its local original data with reference to the encryption processing method of the local original data of the first terminal to obtain second encrypted data, and packs the second encrypted data into a second sample data set and uploads it to the server.
10. A device for jointly expanding sample data volume Characterized in that It includes: A data receiving module, configured to receive a first sample data set uploaded by a first terminal and second sample data sets uploaded by at least one second terminal, wherein the first sample data set includes at least one first encrypted data, and the second sample data set includes a plurality of second encrypted data; A data screening module, configured to screen out second encrypted data similar to the first encrypted data from the second sample data set, and add the second encrypted data to the first sample data set to expand the data volume of the first sample data set; Wherein, the first terminal acquires original data, generates first time-series data according to the original data; generates a first feature vector according to all or part of the data in the first time-series data; encrypts the first feature vector to obtain first encrypted data, and uploads a first sample data set containing the first encrypted data to the server; Each of the second terminals encrypts its local original data with reference to the encryption processing method of the local original data of the first terminal to obtain second encrypted data, and packs the second encrypted data into a second sample data set and uploads it to the server.
11. A system for jointly expanding sample data volume Characterized in that It includes: A server, the server includes the device for jointly expanding sample data volume according to claim 10; A first terminal communicatively connected to the server, the first terminal includes the device for jointly expanding sample data volume according to claim 9; And At least one second terminal communicatively connected to the server.
12. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, Characterized in that When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
13. A computer-readable storage medium, the computer-readable storage medium stores a computer program, Characterized in that When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Spectral clustering method, device and system, computer equipment and storage medium
CN111310817A
Data processing system and method
CN112989399A