A system, a method executed by a computer, a computer program (data generation and classification based on a quantum kernel)
The system addresses the challenge of class imbalance in quantum machine learning by using quantum kernels for data balancing and classification, achieving superior accuracy and efficiency in dataset transformation and model training.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Existing data classification methods, particularly in quantum machine learning, face challenges in balancing datasets with unequal class distributions, leading to biased models that neglect minority classes, and classical techniques like SMOTE are not well-suited for quantum computing.
A system and method utilizing quantum kernels to automatically select and apply quantum kernels for oversampling or undersampling data, generating new datasets, and performing data classification tasks, leveraging quantum computing's capabilities for efficient data distribution transformation and classification.
The approach provides superior classification accuracy by identifying optimal quantum kernels for data generation and classification, balancing datasets, and enhancing the performance of quantum machine learning models.
Smart Images

Figure 2026064235000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to quantum computing, and more specifically, to data classification based on quantum kernels.
Summary of the Invention
Problems to be Solved by the Invention
[0002]
Means for Solving the Problems
[0003] Hereinafter, an overview is presented to provide a basic understanding of one or more embodiments described in this specification. This overview is not intended to identify key or important elements, nor to define the scope or claims of a particular embodiment. Its sole purpose is to present concepts in a simplified form as a prelude to the more detailed description that follows. In one or more embodiments described in this specification, systems, computer-executed methods, apparatuses, and / or computer program products that enable data classification based on quantum kernels are discussed.
[0004] According to one embodiment, a system is provided. The system may include memory capable of storing computer executable components. The system may further include a processor capable of executing computer executable components stored in memory, wherein the computer executable components may include an access component that can access an input dataset. The computer executable components may further include a data generation component that can generate a plurality of new datasets using a plurality of quantum kernels based on the input dataset. As used herein, the terms “new dataset” and “new input dataset” mean versions of the input dataset modified by overbalancing or underbalancing.
[0005] According to various embodiments, the above system can be implemented as a method executed by a computer or as a computer program product. [Brief explanation of the drawing]
[0006] One or more embodiments are described below in the section on embodiments for carrying out the invention with reference to the following drawings.
[0007] [Figure 1] The block diagram shows an exemplary, non-limiting system, according to one or more embodiments described herein, that can generate novel data and perform data classification tasks using a quantum kernel.
[0008] [Figure 2]Another block diagram of an exemplary, non-limiting system capable of generating novel data and performing data classification tasks using a quantum kernel, according to one or more embodiments described herein, is shown.
[0009] [Figure 3] This specification shows a flow diagram of an exemplary, non-limiting method that can generate novel data and perform data classification tasks using a quantum kernel, according to one or more embodiments described herein.
[0010] [Figure 4] This provides an example of a non-limiting process for describing oversampling techniques.
[0011] [Figure 5] This specification provides a non-limiting table illustrating the experimental results obtained by one or more embodiments described herein.
[0012] [Figure 6] This specification shows an illustrative, non-limiting graph illustrating experimental data clustering results according to one or more embodiments described herein.
[0013] [Figure 7] This specification shows non-limiting illustrative graphs illustrating experimental training and test results according to one or more embodiments described herein.
[0014] [Figure 8] This specification shows a flow diagram of an exemplary, non-limiting method for generating novel data using a quantum kernel, according to one or more embodiments described herein.
[0015] [Figure 9] This specification shows a flow diagram of an exemplary, non-limiting method that can be used to select a pair of quantum kernels for performing data generation and classification, according to one or more embodiments described herein.
[0016] [Figure 10] A block diagram of an exemplary, non-limiting operating environment that can facilitate one or more embodiments described herein is shown.
BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The following detailed description is merely exemplary and is not intended to limit the embodiments and / or the application or use of the embodiments. Further, it is not intended to be constrained by any of the information presented in the foregoing Background or Summary of the Invention sections or the Detailed Description of the Invention section.
[0018] Reference will now be made to the drawings, in which one or more embodiments will be described, and throughout the drawings, like reference numerals are used to refer to like elements. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a more thorough understanding of one or more embodiments. However, it will be apparent that in various instances, one or more embodiments may be practiced without these specific details.
[0019] According to one embodiment, a system is provided. The system can include a memory capable of storing computer-executable components. The system can further include a processor capable of executing the computer-executable components stored in the memory, where the computer-executable components can include an access component capable of accessing an input data set, and the computer-executable components can further include a data generation component capable of generating a plurality of new data sets using a plurality of quantum kernels based on the input data set.
[0020] Such an embodiment of the system can provide the advantage of generating a new dataset that can be used in quantum machine learning (QML)-based data classification tasks using quantum computing.
[0021] In one or more embodiments of the above system, the data generation component can further transform the data distribution of the input dataset into respective new data distributions using each quantum kernel of a plurality of quantum kernels. The data generation component can further generate respective new datasets of the plurality of new datasets by oversampling or undersampling data from each new data distribution. The data generation component can further generate respective classification datasets based on each new dataset, where the classification dataset based on the new dataset can include the input dataset and the new dataset.
[0022] Such an embodiment of the system can provide multiple advantages, including capturing the data distribution in the input dataset using quantum kernels and generating a plurality of new data distributions, generating a plurality of new datasets based on the data distributions, and generating a plurality of classification datasets that can be used for QML-based data classification tasks and for training different machine learning models for QML-based data classification tasks.
[0023] In one or more embodiments of the above system, the quantum kernel selection component can randomly select a quantum kernel from a plurality of quantum kernels, where the quantum kernel can be used to generate a new dataset of the plurality of new datasets.
[0024] Such embodiments of the system can offer several advantages, including capturing the data distribution in an input dataset using a quantum kernel and generating multiple novel data distributions that can be used to generate classification datasets for QML-based data classification.
[0025] In one or more embodiments of the above system, the data classification component can classify each classification dataset using multiple quantum kernels.
[0026] Such embodiments of the system can leverage the capabilities of quantum computing to offer the advantage of classifying data and performing computations more efficiently than classical systems, potentially leading to an acceleration of quantum kernel computation and classification tasks.
[0027] In one or more embodiments of the above system, the storage component can store the classification results of each classification dataset in storage.
[0028] Such an embodiment of the system can offer the advantage of making the results accessible for further analysis.
[0029] In one or more embodiments of the above system, an analysis component can analyze the results. Based on the analysis of the results, the analysis component can further generate classification scores corresponding to the quantum kernels contained in the multiple quantum kernels.
[0030] Such embodiments of the system can offer several advantages, including determining the accuracy of each execution of each quantum kernel for data generation and data classification tasks, and comparing their respective accuracies.
[0031] In one or more embodiments of the above system, a quantum kernel identification component may select a first quantum kernel based on a classification score, where the first quantum kernel may be made available for generating a new dataset based on the original dataset through balancing. The quantum kernel identification component may further select a second quantum kernel based on a classification score, where the second quantum kernel may be made available for classifying the data contained in the new dataset.
[0032] Such embodiments of the system can offer several advantages, including identifying the best quantum kernel for data generation from a set of quantum kernels, and identifying the best quantum kernel for data classification from a set of quantum kernels.
[0033] In one or more embodiments of the above system, the multiple quantum kernels can be quantum feature maps.
[0034] Such embodiments of the system can offer the advantage of generating novel datasets using quantum computing and performing QML-based data classification tasks.
[0035] One embodiment of a data generation component that can use the respective quantum kernels of multiple quantum kernels to transform the data distribution of an input dataset into a separate new data distribution, oversample data from each new data distribution (e.g., generate synthetic data) to generate each of the multiple new datasets, and generate each classification dataset based on each new dataset, can provide several advantages, including capturing the data distribution in an input dataset using quantum kernels and generating multiple new data distributions, generating multiple new datasets based on data distributions, and generating multiple classification datasets that can be used for QML-based data classification tasks and for training different machine learning models for QML-based data classification tasks, where the classification datasets may include balanced data of majority and minority classes.
[0036] One embodiment of a data generation component that uses the quantum kernels of multiple quantum kernels to transform the data distribution of an input dataset into a new data distribution, undersamples the data from each new data distribution (e.g., by clustering or otherwise reducing the data) to generate each of the multiple new datasets, and generates each classification dataset based on each new dataset, can provide several advantages, including capturing the data distribution in an input dataset using quantum kernels and generating multiple new data distributions, generating multiple new datasets based on the data distribution, and generating multiple classification datasets that can be used for QML-based data classification tasks and for training different machine learning models for QML-based data classification tasks, where the classification datasets may include balanced data of majority and minority classes.
[0037] Therefore, one or more embodiments can select quantum kernels from a set in the outer loop to be used for data generation for oversampling of minority classes or for data clustering for undersampling of majority classes. Then, in the inner loop, along with this quantum kernel from the outer loop, the embodiment can use the same set of quantum kernels to take each one from the set, classify them, and store the results. Once the quantum kernels in the inner loop are exhausted, the embodiment can return to the outer loop and select the next quantum kernel there for balancing, and this method can continue until the quantum kernels in the outer loop are exhausted. The embodiment can then select the best result and choose a pair of quantum kernels (e.g., outer for balancing, inner for classification). As a result, this embodiment provides a fully automated approach to a favorable method of quantum kernel balancing and classification.
[0038] In one embodiment, the system can be used for data classification tasks such as binary classification, multi-class classification, or other types of data classification tasks. For example, the system can be used to train a machine learning model on balanced data generated via a quantum kernel, and the trained machine learning model can be used in data classification tasks to detect anomalies related to financial transactions, data breaches, etc.
[0039] According to various embodiments, the system can be implemented as a method executed by a computer or as a computer program product.
[0040] Key applications of QML include kernel-based methods for data classification for multiple use cases, particularly in tasks such as anomaly detection (e.g., related to financial fraud, cybersecurity, etc.). Datasets used in such use cases can typically include a target class and a non-target class, where the target class typically contains significantly fewer data points than the non-target class. For example, in the category of financial transactions, the target class might contain data points related to a small number of undesirable or suspicious transactions, while the non-target class might contain data points related to a large number of normal or non-suspicious transactions. For instance, there might be no suspicious transactions among 100 transactions, but there might be one or two suspicious transactions among 10,000 transactions. Therefore, the target class can be referred to as the minority class, and the non-target class as the majority class. If a machine learning model is trained on a dataset containing such minority and majority classes, the model may exhibit bias towards the majority class, and the minority class may be completely ignored by the model when generating predictions. Therefore, before using such a dataset in a data classification task, it is desirable to balance the data for the minority and majority classes in the dataset so that the sizes of the two classes (i.e., the number of data points) are equal. This can be achieved either by oversampling the data points in the minority class or by undersampling the data points in the majority class.
[0041] Existing approaches typically use an oversampling technique called Synthetic Minority Oversampling Technique (SMOTE) to balance data across different classes within a dataset. The mechanism of SMOTE is described in more detail with reference to Figure 4. However, SMOTE is a classical technique that can be used for synthetic data generation / data clustering or data balancing, and using SMOTE to generate new data and then further using quantum kernels to perform data classification based on that new data can be a mismatch due to the fundamental differences between classical and quantum processing. For example, classically balancing data may negate the computational advantages such as superposition, interference, and entanglement provided by quantum computing, or the classical techniques used to balance data may not be well-suited to how data is processed in quantum computing. In other words, when using quantum methods to classify data contained in a balanced classification dataset (i.e., a dataset in which the number of data points in minority and majority classes are equal), it may be desirable to generate the classification dataset via quantum methods rather than classical methods (for example, by oversampling or undersampling existing data points).
[0042] Accordingly, embodiments described herein include systems, computer-executed methods, and computer program products that can use quantum kernel-based methods instead of classical techniques such as SMOTE for balancing data in datasets containing majority and minority classes. For example, in one or more embodiments, a set of quantum kernels can be used to generate synthetic data points by oversampling data, and another set of quantum kernels can be used to compare test results of data classification based on the synthetic data points. Alternatively, in one or more embodiments, a set of quantum kernels can be used to perform undersampling based on clustering based on K-means, K-medoid, or other techniques, and another set of quantum kernels can be used to compare test results of data classification based on the clustered data points. Furthermore, various embodiments herein can automatically determine the most appropriate quantum kernel to be used to generate new data by oversampling or undersampling, and then perform data classification based on the new data to produce the best test results. In this regard, the experimental results presented with reference to Figures 5 to 7 demonstrate that the quantum kernels selected by the methods and techniques of various embodiments can perform optimal data classification with superior classification accuracy compared to existing techniques.
[0043] In various embodiments, a data classification model can automatically select the most appropriate quantum kernel for oversampling or undersampling data, and can select the most appropriate quantum kernel for performing data classification. In various embodiments, a data classification component can use an access component to access an input dataset containing majority and minority class data. In various embodiments, a quantum kernel selection component can randomly select a quantum kernel from a set of quantum kernels, and a data generation component can use a quantum kernel to oversample or undersample the data contained in the input dataset. As a result, a new dataset can be generated containing synthetic data points obtained as a result of oversampling the data, or clustered data points obtained as a result of undersampling the data. The process of generating the new dataset can be performed by a data generation component on a quantum computer or a classical simulator of a quantum computer, where the quantum kernel can process the input dataset via quantum circuits. The data generation component can further generate a classification dataset based on the new dataset, where the classification dataset may include the input dataset and the new dataset. In various embodiments, the data generation component can generate the data contained in the classification dataset by sequentially applying quantum kernels from a set of quantum kernels. Thus, the data classification component can perform multiple consecutive data classification tasks, where different quantum kernels can be applied to the classification dataset during each data classification task. In various embodiments, once all quantum kernels applicable to the classification dataset have been exhausted, the quantum kernel selection component can randomly select a different quantum kernel from its set, and the data generation component can generate another new dataset using the selected quantum kernel.In various embodiments, the process can continue as described until all quantum kernels in the set of quantum kernels are used to generate new data and to perform data classification based on the new data. In various embodiments, a storage component can store the results of the classification task performed by the data classification component throughout the entire process, and an analysis component can analyze the results. The analysis component can analyze the results and generate classification scores for the quantum kernels included in the set of quantum kernels. In various embodiments, a quantum kernel identification component can, based on the analysis, identify a first kernel that can most accurately generate a new dataset based on the original dataset through balancing, and further identify a second kernel that can most accurately classify the data included in the new dataset. For example, QML classification can be performed on an input dataset by a method performed by a computer. Based on the quantum kernel selected for QML classification, a new dataset containing synthesized data or undersampled data can be generated, respectively, by oversampling or undersampling data from the input dataset via the first quantum kernel. Subsequently, QML classification can be performed on the input dataset and the new dataset via the second quantum kernel.
[0044] The embodiments shown in one or more drawings described herein are merely illustrative, and therefore the architecture of the embodiments is not limited to the systems, devices, and / or components shown therein, nor to any particular order, connection, and / or combination of the systems, devices, and / or components shown therein. For example, in one or more embodiments, non-limiting systems described herein, such as the non-limiting system 100 shown in Figure 1, and / or such systems may further include, be associated with, and / or be combined with, one or more computer and / or computing-based elements described herein, with reference to an operating environment such as the operating environment 1000 shown in Figure 10. For example, the non-limiting system 100 may be associated with the computing environment 1000, such as being accessible via the computing environment 1000 described below with reference to Figure 10, so that aspects of processing can be distributed between the non-limiting system 100 and the computing environment 1000. In one or more of the embodiments described, a computer and / or computing-based element can be used to perform one or more systems, devices, components and / or computer operations shown and / or described with respect to Figure 1 and / or other figures described herein.
[0045] Figure 1 shows a block diagram of an exemplary, non-limiting system 100, which can generate synthetic data and perform data classification tasks using a quantum kernel, according to one or more embodiments described herein.
[0046] The non-limiting system 100 and / or components of the non-limiting system 100 can be used to use hardware and / or software and can solve problems of a highly technical nature (e.g., related to QML, data classification, quantum kernels, etc.) that are not abstract and cannot be performed as a set of mental acts by humans. Furthermore, part of the process to be performed may be performed by a specialized computer for performing defined tasks related to data classification based on quantum kernels. The non-limiting system 100 and / or components of the system can be used to solve novel problems arising from advances in the technologies and / or similar technologies described above. The non-limiting system 100 can provide improvements to QML procedures by identifying quantum kernels that can most accurately generate new data by oversampling data in minority datasets or undersampling data in majority datasets, and by identifying quantum kernels that can most accurately classify data. In various embodiments, the methods and techniques disclosed herein can produce more accurate classification results compared to certain existing technologies.
[0047] As shown in Figure 1, the non-limiting system 100 may include a classical system 102 and a quantum system 112. The classical system 102 may be coupled to the quantum system 112 in operational, communicative, electrical and / or similar functions. The quantum system 112 may include at least one quantum processor, such as a quantum processor 114. The classical system 102 may include one or more components, such as a memory 106, a processor 104, a bus 108 and / or a data classification model 110. In one embodiment, the data classification model 110 may be at least partially included by the quantum system 112. The quantum processor 114 may include a quantum logic circuit including one or more qubits, e.g., qubit 114A, qubit 114B, ..., qubit 114n, etc., where n represents a positive integer. The quantum processor 114 can be any suitable processor. The quantum processor 114 may generate one or more instructions for controlling the quantum logic circuit.
[0048] The discussion will briefly move to the processor 104, memory 106, and bus 108 of the non-limiting system 100. For example, in one or more embodiments, the non-limiting system 100 may include a processor 104 (e.g., a computer processing unit, a microprocessor, a classical processor, and / or similar processor). In one or more embodiments, components associated with the non-limiting system 100 described herein with or without reference to one or more figures of one or more embodiments may include one or more computer and / or machine-readable writable and / or executable components and / or instructions that can be executed by the processor 104 to enable the execution of one or more operations defined by such components and / or instructions.
[0049] In one or more embodiments, the non-limiting system 100 may include computer-readable memory (e.g., memory 106) that can be operably connected to the processor 104. Memory 106 can store computer-executable instructions that, when executed by the processor 104, cause the processor 104 and / or one or more other components of the non-limiting system 100 (e.g., a data classification model 110, an access component 202, a quantum kernel selection component 204, a data generation component 206, a data classification component 208, a storage component 210, an analysis component 212, a quantum kernel identification component 214, and / or a training component 216) to perform one or more actions. In one or more embodiments, memory 106 can store computer-executable components (e.g., a data classification model 110, an access component 202, a quantum kernel selection component 204, a data generation component 206, a data classification component 208, a storage component 210, an analysis component 212, a quantum kernel identification component 214, and / or a training component 216).
[0050] The non-limiting system 100 and / or its components described herein can be coupled to one another via bus 108 in communicative, electrical, operational, optical and / or other ways. Bus 108 may include one or more of other types of buses that can use a memory bus, a memory controller, a peripheral bus, an external bus, a local bus and / or one or more bus architectures. One or more of these examples of bus 108 can be used. In one or more embodiments, the non-limiting system 100 can be coupled to one or more external systems (e.g., an electrical output generation system not shown, one or more output targets, an output target controller and / or similar), sources and / or devices (e.g., classical computing devices, communication devices and / or similar devices) via a network and the like (e.g., communicative, electrical, operational, optical and / or similar functionally). In one or more embodiments, one or more components of the non-limiting system 100 may reside in the cloud and / or locally within a local computing environment (e.g., at a specific location).
[0051] In various embodiments, the classical system 102 may include a data classification model 110. As shown in Figure 2, the data classification model 110 may further include an access component 202, a quantum kernel selection component 204, a data generation component 206, a data classification component 208, a storage component 210, an analysis component 212, a quantum kernel identification component 214, and a training component 216. The classical system 102 can be coupled to the quantum system 112 (operationally, communicatively, electrically and / or similarly functionally) and perform the operations described herein by various embodiments.
[0052] The input dataset 120 can contain data from different classes with unequal numbers of data points. For example, the input dataset 120 can be a dataset that can be used to train a machine learning model to perform data classification tasks such as binary classification, multi-class classification, or multi-label classification. For instance, the input dataset 120 can be used to train a machine learning model to detect anomalies related to financial transactions. Thus, the input dataset 120 could contain 100,000 data points belonging to a non-suspicious transaction class and 100 data points belonging to a suspicious transaction class. However, training a machine learning model for a data classification task using data from classes with unequal numbers of data points can lead to the machine learning model being biased towards the majority class if the number of data points in one class (majority class) is significantly greater than the number of data points in another class (minority class). Therefore, it may be desirable to balance the number of data points in different classes of data used to train a machine learning model. Furthermore, in the case of QML, which classifies data using quantum computing techniques, it may be even more desirable to balance the data using quantum computing techniques. Therefore, the various embodiments described herein can utilize the data classification model 110 to identify the best quantum kernel that can be used to balance the number of data points in different classes contained in the input dataset and to further classify the balanced data.
[0053] For example, in various embodiments, the access component 202 can access the input dataset 120. In various embodiments, the data generation component 206 can generate multiple new datasets based on the input dataset 120 using multiple quantum kernels. For example, in various embodiments, the data generation component 206 can transform the data distribution of the input dataset 120 into a new data distribution using each of the quantum kernels contained in a plurality of quantum kernels 122. The data generation component 206 can then generate each of the multiple new datasets by oversampling or undersampling data from each new data distribution. Finally, the data generation component 206 can generate a classification dataset based on each new dataset, where the classification dataset based on the new datasets may include the input dataset 120 and the new datasets. Each new dataset may include either composite data points, which can typically be obtained as a result of oversampling the data, or undersampled data points, which can typically be obtained as a result of clustering the data points. The quantum kernels used by the data generation component 206 to generate the new data distribution and then generate the composite datasets can be randomly selected from a plurality of quantum kernels 122 by the quantum kernel selection component 204. In various embodiments, the data classification component 208 can classify each classification dataset generated by the data generation component 206 using a plurality of quantum kernels 122. In various embodiments, the storage component 210 can store the classification results of each classification dataset in storage.
[0054] More specifically, in various embodiments, the quantum kernel selection component 204 can randomly select a quantum kernel from a plurality of quantum kernels 122. Each quantum kernel included in the plurality of quantum kernels 122 can be defined as a feature map or a unitary that performs a data point-dependent unitary transformation on n qubits. For example, the quantum kernels included in the plurality of quantum kernels 122 can be a ZZ feature map that can be constructed via a Pauli expansion circuit. In quantum computing, quantum kernels can be used to encode classical data having a set of classical features and to obtain a different set of classical features based on the encoded classical data. Thus, in various embodiments described herein, a quantum kernel can be used to transfer one set of classical features to another. The resulting form of classical features can be made more separable via classical techniques. For example, the data generation component 206 can use a quantum kernel to transform the data distribution of an input dataset 120 into a new data distribution. For example, the data generation component 206 can apply a quantum kernel to the input dataset 120 by running a quantum circuit on a quantum system 112 or a classical simulator of a quantum system. As a result, the classical data distribution of the input dataset 120 can be transformed into a novel classical data distribution based on a quantum kernel.
[0055] In various embodiments, the data generation component 206 can generate new data using a new data distribution by oversampling or undersampling data points organized according to the new data distribution. Oversampling refers to increasing the number of data points in minority classes by generating copies of existing data points (e.g., synthetic data) to make the number of data points in minority classes equivalent to those in majority classes. Undersampling refers to decreasing the number of data points in majority classes using a clustering-based approach to make the number of data points in majority classes equivalent to those in minority classes. The data generation component 206 can decide whether to oversample or undersample data points based on a specific use for the input dataset 120, and can automatically oversample or undersample the data according to that decision. The process of generating the new dataset can be performed on a classical system 102. Based on the new dataset, the data generation component 206 can further generate classification datasets by combining the input dataset 120 and the thus generated new dataset. The classification datasets may contain an equivalent number of data points in majority and minority classes.
[0056] In various embodiments, the data classification component 208 can classify data contained in a classification dataset using multiple quantum kernels contained in multiple quantum kernels 122. In some embodiments, the data classification component 208 can use all quantum kernels contained in multiple quantum kernels 122 to classify data in a classification dataset. In one or more embodiments, the data classification component 208 can sequentially apply quantum kernels to a classification dataset. For example, the multiple quantum kernels 122 may include feature maps such as Z, ZZ, XY, and XZ, where X, Y, and Z represent Pauli-X gates, Pauli-Y gates, and Pauli-Z gates, respectively. The data classification component 208 can perform multiple data classification tasks on a classification dataset by using different feature maps in each data classification task. For example, the data classification component 208 may use the XY feature map during the first data classification and the ZZ feature map during the second data classification, and so on. The data classification task can be performed on a classical system 102, for example, on a classical simulator of a quantum system, where the data classification component 208 can apply quantum kernels to the classification dataset by implementing each quantum kernel as its respective quantum circuit.
[0057] Data classification using quantum kernels:
[0058] In one or more embodiments, classical data can be encoded into quantum states via quantum feature maps for classifying data using quantum kernels. Quantum states can be obtained by applying the feature maps to the classical data via quantum circuits that can be executed on a quantum computer or a classical simulator of a quantum computer. Quantum kernel values can then be computed for pairs of data points, and a kernel matrix can be computed based on the quantum kernel computation. The kernel matrix can be used to train a machine learning model (e.g., data classification component 208), such as an SVM model or another machine learning model. The trained machine learning model can be deployed to classify novel data, such as data contained in a classification dataset, where the trained machine learning model can encode classical data points, such as data contained in a classification dataset, by using quantum feature maps. Furthermore, the trained machine learning model can compute kernel values for classical data points, compute a kernel matrix based on the kernel values, and perform data classification using the kernel matrix.
[0059] The data classification process performed by the data classification component 208 can continue until the list of quantum kernels applicable to the classification is exhausted. Thereafter, the quantum kernel selection component 204 can select another quantum kernel from the multiple quantum kernels 122, and the data generation component 206 can automatically generate a new dataset based on the quantum kernel. The data contained in the new dataset can be further classified by the data classification component 208 using multiple quantum kernels from the multiple quantum kernels 122. The cycle of generating classified data based on the input dataset 120 and classifying the classified data can continue until each quantum kernel from the multiple quantum kernels 122 is used to generate and classify a classified dataset. In various embodiments, the storage component 210 can store the results of various classifications in storage (e.g., memory 106 or another form of storage).
[0060] In various embodiments, the analysis component 212 can analyze the results. The analysis component 120 can analyze the results using various classical metrics or the average of a set of metrics. For example, for binary classification, the analysis component 120 can use metrics such as Area Under the Curve (AUC), precision, or F1 score, or the average of such metrics. In one or more embodiments, the results of the analysis can be displayed on the user interface (UI) of a device (e.g., a desktop computer, laptop, tablet, smartphone, etc.) for entities (e.g., hardware, software, machine, artificial intelligence (AI), neural network, and / or user).
[0061] Based on the analysis, the analysis component 120 can generate classification scores corresponding to the quantum kernels contained in the multiple quantum kernels 122. For example, in some embodiments, the analysis component 120 can generate a classification score corresponding to each pair of quantum kernels, where the first quantum kernel in the pair of quantum kernels may be the quantum kernel used to generate a classification dataset, and the second quantum kernel in that pair may be a different quantum kernel used to classify the data within the classification dataset. For example, the classification scores corresponding to a pair of quantum kernels may be the result of a metric or a set of metrics (e.g., AUC, F1 score, etc.) that can indicate the data classification performance of the data classification component 208 based on the pair of quantum kernels. In other embodiments, the analysis component 120 can generate a first score and a second score for each quantum kernel. Each first score may correspond to data generation, and each second score may correspond to data classification. For example, the first score for a quantum kernel may be the result of a metric or a set of metrics (e.g., AUC, F1 score, etc.) that can indicate the data generation performance of the data generation component 206 based on the quantum kernels. Similarly, the second score for the quantum kernel can be the result of a metric or multiple metrics (e.g., AUC, F1 score, RMS, etc.) that can demonstrate the data classification performance of the quantum kernel-based data classification component 208.
[0062] In various embodiments, the quantum kernel identification component 214 can select a first quantum kernel and a second quantum kernel from a plurality of quantum kernels 122 based on their respective classification scores, where the first quantum kernel can be made available to generate a new dataset based on the original dataset through balancing, and the second quantum kernel can be made available to classify the data contained in the new dataset. For example, in some embodiments, the quantum kernel identification component 214 can select a pair of quantum kernels whose classification score is greater than a defined precision threshold. In other embodiments, the quantum kernel identification component 214 can select a quantum kernel having a first score greater than a first defined precision threshold, and another quantum kernel having a second score greater than a second defined precision threshold. The pair of quantum kernels thus selected by the quantum kernel identification component 214 can be used by the data generation component 206 to generate a new dataset, and further used by the data classification component 208 to classify the data in the new dataset. For example, the input dataset 120 may include data related to financial transactions. When the quantum kernel identification component 214 selects a pair of quantum kernels, the data classification model 110 can perform data classification on a new input dataset using the pair of quantum kernels, where the new input dataset may contain different data related to financial transactions.
[0063] Various embodiments described herein can represent a QML-based test process or quantum kernel test procedure in which different feature engineering methods are tested to identify the best quantum kernel for oversampling or undersampling data and performing data classification. For example, various embodiments described herein can be used to identify a quantum kernel that can most accurately map a classical data distribution to a novel classical data distribution and to generate a classification dataset based on the novel classical data distribution. Various embodiments described herein can further be used to identify another quantum kernel that can most accurately classify the data contained in the classification dataset. Various embodiments described herein can also represent a QML-based test process for machine learning models. For example, in one or more embodiments, the data classification component 208 may be a machine learning model that can be trained to use a quantum kernel to classify data. Since the data is classified by the data classification component 208 via different quantum kernels, the performance of the data classification component 208 can also be evaluated via classification scores generated by the analysis component 120. In one or more embodiments, the training component 216 can train one or more machine learning models for various data classification tasks using the classification dataset generated by the data generation component 206.
[0064] Figure 2 shows a block diagram of an exemplary, non-limiting system 200 capable of generating novel data and performing data classification tasks using a quantum kernel, according to one or more embodiments described herein. Descriptions of similar elements and / or repetitions of processes used in each embodiment are omitted for brevity.
[0065] Figure 2 shows the system of the data classification model 110 in Figure 1. As previously mentioned, the data classification model 110 may include a further access component 202, a quantum kernel selection component 204, a data generation component 206, a data classification component 208, a storage component 210, an analysis component 212, a quantum kernel identification component 214, and a training component 216. In one or more embodiments, one or more components included in the data classification model 110 may be machine learning models. In some embodiments, the data classification model 110 may be a multi-stage machine learning model, where each machine learning model can perform the operations described with reference to Figure 1 to identify the pair of quantum kernels that can most accurately classify the input dataset, where the results of one machine learning model can be used by another machine learning model to identify the pair of quantum kernels. The pair of quantum kernels can be selected from a plurality of quantum kernels 122 and tested based on the input dataset 120.
[0066] Figure 3 shows a flow diagram of an exemplary, non-limiting method 300, which, according to one or more embodiments described herein, can generate novel data and perform data classification tasks using a quantum kernel. For brevity, descriptions of similar elements and / or process repetitions used in each embodiment are omitted.
[0067] The non-restrictive method 300 summarizes an automated quantum process / program flow involving quantum kernel-based data balancing and classification, as described with reference to Figures 1 and 2.
[0068] In 302, a non-restrictive method 300 may include selecting a quantum kernel (e.g., a feature map such as Z, ZZ, XY, XZ, or another quantum kernel) from a plurality of quantum kernels (e.g., by a quantum kernel selection component 204) and generating new data (e.g., synthetic data) (e.g., by oversampling or undersampling).
[0069] In 304, the non-limiting method 300 may further include generating new data based on a selected quantum kernel (e.g., by the data generation component 206) and generating a classification dataset based on the new data, where the classification dataset may include balanced data (i.e., an equal number of data points in minority and majority classes).
[0070] In 306, the non-restrictive method 300 may include classifying the data contained in the classification dataset through multiple quantum kernels (e.g., by the data classification component 208) that are contained within multiple quantum kernels until all quantum kernels applicable to the classification dataset have been applied.
[0071] For example, the non-restrictive method 300 may include performing multiple data classification tasks on data contained in a classification dataset, using different quantum kernels during each data classification task (e.g., by the data classification component 208). For example, the data classification component 208 may selectively use a Z-feature map during the first data classification and a ZZ-feature map during the second data classification, and so on. In the non-restrictive method 300, the process of successively applying different quantum kernels to classify the data contained in the classification dataset is shown as a nested inner loop in 308. In this view, a rectangle with a dashed line may represent a portion of the non-restrictive method 300 where automatic data generation (e.g., via oversampling or undersampling) and classification testing can occur.
[0072] Once the list of quantum kernels used to classify the data is exhausted, in which case, at 310, the non-restrictive method 300 can exit the nested inner loop and return to 302, where the non-restrictive method 300 may include selecting a different quantum kernel (e.g., a feature map such as Z, ZZ, XY, XZ, or another quantum kernel) from a group of quantum kernels (e.g., by the quantum kernel selection component 204) to generate a new dataset (e.g., by oversampling or undersampling). This step represents the outer loop of the non-restrictive method 300.
[0073] In 312, the non-restrictive method 300 may include storing the results of a data classification task (e.g., by a memory component 210) and selecting the best quantum kernel for new data generation and data classification (e.g., by a quantum kernel identification component 214). For example, based on the results of data classification, the quantum kernel identification component 214 may select a pair of quantum kernels that can produce the most accurate data classification (in a manner similar to a method performing a grid search), where the first quantum kernel of the pair of quantum kernels may correspond to the outer loop process of the non-restrictive method 300, and the second quantum kernel of the pair of quantum kernels may correspond to the nested inner loop process of the non-restrictive method 300.
[0074] Figures 4 to 7 are intended to illustrate the advantages of the embodiments of this disclosure over existing data classification techniques.
[0075] Figure 4 shows illustrative, non-limiting processes 400 and 410 that describe oversampling techniques. Descriptions of similar elements and / or repetitions of processes used in each embodiment are omitted for brevity.
[0076] Non-restrictive processes 400 and 410 describe the SMOTE approach. SMOTE is a classical oversampling technique in which new synthetic samples (observations or data points) are created using existing samples of data from a minority class. For example, a synthetic (virtual) training record is generated by linear interpolation of samples included in the minority class. The synthetic training record is generated by randomly selecting one or more of the k nearest neighbors for each example in the minority class. After the oversampling process, the data is reconstructed, and multiple classification models can be applied to the processed data. SMOTE includes the following steps:
[0077] Step 1: Set up a minority class set A. For each data point x in A($x\in A), the k nearest neighbors of x are obtained by calculating the Euclidean distance between x and all other samples in the minority class set A.
[0078] Step 2: Set the sampling rate N according to the proportion of imbalanced data. For each x in A($x\in A), N examples (i.e., x1, x2, ..., xn) are randomly selected from its k nearest neighbors and used to construct the set A1 ($A_1$).
[0079] Step 3: For each example x_k in A_1 ($x_k \in A_1$), where k = 1, 2, 3, ..., N, a new example is generated using the formula x' = x + rand(0,1) * (mid(x - x_k)) (or $x' = x + rand(0,1) * \mid x - x_k \mid$), where rand(0,1) represents a random number between 0 and 1.
[0080] If the value of the k nearest neighbor is 2, then each data point finds its 2 nearest neighbor (e.g., using Euclidean distance). For example, if we first consider data point a, then data points b and c can be a's 2 nearest neighbors. By using step 3 of the algorithm, new composite points can be generated. One or more composite samples can be generated on each line, depending on the desired number of composite samples. A single line can correspond to multiple composite points. In SMOTE, all data points of minority class A are considered, and composite observations are similarly generated for each data point.
[0081] SMOTE generates an average of different data points by considering the differences between different data points of the minority class data. For example, as described herein, a new data point can be generated by averaging two existing data points in the minority class. In an unrestricted process 400, graph 402 shows the initial distribution of data points, including the majority class data (smaller points) and the minority class data (larger points enclosed within the rectangle). In SMOTE, the minority sample from graph 402 can be considered for oversampling, as shown by graph 404. Moving to an unrestricted process 410, graph 412 shows the data points a, b, and c of the minority sample shown in graph 404, where b and c can represent the two k nearest neighbors of a. Graph 414 shows the composite data points m and n that can be generated via SMOTE between a and b and c, which are the two k nearest neighbors of a, respectively. SMOTE can generate composite data points through oversampling, but the pairs of points used to generate these composite data points may be randomly selected, making it not the most accurate technique for oversampling data.
[0082] In some other existing techniques, data points within the majority class of data can be clustered to generate different chunks or clusters of data points, and new data points can be generated using each cluster of data points, where the new data points based on clusters can be the midpoints of clusters generated by techniques such as K-means. For example, to undersample data points in the initial distribution, such existing techniques can replace clusters with their midpoints or centroids. Clustering can generate a better data distribution or better data than randomly selecting and averaging data points, as in SMOTE, because clustering considers the original data distribution of the data, rather than randomly selecting and averaging data points. By considering how the data is distributed or organized in the initial distribution, a better data distribution can be generated for new data points generated by oversampling or undersampling.
[0083] Generally, oversampling and undersampling can be performed by randomly selecting or removing data points. However, such techniques do not take into account the structure and distribution of the data. Embodiments of the present disclosure can eliminate the randomness inherent in SMOTE and / or other techniques and use a more strategic and efficient approach to generating new data. In addition, embodiments of the present disclosure can use quantum computing to generate new data. For example, the various methods and techniques described herein can take into account the distribution of data in the initial distribution of data points by using quantum kernels to generate new data by oversampling or undersampling. For example, each quantum kernel (e.g., a quantum feature map or another quantum kernel) can generate a different data distribution based on the data distribution of the input dataset 120, and each different data distribution can result in a new dataset. Thus, each new dataset thus generated can contain data with its own unique pattern, and different results can be produced by using new datasets generated via different kernels in a data classification task. In some embodiments, different new datasets with different data patterns can be used to train different machine learning models that can be used for data classification (e.g., by a training component 216). By using quantum kernels to generate synthetic data based on embedded patterns in the initial distribution of data and to perform data classification, various embodiments of this specification can provide more efficient methods and techniques for QML. Theoretically, classical functions can be described as a subset of quantum functions. Quantum functions can generate any function that may be more efficient than classical functions, considering the nature of entanglement in quantum space. Thus, in experiments performed with synthetic datasets, quantum kernels have been observed to perform better than classical kernels in terms of accuracy.
[0084] Figure 5 shows an illustrative, non-limiting Table 500 illustrating experimental results from one or more embodiments described herein. Descriptions of similar elements and / or process repetitions used in each embodiment are omitted for brevity.
[0085] Referring to Figures 1-4, the non-restrictive Table 500 shows a comparison of experimental results generated by clustering ad-hoc data using different techniques. In the non-restrictive Table 500, column 502 shows results based on the Euclidean distance technique, column 504 shows results based on a Radial Basis Function (RBF) kernel, and column 506 shows results based on a ZZ feature map generated via a Pauli extension circuit (an embodiment of the disclosure). The Euclidean distance technique and the RBF kernel technique are classical techniques. To generate the experimental results shown in the non-restrictive Graph 500, each selected technique (i.e., Euclidean distance, RBF kernel, and quantum kernel) was used to transform the input data from an existing distribution to a novel distribution, and the K-medoid technique was used to perform data clustering based on the novel distribution. In each scenario, data transformation and clustering were performed by a machine learning model trained on ground truth data. The K-medoid technique is also a classical technique that uses the K-medoid clustering algorithm, which considers actual data points as the centers (i.e., medoids) of data point clusters. Calculations for the quantum K-medoid technique (i.e., K-medoids using Pauli extension circuits) were performed using "affinity" as the distance metric.
[0086] Each technique (i.e., Euclidean distance, RBF kernel, and ZZ feature map generated via Pauli extension circuits) was executed on a classical simulator of a quantum computer. A non-restrictive Table 500 shows a comparison of the training, validation, and test results corresponding to each technique. 42 random seed values and a random state value of 0 were defined for the experiment. The random seed and random state values can be values that can be provided to the random number generator to select different parts of a larger dataset for the experiment. Different random seed and random state values can result in different distributions of data selected from a larger data set. This is because, in quantum computing, a large dataset (e.g., 100,000 data points) is a massive amount of data, particularly a massive amount of data that a quantum computing system (e.g., a quantum computer or a classical simulator of a quantum computer) processes while that data is also being processed elsewhere, and therefore, a portion of a large dataset is typically selected for the experiment. The random seed and random state values can also be used to replicate the experiment.
[0087] To evaluate different clustering techniques, mutual information (MI), random index, and precision metrics were used. According to the MI metric, a higher score indicates better performance. In the unrestricted Table 500, training, validation, and test results are organized as MI / random index / precision value. Clearly, the test scores demonstrate that using the quantum kernels described in the various embodiments described herein can produce better clustering than classical techniques. For example, the random index value of 0.50 for the test scores in columns 502 and 504 indicates that classical techniques (i.e., Euclidean distance techniques and RBF kernel techniques) can produce at best random guesses (50%), while the K-medoid using the Pauli extension circuit can produce a random index value of 0.61 for the test score, as seen in column 506. Similarly, the test scores show a precision value of 73% for the K-medoid using the Pauli extension circuit technique, and 52% and 56% for the Euclidean distance technique and RBF kernel technique, respectively.
[0088] Therefore, quantum kernel-based techniques behave better than classical techniques. Specifically, the K-medoid using the Pauli extension circuit produced more accurate results than the K-medoid using Euclidean distance and the K-medoid using the RBF kernel.
[0089] In QML, quantum computation performed on a quantum computer is typically followed by the use of classical classifiers. For example, feature maps can be used to transform data from one classical distribution to another, and then the data can be classified using a Support Vector Machine (SVM), Random Forest, XGBoost, or another machine learning classifier. While new data can be generated using classical techniques through oversampling or undersampling, if subsequent data classification is performed by a quantum classifier, it is desirable to generate the new data using quantum computing techniques to avoid mismatches in data processing techniques stemming from the fundamental differences between classical and quantum processing.
[0090] Figure 6 shows illustrative, non-limiting graphs 600, 610, 620, and 630 illustrating experimental data clustering results according to one or more embodiments described herein. Descriptions of similar elements and / or repetitions of processes used in each embodiment are omitted for brevity.
[0091] The experimental results shown in Figure 6 correspond to the clustering techniques and experimental results described with reference to the unrestricted Table 500 in Figure 5. The experimental results shown in Figure 6 were generated for ad hoc data (Δ=0.5) with 625 data points (N=625). Regarding the clustering techniques, the data distribution can be determined by the clustering accuracy. Unrestricted graph 600 shows the ground truth data, and unrestricted graphs 610, 620, and 630 show the respective clustering datasets generated using the K-medoid with Euclidean distance, the K-medoid with RBF kernel, and the K-medoid with Pauli extension circuit based on the ground truth data. As can be seen from unrestricted graphs 610 and 620, by applying the K-medoid technique to a data distribution generated using the Euclidean distance technique or the RBF kernel technique, it is possible to generate clusterings in which data points from different classes in the ground truth data can be divided in half. Conversely, applying the K-medoid technique to a data distribution generated using a Pauli extension circuit-based ZZ feature map can produce different and more balanced data clusterings with a mixture of data points, contrary to the ground truth data being split into two halves. As can be seen in Figure 7, the clustering generated by the K-medoid using the Pauli extension circuit can result in better data classification.
[0092] Figure 7 shows illustrative, non-limiting graphs 700, 710, and 720 illustrating experimental training and test results according to one or more embodiments described herein. Descriptions of similar elements and / or repetitions of processes used in each embodiment are omitted for brevity.
[0093] Continuing to refer to Figures 5 and 6, the unrestricted graphs 700, 710, and 720 correspond to the K-medoid using Euclidean distance, the K-medoid using the RBF kernel, and the K-medoid using the Pauli extension circuit technique, respectively. In each graph in Figure 7, the Y-axis corresponds to the mutual information value, and the X-axis corresponds to the size of the considered dataset. To generate the unrestricted graphs 700, 710, and 720, each machine learning model was trained on the ground truth data shown by unrestricted graph 600, and the trained machine learning models were used to generate the shown training and test curves. The learning curves shown in Figure 7 describe transformations based on different metrics for the number of data points. Clearly, in each scenario, the calculations converge at approximately 800 for each technique for the experimentally used dataset, despite the large error bars.
[0094] Figure 8 shows a flow diagram 800 of an exemplary, non-limiting method for generating novel data using a quantum kernel, according to one or more embodiments described herein. For brevity, descriptions of similar elements and / or process repetitions used in each embodiment are omitted.
[0095] In 802, a non-limiting method 800 may include accessing the input dataset (for example, by an access component 202) by a system operationally coupled to the processor.
[0096] In 804, the non-restrictive method 800 may include the system generating multiple novel datasets (for example, by the data generation component 206) based on an input dataset by using multiple quantum kernels.
[0097] Figure 9 shows a flow diagram of an exemplary, non-limiting method that can be used to select a pair of quantum kernels for performing data generation and classification according to one or more embodiments described herein. Descriptions of similar elements and / or repetitions of processes used in each embodiment are omitted for brevity.
[0098] Steps 902–906 of the non-restrictive method 900 extend step 804 of the non-restrictive method 800, where a system operationally coupled to a processor can generate over- or under-balanced versions of the input dataset using multiple quantum kernels based on the input dataset. The multiple novel datasets may include synthetic data generated from the input dataset.
[0099] In 902, the non-restrictive method 900 may include the system transforming the data distribution of an input dataset into a new data distribution (for example, by the data generation component 206) by using each of the quantum kernels of a plurality of quantum kernels.
[0100] In 904, the non-limiting method 900 may include the system generating each new dataset of multiple new datasets (for example, by the data generation component 206) by oversampling or undersampling data from each new data distribution.
[0101] In 906, the non-limiting method 900 may include the system generating each classification dataset based on each new dataset (for example, by the data generation component 206), where the classification dataset based on the new dataset includes the input dataset and the new dataset.
[0102] In 908, a non-restrictive method 900 may include the system classifying each classification dataset (for example, by a data classification component 208) by using multiple quantum kernels.
[0103] In 910, the non-limiting method 900 may include the system storing the classification results of each classification dataset in storage (for example, by the storage component 210).
[0104] In 912, the non-restrictive method 900 may include the system analyzing the classification results of each classification dataset (for example, by the analysis component 212).
[0105] In 914, the non-restrictive method 900 may include the system determining (e.g., by the quantum kernel discrimination component 214) whether a pair of quantum kernels has a classification score exceeding a defined precision threshold. If so, in 916, the non-restrictive method 900 may include the system selecting (e.g., by the quantum kernel discrimination component 214) a pair of quantum kernels for performing data classification on a new input dataset. Otherwise, in 918, the non-restrictive method 900 may include the system selecting (e.g., by the quantum kernel discrimination component 214) a different pair of quantum kernels for performing data classification on a new input dataset (i.e., having a classification score exceeding a defined precision threshold).
[0106] For the sake of brevity, the computer implementations and non-computer-based methodologies provided herein are presented and / or described as a series of actions. It should be understood that the innovations of the subject matter are not limited by the actions and / or the order in which they are presented. For example, actions may occur in one or more sequences, and / or simultaneously with and together with other actions not presented and described herein. Furthermore, not all presented actions can be used to implement the computer-based methodologies and non-computer-based methods of the described subject matter. Additionally, the computer-based methodologies described below and throughout this specification may be stored in a product, thereby enabling the transport and transfer of computer-based methods to a computer. Where used herein, the term "product" is intended to encompass computer programs accessible from any computer-readable device or storage medium.
[0107] Systems and / or devices are described herein (and / or further described herein) with respect to interactions between one or more components. Such systems and / or components may include one or more of those components or subcomponents, specified components and / or subcomponents, and / or additional components. Subcomponents may be implemented not within a parent component, but as components communicatively coupled to other components. One or more components and / or subcomponents may be combined into a single component that provides an aggregation function. Components may interact with one or more other components that are not specifically described herein for brevity but are known to those skilled in the art.
[0108] Figure 10 shows a block diagram of an exemplary, non-limiting operating environment that can facilitate one or more embodiments described herein. Figure 10 and the following discussion are intended to provide a general description of a preferred operating environment 1000 in which one or more embodiments described herein can be carried out in Figures 1-9.
[0109] Various aspects of this disclosure are described by explanatory text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of computer program products (CPPs). With respect to any flowchart, operations may be performed in a different order than those shown in a given flowchart, depending on the technology involved. For example, also depending on the technology involved, two operations shown in consecutive blocks of a flowchart may be performed in reverse order, as a single integrated step, simultaneously, or with at least partial time overlap.
[0110] Embodiments of a computer program product ("CPP Embodiment" or "CPP") are terms used in this disclosure to describe any set of one or more storage media (also called "mediums") that collectively comprise a set of one or more storage devices that collectively comprise machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. Computer-readable storage media may, but are not limited to, electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any suitable combination thereof. Some known types of storage devices, including these media, include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital purpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices (such as pits / lands formed on the main surface of punch cards or disks), or any suitable combination of those described above. When the term "computer-readable storage medium" is used in this disclosure, it shall not be construed as storage in the form of transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses passing through optical fiber cables, electrical signals communicated through wires, and / or other transmission media. As will be understood by those skilled in the art, data is moved at several intermittent points during the normal operation of the storage device, such as during access, defragmentation, or garbage collection; however, data is not transient while it is stored, so the foregoing does not make the storage device transient.
[0111] The computing environment 1000 includes an example of an environment for executing at least a portion of computer code involved in performing an inventive method, such as a quantum kernel-based data classification code 1026. In addition to block 1026, the computing environment 1000 includes, for example, a computer 1001, a wide area network (WAN) 1002, an end-user device (EUD) 1003, a remote server 1004, a public cloud 1005, and a private cloud 1006. In this embodiment, the computer 1001 includes a processor set 1010 (including processing circuits 1020 and a cache 1021), a communication fabric 1011, volatile memory 1012, persistent storage 1013 (including the operating system 1022 and block 1026 identified above), a peripheral device set 1014 (including a user interface (UI), a device set 1023, storage 1024, and an Internet of Things (IoT) sensor set 1025), and a network module 1015. The remote server 1004 includes a remote database 1030. Public cloud 1005 includes gateway 1040, cloud orchestration module 1041, host physical machine set 1042, virtual machine set 1043, and container set 1044.
[0112] Computer 1001 may be a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer or other form of computer or mobile device, currently known or to be developed in the future, capable of executing programs, accessing networks or querying databases such as remote database 1030. As is well understood in the field of computer technology, and depending on the technology, the implementation of methods performed by computers may be distributed among multiple computers and / or multiple locations. On the other hand, in presenting this computing environment 1000, the detailed discussion will focus on a single computer, specifically computer 1001, in order to keep the presentation as simple as possible. Computer 1001 may be located in the cloud, although it is not shown in the cloud in Figure 10. On the other hand, computer 1001 is not required to be in the cloud, except to the extent that it may be shown positively.
[0113] The processor set 1010 includes one or more computer processors of any type currently known or to be developed in the future. The processing circuitry 1020 may be distributed across multiple packages, for example, multiple collaborative integrated circuit chips. The processing circuitry 1020 may implement multiple processor threads and / or multiple processor cores. The cache 1021 is memory located in the processor chip package and is typically used for data or code that should be available for high-speed access by threads or cores running on the processor set 1010. The cache memory is usually organized into multiple levels depending on its relative proximity to the processing circuitry. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, the processor set 1010 may operate using qubits and be designed to perform quantum computing.
[0114] Computer-readable program instructions are typically loaded into computer 1001, causing a series of operational steps to be executed by the processor set 1010 of computer 1001, thereby realizing a method to be executed by the computer, and the instructions thus executed instantiate the methods (collectively referred to as “inventive methods”) specified in the flowcharts and / or descriptions of computer-executed methods contained herein. These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 1021 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 1010 to control and direct the execution of the methods of the present invention. In computing environment 1000, at least some of the instructions for executing the inventive methods may be stored in block 1026 in persistent storage 1013.
[0115] The communication fabric 1011 is a signal conduction path that enables various components of the computer 1001 to communicate with one another. Typically, this fabric is made up of switches and conductive paths, such as buses, bridges, physical input / output ports, and similar components. Other types of signal communication paths, such as optical fiber communication paths and / or wireless communication paths, may be used.
[0116] The volatile memory 1012 may be any type of volatile memory currently known or to be developed in the future. Examples include dynamic type random access memory (RAM) or static type RAM. While volatile memory typically features random access, this is not required unless explicitly stated. In computer 1001, the volatile memory 1012 is located in a single package and is internal to computer 1001, or / or additionally, the volatile memory may be distributed across multiple packages and / or located externally to computer 1001.
[0117] Persistent storage 1013 is any form of non-volatile storage for a computer, currently known or to be developed in the future. Non-volatility of this storage means that the stored data is retained regardless of whether the computer 1001 and / or the persistent storage 1013 are directly powered. Persistent storage 1013 may be read-only memory (ROM), but typically, at least a portion of the persistent storage allows data to be written, deleted, and rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 1022 may take multiple forms, such as various known proprietary operating systems using a kernel or open-source portable operating system interface type operating systems. The code contained in block 1026 typically includes at least some computer code involved in the execution of the inventive method.
[0118] The peripheral device set 1014 includes a set of peripheral devices for the computer 1001. Data communication connections between the peripheral devices and other components of the computer 1001 may be implemented in various ways, such as Bluetooth® connections, near-field communication (NFC) connections, connections made by cables (e.g., Universal Serial Bus (USB) type cables), insert-type connections (e.g., Secure Digital (SD) cards), connections made by local area communication networks, and connections made via wide area networks such as the Internet. In various embodiments, the UI device set 1023 may include components such as display screens, speakers, microphones, wearable devices (e.g., goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 1024 is external storage such as an external hard drive, or insertable storage such as an SD card. Storage 1024 may be persistent and / or volatile. In some embodiments, storage 1024 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 1001 requires a large amount of storage (for example, when computer 1001 locally stores and manages a large database), this storage may be provided by peripheral storage devices designed to store very large amounts of data, such as a Storage Area Network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 1025 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another may be a motion detector.
[0119] The network module 1015 is a collection of computer software, hardware, and firmware that enables computer 1001 to communicate with other computers through the WAN 1002. The network module 1015 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depackaging data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of the network module 1015 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control and forwarding functions of the network module 1015 are performed on physically separate devices, so that the control function manages multiple different network hardware devices. Computer-readable program instructions for performing the inventive method can typically be downloaded from an external computer or external storage device to computer 1001 via a network adapter card or network interface included in the network module 1015.
[0120] WAN1002 is any wide area network (e.g., the Internet) that can transmit computer data over non-local distances using any currently known or future-developed technology for transmitting computer data. In some embodiments, the WAN may be replaced and / or complemented by a local area network (LAN), such as a Wi-Fi® network, designed to transmit data between devices located in a local area. The WAN and / or LAN typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0121] The end-user device (EUD) 1003 is any computer system used and controlled by an end-user (e.g., an enterprise customer operating computer 1001) and may take any of the forms discussed above with respect to computer 1001. Typically, EUD 1003 receives useful and valuable data from the operation of computer 1001. For example, in a hypothetical case where computer 1001 is designed to provide recommendations to an end-user, these recommendations are typically communicated from computer 1001's network module 1015 to EUD 1003 via WAN 1002. In this way, EUD 1003 can display or otherwise present the recommendations to the end-user. In some embodiments, EUD 1003 may be a client device such as a thin client, heavy client, mainframe computer, or desktop computer.
[0122] The remote server 1004 is any computer system that provides at least some data and / or functionality to computer 1001. The remote server 1004 may be controlled and used by the same entity that operates computer 1001. The remote server 1004 represents a machine that collects and stores useful and effective data for use by other computers, such as computer 1001. For example, in a hypothetical case, if computer 1001 is designed and programmed to provide recommendations based on historical data, then this historical data may be provided to computer 1001 from the remote database 1030 of the remote server 1004.
[0123] The public cloud 1005 is any computer system available for use by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing capabilities, without direct active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct and active management of the computing resources of the public cloud 1005 is performed by the computer hardware and / or software of the cloud orchestration module 1041. The computing resources provided by the public cloud 1005 are typically implemented by virtual computing environments running on various computers that make up the host physical machine set 1042, which is a universe of physical computers within and / or available to the public cloud 1005. Virtual computing environments (VCEs) typically take the form of virtual machines in the virtual machine set 1043 and / or containers in the container set 1044. These VCEs may be stored as images and are understood to be transferable either as images or after VCE instantiation, among and between various physical machine hosts. The cloud orchestration module 1041 manages the transfer and storage of images, deploys new VCE instantiations, and manages the active instantiation of VCE deployments. The gateway 1040 is a collection of computer software, hardware, and firmware that enables the public cloud 1005 to communicate over the WAN 1002.
[0124] Here, some further explanation of virtualized computing environments (VCEs) is provided. A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from an image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature in which the kernel allows for the existence of multiple isolated user-space instances called containers. These isolated user-space instances typically behave as actual computers in terms of the programs running in them. Computer programs running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and the devices allocated to that container; this feature is known as containerization.
[0125] Private Cloud 1006 is similar to Public Cloud 1005, except that its computing resources are available for use by a single enterprise only. While Private Cloud 1006 is illustrated as being in communication with WAN 1002, in other embodiments, the private cloud may be completely isolated from the internet and accessible only through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a distinct, discrete entity, but the larger hybrid cloud architecture is coupled together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple configuration clouds. In this embodiment, both Public Cloud 1005 and Private Cloud 1006 are parts of the larger hybrid cloud.
[0126] The embodiments described herein may cover one or more systems, methods, apparatus and / or computer program products at any possible level of technical detail of integration. A computer program product may include a computer-readable storage medium (or more mediums) having computer-readable program instructions for causing a processor to execute an aspect of one or more embodiments described herein. Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by instruction-executing devices. Computer-readable storage media may not be limited to, for example, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, superconducting storage devices, and / or any preferred combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media may also include portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanical encoding devices such as punched cards or grooved structures on which instructions are recorded, and / or any preferred combination thereof. The computer-readable storage media used herein shall not be interpreted as transient signals themselves, such as radio waves and / or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides and / or other transmission media (e.g., optical pulses passing through optical fiber cables), and / or electrical signals transmitted through wiring.
[0127] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device and / or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in the computer-readable storage medium within each computing / processing device. Computer-readable program instructions for performing the operation of one or more embodiments described herein may include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, and / or source code and / or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Python®, R, Smalltalk®, C++ or similar, and / or procedural programming languages such as the "C" programming language and / or similar programming languages. Computer-readable program instructions may be executed as a whole on a computer, partially on a computer, as a standalone software package, partially on a computer and / or partially on a remote computer, or entirely on a remote computer and / or server. In the latter scenario, the remote computer may be connected to the computer through any type of network, including a local area network (LAN) and / or a wide area network (WAN), or a combination thereof, and / or may be connected to an external computer (e.g., via the Internet using an Internet service provider).In one or more embodiments, an electronic circuit including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), and / or a programmable logic array (PLA) may execute computer-readable program instructions by personalizing the electronic circuit using state information of computer-readable program instructions in order to perform an aspect of one or more embodiments described herein.
[0128] Aspects of one or more embodiments described herein are described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to one or more embodiments described herein. It will be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions may be provided to the processors of general-purpose computers, dedicated computers, and / or other programmable data processing devices to create machines, and as a result, instructions executed via the processors of the computers or other programmable data processing devices may form means for implementing the functions / actions specified in the blocks or combinations of blocks in the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored in computer-readable storage media that can instruct computers, programmable data processing devices, and / or other devices to function in a particular manner. Thus, computer-readable storage media storing instructions may include products containing instructions that can implement the modes of functions / actions specified in the blocks or combinations of blocks in the flowcharts and / or block diagrams. Computer-readable program instructions can also be loaded into a computer, other programmable data processing device and / or other device, thereby causing a series of actions to be executed on the computer, other programmable device and / or other device, creating a computer implementation process, which in turn allows instructions executed on the computer, other programmable device and / or other device to implement functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0129] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and / or operation of a system, a computer implementable method according to one or more embodiments described herein, and / or a computer program product. In this regard, each block in a flowchart or block diagram may represent a module, segment, and / or a portion of instructions containing one or more executable instructions for implementing a specified logical function. In one or more alternative implementations, the functions described in a block may occur in an order different from that shown in the drawings. For example, two consecutively shown blocks may be executed substantially simultaneously depending on the functionality involved, and / or blocks may, in some cases, be executed in reverse order. It should also be noted that each block in a block diagram and / or flowchart, and / or a combination of blocks in a block diagram and / or flowchart, may be implemented by a special-purpose hardware-based system capable of performing a specified function and / or action and / or one or more combinations of special-purpose hardware and / or computer instructions.
[0130] While the subject matter has been described above in the general context of computer executable instructions for computer program products running on a computer and / or multiple computers, those skilled in the art will recognize that one or more embodiments described herein can be implemented at least partially in parallel with one or more other program modules. Generally, a program module includes routines, programs, components, and / or data structures that perform a particular task and / or implement a particular abstract data type. Furthermore, the computer-based methods described herein can be implemented using single-processor and / or multi-processor computer systems, minicomputing devices, mainframe computers, and other computer system configurations including computers, handheld computing devices (e.g., PDAs®, telephones), and / or microprocessor-based or programmable consumer electronics and / or industrial electronics. The embodiments described may also be implemented in a distributed computing environment in which tasks are performed by remote processing devices linked over a communication network. However, one or more, but not all, of the embodiments described herein can be implemented on a standalone computer. In a distributed computing environment, program modules may reside in both local and remote memory storage devices.
[0131] As used herein, the terms “component,” “system,” “platform,” and / or “interface” may refer to and / or include computer-related entities or entities relating to operating machines having one or more inherent functionalities. Entities described herein may be hardware, a combination of hardware and software, software, or running software. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or a computer. Exemplarily, an application running on a server and both the application and the server may be components. One or more components may exist in a process and / or thread of execution, and components may be localized in one computer and / or distributed among two or more computers. In another example, each component may run from various computer-readable media storing various data structures. Components may communicate via local and / or remote processes according to signals, etc., having one or more data packets (e.g., data from one component interacting with another component across a network such as the Internet, in a local system, a distributed system, and / or with other systems via signals). As another example, a component may be a device having inherent functionality provided by mechanical parts operated by electrical or electronic circuits, which are operated by software and / or firmware applications executed by a processor. In such a case, the processor may be inside and / or outside the device and may execute at least part of the software and / or firmware application. As yet another example, a component may be a device that provides inherent functionality through electronic components without using mechanical parts, and such electronic components may include a processor and / or other means for executing software and / or firmware that at least partially provides the functionality of the electronic components.In one embodiment, the component may emulate an electronic component, for example, via a virtual machine in a cloud computing system.
[0132] In addition, the term “or” is intended to mean an inclusive “or,” not an exclusive “or.” That is, unless otherwise specified or it is clear from the context, “X uses A or B” is intended to mean any of the natural inclusive substitutions. That is, “X uses A or B” is satisfied under any of the aforementioned cases: when X uses A; when X uses B; or when X uses both A and B. Furthermore, the articles “a” and “an” used herein and in the accompanying drawings should generally be interpreted as meaning “one or plural,” unless otherwise specified or it is clear from the context that they refer to a singular form. When used herein, the terms “example” and / or “exemplary” are used to mean serving as an example, instance, or illustration. To avoid doubt, the subject matter described herein is not limited to such examples. In addition, any embodiment or design described herein as “example” and / or “exemplary” is not necessarily construed as being preferable or advantageous to other embodiments or designs, and is not intended to exclude equivalent exemplary structures and techniques known to those skilled in the art.
[0133] As used herein, the term “processor” may refer to substantially any computing processing unit and / or device, including, but not limited to, single-core processors, single-processors with software multithreading capability, multi-core processors, multi-core processors with software multithreading capability, multi-core processors with hardware multithreading technology, parallel platforms, and / or parallel platforms with distributed shared memory. Additionally, a processor may refer to integrated circuits, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), programmable logic controllers (PLCs), complex-programmable logic devices (CPLDs), discrete gate or transistor logic, discrete hardware components, and / or any combination thereof designed to perform the functions described herein. Furthermore, processors may utilize nanoscale architectures such as molecular and quantum dot-based transistors, switches, and / or gates, but not limited to, to optimize space utilization and / or enhance the performance of associated equipment. Processors may be implemented as combinations of computing processing units.
[0134] In this specification, terms such as “memory,” “storage,” “datastore,” “data storage,” “database,” and substantially any other information storage component relating to the operation and functionality of a component are used to refer to “memory” or “memory component” entities embodied in a component containing memory. The memory and / or memory components described herein may be either volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may include, but not limit, read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, and / or non-volatile random-access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). Volatile memory may include RAM that can act as, for example, external cache memory. As examples, not limitations, RAM may be available in many forms, including synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), extended SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and / or Rambus dynamic RAM (RDRAM). Additionally, the memory components described herein in relation to the systems and / or methods performed by computers are intended to include, but are not limited to, these and / or any other suitable types of memory.
[0135] The foregoing merely includes examples of systems and methods performed by computers. Naturally, for the purpose of describing one or more embodiments, it is impossible to describe every conceivable combination of components and / or methods performed by computers; however, those skilled in the art will recognize that many further combinations and / or substitutions of one or more embodiments are possible. Furthermore, where terms such as “includes,” “having,” and “equipped with” are used in the detailed description, claims, appendices, and / or drawings, such terms are intended to be inclusive in the same manner as the term “equipped with” is interpreted when used as a transitional term in a claim.
[0136] While descriptions of various embodiments have been presented for illustrative purposes, they are not intended to be exhaustive or to limit oneself to the embodiments described herein. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described. The terminology used herein has been chosen to best describe the principles, practical applications, and / or technical improvements to the technologies available on the market, and / or to enable other those skilled in the art to understand the embodiments described herein.
Claims
1. Memory for storing computer executable components; and Processor that executes the computer executable components stored in the memory The computer executable component comprises, Access components that access input datasets; and A data generation component that generates multiple new datasets using multiple quantum kernels based on the aforementioned input dataset. A system that includes this.
2. The aforementioned data generation component further: Using each of the quantum kernels of the plurality of quantum kernels, the data distribution of the input dataset is transformed into a new data distribution; Each of the aforementioned new data sets is generated by oversampling or undersampling data from each of the aforementioned new data distributions; and Based on each of the aforementioned new datasets, a classification dataset is generated, where the classification dataset based on the new dataset includes the input dataset and the new dataset. The system according to claim 1.
3. A quantum kernel selection component randomly selects a quantum kernel from the plurality of quantum kernels, wherein the quantum kernel is used to generate a new dataset from the plurality of new datasets. The system according to claim 1, further comprising:
4. A data classification component that classifies each of the classification datasets using the aforementioned plurality of quantum kernels. The system according to claim 2, further comprising:
5. A storage component that stores the classification results of each of the aforementioned classification datasets in storage. The system according to claim 4, further comprising:
6. Analyze the results above; and Based on the analysis of the above results, a classification score corresponding to the quantum kernels included in the plurality of quantum kernels is generated. Analysis Components The system according to claim 5, further comprising:
7. From a set of multiple quantum kernels in the outer loop, select a quantum kernel to be used for data generation for oversampling of minority classes or for data clustering for undersampling of majority classes; Using the same set of quantum kernels in the inner loop, select one of the same set of quantum kernels, classify the original dataset using the selected quantum kernel, and store the result of the classification; In the inner loop, select one quantum kernel after another quantum kernel, classify until the quantum kernels in the inner loop are exhausted, and store the results, then select the next quantum kernel in the outer loop and continue performing balancing; Iteratively select the next quantum kernel in the outer loop until the quantum kernel in the outer loop is exhausted, and perform classification until the inner loop is exhausted; and Once the quantum kernels in the outer loop are exhausted, the best result is selected, and a pair of quantum kernels that produced this best result is chosen, where the pair includes the selected quantum kernel in the outer loop for balancing and the selected quantum kernel in the inner loop for classification, where the best result is defined as the result having the highest under-curve area score. Quantum kernel identification component The system according to claim 6, further comprising:
8. The system according to any one of claims 1 to 7, wherein the plurality of quantum kernels are quantum feature maps.
9. The step of accessing the input dataset by a system operationally coupled to the processor; and The system generates multiple new datasets using multiple quantum kernels based on the input dataset. A method performed by a computer, comprising the following:
10. The system performs the step of transforming the data distribution of the input dataset into a new data distribution using each of the quantum kernels of the plurality of quantum kernels; The steps include: generating each of the multiple new datasets by oversampling or undersampling data from each of the new data distributions using the system; and The system generates a classification dataset based on each of the new datasets, where the classification dataset based on the new dataset includes the input dataset and the new dataset. The computer-based method according to claim 9, further comprising the following:
11. The system performs the step of selecting a quantum kernel from the plurality of quantum kernels, where the quantum kernel is used to generate a new dataset from the plurality of new datasets, and where the quantum kernel is selected randomly. The computer-based method according to claim 9, further comprising the following:
12. The system performs the step of classifying each of the classification datasets using the multiple quantum kernels. A computer-based method according to claim 10, further comprising the following:
13. The system then stores the classification results of each of the classification datasets in storage. A computer-based method according to claim 12, further comprising the following:
14. The system performs the steps of analyzing the results; and The system generates classification scores corresponding to the quantum kernels included in the plurality of quantum kernels based on the analysis of the results. A computer-based method according to claim 13, further comprising the following:
15. The system includes the step of selecting a first quantum kernel based on the classification score, wherein the first quantum kernel can be used to generate a new dataset based on the original dataset through balancing; and The system then selects a second quantum kernel based on the classification score, wherein the second quantum kernel is available for classifying the data contained in the new dataset. A computer-based method according to claim 14, further comprising the following:
16. The method executed by a computer according to any one of claims 9 to 15, wherein the plurality of quantum kernels are quantum feature maps.
17. The program instructions are provided to the processor: The processor allows access to the input dataset; and The processor generates multiple new datasets using multiple quantum kernels based on the input dataset. Thus, a computer program for quantum kernel selection, executable by the aforementioned processor.
18. The aforementioned program instructions are sent to the processor: The processor uses each of the quantum kernels of the plurality of quantum kernels to transform the data distribution of the input dataset into a new data distribution; The processor generates each of the multiple new datasets by oversampling or undersampling data from each of the new data distributions; and The processor causes each of the new datasets to generate each classification dataset, where the classification dataset based on the new dataset includes the input dataset and the new dataset. The computer program according to claim 17, further executable by the processor as described above.
19. The aforementioned program instructions are sent to the processor: The processor uses the multiple quantum kernels to classify each of the classification datasets. The computer program according to claim 18, further executable by the processor as described above.
20. The computer program according to any one of claims 17 to 19, wherein the plurality of quantum kernels are quantum feature maps.