Method and system for identifying a molecule of artificial deoxyribonucleic acid (DNA) nucleobase

IN598754BActive Publication Date: 2026-08-11INDIAN INST OF TECH INDORE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
IN202521002474
Authority / Receiving Office
IN · IN
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2026-08-11
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Conventional methods for identifying artificial DNA nucleobases suffer from limitations such as primer design issues, contamination risks, high costs, complex data analysis, labor-intensiveness, and sensitivity challenges, making them inefficient for precise quantification and identification.

Method used

A method and system utilizing quantum tunnelling nanogap junctions and machine learning (ML) to measure and analyze quantum tunnelling transmission parameters of artificial DNA nucleobases, extracting features through a machine learning model for accurate identification.

Benefits of technology

Enables high-precision identification of artificial DNA nucleobases with improved sensitivity and accuracy, overcoming the drawbacks of traditional techniques.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A method (400) for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning (ML), is disclosed. The method (400) includes positioning (402) at least one artificial DNA nucleobase within a solid-state nanogap device comprising a left electrode and a right electrode. The method (400) includes measuring (404) quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase. The method (400) further includes determining (406) one or more features from the quantum tunnelling transmission parameters. Furthermore, the method (400) includes predicting (408) an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a ML model.
Need to check novelty before this filing date? Find Prior Art

Description

FIELD OF THE INVENTIONThe present disclosure relates to artificial learning techniques, and more particularly to methods and systems for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning.BACKGROUNDDeoxyribonucleic Acid (DNA) is a molecule that carries the genetic instructions used in the growth, development, functioning, and reproduction of all living organisms and many viruses. DNA has remained the unvarying nucleus of terrain life, distinguished by its consistent structure and its unique ability to store, transmit, and undergo evolutionary processes among molecules. Further, recent advances in exploratory microbiology have uncovered life forms that utilize DNA constructed with nucleotides divergent from adenine (A), guanine (G), cytosine (C), and thymine (T). For example, Artificial / Benzo-homologated / expanded DNA (xDNA and yDNA) is the synthesized nucleotide system that is expanded in size by fusing a benzene ring with one of the four natural DNA bases and has potential implications for DNA data storage and biological processes. Further, artificial DNA design has revolutionized high-density information storage and transfer with sequence complexity scaling as 8^n compared to 4^n for natural DNA, where n represents sequence length. This increased complexity offers unparalleled advantages for data encryption, processing, and the development of more intricate structural designs. The synthesized DNA play a crucial role in gene regulation by interacting with specific proteins, potentially offering better stability in the DNA damage response compared to natural DNA. Further, beyond gene regulation, artificial DNA holds promise as a diagnostic tool for probing steric effects in the active sites of polymerase enzymes. The increased base surface area of artificial DNA helices increases their thermal stability through stronger hydrophobic stacking interactions. Moreover, the fluorescent properties of xDNA and yDNA oligomers make them ideal for detecting and imaging natural nucleic acid sequences. Both xDNA and yDNA feature high melting point and smaller HOMO-LUMO gaps than the natural DNA, enabling them to function as molecular nanowires. Therefore, the identification of artificial DNA is essential for advancing the understanding of genetics, evolution, forensic analysis, and DNA data storage applications. To identify and analyze molecules of artificial DNA (also referred to as synthetic DNA) involves various techniques that is used to confirm its presence, structure, sequence, and function, such as Polymerase Chain Reaction (PCR), Gel Electrophoresis, DNA Sequencing, Fluorescent In Situ Hybridization (FISH), Southern Blotting, Mass Spectrometry (MS) and so on. However, these techniques suffer from major drawbacks, such as:PCR: Primer design issues, contamination risks, limited to known sequences, amplification bias.Gel Electrophoresis: Resolution limits, time-consuming, low sensitivity, no precise quantification.DNA Sequencing: High cost, errors, complex data analysis, short read lengths.FISH: Limited resolution, probe design issues, quantification challenges, labour-intensive.Southern Blotting: Labor-intensive, less sensitive, limited resolution requires specialized equipment.Mass Spectrometry: Complex sample prep, sensitivity issues for large molecules, expensive equipment, data interpretation complexity.Thus, there is a need to overcome at least the above challenges associated with conventional techniques of identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase. SUMMARYThis summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended for determining the scope of the invention.According to one embodiment of the present disclosure, a method for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning (ML). The method includes positioning at least one artificial DNA nucleobase within a solid-state nanogap device comprising a left electrode and a right electrode. The left electrode serves as an electron source and the right electrode serves as an electron drain. The method includes measuring quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase. The method further includes determining one or more features from the quantum tunnelling transmission parameters. The one or more features indicate an overlap of one or more electric readouts associated with the at least one artificial DNA nucleobase. Furthermore, the method includes predicting an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a machine learning (ML) model.According to another embodiment of the present disclosure, a system for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning is disclosed. The system includes a nanogap device and a server in communication with the nanogap device. The server includes a memory and at least one processor in communication with the memory. The at least one processor is configured to position at least one artificial DNA nucleobase within a solid-state nanogap device comprising a left electrode and a right electrode. The left electrode serves as an electron source and the right electrode serves as an electron drain. The at least one processor is configured to measure quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase. Further, the at least one processor is configured to determine one or more features from the quantum tunnelling transmission parameters. The one or more features indicate an overlap of one or more electric readouts associated with the at least one artificial DNA nucleobase. Furthermore, the at least one processor is configured to predict an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a machine learning (ML) model.To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting to its scope. The invention will be described and explained with additional specificity and detail with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGSThese and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:Figure 1 illustrates a schematic block diagram depicting an environment for the implementation of a system for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning (ML), according to an embodiment of the present invention;Figure 2 illustrates a schematic block diagram of modules components of the system, according to an embodiment of the present invention;Figure 3 illustrates a process flow associated with the system, according to an embodiment of the present invention; andFigure 4 illustrates an exemplary process flow comprising a method for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning (ML), according to an embodiment of the present invention.Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help improve understanding of aspects of the present invention. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein. DETAILED DESCRIPTION For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the various embodiments, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein being contemplated as would normally occur to one skilled in the art to which the invention relates.It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the invention and are not intended to be restrictive thereof. Reference throughout this specification to "an aspect", "another aspect" or similar language means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrase "in an embodiment", "in another embodiment" and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.The terms "comprise", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components preceded by "comprises... a" does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.Hereinafter, it is understood that terms including "unit" or "module" at the end may refer to the unit for processing at least one function or operation and may be implemented in hardware, software, or a combination of hardware and software.Unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted to not unnecessarily obscure the embodiments herein. Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. The term "or" as used herein, refers to a non-exclusive or unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein. As is traditional in the field, embodiments may be described and illustrated in terms of blocks that carry out a described function or functions. These blocks, which may be referred to herein as units or modules or the like, are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure. The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.For the sake of clarity, the first digit of a reference numeral of each component of the present disclosure is indicative of the Figure number, in which the corresponding component is shown. For example, reference numerals starting with digit "1" are shown at least in Figure 1. Similarly, reference numerals starting with digit "2" are shown at least in Figure 2.An object of the present disclosure is to identify a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning (ML).Another object of the present disclosure is to utilize machine learning for feature extraction, wherein the extracted features are fed into a descriptor matrix which undergoes model training, and evaluation. Yet another object of the present disclosure is to examine electronic transmission pathways with molecular orbitals (MOs) analysis of the quantum tunnelling transmission parameters to reveal the role of chemical and electronic signatures of artificial DNA nucleobases on their tunnelling conductance and electric readouts.Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings.Figure 1 illustrates a schematic block diagram depicting an environment 100 for the implementation of a system 102 for identifying a molecule of artificial DNA nucleobase using quantum tunnelling nanogap junction and ML, according to an embodiment of the present invention.In an embodiment, referring to Figure 1, the system 102 may be implemented in a server 106 or any user equipment (UE) (not shown) via an application installed in the server 106 and running on an operating system (OS) that generally defines a first active user environment. The OS typically presents or displays the application through a graphical user interface ("GUI") of the OS. In a non-limiting example, the server 106 may be a laptop computer, a desktop computer, a Personal Computer (PC), a notebook, a smartphone, a tablet, a smartwatch, or any device capable of displaying electronic media.In an embodiment, the system 102 may include a Machine Learning (ML) model 108 trained using data collected for a set of known xDNA and yDNA. The ML model 108 is configured to capture one or more electric readouts, extract one or more features, train the ML model based on the extracted one or more features and consequently predict an identity of the molecule of artificial DNA nucleobase. Thus, the ML model 108 is configured to display 110 as the prediction.In an embodiment, the system 102 may include a nanogap device 104 comprising a left electrode and a right electrode. The left electrode serves as an electron source and the right electrode serves as an electron drain measuring quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase.Thus, the Figure 1 illustrates the system 102, comprising the nanogap device 104 in communication with the server 106. The server 106 incorporates the ML model 108 configured for the prediction of the identity of a molecule of artificial DNA nucleobase using quantum tunnelling nanogap junction and ML.In an embodiment, the system 102 is configured through modelling and density functional theory (DFT) calculations, to investigate the structural and electronic properties. The structures of all eight isolated artificial DNA (xDNA and yDNA) nucleobases are initially optimized. For the quantum tunnelling transmission of xA, xT, xG, xC, yA, yT, yG, and yC nucleobases, the graphene electrode-based nanogap device is modelled. The electrode edges of the nanogap device are terminated with nitrogen atoms, which significantly enhances its sensitivity for nucleobase identification. The designed graphene nanogap is preferably modelled with a supercell size of 24.00 x 17.16 x 42.49 Å3 in x, y, and z-direction and with an N-N gap distance of 14 Å two (left / right) electrodes. In an embodiment, the server 106 is configured to position at least one artificial DNA nucleobase within a solid-state nanogap device comprising a left electrode and a right electrode. The left electrode serves as an electron source and the right electrode serves as an electron drain.In an embodiment, the server 106 is configured to measure quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase. The quantum tunnelling transmission parameters used for training the ML model comprises one or more fingerprint transmission functions to analyze influence of electronic coupling and molecular orbitals delocalization on the quantum tunnelling transmission parameters and one or more electric current readout of eight artificial DNA nucleobases which includes xDNA (xA, xT, xG, and xC) and yDNA (yA, yT, yG, and yC) located inside the solid-state nanogap device.In an embodiment, the server 106 is configured to determine one or more features from the quantum tunnelling transmission parameters. The one or more features indicate overlap of one or more electric readouts associated with the at least one artificial DNA nucleobase. In an embodiment, the server 106 is configured to predict an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a machine learning (ML) model 108.The ML model's 108 performance was evaluated using accuracy, precision, recall, and F1 score. In an embodiment, the server 106 is configured to train the ML model 108 by passing a set of xDNA and a set of yDNA through the nanogap device. Then the server 106 is configured to capture one or more electric readout in response to passing the set of xDNA and the set of yDNA through the nanogap device and extract one or more features based on the one or more electric readouts. The server 106 is configured to train the ML model 108 based on the extracted one or more feature.In an example embodiment, to train the ML model, the present invention collects transmission function datasets of xDNA and yDNA nucleobases, respectively. In an example embodiment, the fingerprint transmission function (TF) may be used as the primary descriptor that captures the essential structural and chemical properties of nucleobases through electrode-nucleobase-electrode electronic coupling. Since TF consists of a range of datasets, the dimensionality and distinguishability among the datasets may be improved by extracting minimum, maximum, and mean features and generating 'Min', 'Max', and 'Mean' descriptors by normalizing the original TF with these features. Further, Random Forest Classifier (RFC) model may be used for the classification of the datasets.Thus, to perform the ML-driven single molecule recognition of artificial nucleobases, separate xDNA and yDNA databases are utilized for training the RFC model. The RFC model is trained on 80% of each dataset, while the remaining 20% is allocated for testing. Figure 2 illustrates a schematic block diagram of modules components of the system 102, according to an embodiment of the present invention.The server 106 may include but is not limited to, a processor 202, memory 204, modules 206, and data 208. The modules 206 and the memory 204 may be coupled to the processor 202.The processor 202 can be a single processing unit or several units, all of which could include multiple computing units. The processor 202 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any devices that manipulate signals based on operational instructions. Among other capabilities, the processor 202 is adapted to fetch and execute computer-readable instructions and data stored in the memory 204.The memory 204 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random-access memory (SRAM) and dynamic random-access memory (DRAM), and / or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes. The memory 204 may alternatively be referred to as the database 204 in the present disclosure, within the scope of the invention. The modules 206, amongst other things, include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types. The modules 206 may also be implemented as signal processor(s), state machine(s), logic circuitries, and / or any other device or component that manipulates signals based on operational instructions.Further, the modules 206 can be implemented in hardware, instructions executed by a processing unit, or by a combination thereof. The processor 202 can comprise a computer, a processor, a state machine, a logic array, or any other suitable devices capable of processing instructions. The processing unit can be a general-purpose processor (e.g., processor 202) which executes instructions to cause the general-purpose processor to perform the required tasks or, the processing unit can be dedicated to performing the required functions. In another embodiment of the present disclosure, the modules 206 may be machine-readable instructions (software) which, when executed by the processor 202 / processing unit, perform any of the described functionalities / methods, as discussed throughout the present disclosure. Furthermore, the modules 206 may be implemented through an artificial intelligence (AI) model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor.The processor 202 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU).The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed and may be implemented through a separate server / system.The AI model may consist of a plurality of neural network layers. Each layer has a plurality of weight values and performs a layer operation through the calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.The learning algorithm is a method for identifying the disaccharide isomers using quantum tunnelling and the Machine Learning (ML) model wherein the ML model is trained using transmission data collected for a set of known xDNA and a set of known yDNA. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Further details of training the ML model and the method for identifying the molecule of artificial DNA nucleobase are explained in the forthcoming paragraphs for Figure 3 and Figure 4.In an embodiment, the data unit 208 serves, amongst other things, as a repository for storing data processed, received, and generated by one or more of the modules 206.In another embodiment of the present disclosure, the processor 202 via the modules 206 is configured to execute machine-readable instructions (software) which perform the working of the system 102 within the scope of the present invention as described in forthcoming paragraphs. In an embodiment, the system 202 is configured to position at least one artificial DNA nucleobase within a solid-state nanogap device comprising a left electrode and a right electrode. The left electrode serves as an electron source and the right electrode serves as an electron drain.In an embodiment, the system 202 is configured to measure quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase. The quantum tunnelling transmission parameters used for training the ML model comprises one or more fingerprint transmission functions for analyzing influence of electronic coupling and molecular orbitals delocalization on the quantum tunnelling transmission parameters and one or more electric current readout of eight artificial DNA nucleobases which includes xDNA (xA, xT, xG, and xC) and yDNA (yA, yT, yG, and yC) located inside the solid-state nanogap device.In an embodiment, the system 202 is configured to determine one or more features from the quantum tunnelling transmission parameters. The one or more features indicate overlap of one or more electric readouts associated with the at least one artificial DNA nucleobase.In an embodiment, the system 202 is configured to predict an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a machine learning (ML) model.In an embodiment, to train the ML model, the system 202 is configured to pass a set of xDNA and a set of yDNA through the nanogap device. The system 202 is configured to capture one or more electric readouts in response to passing the set of xDNA and the set of yDNA through the nanogap device and extract one or more features based on the one or more electric readouts. Further, the system 202 is configured to train the ML model based on the extracted one or more features.In an embodiment, to determine one or more features, the system 202 is configured to extract minimum, maximum, and mean descriptors from the analyzed quantum tunnelling transmission parameters. Further, the system 202 is configured to generate at least one of 'Min', 'Max', and 'Mean' features by normalizing the extracted features to determine one or more features.In an embodiment, the system 202 is configured to utilize explainable ML techniques, thereby providing a degree of influence for each of the one or more features, wherein explainable AI techniques comprise SHapley Additive exPlanations (SHAP).In an embodiment, to predict, the system 202 is configured to determine each of the nucleotides of artificial DNA nucleobase from the one or more electric readouts in a sequence and predict the identity of the artificial DNA nucleobase by correlating the extracted one or more features with the determined each of the nucleotides of artificial DNA nucleobase.In an embodiment, the system 202 is configured to compute prediction performance metrics, including accuracy, precision, recall, and F1 score, by the ML model based on the prediction. In an embodiment, the system 202 is configured to reanalyse remaining of the quantum tunnelling transmission parameters iteratively, as an input in the ML model. In an embodiment, the system 202 is configured to prepare one or more multi-class classification models for nucleobase recognition to predict one or more artificial DNA nucleobases.Figure 3 illustrates a process flow associated with the system, according to an embodiment of the present invention.The process flow 300 illustrates exemplary steps for training the ML model 108 using the transmission data collected for the set of known xDNA and the set of known yDNA.At block 302, data collection and preprocessing are performed. In an embodiment, transmission data is acquired based on measuring the quantum tunnelling transmission parameters for the set of known xDNA and yDNA. Further, in an embodiment, filtering or smoothing techniques are applied to remove noise and outliers from the measured quantum tunnelling transmission parameters.In an embodiment, the one or more features are extracted from the transmission data (e.g., the quantum tunnelling transmission parameters), such as Minima, maxima, and mean descriptors of transmission parameters. In an example, the extracted one or more features are normalized to ensure uniform scaling for effective model training.At block 304, model selection is performed i.e., an appropriate machine learning model is selected, such as gradient-boosted decision trees, neural networks, or support vector machines (SVM).In an embodiment, the ML model 108 is selected to handle multi-class classification for nucleobase recognition to predict one or more artificial DNA nucleobases.At block 306, the ML model 108 is trained. In an embodiment, a training dataset may be split into training, validation, and test sets. In an embodiment, the ML model 108 is thus trained by passing a set of xDNA and a set of yDNA through the nanogap device. Then the ML model 108 is configured to capture one or more electric readout in response to passing the set of xDNA and the set of yDNA through the nanogap device and extract one or more features based on the one or more electric readouts. The ML model 108 is then trained based on the extracted one or more feature.In an example embodiment, to train the ML model, the present invention collects transmission function datasets of xDNA and yDNA nucleobases, respectively. In an example embodiment, the fingerprint transmission function (TF) may be used as the primary descriptor that captures the essential structural and chemical properties of nucleobases through electrode-nucleobase-electrode electronic coupling. Since TF consists of a range of datasets, the dimensionality and distinguishability among the datasets may be improved by extracting minimum, maximum, and mean features and generating 'Min', 'Max', and 'Mean' descriptors by normalizing the original TF with these features. Further, Random Forest Classifier (RFC) model may be used for the classification of the datasets.Thus, to perform the ML-driven single molecule recognition of artificial nucleobases, separate xDNA and yDNA databases are utilized for training the RFC model. The RFC model is trained on 80% of each dataset, while the remaining 20% is allocated for testing. At block 308, evaluation for the training is performed. In an embodiment, the system 102 may be configured to compute classification performance metrics such as accuracy, precision, recall, F1 score, sensitivity, and specificity.Further, in an embodiment, explainable AI techniques may be used to validate that the ML model's 108 decisions align with the expected influence of features.At block 310, iterative refinement is performed. In an embodiment, the system 102 may be configured to identify underperforming features and refine them (e.g., add new statistical metrics or remove redundant ones). Additionally, the additional transmission data may be simulated for known isomers (e.g., through molecular modeling) to expand / increase the training set.Consequently, refinements may be incorporated for retraining the ML model 108 to improve classification accuracy and robustness.At block 312, the trained ML Model may be deployed upon achieving satisfactory performance, for real-time or batch processing of new quantum tunnelling data to predict the molecule of artificial DNA nucleobase.These steps collectively enable the development of a robust and accurate ML model that leverages quantum tunnelling parameters for isomer identification.Figure 4 illustrates an exemplary process flow comprising a method for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning (ML), according to an embodiment of the present invention. The method 400 may be a computer-implemented method executed, for example, by the server 106 and the modules 206. For the sake of brevity, constructional and operational features of the system 102 that are already explained in the description of Figure 1, Figure 2, and Figure 3 are not explained in detail in the description of Figure 4.At step 402, the method 400 may include positioning at least one artificial DNA nucleobase within a solid-state nanogap device 104 comprising a left electrode and a right electrode. The left electrode serves as the electron source and the right electrode serves as the electron drain. At step 404, the method 400 may include measuring quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase. The quantum tunnelling transmission parameters used for training the ML model comprises one or more fingerprint transmission function for analyzing influence of electronic coupling and molecular orbitals delocalization on the quantum tunnelling transmission parameters and one or more electric current readout of eight artificial DNA nucleobases which includes xDNA (xA, xT, xG, and xC) and yDNA (yA, yT, yG, and yC) located inside the solid-state nanogap device.At step 406, the method 400 may include determining one or more features from the quantum tunnelling transmission parameters, wherein the one or more features indicate overlap of one or more electric readouts associated with the at least one artificial DNA nucleobase.At step 408, the method 400 may include predicting an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a ML model 108. In an embodiment, for training the ML model, the method 400 may include passing a set of xDNA and a set of yDNA through the nanogap device. The method 400 may include capturing one or more electric readout in response to passing the set of xDNA and the set of yDNA through the nanogap device and extracting one or more features based on the one or more electric readouts. The method 400 may include training the ML model based on the extracted one or more feature.In an embodiment, the method 400 may include utilizing explainable ML techniques, thereby providing a degree of influence for each of the one or more features, wherein explainable AI techniques comprise Local Interpretable Model-agnostic Explanations (LIME) with SHapley Additive exPlanations (SHAP) analysis.In an embodiment, for determining one or more features, the method 400 may include extracting minimum, maximum, and mean descriptors from the analyzed quantum tunnelling transmission parameters. The method 400 may further include generating at least one of 'Min', 'Max', and 'Mean' features by normalizing the extracted features to determine one or more features.In an embodiment, for predicting, the method 400 may include determining each of the nucleotides of artificial DNA nucleobase from the one or more electric readouts in a sequence. The method 400 may further include predicting the identity of the artificial DNA nucleobase by correlating the extracted one or more features with the determined each of the nucleotides of artificial DNA nucleobase.In an embodiment, the method 400 may include computing prediction performance metrics, including accuracy, precision, recall, and F1 score, by the ML model based on the prediction.In an embodiment, the method 400 may include reanalysing remaining of the quantum tunnelling transmission parameters iteratively, as an input in the ML model.In an embodiment, the method 400 may include preparing one or more multi-class classification model for nucleobase recognition to predict one or more artificial DNA nucleobases.Further, the present disclosure provides the following technical advantages:The present disclosure enables identification of a molecule of artificial DNA nucleobase with high precision accuracy.The present disclosure enables identification of over a broad range of natural, mutated and artificial DNA.The present disclosure uses the best-fitted ML model for the classification of artificial DNA nucleobases. While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein. The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.

Claims

1. A method (400) for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleobase using quantum tunnelling nanogap junction and machine learning (ML), the method comprising: positioning (402) at least one artificial DNA nucleobase within a solid-state nanogap device comprising a left electrode and a right electrode, wherein the left electrode serves as an electron source and the right electrode serves as an electron drain; measuring (404) quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase; determining (406) one or more features from the quantum tunnelling transmission parameters, wherein the one or more features indicate overlap of one or more electric readouts associated with the at least one artificial DNA nucleobase; and predicting (408) an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a machine learning (ML) model.

2. The method (400) as claimed in claim 1, wherein the quantum tunnelling transmission parameters used for training the ML model comprises one or more fingerprint transmission function for analyzing influence of electronic coupling and molecular orbitals delocalization on the quantum tunnelling transmission parameters and one or more electric current readout of eight artificial DNA nucleobases which includes xDNA (xA, xT, xG, and xC) and yDNA (yA, yT, yG, and yC) located inside the solid-state nanogap device.

3. The method (400) as claimed in claim 1, comprising training the ML model comprises passing a set of xDNA and a set of yDNA through the nanogap device; capturing one or more electric readout in response to passing the set of xDNA and the set of yDNA through the nanogap device; extracting one or more features based on the one or more electric readouts; and training the ML model based on the extracted one or more feature.

4. The method (400) as claimed in claim 1, wherein determining one or more features comprises: extracting minimum, maximum, and mean descriptors from the analyzed quantum tunnelling transmission parameters; and generating at least one of 'Min', 'Max', and 'Mean' features by normalizing the extracted features to determine one or more features.

5. The method (400) as claimed in claim 1, comprising utilizing explainable ML techniques, thereby providing a degree of influence for each of the one or more features, wherein explainable AI techniques comprise Local Interpretable Model-agnostic Explanations (LIME) with SHapley Additive exPlanations (SHAP) analysis.

6. The method (400) as claimed in claim 1, wherein the prediction comprises: determining each of the nucleotides of artificial DNA nucleobase from the one or more electric readouts in a sequence; and predicting the identity of the artificial DNA nucleobase by correlating the extracted one or more features with the determined each of the nucleotides of artificial DNA nucleobase.

7. The method (400) as claimed in claim 1, comprises: computing prediction performance metrics, including accuracy, precision, recall, and F1 score, by the ML model based on the prediction.

8. The method (400) as claimed in claim 1 comprises: reanalysing remaining of the quantum tunnelling transmission parameters iteratively, as an input in the ML model.

9. The method (400) as claimed in claim 1, further comprises: preparing one or more multi-class classification model for nucleobase recognition to predict one or more artificial DNA nucleobases.

10. A system (102) for identifying a molecule of artificial Deoxyribonucleic acid (DNA) nucleabase using quantum tunnelling nanogap junction and machine learning (ML), the system comprising: a nanogap device (104); a server (106) in communication with the nanogap device (104), the server (106) comprising: a memory (204); at least one processor (202) in communication with the memory (204), the at least one processor (202) configured to: position at least one artificial DNA nucleobase within a solid-state nanogap device comprising a left electrode and a right electrode, wherein the left electrode serves as an electron source and the right electrode serves as an electron drain; measure quantum tunnelling transmission parameters of the at least one artificial DNA nucleobase; determine one or more features from the quantum tunnelling transmission parameters, wherein the one or more features indicate overlap of one or more electric readouts associated with the at least one artificial DNA nucleobase; and predict an identity of the molecule of artificial DNA nucleobase based on the extracted one or more features using a machine learning (ML) model.

11. The system (102) as claimed in claim 10, wherein the quantum tunnelling transmission parameters used for training the ML model comprises one or more fingerprint transmission function to analyze influence of electronic coupling and molecular orbitals delocalization on the quantum tunnelling transmission parameters and one or more electric current readout of eight artificial DNA nucleobases which includes xDNA (xA, xT, xG, and xC) and yDNA (yA, yT, yG, and yC) located inside the solid-state nanogap device.

12. The system (102) as claimed in claim 10, wherein to train the ML model, the one or more processors (202) are configured to: pass a set of xDNA and a set of yDNA through the nanogap device; capture one or more electric readout in response to passing the set of xDNA and the set of yDNA through the nanogap device; extract one or more features based on the one or more electric readouts; and train the ML model based on the extracted one or more feature.

13. The system (102) as claimed in claim 10, wherein to determine one or more features, the one or more processors (202) are configured to: extract minimum, maximum, and mean descriptors from the analyzed quantum tunnelling transmission parameters; and generate at least one of 'Min', 'Max', and 'Mean' features by normalizing the extracted features to determine one or more features.

14. The system (102) as claimed in claim 10, the one or more processors (202) are configured to: utilize explainable ML techniques, thereby providing a degree of influence for each of the one or more features, wherein explainable AI techniques comprise Local Interpretable Model-agnostic Explanations (LIME) with SHapley Additive exPlanations (SHAP) analysis.

15. The system (102) as claimed in claim 10, wherein to predict, the one or more processors (202) are configured to: determine each of the nucleotides of artificial DNA nucleobase from the one or more electric readouts in a sequence; and predict the identity of the artificial DNA nucleobase by correlating the extracted one or more features with the determined each of the nucleotides of artificial DNA nucleobase.

16. The system (102) as claimed in claim 10, the one or more processors (202) are configured to: compute prediction performance metrics, including accuracy, precision, recall, and F1 score, by the ML model based on the prediction.

17. The system (102) as claimed in claim 10, the one or more processors (202) are configured to: reanalyse remaining of the quantum tunnelling transmission parameters iteratively, as an input in the ML model.

18. The system (102) as claimed in claim 10, the one or more processors (202) are configured to: prepare one or more multi-class classification model for nucleobase recognition to predict one or more artificial DNA nucleobases.