Methods and apparatus for generating data having characteristics associated with random numbers

By harnessing the entropy from macromolecule characterization instruments like nanopore sequencers and applying statistical and cryptographic methods, the method generates high-quality random numbers efficiently, addressing implementation challenges of TRNGs and meeting security demands.

WO2026030304A1PCT designated stage Publication Date: 2026-02-05VEIOVIA LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039642
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-29
Filing Date
2025-07-29
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing random number generators, particularly true random number generators (TRNGs), face challenges in implementation complexity and speed, making it difficult to meet the demand for high-quality random numbers required in secure applications.

Method used

Utilize the inherent entropy from instruments characterizing macromolecules, such as nanopore sequencers, to generate random numbers by processing data output from these instruments, applying statistical tests to extract and clean the bitstream to remove structured data, and using cryptographic functions to enhance randomness.

Benefits of technology

Provides a high-quality source of random numbers suitable for secure applications, leveraging existing sequencing technology to generate random numbers efficiently and at scale without dedicated hardware, ensuring statistical randomness and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000018_0001
    Figure IMGF000018_0001
  • Figure IMGF000019_0001
    Figure IMGF000019_0001
  • Figure 00000024_0000
    Figure 00000024_0000
Patent Text Reader

Abstract

Disclosed processing circuitry: receives data output from an instrument characterising macromolecules in a sample, the data including a series of time ordered output signals indicative of measurements of macromolecules recorded over time; extracts, from the output signals, a bitstream, the bitstream comprising time ordered bits extracted from a predetermined common position within each output signal; and uses the extracted bitstream as a source of entropy for generating data having characteristics associated with random numbers.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND APPARATUS FOR GENERATING DATA HAVING CHARACTERISTICS ASSOCIATED WITH RANDOM NUMBERS CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent No.63 / 676,897 as filed on July 29, 2024, which is incorporated by reference herein in its entirety for all purposes. TECHNICAL FIELD

[0002] The present disclosure relates to generating random numbers, or data having characteristics associated therewith. In particular, the disclosure provides methods, apparatuses and computer program products for random number generation, or for generation of data having characteristics associated with random numbers based on data output by electronic apparatus, for example sequencing instrument data. BACKGROUND

[0003] Randomness plays an important role in technical operations and protocols including cryptography, communication, secure public lotteries and / or voting, gaming and the like. However, meeting the demand for high quality random numbers can be a significant challenge.

[0004] Random number generators can be divided into two broad categories: Pseudo random number generators (PRNG), and ‘true’ random number generators (TRNG). PRNGs use algorithms and some system-supplied entropy to generate random numbers by an unpredictable process. For example, Random, the Linux staple PRNG, is a cryptographically secure PRNG that uses a locally stored entropy pool to seed random number generation through a process of cryptographically secure hashing.

[0005] A true random number generator (TRNG) is a device or process that generates random numbers using a source of true randomness from the physical world. Unlike pseudo-random number generators (PRNGs), which produce sequences of numbers that appear random but are determined by a starting value or ‘seed’, TRNGs rely oninherently unpredictable physical processes or phenomena to generate random data in a non-deterministic manner.

[0006] The key characteristic of a TRNG is its ability to provide truly unpredictable and statistically random numbers that cannot be feasibly reproduced, even with complete knowledge of the generator's internal state or previous output. These random numbers are said to exhibit "entropy," which is a measure of the unpredictability and randomness of a data source.

[0007] TRNGs generally exploit some naturally unpredictable process, which supplies entropy for example for use in either direct conversion to binary values or as raw material for a cryptographic or hash function. TRNGs can utilize various physical phenomena as sources of entropy, such as: ● Thermal Noise: Random fluctuations in electrical voltages or currents due to thermal effects in electronic components. ● Radioactive Decay: The timing of radioactive decay events, which is a fundamentally random process. For example, Americium and unstable Nickel isotopes are used in a variety of random number generators. ● Atmospheric Noise: Random variations in radio frequencies caused by atmospheric conditions. ● Electronic Component Noise: Random fluctuations in electronic components like resistors or diodes. One example is ChaosKey, which uses ring oscillators to provide randomness as expressed through the unpredictable phase and frequency values of ring oscillators over short time periods. ● Quantum effects: Quantum Phenomena may be used by a special class of TRNGs. Certain quantum processes, like photon emission, may be harnessed that are inherently random and unpredictable.

[0008] TRNGs have some hardware element to them to harness and process the raw entropy, and this hardware is typically specifically designed for the purposes of random number generation. The output of a TRNG is not generally influenced by any external factors or seed values, making it of use for applications that require high levels of security, cryptographic key generation, secure communication, and unbiased statistical sampling. However, TRNGs can be more challenging to implement and are often slower than PRNGs.SUMMARY OF THE DISCLOSURE

[0009] In some aspects, methods and apparatus for generating random numbers, or data having at least some of the characteristics of random numbers, are described. It may be appreciated that it is, in a practical sense, not possible to know for certain that a number generated is truly and completely random. Rather, a number may be described as ‘random’ (or, equivalently herein, as having at least some of the characteristics of random numbers) if structure has not been detected therein, for example through use of statistical tests. Thus, the terms ‘random number’ and ‘data having at least some of the characteristics of random numbers’ are used at least somewhat interchangeably herein. A random number may, in this context, comprise a binary number, i.e. a series of bits which may represent a “1” or a “0”, wherein the sequence of the 1s and 0s is unpredictable. A data output may be received from an instrument characterising macromolecules (such as DNA, RNA, or protein sequences) in a sample. The received data may include measurement event information relating to measurements of individual macromolecules recorded over time.

[0010] According to a first aspect, a method comprises, by processing circuitry, receiving data output from an instrument characterising macromolecules in a sample, the data comprising a series of time ordered output signals indicative of measurements of macromolecules recorded over time; and extracting, from the output signals, a bitstream, the bitstream comprising time ordered bits extracted from a predetermined common position within each output signal (e.g. a bit position); and using the extracted bitstream as a source of entropy for generating data having characteristics associated with random numbers.

[0011] In embodiments, the received data output from the instrument may be genomics data from a sequencer characterising DNA, RNA, or proteomics data output from an instrument characterising biopolymers in a biological sample. In embodiments, the received data output from the instrument may be data representative of a current, for example measured across a nanopore as a nucleic acid sequence passes through the nanopore. The time ordered output signals may be generated and / or output consecutively. The output signals may be generated and / or output at regular or irregular intervals.

[0012] In this way, the operation of instruments such as nanopore sequencers for generating data that can be processed to sequence DNA from a biological sample, may also be processed to generate random numbers by exploiting the entropy inherent in the measurement events relating to measurements of individual macromolecules recordedover time. In some examples, such instruments may thereby effectively provide a source of entropy as a by-product of their primary purpose, in some examples in relatively large volumes.

[0013] In embodiments, the source of entropy may be, at least in part, the physical randomness arising from the interaction between the macromolecules and the environment in which the macromolecules may be measured by the instrument. In embodiments, the source of entropy may be, at least in part, the physical randomness arising from the fluid dynamics of the passage of the macromolecules through a microfluidic structure during measurement by the instrument. In embodiments, the source of entropy may be, at least in part, the physical randomness arising from electronic noise from within components of the apparatus. In some examples, these and other phenomena may contribute alone or in combination to the random qualities of the extracted bitstream.

[0014] Through rigorous statistical analysis and entropy estimation, it has been identified that the entropy arising from the operation of instruments characterising macromolecules is usable as a source of randomness for a hardware random number generator. For example, the data processing method disclosed herein has been shown to process data output from a nanopore sequencer to generate random numbers which pass statistical tests of randomness.

[0015] The methods disclosed herein provide a mechanism developed for extracting entropy from data generated by such macromolecule characterisation instruments, which may then be used to either seed a PRNG or as random output in its own right.

[0016] The patterning of the biological macromolecule itself (for example the sequence of individual nucleotides in DNA or RNA) has been shown to have too much inherent structure to use as a source of random numbers, leading to highly correlated bitstreams when used as a direct source of randomness even after a multiplexing unless said multiplexing involves use of an independent source of random numbers. In contrast, the inventors have found that the output read data can be useful when processed to provide truly random numbers. That is, the methods described herein use technical data derived from sequencing technologies (rather than the sequenced biological data itself), as an entropy pool, from which high-quality entropy is drawn for use in the subsequent random number generation.

[0017] Further, given that instruments characterising macromolecules may read macromolecules at a very fast rate, and often many instruments characterise macromolecules in parallel, the rate of random number generation from the methodsdisclosed herein depends at least in part on the number of parallel devices from which it can draw data. For example, a nanopore sequencer currently generates 450 base reads a second on average, with each base being read multiple times, for example around 5- 10 times. Further, given there are large numbers of such instruments being operated on a continuous basis globally, generating useful work for a range of biomedical and research applications, the entropy in these processes can be exploited to generate in parallel a source of truly random numbers, without requiring dedicated hardware that needs to be specifically designed, deployed, operated, and maintained.

[0018] In some examples, the method further comprises, by processing circuitry: dividing the extracted bitstream into a plurality of datasets, each dataset comprising a plurality of consecutive bits from the extracted bit stream; applying a statistical test of randomness to each dataset; and reconstituting the bitstream using the datasets which pass the test (and not the datasets which do not pass, or which fail, the test), wherein the determined reconstituted bitstream represents the source of entropy. In other words, a test may be applied to identify a structure within each dataset, and datasets which exhibit structure may be excluded from a reconstituted bitstream. Removing datasets which fail the test can effectively ‘clean’ the bitstream of structure, resulting in a bitstream having improved characteristics of randomness.

[0019] The datasets may be non-overlapping and / or contiguous ‘chunks’ or ‘windows’ of the dataset. The test indicative of randomness may for example comprise a Chi-Square (chi-sq) test. While passing this test is not indicative of true randomness, failing the test may be indicative of underlying structure in the data. Therefore, removing datasets which fail the test can effectively ‘clean’ the bitstream.

[0020] The size of each dataset may be predetermined, noting that larger window size may tend to increase the likelihood of neighbouring structured sequences which may pass the chi-sq test, thereby potentially hiding structured sequences. Therefore, the size of the dataset may be selected to be sufficiently small to expose structured portions of the bitstream while not imposing undue processing requirements on performance of the host system or causing an undue level of data rejection. This value (i.e. the size of the data set) may be tuned, for example for a given system, so as to result in an appropriate balance between any of accuracy, data loss and system requirements. For example, the effects of differing window sizes on the relative entropy in the output bitstreams may be modelled. Such modelling would allow the determination of a single “best” window size for a given input data type.

[0021] In some examples, such tests may be carried out in-line with signal output, such that data is tested substantially as it is produced and / or while a sample is being read.

[0022] Moreover, it may be noted that extracting bits based on their position within the output, and / or the process of ‘cleaning’ the data, may obfuscate any underlying data indicative of the sample under test. There may be a desire to keep this data confidential, for example where this sample is taken from a human subject.

[0023] The method may further comprise, by processing circuitry, applying at least one further statistical test of randomness to the bitstream or the reconstituted bitstream. In some examples, a plurality of reconstituted bitstreams may be concatenated before application of such tests. For example, these may comprise reconstituted bitstreams relating to the same bit position produced by different apparatus. While in principle, bitstreams relating to different bit positions may be concatenated, however in some cases this may be avoided, in particular if bitstreams relating to different bit positions are, or could be, correlated. The method of concatenation may be naive (for example, end to end addition of bitstreams), interleaved, or shuffled.

[0024] Such statistical tests may be applied to the whole or to part of the reconstituted bitstream. Examples of such tests / test batteries include any or any combination of: FIPS 140-2 (rngtest), ent, TestU01, Alphabits and Rabbit. This can increase the confidence that the bitstream or the reconstituted bitstream has the characteristics of a random number. In some examples, the bitstream or the reconstituted bitstream may only be used as a source of entropy if such further tests are passed.

[0025] The method may further comprise performing a cryptographic function or a hash function on the reconstituted bitstream. For example, this may comprise performing a hash function on the reconstituted bitstream, wherein the hash function has an input size equal to the output size. Further, in some examples, the method may comprise determining whether collisions occur in the output. Performing such a hash function may operate either to enhance the properties of the reconstituted bitstream and / or to test the bitstream for characteristics which make it suitable for use as random data, noting a detection of collisions may discount the data from use in this manner.

[0026] In some examples, a salt (high entropy values appended to input sequences to a hash function) may be used to further diversify output and prevent collisions. Such salts may provide a high level of randomness as may be provided by a PRNG such as the Linux dev / random function, or any other high quality source or randomness. For example, the source of randomness may be that described in US11907686B1 “True random number generation based on instrument data”. Such a salt may be concatenatedwith the reconstituted bitstream. In some examples, the size of the bitstream may be equal to the size of a hash output. For example, hash function SHA-3-4096 may receive an input of salt + 4096 bits. The hash function may then be performed to produce a diversified hash, which may increase the entropy of the bitstream and provide an output number having the characteristics of a random number.

[0027] In some examples, a plurality of bitstreams may be extracted from the output signal, each bitstream comprising time ordered bits extracted from a different predetermined common position within each output signal. The bit position is indicative of the location of the bit within a signal. For example, the signal may be a 16 bit signal, having bits in position 1 to 16, wherein position 1 is the first bit of the signal, position 2 is the second bit of signal, and so on. Such extracted bitstreams may be divided into a plurality of datasets, each dataset comprising a plurality of consecutive bits from the extracted bit stream; and a statistical test of randomness may be applied to each dataset. In some examples, each bitstream is reconstructed using the datasets of that bitstream which pass the test. In some examples, the datasets may be converted into bytes before testing / hashing. In other words, a first bitstream may be extracted based on or comprising the bits in position x in the output signal and processed to provide a first reconstituted bitstream, a second bitstream may be extracted based on or comprising the bits in position y in the output signal and processed to provide a second reconstituted bitstream, and so on. In some examples, only a subset of the bit positions may be used to provide bitstreams. This may be because, for example in previous testing of an output signal from a particular apparatus, some bit positions have been shown to produce bitstreams which exhibit high amounts of structure and / or lack qualities which make them well suited for use in generating data having the characteristics of random numbers. Such bit positions may be ignored when generating bitstreams for use in methods as set out herein.

[0028] In some examples, the method further comprises sending the extracted bitstream or data derived therefrom (e.g. a reconstituted bitstream, or a hashed reconstituted bitstream) to a recipient responsive to a request for a random number. This may comprise a convenient source of random data from an apparatus which is performing a useful function in characterising molecules.

[0029] The output data representative of a random number (e.g. having at least some characteristics of a random number and / or having passed tests intended to identify and reject structures within candidate data) may be used in cryptography, secure authentication, secure communications, lotteries, e-voting, statistical simulations,random sampling, randomized algorithms, or as a seed for a pseudo random number generator. In some examples, the method may further comprise using the output data in data representative of a random number in one or more of such processes.

[0030] In some examples, the method further comprises acquiring data to form the output signals using the instrument characterising macromolecules. In some examples the method further comprises characterising macromolecules using the instrument.

[0031] In some examples, the method may be operated in parallel, with multiple apparatus analysing samples and providing a plurality of data sources. The method may be carried out by processing circuitry comprising one or more processors. In some examples, some method steps are carried out by different computing apparatus / processing circuitry / processors from other method steps, although in other examples the method steps may be executed by the same computing apparatus / processing circuitry / processor(s).

[0032] According to a second aspect, there is provided an apparatus for generating data having characteristics associated with random numbers, comprising processing circuitry; and a memory storing instructions that, when executed by the processor, cause the processing circuitry to carry out the method of the first aspect of the invention. In some examples the apparatus further comprises at least one instrument for characterising macromolecules (e.g. a sequencer or the like).

[0033] According to a third aspect, there is provided a plurality of the apparatuses as described in the second aspect which are communicatively coupled via a network; wherein each apparatus further comprises instructions that further configure each apparatus to generate random numbers from data output from an instrument characterising macromolecules in a sample received at that computing apparatus; and cause the plurality of apparatuses to operate together to provide a randomness pool sending random numbers responsive to requests from computing apparatus.

[0034] According to a fourth aspect, there is provided a non-transitory computer- readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to carry out the method of the first aspect of the invention. In some examples, the computer-readable storage medium further includes instructions that when executed by a computer, cause the computer to control at least one instrument for characterising macromolecules to provide output signals indicative of measurements of macromolecules.

[0035] Further features of the aspects and features of the invention are set out below and in the accompanying claims.

[0036] US11907686B1 is incorporated herein by reference to the fullest extent permissible. “Extracting Randomness from Nucleotide Sequencers for use in a Decentralised Randomness Beacon” by Hurley-Smith et al, which is part of the proceedings of ARES 2024, July 30-August 2, 2024, Vienna, Austria provides additional information in relation to the concepts set out herein, and is incorporated herein by reference to the fullest extent permissible. BRIEF DESCRIPTION OF THE FIGURES

[0037] The Figures depict various embodiments for purposes of illustration only. Alternative embodiments of the apparatus and methods illustrated herein may be employed without departing from the principles described herein.

[0038] FIG.1 shows an example of apparatus for generating data having at least some characteristics of a random number;

[0039] FIG.2 shows an example of output signals;

[0040] FIG. 3 shows an example method for generating data having at least some characteristics of random numbers;

[0041] FIG. 4 shows an example method for distributing data having at least some characteristics of random numbers; and

[0042] FIG.5 shows an example of a machine readable medium in communication with a processor. DETAILED DESCRIPTION

[0043] FIG.1 shows an example of apparatus 100 for generating random numbers (or data having characteristics thereof) in accordance with aspects of the present disclosure.

[0044] The apparatus 100 comprises a computing apparatus 110 for extracting a bit stream for generating data having characteristics associated with random numbers. The computing apparatus 110 is shown as a laptop computer in this example, but other types of computing apparatuses may be used in other examples.

[0045] In this example, the apparatus 100 further comprises an instrument 102 for characterising macromolecules in a sample, communicatively coupled to provide data tothe computing apparatus 110 for extracting random numbers, or data having characteristics thereof. A macromolecule is a large, complex, molecule typically composed of a long chain of smaller molecular subunits (‘monomers’), covalently bonded together. These subunits can be identical or different and are repeated in a repeating pattern throughout the macromolecule's structure. The four primary types of macromolecules found in living organisms include proteins, nucleic acids, carbohydrates and lipids.

[0046] The instrument 102 generates data including measurement event information relating to measurements of individual macromolecules recorded over time. These events are output as output signals. To that end, the instrument 102 includes sensors and operates in a physical environment in which the sensors capture measurement events in the form of reads characterising macromolecules from a sample in the physical environment. The sensors and / or the sensor read apparatus may introduce a degree of noise. Therefore, the instrument 102, in the act of measuring the macromolecules in the sample in order to characterise them, provides a source of entropy from the physical randomness arising from the interaction between the macromolecules and the physical environment in which the macromolecules may be measured by the instrument, and / or from the measurement process.

[0047] In the example described herein in relation to FIG. 1, the instrument 102 is a sequencer for extracting information from a sample 104 of macromolecules, in the example taken from a biological organism, for example a human or animal subject. The sample 104 taken from the biological organism is contained in a sample tube 106 containing polynucleotide strands from the biological organism. That is, the sample tube 106 may contain a sample 104 of the DNA or RNA of the biological organism suitably prepared for sequencing by the instrument 102.

[0048] In some examples, the instrument 102 comprises a third generation nanopore sequencer such as those available from Oxford Nanopore Technologies plc of Oxford, UK (https: / / nanoporetech.com / ). However, the data used in the computing apparatus 110 for generating random numbers can come from any suitable source and is not limited to this kind of instrument. For example, the computing apparatus 110 can use data output from other types of sequencing technologies such as next generation sequencing technologies e.g. massively parallel sequencing by synthesis or Single-Molecule Real- Time (SMRT) sequencing. Further, any other type of instrument that characterise macromolecules in sequence, that generate data containing measurement event information relating to measurements of individual macromolecules recorded over timethat is representative of a source of physical entropy in the instrument, is suitable for use in conjunction with the computing apparatus 110 to extract random numbers. For example, instruments for sequencing genomes or transcriptomes are suitable, and instruments for characterising proteomes, such as mass spectrometers, may also be suitable.

[0049] In the instrument 102 shown in FIG. 1, which may be a nanopore sequencer instrument, a transmembrane pore 108 (e.g. a nanopore) is used as an electrical biosensor for sensing genetic information in the form of the polynucleotides in a sequence in strands of DNA or RNA from the biological sample contained in the sample tube 106. Transmembrane pores, such as transmembrane pore 108, can be used to identify small molecules or folded proteins and to monitor chemical or enzymatic reactions at approximately the single molecule level by means of sending an ion flow across the transmembrane pore 108, for example, as the strand of DNA / RNA passes through the transmembrane pore 108. Interaction of an analyte with the transmembrane pore 108 can give rise to a characteristic change in ion flow (for example, a characteristic current profile) as the analyte translocates through the transmembrane pore 108. That is, the ion flow (for example, electron flow / current) through a transmembrane pore 108 may be measured under a potential difference applied across the transmembrane pore 108.

[0050] The instrument 102, which may comprise a nanopore sequencer instrument, may comprise an array of nanopores, each under potentially different thermal conditions, electro-osmotic pressure conditions, or electronic resistance conditions. Furthermore, the duration for which a nanopore remains functional in any device is variable. Moreover, electronic components are subject to noise. These are all factors that may contribute hypothetically random elements to the raw signal output produced by such devices.

[0051] It will be appreciated that the apparatus 100 of FIG.1 may not be to scale. For example, instrument 102, which may comprise a nanopore sequencer instrument, may be considerably smaller than laptop computers. In other examples, a plurality of nanopore sequencer instruments, such as a plurality of instrument 102, may be networked into an array, which may be larger than a laptop computer. Moreover, while in this example the instrument 102, which may be a nanopore sequencer instrument, and the computing apparatus 110 are separate components, in other examples, one may be integral to the other. Moreover, the computing apparatus 110 may not be in direct communication with the sequencer instrument 102, and may receive signals from remote sequencer instruments.

[0052] In use of the apparatus 100, electrical signal data is output from the instrument 102, for example current data, which may be in the picoamp (pA) range, and is provided to the computing apparatus 110. This data is noisy and varies in its level in a way that is characteristic of the DNA nucleobases transiting through the transmembrane pore 108 at that time. For example, as a polynucleotide strand such as DNA passes through the transmembrane pore 108, the nucleobases of the DNA (i.e. adenine (A), cytosine (C), guanine (G), and thymine (T)) segment that passes through the transmembrane pore 108 produce a resultant characteristic current profile depending on which of the nucleobases is passing through the sequencer at any given moment. This produces a current signal having a level that indicates with the sequence of nucleobases passing through the transmembrane pore 108. A computing apparatus (which may be the computing apparatus 110 or a different computing apparatus) may operate on the raw current data output using an often computationally-intensive so-called ‘base calling’ application that typically uses a trained deep neural network to extract the progressing sequence of called bases.

[0053] While this signal varies as a result of the substance under analysis, the signal is also subject to noise, for example due to thermal, micro-fluidic phenomena, electro- osmotic pressure, or electronic resistance conditions, jitter in the current read, noise due to the components themselves (e.g. ASIC noise), and the like.

[0054] The electrical signal data may be output to the computing apparatus 110 as a series of numbers, which are shown in FIG.2 as 16 bit numbers (but may in principle be any length). The signals may be indexed with an index i, based on the order on which they are received or generated. In some examples, signals may be collected at regular intervals, while in other examples an idle time between successive signals or ‘reads’ may vary.

[0055] In some examples described herein (for example with reference to FIG.3 below), the output signals are structured into bitstreams, wherein each bit of a signal may contribute to a bitstream based on its position within the read signal. Thus, bits in position 1 may contribute to a first bitstream, wherein the bits in the first bitstream are ordered in the order received / generated, bits in position 2 may contribute to a second bitstream, wherein the bits in the second bitstream are ordered in the order received / generated, and so on. FIG. 2 highlights a particular bitstream associated with bits in position 7, wherein the bits in the seventh bit stream are ordered in the order in which those bits were received or generated.

[0056] FIG.3 shows a method, which may comprise a method of generating data having characteristics associated with random numbers. The method may for example be carried out by processing circuitry or computing apparatus, for example the computing apparatus 110 of FIG. 1.

[0057] In block 302, a plurality of output signals, or ‘reads’, are received from an instrument characterising macromolecules in a sample. The output signals comprise a series of time ordered output signals indicative of measurements of macromolecules recorded over time. In this example, each signal or read is an n-bit number, and are used to populate a one dimensional array of n-bit numbers. In examples herein, and as shown in FIG.2, n=16 but this need not be the case in all examples.

[0058] In block 304, this one-dimensional array is split into a two-dimensional array with n columns, one for each bit position. In some examples, not all bit positions may provide an associated bitstream, as set out in greater detail below.

[0059] In examples herein, each of the bitstreams may be tested, in whole and / or in part(s), to determine if each individual bitstream may have attributes which make is suitable as a source of a random number. Moreover, in some examples, the bitstream may be ‘cleaned’ by testing parts thereof and removing the parts which exhibit poor characteristics.

[0060] In the example of FIG. 3, each bitstream is divided into datasets, or ‘windows’ (block 306). For example, the datasets may comprise 1024, or 4096 bit signals, although in principle other datasets sizes may be used. This provides ‘chunks’ of each bitstream, and a statistic indicative of structure, or conversely a lack thereof, may be generated therefrom. The datasets in this example are non-overlapping windows. Each dataset may be constructed from bits taken from consecutive signals in the time ordered series of output signals, and preserve the time order of the time ordered series.

[0061] A statistical test is applied to each dataset in block 308 to identify any underlying structure. In this example, the Chi-square (hereinafter, chi-sq) test is used, but other tests may be used in other examples. This is a statistical test used to examine the differences between variables from a sample in order to judge the goodness of fit between expected and observed results. Failure of the chi-sq test indicates that a bit-level bias is present (i.e. too many 1s or 0s). In some examples, smaller windows may be preferred as larger window size may tend to increase the likelihood of neighbouring structured sequences which may pass the chi-square test. The size of the window or dataset may therefore vary for given circumstances and may be tailored to particular apparatus / data processingspecifications and the like. In some examples, tests may be carried out to determine an appropriate window size for a given apparatus and / or bitstream.

[0062] It may be noted that passing the chi-sq test is not an indication that the bitstream is random, but failing the chi-sq test is an indication that the bitstream is not. Technically, rather than testing for randomness, tests are carried out to identify structure, and if structure is identified, it may be determined that the data under test is not random.

[0063] In this example it is determined if a dataset passes the test in block 310. For example, a data set may pass the test if it provides a chi-sq P value of 0.01<x<0.99. This range may be different in other examples but this is a sensitive test with a large response to any skew in data. Thus, in some examples, there may be a pass range of 0.05<x<0.95, or wider. If the test is passed, the dataset is retained in block 312. Otherwise, the dataset is discarded in block 314. In block 316, a bitstream is reconstituted using only the retained datasets. In some examples, the retained datasets may be converted to bytes prior to the bitstreams being reconstituted and / or prior to the testing described herein after.

[0064] It may be the case that, for some bitstreams, no or few datasets pass the test. For example, it may be the case that the higher and / or lower bit positions may be more subject to bias than the mid-range of bit positions. This may for example be because one or more such bits are indicative of a status of the device, which may be static or change slowly. In some examples, particular bit positions may be part of a header, and these may tend to be at the beginning of the signal. In other examples, at least one bit (which may be at an end of the signal) may be indicative of a reading which is outside of a typical operational range (for example, the measurement capability may be oversized for a typical application, in case of an unusual event). Such bits may rarely change during use, and therefore be low in entropy.

[0065] In a particular example using Oxford Nanopore Technologies MinIon apparatus, which produces a 16-bit output signal indicative of pA current reading, bitstreams 8 to 16 were effectively discarded in this process. However, the remaining bit streams may be considered to be effectively ‘cleaned’ by this process by removal of data chunks which fail the chi-square test. Thus, in some examples, only a subset of the available bitstreams may be subjected to ‘cleaning’ in this manner. For example, when using Oxford Nanopore Technologies MinIon apparatus, bitstreams 8-16 may be disregarded.

[0066] As is further set out below, other bitstreams may be disregarded for other reasons, for example if they consistently fail the tests described in relation to block 318.

[0067] In block 318, further tests / test batteries may then be applied to at least a subsection of the remaining parts of aggregated bit streams. For example, the tests may comprise any or any combination of: FIPS 140-2 (rngtest), ent, TestU01, Alphabits and Rabbit. In some examples, a plurality of bitstreams, which may be for the same bit position but from different apparatus may be concatenated and tested.

[0068] While it could be technically possible to combine bitstreams from different bit positions, this may be avoided in some cases. For example, bitstreams from different bit positions from a signal from the same source (e.g. the same 16 bit output signal) are likely to be correlated. For example, in the case of Oxford Nanopore Technologies MinIon apparatus, which produces a 16-bit output signal, as is discussed in detail herein after, bits 5, 6, and 7 have been shown to have repeating patterns equivalent to the source signal values. The concatenation may be naïve (end to end), or make use of techniques such as interleaving or shuffling.

[0069] Federal Information Processing Standard (FTPS) 140-2 is insufficient to prove randomness (https: / / ieeexplore.ieee.org / document / 9069949) but is another lightweight, fast method of testing attributes of sequences. In this case, it provides both tests of distribution and tests of independence.

[0070] Ent is a pseudorandom number sequence test program, which performs a variety of tests on a stream of input data. It estimates entropy, optimal compression, chi-square distribution program, the arithmetic mean, Monte Carlo value for Pi and Serial Correlation Coefficient. It will be noted that a chi-sq test may therefore be repeated, but in this case over a larger chunk or over the whole stream, as opposed to the smaller datasets as described above.

[0071] TestU01 is a more robust library of statistical tests. Alphabits and Rabbit are lightweight (but more rigorous than FIPS 140-2 and ent) tests of randomness. Such batteries or libraries of tests may be performed over all or a subset (e.g. the first 1x106) of bits of each aggregated stream. Limiting the test to a subset of the available bits reduces processing requirements without unduly sacrificing accuracy. In some examples, reconstituted bitstreams may be hashed, for example with an input size equal to the hash output size. This generates new bitstreams of equal size to the originals, but run through sha3 with a 256-bit block size. If a collision occurs, this is indicative of failure of a randomness test. In some examples, the hashed bitstreams may provide the data having the characteristics of a random number.

[0072] While the blocks of FIG.3 are shown in a particular order in this example, in other examples they may be carried out in a different order. In some examples, the blocks 308-314 may be omitted from the method.

[0073] FIG.4 shows an example of how the data may be utilised. For example, in block 402, it may be determined if the reconstituted bitstream has passed the statistical tests of randomness described in relation to block 310. If not, in block 404, the bitstream is disregarded but otherwise the bitstream is retained (block 406). In block 408, a request for a random number is received and in block 410, the reconstituted bitstream or data derived therefrom (e.g. a hashed version thereof) is sent to a recipient. This may be used by the recipient in a number of ways, for example in cryptography, secure authentication, secure communications, lotteries, statistical simulations, random sampling, randomized algorithms, or as a seed for a pseudo random number generator. In other examples, the reconstituted bitstream or data derived therefrom may be used by the processing circuitry or computing apparatus which performs the method of FIG.3 FIG. 5 shows an example of a machine readable medium 500 (MRM) in conjunction with a processor 502. The MRM 500 stores instructions which cause the processor to carry out one or more tasks. In particular, the MRM 500 stores instructions which cause the processor 502 to, on receipt of data output from an instrument characterising macromolecules in a sample, the data comprising a series of time ordered output signals indicative of measurements of macromolecules recorded over time, extract, from the output signals, a bitstream, the bitstream comprising time ordered bits extracted from a predetermined common position within each output signal. The MRM 500 stores instructions which cause the processor 502 to use the extracted bitstream as a source of entropy for generating data having characteristics associated with random numbers. The MRM 500 may further store instructions which cause the processor to carry out any of the blocks of FIG. 3 or FIG. 4, or the methods described above in association therewith. Moreover, the MRM 500 may store instructions which cause the processor 502 to act as the computing apparatus 110 of FIG.1, or as a controller of the instrument 102.EXAMPLES

[0074] There follows some examples of results produced using the PromethION Oxford Nanopore Technologies instruments on mouse WGS data.1024-bit datasets were used.

[0075] Table 1 demonstrates that this files output, after removal of 1024-bit datasets which failed a chi-sq test, have a ‘pass’ rate of 0.99 for bitstreams constituted from bit positions 5,6 and 7 in the 16-bit output signals. In other words, approximately 99 out of 100 of these bit streams, once ‘cleaned’, passed the FIPS tests. Bit positions 6 and 7 produced orders of magnitude less viable bits than bit 5, indicating that bit 5 may be a least significant bit for an integer or float representation in the raw signal data.a e s ows n o a sreams , , an aso perorm we or and byte level analyses.

[0077] However, bit 5 is highly biased at the byte level, potentially indicating uneven distribution of 1s and 0s throughout that bitstream. This is reinforced by the high, but not critical, bit level results. Bit level analyses are indicated by the prefix “ent_b...”. While in this example both, tests were carried out both on the bit sequences, and on the bit sequences arranged into bytes, this need not be the case in all examples, and bit level testing is sufficient. The byte level testing may provide an additional degree of assurance that the data lacks structure.

[0078] One can also observe the impact of using sha3 over the bitstreams (see the file names having a _sha3 notation within the name). The significant improvement in statistical characteristics is mirrored by an improvement in pass rate for all bitstreams forFIPS as well as ent. However, this is expected behaviour and does not reduce the possibility of hash collisions if the provided data is sufficiently biased. As a result, it may be preferable to only use high-quality bitstreams as input to hashing functions. In some examples, a salt (high entropy values appended to input sequences to a hash function) may be used to further diversify output and prevent collisions. Bits 5, 6, and 7 fail the Alpha and Rabbit batteries of TestU01. They fail HammingCorrelation, Randomwalk, and Monobits tests, indicating underlying patterns over periods of 1024 and 100016 bits. Cleaning bitstreams further by hashing as previous described results in sequences that pass these more robust tests of randomness. It should be noted that collisions may be indicative that a bit-stream is unusable for cryptographic or high-trust applications. No collisions were evident for this sample.

[0079] Any logic or application described herein that comprises software or code can be embodied in any non-transitory machine- readable or computer-readable medium for use by or in connection with an instruction execution system (e.g. one or more processors) in a computer system or other system. In this sense, the logic may comprise, for example, statements including instructions and declarations that can be fetched from the computer-readable medium and can be executed by the instruction execution system. In the context of the present disclosure, a machine-readable medium or computer-readable medium can be any medium that can contain, store, or update the logic or application described herein for use by or in connection with the instruction execution system. For example, the machine-readable medium or computer-readable medium may comprise one or more of random-access memory (RAM), read-only memory (ROM), hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as compact disc (CD) or digital versatile disc (DVD), magnetic tape, and other memory components. For example, the RAM may comprise one or more of static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), and other forms of RAM. For example, the ROM may comprise one or more of programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and other forms of ROM.

[0080] While the above representative embodiments have been described with certain components in exemplary configurations, it will be understood by one of ordinary skill in the art that other representative embodiments can be implemented using different configurations and / or different components. For example, it will be understood by one ofordinary skill in the art that the order of certain steps and certain components can be altered without substantially impairing the functioning of the invention.

[0081] The representative embodiments and disclosed subject matter, which have been described in detail herein, have been presented by way of example and illustration and not by way of limitation. It will be understood by those skilled in the art that various changes may be made in the form and details of the described embodiments resulting in equivalent embodiments that remain within the scope of the invention. It is intended, therefore, that the subject matter in the above description shall be interpreted as illustrative and shall not be interpreted in a limiting sense.

Claims

Claims:

1. A method comprising, by processing circuitry, receiving data output from an instrument characterising macromolecules in a sample, the data comprising a series of time ordered output signals indicative of measurements of macromolecules recorded over time; extracting, from the output signals, a bitstream, the bitstream comprising time ordered bits extracted from a predetermined common position within each output signal; and using the extracted bitstream as a source of entropy for generating data having characteristics associated with random numbers.

2. The method of claim 1 further comprising, by processing circuitry: dividing the extracted bitstream into a plurality of datasets, each dataset comprising a plurality of consecutive bits from the extracted bitstream; applying a statistical test of randomness to each dataset; and reconstituting the bitstream using the datasets which pass the test, wherein the reconstituted bitstream provides the data having characteristics associated with random numbers.

3. The method of claim 2 further comprising, by processing circuitry: applying at least one further statistical test of randomness to the reconstituted bitstream.

4. The method of claim 2 further comprising, by processing circuitry: performing a cryptographic function or a hash function on the reconstituted bitstream.

5. The method of claim 4 comprising performing a hash function on the reconstituted bitstream, wherein the hash function has an input size equal to an output size.

6. The method of claim 4 further comprising, by processing circuitry:determining whether collisions occur in the output of the hash function.

7. The method of claim 1 further comprising extracting, by processing circuitry, a plurality of bitstreams from the output signal, each bitstream comprising time ordered bits extracted from a different predetermined common position within each output signal; dividing each extracted bitstreams into a plurality of datasets, each dataset comprising a plurality of consecutive bits from the extracted bitstream; applying a statistical test of randomness to each dataset; and reconstituting each bitstream using the datasets of that bitstream which pass the test.

8. The method of claim 1, wherein the received data output from the instrument is genomics data from a sequencer characterising DNA, RNA, or proteomics data output from an instrument characterising biopolymers in a biological sample.

9. The method of claim 1, wherein the received data output from the instrument is data representative of a current measured across a nanopore as a nucleic acid sequence passes through the nanopore.

10. The method of claim 1, further comprising sending, by processing circuitry, the extracted bitstream or data derived therefrom to a recipient responsive to a request for a random number.

11. The method of claim 1, comprising using, by processing circuitry, the output data representative of a random number in cryptography, secure authentication, secure communications, lotteries, statistical simulations, random sampling, randomized algorithms, or as a seed for a pseudo random number generator.

12. The method of claim 1, wherein the source of entropy is physical randomness arising from interactions between the macromolecules and an environment in which the macromolecules are measured by the instrument.

13. The method of claim 1, wherein the source of entropy is physical randomness arising from fluid dynamics measured by observing a passage of the macromolecules through a microfluidic structure during measurement by the instrument.

14. An apparatus for generating data having characteristics associated with random numbers, comprising: processing circuitry; and a memory storing instructions that, when executed by the processing circuitry, cause the processing circuitry to carry out a method comprising: receiving data output from an instrument characterising macromolecules in a sample, the data comprising a series of time ordered output signals indicative of measurements of macromolecules recorded over time; extracting, from the output signals, a bitstream, the bitstream comprising time ordered bits extracted from a predetermined common position within each output signal; and using the extracted bitstream as a source of entropy for generating data having characteristics associated with random numbers.

15. A plurality of apparatuses, where each apparatus in the plurality of apparatuses is as claimed in claim 14, communicatively coupled via a network; wherein the method further comprises: generate random numbers from data output from an instrument characterising macromolecules in a sample received at that computing apparatus; and cause the plurality of apparatuses to operate together to provide a randomness pool sending random numbers responsive to requests from computing apparatus.

16. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to carry out a method comprising: receiving data output from an instrument characterising macromolecules in a sample, the data comprising a series of time ordered output signals indicative of measurements of macromolecules recorded over time; extracting, from the output signals, a bitstream, the bitstream comprising time ordered bits extracted from a predetermined common position within each output signal; and using the extracted bitstream as a source of entropy for generating data having characteristics associated with random numbers.