Framework identical pseudo random number generation for compiled ai and scientific tensor computation graphs

By storing and forwarding generator states within compiled libraries, the method ensures identical pseudo random number generation across AI frameworks, addressing reproducibility and performance issues.

WO2025133717A1PCT designated stage expired Publication Date: 2025-06-26NEC LAB EURO GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/054073
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-04-26
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing AI and scientific compute frameworks face challenges in providing identical pseudo random number generation across different frameworks and deployed libraries, leading to issues with reproducibility and performance.

Method used

The method involves storing a generator state as a static object within a compiled library or runtime library, scheduling random number generations using this state, and forwarding the generator state based on the number of random numbers drawn from previous layers, ensuring identical random number generation across frameworks.

Benefits of technology

This approach guarantees identical random numbers across different frameworks and hardware platforms, enhancing reproducibility and performance by decoupling pseudo random number generator states from frameworks and enabling parallel computation of random layers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024054073_26062025_PF_FP_ABST
    Figure IB2024054073_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for providing for artificial intelligence (AI) framework identical pseudo random number generation includes storing a generator state as a static object within a compiled library having a plurality of layers, or in a runtime library associated with the compiled library, the generator state corresponding to a random state associated with a pseudo random number generation algorithm. Random number generations are scheduled using in each case the generator state, wherein the generator state is forwarded by an amount of random numbers drawn from previous layers of the compiled library executed prior to a current layer. The method has applications including, but not limited to, use cases in machine learning, computational biology, medical AI and healthcare, chemistry, physics, electrical or mechanical engineering.
Need to check novelty before this filing date? Find Prior Art

Description

FRAMEWORK IDENTICAL PSEUDO RANDOM NUMBER GENERATION FORCOMPILED Al AND SCIENTIFIC TENSOR COMPUTATION GRAPHSCROSS-REFERENCE TO PRIOR APPLICATION

[0001] Priority is claimed to U.S. Provisional Application Serial No. 63 / 612,466 filed on December 20, 2023, the entire contents of which is hereby incorporated by reference herein. FIELD

[0002] The present disclosure relates to Artificial Intelligence (Al) and machine learning, and in particular to a method, system, data structure, computer program product and computer- readable medium for ensuring identical random or pseudo random number generation for computation graphs across different frameworks.BACKGROUND

[0003] Existing Al and scientific compute frameworks provide compilers (e.g., 'torch.compile(...)' or 'keras.compile(..., jit_compile=True)') to optimize the execution of the underlying models. The frameworks are facing the technical problem that they need to provide identical random numbers when being compiled or not compiled. State of the art (STAR) frameworks either don't use compilation for random layers at all and instead fallback to their default runtime kernel implementation or use a different algorithm / seeding mechanism and warn the user about this behavior (either through runtime warning message or within documentation). SUMMARY

[0004] In an embodiment, the present disclosure provides a computer-implemented method for providing for artificial intelligence (Al) framework identical pseudo random number generation. The method includes storing a generator state as a static object within a compiled library having a plurality of layers, or in a runtime library associated with the compiled library, the generator state corresponding to a random state associated with a pseudo random number generation algorithm. Random number generations are scheduled using in each case the generator state, wherein the generator state is forwarded by an amount of random numbers drawn from previous layers of the compiled library executed prior to a current layer. The method has applications including, but not limited to, use cases in machine learning, computational biology, medical Al and healthcare, chemistry, physics, electrical or mechanical engineering.BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Embodiments of the present disclosure will be described in even greater detail below based on the exemplary figures. The present disclosure is not limited to the exemplary embodiments. All features described and / or illustrated herein can be used alone or combined indifferent combinations in embodiments of the present disclosure. The features and advantages of various embodiments of the present disclosure will become apparent by reading the following detailed description with reference to the attached drawings which illustrate the following:

[0006] FIG. 1 illustrates an example of a model that executes directly within the framework;

[0007] FIG. 2 illustrates the same example, executing the STAR using the framework's random kernels and only compiling pure computational layers;

[0008] FIG. 3 illustrates the same example using a method according to an embodiment of the present disclosure, that places the Generator object within the compiled module;

[0009] FIG. 4 illustrates the same example using a method according to an embodiment of the present disclosure, that uses a shared runtime library to share Generators across different compiled modules;

[0010] FIG. 5 illustrates an example of three random layers that need to be executed sequentially because of control dependencies of the shared Generator object;

[0011] FIG. 6 illustrates a parallelized version of three random layers depending on one shared Generator object; and

[0012] FIG. 7 is a block diagram of an exemplary processing system, which can be configured to perform any and all operations disclosed herein.DETAILED DESCRIPTION

[0013] Pseudo random number generators play an important role in Al training and other scientific computations as layers such as Random, (Alpha-)Dropout, Bernoulli, RReLU, and others (in the following referred to as “random layers”). Al frameworks use well known pseudo random number generators (e.g., Philox or MT19937) which provide enough randomness while being easy to implement and efficient to compute on various hardware platforms. However, a technical problem is that these generators cannot guarantee reproducibility across versions or platforms. Embodiments of the present disclosure provide solutions to this technical problem by a method and system that guarantees identical random numbers across different frameworks and deployed Al libraries.

[0014] In a first aspect, the present disclosure provides a computer-implemented method for providing for artificial intelligence (Al) framework identical pseudo random number generation. The method includes storing a generator state as a static object within a compiled library having a plurality of layers, or in a runtime library associated with the compiled library, the generator state corresponding to a random state associated with a pseudo random number generation algorithm. Random number generations are scheduled using in each case the generator state, wherein the generator state is forwarded by an amount of random numbers drawn from previous layers of the compiled library executed prior to a current layer.

[0015] In a second aspect, the present disclosure provides the computer-implemented method according to the first aspect, wherein scheduling the random number generations includes creating a clone of the generator state and forwarding the clone of the generator state by the number of random numbers drawn from the previous layers executed prior to the current layer.

[0016] In a third aspect, the present disclosure provides the computer-implemented method according to the first or second aspect, further comprising: generating a plurality of clones of the generator state; and transmitting each clone of the plurality of clones to a separate core / device, wherein the random number generations are performed in parallel by each separate core / device executing one of the layers.

[0017] In a fourth aspect, the present disclosure provides the computer-implemented method according to the first to third aspects, further comprising maintaining, by each separate core / device, a map for the generator state of the corresponding clone, wherein the map identifies a mapping between a device type and device identifier for the device to a device owned generator state.

[0018] In a fifth aspect, the present disclosure provides the computer-implemented method according to the first to fourth aspects, wherein each separate core / device implements the same pseudo random number generation algorithm.

[0019] In a sixth aspect, the present disclosure provides the computer-implemented method according to the first to fifth aspects, wherein each separate core / device sends the amount of random numbers to another core / device that executes a subsequent one of the layers for forwarding the clone of the generator state, and wherein the core / device that executes a final one of the layers updates the generator state.

[0020] In a seventh aspect, the present disclosure provides the computer-implemented method according to the first to sixth aspects, wherein an Al framework using the method for the random number generations is decoupled from the compiled library, and wherein the Al framework is designed to use the random number generations in a medical or healthcare domain.

[0021] In an eighth aspect, the present disclosure provides the computer-implemented method according to the first to seventh aspects, wherein a reseeding method is attached to the Al framework.

[0022] In a ninth aspect, the present disclosure provides the computer-implemented method according to the first to eighth aspects, further comprising parsing the Al framework to identify a generator object and a local value, wherein the generator state corresponds to a generator state object that is generated for each instance of the generator object parsed from the Al framework.

[0023] In a tenth aspect, the present disclosure provides the computer-implemented method according to the first to ninth aspects, wherein the generator state is read and updated after execution of each of the layers by the amount of random numbers in a case of control flow dependencies.

[0024] In an eleventh aspect, the present disclosure provides the computer-implemented method according to the first to tenth aspects, further comprising invoking an initialization function to initialize the random state using a user defined seed value.

[0025] In a twelfth aspect, the present disclosure provides the computer-implemented method according to the first to eleventh aspects, wherein the compiled library is associated with the runtime library, which contains the generator state, and wherein the runtime library is associated with further compiled libraries used for scheduling the random number generations.

[0026] In a thirteenth aspect, the present disclosure provides the computer-implemented method according to the first to twelfth aspects, wherein a reseeding method is attached to different Al frameworks that use the method for the random number generations, and wherein the random number generations generated in the different Al frameworks are ensured to be identical by using the same generator states.

[0027] In a fourteenth aspect, the present disclosure provides a computer system for providing for artificial intelligence (Al) framework identical pseudo random number generation, the computer system comprising one or more hardware processors which, alone or in combination, are configured to provide for execution of the method according to any of the first to thirteenth aspects.

[0028] In a fifteenth aspect, the present disclosure provides tangible, non-transitory computer-readable medium having instructions thereon which, upon being executed by one or more processors, provide for providing for artificial intelligence (Al) framework identical pseudo random number generation by execution of the method according to any of the first to thirteenth aspects.

[0029] In particular, the technical problem faced by existing Al and scientific compute frameworks is not the pseudo random number generating algorithm, because such algorithms can easily be also implemented in compiled modules (which can also be referred to as compiled libraries), but rather the storage of the state object of these random generators. This state can either be global for the entire application and used by all random layers or attached to a so called 'Generator' object (e.g., 'torch.Generator' <<https: / / pytorch.org / docs / stable / generated / torch.Generator.html>> or 'tf.random.Generator' <<https: / / www.tensorflow.org / api_docs / python / tf / random / Generator>>). There are special cases, e.g., 'tf.random.uniform(...)' that use a local Generator object for every instance oftf.random.uniform(...)' within the computation graph. FIG. 1 illustrates an example of a model that executes directly within a framework. The example model of FIG. 1 depicts the generator 102 which stores the state of the random generation (e.g., consumes a seed to initialize the state). The random 100 uses the random state from generator 102 to fdl a tensor with random numbers. Input 106 may represent input provided by a user and can include video, audio, text, etc. Layers 104 may represent any layer of an Al framework but may include Bernoulli layers that check whether “layer_l[...] = input[...] > randomf...].” Output 108 represents the output of the neural network of the Al framework that is provided back to a user.

[0030] Another technical problem is that the compilers are separate projects (e.g., Triton, OpenXLA or TVM) and do not have a direct connection to the framework’s proprietary implementation of the random state objects. The framework is calling into the compilers, not the other way around, which would require a circular connectivity, which is not implemented by any of the projects. For example, in Triton, the random number generator needs a random state provided by the user himself / herself, and is not managed at all within the compiler (see <<https: / / github.com / openai / triton / blob / 440fdlbf20697b0961ddb0822de86dl51c58dd36 / python / triton / language / random.py#L16»). For example, FIG. 2 illustrates the same example as depicted in FIG. 1, executing a STAR framework using the framework's random kernels and only compiling pure computational layers. As described above, the generator 200 stores the state of the random generation and consumes a seed to initialize the state. The random 206 uses the random state from generator 200 to fdl a tensor with random numbers. The input 202 and tensor fdled with random numbers of random 206 are provided to the compiler 204. This guarantees that the random numbers stay within the Al framework and that they are identical. The nonrandom layers 208 are optimized by the compiler 204 to generate output 210. This type of split architecture guarantees that the output 210 is identical between non-compiled and compiled layers.

[0031] Therefore, the STAR is that either the frameworks don’t compile random layers at all and fallback to their normal runtime-kernel implementations or provide random numbers in the compiler’s native random algorithm, that do not match with the ones of the framework. An example implementation for such a fallback be found here: <<https: / / github.com / pytorch / pytorch / blob / e9ebda29d87ce0916ab08c06ab26fd3766a870e5 / torc h / _inductor / lowering.py#L 1032».

[0032] The following code shows the described STAR in PyTorch v2.0.0: import torch class Model ( torch . nn . Module ) : def forward ( self , x ) :return torch . rand ( *x . shape ) + x inp = torch . zeros ( 1 , 2, 3) eager = Model ( ) compiled = torch. compile (eager ) torch. manual seed (123) print (eager (inp) )# tensor ([ [ [0.2961, 0.5166, 0.2517] ,# [0.6886, 0.0740, 0.8665]] ]) torch. manual seed (123) print (compiled (inp) )# [2023-05-31 14:14:33,584] torch ._inductor . utils : [WARNING] using triton random, expect difference from eager# tensor ([ [ [0.6011, 0.4682, 0.5422] ,# [0.1464, 0.2591, 0.0154]] ])

[0033] This results in either lower performance or verification problems (which especially in the medical domain is very important), because running the identical Al training with and without the compiler will result in different trained parameter values.

[0034] A similar problem in Numpy is described here:<<https: / / stackoverflow.com / questions / 55911707 / how-do-i-share-the-random-number- generator-in-numpy-c-api».

[0035] Embodiments of the present disclosure provide solutions to the following technical problems:1. Whenever the user resets the seed ('torch.manual_seed(...)'all Generator's are to be reseeded.2. Depending on the used algorithm, the Generators can either: a. Use the seed value, provided by the user (PyTorch's default random algorithm), b. Use a local seed value specified by the program (e.g. 'tf.random.uniform(..., seed=LOCAL_SEED)' ), or c. Use both (TF's default random algorithm).

[0036] Embodiments of the present disclosure enable framework identical random number generation in compiled models and solve the technical problems 1 and 2.

[0037] It is also a technical problem that, for performance reasons, Al frameworks often use different random number algorithms on different hardware. Technically it is possible to haveidentical random number algorithms used across all possible device types, however, they might not run as performant as on other hardware architectures.Step 1 : Parsing of the Al model

[0038] Al frameworks provide methods to serialize their Al models into intermediate representations (IR) such as TorchScript or tf. Graph. These methods allow to capture attributes such as the used Generator object and the local seed (if available).

[0039] There are some special cases:In TensorFlow, if the model is executed in eager mode, calls to tf.random.uniform(...)' and similar functions differentiate the Generator object by the value of the local seed value.In TensorFlow, if the model is executed in graph mode, all calls to tf.random.uniform(...)' and similar functions use different Generator objects for every function call, independent of the local seed value (see TensorFlow’s introduction guide on random number generation at <<https: / / www.tensorflow.org / guide / random_numbers>>).There are implementation issues that currently prevent to use torch.Generatof objects within TorchScript. However, technically they are possible, it’s an implementation problem (see <<https: / / github.com / pytorch / pytorch / issues / 64005>>; «https : / / github . com / pytorch / pytorch / issues / 88574»; <<https: / / github.com / pytorch / pytorch / issues / 71398>>; and <<https: / / github.com / pytorch / pytorch / issues / 69006>>.

[0040] Embodiments of the present disclosure provide to create a unique identifier for each Generator object, independent if it is the global, a user provided, or a local / shared Generator object, obeying the above-described rules. Further, the generator might contain an initial state and a specific algorithm assigned. For example, in PyTorch the CPU by default uses MT19937, and in TensorFlow it is Philox or ThreeFry for all devices. To continue the example, PyTorch by default utilizes a global generator object. This would ensure that all random layers share an identical generator object that generates the same random state. It is important to keep the exact same execution order of the random layers so that the random state gets incremented in the correct order. In TensorFlow each random layer has its own random state. As such, each generator object would need its own unique identifier. However, TensorFlow also includes an option to set a local seed which if the local seeds are identical can cause the multiple random layers to use the same random state. Thus, the present disclosure describes situations when using a global random state, which requires ensuring correct execution order between the random layers, situations where each random layer has its own state, which requires unique identifiers for each generator object but the order of execution is irrelevant, and some amount of sharing ofthe random state which requires unique identifiers for the generator objects but maintaining the execution order of layers sharing the same state is relevant as well.Step 2: Generators within Compiled Objects

[0041] One technical problem the Al frameworks have with using random generators within compiled objects is that these need to be compatible to their Al framework's implementation. While implementing the identical algorithm is trivial and can, e.g., rely on the same source code or external libraries, the handling of the generator states requires some special handling, to be easily injectable into the Al framework.

[0042] FIG. 3 illustrates the same example as depicted in FIG. 1 but using a method according to an embodiment of the present disclosure that places the Generator object 300 within the compiled module 302. As used herein, the generator, generator object, and generator state refer to an object that contains the random state (e.g., consumes a seed to initialize the state (random state)). In embodiments the random state depends on the algorithm used as it can be a single integer number, multiple integer numbers, or an entire array of numbers. In the method according to an embodiment of the present disclosure, one static generator state object instance is generated for every Generator object 300 captured during parsing the model. The random 304 may represent a tensor filled with numbers that uses the random state from generator object 300. This generator state object gets compiled as a static object into the compiled library and can therefore be used within this compiled library by any random layer 306 that is supposed to use the specific random generator state. For example, in C / C++ this can be done by code such as: struct random algorithm state {} ; random algorithm state unique state name 1 ;

[0043] When the framework loads the compiled library, the method calls an init function, that initializes all random states with the current user defined seed value, that was defined by a previous call to 'torch.manual_seed(...)' or similar. This function is further executed whenever the user resets the seed. As all generator objects 300 within the neural network are known, it is easy to place them within the compiled library. The random state is increased whenever random numbers are generated by the random layer (e.g., 304 and 306). A simple example would be represented by init state = 0 (the initial state is zero); random (&state, 128) (this generates 128 numbers). By executing this pseudo code the random state will have increased by 128. extern random algorithm state state 1 ; extern random algorithm state state 2 ;extern random algorithm state state 3; void random set seed (const int64 t seed) { random algorithm init seed(Sstate 1, seed) ; random algorithm init seed(Sstate 2, 567) ; / / uses only local seed random algorithm init seed(Sstate 3, 123, seed) ; / / uses local and global seed }

[0044] Next, it is ensured that the function 'random_set_seed(...)' gets called. One easy implementation is to use a class wrapper that abstracts the functionality and registers itself to a global state registry. When returning the compiled model to the Al framework 310, a class such as the following can be used: class CompiledModel: seed = 0 states = set ( ) def init (self, . . . ) : self. library = load library ( ... )CompiledModel . instances . add ( self) self. set seed (CompiledModel . seed) # init to last seed def del (self) :CompiledModel . instances . remove ( self ) def set seed(self, seed) : self . library . random set seed(seed)

[0045] Further, it is possible to then modify the Al framework's 310 seeding method to use this global registry: torch manual seed = torch. manual seed def manual seed(seed) : torch manual seed (seed)CompiledModel . seed = seed for x in CompiledModel . instances : x . set seed ( seed) torch . manual seed = manual seed

[0046] This method ensures that the state is always set correctly, as long as the compiled model is loaded to generate output 312 from input 308.

[0047] FIG. 4 illustrates the same example depicted in FIG. 1 but using a method according to an embodiment of the present disclosure, that uses a shared runtime library 400 to share Generators 402, that generate random numbers, across different compiled modules 404 executing multiple layers 406 to generate output 408 using input 410 provided within the Al Framework 412. In case that Generators 402 shall be shared across models, it is possible to store these in a shared runtime library 400 that maintains the different generator states. Random state 414 would use the random state from generator 402 to fill a tensor with random numbers which can be used by layers 406. An example would be the following code: model a = torch . compile (Model ( ) ) model b = torch . compile (ModelB ( ) ) gen = torch . Generator ( ) out a = model a ( input , gen ) out b = model b ( input , gen )Step 3: Parallelization of Random Number Generations Using the Same State Generator

[0048] FIG. 5 illustrates an example of three random layers 502, 504 that need to be executed sequentially because of control flow dependencies of the shared Generator object 500. When relying on the same state generator object 500, algorithmically random layers 502, 504 need to be executed in the correct order to produce the identical sequences of numbers in every run of the model. Within a compiler, the order of these random layers 502, 504 and the exact number of random numbers to be drawn is known or can be computed at runtime in case these layers 502, 504 use variable dimensions. FIG. 5 depicts an example where the Al framework uses a global random state, generated by generator 500, or an example where multiple random layers (502 and 504) share the same random state. In these scenarios it is necessary to execute the random layers 502 and 504 in the correct order to guarantee the correct increase of the random state.

[0049] With this information (knowing the order of the random layers), embodiments of the present disclosure can temporarily create a clone of the generator state and fast-forward through it using the count of random numbers of all layers 502, 504 that should be executed before. Asused herein fast-forwarding may refer to skipping the random state without actually generating numbers. For example, in a scenario such as: init state = 0; random (&state, 128) the state would be 128. Now, in order to get the same result the random layer is not required, instead the pseudo code would be init state = 0; state += 128; which would result in the state being 128 and one would get the same result. However, depending on the random number algorithm this can be more complicated than simply adding 128. For example, Philox requires the increasing of multiple counters. The efficiency of this fast-forward depends on the algorithm itself, however especially for algorithms such as Philox (used in TensorFlow), this can be done very easily, as shown be the following code snippet: class Philox : def forward ( self , ent ) : ent lo , ent hi = ent & OxFFFFFFFF, ent >> 32 ent hi = ent >> 32 self . base

[0000] += ent lo if self . base

[0000] < ent lo : # checks for integer overflow ++cnt hi self . base

[0001] += ent hi ; if self . base

[0001] < ent hi : # checks for integer overflow if ++self . base

[0002] == 0 :++self . base

[0003]

[0050] This allows to execute multiple random layers at the same time on different cores or devices while maintaining the identical random number generation sequence. This is especially advantageous in TensorFlow's RNN layers that use the same random generator for computing the values of its dropout for every sequence. For example, the input data for RNN layers may be required to have three dimensions: [Sequence] [Batchsize] [Channels]. The RNN layers may compute this as: def mn(input): hidden_state = 0 for sequence in input: hidden_state = mnLayer(sequence, hidden_state) return hidden state.

[0051] In scenarios where dropout is used the inner loop of the above would look like: for sequence in input: dropout = random(sequence. shape) hidden state = mnLayer(sequence, hidden State, dropout).

[0052] As such a random execution is required in every single iteration which is not optimal for performance. One solution is to do: dropouts = [] for I, sequence in enumerate(input): dropouts .append(random(sequence . shape, fast_forward=i)) for sequence, dropout in zip(input, dropouts): hidden state = mnLayer(sequence, hidden state, dropout).

[0053] The first loop above can be done in parallel because with the information of the “fast_forward=i” the system knows how many states it needs to skip to compute the correct random numbers.

[0054] As the number of elements within a sequence can be rather small (Batch Size * Channels), it is possible to parallelize this among all resources of the device, resulting in significant speedup.

[0055] At the end of the computation, care is taken that the generator 500 state is set correctly for another execution run of the model. For this, the last scheduled random layer 504 updates the generator 500 after it has been executed.

[0056] FIG. 6 illustrates a parallelized version of three random layers 600, 602, and 604 depending on one shared Generator object 606. Assuming that cloning and forwarding the random state takes significantly less time than generating the random numbers, an up to ~3x speedup can be achieved. In FIG. 6 “mimel” represents the number of elements within a tensor. Step 4: Multi -Device Support

[0057] In Al training, usually multiple computational devices are used. Each of them maintain its own seed state, which can be updated by functions such as 'torch.cuda.seed(...)' (see <<https: / / pytorch.org / docs / stable / generated / torch.cuda.seed.html>>). In an embodiment, the present disclosure only stores a single state per Generator, which is sufficient for devices where the state is stored on the same device as the computation happens (e.g. CPUs).

[0058] However, in hybrid multi-device setups (e.g. multi-GPU), each of the devices 608, 610, 612 needs to maintain its own state. As the containing library is only loaded once, but reused for every device 608, 610, and 612, the library internally needs to maintain differenceinstances of the Generators 614, 616, and 618. In embodiments the clones 614-618 of generator 606 are identical copies of generator 606. In the forward layers 620 and 622 the state is increased accordingly. As an example, the first forward layer 620 would increase the random state from generator clone 616 by a random number of elements that corresponds to random 600. The second forward layer 622 would increase the random state from generator clone 618 by a random number of elements that correspond to random 600 plus the random number of elements that correspond to random 602, etc. For this, within each Generator 614, 616, and 618, an embodiment of the present disclosure maintains a map, that maintains the device’s 608, 610, 612 random state, which can be distinguished by type or id of the device. This way, each device 608, 610, 612 maintains its own random state and does not interfere when computing in parallel. This also prevents race conditions, which could occur if one device reaches a random layer before the other. In a race condition the output depends on the execution order of the devices. So in a scenario with device 1, 2, and 3, the system has a specific sequence of random numbers. If the system needs to execute devices in an order of 1, 3, and 2, one would end up with a different sequence of random numbers. By using cloning and forwarding of the random states the system guarantees that the devices work on the correct random state and the execution is independent of the execution order of the devices. Alternatively, the race conditions could be prevented by device synchronization, but would yield in enormous performance penalties.

[0059] In an embodiment, the present disclosure can be applied for random number reproducibility across framework versions or hardware platforms. Existing Al frameworks don’t guarantee reproducibility of random numbers across different framework versions or hardware platforms. Especially in areas such as the medical domain and healthcare, it is especially important to produce identical results within the certification processes, which can take several years. Embodiments of the present disclosure allow to fix the random number generator independent of the used Al framework, to produce always identical numbers, irrespective of used framework version or hardware platform. This allows to make use of bugfixes within newer Al frameworks versions, while not risking losing the product’s certification.

[0060] In another embodiment, the present disclosure can be applied for application acceleration. Dropout, Bernoulli, RReLU, and other pseudo random layers are heavily used in Al training to prevent overfitting of Al models. However, with increasing interest of reproducibility, explainability and increasing governmental legal obligations, data scientists and Al devops are required to build their training pipelines in a reproduceable fashion, with fixed initial seed values and fixed pseudo random number generating algorithms. Especially for governmental obligations it can be that experiments need to be repeated years after the initial development of the code. On the other side, there are the ever-growing computational demandsof the computations, currently with bigger and more powerful large language models. Also, other scientific computation domains (including, but not limited to 3D graphics, biology, chemistry, mechanical engineering, physics, etc.) use Monte Carlo Sampling, which repeats simulations using different randomized starting points. Genetic algorithms are another example that employ random numbers for mutating its search population. In all these applications, performance and reproducibility are strong requirements, and can be provided for by embodiments of the present disclosure. These two demands contradict each other in current Al frameworks, where either the computation can be optimized by compiling the model, but using different / non-reproduceable random numbers, or to be reproduceable, but not both. Embodiments of the present disclosure bridge this gap.

[0061] In a further embodiment, the present disclosure can be applied for neural network deployment for edge training. Training networks on the edge is an increasing topic. Because of the limited hardware capabilities of edge hardware, it’s utterly important to be as efficient as possible. Still, the training itself should not deviate from the training process running on big servers. As randomness is an important part of training, it is a strong requirement to leverage the same framework’s random number generators. STAR is to use specialized versions of TensorFlow or libraries such as libtorch (pure C / C++-API of PyTorch), which are however based on the original frameworks, with their limited performance. The method according to an embodiment of the present disclosure is advantageously able to provide hardware optimized execution libraries, featuring identical training capabilities, but tailor-made for the target hardware, achieving peak performance.

[0062] Embodiments of the present disclosure provide improvements to computer functionality in Al frameworks in general by enabling identical random number generation across the different frameworks. Moreover, embodiments of the present disclosure can be practically applied in a number of technical fields and scientific areas / applications to effect further improvements in those fields. For example, embodiments of the present disclosure can be especially advantageously applied in the medical and health care domains (e.g., drug or vaccine development, therapies and treatments, diagnosis, etc.) to provide for reproducibility and verification.

[0063] Embodiments of the present disclosure provide an improved Al platform with enhanced computer functionality and apply to a very wide area of applications. In general, embodiments of the present disclosure are applicable to any application that can be implemented using popular frameworks such as (but not limited to) PyTorch, TensorFlow, Numpy, JAX, Scikit-Leam, from various application domains such as (but not limited to) machine learning,computational biology, medical Al and healthcare, chemistry, physics, electrical or mechanical engineering.

[0064] In an embodiment, the present disclosure provides a method for framework identical pseudo random number generation in compiled tensor computation graphs, the method comprising the steps of:1) Implement identical pseudo random number generation algorithms.2) Store generator states as static objects within a compiled library (see improvement 1 below).3) Attach a reseeding method to the framework’s seeding method (see improvement 1 below). In embodiments this guarantees that in a scenario where a user resets the seed within an Al framework that the compiler also resets the seed accordingly. Without this mechanism a user would be required to manually reset the seeds in the compiler, which is inconvenient, resource and time intensive and prone to error, and without doing this would result in wrong outputs.4) Schedule random number generations depending on the same random generator state in parallel and compute correct random generator state by fast forwarding it by the exact number of random numbers drawn in previous layers (see improvement 2 below). This is associated with the execution order of the random layers. The random layers can only produce the correct outputs when the generator state is correct. If the random layers are executed in a different order then the random state is increased differently which would result in different random numbers being output.5) Maintain a map for per-device owned generator states (see improvement 3 below). In scenarios where the system has multiple devices (e.g., GPU 1, GPU 2, GPU 3, ...) each device maintains their own copy of the random states. As GPUs don’t operate on their own but are controlled by the host system the map is provided for distinguishing between the different devices or GPUs.

[0065] Embodiments of the present disclosure provide for the following improvements and technical advantages over existing technology:1. Providing to decouple pseudo random number generator states from the frameworks by placing them into the compiled modules.2. Using random state forwarding to enable parallel computation of random layers on different cores / devices.3. Using mapping of device type and id to a device owned state to enable multi -device computations.

[0066] Embodiments of the present disclosure enable to integrate framework identical random number generators into compiled execution modules. It not only provides betterperformance through specialization and parallelization of multiple random number layers than precompiled generic random number kernels, but also adds reproducibility which is not possible in STAR frameworks.

[0067] Referring to FIG. 7, a processing system 700 can include one or more processors 702, memory 704, one or more input / output devices 706, one or more sensors 708, one or more user interfaces 710, and one or more actuators 712. Processing system 700 can be representative of each computing system disclosed herein.

[0068] Processors 702 can include one or more distinct processors, each having one or more cores. Each of the distinct processors can have the same or different structure. Processors 702 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuitry (e.g., application specific integrated circuits (ASICs)), digital signal processors (DSPs), and the like. Processors 702 can be mounted to a common substrate or to multiple different substrates.

[0069] Processors 702 are configured to perform a certain function, method, or operation (e.g., are configured to provide for performance of a function, method, or operation) at least when one of the one or more of the distinct processors is capable of performing operations embodying the function, method, or operation. Processors 702 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 704 and / or trafficking data through one or more ASICs. Processors 702, and thus processing system 700, can be configured to perform, automatically, any and all functions, methods, and operations disclosed herein. Therefore, processing system 700 can be configured to implement any of (e.g., all of) the protocols, devices, mechanisms, systems, and methods described herein.

[0070] For example, when the present disclosure states that a method or device performs task “X” (or that task “X” is performed), such a statement should be understood to disclose that processing system 700 can be configured to perform task “X”. Processing system 700 is configured to perform a function, method, or operation at least when processors 702 are configured to do the same.

[0071] Memory 704 can include volatile memory, non-volatile memory, and any other medium capable of storing data. Each of the volatile memory, non-volatile memory, and any other type of memory can include multiple different memory devices, located at multiple distinct locations and each having a different structure. Memory 704 can include remotely hosted (e.g., cloud) storage.

[0072] Examples of memory 704 include a non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any kind of optical storage disk such as a DVD, a Blu-Ray® disc, magnetic storage, holographic storage, a HDD, a SSD, any medium that can be used to store program code in the form of instructions or data structures, and the like. Any and all of the methods, functions, and operations described herein can be fully embodied in the form of tangible and / or non-transitory machine-readable code (e.g., interpretable scripts) saved in memory 704.

[0073] Input-output devices 706 can include any component for trafficking data such as ports, antennas (i.e., transceivers), printed conductive paths, and the like. Input-output devices 706 can enable wired communication via USB®, DisplayPort®, HDMI®, Ethernet, and the like. Input-output devices 706 can enable electronic, optical, magnetic, and holographic, communication with suitable memory 706. Input-output devices 706 can enable wireless communication via WiFi®, Bluetooth®, cellular (e.g., LTE®, CDMA®, GSM®, WiMax®, NFC®), GPS, and the like. Input-output devices 706 can include wired and / or wireless communication pathways.

[0074] Sensors 708 can capture physical measurements of environment and report the same to processors 702. User interface 710 can include displays, physical buttons, speakers, microphones, keyboards, and the like. Actuators 712 can enable processors 702 to control mechanical forces.

[0075] Processing system 700 can be distributed. For example, some components of processing system 700 can reside in a remote hosted network service (e.g., a cloud computing environment) while other components of processing system 700 can reside in a local computing system. Processing system 700 can have a modular design where certain modules include a plurality of the feature s / functions shown in FIG. 7. For example, I / O modules can include volatile memory and one or more processors. As another example, individual processor modules can include read-only-memory and / or local caches.

[0076] While subject matter of the present disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. Any statement made herein characterizing the invention or disclosure is also to be considered illustrative or exemplary and not restrictive as the invention is defined by the claims. It will be understood that changes and modifications may be made, by those of ordinary skill in the art, within the scope of the following claims, which may include any combination of features from different embodiments described above.

[0077] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality ofelements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that the recitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method for providing for artificial intelligence (Al) framework identical pseudo random number generation, the computer-implemented method comprising: storing a generator state as a static object within a compiled library having a plurality of layers, or in a runtime library associated with the compiled library, the generator state corresponding to a random state associated with a pseudo random number generation algorithm; and scheduling random number generations using in each case the generator state, wherein the generator state is forwarded by an amount of random numbers drawn from previous layers of the compiled library executed prior to a current layer.

2. The computer-implemented method according to claim 1, wherein scheduling the random number generations includes creating a clone of the generator state and forwarding the clone of the generator state by the number of random numbers drawn from the previous layers executed prior to the current layer.

3. The computer-implemented method according to claim 2, further comprising: generating a plurality of clones of the generator state; and transmitting each clone of the plurality of clones to a separate core / device, wherein the random number generations are performed in parallel by each separate core / device executing one of the layers.

4. The computer-implemented method according to claim 3, further comprising maintaining, by each separate core / device, a map for the generator state of the corresponding clone, wherein the map identifies a mapping between a device type and device identifier for the device to a device owned generator state.

5. The computer-implemented method according to claim 4, wherein each separate core / device implements the same pseudo random number generation algorithm.

6. The computer-implemented method according to claim 4, wherein each separate core / device sends the amount of random numbers to another core / device that executes a subsequent one of the layers for forwarding the clone of the generator state, and wherein the core / device that executes a final one of the layers updates the generator state.

7. The computer-implemented method according to any of the preceding claims, wherein an Al framework using the method for the random number generations is decoupled from the compiled library, and wherein the Al framework is designed to use the random number generations in a medical or healthcare domain.

8. The computer-implemented method according to claim 7, wherein a reseeding method is attached to the Al framework.

9. The computer-implemented method according to claim 7, further comprising parsing the Al framework to identify a generator object and a local value, wherein the generator state corresponds to a generator state object that is generated for each instance of the generator object parsed from the Al framework.

10. The computer-implemented method according to any of the preceding claims, wherein the generator state is read and updated after execution of each of the layers by the amount of random numbers in a case of control flow dependencies.

11. The computer-implemented method according to any of the preceding claims, further comprising invoking an initialization function to initialize the random state using a user defined seed value.

12. The computer-implemented method according to any of the preceding claims, wherein the compiled library is associated with the runtime library, which contains the generator state, and wherein the runtime library is associated with further compiled libraries used for scheduling the random number generations.

13. The computer-implemented method according to any of the preceding claims, wherein a reseeding method is attached to different Al frameworks that use the method for the random number generations, and wherein the random number generations generated in the different Al frameworks are ensured to be identical by using the same generator states.

14. A computer system for providing for artificial intelligence (Al) framework identical pseudo random number generation, the computer system comprising one or more hardware processors which, alone or in combination, are configured to provide for execution of the following steps: storing a generator state as a static object within a compiled library, or in a runtime library associated with the compiled library, the generator state corresponding to a random state associated with a pseudo random number generation algorithm; and scheduling random number generations using the generator state, wherein the generator state is determined to be correct by forwarding the generator state by a number of random numbers drawn from previous layers executed prior to a current layer, the generator state being a same generator state for the random number generations.

15. A tangible, non-transitory computer-readable medium having instructions thereon which, upon being executed by one or more processors, provide for providing for artificial intelligence (Al) framework identical pseudo random number generation by execution of the following steps:storing a generator state as a static object within a compiled library, or in a runtime library associated with the compiled library, the generator state corresponding to a random state associated with a pseudo random number generation algorithm; and scheduling random number generations using the generator state, wherein the generator state is determined to be correct by fast forwarding the generator state by a number of random numbers drawn from previous layers executed prior to a current layer, the generator state being a same generator state for the random number generations.

Citation Information

Patent Citations

  • Flexible SPI Backlight Controller and Compiler

    US62636124P0

  • Generation of distinct pseudorandom number streams based on program context

    US9128791B1