Physics-aware training for deep physics neural networks

The physics-aware neural network system addresses the energy limitations of deep neural networks by using a hybrid physical-digital approach with backpropagation, resulting in faster and more energy-efficient machine learning hardware.

JP7681125B2Active Publication Date: 2025-05-21NTT RESEARCH INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023564492
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-22
Filing Date
2022-04-21
Publication Date
2025-05-21
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

Deep neural networks are increasingly limited by their growing energy requirements, which outpace Moore's Law and hinder their scalability and widespread use.

Method used

A physics-aware neural network system that combines a physical component and a digital component to perform a physics-aware training process, using backpropagation to train sequences of controllable physical systems as deep neural networks, facilitating non-conventional machine learning hardware that is faster and more energy-efficient.

Benefits of technology

This approach enables efficient training of deep neural networks using significantly less energy and faster processing, overcoming the limitations of conventional electronic processors and achieving orders of magnitude improvement in energy efficiency and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681125000029
    Figure 0007681125000029
  • Figure 0007681125000030
    Figure 0007681125000030
  • Figure 0007681125000031
    Figure 0007681125000031
Patent Text Reader

Abstract

The physical neural network system includes a physical component and a digital component. The digital component includes a computing system. The physical component and the digital component work together to perform a physics-aware training process. The physics-aware training process includes generating, by the digital component, an input data set for input to the physical component, applying, by the physical component, one or more transformations to the input data set to generate an output for a forward pass of the physics-aware training process, comparing, by the digital component, the generated output to a standard output to determine an error, generating, by the digital component, a loss gradient using a differentiable digital model for a backward pass of the physics-aware training process, and updating, by the digital component, training parameters for subsequent input to the physical component based on the loss gradient.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 178,318, filed April 22, 2021, which is incorporated by reference in its entirety.

[0002] The present disclosure relates generally to deep physical neural networks, and more particularly to deep physical neural networks trained using the backpropagation method for any physical system. [Background technology]

[0003] Deep neural networks are growing in their applications across business, technology, and science. As deep neural networks continue to grow in size, so does the energy they consume. While the hardware field was initially able to keep up with the development of deep neural network technology, advances in deep learning are occurring so rapidly that they are outpacing Moore's Law. Summary of the Invention

[0004] In some embodiments, a physics-aware neural network system is disclosed herein. The physics-aware neural network system includes a physical component and a digital component. The digital component includes a computing system. The physical component and the digital component work together to perform a physics aware training process. The physics aware training process includes generating, by the digital component, an input data set for input to the physical component. The physics aware training process further includes applying, by the physical component, one or more transformations to the input data set to generate an output for a forward pass of the physics aware training process. The physics aware training process further includes comparing, by the digital component, the generated output to a standard output based on the generated output to determine an error. The physics aware training process further includes generating, by the digital component, a loss gradient using a differentiable digital model for a backward pass of the physics aware training process. The physics aware training process further includes updating, by the digital component, training parameters for subsequent input to the physics component based on the loss gradient.

[0005] In some embodiments, a method of training a physical neural network is disclosed herein. A digital component of the physical neural network generates an input data set for input to the physical component. The physical component of the physical neural network applies one or more transformations to the input data set to generate an output for a forward pass of training. Based on the generated output, the digital component compares the generated output to a canonical output to determine an error. The digital component generates a loss gradient using a differentiable digital model for a backward pass of training. The digital component updates training parameters for subsequent input to the physical component based on the loss gradient.

[0006] In some embodiments, a non-transitory computer-readable medium is disclosed herein. The non-transitory computer-readable medium includes one or more sequences of instructions that, when executed by one or more processors, cause a computing system to perform operations. The operations include generating an input data set for input to a physical component of a physical neural network. The operations further include causing the physical component of the physical neural network to apply one or more transformations to the input data set to generate an output for a forward pass of the training process. The operations further include comparing the generated output to a standard output based on the generated output to determine an error. The operations further include generating a loss gradient using a differentiable digital model for a backward pass of the training process. The operations further include updating training parameters for subsequent input to the physical component based on the loss gradient.

[0007] In a manner that the above-mentioned features of the present disclosure may be understood in detail, a more particular description of the present disclosure briefly summarized above can be had by reference to embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical embodiments of the present disclosure and therefore should not be considered as limiting the scope of the present disclosure, since the present disclosure may admit of other equally effective embodiments. [Brief description of the drawings]

[0008] [Figure 1A] FIG. 1 is a block diagram illustrating an example physical neural network, in accordance with an example embodiment. [Figure 1B] FIG. 2 is a block diagram illustrating a physical neural network in greater detail, in accordance with an example embodiment; [Diagram 2] FIG. 1 is a block diagram illustrating a general physical neural network system, in accordance with an example embodiment. [Figure 3A] FIG. 1 is a block diagram illustrating a conventional backpropagation process, in accordance with an example embodiment. [Figure 3B] FIG. 1 is a block diagram illustrating a physics-aware backpropagation process in accordance with an example embodiment. [Figure 4] FIG. 11 is a block diagram illustrating a detailed physics awareness training process in accordance with an exemplary embodiment; [Diagram 5] FIG. 1 is a flow diagram illustrating a method for training a physics neural network using physics perception training, according to an example embodiment. [Figure 6] FIG. 1 is a block diagram illustrating an example physical neural network, in accordance with an example embodiment. [Figure 7A] FIG. 2 illustrates a system bus architecture of a computing system according to an exemplary embodiment. [Figure 7B] FIG. 1 illustrates a computer system having a chipset architecture, according to an exemplary embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] For ease of understanding, the same reference numerals have been used, wherever possible, to designate identical elements common to each of the figures, and it is contemplated that elements disclosed in one embodiment may be beneficially utilized on other embodiments without specific recitation.

[0010] Deep neural networks have become a pervasive tool in science and engineering. However, the growing energy requirements of modern deep neural networks increasingly limit their scalability and widespread use. To address this, one or more techniques described herein propose a radical alternative to implementing deep neural network models via physical neural networks. For example, disclosed herein is a hybrid physical-digital algorithm, referred to as "physics-aware training," that efficiently trains sequences of controllable physical systems to function as deep neural networks. Such an approach uses backpropagation, the same technique used for modern deep neural networks, to directly and automatically train the functions of any sequence of real physical systems. Physical neural networks can facilitate non-conventional machine learning hardware that is orders of magnitude faster and more energy efficient than conventional electronic processors.

[0011] Like many historical developments in artificial intelligence, the widespread adoption of deep neural networks (DNNs) was made possible in part by synergistic hardware. In 2012, Krizhevsky et al., building on numerous previous works, showed that the stochastic gradient descent (SGD) backpropagation algorithm can be efficiently executed on graphics processing units to train large convolutional DNNs to perform accurate image classification. Since 2012, the breadth of applications of DNNs has expanded, but so has their typical size. As a result, the computational requirements of DNN models have grown rapidly, outpacing Moore's Law. Currently, DNNs are increasingly limited by the energy efficiency of hardware.

[0012] The emerging DNN energy problem has implications for special-purpose hardware: DNN "accelerators." Several proposals go beyond conventional electronics, using alternative physics platforms such as optical elements or memristor crossbar arrays. These devices typically rely on approximate similarities between the hardware physics and the mathematical operations in DNNs. Their success therefore hinges on intensive engineering to push device performance closer to the limits of the hardware physics, while carefully suppressing parts of the physics that violate the similarity, such as unintended nonlinearities, noise processes, and device variations.

[0013] More generally, however, the controlled evolution of physical systems is well suited to realizing deep learning models. DNNs and physical processes share many structural similarities, such as hierarchy, approximate symmetry, redundancy, and nonlinearity. These structural commonalities explain much of the success of DNNs that operate soundly on data from the natural physical world. As the physical system evolves, it effectively performs mathematical operations within the DNN, such as controlled convolutions, nonlinearities, and matrix-vector operations. These physical computations can be exploited by encoding input data into the initial conditions of the physical system and then performing measurements and reading the results after the system has evolved. The physical computations can be controlled by adjusting physical parameters. By cascading such controlled physical input-output transformations, trainable hierarchical physical computations can be realized. As anyone who has simulated the evolution of a complex physical system knows, physical transformations are typically faster and consume less energy than digital emulations, and processes that require nanoseconds or nanojoules often require seconds or joules to simulate digitally. PNNs are therefore a route to scalable, energy-efficient and fast machine learning.

[0014] Theoretical proposals for physics-learning hardware have recently emerged in a variety of fields, including optics, spintronic nano-oscillators, nanoelectronic devices, and small-scale quantum computers. A related trend is physics-reservoir computing, where the information transformations of physical systems "reservoirs" are not trained but instead linearly combined by a trainable output layer. Reservoir computing exploits general physical processes for computation, but its training is inherently shallow and does not allow for the hierarchical process learning that characterizes state-of-the-art deep neural networks. In contrast, new proposals for physics-learning hardware overcome this problem by training the physical transformations themselves.

[0015] However, experimental studies on physical learning hardware are scarce, and those that do exist rely on gradient-free learning algorithms. While these studies have made important steps, it is now recognized that gradient-based learning algorithms, such as the backpropagation algorithm, are essential for efficient training and good generalization of large-scale DNNs. To solve this problem, proposals have emerged to realize backpropagation in physical hardware. These proposals are inspirational, but they often rely on limiting assumptions, such as linearity and dissipation-free evolution. The most common proposals may overcome such constraints, but still rely on in silico training, i.e., performing training entirely within numerical simulations. Therefore, to realize it experimentally and on scalable hardware, based on mathematical similarities, we will face the same challenge as hardware: dedicated engineering efforts to accurately match the hardware to the ideal simulation.

[0016] FIG. 1A is a block diagram illustrating an exemplary physical neural network 100 according to an exemplary embodiment. To facilitate the discussion of the physical neural network 100, the following is a brief discussion of conventional artificial neural networks. Conventional artificial neural networks (ANNs) are typically built from computational units (layers) that include trainable matrix-vector multiplications followed by element-wise nonlinear activations such as rectifiers (ReLUs). The weights of the matrix-vector multiplications can be adjusted during training so that the ANN can perform the desired mathematical operations. Deep neural networks (DNNs) can be created by cascading these computational units together. DNNs can be trained to learn multi-step (hierarchical) computations. As the physical system evolves, it can perform the computations. The controllable properties of the system can be divided into input data and control parameters. By varying the control parameters, the physical transformations performed on the input data can be altered.

[0017] As illustrated in FIG. 1A, the physical neural network 100 may receive as inputs data inputs 102 and parameters 104. Based on the data inputs 102 and parameters 104, the physical neural network 100 may generate outputs 106. Just as hierarchical information processing is realized in DNNs by a sequence of trainable nonlinear functions, the deep physical neural network 100 may be created by cascading layers of trainable physical transformations. In the physical neural network 100, each physical layer implements a controllable function, which may be of a more general form than that of a traditional DNN layer.

[0018] 1B is a block diagram illustrating physical neural network 100 in more detail, according to an exemplary embodiment. As shown, physical neural network 100 may be comprised of multiple layers 112. For example, as shown, physical neural network 100 may be comprised of layers 112a, 112b, 112c, 112d, and 112n. In this particular example, layer 112a may be the first layer in physical neural network 100. As shown, layer 112a may receive data input 114 and parameters 116 as inputs. Based on data input 114 and parameters 116, layer 112a may generate output 118. Data input 114, parameters 116, and output 118 may be provided as inputs to the next layer, i.e., layer 112b. Based on data input 114, parameters 116, and output 118, layer 112b may generate output 120. The data input 114, parameters 116, and output 120 may be provided as inputs to the next layer, i.e., layer 112c. Such a process may continue for each of layers 112d through 112n to produce a final output 122.

[0019] As described above, the physical neural network 100 may exemplify a universal framework for directly training any real physical system to perform deep neural networks using backpropagation. The trained hierarchical physical computation may be referred to as a physical neural network (PNN). A hybrid physical-digital algorithm, i.e., physical awareness training (PAT), allows the backpropagation algorithm to be efficiently and accurately performed on the fly directly on sequences of physical input-output transformations. PNNs represent a radical departure from traditional hardware, yet are easily integrated into modern machine learning. For example, PNNs can be seamlessly combined with traditional hardware and neural network methods via a physical-digital hybrid architecture, where traditional hardware learns to work conveniently with non-traditional physical resources using PAT. Ultimately, PNNs provide a basis for hardware-physical-software co-design in artificial intelligence, a pathway to improve the energy efficiency and speed of machine learning by orders of magnitude, and a path to automatically design complex functional devices such as functional nanoparticles, robots, smart sensors, and the like.

[0020] PAT may be used to train the parameters of the physical neural network 100. PAT is an algorithm that allows the backpropagation algorithm of stochastic gradient descent (SGD) to be performed directly on a sequence of physical input-output transformations. In some embodiments, automatic differentiation in the backpropagation algorithm may efficiently determine the gradient of the loss function with respect to the trainable parameters. This makes the algorithm about N times more efficient than finite difference methods of gradient estimation, where N is the number of parameters. PAT may have some similarities to quantized cognitive training algorithms used to train neural networks for low-precision hardware and feedback alignment. PAT may be viewed as solving a problem similar to the "simulation-reality gap" in robotics, which is incrementally addressed by physical-digital hybrid techniques.

[0021] As mentioned above, Physics Awareness Training (PAT) is a gradient-based learning algorithm. The algorithm may calculate the gradient of the loss with respect to the parameters of the network. The gradient of the loss is then used to update the parameters of the network, since the loss may indicate how well the network is performing at its machine learning task. The gradient may be efficiently calculated via a backpropagation algorithm.

[0022] The backpropagation algorithm is commonly applied to neural networks composed of differentiable functions. It involves two key steps: a forward pass to compute a loss, and a backward pass to compute a gradient with respect to the loss. The mathematical technique behind this algorithm may be referred to as reverse-mode autodiff. In some embodiments, each differentiable function in the network may be an autodiff function, which may specify how signals propagate forward through the network and how error signals propagate backward. Given the constituent autodiff functions and a specification of how these different functions are interconnected, i.e., the network architecture, reverse-mode autodiff may be able to iteratively compute the desired gradient from the output towards the input and parameters (heuristically called "backwards") in an efficient manner. For example, the output of a conventional deep neural network may be expressed as f(f(...f(f(x,θ 1 ),θ 2 )...,θ [N-1] ),θ N ), where f may represent the constituent autodiff functions. For example, f may be given by f(x,θ)=Relu(Wx+b), where the weight matrix W and bias b may represent the parameters of a given layer, and Relu may be the activation function of a rectified linear unit (although other activation functions may be used). With a specification of how the forward and backward passes of f are performed, the autograd algorithm may be able to compute the loss of the entire network.

[0023] In physics recognition training, an alternative implementation of the traditional backpropagation algorithm may be used. For example, this variation may employ an autodiff function that may utilize different functions for the forward and backward passes.

[0024] FIG. 2 is a block diagram illustrating a general physical neural network system 200, according to an example embodiment. As shown, the physical neural network system 200 may include a digital component 202 and a physical component 204. The digital component 202 may be configured to generate an input data set for the physical component 204. In general, the digital component 202 may represent a computing system as described below in connection with FIG. 7A and FIG. 7B. The physical component 204 may represent a physical system, such as, but not limited to, a mechanical system, an optical system, an electronic system, etc. More generally, the physical component 204 may be configured to perform what would traditionally be a forward pass of a backpropagation process to generate an output. The digital component 202 may be configured to perform what would traditionally be a backward pass of a backpropagation process based on the output generated from the forward pass. The overall training process is further described below.

[0025] Figure 3A is a block diagram illustrating a conventional backpropagation process 300, according to an example embodiment. Figure 3B is a block diagram illustrating a PAT backpropagation process 350, according to an example embodiment.

[0026] 3A, a conventional backpropagation algorithm uses the same function for the forward pass 302 and the backward pass 304. Mathematically, the forward pass 302, which maps the inputs and parameters to the output, is given by y=f(x,θ), where:

number

number

number

number

[0027] The backward pass 304 that maps the gradient with respect to the outputs to the gradient with respect to the inputs and parameters may be given by the Jacobian vector product:

number

number

number

number

number

number

number

[0028] The traditional backpropagation algorithm illustrated in FIG. 3A may work well for deep learning using digital computers, but it cannot be simply applied to physical neural networks. It is based on the analytical gradient (

number

[0029] In contrast, FIG. 3B illustrates an alternative backpropagation algorithm suitable for any practical physical input-output transformation. The PAT backpropagation process 350 may employ an autofiff function that may utilize different functions for the forward pass 352 and the backward pass 354. As illustrated diagrammatically in FIG. 3B, a non-differentiable transformation (f p ) may be used in the forward path 352 to obtain a differentiable digital model of the physical transformation (f m ) may be used in the backward pass 354. Physics-aware training is a backpropagation algorithm, which can efficiently compute gradients iteratively and thus efficiently train physics systems. Furthermore, because physics-aware training is formulated in the same paradigm of reverse-mode automatic differentiation, users can define these custom autodiff functions in any traditional machine learning library (such as PyTorch) to design and train physics neural networks using the same workflow employed for traditional neural networks.

[0030] For a physical transformation of a component in the overall physical neural network, the forward pass operation of this component is y=f p(x,θ). Because a different function is used in the backward pass 354 than in the forward pass 352, the autodiff function may not be able to backpropagate the gradients at the output layer to the exact gradients at the input layer. Instead, it may try to approximate the backpropagation of the exact gradients. Thus, the backward pass 354 is

number

number

number

number

[0031] In other words, in PAT training, the backward path 354 is a differentiable digital model, i.e., f m (x,θ), while the forward path 352 can be estimated using a physical system, i.e., f p (x, θ).

[0032] FIG. 4 is a block diagram illustrating a detailed physics awareness training process 400 according to an exemplary embodiment. As shown, FIG. 4 may illustrate a complete training loop for physics awareness training. Such discussion may utilize a feed-forward PNN architecture because it is the PNN equivalent of a standard feed-forward neural network (multilayer perceptron) in deep learning. Physics awareness training is specifically designed to avoid the shortcomings of both the ideal backpropagation algorithm and in silico training. It is a hybrid algorithm that includes computations in both the physical and digital domains.

[0033] More specifically, the physical system is used to perform the forward pass, thereby alleviating the burden of making the differential digital model very accurate (as in in silico training). The differentiable digital model may only be utilized in the backward pass to complement the parts of the training loop that the physical system cannot perform. Physics-aware training can be formalized by using custom component autodiff functions in the overall network architecture. For feedforward PNNs, the autodiff algorithm with these custom functions results in the following simplified training loop: Execute the forward pass. x [l+1] =y [l] =f p (x [l] ,θ [l] ) Calculate the error vector.

number

number

number

number

number

number

number

number

number

number

[0034] 5 is a flow diagram illustrating a method 500 for training a physical neural network (PNN) using physics perception training, according to an example embodiment. The method 500 may begin as step 502.

[0035] As step 502, a computing system may provide input to the PNN. In some embodiments, the input may include training data and trainable parameters. In some embodiments, the training data and trainable parameters may be encoded prior to input to the PNN. Using a specific example, the input data and parameters may be encoded into a time-dependent force applied to a suspended metal plate.

[0036] In step 504, the PNN may use the transformation to generate outputs in the forward pass. For example, as described above, in the forward pass, the PNN may [l+1] =y [l] =f p (x [l] ,θ [l] ) to generate the output in the forward pass.

[0037] In step 506, the computing system may generate or calculate an error. For example, the computing system may compare the actual physical output to a standard or expected physical output. The difference between the actual physical output and the standard or expected physical output may represent the error. In some embodiments, the error vector is:

number

[0038] In step 508, the computing system may use a differentiable digital model to generate a loss gradient. For example, using a differentiable digital model to estimate the gradient of the PNN, the computing system may generate a gradient of the loss with respect to the controllable parameters. The backward pass is

number

[0039] In step 510, the computing system may update the parameters. For example, the computing system may update the parameters of the system based on the estimated gradient.

[0040] Such a process may continue until transformation occurs.

[0041] 6 is a block diagram illustrating an example physical neural network (PNN) 600, according to an example embodiment. As shown, the PNN 600 may represent an oscillating plate experimental setup. For example, the PNN 600 may include an audio amplifier 602, a commercial full-range speaker 604, a microphone 606, and a computer 608 that controls the setup. The computer 608 may provide an input signal 610 to the amplifier 602. In some embodiments, the input signal 610 may include a physical input encoded in the time domain. For example, the computer 608 may encode the input with a time series of square pulses that may be converted to an analog signal by a digital-to-analog converter of the computer 608.

[0042] The amplifier 602 may be configured to amplify an input signal received from the computer 608 and apply the amplified input signal to a mechanical oscillator realized by the voice coil of an acoustic speaker. For example, the speaker may be used to drive the mechanical vibration of a titanium plate attached to the voice coil of the speaker.

[0043] The microphone 606 may be configured to record the sound generated by the diaphragm. In some embodiments, the sound recorded by the microphone 606 may be converted back to digital. The recorded sound may represent an output signal 612 that is returned to the computer 608. The computer 608 may be configured to compare the output signal 612 to an expected output signal to generate an error. The computer 608 may further be configured to evaluate a loss slope with respect to the controllable parameters using the digital model. Based on the generated slope, the computer 608 may update or change the parameters passed to the amplifier 602.

[0044] FIG. 7A illustrates a system bus architecture of a computing system 700 according to an exemplary embodiment. The system 700 may represent at least a portion of the digital component 202. One or more components of the system 700 may be in electrical communication with each other using a bus 705. The system 700 may include a processing unit (CPU or processor) 710 and a system bus 705 that couples various system components to the processor 710, including a system memory 715 such as a read-only memory (ROM) 720 and a random access memory (RAM) 725. The system 700 may include a cache of high-speed memory directly connected to the processor 710, closely connected to the processor 710, or integrated as part of the processor 710. The system 700 may copy data from the memory 715 and / or the storage device 730 to the cache 712 for quick access by the processor 710. In this manner, the cache 712 may provide a performance boost that avoids delays to the processor 710 while waiting for data. These and other modules may control or be configured to control the processor 710 to perform various operations. Other system memory 715 may be available for use as well. The memory 715 may include multiple different types of memory with different performance characteristics. The processor 710 may include any general purpose processor and hardware or software modules, such as service 1 732, service 2 734, and service 3 736 stored in the storage device 730, configured to control the processor 710 and special purpose processors, where the software instructions are embedded in the actual processor design. The processor 710 may be essentially a fully self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0045] To enable user interaction with the computing system 700, the input device 745 may represent any number of input mechanisms, such as a microphone for audio, a touch-sensitive screen for gesture or graphic input, a keyboard, a mouse, motion input, voice, etc. The output device 735 may be one or more of many output mechanisms known to those skilled in the art. In some cases, a multi-modal system may allow a user to provide multiple types of input to communicate with the computing system 700. The communication interface 740 may generally govern and manage user input and system output. There is no limitation to operation in any particular hardware configuration, so the basic functions herein may be easily replaced with improved hardware or firmware configurations developed.

[0046] The storage device 730 may be a non-volatile memory, such as a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a random access memory (RAM) 725, a read only memory (ROM) 720, and hybrids thereof, which may store data accessible by a computer, such as a hard disk or other type of computer readable medium.

[0047] The storage device 730 may include services 732, 734, and 736 for controlling the processor 710. Other hardware or software modules are also contemplated. The storage device 730 may be connected to the system bus 705. In one aspect, hardware modules that perform certain functions may include software components stored on computer readable media that connect with necessary hardware components such as the processor 710, the bus 705, output devices 735 (e.g., displays), etc. to perform the functions.

[0048] FIG. 7B illustrates a computer system 750 having a chipset architecture that may represent at least a portion of the digital component 202. The computer system 750 may be an example of computer hardware, software, and firmware that may be used to implement the disclosed techniques. The system 750 may include a processor 755, which represents any number of physically and / or logically distinct resources capable of executing software, firmware, and hardware configured to perform the specified computations. The processor 755 may communicate with a chipset 760, which may control input to and output from the processor 755. In this example, the chipset 760 may output information to an output 765, such as a display, and may read and write information to a storage device 770, which may include, for example, magnetic media and solid-state media. The chipset 760 may also read data from and write data to the storage device 775 (e.g., RAM). A bridge 780 for interfacing with various user interface components 785 may be provided for interfacing with the chipset 760. Such user interface components 785 may include a keyboard, a microphone, touch detection and processing circuitry, a pointing device such as a mouse, etc. In general, input to the system 750 may come from any of a variety of sources, including machine-generated and / or human-generated.

[0049] The chipset 760 may also interface with one or more communication interfaces 790, which may have different physical interfaces. Such communication interfaces may include interfaces for wired and wireless local area networks, broadband wireless networks, and personal area networks. Some applications of the methods for generating, displaying, and using GUIs disclosed herein may include receiving an ordered data set via a physical interface, or may be generated by the machine itself by the processor 755 analyzing data stored in the storage device 770 or storage device 775. Additionally, the machine may receive inputs from a user via the user interface component 785 and perform appropriate functions, such as browsing functions, by interpreting these inputs using the processor 755.

[0050] It may be appreciated that the exemplary systems 700, 750 may have one or more processors 710, or may be part of a group or cluster of computing devices networked together to provide greater processing power.

[0051] While the above is directed to the embodiments described herein, other and further embodiments may be devised without departing from the basic scope thereof. For example, aspects of the disclosure may be implemented in hardware or software, or a combination of hardware and software. An embodiment described herein may be implemented as a program product for use with a computer system. The programs of the program product define the functions of the embodiments (including the methods described herein) and may be stored on various computer-readable storage media. Exemplary computer-readable storage media include, but are not limited to, (i) non-writeable storage media in which information is permanently stored (e.g., a read-only memory (ROM) device in a computer, such as a CD-ROM disk readable by a CD-ROM drive, a flash memory, a ROM chip, or any type of solid-state non-volatile memory), and (ii) writeable storage media in which changeable information is stored (e.g., a floppy disk in a diskette drive, a hard disk drive, or any type of solid-state random access memory). Such computer-readable storage media, when having computer-readable instructions that direct the functions of the disclosed embodiments, are embodiments of the disclosure.

[0052] Those skilled in the art will recognize that the foregoing examples are illustrative and not limiting. All permutations, extensions, equivalents, and improvements thereto will be apparent to those skilled in the art upon reading the specification and studying the drawings, and are intended to be included within the true spirit and scope of the present disclosure. Accordingly, it is intended that the following appended claims include all such modifications, permutations, and equivalents that are included within the true spirit and scope of these teachings.

Claims

1. 1. A physical neural network system, comprising: A physical component; a digital component including a computing system, the physical component and the digital component operating in conjunction to execute a physics awareness training process; wherein the physics awareness training process comprises: encoding, by the digital component, input data and initial parameters to generate an input data set for input to the physical component; applying, by the physics component, one or more transformations to the input data set to generate outputs for a forward pass of the physics perception training process; comparing the generated output by the digital component to a standard output to determine errors based on the generated output; generating, by the digital component, a loss gradient using a differentiable digital model for a backward pass of the physics perception training process; updating, by the digital component, training parameters for subsequent input to the physical component based on the loss gradient; A physical neural network system, including:

2. applying, by the physics component, the one or more transformations to the input data set to generate the output for the forward pass of the physics perception training process; employing at least one non-differentiable transformation function on the input data set in the forward pass; 2. The physical neural network system of claim 1, comprising:

3. generating, by the digital component, the loss gradient using the differentiable digital model for the backward pass of the physics perception training process, 3. The physical-neural-network system of claim 2, further comprising approximating the loss gradient using the differentiable digital model, wherein the differentiable digital model in the backward path is different from the at least one non-differentiable transformation function in the forward path.

4. The physical neural network system of claim 1 , wherein the physical component comprises a first layer and a second layer.

5. applying, by the physics component, the one or more transformations to generate the output for the forward pass of the physics perception training process, applying, by the first layer, a first transformation to the input data set to generate an intermediate output; applying, by the second layer, a second transformation to the input data set and the intermediate output to generate the output; 5. The physical neural network system of claim 4, comprising:

6. generating, by said digital component, an updated input data set by encoding said input data and said updated training parameters; 10. The physical neural network system of claim 1 further comprising:

7. 1. A method for training a physical neural network, comprising: encoding, by a digital component of the physical neural network, input data and initial parameters to generate an input data set for input to a physical component of the physical neural network; applying, by the physics component, one or more transformations to the input dataset to generate outputs for a forward pass of the training; based on the generated output, comparing the generated output with a standard output by the digital component to determine an error; generating, by the digital component, a loss gradient for a backward pass of the training using a differentiable digital model; updating, by the digital component, training parameters for subsequent input to the physical component based on the loss gradient; A method comprising:

8. applying, by the physics component, the one or more transformations to the input dataset to generate the outputs for the forward pass of the training, employing at least one non-differentiable transformation function on the input data set in the forward pass; The method of claim 7, comprising:

9. generating, by the digital component, the loss gradient using the differentiable digital model for the backward pass of the training, 9. The method of claim 8, comprising approximating the loss gradient using the differentiable digital model, the differentiable digital model in the backward path being different from the at least one non-differentiable transformation function of the forward path.

10. The method of claim 7 , wherein the physical component comprises a first layer and a second layer.

11. applying, by the physics component, the one or more transformations to generate the outputs for the forward pass of the training, applying, by the first layer, a first transformation to the input data set to generate an intermediate output; applying, by the second layer, a second transformation to the input data set and the intermediate output to generate the output; The method of claim 10, comprising:

12. generating, by said digital component, an updated input data set by encoding said input data and said updated training parameters; The method of claim 7 further comprising:

13. A non-transitory computer-readable medium comprising one or more sequences of instructions that, when executed by one or more processors, cause a computing system to: encoding the input data and the initial parameters to generate an input data set for input to a physical component of the physical neural network; causing the physics component to apply one or more transformations to the input data set to generate outputs for a forward pass of a training process; based on the generated output, comparing the generated output with a standard output to determine errors; generating a loss gradient using a differentiable digital model for a backward pass of the training process; updating training parameters for subsequent inputs to the physics component based on the loss gradient; A non-transitory computer-readable medium for causing operations to be performed, including:

14. applying, by the physics component, the one or more transformations to the input data set to generate the outputs for the forward pass of the training process; employing at least one non-differentiable transformation function on the input data set in the forward pass; 14. The non-transitory computer readable medium of claim 13, comprising:

15. generating the loss gradient using the differentiable digital model for the backward pass of the training process, 15. The non-transitory computer-readable medium of claim 14, further comprising approximating the loss gradient using the differentiable digital model, wherein the differentiable digital model in the backward path differs from the at least one non-differentiable transformation function of the forward path.

16. The non-transitory computer-readable medium of claim 13 , wherein the physical component comprises a first layer and a second layer.

17. applying the one or more transformations to the physical components to generate the outputs for the forward pass of the training process, causing the first layer to apply a first transformation to the input data set to generate an intermediate output; causing the second layer to apply a second transformation to the input data set and the intermediate output to generate the output; 20. The non-transitory computer readable medium of claim 16, comprising:

Citation Information

Patent Citations

  • Semiconductor device and learning method therefor

    JP2004030624A

  • Analog Electronic Neural Networks

    JP2019511799A

  • Classification device, classification method and classification program

    JP2020173624A

  • Resistive processing unit architecture with separate weight update and inference circuitry

    US20190318239A1