Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

125 results about "Backward propagation" patented technology

Backward propagation: We can define a cost function that measures how good our neural network performs. where k stands for the training example and the output is assumed to be the activation of the output neuron, and y is the actual desired output.

Quantitative perception training method and device of neural network model, electronic equipment and storage medium

The invention relates to a quantitative perception training method and device of a neural network model, electronic equipment and a storage medium. The method comprises the steps of obtaining a to-be-trained first neural network model; an operator pair and a non-module operator in the first neural network model are identified, the non-module operator is an operation which is not realized based on a module class, and the operator pair comprises a convolution operator and a batch normalization operator which are connected; the non-module operators are packaged into module operators, the operator pairs are packaged into new convolution operators, a second neural network model is obtained, the new convolution operators run the fusion process of the convolution operators and batch normalization operators during forward propagation, and the parameters of the convolution operators and the batch normalization operators are updated during back propagation; and performing quantitative perception training on the second neural network model. By adopting the method, the reasoning precision of the quantitative model can be guaranteed, and the deployment efficiency of the quantitative model is improved.
Owner:GUANGZHOU XIAOMA HUIXING TECH CO LTD

Teacher tensor knowledge distillation-driven anti-violation target detection method in electric power safety supervision scene

The invention belongs to the technical field of target detection, and particularly relates to a teacher tensor knowledge distillation-driven anti-violation target detection method in an electric power safety supervision scene, and the method comprises the following steps: inputting image data to a student model and a pre-trained teacher model; performing forward propagation on the student model to obtain a student prediction tensor; associating differences between the real tags based on the student prediction tensor and the input image data; acquiring an original teacher output tensor by using the same input image data by calling a bottom layer forward propagation method of a teacher model; calculating knowledge distillation loss based on the difference between the student prediction tensor and the original teacher output tensor; the standard detection loss and the knowledge distillation loss are combined to form total loss; and performing back propagation updating on parameters of the student model based on the total loss so as to complete model training and real-time target detection and detection method updating optimization. According to the invention, the performance of the student model can be significantly improved, and the high efficiency of the model is maintained.
Owner:HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER +2

Slow node detection method and device, equipment, storage medium and program product

Embodiments of the invention disclose a slow node detection method and apparatus, a device, a storage medium and a program product. The method comprises the steps of obtaining respective first duration data of a plurality of processing nodes participating in distributed model training; the first duration data represents the calculation time consumption of the corresponding processing node in the current iteration; based on a parallel strategy trained by the distributed model, grouping the first duration data to obtain a first duration set of multiple groups; the parallel strategy represents an association relationship of the plurality of processing nodes in a forward and backward propagation process and training tasks undertaken by the plurality of processing nodes; and based on the respective probability density function and a first threshold value of the plurality of groups, respectively analyzing and processing the first duration sets of the plurality of groups to obtain slow nodes in the plurality of processing nodes. Thus, the slow nodes in distributed model training can be intelligently and accurately detected, the efficiency and accuracy of slow node detection are improved, and the method has high compatibility.
Owner:MOORE THREAD INTELLIGENT TECHNOLOGY (HANGZHOU) CO LTD

Image-text pedestrian retrieval method based on image block replacement and cross-modal identity alignment

The invention belongs to the field of computer vision and cross-modal retrieval, and particularly relates to an image-text pedestrian retrieval method based on image block replacement and cross-modal identity alignment, which comprises the following steps: acquiring a public data set of image and text description, and constructing a pedestrian re-recognition model PRCIA; inputting the data set into a PRCIA model for training and verification, and performing iterative updating on a training weight file through forward and backward propagation to obtain a trained PRCIA model; constructing a reasoning stage model, reserving double encoders and fusing global features in a reasoning stage, and ensuring the calculation efficiency; and inputting the test set into the reasoning stage model to obtain a detection result, thereby realizing text-based pedestrian detection. According to the method, fine-grained association between the image blocks and the text phrases is established through the PR module, cross-modal identity feature expression is enhanced through the CIA module, the accuracy and robustness of text-to-image pedestrian retrieval are remarkably improved, and the requirements of actual scenes are more easily met.
Owner:JIANGSU UNIV

Distributed AI training-oriented RDMA transmission mode intelligent selection method

The invention discloses a distributed AI training-oriented RDMA transmission mode intelligent selection method, and the method comprises the steps: obtaining multi-dimensional data features in a deep learning training process, the multi-dimensional data features comprising a tensor structure feature, a data access mode feature, a data timeliness feature, a data dependency feature and a data change rate feature; identifying a current training stage, wherein the training stage comprises a forward propagation stage, a back propagation stage and a parameter updating stage; based on the multi-dimensional data features and the training stage, an optimal RDMA transmission mode matched with the current training scene is selected through an adaptive decision matrix, and the optimal RDMA transmission mode comprises one or more combinations of RDMA Write operation, RDMA Read operation, Send / Receive operation and Atomic operation; and switching the RDMA transmission mode in real time according to the dynamic change of a network environment, ensuring that a training data flow is not interrupted in a switching process by adopting a progressive flow migration strategy, and optimizing a subsequent mode selection decision through a historical performance learning mechanism.
Owner:JINAN INSPUR DATA TECH CO LTD

Bidirectional adaptive video super-resolution method based on frame difficulty index

The invention discloses a bidirectional adaptive video super-resolution method based on frame difficulty index, and belongs to the field of computer vision. The invention provides a frame-level dynamic reconstruction method for solving the problems that simple frame calculation is redundant and difficult frame reconstruction is insufficient due to the fact that an existing model adopts a fixed calculation strategy for video frames with different difficulties. According to the method, a motion detail decoupling propagation network is constructed, motion information is efficiently transmitted by utilizing a shallow forward propagation branch, and a deep backward propagation branch focuses on recovering texture details so as to decouple a time sequence propagation task; meanwhile, a frame reconstruction difficulty evaluation network is introduced to generate a global difficulty index, so that the receptive field weight of the adaptive time sequence fusion network and the refining depth of the dynamic refining network are regulated and controlled. According to the method, through explicit modeling frame-level reconstruction difficulty, adaptive matching of the model capacity and the video frame feature complexity is realized, and the video reconstruction performance is remarkably improved under limited computing power.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Rural domestic sewage treatment system and method for controlling A2O reaction process based on energy detection

The invention discloses a rural domestic sewage treatment system and method for controlling an A2O reaction process based on energy detection. The system is provided with a grating well, a regulating tank, an anaerobic-anoxic zone, an aerobic zone and a sedimentation tank, and is equipped with a monitoring module for monitoring the temperature and water level of sewage in each section so as to obtain volume information. The control method adopts a long short-term memory (LSTM) neural network model to perform reaction process prediction, and specifically comprises the following steps: firstly, performing normalization processing on monitored temperature and water level data, and constructing an input feature vector; defining an LSTM model structure and dividing a data set; model training is completed through forward propagation calculation, loss function calculation and back propagation updating parameters; and finally, predicting anaerobic, anoxic and aerobic zone reaction processes by using the trained model, and adjusting parameters such as an internal reflux ratio according to a prediction result. The method can accurately grasp the reaction process, is suitable for the characteristic of large fluctuation of the quality and quantity of rural sewage, improves the sewage treatment effect, reduces the operation and maintenance cost, and achieves data-driven intelligent decision making.
Owner:YANGTZE ECOLOGY & ENVIRONMENT CO LTD

Reinforcement learning

Method comprising: monitoring whether a MTLF receives a first state of an environment on which a RL training is to be performed; performing a ML model forward propagation on a first model of the environment having the first state for each of plural actions to obtain a respective expected reward for each of the plural actions: informing a service consumer on the plural actions and their respective expected reward; supervising whether the MTLF receives a RL training result information after the informing the service consumer on the plural actions, wherein the RL training result information comprises an indication of one of the plural actions, a second state of the environment, and a reward feedback; conducting a ML model backward propagation on the first model of the environment having the second state for the one of the plural actions using the reward feedback to obtain a second model of the environment.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Body-aware robot data closed loop system and method

This invention relates to a closed-loop data system and method for embodied intelligent robots. The system includes: a data acquisition module for acquiring multimodal operational data of the embodied intelligent robot during task execution and constructing multiple embodied state nodes based on the multimodal operational data; a construction module for constructing event correlation relationships between embodied state nodes based on temporal and physical constraints, forming an embodied state event graph; an anomaly identification module for identifying anomalous embodied state nodes in the embodied state event graph and marking them as failure source nodes; and a propagation analysis module for determining failure propagation paths forward or backward from the failure source nodes based on event correlation relationships. Using the above scheme, the operational data during the task execution of the embodied intelligent robot can be effectively analyzed, anomalous states can be accurately identified, and data-driven closed-loop optimization can be achieved.
Owner:KUNHUA TECHNOLOGY (GUANGZHOU) CO LTD

Model interpretation method and apparatus, storage medium, and computer program product

Embodiments of the present application provide a model interpretation method and apparatus, a storage medium, and a computer program product. The method comprises: determining an attention weight value of first multimedia data in a forward propagation process in a first model, and determining a gradient value of the first multimedia data in a backward propagation process in the first model, wherein the first multimedia data is any input data input to the first model; using the attention weight value and the gradient value to determine an attention value of the first multimedia data; and using the attention value as interpretability information of the first multimedia data, so as to use the interpretability information to interpret a decision-making process of the first model.
Owner:CHINA MOBILE COMM LTD RES INST +1

An end-side fingerprint representation identification method based on multi-task learning

The application discloses an end-side fingerprint representation recognition method based on multi-task learning. First, the fingerprint picture is preprocessed, and then the preprocessed training data is input into a backbone neural network to obtain basic features; then the basic features are input into a minutia extraction network, and texture information maps are generated through multiple layers of convolution and deconvolution. The basic features are input into a multi-layer perceptron to generate topological information and obtain corresponding category information. Finally, the basic features, the texture information and the topological information are updated through a joint loss function and backward propagation in three network modules to achieve the purpose of converting the basic features into fingerprint representation information with the assistance of the texture information and the topological information. The application uses a multi-task learning method to learn fingerprint feature information in multiple scales, effectively improving the fingerprint recognition accuracy. The application fuses fingerprint multi-scale information into a fingerprint representation, making the robustness stronger. The application adopts a lightweight network throughout, saving hardware resource overhead and being more suitable for end-side devices.
Owner:HANGZHOU DIANZI UNIV

Handling early exit in a pipelined query execution engine via backward propagation of early exit information

In some implementations, there is provided receiving a query request including a top k query operator for query plan generation, optimization, and execution; generating a query plan that includes at least one pipeline of a plurality of operators, wherein the at least one pipeline is associated with a directed acyclic graph; detecting in the query plan a first operator in the at least one pipeline that causes an early exit; and in response to the early exit by the first operator during query execution, processing back through the directed acyclic graph to identify at least one preceding operator that should or should not run given the early exit of the first operator.
Owner:SAP SE

Reconvergent clock mesh for on-chip compute cluster

A compute cluster includes a compute mesh of interconnected compute elements. A clock mesh includes reconvergent elements, and the reconvergent clock mesh distributes a clock signal to the compute elements. The clock mesh has a topology of nodes connected by branches. Repeaters are located on the branches between nodes. As a result, the reconvergent clock mesh is unidirectional. The repeaters limit the clock signal to forward propagation through the clock mesh and prevent backward propagation through the clock mesh. The clock mesh is reconvergent, in that the clock signal may branch out along different paths from one node and then these paths reconverge at a later node.
Owner:SIMA TECHNOLOGIES INC

Systems and methods for accelerating neural network convolution and training

A specialized integrated circuit for artificial neural networks is integrated with high bandwidth memory. The neural network includes a systolic array of interconnected processing elements, including upstream processing elements and downstream processing elements. Each processing element includes a pair of input / output ports for concurrent forward and backward propagation. The processing elements can be used for convolution, in which case the pair of input / output ports can support fast and efficient scanning of a kernel over activations.
Owner:RAMBUS INC

Medical image class incremental learning method and system based on feature principal direction guided projection

PendingCN122637089AHat matrixIncremental learning
The application discloses a medical image class incremental learning method and system based on feature principal direction guided projection, and belongs to the technical field of medical image analysis; the method acquires medical image features through a pre-trained feature extractor; random projection is performed on the features to obtain first-view features; a covariance matrix is calculated based on a feature cache area and is decomposed to extract the first K principal directions, construct a guided projection matrix, and obtain second-view features; the double-view features are spliced to obtain final representation; the classifier weight is directly calculated through a ridge regression closed-form solution without backward propagation training; the application adaptively captures the core structure of the feature space through data-driven principal direction projection, retains the detailed information in combination with random projection, effectively improves the feature separability of fine-grained classes, adopts an analytical classifier updating mechanism, realizes stable class incremental learning without storing historical data, improves the model updating efficiency, and is suitable for efficient continuous learning tasks in various medical image scenes.
Owner:XI AN JIAOTONG UNIV

Potential function model construction and training method, system and device and storage medium

The invention relates to a potential function model construction method, system and device and a storage medium, and belongs to the technical field of digital data processing. The construction method of the potential function model comprises the following steps: determining a neighbor atom set for each atom in an atom system according to a truncation radius; constructing a neighbor environment descriptor of each atom based on the neighbor atom set of each atom; inputting the neighbor environment descriptor of each atom into an independently configured neural network corresponding to the element type of the neighbor environment descriptor, and performing forward propagation calculation from an input layer to an output layer of the neural network to obtain energy of each atom; performing back propagation calculation and descriptor back calculation on the energy of each atom from the output layer to the input layer through each independently configured neural network to obtain partial derivative of the energy of each atom to the position of the adjacent atom; and obtaining the stress of each atom based on the partial derivative of the energy of each atom to the position of the adjacent atom. The potential function model has both precision and calculation efficiency.
Owner:NINGBO INST OF MATERIALS TECH & ENG CHINESE ACAD OF SCI

Construction method of reversible spiking neural network supporting parallel computing

The invention belongs to the technical field of spiking neural networks. The invention provides a construction method of a reversible pulse neural network supporting parallel computing. According to the embodiment of the invention, the reversible calculation structure is introduced, the neuron state of an intermediate layer does not need to be explicitly stored in the forward propagation process of the network, and the reconstruction of the activation value can be realized through the reversible mapping relation in the backward propagation stage. In the design stage of the reversible structure, hardware mapping characteristics are fully considered, and forward calculation and reverse reconstruction processes of reversible mapping are simultaneously incorporated into a parallel execution framework of a configurable calculation array. Through collaborative design of a data path and control logic, forward propagation and reverse reversible reconstruction can be efficiently and parallelly operated in the same array, and inter-stage calculation idle time and data waiting time are remarkably shortened. On the premise of ensuring reversibility, streamlined and parallel execution of forward and reverse calculation is realized, and the throughput rate and the training speed of the system are integrally improved.
Owner:XIDIAN UNIV

Machine learning robustness through sensible decision boundaries

Computer systems and computer-implemented methods modify a machine learning network, such as a deep neural network, to introduce judgment to the network. A “combining” node is added to the network, to thereby generate a modified network, where activation of the combining node is based, at least in part, on output from a subject node of the network. The computer system then trains the modified network by, for each training data item in a set of training data, performing forward and back propagation computations through the modified network, where the backward propagation computation through the modified network comprises computing estimated partial derivatives of an error function of an objective for the network, except that the combining node selectively blocks back-propagation of estimated partial derivatives to the subject node, even though activation of the combining node is based on the activation of the subject node.
Owner:D5AI LLC

A machine vision-based monitoring system and equipment for riverbank collapse management

This invention belongs to the field of bank erosion control, specifically a machine vision-based bank erosion control monitoring system and equipment. The monitoring system includes the following steps: Step 1: Building an outdoor intelligent data acquisition platform, its network equipment, and a GPS positioning module; remotely controlling the monitoring equipment to monitor the entire construction process, collecting construction videos and location information, and uploading them to a data center; Step 2: Preprocessing the construction videos, constructing a construction material dataset, inputting the dataset into a YOLOv5-based neural network for training, and optimizing the cost function through forward and backward propagation algorithms; Step 3: Using the trained neural network, identifying important process events, recording key frames of important events occurring at key embankment sections using time as an index, and saving them to the data center; Step 4: Saving the construction videos and key frames collected in Steps 1 and 3 in the database, reading key frames and related information from the data center, and displaying them in the form of charts and reports.
Owner:HUNAN INSTITUTE OF SCIENCE AND TECHNOLOGY +1

An online task screening based meta-learning training method and system

The application discloses an online task screening-based meta-learning training method and system for small sample image classification model training. The existing task-level self-paced meta-learning training adopts a double-pass process, which needs to traverse the complete training round to count all task losses and then replay the training, causing the problems of repeated traversal and high training time consumption. The method of the application comprises the following steps: sampling a main task and performing one gradient-free detection forward propagation to obtain a detection loss; maintaining a sliding loss buffer and an exponential moving average value through an online loss tracker; determining a target retention rate and constructing a dynamic threshold through a course scheduler; comparing the detection loss with the threshold to decide whether to retain or skip; and performing only one backward propagation and optimizer update for the retained task. The double-pass process is transformed into online single-pass screening, the repeated traversal and replay overhead in the same training round is reduced, the additional training time consumption is reduced, and the additional propagation and parameter update overhead of the old multi-step update is reduced in consistency stabilization.
Owner:HAINAN UNIV

A keyword reverse propagation algorithm based on citation network structure

The application provides a keyword reverse propagation algorithm based on a citation network structure, and comprises the following steps: a spring charge model is established, force-directed layout processing is performed, and a force-directed layout graph is established; a keyword propagation model is established by using a reverse propagation algorithm, and a keyword weight change contrast curve is obtained; a citation network model is constructed by using the force-directed layout graph; iterative calculation is performed on the force-directed layout graph until the energy state in the force-directed layout graph reaches a minimum value; while the citation network is iteratively calculated, the keyword weight in the citation network model is adjusted, and a converged citation network layout graph is calculated. The application has the beneficial effects that the entire network is taken as a main body, keywords owned by cited documents are selected at a certain probability and are propagated backward along the network to the citing documents, the weight of the keywords is increased while the text clustering idea is retained, and clear visual data effects are obtained through force-directed layout.
Owner:UNICLOUD TECH CO LTD

Split federated learning method and device based on adaptive pipeline parallelism

The application discloses a kind of based on adaptive flow parallel segmented federal learning method and device. The method comprises: for each device: the device divides bottom submodel into multiple segmented models, sequentially respectively to each segmented model is carried out forward propagation, and the intermediate activation data corresponding to the segmented model after forward propagation is uploaded to server;After server receives intermediate activation data, forward propagation and back propagation are carried out using top submodel, and the gradient of corresponding segmented model is obtained and is issued to corresponding device;For each device: after the gradient is received, the gradient is updated using its set weight, and the set weight is inversely proportional to the update frequency of the device in the current round, and the local bottom submodel is updated using the updated gradient;Forward and backward propagation, data upload, gradient issuing, model updating are carried out in the way of pipeline parallelly. The application accelerates the training of large model on resource-limited edge device and maintains accuracy.
Owner:SUZHOU INST FOR ADVANCED STUDY USTC +1

Three-stage semi-supervised instance segmentation training method and system

A three-stage semi-supervised instance segmentation training method and a system. In a first stage, a teacher model and a student model are trained based on labeled data. In a second stage, the teacher model performs prediction on unlabeled data and generates pseudo labels according to a prediction result. The student model learns labeled data and unlabeled data based on the pseudo labels. A soft label filter is used to filter for obtaining high-quality pseudo labels, a positive sample loss function is used to eliminate a problem of incorrect model convergence due to the pseudo labels with incomplete information, and the parameters of the student model are updated via a backward propagation process. In a third stage, the parameters of the student model are transferred to the teacher model through an exponential moving average operation, so that the teacher model and the student model are under a same architecture.
Owner:LITE ON TECH CORP

Task execution method, visual reasoning method, device and equipment of multi-modal model

This application discloses a multimodal model task execution method, visual reasoning method, apparatus, and device. The method includes: performing a visual reasoning task on first visual data and first text data using a first multimodal model to obtain a first hybrid feature representation; determining a text loss based on the discrete text feature representation; determining a visual loss based on the continuous visual feature representation; and training the first multimodal model based on the text loss and the visual loss to obtain a second multimodal model. Specifically, the forward inference process retains the continuous visual feature branch, eliminating the need to forcibly compress the visual signal into discrete text tokens, thus improving the prediction accuracy of the trained multimodal model in the visual reasoning task. The backward propagation process decouples the learning of the discrete text feature representation from the learning of the continuous visual feature representation, effectively stabilizing the update of the continuous visual features without hindering text convergence, thereby improving the training stability of the multimodal model.
Owner:TENCENT TECH (BEIJING) CO LTD

A flow field reconstruction and deduction method based on K-A theorem

The application provides a flow field reconstruction and deduction method based on K-A theorem, and belongs to the technical field of ship hydrodynamics. First, multi-dimensional data such as the direction, amplitude, period, frequency and flow velocity of irregular waves are collected and integrated into high-dimensional wave vectors. Then, the KAN network with B-spline curve as a learnable function is used to map the high-dimensional input into a low-dimensional parameter list, obtain the amplitude, angular frequency and initial phase of the dominant equivalent regular wave, calculate the decoupling target function by constructing an equivalent wave superposition model, compare the decoupling target function with the true value to obtain a loss function, and then optimize the network parameters by backward propagation until convergence. Finally, based on the network output parameters, a wave superposition model is constructed to represent the real wave in the form of a series of interpretable regular waves superimposed in time sequence, so that efficient and accurate flow field reconstruction and deduction are realized.
Owner:QINGDAO INNOVATION & DEV CENT OF HARBIN ENG UNIV +1

A recalculation control method, device, storage medium and program product

Embodiments of the present application provide a recalculation control method and device, a storage medium and a program product, relating to the technical field of artificial intelligence chips. The method comprises: obtaining a first data amount of activation values generated in a forward propagation process of a to-be-tested model, and obtaining an expected offloading data amount of the activation values that can be offloaded through a target bandwidth between a central processing unit and an artificial intelligence chip; when the first data amount is greater than the expected offloading data amount, it indicates that the target bandwidth cannot offload all the activation values to the memory, so a difference data amount between the first data amount and the expected offloading data amount is used to select part of candidate operators from a plurality of candidate operators included in the to-be-tested model to perform recalculation in a backward propagation process; and the other part of the candidate operators performs activation value offloading in the forward propagation process, which greatly reduces the memory requirement of model training, and avoids bandwidth shortage caused by activation value offloading, thereby improving the overall performance of the artificial intelligence chip during model training.
Owner:SHANGHAI BIREN TECH CO LTD

Method and apparatus with memory management and neural network operation

A processor-implemented memory management method includes: receiving a parameter of a neural network and information of a device configured to perform an operation using the neural network; storing a result of an operation by at least one of layers included in the neural network in a first memory of the device, during a forward propagation operation performed for the neural network based on the parameter; storing a gradient of a layer included in the neural network in a second memory of the device, during a backward propagation operation performed for the neural network based on the parameter and the result of the operation by the at least one layer; and managing the first memory and the second memory based on the information, the result of the operation by the at least one layer, and the gradient.
Owner:SAMSUNG ELECTRONICS CO LTD

A deep learning-based 3d object detection method

The application discloses a 3D target detection method based on deep learning, which comprises the following steps: pre-processing loaded training sample images, calculating a 3D center point of a target, a projection point of the 3D center point on an image, eight corner point positions and a Gaussian distribution of the target center point; constructing a deep learning convolutional neural network, including a main network and two branch networks; loading a data set as a training set, and obtaining an output of the deep learning convolutional neural network through forward propagation of data, calculating a loss degree, performing backward propagation, updating network parameters and obtaining a trained neural network model; in a use stage, receiving test set image data, feeding the image into the pre-trained neural network model, obtaining an output corresponding target, and calculating a 3D position and a category of each target. The 3D target detection method can improve the environmental perception ability of a vehicle in automatic driving.
Owner:AUTOCORE INTELLIGENT TECH (NANJING) CO LTD

Neural network activation compression with non-uniform mantissa

Apparatuses and methods for training a neural network accelerator using quantization precision data formats are disclosed, and in particular for storing activation values from a neural network in a compressed format having lossy or non-uniform mantissa for use during forward and backward propagation training of the neural network. In certain examples of the disclosed technology, a computing system includes a processor, a memory, and a compressor in communication with the memory. The computing system is configured to perform forward propagation for a layer of a neural network to produce first activation values in a first block floating point format. In some examples, the activation values generated by the forward propagation are converted by the compressor to a second block floating point format having a non-uniform and / or lossy mantissa. The compressed activation values are stored in the memory, where they can be retrieved for use during backward propagation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method for modeling training, host and storage apparatus

The present disclosure provides a method for model training, a host, and a storage apparatus. The method may include instructing, by a host apparatus to a graphics processing unit (GPU), to load first intermediate data of a current layer of a model from a dynamic random access memory (DRAM) of a storage apparatus into a memory of the GPU, wherein the current layer is a layer of the model for which computation is being performed; and notifying, the host apparatus to the storage apparatus, to prefetch second intermediate data of a layer of the model to be computed from NAND of the storage apparatus into the DRAM of the storage apparatus based on a space capacity of the DRAM, wherein the first intermediate data of the current layer is used by the GPU in backward propagation of the model to perform the computation for the current layer.
Owner:SAMSUNG ELECTRONICS CO LTD