Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

447 results about "Counter propagation" patented technology

User behavior prediction system and method based on multi-modal data fusion

The invention discloses a user behavior prediction system and method based on multi-modal data fusion, and particularly relates to the field of user behavior prediction, and the system comprises a multi-modal data collection module, a preprocessing and feature extraction module, a cross-modal fusion module, a user behavior prediction module, and a model optimization and feedback module. According to the system, multi-dimensional original data such as visual sense, auditory sense, text, physiological signals and environment context of a user are acquired in real time through a multi-modal data acquisition module; then, deep networks such as ResNet, VGGish and BERT are adopted to extract high-dimensional feature vectors of all modals, and contribution weights of features of different modals are dynamically learned through an attention mechanism; and finally, based on a time sequence model of Transform and LSTM, analyzing fusion features, and outputting probability distribution of future behavior intentions. And parameter joint optimization and continuous learning are realized through a multi-objective loss function and end-to-end back propagation.
Owner:BEIJING DATA100 INFORMATION TECH CO LTD

Assembly line optimization method and device for multi-modal large model training

The embodiment of the invention provides an assembly line optimization method and device for multi-modal large model training, and relates to the technical field of artificial intelligence. According to the method, each model slice of a to-be-trained multi-modal large model is deployed to different training devices, each training batch is divided into a plurality of micro-batches, for each training batch, training sequences corresponding to different training devices are determined according to a parallel scheduling strategy, and the training sequences are used for training the multi-modal large model. The training sequence comprises a forward propagation position, an input back propagation position and a weight back propagation position of each micro-batch, and finally, on each training device, a training process of a corresponding training batch is executed based on the training sequence until training is completed. A back propagation process is divided into back propagation of an input matrix and back propagation of a weight matrix, three calculation stages are jointly formed by the back propagation process and forward propagation, peak shifting calculation can be carried out, bubble time can be effectively filled, idle duration can be reduced, the bubble proportion of an assembly line can be remarkably reduced, and training throughput can be greatly improved.
Owner:PENG CHENG LAB

Robot decision control method based on gradient rarefaction and robot

The invention relates to a robot decision control method based on gradient rarefaction and a robot. The method comprises the following steps: acquiring multi-modal sensor data of a robot; based on a preset sparsification strategy, generating a dynamic mask corresponding to the gradient matrix of the multi-modal large model; based on the generated dynamic mask, screening an effective gradient in a back propagation process of the dynamic mask; updating parameters corresponding to the effective gradient in real time, and obtaining the output of the multi-modal large model based on the updated parameters; according to the obtained multi-modal sensor data and the output of the multi-modal large model based on the updated parameters, feature fusion is carried out, and a combined state code including an environment state, a robot body state and historical decision information is generated; and according to the determined joint state code and based on a time sequence model, generating an action sequence, a force control parameter and a path planning dynamic decision instruction of the robot, so that the robot can act based on the generated dynamic decision instruction, thereby realizing decision control of the robot.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Multivariable time sequence prediction method and system based on implicit neural network

The invention discloses a multivariable time sequence prediction method and system based on an implicit neural network. The method comprises the following steps: 1) collecting data and preprocessing the data; 2) performing window division on the standardized or normalized multivariable time sequence and determining the length of a to-be-predicted window; 3) performing variable correlation coding on the input window to obtain a variable-level feature vector; 4) the implicit neural network based on time attention predicts target parameters by using the variable features in the step 3), and implicit neural representation of the target sequence is modeled through the parameters; 5) taking the output of the implicit nerve representation and the original input window as the input of the multi-head attention predictor, and obtaining a prediction result through cross-sequence cross attention calculation performed in the implicit space and multi-layer perceptron conversion output dimension; and 6) training and optimizing model parameters, calculating a mean square error of a prediction result and a real result, taking the mean square error as a loss function, carrying out back propagation to optimize trainable parameters of the variable correlation coding module, the implicit neural network module multi-head attention predictor and the multi-layer perceptron, and then repeating the steps 3) to 6) to obtain the multi-head attention predictor. Until the preset number of iterations is reached or the error of the model on the verification set meets the requirement of early stop; and 7) performing prediction by using a model of training convergence, and performing reverse normalization on a prediction result to obtain a final prediction result. The method has good generalization, and meanwhile, the interpretability of the attention mechanism is remarkably improved by generating the hidden space characteristics of the trend component and the season component.
Owner:ZHEJIANG UNIV

Vision and GNSS fusion displacement monitoring method and system based on deep learning

The invention relates to a vision and GNSS fusion displacement monitoring method and system based on deep learning. The method comprises the following steps: respectively acquiring visual data and GNSS data, preprocessing the visual data and the GNSS data, and synchronizing the visual data and the GNSS data to obtain synchronized visual data and synchronized GNSS data; performing feature extraction and combination to obtain a modal combination feature sequence; constructing a dynamic weight adaptive fusion network; and training the dynamic weight adaptive fusion network to obtain a displacement monitoring model, and inputting the modal joint feature sequence into the displacement monitoring model to complete displacement monitoring. According to the invention, through a feature extraction network of joint learning, time sequence correlation modeling of vision and GNSS in the same embedding space is realized. And dynamic self-adjustment of the modal weight is realized through a dynamic weight adaptive fusion network. All modules can be micro, and end-to-end back propagation training is supported. And when the visual signal is interfered or the GNSS is interrupted for a short time, the system can carry out self-adaptive compensation, and continuous and stable displacement output is kept.
Owner:SOUTH SURVEYING & MAPPING INSTR

Method and system for detecting residual stress of drawn copper bar

The invention relates to the technical field of metal processing detection, in particular to a copper bar drawing residual stress detection method and system, and the method comprises the steps: collecting production line process state parameters in real time, and mapping the parameters to a predefined multi-dimensional process parameter space to generate a dynamic state vector; adjusting a key coefficient of a material constitutive equation based on the dynamic state vector, and generating a reference stress field; synchronously acquiring multi-physical field data on the surface of the copper bar; constructing a physically constrained machine learning model, and generating a cross-scale fusion stress field; an adjoint optimization equation is constructed based on dynamic process constraints, and model weight parameters are updated through real-time back propagation gradient; and outputting a new cross-scale fusion stress field according to the updated machine learning model, and obtaining related indexes of residual stress detection. According to the invention, the defects of static stiffness and insufficient data fusion of the model in the prior art are effectively overcome, and the real-time detection of the residual stress with adaptive technological parameters is realized.
Owner:扬中凯悦铜材有限公司

VLA model training method and device

The invention provides a training method and device for a VLA model, and relates to the technical field of intelligent robots with bodies, and the method comprises the steps that the VLA model outputs a joint angle sequence of a mechanical arm of a robot with a body based on visual input and language input; the differentiable forward kinematics module outputs a real-time pose of an end effector of the mechanical arm in a task space; based on joint space loss formed by the joint angle sequence and the joint angle in the teaching data, task space loss formed by the real-time pose and the tail end pose in the teaching data, and task space constraint loss, a multi-objective loss function is constructed, and total loss is calculated; calculating the gradient of the total loss to the joint angle sequence and the real-time pose through back propagation, and optimizing preset parameters of the VLA model; the above steps are repeatedly executed until the total loss meets the expectation, and a trained target VLA model is obtained; the technical problems that when a VLA model is trained in a joint space, the hardware coupling performance is high, and the generalization ability is insufficient are solved.
Owner:ANHUI KAIYANG TECHNOLOGY CO LTD +1

Online video instance segmentation method and system based on mask propagation

The invention provides an online video instance segmentation method and system based on mask propagation, and the method comprises the following steps: S1, achieving the cross-frame propagation of a target mask through a mask propagation model, obtaining a prediction mask of a current frame, and achieving the correlation of a target through the calculation of the intersection-union ratio of the prediction mask; and S2, performing back propagation on the target by using the mask propagation model so as to complement the missed target mask and enhance the continuity of the track. Objects between different frames are associated on the basis of a mask propagation mechanism and in combination with the intersection-to-union ratio of masks, and meanwhile, in order to relieve the scene that a segmentation model fails to segment shielded objects, non-significant objects and the like, a reverse mask propagation technology is provided to complement the failed frames. The method aims at performing fine segmentation and continuous tracking on various scenes such as intelligent monitoring, automatic driving, augmented reality and video editing on the target in the video, and has wide research and application prospects.
Owner:FUDAN UNIVERSITY

Segmentation federal learning method and system based on aggregation gradient broadcast

The invention provides a segmentation federal learning method and system based on aggregation gradient broadcast, and the method comprises the steps: initializing a global model, and segmenting the global model into a server model and a client model according to the capability limitation of a local device; issuing the client model to all local equipment terminals; the multiple local devices execute forward propagation at the same time, and shredded data are obtained through calculation; transmitting the shredded data and the training data label to an edge server; performing forward propagation in parallel according to the shredded data, calculating the loss of each local equipment end in combination with a corresponding training data label, and performing back propagation to obtain a loss function and a gradient of the shredded data; aggregating the gradient of the shredded data and broadcasting to all local equipment; updating the server side model and the client side model, and aggregating the server side model; and the iterative learning loop is repeatedly executed until convergence or the maximum communication round is reached. According to the invention, the communication overhead in segmentation federal learning is saved, and the model training efficiency is effectively improved.
Owner:WUHAN UNIV

Large model training method and device, electronic equipment, storage medium and program product

The invention relates to a large model training method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: for any item of training data of a target large model, segmenting the training data into a plurality of parts of segmented data, storing the plurality of parts of segmented data in a nonvolatile memory, and sequentially performing forward propagation calculation and back propagation calculation on the plurality of parts of segmented data; for any part of segmented data, reading the segmented data from a nonvolatile memory to a video memory, and executing forward propagation calculation on the segmented data through a GPU (Graphics Processing Unit) to obtain an activation value corresponding to the segmented data; and for any part of segmented data, executing back propagation calculation based on the activation value corresponding to the segmented data through the GPU to obtain gradient data corresponding to the segmented data, and moving the gradient data corresponding to the segmented data from the video memory to a nonvolatile memory or a CPU memory. The display memory occupation of the activation value can be reduced.
Owner:MOORE THREADS TECH CO LTD

Crystal oscillator circuit fault classification method based on multi-source information fusion network

The invention discloses a crystal oscillator circuit fault classification technology based on a multi-source information fusion network, and belongs to the technical field of analog circuit fault diagnosis. Firstly, a multi-source data set of different measurement points of the crystal oscillator circuit is acquired; carrying out conversion from a time domain to a frequency domain on the data set by utilizing fast Fourier transform; extracting data fault features by using a convolutional neural network, and performing a trust distribution function under each piece of source data by using a softmax classifier; fusing different trust distribution functions under the multi-source data by adopting a D-S evidence theory to obtain a final diagnosis result and diagnosis probability output, and calculating cross entropy loss; and finally, training model parameters through back propagation to obtain a final diagnosis model. According to the method, the time-frequency transformation algorithm, the deep neural network algorithm and the information fusion algorithm are combined, a multi-source information fusion network is constructed, the defects of an existing diagnosis model in crystal oscillator circuit fault classification are overcome, and the accuracy and stability of crystal oscillator circuit fault diagnosis are remarkably improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Multi-modal data feature representation optimization method based on comparative learning and two-stage mask

The invention discloses a multi-modal data feature representation optimization method based on comparative learning and two-stage masks. The method comprises the following steps of: firstly, aiming at an input single image-text pair, generating two groups of heterogeneous image-text data with mode missing as two inputs of a model by applying two random mask strategies with different mask areas and proportions; wherein the first group of data applies a high-proportion mask to the image and applies a low-proportion mask to the text; the second group applies a low-scale mask to the image and a high-scale mask to the text. Then, the two groups of data respectively pass through an image encoder and a text encoder which share parameters, and two different multi-modal fusion feature vectors are generated through a cross-modal fusion encoder; according to the method, masked modal information is recovered through a decoder, and the reconstruction loss of the difference between a recovery result and original data is calculated. Meanwhile, two multi-modal fusion feature vectors generated twice are subjected to comparative learning, and the comparative loss of feature representation distances under different enhanced views for approaching the same image-text pair is calculated. And finally, performing weighted summation on the reconstruction loss and the comparison loss to form an overall loss function, and optimizing model parameters through back propagation. When the trained encoder is used for a downstream multi-modal classification task, the classification precision and generalization ability of the model can be effectively improved.
Owner:NORTHWEST UNIV

Teacher tensor knowledge distillation-driven anti-violation target detection method in electric power safety supervision scene

The invention belongs to the technical field of target detection, and particularly relates to a teacher tensor knowledge distillation-driven anti-violation target detection method in an electric power safety supervision scene, and the method comprises the following steps: inputting image data to a student model and a pre-trained teacher model; performing forward propagation on the student model to obtain a student prediction tensor; associating differences between the real tags based on the student prediction tensor and the input image data; acquiring an original teacher output tensor by using the same input image data by calling a bottom layer forward propagation method of a teacher model; calculating knowledge distillation loss based on the difference between the student prediction tensor and the original teacher output tensor; the standard detection loss and the knowledge distillation loss are combined to form total loss; and performing back propagation updating on parameters of the student model based on the total loss so as to complete model training and real-time target detection and detection method updating optimization. According to the invention, the performance of the student model can be significantly improved, and the high efficiency of the model is maintained.
Owner:HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER +2

Coupling system parameter prediction method based on physical-data co-driven deep neural network CMT-NN

The invention relates to the technical field of optics and the like, in particular to a coupled system parameter prediction method based on a physical-data co-driven deep neural network CMT-NN. The method comprises the following steps: taking optical response spectral lines of a double-micro-ring coupling system under different physical parameters as the input of a CMT-NN model, and taking the output as the physical parameters of a coupling resonance system; the step of constructing the trained CMT-NN model specifically comprises the following steps: S1, constructing a physical model of a target coupling system; s2, constructing a CMT-NN model, bringing physical constraints and physical parameters of the coupling model into a loss function of CMT-NN at the same time, calculating physical loss in a back propagation process through the loss function through iterative training and back propagation, and if a preset condition is met, completing training to obtain the CMT-NN model for predicting the physical parameters of the coupling resonance system; and S3, verification of the CMT-NN model is completed, and a trained CMT-NN model is obtained.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Dynamic risk prediction system

The invention relates to the field of constructional engineering, and discloses a dynamic risk prediction system, which generates space-time alignment input through multi-source data fusion, adopts tensor field modeling to embed contract constraint to construct a risk dynamic model, and solves and outputs a continuous risk field through a partial differential equation; a propagation path is analyzed in combination with asymmetric causal analysis, model parameters are adjusted in real time through a dynamic optimization algorithm, and closed-loop optimization of a risk field is achieved; and finally, through four-dimensional thermodynamic diagram interaction early warning and resource intelligent scheduling, a whole-process closed-loop system of data modeling-causal analysis-dynamic optimization-visual management and control is formed. According to the method, dynamic optimization of risk field parameters is realized through adjoint equation back propagation, and the modeling precision of a complex scene is improved; a four-dimensional space-time thermodynamic diagram rendering technology is innovated to solve the problem of fragmentation of multi-modal information expression, and risk disposal response is accelerated; key task resource supply is guaranteed by combining video memory preemption and containerization scheduling strategies, and the system stability bottleneck in a high-load scene is overcome.
Owner:BEIJING NUO SHICHENG INT ENG PROJECT MANAGEMENT CO LTD

Error adaptive optical diffraction neural network in-situ training method

The invention belongs to the technical field of optical calculation, and discloses an error-adaptive optical diffraction neural network in-situ training method, which comprises the following steps: based on the phase of each phase modulation layer, calculating the light field complex amplitude Uk of the input sample complex amplitude on the rear surface of each phase modulation layer and the light field complex amplitude Uo of the input sample complex amplitude on an output plane after forward propagation; calculating an error light field P based on Uo, carrying out three-dimensional central symmetry on a back propagation light path of the error light field P without changing the phase distribution of each phase modulation layer, and calculating the light field complex amplitude of the complex amplitude of the symmetrical error light field on the front surface of each phase modulation layer after forward propagation; the method comprises the steps of obtaining the light field complex amplitude Un + 1-k'of P back propagation to the rear surface of each phase modulation layer, calculating the gradient related to a loss function based on Uk and Un + 1-k 'so as to carry out updating, and carrying out the next round of training until a preset training round or loss convergence is reached. The method can adapt to errors existing in an optical system.
Owner:HUAZHONG UNIV OF SCI & TECH

Point supervision-based high-resolution remote sensing image small target detection method, system and product

The invention discloses a high-resolution remote sensing image small target detection method, system and product based on point supervision, in a forward propagation stage, point labeling is used as supervision information, a feature aggregation strategy based on a gating mechanism is combined, shallow small target detail features are effectively reserved, feature response scores and edge response scores are combined at the same time, and the detection accuracy is improved. And generating a preliminary pseudo tag through a discrete second derivative method. Through a dynamic multi-instance learning and feature enhancement technology, the distinction degree between candidate frames is improved, and high-quality pseudo labels are screened out for training, so that the performance of a detector is improved. In the training stage, back propagation is carried out, multi-instance learning loss and intersection-parallel ratio loss joint optimization is adopted, and a dynamic weight distribution mechanism is introduced to balance various types of loss. According to the point supervision-based high-resolution remote sensing image small target detection method provided by the invention, the learning ability of a network to a small target sample under a point supervision normal form can be effectively enhanced, and the overall detection performance of a model to a small target is effectively improved.
Owner:WUHAN UNIV

Wind power plant short-term wind power prediction method based on space-time attention mechanism

The invention belongs to the technical field of wind power plant power prediction and intelligent operation and maintenance, and particularly relates to a wind power plant short-term wind power prediction method based on a space-time attention mechanism. According to the method, a dynamic graph evolved along with time is constructed through multi-source heterogeneous data fusion, and a time-varying adjacency matrix is output; obtaining space-time coupling characteristics through disturbance-sensitive space attention coding and wind speed-power coupling sensitivity driven time attention decoding; multi-scale features are output through dual-scale adaptive fusion; and finally, forming a continuously evolved prediction chain through edge end rolling learning and topological closed-loop correction by combining a loss back propagation optimization model of physical consistency constraint. According to the method, the problems of lack of physical constraints, insufficient dynamic adaptation and low efficiency of edge deployment in the prior art are solved, the prediction physical reasonability and dynamic adaptation are improved, efficient edge deployment is realized, and reliable support is provided for wind power plant scheduling.
Owner:XINJIANG YUANXIAO TECHNOLOGY INNOVATION CO LTD

Generative video compression with a transformer-based discriminator

A method, an apparatus, and a non-transitory computer-readable storage medium for video compression using a generative adversarial network (GAN) are provided. The method includes obtaining, by a generator of the GAN, a reconstructed target frame based on a reference frame and a raw target frame to be reconstructed; concatenating, by a transformer-based discriminator of the GAN, the reference frame, the raw target frame and the reconstructed target frame to obtain a paired data; determining, by the transformer-based discriminator of the GAN, whether the paired data is real or fake to guide reconstruction of the raw target frame; and determining a generator loss and a transformer-based discriminator loss, and performing gradient back propagation and updating network parameters of the GAN based on the generator loss and the transformer-based discriminator loss.
Owner:SANTA CLARA UNIVERSITY +1

Text-to-structured query statement method based on large language model and reinforcement learning

The invention provides a text-to-structured query statement method based on a large language model and reinforcement learning, and relates to the technical field of information, the method comprises the following steps: integrating a natural language problem and a database mode into a unified prompt template to obtain a candidate structured query, and generating a reference structured query by a reference model; constructing a multi-dimensional reward framework, and performing weighted aggregation on multi-dimensional rewards to obtain corresponding final reward scores; calculating a KL penalty value between the output of the strategy model and the output of the reference model based on the candidate structured query and the reference structured query to obtain a constraint of stable strategy update; updating the strategy model parameters through back propagation iteration to obtain an updated strategy model; and analyzing the prompt template by using the updated strategy model to obtain a text-to-structured query statement processing result, and completing the process from the text to the structured query statement. The problems that existing structured query generation is low in accuracy and semantic consistency is difficult to guarantee are solved.
Owner:CHENGDU UNIV OF INFORMATION TECH

Model quantitative perceptual training system on edge device

The invention belongs to the technical field of edge devices, and particularly relates to a model quantitative perception training system on edge devices, which comprises a data acquisition module used for acquiring data from various edge devices; the data preprocessing module is used for cleaning, normalizing and enhancing preprocessing operation on the collected data so as to improve the quality and diversity of the data and provide a better data basis for subsequent model training; and the quantitative perception training module is used for applying a quantitative perception training technology in a training process so as to simulate a quantization operation in forward propagation of the model, performing quantization and inverse quantization on the weight and the activation value, and calculating a gradient by adopting a gradient approximation method in back propagation so as to optimize model parameters. According to the method, through quantitative perception training, the weight and the activation value of the model can be converted into low-precision representation, and the calculation amount when the model runs on edge equipment is greatly reduced.
Owner:RUIBO (BEIJING) ARTIFICIAL INTELLIGENCE TECH CO LTD

GNN-PINN coupled multi-lane heterogeneous traffic joint modeling method and system

The invention discloses a multi-lane heterogeneous traffic joint modeling method and system based on a GNN-PINN coupling architecture. The method comprises the following steps: constructing a space-time dynamic graph coding lane changing rule through a GNN, embedding a traffic flow physical constraint through a PINN, and solving a source item; according to the method, CAV behaviors are described by using improved IDM in the mid-microscopic view, an LWR equation subjected to GNN lane change flow correction is solved by using PINN in the macro view, GNN and PINN closed-loop optimization is realized by combining a double-path back propagation mechanism, and a model is optimized based on a residual loss function fusing physical loss, data driving items and boundary constraints. The system comprises a dynamic graph topology module, a GNN coding module, a PINN solving module and the like, the multi-lane traffic dynamic description precision can be improved, the multi-scale characteristic is considered, real-time simulation is supported, and the system is suitable for heterogeneous traffic modeling.
Owner:CHENGDU JIAOTOU INTELLIGENT TRANSPORTATION TECHNOLOGY SERVICE CO LTD

Target detection method and device based on domain self-adaption, equipment and medium

The invention discloses a target detection method and device based on domain self-adaption, equipment and a medium. The target detection method based on domain adaptation is realized through a target detection model, the target detection model comprises a dynamic domain adaptation module, and the dynamic domain adaptation module comprises an image level domain classifier with a first dynamic confrontation gradient inversion layer and an object level domain classifier with a second dynamic confrontation gradient inversion layer. In the back propagation of the training process of the target detection model, the gradient inversion intensity of the first dynamic confrontation gradient inversion layer is dynamically adjusted according to the image level domain classification loss, and the gradient inversion intensity of the second dynamic confrontation gradient inversion layer in the back propagation is dynamically adjusted according to the object level domain classification loss; according to the method, the target detection precision under the severe weather condition can be improved, the target missing detection rate and the target false detection rate under the severe weather condition are reduced, and then the safety and the reliability of the automatic driving system are improved.
Owner:TIANJIN PORT (GROUP) COMPANY

Wind power plant short-term wind power prediction method based on space-time attention mechanism

The invention belongs to the technical field of wind power plant power prediction and intelligent operation and maintenance, and particularly relates to a wind power plant short-term wind power prediction method based on a space-time attention mechanism. According to the method, a dynamic graph evolved along with time is constructed through multi-source heterogeneous data fusion, and a time-varying adjacency matrix is output; obtaining space-time coupling characteristics through disturbance-sensitive space attention coding and wind speed-power coupling sensitivity driven time attention decoding; multi-scale features are output through dual-scale adaptive fusion; and finally, forming a continuously evolved prediction chain through edge end rolling learning and topological closed-loop correction by combining a loss back propagation optimization model of physical consistency constraint. According to the method, the problems of lack of physical constraints, insufficient dynamic adaptation and low efficiency of edge deployment in the prior art are solved, the prediction physical reasonability and dynamic adaptation are improved, efficient edge deployment is realized, and reliable support is provided for wind power plant scheduling.
Owner:LANXIAN HUYUETONG DASHETOU WIND POWER CO LTD +1

Memory efficient neural network training method and system

A memory efficient neural network training method and system. A forward pass is performed by inputting a batch of training data to a neural network. A loss is determined from an output of the neural network resulting from the forward pass, and back propagation is performed as part of the training on the neural network. Performing the back propagation involves, for each layer of the neural network, determining a gradient or optimizer state for the layer of the neural network, and compressing the gradient or optimizer state by performing a random down-projection on the gradient. Following determining and down projecting the gradients or optimizer states for the layers of the neural network, the gradients or optimizer states based on the gradients are decompressed, and the weights of the neural network are updated based on the decompressed gradients or optimizer states. A different random down-projection is used for each layer.
Owner:ROYAL BANK OF CANADA

Chip working temperature and thermal stress prediction method based on physical information neural network

The invention discloses a chip working temperature and thermal stress prediction method based on a physical information neural network. The invention aims to solve the technical problems that the chip temperature and thermal stress prediction precision is low, the data-driven network interpretability is poor, and a large amount of training data is needed. The method comprises the following steps of: 1, simulating and establishing a chip working temperature and thermal stress prediction data set through a high-precision chip simulation model, taking a space-time coordinate of the chip as an input parameter, and taking the working temperature and thermal stress as labels; 2, establishing a deep neural network containing a full connection layer and embedded physical information in a normalization mode, adopting an SELU activation function and an SGD optimizer to perform back propagation, and combining a mean square error, a heat conduction equation and a thermal stress calculation equation as a loss function; 3, training the deep neural network; and 4, efficiently predicting the temperature and the thermal stress of the actual chip.
Owner:SHANGHAI UNIV

TEE-GPU collaborative model credible training method and device based on parameter confusion

The invention discloses a TEE-GPU collaborative model credible training method and device based on parameter confusion, and the method comprises the steps: enabling a client to encrypt training data and a model architecture in a preprocessing stage, and uploading the encrypted training data and model architecture to a server; and the trusted execution environment of the server side decrypts the model architecture, carries out confusion processing on a linear layer and then deploys the linear layer on an external GPU (Graphics Processing Unit). In the training stage, initial forward propagation calculation of training data is firstly completed in the TEE, and then intermediate results are confused and then transmitted to the GPU so as to execute subsequent forward propagation calculation. In the back propagation process, extra confusion is applied to the gradient by the TEE, and the gradient is issued to the GPU to calculate a new gradient; and after receiving the confusion gradient, the GPU updates the parameters in a confusion form. And randomly sampling calculation data in the TEE, and carrying out integrity verification to ensure the integrity of a calculation result. Compared with an existing credible training scheme, the method has the advantages that the time overhead can be remarkably reduced while the calculation accuracy, integrity and privacy of the model are ensured.
Owner:WUHAN UNIV

Film time sequence large model data prediction method based on ridge regression constraint attention

The invention discloses a slice time sequence large model data prediction method and system based on ridge regression constraint attention, a medium and equipment, and the method comprises the steps: carrying out the normalization processing of multivariate time sequence data of each channel, segmenting the multivariate time sequence data into a plurality of overlapped slice time sequences according to a fixed length and a step length, and obtaining the slice time sequence of each channel; lLM recoding: embedding and recoding a slice time sequence into a text prototype space by using a multi-head cross attention mechanism, and splicing a text and data to obtain time sequence data with prompts and attention scores; constructing an interaction channel code to model an interdependency relationship between channels, and obtaining potential feature representation based on time sequence data with prompts; the LLM outputs projection, the potential feature representation is sent into the LLM for prediction, and a prediction output sequence is generated through linear projection; and carrying out ridge regression attention regularization, including the attention score into a loss function, and optimizing the model through back propagation.
Owner:XI AN JIAOTONG UNIV

Large language model control fine tuning method and system based on multi-task cooperative regulation and control

The invention discloses a large language model control fine tuning method and system based on multi-task cooperative regulation and control, and belongs to the technical field of large language models. Generating a gating coefficient and an initial dynamic evaluation signal through the intelligent regulation and control network; a task difficulty index is obtained by combining multi-index fusion and historical moving average, and the sampling probability and the exclusive learning rate are dynamically adjusted; weighting the fusion gradient and carrying out back propagation to update parameters; and closed-loop feedback monitoring is carried out and parameters of the regulation and control network and the scheduling policy device are optimized. The system comprises a data coding module, a collaborative intelligent regulation and control module, a dynamic balance control module, a joint optimization module and a closed-loop feedback module. According to the method, gradient conflicts among tasks are relieved, the problems of convergence instability and performance imbalance are solved, the multi-task training efficiency, convergence stability and generalization ability of a large language model are improved, and the method can be widely applied to multi-class multi-task learning scenes.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Model training method and device, storage medium and program product

The invention relates to a model training method and device, a storage medium and a program product. The method comprises the following steps: in a forward propagation stage, for any check point module, storing input and output of the check point module in a video memory, and releasing an intermediate activation value of the check point module in the video memory; the input of the check point module is used for recalculating an intermediate activation value, and the output of the check point module is used for forward calculation of subsequent modules of the check point module; and in a back propagation stage, for any check point module, responding to the condition that the last layer of the check point module is a linear layer, skipping forward calculation of the last layer, calculating a gradient according to a gradient formula corresponding to the last layer, and completing back propagation of the check point module according to the gradient of each layer in the check point module. According to the method and the device, the calculation cost can be remarkably reduced while the calculation precision and the video memory saving amount are the same as those of a standard re-calculation scheme.
Owner:MOORE THREADS TECH CO LTD