Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

828 results about "Forward propagation" patented technology

Engineering safety progress intelligent monitoring method based on multi-source data collaboration

The invention discloses an engineering safety progress intelligent monitoring method based on multi-source data collaboration, which relates to the technical field of intelligent engineering monitoring, and comprises the following steps of: mapping multi-source engineering monitoring data into nodes and edges of a graph in real time by utilizing an incremental graph updating algorithm, generating a dynamic knowledge graph, and generating a dynamic mapping result; performing graph traversal on the dynamic knowledge graph through an association subgraph extraction algorithm, extracting a security event matrix and a project progress matrix, calculating an SPI index value by using a dynamic weighted fusion algorithm, synchronously performing multi-threshold interval grading on the SPI index value, generating an SPI early warning level, performing state coding on the SPI early warning level, and generating a comprehensive feature vector; according to the method, data of different types can be effectively integrated and a basis is provided for formulating a targeted engineering safety progress solution through an incremental graph updating algorithm and a Bayesian causal graph model, and node probability distribution characteristics are extracted by using a forward propagation layer. The method is advantaged in that the incremental graph updating algorithm and the Bayesian causal graph model are utilized to effectively integrate data of different types and provide a basis for formulating a targeted engineering safety progress solution.
Owner:SHAANXI HUISHENG SPACE-TIME INFORMATION TECH CO LTD

Training method, data processing method, electronic equipment and computer readable storage medium

The invention relates to a training method, a data processing method, electronic equipment and a computer readable storage medium, and relates to the technical field of computers. The training method comprises the steps that a training sample is input into a to-be-trained model, forward propagation is executed, output of the to-be-trained model is obtained, the to-be-trained model comprises a plurality of expert layers, and a specific activation value is not stored in the forward propagation process; determining the value of a loss function according to the output of the to-be-trained model; according to the value of the loss function and the loss function, back propagation is executed, in the back propagation process of each expert layer, a recalculation task of a specific activation value and a first communication task are executed in an overlapping mode, and for each expert layer, the first communication task comprises a communication task of data interaction among multiple devices corresponding to the expert layer.
Owner:BYTEDANCE TECHNOLOGY CO LTD +1

Semantic alignment-based large language model equipment life prediction method and system

The invention belongs to the technical field of industrial equipment life prediction, and discloses an equipment life prediction method and system of a large language model based on semantic alignment, and the method comprises the steps: obtaining original multi-dimensional sensor time sequence data, and obtaining an embedded matrix after preprocessing; constructing a prompt text with domain semantics to obtain a natural language embedded representation; constructing a semantic text prototype, and realizing alignment of the embedding matrix and the semantic text prototype to obtain a patch embedding sequence; and splicing the patch embedding sequence and the natural language embedding representation, inputting the spliced patch embedding sequence and the natural language embedding representation into a pre-trained large language model for forward propagation, extracting hidden vectors output corresponding to the patch embedding sequence, splicing and flattening the hidden vectors into a single vector, and outputting to obtain an equipment life prediction result. According to the method, cross-modal knowledge learned by the LLM in large-scale pre-training and the powerful reasoning ability are fully utilized, and accurate prediction of the residual life of the equipment is achieved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Method for predicting workability of crane ship based on BP (Back Propagation) neural network

The invention relates to the technical field of crane ship operation prediction, and discloses a BP neural network-based crane ship operability prediction method. The method comprises the following steps: collecting and preprocessing environmental data of a crane ship operation sea area, classifying and extracting features according to a preset operation state, and dividing a training test set; detecting sample data quality, and marking abnormal samples; using an input feature set of a non-abnormal sample and a historical operation state label matching result to train a model, and generating a BP neural network prediction model; the model parameters are optimized and corrected based on the incidence relation between the weight offset parameters and the prediction errors; and inputting real-time environment data, calculating an output probability along a network forward propagation path, determining the operability of the crane ship, and generating a prediction result. The method can effectively process the non-linear relationship between environmental factors, reduce abnormal data interference, improve the accuracy, objectivity and consistency of prediction, and provide a reliable basis for operation decision making of the crane ship.
Owner:CCCC THIRD HARBOR ENGINEERING CO LTD

Fine tuning system of large-scale pre-training model in federated learning environment and application thereof

The invention discloses a fine tuning system of a large-scale pre-training model in a federated learning environment and application thereof. The system comprises a local disturbance gradient estimation module, a differential privacy protection module and a global model aggregation and update module. The local disturbance gradient estimation module is used for calculating a global model loss value by combining forward propagation with a zero-order optimization method so as to estimate a gradient and realize all-parameter fine tuning; the differential privacy protection module performs differential privacy protection processing on the estimated disturbance gradient to prevent gradient information from leaking user sensitive data; and the global model aggregation and update module reconstructs a disturbance vector and completes global model update based on a random seed and a scalar gradient uploaded by a client. Compared with the prior art, on the premise of not depending on back propagation, all-parameter fine tuning of a large-scale pre-training model is achieved, data privacy is guaranteed, meanwhile, calculation and memory expenses are remarkably reduced, and the method is suitable for a resource-limited distributed calculation environment.
Owner:SHENZHEN MSU-BIT UNIVERSITY

Marine main engine power real-time optimization method based on hybrid network model

The invention provides a ship main engine power real-time optimization method based on a hybrid network model, and belongs to the technical field of ship energy saving.A hybrid neural network model is constructed, the hybrid neural network model is based on a physical information neural network, a KAN network is introduced to serve as a front-end network structure, and the power of a ship main engine is optimized in real time; the high-dimensional nonlinear mapping module is used for establishing high-dimensional nonlinear mapping from navigational speed, a ship type parameter matrix, propulsive efficiency, fuel conversion efficiency and environmental factors to the minimum power of a main engine; a composite loss function is adopted to train the hybrid neural network model, wherein the composite loss function is formed by weighted summation of mean square error loss, dynamics constraint loss, propulsive efficiency constraint loss and fuel consumption constraint loss; using the trained hybrid neural network model to receive ship operation parameters and environment parameters collected in real time, and outputting a minimum power prediction value of the ship main engine through one-time forward propagation calculation to realize real-time optimization of the main engine power.
Owner:QINGDAO INNOVATION & DEV CENT OF HARBIN ENG UNIV +1

Assembly line optimization method and device for multi-modal large model training

The embodiment of the invention provides an assembly line optimization method and device for multi-modal large model training, and relates to the technical field of artificial intelligence. According to the method, each model slice of a to-be-trained multi-modal large model is deployed to different training devices, each training batch is divided into a plurality of micro-batches, for each training batch, training sequences corresponding to different training devices are determined according to a parallel scheduling strategy, and the training sequences are used for training the multi-modal large model. The training sequence comprises a forward propagation position, an input back propagation position and a weight back propagation position of each micro-batch, and finally, on each training device, a training process of a corresponding training batch is executed based on the training sequence until training is completed. A back propagation process is divided into back propagation of an input matrix and back propagation of a weight matrix, three calculation stages are jointly formed by the back propagation process and forward propagation, peak shifting calculation can be carried out, bubble time can be effectively filled, idle duration can be reduced, the bubble proportion of an assembly line can be remarkably reduced, and training throughput can be greatly improved.
Owner:PENG CHENG LAB

Model training method and device, electronic equipment and storage medium

The invention discloses a model training method and device, electronic equipment and a storage medium. Comprising the following steps: sampling an original video frame sequence based on a target training temporal-spatial resolution to obtain target training data; inputting the target training data into a teacher model, generating a target token set through dynamic token selection, and generating teacher training features through forward propagation; performing multi-scale cutting on the target token set according to a target self-attention weight of the teacher model to generate at least three student training masks with different token numbers; inputting the target training data and the different student training masks into a student model for forward propagation, and generating student training features; and performing alignment distillation on the student training features and the teacher training features to obtain a target student model. The defect of a video understanding model in downstream flexible reasoning is overcome, and dynamic token selection and multi-scale mask training under high temporal-spatial resolution are utilized, so that the model can obtain better performance under various downstream calculated amount limits.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Digital twinborn calibration framework construction method based on Lyapunov strategy

The invention discloses a digital twin calibration framework construction method based on a Lyapunov strategy, and the method comprises the steps: constructing a digital twin model containing a proxy network to simulate the dynamic state of a system, and defining a state, an action space and a reward function through a Markov decision process. A constrained Lyapunov action-commentator (CLAC) algorithm is introduced, a strategy network and a Lyapunov network are optimized, and the algorithm can stably act under high noise and deviation. Real-time parameter optimization is realized by means of single-time neural network forward propagation, an experience playback pool and the like. Each network structure is clear, and a specific initialization and optimization method is adopted. According to the method, a calibration problem can be converted into a parameter tracking task, efficient training is performed under unmarked data, constraint requirements such as stability can be met, and the method is suitable for scenes such as industrial robot joint control.
Owner:CHINA YANGTZE POWER

Underwater acoustic target recognition method based on effective receptive field regulation

Disclosed in the present invention is an underwater acoustic target recognition method based on effective receptive field regulation. A proposed AEU-Net model has a plurality of resolution branches, with each resolution branch having independent convolution kernels, which adaptively adjust their sizes during training while aligning with their respective resolutions. The AEU-Net model comprises an ERF-Server, which can regulate an effective receptive field during forward propagation, and has two operations of increasing or decreasing the effective receptive field of a feature map of any designated module. In the present invention, acoustic physical information of an underwater acoustic target sample can be captured in a plurality of interaction dimensions, and feature fusion of information from effective receptive fields across a plurality of scales is performed, so as to adapt to sonar image targets of different sizes, thereby improving the accuracy and recognition speed in the recognition of an underwater acoustic target.
Owner:ZHEJIANG UNIV

STFT dimension transformation-based spiking neural network mechanical fault diagnosis method

The invention is applied to the field of mechanical fault diagnosis signal processing, and particularly provides a pulse neural network mechanical fault diagnosis method based on STFT dimension transformation, and the method comprises the steps: collecting a one-dimensional mechanical vibration signal, carrying out the wavelet decomposition, carrying out the wavelet reconstruction of a low-frequency component and a denoised high-frequency component, and carrying out the wavelet reconstruction of the low-frequency component and the denoised high-frequency component; obtaining a denoised one-dimensional vibration signal; performing short-time Fourier transform, and converting the time-frequency two-dimensional matrix into a time-frequency two-dimensional matrix; inputting the time-frequency two-dimensional matrix into an improved HH threshold neuron model, carrying out Poisson sparse coding on the time-frequency two-dimensional matrix, and only carrying out pulse response on signal significant features; constructing a suprathreshold coding convolutional network with residual connection, inputting a sparse coding matrix, training by adopting an unsupervised learning rule based on STDP, and adaptively adjusting a network synaptic weight; and inputting to a trained above-threshold coding convolutional network, and obtaining pulse emission activity of neurons of an output layer through network forward propagation to determine a fault diagnosis result.
Owner:WESTLAKE INSTITUTE FOR OPTOELECTRONICS

Large language model weight and activation combined quantification method and system

The invention discloses a large language model weight and activation combined quantification method and system, and belongs to the technical field of model quantification. The method comprises the following steps: collecting and preprocessing a calibration set, inputting a large language model to execute forward propagation, and recording an activation matrix; for each embedding dimension, counting maximum activation absolute values of all lexical elements on the dimension; generating a global threshold value by combining a quantile statistical method with the global sensitivity coefficient, and determining that the dimension with the maximum activation absolute value exceeding the global threshold value in the dimensions is an outlier dimension; respectively designing scaling factors for the normal dimension and the outlier dimension, and generating a reconstruction weight matrix; using Bayesian-gradient joint optimization to reconstruct a truncation threshold value of the weight; calculating a scaling factor of the reconstructed weight matrix to obtain a reconstructed quantized weight matrix; quantizing the activation matrix of the current layer according to the embedding dimension by applying a scaling factor; and performing multiplication calculation with the reconstructed quantization weight matrix to obtain a multiplication output result in an integer domain, and then performing unified inverse quantization recovery.
Owner:NANJING UNIV OF POSTS & TELECOMM

Intelligent management platform system for whole-process data integration of seawall engineering construction

The invention relates to the technical field of intelligent construction and digital twinning, and discloses an intelligent management platform system for seawall engineering construction whole-process data integration, and the system comprises a multi-source heterogeneous data collection module which is used for obtaining and standardizing construction data; the core data fusion and reasoning module is used for constructing a construction process causal atlas representing the relationship between elements and calculating the credibility of each state node through a multi-source evidence theory fusion engine; the intelligent application and evolution module is used for carrying out risk forward propagation early warning and root cause backward tracing diagnosis based on the construction process causal atlas; and the comprehensive visualization and decision support module is used for performing fusion display on the analysis result and the BIM / GIS model. According to the method, the causal knowledge graph and evidence fusion reasoning mechanism is constructed, the problems of data islands and information uncertainty are solved, active early warning and accurate traceability of construction risks and dynamic self-adaption of the system are achieved, and the intelligent management level of seawall engineering is remarkably improved.
Owner:ZHEJIANG HYDROPOWER CONSTR & INSTALLATIONCO

Edge end spiking neural network compression and deployment method and system

The invention relates to an edge end spiking neural network compression and deployment method and system, and belongs to the technical field of machine learning, the edge end spiking neural network compression and deployment method performs forward propagation on an initial spiking neural network model based on acquired multi-modal data, obtaining a memory use state in a forward propagation process based on hardware sensing initialization, and performing dynamic sparsification on the initial pulse neural network model based on the memory use state to construct a sparsified pulse neural network; based on the acquired activation statistical information of the sparse spiking neural network, performing dynamic pruning on the sparse spiking neural network by adopting a pruning strategy based on neuron activeness so as to realize compression of a spiking neural network model; the compressed spiking neural network model is optimized according to the edge end configuration, and the optimized spiking neural network model is deployed to the edge end to execute the reasoning task, so that the storage requirement and the calculation complexity are reduced.
Owner:HUBEI ENG UNIV

Neural network global one-time structured pruning method, system and device and medium

The invention relates to a neural network global one-time structured pruning method, system and device and a medium, and the method comprises the steps: carrying out the parallel capturing of the input activation tensors of all target layers in a to-be-pruned neural network through a calibration data set through single-time forward propagation; according to an input activation tensor, synchronously calculating a difference entropy index and amplitude response intensity for an intermediate neuron weight group of each target layer, and performing normalization and fusion to form a static global importance map; determining an importance threshold according to a preset pruning rate, and generating a global to-be-pruned index set of which the mixed importance score is lower than the importance threshold at one time based on the global importance map; and on the basis of the index set, performing one-time physical structured pruning on the weight matrix of each target layer. Therefore, the static global importance map is generated through single forward propagation and parallel capture of the activation tensor, and maskless one-time pruning is completed through physical structured pruning.
Owner:SHANGHAI BANGTU INFORMATION TECH CO LTD

Deterministically defined, differentiable, neuromorphically-informed i / o-mapped neural network

A system includes a neural network architecture. It a new type of neural network able to process statically mapped as well as temporally sequenced information with much better power utilization, data requirements and operational efficiencies. Unlike prior artificial neural network approaches, the present invention includes uniquely defined sets of relationships. The unique use of non-linear input-output mapping functions combined with a time-variant pilot function, and a deterministically bounded, fully-differentiable, nonlinear resonance field subsystem allows the present invention to be readily deployed to work with virtually any neural network architecture / implementation including photonic, opto-acoustic or other variants. This dramatically reduces the size and complexity of virtually any neural network architecture because it offloads what would otherwise need to be done in the form of back / forward propagation trained weights and biases to much simpler, more scalable differentiable input / output mapping functions.
Owner:ZON GLOBAL IP INC

Out-of-distribution prediction

A set of features of a training document are identified in a training document for training a machine learning model. A subset of the features is selected to be omitted from a training forward propagation. As a result of omitting the subset of the set of features, a different subset of the set of features is used to train the machine learning model to classify documents and distinguish between an out-of-domain document and in-domain document.
Owner:CITIGROUP

Model training method and device, equipment and storage medium

The invention provides a model training method and device, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of neural network models and model training. The specific implementation scheme is as follows: a calculation unit executes quantization matrix multiplication based on Hadamard pre-transformation on an activation tensor and a weight tensor of a target model stored in a memory so as to generate an output tensor of a linear layer based on a low-precision tensor with smaller data bit width; using the output tensor and a subsequent network layer of the target model to complete forward propagation so as to obtain a loss value; and according to the loss value, updating model parameters of the target model stored in a memory through a back propagation algorithm. By means of the technical scheme, on the premise that the model training precision is guaranteed, memory resource occupation and the calculation amount in the calculation process can be remarkably reduced, and the training cost is reduced.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Elevator fault diagnosis method and system

The invention relates to the technical field of elevator fault diagnosis, and discloses an elevator fault diagnosis method and system.The elevator fault diagnosis method comprises the steps that elevator operation data are collected, operation stages are divided for the elevator operation process according to the elevator operation data, and the residual error of the elevator operation data is calculated in the operation stages; based on an elevator structure, a causal chain graph reflecting the relation between components is established through fault mode analysis, and a direct upstream node set of downstream nodes in the causal chain graph serves as a parent set of the downstream nodes; in the operation stage, performing block replacement processing for keeping time sequence characteristics on an upstream node residual sequence of each edge in the causal chain graph to generate a contrast residual sequence, and calculating the propagation intensity of the edge according to the contrast residual sequence; effective propagation paths are screened according to the propagation intensity of the edges, the effective propagation paths are sorted to obtain candidate nodes, residual errors of the candidate nodes are set to be zero through virtual pinch-off, forward propagation is carried out through a semi-physical model, and fault nodes are determined in the candidate nodes based on forward propagation results.
Owner:HUNAN ELECTRICAL COLLEGE OF TECH

Medical image segmentation method based on multi-branch distillation

The invention belongs to the field of image processing, and discloses a medical image segmentation method based on multi-branch distillation, which comprises the following steps: constructing and training a multi-branch collaborative distillation image segmentation model based on uncertainty perception, and establishing a network architecture comprising a public encoder, a first decoder and a second decoder, including a guiding decoder and a guided decoder; dropout disturbance is differentiated on the output of the public encoder, multiple features are generated, multi-branch forward propagation is executed, and a virtual label is generated through fusion of a CTF module; calculating supervision loss and distribution alignment loss of the original feature map through a guided branch; performing multi-level consistency constraint on the output of the disturbance characteristic graph by the guided model to realize multi-view structure consistency, automatically balancing parameter weights through an adaptive task balancing mechanism, and constructing a total loss function; and obtaining a to-be-segmented medical image, and inputting the to-be-segmented medical image into the trained image segmentation model to obtain a segmentation result. According to the invention, high-quality medical image segmentation under a low marking rate can be realized.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Tensor segmentation and mapping method of neural network operator on wafer chip

The invention provides a tensor segmentation and mapping method of a neural network operator on a wafer chip, and the method comprises the steps: expanding a forward calculation graph into a training calculation graph containing forward propagation, back propagation and gradient updating, and employing different operators in the calculation of these stages; carrying out dimension segmentation on the input tensor of each operator in the training calculation graph to generate a plurality of global candidate segmentation strategies; aiming at each candidate segmentation strategy, enumerating feasible resource mapping strategies of all operators, and combining to generate candidate design points; and screening an optimal design point through cost evaluation, and outputting a corresponding candidate segmentation strategy and a resource mapping strategy. According to the method, forward and backward operators in a neural network training process are decoupled through expansion of a calculation graph, and a segmentation scheme for avoiding tensor copying in forward and backward calculation in the training process and a resource mapping strategy corresponding to the segmentation scheme are explored, so that tensor copying is avoided, global resource occupation is reduced, and the training efficiency is improved. Therefore, the same wafer-level chip can support larger model training.
Owner:TSINGHUA UNIVERSITY

Method and device for predicting multi-working-condition flow field of wind driven generator based on neural network

The invention relates to the technical field of wind power generation, artificial intelligence and fluid mechanics, and discloses a prediction method and device for a multi-working-condition flow field of a wind driven generator based on a neural network, and the prediction method comprises the steps: obtaining a sample data set which comprises multiple groups of multi-working-condition data; and inputting the sample data set into a physical information field adversarial neural network model for training to obtain the total loss of forward propagation of the sample data in the field adversarial neural network model. According to the total loss, determining whether training of the wake flow field prediction model is completed; and under the condition that a new round of training is carried out on the wake flow field prediction model, back propagation is carried out on the total loss, and neural network parameters are optimized. In this way, a dual-branch loss collaborative optimization mechanism is formed. A physical information neural network and a domain adversarial neural network are combined, and complementary advantages of the two are fully exerted. When the data is limited or the distribution difference is large, high-precision and physically consistent wake flow field prediction can be realized.
Owner:OCEAN UNIV OF CHINA

Model reinforcement learning method, device and equipment

The embodiment of the invention provides a model reinforcement learning method, device and equipment. The scheme comprises the following steps: in a sampling stage, using an inference engine and adopting a to-be-trained target model to generate an output sequence for an input sequence under a first strategy parameter, recording a first probability value of each lexical element in the output sequence, and calculating a dominant value of each lexical element, in a training stage, after a training engine is used for forward propagation to obtain a second probability value of each lexical element generated by a target model under a first strategy parameter, a first ratio of the second probability value to the first probability value of each lexical element can be calculated, and then the lexical elements with the first ratios within a preset numerical range are screened out to participate in calculation of a target function; and optimizing the target function to update the parameters of the target model.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Construction method of coefficient prediction model, coefficient prediction method and storage medium

The invention relates to a construction method of a coefficient prediction model, a coefficient prediction method and a storage medium, and the construction method comprises the steps: carrying out the modeling of a bridge prestressed beam through finite element numerical software, simulating a tensioning process, constructing a bridge parameter coefficient array, dividing the bridge parameter coefficient array into a training data set and a verification data set, performing data standardization and tensor data conversion, and constructing a data loader; an AND-based full-connection multi-layer multi-task neural network is constructed, and a standardizing device, a forward propagation function, a loss function, a learning rate scheduler and an optimizer are arranged; and carrying out model training, packaging a prediction function, and integrating a standardizing device and the trained prediction model. The prediction method comprises the step of predicting the friction resistance loss coefficient and the pipeline deviation coefficient according to the prediction model. According to the method and the device, coefficient prediction can be quickly, efficiently and accurately realized by only needing several on-site easy-to-measure parameters.
Owner:JILIN JIANZHU UNIVERSITY

Teaching object interaction method, system and device

The invention discloses a teaching object interaction method, system and device. The teaching object interaction method comprises the steps of obtaining a plurality of interaction behavior characteristics of a real-time user at a current knowledge point; constructing a plurality of interaction behavior characteristics of the real-time user at the current knowledge point into a real-time interaction vector; inputting the real-time interaction vector into a learning recognition model, performing forward propagation after pre-training in the learning recognition model, and generating a mastering state label of a real-time user at the current knowledge point; according to the mastering state label of the real-time user at the current knowledge point, constructing a teaching interaction strategy of a plurality of to-be-learned knowledge points in a future time axis; wherein the mastering state labels of the plurality of knowledge points to be learned are inherited in a training sample of the learning recognition model; according to the method, the accuracy of learning path deduction and the individuation of learning strategies are improved, and the user learning experience of a teaching system is remarkably improved by combining real-time mastering state prediction and deduction based on a historical behavior mode.
Owner:JIAN COLLEGE

Data parallel communication method and device in distributed training, storage medium and program product

The invention relates to a data parallel communication method and device in distributed training, a storage medium and a program product. The method comprises the following steps: for any computing device in a computing device cluster participating in distributed training of a target model, obtaining global latest parameters of the target model through global collection operation in a forward propagation process, according to the training data subsets corresponding to the computing devices and the global latest parameters of the target model, forward propagation calculation is carried out, loss values corresponding to the computing devices are obtained, and different computing devices in the computing device cluster carry out forward propagation calculation based on different training data subsets; and for any computing device, calculating a gradient corresponding to the computing device based on the loss value corresponding to the computing device in a back propagation process, and transmitting the gradient corresponding to the computing device to other computing devices in the computing device cluster through reduction and dispersion operation. The communication overhead can be reduced, and the efficiency of large-scale training can be improved.
Owner:MOORE THREADS TECH CO LTD

Segmentation federal learning method and system based on aggregation gradient broadcast

The invention provides a segmentation federal learning method and system based on aggregation gradient broadcast, and the method comprises the steps: initializing a global model, and segmenting the global model into a server model and a client model according to the capability limitation of a local device; issuing the client model to all local equipment terminals; the multiple local devices execute forward propagation at the same time, and shredded data are obtained through calculation; transmitting the shredded data and the training data label to an edge server; performing forward propagation in parallel according to the shredded data, calculating the loss of each local equipment end in combination with a corresponding training data label, and performing back propagation to obtain a loss function and a gradient of the shredded data; aggregating the gradient of the shredded data and broadcasting to all local equipment; updating the server side model and the client side model, and aggregating the server side model; and the iterative learning loop is repeatedly executed until convergence or the maximum communication round is reached. According to the invention, the communication overhead in segmentation federal learning is saved, and the model training efficiency is effectively improved.
Owner:WUHAN UNIV

Neural network training method based on back propagation algorithm

The invention discloses a neural network training method based on a back propagation algorithm, and relates to the technical field of network training, and the method comprises the steps: obtaining the neuron output data of each neural network layer through a forward propagation process, and constructing the output distribution data reflecting the activation features of each layer of neuron; obtaining error data based on a back propagation process, establishing error term distribution of each neural network layer, and constructing an entropy change estimation factor for measuring network information state change; performing disturbance control on the error data according to the value trend of the entropy change estimation factor, and dynamically adjusting the parameter updating amplitude and direction of the neural network by the processed error data and the output distribution data; and weight iteration of the neural network is guided through the parameter updating behavior. According to the method, error disturbance control and dynamic parameter adjustment are combined, a whole set of mechanism of information perception-disturbance control-dynamic optimization is formed, and dynamic perception and accurate guidance of information state evolution in the neural network training process are achieved.
Owner:NANJING DANIU INFORMATION TECH CO LTD

Confrontation sample generation method and device, storage medium and program product

The invention provides a migration scene-oriented adversarial sample generation method and device, a storage medium and a program product, and the method comprises the steps: receiving a to-be-attacked agent model, selecting a plurality of network layers in the agent model, and generating a random mask matched with the parameter shape of the selected network layer; shielding the parameters of the selected network layer based on the random mask to generate a multi-version proxy model; an original input sample is input into the multi-version agent model for forward propagation, gradient information of each agent model for a current adversarial sample is calculated, and a gradient direction is generated; and iteratively updating disturbance by using the gradient direction, and superposing the updated disturbance to the original input sample to generate an adversarial sample. According to the method, high-quality adversarial samples can be generated, and the attack efficiency is remarkably improved.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Quantitative perception training method and device of neural network model, electronic equipment and storage medium

The invention relates to a quantitative perception training method and device of a neural network model, electronic equipment and a storage medium. The method comprises the steps of obtaining a to-be-trained first neural network model; an operator pair and a non-module operator in the first neural network model are identified, the non-module operator is an operation which is not realized based on a module class, and the operator pair comprises a convolution operator and a batch normalization operator which are connected; the non-module operators are packaged into module operators, the operator pairs are packaged into new convolution operators, a second neural network model is obtained, the new convolution operators run the fusion process of the convolution operators and batch normalization operators during forward propagation, and the parameters of the convolution operators and the batch normalization operators are updated during back propagation; and performing quantitative perception training on the second neural network model. By adopting the method, the reasoning precision of the quantitative model can be guaranteed, and the deployment efficiency of the quantitative model is improved.
Owner:GUANGZHOU XIAOMA HUIXING TECH CO LTD