Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

93 results about "Batch training" patented technology

Batch Training. Running algorithms which require the full data set for each update can be expensive when the data is large. In order to scale inferences, we can do batch training. This trains the model using only a subsample of data at a time.

Digital twin model incremental learning method and system and equipment state prediction method

The invention provides a digital twin model incremental learning method and system and an equipment state prediction method.The digital twin model incremental learning method introduces an incremental learning mechanism based on elastic weight merging to train an industrial equipment state parameter prediction model; the current working condition samples are stored in an experience playback buffer area, the experience playback buffer area adopts a circular queue structure, the samples stored in the buffer area are used in a mixed mode for training when a new working condition is learned, and mixed loss optimization is executed on a current data batch; after batch training is completed, Fisher information matrix diagonal terms, about historical data, of current parameter configuration are calculated firstly, then global parameter importance is updated until Neural ODE incremental training converges, and finally a complete model state tuple fusing multi-working-condition knowledge is output; the incremental Neural ODE can continue to learn knowledge of new working conditions on the premise that knowledge of old working conditions is not forgotten, and the sustainable evolution ability of the digital twin under the limited resource condition is achieved.
Owner:XI AN JIAOTONG UNIV

Satellite orbit forecasting method based on deep learning physical constraint loss

The invention discloses a satellite orbit forecasting method based on deep learning physical constraint loss, and the method comprises the following steps: 1, carrying out the normalization preprocessing of input data, forming a training data set and a test data set, and constructing batch processing training data; and 2, performing dimension expansion on sample data points in each window in the batch processing data formed in the step 1, constructing a multi-dimensional feature space of the sample points, and forming a batch processing input data format capable of being introduced into the model. And 3, performing forward reasoning on the batch data formed in the step 2 by using a model, and obtaining a batch processing orbit prediction value output by the model at the next moment through a CNN lightweight spatial-temporal feature extraction module and a BiLSTM bidirectional time sequence neural network module. And 4, taking the track prediction value obtained in the step 3 and the truth value label in the training set obtained in the step 1 as input, calculating to obtain a loss value of a current training iteration batch through a multi-random learning loss module fusing physical constraints, and performing reverse updating of model parameters to complete model training. And step five, through the steps two to four, performing reasoning verification on the model by using the test set formed in the step one, and comparing with a truth value in the test set to obtain a model test result.
Owner:CHINA ACADEMY OF SPACE TECHNOLOGY +1

Large language model adaptive fusion method, device and equipment

The invention relates to the technical field of large language model fusion, and discloses a large language model adaptive fusion method, device and equipment. The method comprises the following steps: performing fine tuning on a base model based on a plurality of vertical domain data sets to obtain a plurality of vertical domain models; calculating an incremental vector of a model parameter of each vertical domain model relative to a base model parameter, and recording the incremental vector as a task vector; sampling in a union set of the plurality of vertical domain data sets to obtain a plurality of batch training sets; fixing the base model, and carrying out batch training on gating parameters; in the training process, for each training sample, respectively extracting a semantic feature vector of an input text, and calculating a gating probability matrix; calculating a fusion weight of each task vector; updating the fusion model based on the base model parameters, the task vectors and the corresponding fusion weights; and finally, through model feedforward and back propagation, gate control parameters are updated, and training is repeated. According to the invention, the fusion model suitable for multiple vertical domains can be obtained.
Owner:INSPUR GENERSOFT CO LTD

Reinforcement learning method and system based on state compression and unlabeled reward

The invention discloses a state compression and unlabeled reward-based reinforcement learning method and system and electronic equipment, and the method comprises the steps: receiving a user prompt outputted by a model, and converting the user prompt into a structured state vector which comprises a task target, a function call sequence and an environment feedback result; mapping the function call sequence into a complex plane vector, and automatically generating an unlabeled reward value based on the complex plane vector; based on the unmarked reward value, a low-rank adaptive network is adopted in a model training engine to carry out parameter fine tuning on the model, and batch training is carried out through a token gradient optimization strategy; and identifying and extracting a function call instruction from a text generated by the model, converting the function call instruction into tool API call, and injecting an execution result of the tool API call into a next round of model input in real time. According to the scheme, the decision quality is improved, the resource efficiency is optimized, the implementation cost is reduced, and the application scene can be expanded.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Distributed dyeing machine anomaly detection method based on federal learning

The invention specifically relates to a federated learning-based distributed dyeing machine anomaly detection method, and relates to the technical field of industrial equipment anomaly detection and federated learning, and the method comprises the steps: compressing an obtained weight difference value, encrypting the weight difference value, and transmitting the weight difference value to a federated server; according to the method, a CNN-LSTM mixed structure is adopted for model training, space and time sequence characteristics are considered, and dyeing machine multi-sensor data association and abnormal hysteresis characteristics are adapted; incremental training and small-batch training strategies are adopted to solve the problem of computing power memory limitation of edge equipment, so that the model can be locally and stably trained in the dyeing machine, and network dependence is avoided; a compression transmission link, a dynamic sparse adaptation process stage and hierarchical quantification distinguishing of information levels are performed, so that the data volume is compressed, the transmission efficiency is guaranteed, key information is reserved, the federal learning process is stably operated in an industrial environment, and a dyeing machine production scene is comprehensively adapted from data processing to model training deployment; and a solid support is provided for abnormal detection landing.
Owner:SHAOXING KEQIAO WEAVING PRINTING & DYEING IND BRAIN OPERATION CO LTD

Spatiotemporal consistency oriented training framework for ai based stable video generation

A method includes identifying at least one point in a set of image frames within a temporal window. The set of image frames within the temporal window forms video content. The method also includes extracting temporal information including movement of the at least one point through the set of image frames within the temporal window based on estimation of a local motion vector and / or a supervised optical flow represented in the set of image frames. The method further includes generating a video portion based on association of the temporal information with the set of image frames within the temporal window. In addition, the method includes inputting the video portion as at least part of batch training data for one or more generative machine learning models, where the one or more generative machine learning models that are configured by being trained with the video portion generate temporally stable video content.
Owner:SAMSUNG ELECTRONICS CO LTD

Adaptation method and device for online test of time sequence prediction

The invention relates to the technical field of artificial intelligence, in particular to an online test adaptation method and device for time sequence prediction. Comprising the following steps: establishing and maintaining a historical sample memory bank, storing historical time sequence data, and updating the memory bank through a first-in first-out strategy; screening a historical sample set in a historical sample memory bank, wherein the similarity between the historical sample set and the test sample in the potential space meets a preset condition; performing frequency domain-based mixed data enhancement on the test sample and the historical sample set to generate an enhanced sample set; inputting the enhanced sample set into a time sequence prediction model for batch training, dynamically adjusting model parameters to adapt to distribution offset, and outputting a final prediction result generated through fusion; and evaluating the performance of the model through the loss function. According to the method, the adaptability of the model to new distribution can be enhanced, the robustness of the model is improved, the noise sensitivity is suppressed, data enhancement is performed on the time sequence, and damage to a time domain of the time sequence is avoided.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Building collapse ground penetrating radar inversion and survival gap detection method based on deep learning

A building collapse ground penetrating radar inversion and survival gap detection method based on deep learning is used for reconstructing dielectric constant distribution in a collapsed building, in a complex disaster environment, when the building collapses due to earthquakes, explosions or structural failures, the internal structure of the building can be influenced by various factors, an irregular supporting space is formed, and the building collapses due to the fact that the building collapses due to the earthquakes, the explosions or the structural failures. Holes may become living spaces of trapped persons, a U-Net structure is introduced to serve as an inversion network, and an SE channel attention mechanism is embedded in the framework, so that the attention degree of the model on key region features is improved, and the inversion precision is improved; meanwhile, the BN is replaced by the GN, so that the stability and generalization ability during small-batch training are enhanced; and finally, carrying out logarithmic normalization processing on the label data, optimizing the numerical learning ability of the model, and carrying out reverse normalization in a post-processing stage to ensure the physical interpretability of a prediction result. According to the method, dielectric constant distribution can be more accurately reconstructed in a complex collapse structure environment, and reliable auxiliary support is provided for disaster rescue.
Owner:CENT SOUTH UNIV

Microphone voice recognition system and method based on multi-mode audio-visual fusion

The invention discloses a microphone voice recognition system and method based on multi-mode audio-visual fusion, and belongs to the technical field of artificial intelligence and voice interaction. Firstly, an audio module collects voice signals through a microphone, voice is converted into texts by means of a cloud voice recognition API, and words are further mapped into 300-dimensional semantic vectors through Word2Vec. A visual module extracts lip movement and log-Mel frequency spectrum features, after being subjected to Dlib detection and normalization processing, lip images are sent into a 3D CNN and a dense space-time CNN to extract space-time features, key areas are highlighted with the assistance of a space attention mechanism, and finally sequence visual features are extracted through bidirectional GRU. Meanwhile, a log-Mel spectrogram is generated from the audio signal, and the perception characteristic is enhanced through Mel filtering and logarithm processing. The audio word vector, the lip movement feature and the log-Mel feature are spliced into a multi-modal fusion vector, the multi-modal fusion vector is sent to a CTC decoder, and a text is predicted through Beam Search decoding. An Adam optimizer and a small-batch training strategy are used in the training process, and the model performance and generalization ability are improved.
Owner:ZHONGSHAN FENGXU ELECTRONIC IND CO LTD

A Model Optimization and Update Method, Device, and Medium for Generative Artificial Intelligence

The present application discloses a method, device and medium for optimizing and updating a model of generative artificial intelligence, which is used to reduce the space occupancy of repeated training data when training behavior analysis models for different behaviors to be analyzed. The method of the present application includes: obtaining engineering data to be trained; extracting initial sample data from the engineering data according to the behavior to be analyzed; performing data cleaning on the initial sample data set to obtain a target sample data set; determining model parameters according to the behavior to be analyzed, and obtaining target model training data from the target sample data set; using a batch training method to train the model corresponding to the behavior to be analyzed through the target model training data, and outputting a basic model; combining the basic model with generative artificial intelligence to obtain a target artificial intelligence; monitoring the dynamics of the analysis samples through the target artificial intelligence, and fine-tuning the basic model; monitoring the convergence state after the basic model is fine-tuned; deploying the updated model as the basic model of the target artificial intelligence.
Owner:SHENZHEN EXTREME VISION TECH CO LTD

A power equipment state prediction method and system based on online test-time adaptation

This invention provides a method and system for predicting the state of power equipment based on online testing adaptation, comprising: collecting power equipment state data in real time through sensors and forming test samples; filtering a set of adapted historical samples from a historical sample memory bank that meet preset conditions in terms of similarity to the test samples in the latent space through a transferable historical sample selection module, wherein the historical sample memory bank stores historical power equipment state data; performing time-frequency domain hybrid data augmentation on the test samples and the adapted historical sample set through a transferable online augmentation module to generate an augmented sample set; inputting the augmented sample set into a pre-trained power equipment state prediction model for batch training, dynamically adjusting the model parameters to adapt to the distribution shift; and fusing the output of the dual-stream predictor of the power equipment state prediction model to generate the power equipment state prediction result for the next time period. This invention can perform power equipment state prediction.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Semi-supervised domain adaptive deep forgery detection method

The invention discloses a semi-supervised domain self-adaptive deep forgery detection method, relates to the technical field of forgery detection, and solves the technical problem that a model cannot fully adapt to data distribution of a new domain due to the fact that a small amount of annotated data and a large amount of unannotated data in a target domain are generally difficult to use at the same time in an existing method. The method comprises the following steps: performing multi-batch training on a deep neural network model by adopting samples, wherein each batch of samples comprise a source domain labeled sample, a target domain labeled sample and a target domain unlabeled sample; in the training process, a joint loss function is optimized by adopting stochastic gradient descent, network parameters are iteratively trained, any input image is detected by using a trained Xception model, and the probability that the image is real or forged is output; according to the method, a semi-supervised field adaptive framework is introduced, so that the feature distribution difference between a source domain and a target domain is effectively reduced, and the cross-domain generalization ability is remarkably enhanced.
Owner:RES INST OF YIBIN UNIV OF ELECTRONIC SCI & TECH

A behavior recognition model training method based on time-frequency fusion enhancement

The application provides a behavior recognition model training method based on time-frequency fusion enhancement, comprising: A1, obtaining a training set, wherein original samples are sensing data, and labels indicate the behavior categories of human bodies when the corresponding sensing data is collected; A2, dividing the training set into multiple batches, and iteratively training a feature extractor of a behavior recognition model in batches, wherein each batch training comprises: A21, respectively performing time domain enhancement and frequency domain enhancement on each original sample and then fusing, A22, based on each original sample and the corresponding time-frequency enhanced sample, training the feature extractor to extract sample features according to the input sample through a contrast learning manner, and narrowing the distance between the sample features of the original sample and the corresponding time-frequency enhanced sample and widening the distance between the sample features of the original sample and other samples; A3, obtaining a behavior recognition model comprising the feature extractor trained in step A2 and a classifier, and performing classification training on the behavior recognition model by using the training set.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Model training data construction method and apparatus

PendingCN122391773ABatch trainingData set
The application discloses a model training data construction method and device, which can be applied to various scenes such as cloud technology, artificial intelligence, intelligent transportation and Internet of Vehicles. The method comprises the following steps: obtaining a set of original object images and a set of object wearing images corresponding to a virtual wearing object in a virtual object display platform; extracting object key point information corresponding to each object wearing image in the set of object wearing images and a wearing object category corresponding to each object wearing image; determining an object wearing image with a matching result representing a successful matching result in the set of object wearing images as a preliminary screening wearing image; screening an image meeting a preset condition from the preliminary screening object wearing image to obtain a target wearing image; and constructing a training data set of the virtual wearing object based on the set of original object images and the target wearing image. The application realizes rapid and accurate construction of a batch training data set of the virtual wearing object.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A graph neural network node classification method fusing meta-learning and small batch training

This paper presents a graph neural network node classification method that integrates meta-learning and mini-batch training, belonging to the field of information technology. First, the method utilizes the METIS algorithm to divide the original large-scale graph data into multiple non-overlapping connected subgraphs. Then, by constructing a hybrid selection mechanism based on label coverage and label entropy, subgraphs with high information content and strong representativeness are selected from the subgraph pool as the meta-learning task. Subsequently, iterative training is performed on the selected subgraphs using the meta-learning framework to capture the general prior features of the graph structure, thereby obtaining a set of initial parameters for the model with rapid adaptability. Finally, these optimized initial parameters are transferred to the mini-batch training stage on the full dataset, guiding the model to achieve rapid convergence through high-quality initialization. On large-scale benchmark datasets, this method significantly reduces the number of iterations required by the model while maintaining the same classification accuracy as current mainstream graph neural network models, thus greatly shortening the overall training time.
Owner:HEFEI UNIV

Relationship extraction model training method, relation extraction method, equipment and storage medium

The invention relates to the technical field of data processing, and provides a relation extraction model training method, a relation extraction method, equipment and a storage medium, and the method comprises the steps: obtaining training data corresponding to a current training batch, the training data comprises a current batch of training samples, a current batch of replay samples and a current batch of enhancement samples which mark entity relationship categories under the current training task, and then training the relationship extraction model of the current training stage by using the training data corresponding to the current training batch to obtain a trained target relationship extraction model. According to the method, the replay sample of the historical task and the enhanced sample generated on the basis of the replay sample are added into the original training sample, so that the performance and robustness of the relation extraction model during similar relation data processing are effectively improved.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Model training method and device, electronic equipment and storage medium

The invention provides a model training method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence system performance optimization and deep learning framework scheduling. The method comprises the following steps: in a current batch training process of a target model, acquiring execution information of a corresponding bottom kernel when a processor executes a plurality of computational operators of the target model; according to the execution information, determining a global key path, related to the total calculation duration, of the target model in the current batch training process, and identifying target operators forming training iteration delay in calculation operators contained in the global key path; and generating a delay optimization instruction for the target operator, and in response to the delay optimization instruction, adjusting training configuration parameters corresponding to the target operator in an upper-layer training framework of the processor so as to apply the adjusted training configuration parameters in a subsequent batch training process of the target model. According to the scheme, the training performance of model training can be improved.
Owner:MOORE THREADS TECH CO LTD

Electroencephalogram signal continuous learning classification method and system based on similarity perception playback

The invention discloses an electroencephalogram signal continuous learning classification method and system based on similarity perception playback, and relates to the crossing field of artificial intelligence and neural engineering technologies. The method is characterized in that EEG data stream segments are continuously input into a trained electroencephalogram classification model in an incremental mode, and dynamic adaptation of personalized electroencephalogram signals is achieved; the training process of the classification model comprises the following steps: constructing a time step driven incremental learning framework by adopting a time sequence cross validation strategy, and constructing a training sample set of each time step; a deep learning model is constructed by using deep convolution and separable convolution, the deep learning model is trained based on the training sample set of each time step, after training of each time step is finished, an experience pool is updated based on a similarity perception mechanism, and the experience pool is used for storing historical data samples; when the deep learning model is trained based on the training sample set from the time step 2 to the time step T, carrying out joint training on the training sample set of the current time step and historical samples randomly retrieved from the experience pool; constructing a loss function of batch training, and optimizing trainable parameters in the classification model; according to the continuous learning classification method and system, the dynamic adaptive capacity of the EEG classification model to new data is remarkably enhanced, and high efficiency and accuracy are kept in a continuously changing clinical environment.
Owner:UNIV OF SCI & TECH OF CHINA

System, method, and computer-readable media for leakage correction in graph neural network based recommender systems

Systems, methods, and computer-readable media provide a graph processing system that incorporates a graph neural network (GNN) based recommender system (RS), as well as a method for training a GNN based RS to address feature leakage that leads to overfitting of the trained GNN based RS. A message correction algorithm is used to modify a user node embedding and a positive item node embedding generated by the graph neural network when generating mini batches of training triples used to train the GNN based RS. The GNN message passing operations are performed on one graph only, in contrast to existing approaches which typically run GNN message passing operations on multiple adjusted input graphs constructed for multiple training triples.
Owner:HUAWEI TECH CO LTD

Long-tail radiation source individual identification method and device based on field generalization and storage medium

The invention discloses a long-tail radiation source individual identification method and device based on field generalization and a storage medium. The core of the method is that a plurality of types of source domain data sets distributed in a long-tail mode are obtained; sampling to obtain batch training data; after extracting feature representation by using a feature extractor, inputting the feature representation into a classifier and a domain discriminator at the same time; based on a classification prediction result, combining a field prediction result of a gradient inversion mechanism and class balance loss to construct an integrated overall objective function; all model parameters are jointly optimized by minimizing the target function, and an optimized recognition model is obtained and is finally used for performing high-precision individual recognition on target domain radiation source signals. Through collaborative optimization of classification precision, category balance and domain generalization ability, the problem of insufficient recognition performance of the model on an unknown target domain in a complex scene with training data long tail distribution and domain offset is effectively solved.
Owner:HANGZHOU DIANZI UNIV

A method for determining the occurrence boundary of liquid column separation and risk level

This invention relates to the field of fluid mechanics and transient process analysis of hydraulic systems, and discloses a method for determining the occurrence boundary and risk level of liquid column separation. The method utilizes a visual test bench to conduct multiple experiments under different combinations of test parameters, collecting data and simultaneously recording the liquid column separation state. Based on Buckingham's π theorem, dimensional analysis is used to construct multiple independent dimensionless parameter sets from the test parameter combinations. These are used as input features, with the separation state as the target variable. A machine learning algorithm is introduced for multiple batch training to calculate importance scores, selecting the dominant dimensionless parameter with the highest and most stable score. Subsequently, regression analysis is used to fit the data boundary, solving for the intersection point and outputting the critical equilibrium point as a quantitative benchmark for the occurrence boundary. After confirming separation, high-scoring dimensionless parameters are selected to construct composite judgment parameters, which are compared with preset thresholds to classify the risk level. This invention can accurately define the critical equilibrium point of cavitation occurrence and quantitatively assess its risk severity.
Owner:ZHEJIANG SCI-TECH UNIV

Multi-teacher mixed distillation method and system and computing equipment

The invention relates to a multi-teacher mixed distillation method and system and computing equipment, and the method comprises the following steps: evaluating and screening teacher models, and obtaining pre-training models with cross-domain migration potential to form a teacher set; selecting a lightweight neural network as a student model, and configuring an independent MLP mapping layer for each teacher model in the teacher set to adapt to an output dimension; sampling is carried out in a pre-training data set of each teacher model, and mixed batch training data containing multi-field samples is constructed; and inputting the mixed batch training data into the student model, outputting a prediction result through an MLP mapping layer, and constructing a loss function in combination with teacher model output to complete distillation training. The student model obtained by the invention is no longer limited to a certain specific scientific field or a certain data feature, but becomes a relatively general scientific time sequence signal analysis model, and shows excellent performance in multiple fields.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Real-time diagnosis method for mechanical failure of twin-screw oil transfer pump on-line monitoring

The application provides a kind of twin-screw oil pump on-line monitoring mechanical fault real-time diagnosis method, comprising: obtaining the fault data of twin-screw oil pump to be trained;Data preprocessing is carried out for each parameter attribute related to the fault;Based on the corresponding input attribute of the corresponding node position of the neural network, the fault related data to be processed is displayed;Get the initial network model trained;Align the fault data node parameters in the target optimization parameters;After the weight is optimized using L-BFGS algorithm, supervised training is carried out, and the BP neural network model trained is obtained;The completed BP neural network model is used to train the fault related data in real-time database in batches, and finally the result of accurate prediction is achieved.The twin-screw oil pump on-line monitoring mechanical fault real-time diagnosis method can greatly reduce the production accident rate of twin-screw oil pump, improve the economic efficiency of enterprise production, and protect the life safety of enterprise workers.
Owner:CHINA PETROLEUM & CHEMICAL CORP +1

Virtual training method, system and device, electronic equipment and storage medium

The invention provides a virtual training method, system and device, electronic equipment and a storage medium. According to the method, the acquired training instruction is responded, the virtual scene corresponding to the training instruction is loaded, the operation instruction is acquired through the input device, the input device can be a mouse or a keyboard or other devices, VR equipment or other high-cost devices are not needed in the training process, and batch training can be achieved; obtaining an operation instruction for any second entity virtual object; wherein the operation instruction comprises an operation position relation between the second entity virtual object and the first entity virtual object; and generating a training report according to the position relationship corresponding to each second entity virtual object and the sequence of obtaining the operation instruction of each second entity virtual object. According to the virtual training party disclosed by the invention, not only is the result monitored, but also the assembly process is finely monitored, so that the training efficiency is improved to a certain extent.
Owner:WEICHAI POWER CO LTD

Pedestrian Re-identification Network Model Data Augmentation and Training Method, Training Device

The present invention relates to the field of artificial intelligence technology, and discloses a method for data augmentation and training of a person re-identification network model, including the following steps: S101: Obtain M training images and the annotation data of the M training images. The M training images include pedestrians, and the annotation data of each training image includes the bounding box where the pedestrian in each training image is located and the pedestrian identity identification information; S102: Apply a set sampling strategy to select a batch of training images from the M training images as a batch of training samples, and apply the horizontal strip segmentation-shuffle method to perform data augmentation on the batch of training samples to obtain data-augmented batch training samples. The extended triple loss function designed for the person re-identification network model in the present invention can handle decimal similarity labels, so that it can be jointly applied to the training of the person re-identification network model with the data augmentation Strip-Cutmix method.
Owner:GUANGDONG GAOHANG INTELLECTUAL PROPERTY OPERATION CO LTD

Differential privacy model training method based on buffer mechanism, medium and system

PendingCN122471054ABatch trainingData set
The application discloses a differential privacy model training method based on a buffer mechanism, a medium and a system, wherein the method comprises the following steps: sampling a training data set to obtain small-batch training data; calculating a gradient value and processing the gradient value to generate a private gradient, which is applied to a current model parameter to obtain a preselected weight; sampling a verification data set to obtain small-batch verification data; calculating a loss change value, and performing clipping and noise adding processing on the loss change value to generate a noisy loss change value; judging whether the noisy loss change value is less than a preset rejection threshold; if yes, adding the candidate weight to the buffer; when the number of candidate weights is equal to a preset number threshold, determining the optimal weight according to the relative noisy loss change value, and updating the current model parameter based on the optimal weight; the method can effectively protect privacy, improve the updating quality of the model in the training process, and improve the accuracy of the final prediction result.
Owner:XIAMEN UNIV OF TECH

A general domain adaptation image classification model implementation method and system

The application discloses a general domain self-adaptive image classification model implementation method and system, wherein the method comprises the following steps: setting a source domain image dataset with classification labels and a target domain image dataset different from the source domain data distribution and without classification labels as image data for classification model training; training closed set and open set classifiers in the source domain classification model, outputting classification probabilities belonging to each category of the source domain through the closed set classifier, outputting in-class and out-of-class classification probabilities and maximum open set entropy through the open set classifier; determining the most confused categories of the target domain from the maximum open set entropy and setting a transition zone, separating simple samples outside the transition zone with a self-confidence higher than a preset first threshold value and difficult samples inside the transition zone with a self-confidence lower than a preset second threshold value; and performing batch training on the classification model by combining a cross-entropy loss, a selection optimization strategy loss and a neighbor aggregation strategy loss function, so as to obtain a final general domain self-adaptive image classification model.
Owner:GUANGZHOU UNIVERSITY

Lightweight model-based forklift tray tracking method, system and equipment and medium

The invention relates to the technical field of artificial intelligence, and particularly provides a forklift tray tracking method, system and device based on a lightweight model and a medium, and the method comprises the steps: constructing a rotating target data set of a standard tray, and carrying out the targeted data enhancement and preprocessing; secondly, performing three-point improvement on a YOLOv12 algorithm: introducing a brand new StarNet backbone network, constructing a high-dimensional implicit feature space through star operation, and improving feature expression capability while reducing parameter quantity; a dynamic hybrid convolution module is designed in the neck network, multi-scale features are adaptively extracted by using multi-branch deep convolution and a dynamic weight fusion mechanism, and the flexibility of the model is enhanced; and a lightweight rotation detection head is provided, and the rotation angle of the tray is efficiently predicted while the parameter quantity is reduced and the small-batch training stability is improved through application group normalization and a shared convolution structure. And finally, combining the improved detection model with a ByteTrack tracking algorithm of the optimized adaptive rotating frame to form a complete identification tracking system.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Pipeline parallel distributed training method, device and system for deep neural network

The application provides a pipeline parallel distributed training method, device and system for a deep neural network. The method comprises: sequentially transmitting each group of small batch training data constituting a current batch of training data to an edge server via an optical network, so that the edge server and the cloud server cooperatively perform asynchronous parallel cooperative training on different sub-task models constituting the deep neural network in a communication and training decoupling manner, and the edge server sequentially outputs the gradients corresponding to each group of small batch training data; and sequentially receiving the gradients of each group of small batch training data. The application can ensure correct transmission of data during model training, reduce communication overhead during training, achieve load balancing between devices, improve model training efficiency, effectiveness and resource utilization of devices participating in training.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Multi-granularity perception and prototype-driven adaptive recognition method for ultra-high-definition image domain

This invention belongs to the field of domain adaptation technology for ultra-high-definition (UHD) images. It proposes a multi-granularity sensing and prototype-driven UHD image domain adaptive recognition method, comprising the following steps: First, acquiring a cross-image domain conventional image dataset and an UHD image dataset, and preprocessing both datasets to obtain a preprocessed batch training image dataset; second, constructing and initializing a multi-granularity feature extractor and a feature classifier; then, training an unsupervised domain adaptive recognition model for UHD images based on the batch training image dataset and combining the multi-granularity feature extractor and the feature classifier; finally, inputting the UHD image to be recognized into the unsupervised domain adaptive recognition model for classification to obtain the recognition result. This invention solves the problems of neglecting fine-grained features in UHD images and the difficulties, limited quantity, and high computational resource consumption of UHD image domain annotation.
Owner:SICHUAN NATIONAL INNOVATION VISION UHD VIDEO TECHNOLOGY CO LTD