Joint Neural Network Training via Weighted Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for training neural networks lack an efficient mechanism for joint training of multiple neural networks, which hinders their performance in multimedia transport and inference tasks, such as image classification and video processing, where coordinated learning and optimal network selection are crucial.
Innovation Solution
An apparatus and method for jointly training multiple neural networks by determining weighted losses based on their performance on training samples, allowing for iterative training and overfitting, and selecting optimal networks for inference using score computations and auxiliary networks, with weight updates and compression for efficient communication between encoder and decoder devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple neural networks are trained separately using conventional methods, then each network can be trained independently, but the coordination and performance in multimedia transport tasks deteriorates
Solution Approach 1:
The patent merges multiple separate neural network training processes into a single joint training framework. Multiple neural networks are trained simultaneously using shared training data and coordinated weight updates, allowing them to learn complementary patterns for multimedia transport tasks while maintaining a unified training objective that improves overall system reliability.
Solution Approach 2:
The joint training apparatus serves multiple functions: it trains multiple neural networks simultaneously, coordinates their learning through shared loss functions, enables weight sharing and synchronization, and supports both encoder and decoder network training. This multi-functional approach improves performance without proportionally increasing complexity.
2Manufacturing precision
If neural networks are trained with coordinated learning and weight sharing, then performance in multimedia processing improves, but the training process complexity increases
Solution Approach 1:
The training process is segmented into distinct phases: initial independent training of individual neural networks, followed by coordinated joint training with weight sharing, and finally inference-time selection. This segmentation allows the system to achieve high accuracy through coordinated learning while managing complexity by breaking the training process into manageable stages.
Solution Approach 2:
The system dynamically changes training parameters including loss function weights, learning rates, and weight sharing ratios during the training process. By adjusting these parameters adaptively, the system achieves high processing accuracy while reducing the effective complexity of coordinated training through automated parameter optimization.
3Reliability
If multiple neural networks are trained and stored, then optimal network selection for inference is possible, but communication and storage requirements increase
Solution Approach 1:
The system extracts and transmits only the essential weight parameters of neural networks during joint training, rather than transmitting complete network models. This extraction approach enables optimal network selection for inference while significantly reducing communication and storage requirements by transmitting only the necessary learning parameters.
Solution Approach 2:
The system trains multiple neural networks with potentially excessive computational resources but only stores and communicates the essential weight parameters. This partial action approach allows optimal network selection capability while minimizing storage and communication data volume by retaining only the necessary learned parameters for inference.
Data Source
AI summary
Various embodiments provide an apparatus, a method, and a computer program product. The apparatus includes at least one processor; and at least one non-transitory memory comprising computer program code; wherein the at least one memory and the computer program code are configured to, with the at least one processor, cause the apparatus at least to perform: determine a plurality of weights used to calculate a weighted loss based at least on a performance of a plurality of neural networks on one or more training samples; and jointly train the plurality of neural networks, wherein at each training iteration the plurality of neural networks are trained based at least on the weighted loss.


