Bi-Modal Generative Model for Neural Architecture Design

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning techniques, particularly in neural architecture search (NAS), are limited to uni-modal learning, are dataset-dependent, computationally complex, and lack understanding of natural language and generative capabilities.

Innovation Solution

A bi-modal generative model is developed for joint learning of natural language (NL) and artificial neural network architectures (NA), enabling tasks such as NA answer generation, text-based NA generation, and multi-modal NA translation assisted by NL information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If neural architecture search (NAS) is used to automate neural network design, then the ability to identify suitable architectures for classification tasks is improved, but computational complexity increases significantly and the system becomes dataset-dependent

Engineering Contradiction:
Improveability to identify suitable neural network architectureVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses a pre-trained generative model that has learned neural network architecture patterns from diverse datasets. Instead of performing computationally expensive NAS for each new dataset, the system copies/adapts the pre-trained model's knowledge to generate architectures for new datasets, significantly reducing computational complexity while maintaining adaptability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary training of the generative model on a large corpus of neural network architectures and their corresponding performance data before deployment. This preliminary action enables the model to make rapid architecture recommendations for new datasets without requiring expensive real-time NAS computations.

Inventive Principle:
Principle #10Preliminary action

2Extent of automation

If existing NAS approaches are used, then neural network architecture selection for classification tasks is automated, but the system becomes limited to specific datasets and cannot generalize to other modalities or tasks

Engineering Contradiction:
Improveautomation of neural network architecture selectionVSAvoidgeneralizability to different datasets and modalities
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The patent develops a universal generative model that can handle multiple inference tasks (classification, generation, translation) and multiple data modalities (images, text, audio). The model uses a unified architecture that automatically adapts to different datasets and task types, eliminating the need for separate NAS systems for each specific application.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a dynamic system where the generative model can adjust its behavior based on the input dataset characteristics and task requirements. The model dynamically selects appropriate architecture patterns and modifies them to suit the specific needs of each dataset, enabling both automation and generalizability.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If multi-modal language models are used to jointly learn from multiple modalities, then understanding of various senses in information processing is improved, but the model requires extensive training data and computational resources

Engineering Contradiction:
Improveunderstanding of various senses in information processingVSAvoidcomputational resources for training
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary pre-training of the multi-modal generative model on large corpora of neural network architectures, natural language descriptions, and task specifications. This preliminary action enables the model to acquire general knowledge about neural network design patterns across multiple modalities, reducing the computational resources needed for subsequent fine-tuning on specific tasks.

Inventive Principle:
Principle #10Preliminary action

4Ease of manufacture

If uni-modal learning is used for generative tasks, then training is simplified to use a single data type, but the model cannot produce plausible generative outputs that capture cross-modal relationships

Engineering Contradiction:
Improvesimplicity of training processVSAvoidplausibility of generative outputs
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent merges multiple modalities (neural network architecture data, natural language descriptions, task specifications) into a unified multi-modal generative model. The model learns cross-modal relationships between these different data types, enabling it to generate plausible outputs that satisfy both architectural constraints and task requirements while maintaining training simplicity through a unified framework.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12321700B2Methods, systems, and media for bi-modal generation of natural languages and neural architectures
Publication Date: 2025.06.03 HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
  • US12321700B2 patent drawing
  • US12321700B2 patent drawing
  • US12321700B2 patent drawing

AI summary

Methods, systems, and computer-readable media for bi-modal generation of natural language (NL) and artificial neural network architectures (NA), with reference to an example implementation framework entitled “ArchGenBERT”. A model and method of training the model for bi-modal generation of NL and NA are described. The model trained for bi-modal generation of NL and NA can be deployed to perform a number of useful tasks to assist with designing, describing, translating, and modifying neural network architectures.