Generative Dialog Model Safety Optimization via Iterative Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generative dialog systems face challenges in ensuring output safety, as they may generate inappropriate or harmful replies due to noise, bias, or errors in language materials, posing risks to user feelings, trust, and legal or moral responsibility, especially when handling diverse task forms and scenarios.

Innovation Solution

A progressively-iterative generative dialog model optimization method is employed, where the safety specification and the generative dialog model are alternately optimized through multiple iterations, using a detection model to ensure that generated replies conform to a target safety specification, aligning with human safety values across various content fields and application scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the generative dialog model is trained with diverse language materials to improve its dialog capability, then the model's versatility and responsiveness are improved, but the output safety deteriorates due to noise, bias, or errors in the training data

Engineering Contradiction:
Improvedialog capabilityVSAvoidoutput safety
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary action by conducting pre-training on diverse language materials to build the model's dialog capability, followed by a dedicated safety alignment phase. The safety specification is established before final training, and the model is specifically trained to conform to this specification, preventing harmful outputs before they occur rather than correcting them after generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback through the safety specification that provides continuous guidance during training. The model generates replies and these replies are evaluated against the safety specification, with the training process adjusting the model based on this feedback to ensure conformity. This creates a closed-loop system where safety is continuously monitored and reinforced.

Inventive Principle:
Principle #23Feedback

2Object-affected harmful factors

If the safety specification is updated to address new safety risks, then the output safety is improved, but the model optimization complexity increases due to additional training iterations

Engineering Contradiction:
Improveoutput safetyVSAvoidoptimization process
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent applies dynamics by making the safety specification adjustable and updatable. When new safety risks are identified, the specification can be dynamically modified without requiring complete retraining from scratch. The system supports iterative optimization where the safety specification evolves alongside the model, allowing flexible adaptation to new safety requirements.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent utilizes parameter changes by modifying the safety specification parameters (such as safety thresholds, evaluation criteria, and constraint parameters) to address new risks. These parameter adjustments guide the training process and allow the model to adapt to updated safety requirements through controlled optimization iterations, balancing safety improvements with training complexity.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If multiple optimization iterations are performed to ensure safety compliance, then the output safety is improved, but the training time and computational resources increase

Engineering Contradiction:
Improveoutput safetyVSAvoidtraining time
Core Design Contradiction:
Object-affected harmful factorsVSLoss of time

Solution Approach 1:

The patent applies continuity of useful action by implementing multiple optimization iterations where the model is continuously trained to conform to the safety specification. Each iteration builds upon the previous one, with the model progressively improving its safety alignment. This continuous process ensures that safety is not a one-time check but an ongoing training objective that permeates the entire optimization process.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS20240338530A1Generative dialog model training method and apparatus as well as generative dialog implementing method and apparatus
Publication Date: 2024.10.10 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20240338530A1 patent drawing
  • US20240338530A1 patent drawing
  • US20240338530A1 patent drawing

AI summary

A generative dialog model training method in the fields of artificial intelligence, such as deep learning, natural language processing, intelligent dialogs, is disclosed. The generative dialog model training method may include: in response to determination of an update of a safety specification, taking an updated safety specification as a target safety specification, and determining a dialog input corresponding to a current optimization according to the target safety specification, the update being performed on a previous safety specification when a generative dialog model after last optimization is determined not to meet a launch requirement; and optimizing the generative dialog model according to the dialog input and a principle that a reply generated by the generative dialog model conforms to the target safety specification, the generative dialog model being configured to generate the reply corresponding to the dialog input.