Diffusion Language Model for Accurate Corpus Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autoregressive language models suffer from error accumulation and lack of global vision during iterative training, leading to poor quality and low accuracy in corpus processing results.

Innovation Solution

A target diffusion language model is trained using a diffusion model approach, where a mask language model is pre-trained and then fine-tuned using corpus samples with varying mask rates to improve training efficiency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If autoregressive language model is used for corpus processing, then processing capability is provided, but accuracy and quality of results deteriorate due to error accumulation and lack of global vision

Engineering Contradiction:
Improveprocessing capabilityVSAvoidaccuracy of corpus processing results
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent changes the fundamental parameter of the language model architecture from autoregressive to diffusion-based. This parameter change enables the model to process corpus data with different mask rates, allowing it to capture global context and reduce error accumulation, thereby improving accuracy while maintaining processing capability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies preliminary action through pre-training the diffusion language model on corpus samples with various mask rates before actual processing. This pre-training phase equips the model with global vision and error correction capabilities, enabling it to handle complex corpus processing tasks with higher accuracy

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If diffusion model approach is used with multiple mask rates, then accuracy and quality improve, but training complexity increases

Engineering Contradiction:
Improveaccuracy of corpus processing resultsVSAvoidtraining complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the training process universal by using corpus samples with multiple mask rates that can serve multiple purposes: teaching the model to handle different levels of data corruption, improving robustness, and enabling flexible processing of various corpus types. This multi-functionality justifies the increased training complexity by delivering broad performance improvements

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent systematically varies the mask rate parameter during training to create diverse training scenarios. By changing this single parameter across multiple training iterations, the model learns to handle various corpus processing challenges, improving accuracy without requiring fundamentally different training approaches for each task type

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250053808A1Method, apparatus, electronic device, and storage medium for data processing
Publication Date: 2025.02.13 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250053808A1 patent drawing
  • US20250053808A1 patent drawing
  • US20250053808A1 patent drawing

AI summary

The embodiment of the disclosure provides a method, apparatus, electronic device, and storage medium for data processing. The method includes: receiving a corpus to be processed; obtaining a target prediction result corresponding to the corpus to be processed by processing the corpus to be processed based on a target diffusion language model, wherein the target diffusion language model is obtained by training based on a plurality of corpus samples, and a mask corpora in the corpus samples corresponds to different mask rates; and displaying the target prediction result. According to the technical solution of the embodiment of the disclosure, the effect of making the target diffusion language model can process the corpus data based on the principle of diffusion model and the obtained corpus data processing results can meet the requirements of corpus processing tasks is implemented.