ML Data Decoder for High-Density Storage Read Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data storage and retrieval methods face interference issues due to the transduction processes used in high-throughput, high-density data storage, which can lead to errors and inefficiencies in decoding data from storage media.

Innovation Solution

The implementation of a machine-learning based approach, specifically using a convolutional neural network (CNN), to decode data from optical storage media by transforming image data from analyzer-camera images into probability arrays that represent the likelihood of actual data values, bypassing the need for explicit computation of intermediate metrics like birefringence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional transduction processes are used for high-throughput data storage, then storage capacity and speed are improved, but interference and errors increase

Engineering Contradiction:
Improvedata storage throughputVSAvoiddata retrieval accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces traditional mechanical/optical transduction processes with a machine-learning-based decoding system. Instead of using conventional physical transduction methods to read and interpret stored data, the system uses trained neural networks to process measurement representations and predict actual data values, thereby maintaining high throughput while reducing interference-related errors through intelligent pattern recognition

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the operational parameters of the data retrieval process by transitioning from direct transduction to probabilistic prediction. The machine learning model outputs probability distributions for each data location, allowing the system to select the most likely data values and thereby improve reliability while maintaining the high throughput enabled by the underlying storage medium

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If explicit computation of intermediate metrics is performed during data decoding, then measurement precision is improved, but processing time and complexity increase

Engineering Contradiction:
Improvedata decoding accuracyVSAvoiddata retrieval time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by training the machine learning model offline before actual data retrieval operations. During the training phase, the system learns optimal decoding strategies and patterns from大量 training data, storing this knowledge in the model's parameters. During actual retrieval, the pre-trained model can quickly predict data values without performing complex intermediate computations, thereby maintaining high precision while reducing processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating a trained machine learning model that captures the complex relationships between measurements and data values. Instead of repeatedly performing explicit computations during each retrieval operation, the system uses the copied knowledge embedded in the model's parameters to rapidly predict data values, achieving both speed and accuracy

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12019705B2Machine-learning optimization of data reading and writing
Publication Date: 2024.06.25 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12019705B2 patent drawing
  • US12019705B2 patent drawing
  • US12019705B2 patent drawing

AI summary

Examples are disclosed that relate to encoding data on a data-storage medium. The method comprises obtaining a representation of a measurement performed on the data-storage medium, the representation being based on a previously recorded pattern of data encoded in the data-storage medium in a layout that defines a plurality of data locations. The method further comprises inputting the representation into a data decoder comprising a trained machine-learning function, and obtaining from the data decoder, for each data location of the layout, a plurality of probability values, wherein each probability value is associated with a corresponding data value and represents the probability that the corresponding data value matches the actual data value in the previously recorded pattern of data at a same location in the layout.