Systems and methods for automatic data analysis, organization, and labelling
An AI pipeline automatically organizes and labels data using machine learning, addressing inefficiencies in user-dependent data management by identifying and grouping relevant keywords, enhancing data utilization and reducing costs.
WO2026106978A1PCT designated stage Publication Date: 2026-05-21SK HYNIX NAND PRODUCT SOLUTIONS CORP
View PDF 5 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SK HYNIX NAND PRODUCT SOLUTIONS CORP
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-21
AI Technical Summary
Technical Problem
Existing data management systems require user intervention for data ranking and reduction, leading to inefficiencies and high costs, and lack a universal definition of valuable data.
Method used
An AI processing pipeline that automatically organizes and labels data using machine learning techniques, including embedding extraction, clustering, and image-to-text techniques, to identify and group relevant keywords without user intervention.
Benefits of technology
Enables efficient data management with minimal user input, allowing for data to be used in various applications such as training models and event detection, while reducing storage and computational costs.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure US2025055013_21052026_PF_FP_ABST
Abstract
Some embodiments are directed to systems and methods for preparing data. IN one aspect, a computer system obtains input images and groups them into image clusters including a first image cluster that includes a first set of input images. The computer system extracts image keywords from each of the first set of input images and groups the image keywords to identify a plurality of cluster keywords of the first image cluster. The computer device determines a plurality of keyword weights, each associated with a respective one cluster keyword of the plurality of cluster keywords based on cluster locations of the first set of input images in the first image cluster. The computer system labels the first set of input images based on the plurality of cluster keywords and the plurality of keyword weights.
Need to check novelty before this filing date? Find Prior Art