数据收集方法、装置、设备及介质
By expanding the seed word list through few-shot learning algorithms and thought chain algorithms, a high-quality dataset is generated, which solves the problem of insufficient data for pre-trained models and achieves diversity and quantity assurance in data collection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2024-12-30
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, the data collection for pre-trained models mainly relies on data accumulated by the manufacturers themselves or low-quality publicly available data crawled from web crawlers, resulting in a small amount of data with low quality, which is difficult to meet the training needs of the model.
We use a few-shot learning algorithm to determine data collection needs, expand the seed word list through a thought chain algorithm, generate prompts based on part-of-speech tags, and combine network retrieval and target pre-trained models to obtain a high-quality target dataset.
It improves the diversity and coverage of data, ensures data quality and quantity, forms a coherent data collection process, and enables autonomous learning and adaptability to complex environments.
Smart Images

Figure CN119719465B_ABST