数据收集方法、装置、设备及介质

By expanding the seed word list through few-shot learning algorithms and thought chain algorithms, a high-quality dataset is generated, which solves the problem of insufficient data for pre-trained models and achieves diversity and quantity assurance in data collection.

CN119719465BActive Publication Date: 2026-07-17INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2024-12-30
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, the data collection for pre-trained models mainly relies on data accumulated by the manufacturers themselves or low-quality publicly available data crawled from web crawlers, resulting in a small amount of data with low quality, which is difficult to meet the training needs of the model.

Method used

We use a few-shot learning algorithm to determine data collection needs, expand the seed word list through a thought chain algorithm, generate prompts based on part-of-speech tags, and combine network retrieval and target pre-trained models to obtain a high-quality target dataset.

Benefits of technology

It improves the diversity and coverage of data, ensures data quality and quantity, forms a coherent data collection process, and enables autonomous learning and adaptability to complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719465B_ABST
    Figure CN119719465B_ABST
Patent Text Reader

Abstract

本发明公开了一种数据收集方法、装置、设备及介质,涉及计算机技术领域,应用于智能体,包括:利用少样本学习算法确定数据收集需求对应的目标主题领域和所需数据量;从种子词语管理器中确定与目标主题领域对应的初始种子词语列表,利用目标预训练模型对初始种子词语列表进行扩充;基于最少到最多提示算法鉴别扩充后种子词语列表中各目标种子词语的词性,利用与目标种子词语的词性对应的提示信息生成模板生成与目标种子词语对应的提示信息;基于提示信息并利用网络检索方式获取与所需数据量对应的目标数据集,其中,采用思维链条算法将当前步骤的输出作为下一步骤的输入。思维链条算法形成一个连贯的思维过程,并使得收集的数据可以保质保量。
Need to check novelty before this filing date? Find Prior Art