Language-Model Data Acquisition for Accurate Source Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The current data acquisition process is costly due to the need for manual setting and frequent adjustment of acquisition conditions, which often have low association with the desired data.

Innovation Solution

A data acquisition method utilizing a pre-trained language model to determine task information from user input, search for associated data sources, and provide metadata, reducing the need for manual condition setting and adjustment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual setting and adjustment of acquisition conditions is used, then data acquisition can be performed, but the process is costly and time-consuming with low association accuracy

Engineering Contradiction:
Improvedata source selection accuracyVSAvoidmanual condition setting time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of setting and adjusting acquisition conditions with an automated language model-based system. The language model processes user queries and automatically determines acquisition conditions, eliminating the need for manual configuration while improving accuracy in data source selection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing users to simply input their data needs as natural language queries, and the language model automatically handles the complex task of condition determination and data source selection without requiring user expertise or manual intervention in the acquisition process.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If manual adjustment of acquisition conditions is required, then data acquisition can be adapted to different needs, but the complexity and cost increase

Engineering Contradiction:
Improvedata acquisition adaptabilityVSAvoidcondition management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent changes the parameters of condition determination from manual configuration to automated language model processing. The language model dynamically adjusts acquisition conditions based on user queries, providing adaptability to different data needs while simplifying the overall system complexity by centralizing condition management in the language model.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If automated language model-based search is used, then manual condition setting is eliminated, but the system complexity increases

Engineering Contradiction:
Improvedata acquisition automation levelVSAvoidlanguage model integration complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent introduces a language model as an intermediary component that bridges user queries and data acquisition operations. This intermediary handles the complexity of condition determination and data source matching, allowing high-level automation while isolating the complexity management to a dedicated component rather than distributing it throughout the entire system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12602374B2Data acquisition method and apparatus, computer device and storage medium
Publication Date: 2026.04.14 BEIJING VOLCANO ENGINE TECH CO LTD
  • US12602374B2 patent drawing
  • US12602374B2 patent drawing
  • US12602374B2 patent drawing

AI summary

The present disclosure relates to the field of data processing technology, and discloses a data acquisition method and apparatus, a computer device and a storage medium. The method includes: acquiring question information; determining task information corresponding to the question information using a pre-trained language model, where the task information comprises a keyword in the question information and a target data source type; obtaining a search result by searching, using the pre-trained language model and according to a prompt corresponding to the task information, an associated data source in at least one candidate data source; and determining at least one data source to be provided from the at least one candidate data source according to the search result, and providing metadata of the at least one data source to be provided to the user.