Semantic Vector Similarity for Untagged Data Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current question-answering systems require significant human resource consumption for tagging data to establish relationships between user queries and knowledge points, making the process inefficient and prone to errors.
Innovation Solution
A method and device for processing untagged data by performing similarity comparisons between semantic vectors of untagged and tagged data, using a trained tagging model to predict tagging feasibility, and dividing data into taggable and non-taggable categories based on similarity and prediction results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging is used to establish relationships between user queries and knowledge points, then the accuracy of establishing relationships is improved, but human resource consumption increases
Solution Approach 1:
The system uses a semantic similarity model to automatically tag user queries without requiring manual human intervention. The model computes semantic vectors for both user queries and knowledge points, calculates similarity scores, and automatically establishes relationships, enabling the system to serve itself rather than relying on manual human tagging operations
Solution Approach 2:
The patent replaces the mechanical manual tagging process with an automated computational system. Instead of humans manually assigning tags, the system uses semantic vector computation, similarity calculation algorithms, and automated decision-making to perform the tagging function, substituting human cognitive work with computational processes
2Measurement precision
If more tagged data is accumulated to train the semantic similarity model, then the model accuracy is improved, but the time required for data preparation increases
Solution Approach 1:
The system automatically generates training data by computing semantic vectors and calculating similarity scores between user queries and knowledge points. The automated process filters and selects appropriate data pairs without requiring manual curation, enabling the system to prepare its own training data independently and efficiently
Solution Approach 2:
The patent performs preliminary computations of semantic vectors and similarity scores during the data processing stage, preparing the data in advance for model training. By pre-computing these representations and filtering data pairs based on similarity thresholds, the system ready's training data before the actual training process begins, saving time in the overall workflow
3Productivity
If automated tagging is used to reduce human resource consumption, then productivity is improved, but the reliability of tagging results may deteriorate
Solution Approach 1:
The system incorporates feedback mechanisms where the semantic similarity model continuously learns from tagged data and refines its performance. The automated tagging process includes evaluation steps that assess the quality of generated tags, and the model adjusts its parameters based on this feedback to improve the reliability of tagging results over time
Solution Approach 2:
The patent uses adjustable parameters in the semantic similarity calculation, such as similarity thresholds and weightings, that can be optimized to balance automation level with result reliability. By tuning these parameters, the system can adapt to different data characteristics and maintain high tagging accuracy while operating in automated mode
Data Source
AI summary
A method for processing untagged data includes: similarity comparison is performed on a semantic vector of untagged data and a semantic vector of each piece of tagged data to obtain similarities corresponding to respective pieces of tagged data; a preset number of similarities are selected according to a preset selection rule; the untagged data is predicted with a tagging model obtained by training through the tagged data, to obtain a prediction result of the untagged data; and the untagged data is divided into untagged data that can be tagged by a device or untagged data that cannot be tagged by the device according to the preset number of similarities and the prediction result.


