Artificial intelligence data annotation and cue word automatic construction engine system

By using an AI-powered data annotation and prompt word automatic construction engine system, and leveraging pre-trained models and reinforcement learning techniques, the system solves the problems of low efficiency, difficulty in ensuring consistency, and reliance on human experience in traditional manual annotation, achieving efficient, low-cost, and highly adaptable data annotation and prompt word generation.

CN120950973APending Publication Date: 2025-11-14UFO TECH (BEIJING) CO LTD

Patent Information

Application Number
CN202511077523.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional AI data annotation is inefficient, costly, and inconsistent. Prompt word construction relies on human experience, lacks adaptability, and is cumbersome.

Method used

An AI-powered data annotation and prompt word automatic construction engine system is adopted, including a data annotation module, a prompt word automatic construction module, a system management module, and a model evaluation module. It utilizes pre-trained models, active learning algorithms, reinforcement learning techniques, and evaluation metrics to optimize the annotation and prompt word generation process.

Benefits of technology

Significantly improves annotation efficiency, reduces costs, ensures annotation consistency, generates scientific and reasonable prompts, adapts to model changes, reduces manual construction time and effort, and enhances system usability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950973A_ABST
    Figure CN120950973A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and particularly relates to an artificial intelligence data annotation and cue word automatic construction engine system, which comprises a data annotation module for firstly carrying out preliminary annotation on data based on a pre-training model, then automatically annotating a data sample through an active learning algorithm, and meanwhile, monitoring the quality of annotated data in real time; and the cue word automatic construction module generates cue words based on task analysis, optimizes the cue words by using a reinforcement learning technology, and performs classified storage and management on the generated cue words. By adopting the semi-automatic labeling and active learning labeling functions, not only can the workload be greatly reduced, but also the unnecessary labeling work can be reduced, so that the labeling efficiency is remarkably improved, the labeling time is shortened, the large-scale data labeling requirement is met, and the problem of efficiency bottleneck caused by slow manual labeling is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to an artificial intelligence data annotation and prompt word automatic construction engine system. Background Technology

[0002] With the rapid development of artificial intelligence (AI) technology, its applications are becoming increasingly widespread and profound across various fields. From image recognition assisting in security monitoring to natural language processing empowering intelligent customer service, AI is constantly changing people's lives and work. Behind this flourishing development, high-quality data annotation and accurate and effective prompt word construction are the key foundations supporting the powerful performance of AI models.

[0003] In terms of data annotation, traditional annotation methods have the following problems:

[0004] 1. Low annotation efficiency: Manual annotation is slow and cannot meet the ever-increasing demand for data. In autonomous driving scenarios, a large amount of driving video needs to be annotated, including information such as the location and category of various target objects such as vehicles, pedestrians, and traffic lights. If manual annotation is used to complete the annotation of massive amounts of video data frame by frame, it will be time-consuming and seriously affect the research and development progress of autonomous driving technology.

[0005] 2. High annotation costs: Building a professional annotation team requires a significant investment in personnel recruitment, training, and the purchase of annotation tools. For example, a medium-sized image annotation project would incur substantial costs if it required hiring dozens of professional annotators, plus training expenses and the cost of using professional annotation software.

[0006] 3. Difficulty in ensuring annotation consistency: Different annotators may have different understandings and implementations of annotation rules. In medical image annotation, for tumor annotation, different annotators may have different experiences and subjective judgments, leading to inconsistent tumor boundaries. This will affect the accuracy of medical image diagnostic models trained based on such annotated data.

[0007] In terms of prompt word construction, traditional construction methods have the following problems:

[0008] 1. Reliance on human experience: When constructing prompts, humans often design them based on their own understanding of the task and the model, lacking a scientific and systematic approach. In intelligent writing assistance tools, in order for the model to generate a logically clear and content-rich paper abstract, humans may conceive prompts based on their past writing experience. However, due to a lack of comprehensive and in-depth analysis of the model's characteristics and task requirements, the generated prompts may not be able to guide the model to generate the ideal abstract content.

[0009] 2. Lack of adaptability: When the task scenario or model changes, the original prompts may no longer be applicable and need to be rebuilt. For example, in the e-commerce field, when the types of goods change or user needs and preferences change, the prompts originally used in the product recommendation system may not be able to accurately guide the model to recommend suitable products to users, requiring manual time and effort to readjust the prompts.

[0010] 3. The construction process is cumbersome: The process of manually trying different combinations of prompts and evaluating their effects is tedious and time-consuming. When performing image generation tasks, in order for the model to generate an image with a specific style and elements, humans need to continuously try different combinations of descriptive prompts, such as combinations of words describing the main subject, colors, style, and other elements of the image, and then observe the effect of the image generated by the model. This process may need to be repeated many times, consuming a lot of time and effort.

[0011] Based on the above, an artificial intelligence data annotation and prompt word automatic construction engine system is invented. Summary of the Invention

[0012] To address the aforementioned technical problems, according to one aspect of the present invention, the present invention provides the following technical solution:

[0013] An AI-powered data annotation and prompt word automatic construction engine system, which includes:

[0014] The data annotation module first performs preliminary annotation on the data based on the pre-trained model, and then automatically annotates the data samples through an active learning algorithm. At the same time, it monitors the quality of the annotated data in real time.

[0015] The automatic prompt word construction module first generates prompt words based on task analysis, then optimizes the prompt words using reinforcement learning techniques, and at the same time, classifies, stores and manages the generated prompt words.

[0016] The system management module is used to first manage user permissions, then manage data annotation tasks and prompt word construction tasks, and at the same time, it can adjust the relevant parameters of data annotation algorithms and prompt word generation algorithms;

[0017] The model evaluation module is used to comprehensively evaluate the performance of the model after data labeling and training using evaluation metrics. At the same time, it analyzes the degree of fit between the generated content and the expected target for the results generated by the prompt word automatic construction module.

[0018] As a preferred embodiment of the AI ​​data annotation and prompt word automatic construction engine system described in this invention, the data annotation module includes:

[0019] The data preprocessing module is used to clean, transform, and extract data features from the raw data.

[0020] Semi-automatic annotation is used to perform preliminary annotation of data using a pre-trained model, which is then reviewed and corrected by humans, thereby improving annotation efficiency.

[0021] The active learning annotation module is used to automatically select the most valuable data samples for annotation through active learning algorithms, so as to improve the quality and representativeness of the labeled data. At the same time, based on the uncertainty of the model and the diversity of the samples, it will select the most needed samples from the massive unlabeled data, allowing the annotators to annotate them, so that the labeled data can better cover various situations and improve the model training effect.

[0022] The annotation quality monitoring module is used to monitor the quality of annotation data in real time, and to promptly identify and correct problems in the annotation process through consistency checks and error rate statistics.

[0023] As a preferred embodiment of the AI ​​data annotation and prompt word automatic construction engine system described in this invention, the specific steps of the data preprocessing module are as follows:

[0024] S1, Data Collection and Integration: First, raw data is obtained from multiple data sources, and then data of different formats and structures are integrated together to form a unified data set;

[0025] S2, data cleaning, firstly identifies and deletes duplicate data records by comparing key features or all attributes of the data; then it processes missing values ​​in the data; finally, it identifies erroneous data based on the business rules and statistical patterns of the data, and corrects it according to the actual situation.

[0026] S3, Data Transformation: First, convert data of different formats into a standard format that the system can process; then, according to the needs of data processing, convert the data from one type to another; then, for categorical data, use an appropriate encoding method to convert each category into a binary vector to meet the input requirements of machine learning algorithms.

[0027] S4, Data Feature Extraction: First, extract more valuable features from the original data; then, select the most relevant and representative features from the extracted features to remove redundant and irrelevant features; finally, verify the data to check its integrity, accuracy, and consistency.

[0028] As a preferred embodiment of the AI ​​data annotation and prompt word automatic construction engine system described in this invention, the prompt word automatic construction module includes:

[0029] The prompt word generation module is used to perform in-depth analysis of the task description input by the user, and generate initial prompt words by combining relevant domain knowledge and model characteristics;

[0030] The prompt word optimization module is used to continuously optimize prompt words based on the model's performance in real-world tasks using reinforcement learning techniques, thereby improving the effectiveness of the prompt words;

[0031] The prompt word library management module is used to classify, store, and manage the generated prompt words, making it convenient for users to query and reuse them.

[0032] As a preferred embodiment of the AI ​​data annotation and prompt word automatic construction engine system described in this invention, the system management module includes:

[0033] The user management module is used to manage user permissions and ensure the security and standardization of system operations.

[0034] The task management module is used to manage the creation, allocation, and progress tracking of data annotation tasks and prompt word construction tasks;

[0035] The settings module is used to adjust the relevant parameters of the data annotation algorithm and the prompt word generation algorithm according to the actual use scenario and data characteristics.

[0036] As a preferred embodiment of the AI ​​data annotation and prompt word automatic construction engine system described in this invention, the model evaluation module includes data annotation model evaluation and prompt word guidance model evaluation.

[0037] As a preferred embodiment of the AI ​​data annotation and prompt word automatic construction engine system described in this invention, the specific steps for evaluating the data annotation model are as follows:

[0038] S1, Determine the evaluation indicators: Select appropriate evaluation indicators based on the type and objectives of the data labeling task;

[0039] S2, Divide the evaluation dataset: Divide a portion of the labeled dataset as the evaluation dataset, including hold-out and cross-validation methods;

[0040] S3, Model Prediction: Use the trained data-labeled model to make predictions on the evaluation dataset to obtain the prediction results;

[0041] S4, Calculate the evaluation index: Compare the model's prediction results with the true labels in the evaluation dataset to calculate the corresponding values ​​according to the selected evaluation index formula;

[0042] S5, Generate Evaluation Report: Organize the calculated evaluation indicators and generate an evaluation report in an intuitive form; the report not only presents the specific values ​​of the evaluation indicators, but also analyzes and interprets the indicator results to explain the model's performance in different aspects; at the same time, it compares with historical evaluation results or the evaluation results of other models to help users more clearly understand the model's performance level and room for improvement.

[0043] As a preferred embodiment of the AI ​​data annotation and prompt word automatic construction engine system described in this invention, the specific steps for prompt word-guided model evaluation are as follows:

[0044] S1. Define the evaluation objectives and dimensions: Based on the application scenarios and task requirements of the prompt words, determine the specific evaluation objectives and dimensions;

[0045] S2, Set evaluation criteria: Develop detailed evaluation criteria and scoring rules for each evaluation dimension;

[0046] S3, Collect Generation Results: Apply the prompt words to the target model and collect the results generated by the model based on the prompt words;

[0047] S4, manual or automated evaluation:

[0048] Manual evaluation: Invite professionals or target user groups to score and evaluate the generated results according to the established evaluation criteria; to ensure the consistency and objectivity of the evaluation, train the evaluators before the evaluation and standardize the evaluation criteria and process.

[0049] Automated evaluation: Develop automated evaluation algorithms using techniques from the fields of natural language processing or computer vision;

[0050] S5, Analysis and Feedback: Perform statistical analysis on the evaluation results, calculate the average score and standard deviation of each evaluation dimension to understand the overall level and fluctuation of the generated results; and based on the analysis results, identify the problems and deficiencies of the prompt words, and then feed the evaluation results back to the prompt word optimization module to provide a basis for the adjustment and improvement of prompt words, forming a closed-loop process of evaluation-optimization.

[0051] Compared with existing technologies:

[0052] 1. Addressing the issue of low data annotation efficiency: By adopting semi-automatic annotation and active learning annotation functions, not only can the workload be significantly reduced, but also unnecessary annotation work can be reduced, thereby significantly improving annotation efficiency, shortening annotation time, meeting the needs of large-scale data annotation, and solving the efficiency bottleneck caused by slow manual annotation.

[0053] 2. Addressing the issue of high data annotation costs: By improving annotation efficiency and reducing manual workload, the need for a large number of annotation personnel can be reduced, thereby reducing expenses such as recruitment, training, and annotation tool procurement, effectively lowering annotation costs and alleviating the financial pressure caused by building an annotation team.

[0054] 3. Addressing the issue of difficulty in ensuring data annotation consistency: The annotation quality monitoring function can not only monitor the annotation data in real time, but also promptly identify annotation differences between different annotators through consistency checks and error rate statistics, reminding them to review, ensuring uniform annotation standards, improving annotation consistency, guaranteeing the accuracy of annotation data, providing high-quality data for model training, and avoiding the impact of inconsistent annotations on model training results.

[0055] 4. Addressing the issue of relying solely on human experience in prompt word generation: The prompt word generation function based on task analysis not only deeply analyzes the user's task description but also generates initial prompt words by combining domain knowledge and model characteristics. This changes the previous reliance on human experience, providing a scientific basis for prompt word generation and improving the accuracy and effectiveness of initial prompt words, thus guiding the model to generate more relevant content.

[0056] 5. Addressing the lack of adaptability in prompt word construction: The prompt word optimization function can utilize reinforcement learning technology to iteratively optimize the model based on its performance in real-world tasks. When the task scenario or model changes, the system can automatically adjust the prompt words to adapt to the new situation without the need for manual reconstruction, saving time and effort and ensuring that the prompt words can play a good role in different situations.

[0057] 6. Addressing the cumbersome process of constructing prompt words: The prompt word library management function categorizes and stores the generated prompt words, which not only facilitates user retrieval and reuse but also allows users to modify and improve based on existing prompt words. This reduces the tedious process of repeatedly trying different combinations of prompt words and improves the efficiency of prompt word construction. For example, in image generation tasks, users can quickly find relevant prompt words and speed up task completion. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the framework of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0060] This invention provides an AI-powered data annotation and prompt word automatic construction engine system. Please refer to [link / reference]. Figure 1 ,include:

[0061] The data annotation module first performs preliminary annotation on the data based on the pre-trained model, and then automatically annotates the data samples through an active learning algorithm. At the same time, it monitors the quality of the annotated data in real time.

[0062] The automatic prompt word construction module first generates prompt words based on task analysis, then optimizes the prompt words using reinforcement learning techniques, and at the same time, classifies, stores and manages the generated prompt words.

[0063] The system management module is used to first manage user permissions, then manage data annotation tasks and prompt word construction tasks, and at the same time, it can adjust the relevant parameters of data annotation algorithms and prompt word generation algorithms;

[0064] The model evaluation module is used to comprehensively evaluate the performance of the model after data labeling and training using evaluation metrics. At the same time, it analyzes the degree of fit between the generated content and the expected target for the results generated by the prompt word automatic construction module.

[0065] By setting up a model evaluation module, quantitative feedback can be provided for data annotation and prompt word construction. Based on the evaluation of the model, deficiencies in the labeled data can be identified, such as inaccurate or insufficient labeling of certain categories, allowing for targeted optimization of the annotation work. For prompt word construction, the evaluation results enable timely adjustments to the prompt word optimization strategy, quickly finding more effective prompt word combinations, avoiding blind attempts, improving the efficiency and quality of prompt word construction, and ultimately enhancing the overall system's usability and user satisfaction.

[0066] The data annotation module includes:

[0067] The data preprocessing module is used to clean, transform, and extract data features from the raw data.

[0068] By setting up a data preprocessing module, the quality of input data can be significantly improved, and labeling errors caused by problems with the data itself can be reduced.

[0069] Semi-automatic annotation is used to perform preliminary annotation of data using pre-trained models. For image classification tasks, pre-trained deep learning models such as ResNet and VGG can be selected; for text classification tasks, pre-trained language models such as BERT and RoBERTa can be selected; for object detection tasks, models such as Faster R-CNN and YOLO are commonly used choices. Then, manual review and correction are performed to improve annotation efficiency.

[0070] The active learning annotation module is used to automatically select the most valuable data samples for annotation through active learning algorithms, so as to improve the quality and representativeness of the labeled data. At the same time, based on the uncertainty of the model and the diversity of the samples, it will select the most needed samples from the massive unlabeled data, allowing the annotators to annotate them, so that the labeled data can better cover various situations and improve the model training effect.

[0071] The annotation quality monitoring module is used to monitor the quality of annotation data in real time, and to promptly identify and correct problems in the annotation process through consistency checks and error rate statistics.

[0072] The specific steps of the data preprocessing module are as follows:

[0073] S1, Data Collection and Integration: First, raw data is obtained from multiple data sources, which may include databases, file systems, web crawlers, etc. Then, data of different formats and structures are integrated to form a unified data set.

[0074] S2, Data Cleaning: First, identify and delete duplicate data records by comparing key features or all attributes of the data; then, handle missing values ​​in the data; if there are few missing values, data records containing missing values ​​can be deleted directly; for numerical data, the mean, median, or mode can be used to fill missing values; for categorical data, the most frequent category can be used to fill missing values; finally, based on the business rules and statistical patterns of the data, identify erroneous data and correct them according to the actual situation.

[0075] S3, Data Transformation: First, convert data of different formats into a standard format that the system can process; then, according to the needs of data processing, convert data from one type to another; for example, convert string-type numeric data into numeric data for mathematical operations; convert date and time data into a specific date and time format for time series analysis; then, for categorical data, use appropriate encoding methods for conversion, such as one-hot encoding, to convert each category into a binary vector to meet the input requirements of machine learning algorithms;

[0076] S4, Data Feature Extraction: First, extract more valuable features from the original data. In image data, features such as image edges and textures can be extracted. In text data, feature vectors can be extracted using methods such as the bag-of-words model and TF-IDF. Then, from the numerous extracted features, select the most relevant and representative features to remove redundant and irrelevant features. Feature selection can be performed using methods such as filtering (e.g., variance-based selection, correlation analysis), wrapping (e.g., recursive feature elimination), and embedding (e.g., Lasso regression) to reduce data dimensionality and improve the efficiency and effectiveness of subsequent processing. Finally, the data is validated to check its integrity, accuracy, and consistency.

[0077] The automatic prompt word construction module includes:

[0078] The prompt word generation module is used to perform in-depth analysis of the task description input by the user, and generate initial prompt words by combining relevant domain knowledge and model characteristics. For example, when a user needs to build a prompt word for generating intelligent customer service responses, the system will analyze common customer service question types, user needs, and language style requirements for customer service responses, and generate initial prompt words containing relevant keywords and sentence structures.

[0079] The prompt word optimization module is used to continuously optimize prompt words based on the model's performance in real-world tasks using reinforcement learning techniques, thereby improving the effectiveness of the prompt words;

[0080] The prompt word library management module is used to classify, store, and manage the generated prompt words, making it convenient for users to query and reuse them. Users can quickly find existing relevant prompt words in the prompt word library based on keywords such as task type and field. If the existing prompt words cannot fully meet the needs, they can be modified and improved to improve the efficiency of prompt word construction.

[0081] The system management module includes:

[0082] The user management module is used for user permission management, including permission settings for different roles such as annotators, prompt word builders, and administrators, to ensure the security and standardization of system operations. For example, annotators can only perform data annotation-related operations, prompt word builders focus on prompt word generation and optimization, and administrators have advanced permissions such as system configuration and user management.

[0083] The task management module is used to manage the creation, allocation, and progress tracking of data annotation tasks and prompt word construction tasks. Administrators can create data annotation tasks according to project requirements, assign tasks to appropriate annotators, and view task progress in real time to ensure that projects are completed on time. For prompt word construction tasks, tasks can also be created and allocated, and the generation and optimization progress of prompt words can be tracked.

[0084] The settings module is used to adjust the relevant parameters of the data annotation algorithm and the prompt word generation algorithm according to the actual use scenario and data characteristics. For example, it can adjust the sample selection threshold in the active learning annotation algorithm or adjust the reward function parameters in the prompt word optimization algorithm, so that the system can better perform.

[0085] The model evaluation module includes data-annotated model evaluation and prompt-word-guided model evaluation.

[0086] The specific steps for evaluating the data annotation model are as follows:

[0087] S1. Determine the evaluation metrics: Select appropriate evaluation metrics based on the type and objective of the data labeling task. For classification tasks, commonly used metrics include accuracy, recall, F1 score, and confusion matrix. For regression tasks, metrics such as mean squared error (MSE), mean absolute error (MAE), and coefficient of determination (R2) can be used. For example, in image classification labeling tasks, accuracy is used to evaluate the proportion of images correctly classified by the model, and the confusion matrix is ​​used to visually display the model's classification errors in different categories.

[0088] S2, Splitting the Evaluation Dataset: A portion of the labeled dataset is partitioned as the evaluation dataset. This can be achieved using hold-out or cross-validation. Hold-out involves dividing the dataset into training and test sets, such as a 7:3 ratio. Cross-validation typically divides the dataset into k parts, using k-1 parts alternately as the training set and one part as the test set, repeating this process k times. The average of the k evaluation results is then used as the model's evaluation metric. It is crucial to ensure that the evaluation dataset is representative and reflects the characteristics and distribution of the overall data.

[0089] S3, Model Prediction: The trained data labeling model is used to make predictions on the evaluation dataset to obtain the prediction results; in the text sentiment classification labeling task, the model predicts the sentiment category of the text in the evaluation dataset and outputs the predicted sentiment label (such as positive, negative, neutral).

[0090] S4, Calculate the evaluation index: Compare the model's prediction results with the true labels in the evaluation dataset to calculate the corresponding values ​​according to the selected evaluation index formula; for example, when calculating the accuracy, divide the number of correctly predicted samples by the total number of samples; when calculating the mean square error, sum the squares of the differences between the predicted value and the true value for each sample and then take the average.

[0091] S5, Generate Evaluation Report: Organize the calculated evaluation indicators and generate an evaluation report in an intuitive form; the report not only presents the specific values ​​of the evaluation indicators, but also analyzes and interprets the indicator results to explain the model's performance in different aspects, such as which categories have high classification accuracy and which categories are prone to misclassification; at the same time, it compares with historical evaluation results or the evaluation results of other models to help users more clearly understand the model's performance level and room for improvement.

[0092] The specific steps for evaluating the prompt-guided model are as follows:

[0093] S1. Define Evaluation Objectives and Dimensions: Based on the application scenario and task requirements of the prompt words, determine the specific evaluation objectives and dimensions. For example, in text generation tasks, evaluation dimensions may include grammatical accuracy, semantic coherence, content relevance, and information completeness. In image generation tasks, evaluation can be conducted from dimensions such as image clarity, composition rationality, element accuracy, and style consistency. If used in intelligent customer service scenarios, the practicality and politeness of the response should also be considered.

[0094] S2, Set Evaluation Criteria: For each evaluation dimension, develop detailed evaluation criteria and scoring rules; for example, for the grammatical accuracy of text generation, it can be stipulated that no grammatical errors will receive full marks, and a certain number of points will be deducted for each grammatical error; for the clarity of image generation, different scoring levels can be set according to image resolution, level of detail, etc.; expert scoring methods or industry standards can be used to ensure the scientific and reasonable nature of the evaluation criteria.

[0095] S3, Collect Generation Results: Apply the prompt words to the target model and collect the results generated by the model based on the prompt words; in the intelligent writing assistance task, collect multiple articles generated using different prompt words; in the image generation task, obtain the image set generated under the corresponding prompt words;

[0096] S4, manual or automated evaluation:

[0097] Manual evaluation: Invite professionals or target user groups to score and evaluate the generated results according to the established evaluation criteria; to ensure the consistency and objectivity of the evaluation, train the evaluators before the evaluation and standardize the evaluation criteria and process.

[0098] Automated evaluation: Utilizing techniques from natural language processing or computer vision, automated evaluation algorithms are developed. In text generation evaluation, content relevance can be assessed by calculating the semantic similarity between the generated text and the reference text (e.g., using indicators such as cosine similarity and BLEU score). In image generation evaluation, image quality assessment indicators (e.g., PSNR, SSIM) can be used to evaluate image sharpness and quality. Automated evaluation can quickly process a large number of generated results, improving evaluation efficiency, but it may have certain limitations. It can be combined with manual evaluation to complement each other.

[0099] S5, Analysis and Feedback: Perform statistical analysis on the evaluation results, calculate the average score and standard deviation of each evaluation dimension to understand the overall level and fluctuation of the generated results; and based on the analysis results, identify the problems and deficiencies of the prompt words, such as which prompt words lead to low scores in certain dimensions, and then feed the evaluation results back to the prompt word optimization module to provide a basis for the adjustment and improvement of prompt words, forming a closed-loop process of evaluation-optimization.

[0100] In practical use, the specific steps are as follows:

[0101] Step 1: The data preprocessing module cleans, transforms, and extracts features from the raw data. After preprocessing, a semi-automatic annotation process using a pre-trained model performs initial annotation, followed by manual review and correction to improve annotation efficiency. Next, the active learning annotation module uses an active learning algorithm to automatically select the most valuable data samples for annotation, improving the quality and representativeness of the labeled data. Simultaneously, considering model uncertainty and sample diversity, the module selects the most critical samples from the massive unlabeled data for annotation, ensuring the labeled data better covers various scenarios and improves model training effectiveness. Finally, the annotation quality monitoring module monitors the quality of the labeled data in real time, using consistency checks and error rate statistics to promptly identify and correct problems during the annotation process.

[0102] Step 2: The prompt word generation module performs in-depth analysis of the user-input task description, combining relevant domain knowledge and model characteristics to generate initial prompt words. After generation, the prompt word optimization module uses reinforcement learning techniques to continuously optimize the prompt words based on the model's performance in actual tasks, thereby improving their effectiveness. At the same time, the prompt word library management module categorizes, stores, and manages the generated prompt words for easy user retrieval and reuse.

[0103] Step 3: Manage user permissions through the user management module to ensure the security and standardization of system operations; after permission management, manage the creation, allocation, and progress tracking of data annotation tasks and prompt word construction tasks through the task management module; at the same time, adjust the relevant parameters of the data annotation algorithm and prompt word generation algorithm through the settings module according to the actual use scenario and data characteristics.

[0104] Step 4: Using the model evaluation module, the model used for training after data labeling is comprehensively evaluated using evaluation metrics. At the same time, the results generated by the automatic prompt word construction module are analyzed to determine the degree of fit between the generated content and the expected goals.

[0105] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. An AI-powered data annotation and prompt word automatic construction engine system, characterized in that, include: The data annotation module first performs preliminary annotation on the data based on the pre-trained model, and then automatically annotates the data samples through an active learning algorithm. At the same time, it monitors the quality of the annotated data in real time. The automatic prompt word construction module first generates prompt words based on task analysis, then optimizes the prompt words using reinforcement learning techniques, and at the same time, classifies, stores and manages the generated prompt words. The system management module is used to first manage user permissions, then manage data annotation tasks and prompt word construction tasks, and at the same time, it can adjust the relevant parameters of data annotation algorithms and prompt word generation algorithms; The model evaluation module is used to comprehensively evaluate the performance of the model after data labeling and training using evaluation metrics. At the same time, it analyzes the degree of fit between the generated content and the expected target for the results generated by the prompt word automatic construction module.

2. The AI ​​data annotation and prompt word automatic construction engine system according to claim 1, characterized in that, The data annotation module includes: The data preprocessing module is used to clean, transform, and extract data features from the raw data. Semi-automatic annotation is used to perform preliminary annotation of data using a pre-trained model, which is then reviewed and corrected by humans, thereby improving annotation efficiency. The active learning annotation module is used to automatically select the most valuable data samples for annotation through active learning algorithms, so as to improve the quality and representativeness of the labeled data. At the same time, based on the uncertainty of the model and the diversity of the samples, it will select the most needed samples from the massive unlabeled data, allowing the annotators to annotate them, so that the labeled data can better cover various situations and improve the model training effect. The annotation quality monitoring module is used to monitor the quality of annotation data in real time, and to promptly identify and correct problems in the annotation process through consistency checks and error rate statistics.

3. The AI ​​data annotation and prompt word automatic construction engine system according to claim 2, characterized in that, The specific steps of the data preprocessing module are as follows: S1, Data Collection and Integration: First, raw data is obtained from multiple data sources, and then data of different formats and structures are integrated together to form a unified data set; S2, data cleaning, firstly identifies and deletes duplicate data records by comparing key features or all attributes of the data; then it processes missing values ​​in the data; finally, it identifies erroneous data based on the business rules and statistical patterns of the data, and corrects it according to the actual situation. S3, Data Transformation: First, convert data of different formats into a standard format that the system can process; then, according to the needs of data processing, convert the data from one type to another; then, for categorical data, use an appropriate encoding method to convert each category into a binary vector to meet the input requirements of machine learning algorithms. S4, Data Feature Extraction: First, extract more valuable features from the original data; then, select the most relevant and representative features from the extracted features to remove redundant and irrelevant features; finally, verify the data to check its integrity, accuracy, and consistency.

4. The AI ​​data annotation and prompt word automatic construction engine system according to claim 1, characterized in that, The automatic prompt word construction module includes: The prompt word generation module is used to perform in-depth analysis of the task description input by the user, and generate initial prompt words by combining relevant domain knowledge and model characteristics; The prompt word optimization module is used to continuously optimize prompt words based on the model's performance in real-world tasks using reinforcement learning techniques, thereby improving the effectiveness of the prompt words; The prompt word library management module is used to classify, store, and manage the generated prompt words, making it convenient for users to query and reuse them.

5. The AI ​​data annotation and prompt word automatic construction engine system according to claim 1, characterized in that, The system management module includes: The user management module is used to manage user permissions and ensure the security and standardization of system operations. The task management module is used to manage the creation, allocation, and progress tracking of data annotation tasks and prompt word construction tasks; The settings module is used to adjust the relevant parameters of the data annotation algorithm and the prompt word generation algorithm according to the actual use scenario and data characteristics.

6. The AI ​​data annotation and prompt word automatic construction engine system according to claim 1, characterized in that, The model evaluation module includes data-annotated model evaluation and prompt-word-guided model evaluation.

7. The AI ​​data annotation and prompt word automatic construction engine system according to claim 6, characterized in that, The specific steps for evaluating the data annotation model are as follows: S1, Determine the evaluation indicators: Select appropriate evaluation indicators based on the type and objectives of the data labeling task; S2, Divide the evaluation dataset: Divide a portion of the labeled dataset as the evaluation dataset, including hold-out and cross-validation methods; S3, Model Prediction: Use the trained data-labeled model to make predictions on the evaluation dataset to obtain the prediction results; S4, Calculate the evaluation index: Compare the model's prediction results with the true labels in the evaluation dataset to calculate the corresponding values ​​according to the selected evaluation index formula; S5, Generate Evaluation Report: Organize the calculated evaluation indicators and generate an evaluation report in an intuitive form; the report not only presents the specific values ​​of the evaluation indicators, but also analyzes and interprets the indicator results to explain the model's performance in different aspects; at the same time, it compares with historical evaluation results or the evaluation results of other models to help users more clearly understand the model's performance level and room for improvement.

8. The AI ​​data annotation and prompt word automatic construction engine system according to claim 6, characterized in that, The specific steps for evaluating the prompt-guided model are as follows: S1. Define the evaluation objectives and dimensions: Based on the application scenarios and task requirements of the prompt words, determine the specific evaluation objectives and dimensions; S2, Set evaluation criteria: Develop detailed evaluation criteria and scoring rules for each evaluation dimension; S3, Collect Generation Results: Apply the prompt words to the target model and collect the results generated by the model based on the prompt words; S4, manual or automated evaluation: Manual evaluation: Invite professionals or target user groups to score and evaluate the generated results according to the established evaluation criteria; to ensure the consistency and objectivity of the evaluation, train the evaluators before the evaluation and standardize the evaluation criteria and process. Automated evaluation: Develop automated evaluation algorithms using techniques from the fields of natural language processing or computer vision; S5, Analysis and Feedback: Perform statistical analysis on the evaluation results, calculate the average score and standard deviation of each evaluation dimension to understand the overall level and fluctuation of the generated results; and based on the analysis results, identify the problems and deficiencies of the prompt words, and then feed the evaluation results back to the prompt word optimization module to provide a basis for the adjustment and improvement of prompt words, forming a closed-loop process of evaluation-optimization.

Citation Information

Patent Citations

  • Small sample text labeling method and device based on active learning

    CN115129872A

  • Language model and active learning-based data annotation method, medium and equipment

    CN117669764A

  • Automatic driving data labeling system and method

    CN117763348A

  • Equipment state prediction method based on weak supervised learning

    CN119179894A

  • Model cue word automatic optimization method and device, equipment and storage medium

    CN119226476A

Cited By

  • Data labeling method and system based on cue word driving

    CN121683727A

  • A prompt word driven data labeling method and system

    CN121683727B