Method, system and device for analyzing annotation data
By analyzing and verifying log data and merging and analyzing multi-source databases, the problems of poor query performance and high data redundancy in existing technologies are solved, accurate evaluation and real-time interactive analysis of labeled data are achieved, and the efficiency and visualization capabilities of data management are improved.
Patent Information
- Application Number
- CN202511106437.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing data annotation and analysis platforms have poor query performance under large data volumes, lack unified granularity standards, have high data redundancy, have limited charting capabilities, and are unable to support real-time interactive analysis, making it difficult for project managers to accurately assess the workload and quality of data annotation personnel.
By parsing and verifying the log data stored in the topic queue, identifying anomalies and merging key fields, annotating result indexes, and quality assessment data, using multi-source heterogeneous database information for analysis and evaluation, and using distributed system clusters and analytical databases for data storage and calculation, real-time interactive analysis capabilities are provided.
It achieves accurate analysis and evaluation of labeled data, improves query efficiency and response speed, reduces data redundancy, supports multi-dimensional cross-analysis, and ensures real-time data updates and visualization depth.
Smart Images

Figure CN120597007A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a method, system, and device for analyzing and annotating data. Background Art
[0002] In the data annotation industry, various data analysis platforms and tools are used in daily operations to effectively manage annotation output, quality inspection results, and workload statistics. The Jishiguan Intelligent Data Annotation Management Platform integrates task management, annotation tools, quality review, and performance statistics to help teams track annotation progress and quality. The platform records the completion status of each annotation task, quality inspection feedback (pass or reject), and the annotator's time spent, and regularly generates statistical reports. The existing platform system provides a "dashboard" that displays key indicators such as total annotation volume, quality inspection rejections, effective working hours, and annotation efficiency, and supports viewing data at the overall, project, and individual levels. Quality inspections typically ensure annotation quality through randomized reviews or two-person audits. The results are recorded in a quality inspection log to calculate each annotator's pass rate and error distribution. Furthermore, workload statistics are typically aggregated on a daily or project basis. For example, each annotator's daily number of annotations (distinguished by pass and reject) and hourly annotation efficiency are counted for performance evaluation. Existing platforms allow switching of statistical calibers, such as counting by "samples" or by "operation units", to reflect workloads of different granularities.
[0003] While existing technical solutions have provided some assistance for data annotation and analysis, they still suffer from several objective drawbacks, primarily manifested in the following areas: Inconsistent data granularity: Different platforms or statistical methods use inconsistent measurement methods for labeled data, potentially resulting in the coexistence of "sample" and "operation unit" units. For example, within the same batch of data, a multi-label annotation task can be treated as either a single sample or multiple records based on the number of labels actually annotated. Without a unified granularity standard, data alignment between different reports is difficult, making interpretation difficult for project managers. Poor query performance: Traditional solutions often rely on relational databases or offline reporting, resulting in slow query response times when data volumes are large. As annotation tasks accumulate to millions of records, complex joined-table queries using standard MySQL databases are time-consuming and struggle to return results in a timely manner. Even some systems that utilize offline aggregation methods can only provide a small number of predefined metrics and lack real-time, interactive analysis capabilities. Overall, existing solutions offer unsatisfactory query performance for large data volumes, lacking the ability to respond in seconds even for petabyte-scale data. High data redundancy: Due to the lack of a unified data flow architecture, existing solutions often store data in multiple locations. For example, in existing annotation platform management systems, in order to increase query speed, multiple redundant tables such as detailed lists, daily summaries, and weekly summaries may be saved at the same time; or data may be exported to Excel for secondary analysis, forming data islands outside the system. This duplication not only increases storage and maintenance costs, but also introduces the risk of data inconsistency when multiple copies of different report data are not synchronized in time due to network delays or processing anomalies. Limited charting capabilities: Many traditional annotation statistics tools only provide preset static charts and lack flexible interactive analysis methods. Common pie charts, bar charts, and line charts are limited in variety, making it difficult for managers to customize analysis dimensions. For example, when cross-analysis is required by multiple dimensions such as project, personnel, and time, existing tools often cannot provide support. In addition, reports are usually not updated in real time, making it difficult for project managers to obtain the latest production status in a timely manner, and the depth and breadth of data visualization are lacking.
[0004] Therefore, how to accurately analyze and evaluate the labeled data completed by the data labelers is a technical problem that needs to be solved. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide a method for analyzing and labeling data. Through the technical solutions of the embodiments of the present application, it is possible to accurately analyze and evaluate the labeling data completed by the labeling data personnel.
[0006] In the first aspect, an embodiment of the present application provides a method for analyzing and annotating data, which is applied to an analysis center, including parsing and verifying log data stored in a subject queue to determine whether the log data has any abnormalities, wherein the log data includes: the annotation content of the data annotator and the corresponding quality indicator data; when it is determined that the log data has an abnormality, according to a preset processing strategy, the log data and the current annotation task and structured data related to the target user are processed and merged to obtain data to be analyzed, wherein the structured data includes: basic task information, basic information of the data annotator and quality verification standards, and the processing strategy includes a discarding strategy and a correction strategy; the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed are merged to obtain merged data; based on the merged data, the annotation data completed by the data annotator within a preset period is analyzed and evaluated to obtain analysis results.
[0007] In the above embodiments of the present application, by detecting anomalies in log data, multi-source heterogeneous database information such as structured data, key fields, annotation result indexes and quality assessment data related to the log data can be obtained. By simultaneously analyzing and evaluating the total of the above-mentioned multiple information, the effect of accurately analyzing and evaluating the annotation data completed by the annotation data personnel can be achieved.
[0008] In some embodiments, before parsing and verifying the log data stored in the subject queue to determine whether there is an abnormality in the log data, it also includes: obtaining the log report corresponding to the current task; preprocessing the data in the log report to obtain log data, wherein the preprocessing includes: cleaning, filtering, supplementing and packaging; and storing the log data in the subject queue.
[0009] In the above-mentioned embodiment of the present application, the influence of irrelevant data can be removed after preprocessing the log data, thereby providing more accurate data for subsequent log data analysis.
[0010] In some embodiments, key fields are extracted by querying unstructured data and semi-structured data related to the log data from a database.
[0011] In the above embodiment of the present application, key fields are extracted from the database query of unstructured data and semi-structured data related to log data, which can be used as an analysis basis for marking data keys.
[0012] In some embodiments, before merging the key fields, annotation result index, quality assessment data and data to be analyzed of the log data to obtain the merged data, it also includes: storing the data to be analyzed and the key fields in the memory; querying the corresponding annotation result index and quality assessment data according to the identification information in the log report, and storing the annotation result index and quality assessment data in the memory.
[0013] In the above embodiments of the present application, by storing the annotation result index, quality assessment data, analysis data and key fields in the memory, the data in the memory can be quickly obtained for data analysis when subsequent data analysis is performed.
[0014] In some embodiments, after analyzing and evaluating the labeled data completed by the data labeling personnel within a preset period based on the merged data and obtaining the analysis results, it also includes: storing the analysis results in a preset database wide table and updating the index analysis results.
[0015] In the above embodiment of the present application, by using the data index stored in the wide table of the database, the query content can be quickly obtained when subsequently querying the labeled data and analysis results.
[0016] In some embodiments, the log data stored in the subject queue is parsed and verified to determine whether the log data is abnormal, including: determining whether each field of the log data conforms to an expected format.
[0017] In the above embodiment of the present application, by determining whether each field of the log data conforms to the expected format, it is possible to accurately determine whether the log data has an anomaly.
[0018] In some embodiments, log data, the current annotation task, and structured data related to the target user are processed and merged according to a preset processing strategy, including: deleting or correcting log data and structured data according to the processing strategy; merging log data and structured data to obtain data to be analyzed.
[0019] In the above embodiments of the present application, by merging the above data, the data to be analyzed can be obtained, which facilitates the evaluation of subsequent labeled data.
[0020] In a second aspect, an embodiment of the present application provides a system for analyzing annotated data, including: Annotation tool system, annotation platform system, message middleware, business database, distributed system cluster, distributed file storage database, analytical database, second message middleware, streaming computing module and front-end query service module; The annotation tool system is used to query business databases, distributed system clusters, and distributed file storage databases to generate log data for the annotation task process: The annotation platform system is used to receive log data from the annotation tool system, parse and verify the log data stored in the subject queue, and determine whether the log data is abnormal. When it is determined that the log data is abnormal, it processes and merges the log data with the current annotation task and the structured data related to the target user according to the preset processing strategy to obtain the data to be analyzed, merges the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed to obtain the merged data, and analyzes and evaluates the annotation data completed by the annotation personnel within the preset period based on the merged data to obtain the analysis results: Message middleware, used for asynchronous decoupling and distributed data processing of various data in the annotation platform system: Analytical database, used to store and index analysis results through database wide tables; The second message middleware is used to send various data to the stream computing module The streaming computing module is used to perform batch processing and data computing on various data.
[0021] The front-end query service module is used to provide users with query and visualization functions of analysis results.
[0022] In a third aspect, an embodiment of the present application provides a device for analyzing labeled data, including: A parsing and verification module is used to parse and verify the log data stored in the subject queue to determine whether the log data is abnormal, wherein the log data includes: the annotation content of the data annotator and the corresponding quality indicator data; a processing module configured to, upon determining that the log data is abnormal, process and merge the log data, the current annotation task, and structured data related to the target user according to a preset processing strategy to obtain data to be analyzed, wherein the structured data includes: basic information of the task, basic information of the person annotating the data, and quality verification standards; and the processing strategy includes a discarding strategy and a correction strategy; A merging module, configured to merge the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed to obtain merged data; The analysis module is used to analyze and evaluate the labeled data completed by the data labeling personnel within a preset period based on the merged data to obtain analysis results.
[0023] Optionally, the device further includes: A pre-processing module is used for obtaining a log report corresponding to the current task before the parsing and verification module parses and verifies the log data stored in the subject queue to determine whether the log data is abnormal; Preprocess the data in the log report to obtain log data, where the preprocessing includes: cleaning, filtering, supplementing and packaging; Store log data in a topic queue.
[0024] Optionally, key fields are extracted by querying unstructured data and semi-structured data related to the log data from a database.
[0025] Optionally, the device further includes: A storage module, configured to store the data to be analyzed and the key fields in a memory before the merging module merges the key fields of the log data, the annotation result index, the quality assessment data, and the data to be analyzed to obtain the merged data; The corresponding annotation result index and quality assessment data are queried according to the identification information in the log report, and the annotation result index and quality assessment data are stored in the memory.
[0026] Optionally, the device further includes: The index module is used for the analysis module to analyze and evaluate the labeled data completed by the labeling personnel within a preset period based on the merged data, and after obtaining the analysis results, store the analysis results in a preset database wide table and update the index analysis results.
[0027] Optionally, the parsing and verification module is specifically used to: Determine whether the fields of the log data conform to the expected format.
[0028] Optionally, the analysis module is specifically used to: Deleting or amending log data and structured data in accordance with processing policies; Merge log data and structured data to obtain data to be analyzed.
[0029] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps of the method provided in the first aspect above are executed.
[0030] In a fifth aspect, an embodiment of the present application provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the method provided in the first aspect above are executed.
[0031] Other features and advantages of the present application will be described in the subsequent description, and in part will become apparent from the description, or will be understood by practicing the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0033] Figure 1 A flowchart of a method for analyzing labeled data provided in an embodiment of the present application; Figure 2 A schematic block diagram of a system for analyzing and annotating data provided in an embodiment of the present application; Figure 3 A schematic block diagram of a device for analyzing and annotating data provided in an embodiment of the present application; Figure 4 A schematic block diagram of the structure of a device for analyzing and annotating data provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0035] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0036] First, some of the terms involved in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.
[0037] An Elasticsearch cluster is a distributed system consisting of multiple nodes that jointly store data and provide search services, achieving high availability and scalability through mechanisms such as sharding and replication.
[0038] MongoDB is a distributed file-based database written in C++. It is designed to provide a scalable, high-performance data storage solution for web applications.
[0039] This application is applied to the scenario of evaluating labeled data. The specific scenario is to aggregate information from multiple heterogeneous databases such as MySQL (database), Elasticsearch and MongoDB, and store it uniformly in a large wide table of SelectDB, which significantly improves the query and analysis efficiency and response speed of massive labeled data, and avoids the performance bottlenecks and complexity problems caused by multi-database joint queries in traditional technologies.
[0040] In the data annotation industry, various data analysis platforms and tools are used in daily operations to effectively manage annotation output, quality inspection results, and workload statistics. The Jishiguan Intelligent Data Annotation Management Platform integrates task management, annotation tools, quality review, and performance statistics to help teams track annotation progress and quality. The platform records the completion status of each annotation task, quality inspection feedback (pass or reject), and the annotator's time spent, and regularly generates statistical reports. The existing platform system provides a "dashboard" that displays key indicators such as total annotation volume, quality inspection rejections, effective working hours, and annotation efficiency, and supports viewing data at the overall, project, and individual levels. Quality inspections typically ensure annotation quality through randomized reviews or two-person audits. The results are recorded in a quality inspection log to calculate each annotator's pass rate and error distribution. Furthermore, workload statistics are typically aggregated on a daily or project basis. For example, each annotator's daily number of annotations (distinguished by pass and reject) and hourly annotation efficiency are counted for performance evaluation. Existing platforms allow for switching statistical calibers, such as counting by "sample" or by "operation unit," to reflect workloads at different granularities. While existing technical solutions have provided some assistance for data annotation and analysis, they still suffer from several objective drawbacks, primarily manifested in the following areas: Inconsistent data granularity: Different platforms or statistical methods do not uniformly measure labeled data, potentially resulting in the coexistence of "sample" and "operation unit" units. For example, within the same batch of data, a multi-label annotation task can be treated as a single sample or counted as multiple records based on the actual number of labels. Without a unified granularity standard, data alignment between different reports is difficult, creating difficulties for project managers in interpreting data. Poor query performance: Traditional solutions often rely on relational databases or offline reporting, resulting in slow query response when dealing with large data volumes. As annotation tasks accumulate to millions of records, complex joined-table queries using standard MySQL systems take a long time and struggle to return results in a timely manner. Even some systems that employ offline aggregation methods can only provide a small number of predefined metrics and lack real-time, interactive analysis capabilities. Overall, the query performance of existing solutions under large data volumes is unsatisfactory, and it lacks support for second-level responses for PB-level data. High data redundancy: Due to the lack of a unified data flow architecture, existing solutions often have multiple data duplicate storage locations. For example, in the existing annotation platform management system, in order to increase query speed, multiple redundant tables such as detailed lists, daily summaries, and weekly summaries may be saved at the same time; or data may be exported to Excel for secondary analysis, forming data islands outside the system. This duplication not only increases storage and maintenance costs, but also introduces the risk of data inconsistency when multiple copies of different report data are not synchronized in time due to network delays or processing anomalies. Limited charting capabilities: Many traditional annotation and statistical tools only provide preset static charts and lack flexible interactive analysis methods.Common pie charts, bar charts, and line charts are limited in variety and struggle to support managers' custom analysis dimensions. For example, existing tools often fail to support cross-analysis by project, personnel, and timeframe. Furthermore, reports often lack real-time updates, making it difficult for project managers to obtain the latest production status. Data visualization lacks depth and breadth.
[0041] To this end, this application parses and verifies the log data stored in the subject queue to determine whether the log data is abnormal, wherein the log data includes: the annotation content of the data annotation personnel and the corresponding quality indicator data; when it is determined that the log data is abnormal, according to the preset processing strategy, the log data and the current annotation task and the structured data related to the target user are processed and merged to obtain the data to be analyzed, wherein the structured data includes: basic information of the task, basic information of the data annotation personnel and quality verification standards, and the processing strategy includes a discarding strategy and a correction strategy; the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed are merged to obtain the merged data; based on the merged data, the annotation data completed by the data annotation personnel within the preset period is analyzed and evaluated to obtain the analysis results. By detecting anomalies in the log data, multi-source heterogeneous database information such as structured data, key fields, annotation result index and quality assessment data related to the log data can be obtained. By simultaneously analyzing and evaluating the sum of the above multiple information, the effect of accurately analyzing and evaluating the annotation data completed by the data annotation personnel can be achieved.
[0042] In the embodiment of the present application, the execution entity may be an analysis and annotation data device in the analysis and annotation data system. In actual applications, the analysis and annotation data device may be an electronic device such as a terminal device and a server, and no limitation is made here.
[0043] The following combination Figure 1 The method for analyzing and annotating data in an embodiment of the present application is described in detail.
[0044] Please see Figure 1 , Figure 1 A flowchart of a method for analyzing labeled data provided in an embodiment of the present application is applied to an analysis center, such as Figure 1 The method for analyzing the labeled data shown includes: Step 110: parse and verify the log data stored in the subject queue to determine whether the log data is abnormal.
[0045] Log data includes the annotations made by the data annotator and the corresponding quality indicator data. Log data can be generated from logs on various data platforms, including but not limited to transaction data, maintenance data, user work data, and quality inspection data within the annotation platform system. Annotation content includes annotation time, annotation label, and annotation type. Quality indicator data includes various quality assessment indicator rules, the specific rules of which can be determined based on the type of annotation platform system.
[0046] Optionally, the present application also provides a system for analyzing and annotating data, which includes an annotation tool system, an annotation platform system, a message middleware, a business database, a distributed system cluster, a distributed file storage database, an analytical database, a message middleware, a stream computing module, and a front-end query service module. Figure 2 The system shown in FIG. Figure 1 Can be combined Figure 2 The system implements a method for analyzing and labeling data. The details can be demonstrated in the subsequent description.
[0047] In some embodiments of the present application, before parsing and verifying the log data stored in the subject queue to determine whether the log data is abnormal, Figure 1 The method shown also includes: obtaining a log report corresponding to the current task; preprocessing the data in the log report to obtain log data, wherein the preprocessing includes: cleaning, filtering, supplementing and packaging; and storing the log data in a topic queue.
[0048] In the above process, the present application pre-processes the log data to remove the influence of irrelevant data, providing more accurate data for subsequent log data analysis.
[0049] The current task represents the annotation task assigned by the data annotator to the annotation platform system. Preprocessing can also include deletion, addition, and modification. The topic queue can be an additional middleware that temporarily stores log messages published by the annotation platform system and delivers them to consumers subscribed to the topic.
[0050] Specifically, the annotation tool system generates workload log data during the daily work of annotators. The annotation tool system has a built-in scheduling mechanism. For example, at fixed intervals or upon task completion, it packages the accumulated annotation workload and quality-related metrics into a log report and the corresponding log data. The annotation platform system consolidates and preprocesses the received raw log data. First, it performs data cleansing to filter out records with illogical formats or missing content. Then, based on the task ID or user ID associated with the log, it retrieves supplementary information (such as task name, user group membership, etc.) from the platform's internal database and appends it to the log data. The annotation platform system encapsulates the processed log data into a message object according to an internal unified format, preparing it for subsequent publishing. Once log data consolidation is complete, the annotation platform system calls the message middleware's publishing interface to publish the log message to a predefined topic queue. The message contains necessary fields, such as the log ID, timestamp, task / user identifier, and the consolidated workload and quality data fragments. After successful publishing, the message middleware stores the message persistently for consumer processing.
[0051] In some embodiments of the present application, the log data stored in the subject queue is parsed and verified to determine whether the log data is abnormal, including: determining whether each field of the log data conforms to the expected format.
[0052] In the above process, the present application can accurately determine whether there is an anomaly in the log data by determining whether each field of the log data conforms to the expected format.
[0053] The expected format can be set according to requirements, including a character string in the expected format.
[0054] Specifically, when a new message arrives, Pulsar immediately pushes it to the message consumption module in the analysis center. Upon receiving the log message, the analysis center system parses and verifies it to ensure message integrity and that all fields conform to the expected format. If the message content contains an anomaly (such as a missing field or incorrect data type), the analysis center system can log the error and, based on policy, decide whether to discard the message or attempt to correct the data before continuing processing.
[0055] Step 120: When it is determined that the log data is abnormal, the log data, the current annotation task, and the structured data related to the target user are processed and merged according to a preset processing strategy to obtain data to be analyzed.
[0056] Structured data includes basic task information, basic information about the data annotator, and quality verification standards. Processing strategies include both discarding and correcting data. Specifically, when log data anomalies occur, error logs are recorded and, based on the strategy, the system decides whether to discard data or attempt to correct the data before continuing processing.
[0057] Optionally, the data to be analyzed may include not only log data and structured data, but also basic information related to the person who annotated the data. By merging the basic information related to the person who annotated the data, the annotated data of different people who annotated the data and the corresponding annotation evaluation and analysis results can be quickly analyzed.
[0058] In some embodiments of the present application, log data, the current annotation task, and structured data related to the target user are processed and merged according to a preset processing strategy, including: deleting or correcting log data and structured data according to the processing strategy; merging log data and structured data to obtain data to be analyzed.
[0059] In the above process, the present application can obtain the data to be analyzed by merging the above data, which facilitates the evaluation of subsequent labeled data.
[0060] Specifically: Querying the MySQL Database: According to the task plan, the analysis center system first establishes a connection to the MySQL database and executes a query to retrieve structured data related to the task or user referenced in the log. The query may include basic task information (project name, task type, etc.), annotator information (username, department, etc.), and task configuration (quality verification criteria, etc.). Pre-optimized SQL queries and appropriate indexes ensure that this step quickly returns the required data. The retrieved data is merged with the original log information in memory, providing a foundation for subsequent steps. Next, the analysis center system invokes the search interface provided by the Elasticsearch (distributed system) cluster to query the corresponding annotation result index and quality assessment data based on the identifying information in the log. For example, the task ID is used to retrieve statistical information such as quality scores and error flags for all annotated items in Elasticsearch for that task. Because Elasticsearch excels at full-text search and aggregate analysis, the analysis center system can execute pre-defined aggregate queries to obtain a summary of quality scores for the task or user. Once the query results are returned, they are also temporarily stored in an in-memory data structure, ready for merging with other data sources.
[0061] Step 130: Merge the key fields of the log data, the annotation result index, the quality assessment data, and the data to be analyzed to obtain merged data.
[0062] The annotation result index can be any symbolic symbol, through which the annotation result can be quickly obtained. The quality assessment data includes the custom set quality assessment rules and the specific annotation content in the corresponding annotation data.
[0063] In some embodiments of the present application, key fields are extracted by querying unstructured data and semi-structured data related to log data from a database.
[0064] In the above process, the present application extracts key fields by querying the unstructured data and semi-structured data related to the log data from the database, which can be used as an analysis basis for marking the key of the data.
[0065] In some embodiments of the present application, before merging the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed to obtain the merged data, Figure 1 The method shown also includes: storing the data to be analyzed and the key fields in the memory; querying the corresponding annotation result index and quality assessment data according to the identification information in the log report, and storing the annotation result index and quality assessment data in the memory.
[0066] In the above process, the present application stores the annotation result index, quality assessment data, analysis data and key fields in the memory, so that the data in the memory can be quickly obtained for data analysis when performing data analysis later.
[0067] Among them, the memory can be a temporary middleware. By storing various data in the memory, it can achieve the effect of quickly querying the data in the memory. At the same time, it can also modify and delete the data in the memory without affecting the original data in the analysis center.
[0068] Step 140: Based on the merged data, the labeled data completed by the data labeling personnel within a preset period are analyzed and evaluated to obtain analysis results.
[0069] Analyze and evaluate the labeled data completed by the data labeling personnel within the preset period. For example, summarize and calculate the total number of annotations completed by the labeling personnel corresponding to the log within the specified period, the average time taken for each annotation, and the comprehensive quality score (calculated in combination with the error rate and manual review results, etc.). If the log involves multiple tasks or multiple labelers, the system will calculate the indicators of each entity separately. In implementation, the analysis center system may use multi-threading or vectorized computing to accelerate processing, splitting the calculation of a large number of records into sub-tasks that can be executed in parallel, thereby making full use of the CPU and memory bandwidth to improve computing efficiency. After the calculation is completed, a structured analysis result data object is obtained, which contains the labeling task / personnel identification and the corresponding workload and quality indicator values. In some embodiments of the present application, after analyzing and evaluating the labeled data completed by the labeling personnel within a preset period based on the merged data and obtaining the analysis results, Figure 1 The method shown also includes: storing the analysis results in a preset database wide table, and updating the index analysis results.
[0070] In the above process, the present application uses the data index stored in the wide table of the database to quickly obtain the query content when subsequently querying the labeled data and analysis results.
[0071] Among them, the database wide table includes various annotation data, information about the person who annotated the data, analysis results, corresponding index fields and other data.
[0072] In the above Figure 1 In the process shown, the present application determines whether the log data is abnormal by parsing and verifying the log data stored in the subject queue, wherein the log data includes: the annotation content of the annotation data personnel and the corresponding quality indicator data; when it is determined that the log data is abnormal, according to the preset processing strategy, the log data and the current annotation task and the structured data related to the target user are processed and merged to obtain the data to be analyzed, wherein the structured data includes: basic information of the task, basic information of the annotation data personnel and quality verification standards, and the processing strategy includes a discarding strategy and a correction strategy; the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed are merged to obtain the merged data; based on the merged data, the annotation data completed by the annotation data personnel within the preset period are analyzed and evaluated to obtain the analysis results. By detecting anomalies in the log data, multi-source heterogeneous database information such as structured data, key fields, annotation result indexes and quality assessment data related to the log data can be obtained. By simultaneously analyzing and evaluating the sum of the above multiple information, the effect of accurately analyzing and evaluating the annotation data completed by the annotation data personnel can be achieved.
[0073] The following combination Figure 2 The system for analyzing and annotating data in an embodiment of the present application is described in detail.
[0074] Please see Figure 2 , Figure 2 A schematic block diagram of a system for analyzing and annotating data provided in an embodiment of the present application is shown in FIG. Figure 2 The system for analyzing labeled data shown includes: Annotation tool system 10, annotation platform system 11, message middleware 12 (Pulasar), business database 13 (MySQL), distributed system cluster 14 (Elasticsearch cluster), distributed file storage database 15 (MongoDB database), analytical database 16 (SelectDB), second message middleware 17 (Kafka), stream computing module 18 (Flink), and front-end query service module 19; The annotation tool system 10 is used to query the business database 13, the distributed system cluster 14 and the distributed file storage database 15 to generate log data of the annotation task process: The annotation platform system 11 is used to receive log data from the annotation tool system, parse and verify the log data stored in the subject queue, determine whether the log data is abnormal, and when it is determined that the log data is abnormal, process and merge the log data with the current annotation task and structured data related to the target user according to the preset processing strategy to obtain data to be analyzed, merge the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed to obtain merged data, and analyze and evaluate the annotation data completed by the annotation personnel within the preset period based on the merged data to obtain the analysis results: Message middleware 12 is used to asynchronously decouple and distribute data processing for various data in the annotation platform system: Analytical database 13, used to store and index analysis results through database wide tables; The second message middleware 17 is used to send various data to the stream computing module The stream computing module 18 is used to perform batch processing and data computing on various data.
[0075] The front-end query service module 19 is used to provide users with query and visualization functions of analysis results.
[0076] Specifically: The annotation tool system 10 is used to generate and collect workload log information of the annotation data. This module is usually composed of a specific annotation client or tool, which regularly summarizes the operation records of the annotation personnel, such as the number of annotations, time, quality assessment results, etc., to form log data through a pre-set time interval or event trigger mechanism. The annotation tool system 10 packages the log data into a structured message and sends it to the annotation platform system 11 through the network interface. In specific implementation, the annotation tool system 10 can call the API interface provided by the annotation platform system 11 to upload the recorded workload log in the form of an HTTP request or a remote procedure call, thereby ensuring that the log data can be delivered to the platform end in real time or quasi-real time.
[0077] The annotation platform system 11 is used to receive log information from the annotation tool system 10 and perform integration processing. The annotation platform system 11 merges, verifies and normalizes the format of workload logs from different sources or different time periods, and aggregates the scattered logs into a data structure that meets the requirements of subsequent processing. For example, the annotation platform system 11 can clean the log data internally, remove redundant or erroneous entries, add necessary task or personnel metadata, and generate the message objects required by the message queue in a predetermined format. When the integration is completed, the annotation platform system 11, as a message producer, pushes the processed workload information through the message middleware Pulsar12. The implementation method can be to call Pulsar's publishing interface to publish the message to the specified topic, including log content and related identifiers such as task record ID, personnel ID, etc., for subscription and consumption by the analysis center system.
[0078] The message middleware Pulsar12 plays the role of asynchronous decoupling and data distribution in this system. Pulsar adopts a publish-subscribe mechanism to temporarily store the log messages published by the annotation platform system 11 and pass them to consumers who subscribe to the topic. In implementation, the analysis center system creates a corresponding subscriber consumer on Pulsar12 in advance to continuously monitor the message queue of a specific topic. When a new annotated log message arrives at Pulsar12, the analysis center system is immediately notified and obtains the data content of the message. By introducing the message middleware 12, the annotation platform and the analysis center are decoupled in processing: the platform can publish messages at its own system processing rhythm, and the analysis center system automatically receives them asynchronously, ensuring that high-throughput log data is transmitted reliably and without loss.
[0079] After receiving the log data from the message middleware Pulsar12, the analysis center system automatically connects to multiple business databases to perform data aggregation and calculation. Specifically, the processing module included in the analysis center system will establish connections with the MySQL database 13, Elasticsearch cluster 14 and MongoDB database 15 on which the annotation platform relies, and query relevant data based on the information contained in the log. The MySQL database 13 usually stores basic structured information of the annotation task, such as project information, user permission configuration, task allocation, etc.; the Elasticsearch cluster 14 usually stores index data related to the annotation content, which can be used to retrieve text or image annotation results and quality indicators; the MongoDB database 15 is usually used to store semi-structured or unstructured annotation data and intermediate results, such as detailed annotation records, comment information, etc. The analysis center system queries MySQL by calling the corresponding data interface MybatisPlus, queries the index using the Elasticsearch REST API, or queries through the MongoDB driver to obtain the necessary information in these databases for enriching and verifying the workload log data. In practice, the analysis center system serially schedules these queries in a predetermined order: first, it queries MySQL13 to obtain basic task information, then queries Elasticsearch14 to retrieve relevant quality scores, and finally, it retrieves detailed annotations or historical records from MongoDB15. The information returned by each data source is then correlated and integrated by the analysis center's internal computing module, ensuring that data from different sources is correctly linked based on key fields such as task record ID or user ID, forming a complete data view.
[0080] After completing multi-source data query and retrieval, the analysis center system performs statistical analysis and quality assessment calculations on the aggregated data to generate the final analysis results. This calculation process includes: accumulating metrics such as the number of annotation tasks and total time taken based on workload data in the logs, combining it with quality metrics such as accuracy and error rate obtained from various databases for comprehensive analysis, and then calculating workload and quality scores at various granularities, such as by user, team, and workflow, according to predefined statistical models or business rules. To improve computational efficiency and accuracy, the analysis center system can employ in-memory caching and batch processing techniques, loading relevant data into memory for processing, or utilizing a distributed computing framework for parallel processing of large batches of logs. The processed analysis results are generated as structured records, each of which contains key fields for workload and quality analysis at a specified granularity, such as annotator ID, number of completed tasks, average time taken, and quality score.
[0081] The analysis center system writes the result data obtained from the above calculations into the wide table of the analytical database SelectDB16. SelectDB16 is a high-performance database for real-time analysis, suitable for storing and querying large-scale structured analytical data. A wide table is a large table structure containing multiple dimensions and multiple indicators, which is used in this system to store the integrated analysis result data. In terms of specific implementation, the analysis center system sends data to the streaming computing module Flink18 through the message middleware Kafka17. After batch processing, the streaming computing module Flink18 calls the SelectDB16 batch processing interface to perform batch import, and appends the new data records obtained from each analysis to the wide table. The design of this wide table covers fields from different data sources, and uses technologies such as redundancy elimination and column storage to improve query performance. After the data is written, SelectDB16 will optimize the index and partition of the data to support subsequent efficient query analysis. By using the analytical database 16, this system can still maintain low query latency and fast response results even when the data volume is huge.
[0082] The front-end query service module 19 provides users with data query and visualization capabilities. This module is directly connected to SelectDB 16 and is responsible for presenting analysis results to end users. The front-end query service 19 retrieves the required data from SelectDB 16 by constructing efficient query statements, such as multi-dimensional aggregation queries, and retrieves a portion of the data in a large result set through paging queries, ensuring timely front-end responses and smooth interface loading. This module also features summary statistics and chart analysis: query results are further grouped and aggregated as needed at the application layer, statistical indicators such as totals and averages are calculated, and the results are rendered into intuitive formats such as line charts and bar charts using a chart library for user presentation. Because SelectDB 16 provides fast OLAP query capabilities, the front-end service 19 is able to complete interactive queries with high performance, achieving sub-second query responses even when faced with massive amounts of analytical data. When a user submits a query request through the front-end interface, the front-end service 19 calls the analysis center system, converting the request into an SQL query targeting a wide table. SelectDB executes the query and returns the result data. The front-end then beautifies and visualizes the results, ultimately presenting the results of the annotation workload and quality analysis intuitively.
[0083] also, Figure 2 The specific methods and steps performed by the system shown can be found in Figure 1 The method shown here will not be described in detail.
[0084] Previous article passed Figure 1 , combined with Figure 3-Figure 4 Describe the apparatus for analyzing labeled data.
[0085] Please refer to Figure 3 , is a schematic block diagram of a device 300 for analyzing and annotating data provided in an embodiment of the present application. The device 300 may be a module, program segment or code on an electronic device. The device 300 is similar to the above Figure 1 The method embodiment corresponds to the embodiment that can be executed Figure 1 The various steps involved in the method embodiment and the specific functions of the device 300 can be found in the description below. To avoid repetition, detailed description is appropriately omitted here.
[0086] Optionally, the device 300 includes: The parsing and verification module 310 is used to parse and verify the log data stored in the subject queue to determine whether the log data is abnormal, wherein the log data includes: the annotation content of the data annotator and the corresponding quality indicator data; Processing module 320 is configured to, upon determining that the log data is abnormal, process and merge the log data with the current annotation task and structured data related to the target user according to a preset processing strategy to obtain data to be analyzed, wherein the structured data includes basic task information, basic information of the person annotating the data, and quality verification standards, and the processing strategy includes a discarding strategy and a correction strategy; A merging module 330 is configured to merge the key fields of the log data, the annotation result index, the quality assessment data, and the data to be analyzed to obtain merged data; The analysis module 340 is used to analyze and evaluate the annotated data completed by the annotator within a preset period based on the merged data to obtain an analysis result.
[0087] Optionally, the device further includes: The preprocessing module is used for the parsing and verification module to parse and verify the log data stored in the subject queue to determine whether the log data is abnormal, obtain the log report corresponding to the current task; preprocess the data in the log report to obtain log data, wherein the preprocessing includes: cleaning, filtering, supplementing and packaging; and storing the log data in the subject queue.
[0088] Optionally, key fields are extracted by querying unstructured data and semi-structured data related to the log data from a database.
[0089] Optionally, the device further includes: A storage module is used for the merging module to store the data to be analyzed and the key fields in the memory before merging the key fields, annotation result index, quality assessment data and data to be analyzed of the log data to obtain the merged data; query the corresponding annotation result index and quality assessment data according to the identification information in the log report, and store the annotation result index and quality assessment data in the memory.
[0090] Optionally, the device further includes: The index module is used for the analysis module to analyze and evaluate the labeled data completed by the labeling personnel within a preset period based on the merged data, and after obtaining the analysis results, store the analysis results in a preset database wide table and update the index analysis results.
[0091] Optionally, the parsing and verification module is specifically used to: Determine whether the fields of the log data conform to the expected format.
[0092] Optionally, the analysis module is specifically used to: According to the processing strategy, the log data and structured data are deleted or modified; the log data and structured data are merged to obtain the data to be analyzed.
[0093] Please refer to Figure 4 This is a schematic block diagram of a device for analyzing and annotating data provided in an embodiment of the present application. The device may include a memory 410 and a processor 420. Optionally, the device may also include: a communication interface 430 and a communication bus 440. The device is similar to the above Figure 1 The method embodiment corresponds to the embodiment that can be executed Figure 1 The various steps involved in the method embodiment and the specific functions of the device can be found in the description below.
[0094] Specifically, the memory 410 is used to store computer-readable instructions.
[0095] Processor 420 is used to process the readable instructions stored in the memory and can execute Figure 1 The steps in the method.
[0096] The communication interface 430 is used for signaling or data communication with other node devices, for example, for communication with a server or terminal, or for communication with other device nodes, but the embodiments of the present application are not limited thereto.
[0097] The communication bus 440 is used to realize direct connection and communication among the above components.
[0098] Among them, the communication interface 430 of the device in the embodiment of the present application is used to communicate signaling or data with other node devices. The memory 410 can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 410 can also be at least one storage device located away from the aforementioned processor. The memory 410 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 420, the electronic device executes the above-mentioned Figure 1 The method process shown. The processor 420 can be used on the device 300 and is used to perform the functions of the present application. For example, the above-mentioned processor 420 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, but the embodiments of the present application are not limited thereto.
[0099] The embodiment of the present application further provides a readable storage medium, wherein when the computer program is executed by a processor, Figure 1 The method process in the illustrated method embodiment is performed by the electronic device.
[0100] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.
[0101] In summary, the embodiments of the present application provide a method, system and device for analyzing and annotating data, the method comprising: parsing and verifying the log data stored in the subject queue to determine whether the log data is abnormal, wherein the log data comprises: the annotation content of the annotating data personnel and the corresponding quality indicator data; when it is determined that the log data is abnormal, according to a preset processing strategy, the log data and the current annotation task and the structured data related to the target user are processed and merged to obtain the data to be analyzed, wherein the structured data comprises: basic information of the task, basic information of the annotating data personnel and quality verification standards, and the processing strategy comprises a discarding strategy and a correction strategy; the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed are merged to obtain the merged data; based on the merged data, the annotation data completed by the annotating data personnel within the preset period are analyzed and evaluated to obtain the analysis results. This method can achieve the effect of accurately analyzing and evaluating the annotation data completed by the annotating data personnel.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0103] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0104] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0105] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0106] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
[0107] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
Claims
1. A method for analyzing labeled data, characterized in that: Applications in analytical centers include: Parsing and verifying the log data stored in the subject queue to determine whether the log data is abnormal, wherein the log data includes: the annotation content of the data annotator and the corresponding quality indicator data; When it is determined that the log data is abnormal, the log data, the current annotation task, and the structured data related to the target user are processed and merged according to a preset processing strategy to obtain data to be analyzed, wherein the structured data includes: basic information of the task, basic information of the person who annotated the data, and quality verification standards; the processing strategy includes a discarding strategy and a correction strategy; Merging the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed to obtain merged data; Based on the merged data, the annotated data completed by the annotating personnel within a preset period is analyzed and evaluated to obtain an analysis result.
2. The method according to claim 1, characterized in that Before parsing and verifying the log data stored in the subject queue to determine whether the log data is abnormal, the method further includes: Obtaining a log report corresponding to the current annotation task; Preprocessing the data in the log report to obtain the log data, wherein the preprocessing includes: cleaning, filtering, supplementing and packaging; The log data is stored in the topic queue.
3. The method according to claim 1 or 2, characterized in that The key fields are extracted by querying unstructured data and semi-structured data related to the log data from a database.
4. The method according to claim 1 or 2, characterized in that Before merging the key fields of the log data, the annotation result index, the quality assessment data, and the data to be analyzed to obtain the merged data, the method further includes: Storing the data to be analyzed and the key fields in a memory; The corresponding annotation result index and the quality assessment data are searched according to the identification information in the log report, and the annotation result index and the quality assessment data are stored in the memory.
5. The method according to claim 1 or 2, characterized in that After analyzing and evaluating the labeled data completed by the data labeling personnel within a preset period based on the merged data to obtain the analysis results, the method further includes: The analysis result is stored in a preset database wide table, and the index of the analysis result is updated.
6. The method according to claim 1 or 2, characterized in that The parsing and verifying of the log data stored in the subject queue to determine whether the log data is abnormal includes: Determine whether each field of the log data conforms to an expected format.
7. The method according to claim 1 or 2, characterized in that The processing and merging of the log data, the current annotation task, and the structured data related to the target user according to a preset processing strategy includes: Deleting or modifying the log data and the structured data according to the processing policy; The log data and the structured data are merged to obtain the data to be analyzed.
8. The method according to claim 1 or 2, characterized in that The step of analyzing and evaluating the labeled data completed by the data labeling personnel within a preset period based on the merged data to obtain analysis results includes: Summarize the task data of the data annotators to obtain the target task, wherein the task data includes: the total number of annotations completed by each data annotator within a specified period, the average time taken for each annotation, and the comprehensive quality score; Using multi-threading or vectorized computing acceleration processing methods to split the target task into multiple subtasks; The plurality of subtasks are analyzed respectively by using a central processing unit and memory bandwidth to obtain the analysis results, wherein the analysis results include: an identifier of a person who annotates the data and corresponding values of various workloads and quality indicators.
9. A system for analyzing labeled data, characterized in that: include: Annotation tool system, annotation platform system, message middleware, business database, distributed system cluster, distributed file storage database, analytical database, second message middleware, streaming computing module and front-end query service module; The annotation tool system is used to query the business database, the distributed system cluster, and the distributed file storage database to generate log data of the annotation task process: The annotation platform system is used to receive the log data of the annotation tool system, parse and verify the log data stored in the subject queue, determine whether the log data is abnormal, and when it is determined that the log data is abnormal, process and merge the log data with the current annotation task and structured data related to the target user according to a preset processing strategy to obtain data to be analyzed, merge the key fields, annotation result index, quality assessment data of the log data and the data to be analyzed to obtain merged data, and analyze and evaluate the annotation data completed by the annotation personnel within a preset period based on the merged data to obtain analysis results: The message middleware is used for asynchronous decoupling and distributed data processing of various data in the annotation platform system: The analytical database is used to store and index the analysis results through a database wide table; The second message middleware is used to send the various data to the stream computing module The stream computing module is used to perform batch processing and data computing on the various data; The front-end query service module is used to provide users with query and visual display functions of the analysis results.
10. A device for analyzing labeled data, characterized in that: include: A parsing and verification module is used to parse and verify the log data stored in the subject queue to determine whether the log data is abnormal, wherein the log data includes: the annotation content of the data annotator and the corresponding quality indicator data; a processing module configured to, upon determining that the log data is abnormal, process and merge the log data, the current annotation task, and structured data related to the target user according to a preset processing strategy to obtain data to be analyzed, wherein the structured data includes: basic information of the task, basic information of the person annotating the data, and quality verification standards; and the processing strategy includes a discarding strategy and a correction strategy; A merging module, configured to merge the key fields of the log data, the annotation result index, the quality assessment data and the data to be analyzed to obtain merged data; The analysis module is used to analyze and evaluate the labeled data completed by the data labeling personnel within a preset period based on the merged data to obtain analysis results.
Citation Information
Patent Citations
Artificial annotation quality evaluation method and device, electronic equipment and storage medium
CN111291567A
Data annotation quality evaluation and improvement system and method
CN117762912A
Deep concentration production scheduling big data analysis system and method
CN120011373A
Data annotation method and system, and electronic device
WO2022063274A1
Low-cost zero-shot online log parsing method based on large language model
WO2025077116A1