Standard annotation data checking and quality evaluation method and system based on multi-dimensional feature extraction

By using a multi-dimensional feature extraction method to comprehensively analyze structured and unstructured data, dynamic evolution features of labeled data are constructed, which solves the problems of misjudgment and omission in the labeling quality assessment of existing technologies and achieves efficient and reliable labeling quality assessment.

CN121935508APending Publication Date: 2026-04-28ZHEJIANG INSTITUTE OF QUALITY SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG INSTITUTE OF QUALITY SCIENCES
Filing Date
2025-11-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods for evaluating the quality of labeled data mainly rely on manual review or rule-based automated checks, ignoring the dynamic evolutionary characteristics during the labeling process, such as the behavior patterns of labelers, semantic changes in labeled content, and version evolution trajectories. This makes it difficult to comprehensively and accurately evaluate the labeling quality, especially in complex labeling tasks or large-scale labeled data, where misjudgments or omissions are prone to occur.

Method used

A multi-dimensional feature extraction method is adopted to extract features from structured and unstructured labeled data, obtain the frequency of flipped labels, structural drift, labeling behavior sequence and contextual semantic association sequence, construct the labeling evolution trajectory, comprehensively calculate the labeling consistency feature value and dynamic evolution feature value, and finally generate a quality assessment score.

Benefits of technology

It enables accurate and reliable quality assessment of labeled data, avoids misjudgment or omission, and improves the accuracy and reliability of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935508A_ABST
    Figure CN121935508A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data annotation, and particularly discloses a standard annotation data checking and quality evaluation method and system based on multi-dimensional feature extraction. According to the method, through multi-dimensional feature extraction and comprehensive analysis by integrating structured and unstructured data, firstly, node division is performed on structured annotation data, an annotation consistency feature value is calculated based on an overturning annotation frequency and a structure drift degree, and meanwhile, an annotation behavior sequence is extracted; generating a check correction backtracking sequence and a context semantic association sequence, further combining the consistency feature and the semantic sequence to construct an annotation evolution track, analyzing the correction backtracking and evolution track to obtain a dynamic evolution feature value, and finally fusing the consistency and dynamic evolution features to generate a quality evaluation score. Accurate and reliable marking quality judgment is achieved, and meanwhile the problem of misjudgment or missed judgment can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data annotation technology, and in particular to a method and system for verifying and evaluating the quality of standard annotated data based on multi-dimensional feature extraction. Background Technology

[0002] In the fields of artificial intelligence, machine learning, and data science, standardized labeled data is fundamental for training high-quality models. The quality of labeled data directly impacts the performance and reliability of the model. Existing methods for assessing the quality of labeled data primarily rely on manual review or rule-based automated checks; However, existing methods typically focus only on the static consistency of annotation results, neglecting the dynamic evolutionary characteristics during the annotation process, such as annotator behavior patterns, semantic changes in the annotated content, and version evolution trajectories. This makes it difficult to comprehensively and accurately evaluate annotation quality when faced with complex annotation tasks or large-scale annotated data, easily leading to misjudgments or omissions. Therefore, a standard annotation data verification and quality assessment method based on multi-dimensional feature extraction is needed to address these issues. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for verifying and evaluating the quality of standard labeled data based on multi-dimensional feature extraction, so as to solve the technical problems mentioned in the background art.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a method for verifying and evaluating the quality of standard labeled data based on multi-dimensional feature extraction, comprising: Obtain the data to be labeled, and extract features from the data to be labeled based on multi-dimensional features to obtain the standard labeled data to be verified, wherein the standard labeled data includes structured labeled data and unstructured labeled data; The structured annotation data is divided into nodes to obtain multiple standard annotation node data. The annotation flipping frequency and annotation structure drift are obtained based on the multiple standard annotation node data. The annotation consistency feature value is obtained based on the annotation flipping frequency and annotation structure drift. Annotation behavior sequence is obtained based on the unstructured annotation data, and verification and correction backtracking sequence and context semantic association sequence are obtained based on the annotation behavior sequence; The annotation evolution trajectory sequence is obtained based on the annotation consistency feature value and the context semantic association sequence, and the annotation dynamic evolution feature value is obtained based on the verification and correction backtracking sequence and the annotation evolution trajectory sequence; The annotation quality assessment score is obtained based on the annotation consistency feature value and the annotation dynamic evolution feature value; The standard annotation data is evaluated based on the annotation quality assessment score to obtain the evaluation result.

[0005] Preferably, the step of dividing the structured annotation data into nodes to obtain multiple standard annotation node data, obtaining the annotation flipping frequency and annotation structure drift degree based on the multiple standard annotation node data, and obtaining the annotation consistency feature value based on the annotation flipping frequency and annotation structure drift degree includes: Annotation logic units are obtained based on the structured annotation data, wherein the annotation logic unit includes a start annotation unit, an annotation attribute unit, and an annotation end unit; The structured annotation data is divided according to the annotation logic unit to obtain multiple standard annotation node data; Obtain multiple timestamp sequences and multiple annotation content change records for each standard annotation node data, and obtain the number of flips for multiple annotation labels based on the multiple annotation content change records; The number of annotation operations within a preset time window is obtained based on the multiple timestamp sequences. Multiple first flip label frequencies are obtained based on the number of times the multiple labeling operations are performed and the number of times the multiple labeling tags are flipped, and an average flip label frequency is obtained based on the multiple first flip label frequencies, and the average flip label frequency is used as the flip label frequency. Obtain the graph structure features of each standard labeled node data, and obtain multiple graph structure feature difference values ​​within a preset adjacent time window based on multiple graph structure features, and obtain the average graph structure feature difference value of the multiple graph structure feature difference values; The difference ratio is calculated based on the average difference value of the structural features of the graph and the difference value of the preset structural features of the graph, and the difference ratio is used as the drift degree of the labeled structure. The connection coefficient in the annotation process is obtained based on the starting annotation unit, the annotation attribute unit, and the annotation ending unit; The annotation consistency feature value is obtained by weighted fusion based on the flipping annotation frequency, the annotation structure drift degree, and the connection coefficient.

[0006] Preferably, the steps of obtaining the annotation behavior sequence based on the unstructured annotation data, and obtaining the verification and correction backtracking sequence and the context semantic association sequence based on the annotation behavior sequence, include: The storage path structure and annotation task association structure of the annotation data are obtained based on the unstructured annotation data, and the annotation organization topology is constructed based on the storage path structure and annotation task association structure. Obtain the corresponding operation logs based on the unstructured labeled data, and generate a labeling behavior sequence based on the operation logs, the labeling organization topology, and preset operation time points; Multiple annotation connection nodes are obtained based on the annotation behavior sequence; The number of verification and correction backtracking steps for each of the multiple labeled connection nodes is obtained, and a verification and correction backtracking sequence is generated based on the multiple verification and correction backtracking steps and a preset time window. Obtain corresponding node semantic data based on multiple labeled connection nodes, and perform semantic vectorization processing on the multiple node semantic data to obtain multiple node semantic vectors; The semantic vectors of the multiple nodes are associated according to the labeled time order to obtain the context semantic association sequence.

[0007] Preferably, the step of obtaining the annotation evolution trajectory sequence based on the annotation consistency feature value and the context semantic association sequence, and obtaining the annotation dynamic evolution feature value based on the verification and correction backtracking sequence and the annotation evolution trajectory sequence, includes: Extract semantic change vectors based on the contextual semantic association sequence; Based on the temporal alignment algorithm, the historical sequence of the annotation consistency feature values ​​and the semantic change vector are used to generate an annotation evolution trajectory sequence, wherein the evolution trajectory includes version switching points and semantic leap degrees; The correction intensity distribution is obtained based on the verification and correction backtracking sequence, and the inter-version correction density is calculated based on the correction intensity distribution and the version switching point; The evolutionary activity is obtained by weighted fusion of the semantic transition degree and the modified density, and the evolutionary activity is subjected to stability analysis within a preset time to obtain multiple fluctuation characteristic values. The fluctuation characteristic coefficient is calculated based on multiple fluctuation characteristic values, and the fluctuation characteristic coefficient is used as the labeling dynamic evolution characteristic value.

[0008] Preferably, the step of obtaining the annotation quality assessment score based on the annotation consistency feature value and the annotation dynamic evolution feature value includes: The first weight value is obtained based on the annotation consistency feature value; The dynamic-static synergy factor is obtained based on the annotation consistency feature value and the annotation dynamic evolution feature value; The annotation quality assessment score is obtained based on the annotation consistency feature value, the first weight value, the dynamic-static synergy factor, and the annotation dynamic evolution feature value, wherein the calculation formula is:

[0009] in, This indicates the quality assessment score. This indicates the consistency feature value of the annotation. This represents the first weight value. This indicates the annotation of dynamically evolving feature values. This represents the dynamic-static synergy factor.

[0010] Preferably, the step of evaluating the quality of the standard annotation data based on the annotation quality assessment score to obtain the evaluation result includes: Determine whether the annotation quality assessment score is greater than a preset threshold; If the quality assessment score of the annotation is not greater than the preset threshold, the standard annotation data is determined to be qualified annotation data; If the quality assessment score of the annotation is greater than the preset threshold, the standard annotation data is determined to be unqualified annotation data.

[0011] This application also provides a standard labeled data verification and quality assessment system based on multi-dimensional feature extraction, including: The first acquisition module is used to acquire the data to be labeled and to extract features from the data to be labeled based on multi-dimensional features to obtain the standard labeled data to be verified, wherein the standard labeled data includes structured labeled data and unstructured labeled data. The partitioning module is used to partition nodes according to the structured annotation data to obtain multiple standard annotation node data, and to obtain the flipping annotation frequency and annotation structure drift degree according to the multiple standard annotation node data, and to obtain the annotation consistency feature value according to the flipping annotation frequency and annotation structure drift degree; The second acquisition module is used to acquire an annotation behavior sequence based on the unstructured annotation data, and to acquire a verification and correction backtracking sequence and a context semantic association sequence based on the annotation behavior sequence. The third acquisition module is used to acquire the annotation evolution trajectory sequence based on the annotation consistency feature value and the context semantic association sequence, and to acquire the annotation dynamic evolution feature value based on the verification correction backtracking sequence and the annotation evolution trajectory sequence; The fourth acquisition module is used to obtain the annotation quality evaluation score based on the annotation consistency feature value and the annotation dynamic evolution feature value; The evaluation module is used to evaluate the quality of the standard annotation data based on the annotation quality evaluation score, and obtain the evaluation result.

[0012] Preferably, the second acquisition module includes: The first acquisition unit is used for the steps of dividing the structured annotation data into nodes to obtain multiple standard annotation node data, obtaining the annotation flipping frequency and annotation structure drift degree based on the multiple standard annotation node data, and obtaining annotation consistency feature values ​​based on the annotation flipping frequency and annotation structure drift degree, including: The second acquisition unit is used to acquire an annotation logic unit based on the structured annotation data, wherein the annotation logic unit includes a start annotation unit, an annotation attribute unit, and an annotation end unit; A partitioning unit is used to partition the structured annotation data according to the annotation logic unit to obtain multiple standard annotation node data; The third acquisition unit is used to acquire multiple timestamp sequences and multiple annotation content change records for each standard annotation node data, and to acquire the number of flips of multiple annotation labels based on the multiple annotation content change records; The fourth acquisition unit is used to acquire the number of annotation operations within a preset time window based on the multiple timestamp sequences; The fifth acquisition unit is used to acquire multiple first flip label frequencies based on the multiple number of labeling operations and the multiple number of flips of the label, and to acquire an average flip label frequency based on the multiple first flip label frequencies, and to use the average flip label frequency as the flip label frequency. The sixth acquisition unit is used to acquire the graph structure features of each standard labeled node data, and to acquire multiple graph structure feature difference values ​​within a preset adjacent time window based on the multiple graph structure features shown, and to acquire the average graph structure feature difference value of the multiple graph structure feature difference values. The calculation unit is used to calculate the difference ratio based on the average difference value of the map structure features and the preset difference value of the map structure features, and to use the difference ratio as the drift degree of the labeled structure. The seventh acquisition unit is used to acquire the connection coefficient in the annotation process based on the starting annotation unit, the annotation attribute unit and the annotation end unit; The fusion unit is used to obtain a label consistency feature value by weighted fusion based on the flip label frequency, the label structure drift degree and the connection coefficient.

[0013] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0014] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0015] The beneficial effects of this application are as follows: This invention extracts features from multiple dimensions and conducts a comprehensive analysis of both structured and unstructured data. First, the structured labeled data is divided into nodes. Based on the frequency of label flipping and the degree of structural drift, the label consistency feature value is calculated. At the same time, the labeling behavior sequence is extracted to generate a verification and correction backtracking sequence and a contextual semantic association sequence. Furthermore, the consistency feature and the semantic sequence are combined to construct the labeling evolution trajectory. By analyzing the correction backtracking and evolution trajectory, the dynamic evolution feature value is obtained. Finally, the consistency and dynamic evolution features are integrated to generate a quality assessment score, achieving accurate and reliable labeling quality judgment, while also avoiding the problems of misjudgment or omission. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of a method flow according to an embodiment of this application.

[0017] Figure 2 This is a schematic diagram of the system structure according to an embodiment of this application.

[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0020] like Figures 1-2 As shown, this application provides a method for verifying and evaluating the quality of standard labeled data based on multi-dimensional feature extraction, including: S1. Obtain the data to be labeled, and extract features from the data to be labeled based on multi-dimensional features to obtain the standard labeled data to be verified, wherein the standard labeled data includes structured labeled data and unstructured labeled data; S2. Nodes are divided according to the structured annotation data to obtain multiple standard annotation node data, and the annotation flipping frequency and annotation structure drift degree are obtained according to the multiple standard annotation node data. Annotation consistency feature value is obtained according to the annotation flipping frequency and annotation structure drift degree. S3. Obtain the annotation behavior sequence based on the unstructured annotation data, and obtain the verification and correction backtracking sequence and the context semantic association sequence based on the annotation behavior sequence; S4. Obtain the annotation evolution trajectory sequence based on the annotation consistency feature value and the context semantic association sequence, and obtain the annotation dynamic evolution feature value based on the verification correction backtracking sequence and the annotation evolution trajectory sequence; S5. Obtain the annotation quality assessment score based on the annotation consistency feature value and the annotation dynamic evolution feature value; S6. Evaluate the quality of the standard annotation data based on the annotation quality assessment score to obtain the evaluation result.

[0021] As described in steps S1-S6 above, existing technologies rely on manual review or simple rule checks, which cannot comprehensively evaluate the quality of labeled data. Therefore, this invention first acquires the data to be labeled and extracts features from the data based on multi-dimensional features to obtain standard labeled data to be verified. The standard labeled data includes structured and unstructured labeled data, and in text labeling tasks, it can be the original text and preliminary entity labeling results. Multi-dimensional feature extraction processes the data from multiple dimensions such as data structure, content attributes, and semantic information. This allows for a comprehensive capture of the features of the labeled data, avoiding the limitations of a single perspective. Structured labeled data includes the hierarchical structure and attributes of the labels, while unstructured labeled data includes labeling behavior logs and semantic information, providing rich data sources for subsequent analysis. Then, the structured labeled data is divided into nodes to obtain multiple standard labeled node data. The label flipping frequency and label structure drift are obtained from these standard labeled node data, and label consistency feature values ​​are obtained based on the label flipping frequency and label structure drift. This node division decomposes complex labeled data into manageable units, facilitating analysis. The frequency of label flipping reflects the stability of label levels, while label structure drift reflects the degree of change in label structure. Combining these two metrics can quantify label consistency and identify inconsistencies in the labeling process. Next, a labeling behavior sequence is obtained from the unstructured labeling data, and a verification and correction backtracking sequence and a contextual semantic association sequence are derived from this sequence. The labeling behavior sequence captures the operational patterns of labelers, the verification and correction backtracking sequence records the history of label corrections, and the contextual semantic association sequence captures semantic changes, providing a foundation for dynamic analysis. Then, a labeling evolution trajectory sequence is obtained from the labeling consistency feature value and the contextual semantic association sequence, and a labeling dynamic evolution feature value is obtained from the verification and correction backtracking sequence and the labeling evolution trajectory sequence. By analyzing the evolution trajectory sequence and the correction backtracking sequence, the dynamic evolution behavior of the labeling data is analyzed, identifying trends and anomalies in labeling quality. The dynamic evolution feature value quantifies the activity and stability of the labeling data over time. Finally, the annotation quality assessment score is obtained based on the annotation consistency feature value and the annotation dynamic evolution feature value, and the standard annotation data is evaluated based on the annotation quality assessment score to obtain the evaluation result. In this way, by combining consistency and dynamic evolution features, a comprehensive quality assessment score is generated, thereby accurately determining whether the annotation data is qualified and improving the accuracy and reliability of the evaluation.

[0022] In one embodiment, step S2, which involves dividing the structured annotation data into nodes to obtain multiple standard annotation node data, obtaining the annotation flipping frequency and annotation structure drift based on the multiple standard annotation node data, and obtaining annotation consistency feature values ​​based on the annotation flipping frequency and annotation structure drift, includes: S201. Obtain annotation logic units based on the structured annotation data, wherein the annotation logic units include a start annotation unit, an annotation attribute unit, and an annotation end unit; S202. Divide the structured annotation data according to the annotation logic unit to obtain multiple standard annotation node data; S203. Obtain multiple timestamp sequences and multiple annotation content change records for each standard annotation node data, and obtain the number of flips of multiple annotation labels based on the multiple annotation content change records; S204. Obtain the number of multiple annotation operations within a preset time window based on the multiple timestamp sequences; S205. Obtain multiple first flip label frequencies based on the multiple number of labeling operations and the multiple number of flips of the label labels, and obtain an average flip label frequency based on the multiple first flip label frequencies, and use the average flip label frequency as the flip label frequency. S206. Obtain the graph structure features of each standard labeled node data, and obtain multiple graph structure feature difference values ​​within a preset adjacent time window based on the multiple graph structure features shown, and obtain the average graph structure feature difference value of the multiple graph structure feature difference values. S207. Calculate the difference ratio based on the average map structural feature difference value and the preset map structural feature difference value, and use the difference ratio as the labeled structural drift degree; S208. Obtain the connection coefficient in the annotation process based on the starting annotation unit, the annotation attribute unit, and the annotation end unit; S209. Obtain the annotation consistency feature value by weighted fusion based on the flipping annotation frequency, the annotation structure drift degree, and the connection coefficient.

[0023] As described in steps S201-S209 above, this invention first obtains annotation logic units based on the structured annotation data. These annotation logic units include a start annotation unit, an annotation attribute unit, and an annotation end unit. By defining these annotation logic units, the structure of the annotation data is standardized, ensuring consistency in node partitioning. The start annotation unit identifies the beginning of annotation, the annotation attribute unit contains annotation attributes, and the annotation end unit identifies the end of annotation, forming a complete annotation logic flow. Structured annotation data typically possesses clear logical connections and structural features. For example, image label annotation data in tabular form includes fixed fields such as annotation ID (start identifier), label category (attribute information), and annotation completion time (end identifier). The core quality of this type of data lies in the completeness of the logical connections between fields and the stability of the annotation results. If the structured annotation data has problems such as missing logic units, frequent changes in annotation results, or deviations from the standard structure, it will directly lead to disordered data input during subsequent model training, affecting the model's accurate learning of features. Then, the structured annotation data is partitioned according to the annotation logic units to obtain multiple standard annotation node data. This node partitioning decomposes large-scale annotation data into smaller units, facilitating refined analysis of the annotation behavior of each node. Next, multiple timestamp sequences and multiple annotation content change records are obtained for each standard annotation node data. Based on these multiple annotation content change records, the number of flips for multiple annotation labels is obtained. The timestamp sequences record the time points of annotation operations, the annotation content change records the history of label changes, and the number of flips quantifies the frequency of label changes, reflecting the stability of the annotator's decisions. Then, based on the multiple timestamp sequences, the number of annotation operations within a preset time window is obtained. This preset time window focuses on annotation activities within a specific time period, and the number of operations reflects the activity level of annotation. Next, multiple first flip annotation frequencies are obtained based on the multiple number of annotation operations and the multiple number of label flips. An average flip annotation frequency is then obtained based on these first flip annotation frequencies. This average flip annotation frequency is used as the overall flip annotation frequency. Thus, the flip annotation frequency is the number of label flips per unit of operation, and the average flip annotation frequency smooths out fluctuations across multiple nodes, providing an overall flipping trend. Then, the graph structure features of each standard labeled node data are obtained, and multiple graph structure feature difference values ​​within a preset adjacent time window are obtained based on multiple graph structure features. The average graph structure feature difference value of the multiple graph structure feature difference values ​​is then obtained. In this way, the graph structure features capture the topological structure of the labeled data, the difference values ​​quantify structural changes, and the average difference value reflects the overall degree of structural drift. Next, the difference ratio is calculated based on the average graph structure feature difference value and the preset graph structure feature difference value, and the difference ratio is used as the labeled structural drift degree. In this way, the difference ratio standardizes the structural drift and makes it comparable to a preset threshold for easy evaluation.Then, the connection coefficient in the annotation process is obtained based on the starting annotation unit, the annotation attribute unit, and the annotation ending unit. This connection coefficient measures the tightness of the connection between annotation logic units and reflects the smoothness of the annotation process. Finally, the annotation consistency feature value is obtained by weighted fusion based on the annotation flipping frequency, the annotation structure drift degree, and the connection coefficient. This weighted fusion comprehensively considers multiple factors to generate a consistency feature value, where the weights can be adjusted according to actual applications to highlight key factors.

[0024] In one embodiment, step S3, which involves obtaining a sequence of annotation behaviors based on the unstructured annotation data and obtaining a verification and correction backtracking sequence and a contextual semantic association sequence based on the annotation behavior sequence, includes: S301. Obtain the storage path structure and annotation task association structure of the annotation data based on the unstructured annotation data, and construct the annotation organization topology based on the storage path structure and annotation task association structure; S302. Obtain the corresponding operation log based on the unstructured annotation data, and generate an annotation behavior sequence based on the operation log, the annotation organization topology, and the preset operation time points; S303. Obtain multiple annotation connection nodes based on the annotation behavior sequence; S304. Obtain the number of verification and correction backtracking steps for each of the multiple labeled connection nodes, and generate a verification and correction backtracking sequence based on the multiple verification and correction backtracking steps and a preset time window; S305. Obtain corresponding node semantic data based on the multiple labeled connection nodes, and perform semantic vectorization processing on the multiple node semantic data to obtain multiple node semantic vectors; S306. Associate the semantic vectors of the multiple nodes according to the labeled time order to obtain the context semantic association sequence.

[0025] As described in steps S301-S306 above, this invention first obtains the storage path structure and annotation task association structure of the annotation data based on the unstructured annotation data, and then constructs an annotation organization topology based on the storage path structure and annotation task association structure. This annotation organization topology maps the storage and task relationships of the annotation data, revealing the organizational pattern of the annotation project. The storage path structure can be parsed from the file directory system of the annotation data management platform, reflecting the physical or logical organization of the data; the annotation task association structure may originate from project management tools, describing the ownership and collaboration relationships between annotation tasks, annotation personnel, and data assets. By integrating these two types of structural information, the constructed annotation organization topology diagram has the physical significance of formally depicting the organizational framework of annotation activities, providing spatial and logical support for understanding the context in which annotation behavior occurs. Then, the corresponding operation logs are obtained based on the unstructured annotation data. Annotation behavior sequences are generated based on these operation logs, the annotation organization topology, and preset operation time points. The operation logs record the specific operations of the annotators. Combined with the organization topology and time points, the generated behavior sequences can capture annotation dynamics. Specifically, the operation logs are typically automatically recorded by the annotation tool or system backend, containing user ID, operation type (e.g., create, modify, delete, review), operation object (e.g., image ID, annotation area ID), and a precise timestamp. Preset operation time points can be used to slice continuous logs, for example, by hour or day. The essence of generating annotation behavior sequences is to reorganize the original, discrete operation records into a coherent chain of operations within the dual dimensions of context and time defined by the organization topology. This makes it possible to analyze the annotator's work rhythm, task switching patterns, and team collaboration processes. Next, multiple annotation connection nodes are obtained based on the annotation behavior sequences. These connection nodes are key points in the behavior sequences, representing the relationships between annotation operations, such as the completion node of an annotation task, the start node of an review task, or the switching node from one annotation object to another. The physical significance of identifying these nodes lies in the fact that they mark noteworthy stages or decision points in the annotation process, serving as anchor points for subsequent more detailed analysis. Then, based on the multiple annotation connection nodes, the number of verification and correction backtracking iterations for each node is obtained. A verification and correction backtracking sequence is generated based on these iterations and a preset time window. The number of iterations reflects the frequency of annotation corrections, and the backtracking sequence records the temporal distribution of correction actions, identifying correction patterns. This is typically achieved by comparing changes in the node's version before and after in the operation log. The preset time window defines the time range for backtracking examination; for example, it only considers correction actions within 24 hours of node completion. Arranging the backtracking iterations of each node in chronological order yields the verification and correction backtracking sequence.The physical significance of this sequence lies in its clear record of the intensity and pattern of re-examination and modification of the annotation results after completion. High-frequency or dense backtracking often indicates uncertainty in the initial annotation quality or stricter implementation of acceptance criteria. Secondly, based on the multiple annotation connection nodes, corresponding node semantic data is obtained, and semantic vectorization is performed on the multiple node semantic data to obtain multiple node semantic vectors. This semantic vectorization transforms textual semantics into numerical vectors, facilitating the calculation of semantic similarity and changes. Node semantic data typically refers to the textual description, label name, or annotation information of the annotation content itself at the corresponding connection node time. Semantic vectorization typically uses pre-trained language models (such as BERT, Word2Vec) to map this textual information into numerical vectors in a high-dimensional space. The technical effect of this step is to transform human-readable semantic content into machine-computable features, laying the foundation for quantifying semantic changes. Finally, the multiple node semantic vectors are associated according to the annotation time order to obtain a contextual semantic association sequence. This association is not merely a simple time ordering; it constructs a semantic evolution trajectory on the time axis. Its physical significance lies in the fact that it makes it possible to analyze the continuous changes, transitions, and even drifts of labeled content in the semantic space. For example, in text sentiment labeling, this sequence can be used to observe whether the judgment of the sentiment polarity of a sentence fluctuates drastically between different proofreading versions (such as a vector transition from "positive" to "negative"), or whether the labeled descriptive terms gradually deviate from the established norms. At the same time, the contextual semantic association sequence captures the changes in semantics over time, revealing the evolution of labeled semantics and providing a basis for subsequent quality assessment.

[0026] In one embodiment, step S4, which involves obtaining the annotation evolution trajectory sequence based on the annotation consistency feature value and the context semantic association sequence, and obtaining the annotation dynamic evolution feature value based on the verification correction backtracking sequence and the annotation evolution trajectory sequence, includes: S401. Extract the semantic change vector based on the context semantic association sequence; S402. Based on the temporal alignment algorithm, the historical sequence of the annotation consistency feature values ​​and the semantic change vector are used to generate an annotation evolution trajectory sequence, wherein the evolution trajectory includes version switching points and semantic transition degrees; S403. Obtain the correction intensity distribution based on the verification correction backtracking sequence, and calculate the inter-version correction density based on the correction intensity distribution and the version switching point; S404. Based on the semantic transition degree and the modified density, the evolutionary activity degree is obtained by weighted fusion, and the evolutionary activity degree is subjected to stability analysis within a preset time to obtain multiple fluctuation characteristic values. S405. Calculate the fluctuation characteristic coefficient based on multiple fluctuation characteristic values, and use the fluctuation characteristic coefficient as the labeling dynamic evolution characteristic value.

[0027] As described in steps S401-S405 above, the present invention first extracts a semantic change vector based on the contextual semantic association sequence, which is composed of semantic vectors from multiple time points in sequence. The extraction of the semantic change vector is typically achieved by calculating the difference between adjacent semantic vectors in the sequence, for example, using vector difference or cosine distance. Its physical significance lies in quantifying the magnitude and direction of semantic changes in the labeled content between consecutive time points, providing a numerical basis for subsequent analysis of semantic evolution. For example, in text classification labeling, this vector can capture the movement trajectory of the semantic descriptions corresponding to the classification labels between different versions in the vector space, thus enabling the semantic change vector to quantify the magnitude and direction of semantic changes. Then, based on a temporal alignment algorithm, a label evolution trajectory sequence is generated from the historical sequence of the label consistency feature values ​​and the semantic change vector, wherein the evolution trajectory includes version switching points and semantic transition degrees. The historical sequence of the label consistency feature values ​​is time-series data collected from the periodic execution results of step S2. Temporal alignment algorithms, such as dynamic time warping, are used to address the potential differences in sampling frequencies or phases between two sequences, ensuring accurate correspondence between them on the time axis. The generated annotation evolution trajectory sequence physically constitutes a joint evolutionary path (consistency and semantics) in two dimensions. Version switching points in this trajectory mark key moments when the annotation state undergoes significant changes, while semantic transitions quantify the intensity of semantic changes near these key points, reflecting the degree of semantic abrupt change in the annotation content during important version iterations. This temporal alignment ensures the consistency of the time series. The evolution trajectory captures changes in annotation data in both time and features, version switching points identify major changes, and semantic transitions measure the degree of semantic change. Temporal alignment algorithms, such as Dynamic Time Warping (DTW), are used to address potential rate mismatches between the two time series, ensuring an accurate correspondence between consistency changes and semantic changes on the time axis. The physical significance of the generated annotation evolution trajectory sequence lies in its depiction of a collaborative evolutionary path of annotation data in both "formal consistency" and "content semantics." Version switching points in this trajectory mark key moments when annotations undergo significant structural or semantic changes. Next, the correction intensity distribution is obtained based on the verification and correction backtracking sequence, and the inter-version correction density is calculated based on the correction intensity distribution and the version switch point. The verification and correction backtracking sequence records the correction frequency for each time period. The correction intensity distribution is obtained by statistically analyzing the concentration of correction operations within different time periods. The inter-version correction density specifically refers to the frequency of correction operations around the version switch point (e.g., within a specific time window before and after the switch point). Its physical significance lies in revealing whether significant changes to the labeled data (version switch) are accompanied by abnormally dense correction activities. High density may indicate that the version transition is unstable or involves significant controversy. Thus, the correction intensity distribution shows the concentration of correction operations, and the inter-version correction density reflects the correction activities between different versions, identifying correction hotspots.Then, the evolutionary activity is obtained by weighted fusion based on the semantic transition degree and the correction density. The evolutionary activity is then subjected to stability analysis over a preset time period to obtain multiple fluctuation characteristic values. The fluctuation characteristic coefficient is typically calculated using the coefficient of variation, i.e., the standard deviation of the fluctuation characteristic value divided by its mean. The physical significance of this step lies in converting the absolute amplitude of the fluctuation into a relative fluctuation intensity relative to the average activity level. A higher fluctuation characteristic coefficient indicates that the evolutionary process of the labeled data is extremely unstable, with large fluctuations in activity; while a lower coefficient indicates that the evolutionary process is stable and controllable. This relative measure makes the evaluation results more comparable to labeled data at different activity levels. Thus, the evolutionary activity, combining semantics and correction behavior, represents the dynamics of the labeled data, and stability analysis identifies the fluctuation patterns of activity. Finally, the fluctuation characteristic coefficient is calculated based on multiple fluctuation characteristic values. The specific implementation process is as follows: First, the average fluctuation characteristic value is calculated based on multiple fluctuation characteristic values ​​and a preset time, where the preset time includes the start and end times. The calculation formula is: ; in, Represents the average fluctuation characteristic value. Indicates the start time. Indicates the end time. Let i represent the fluctuation characteristic value, where i = 1, 2, 3...n, and n represents the number of fluctuation characteristic values. The standard volatility characteristic value is calculated based on multiple volatility characteristic values ​​and the average volatility characteristic value, wherein the calculation formula is: ; in, Represents the standard fluctuation characteristic value. This represents the k-th fluctuation characteristic value. This represents the number of fluctuation characteristic values, where k represents the index of the fluctuation characteristic value. This represents the average fluctuation characteristic value; The volatility characteristic coefficient is calculated based on the standard volatility characteristic value and the average volatility characteristic value, wherein the calculation formula is as follows: ; in, Represents the volatility characteristic coefficient. Indicates the average gas temperature. Indicates the standard fluctuation characteristic value; Using the aforementioned fluctuation characteristic coefficients as feature values ​​for the dynamic evolution of annotations, and incorporating a series of specific algorithms such as temporal alignment, cross-sequence association, density analysis, and stability quantification, an evaluation module capable of deeply analyzing the dynamic evolution patterns of annotation data was constructed. Compared to existing technologies that typically examine version differences in isolation or simply count the number of corrections, this invention creatively correlates and integrates consistency changes, semantic evolution, and correction behavior within a unified time framework. This not only identifies obvious quality fluctuation events but also finely characterizes the dynamic stability of annotation data throughout its entire lifecycle. This deep insight into the evolutionary process enables this invention to effectively distinguish between healthy iterative optimization and chaotic annotation fluctuations when dealing with large-scale annotation data with complex version histories and collaborative backgrounds. It provides crucial dynamic evidence for the final quality assessment, directly contributing to improving the accuracy of early warnings of potential quality risks and significantly reducing the risk of batch misjudgments or omissions caused by uncontrolled evolution processes.

[0028] In one embodiment, step S5, which involves obtaining the annotation quality assessment score based on the annotation consistency feature value and the annotation dynamic evolution feature value, includes: S501. Obtain the corresponding first weight value based on the annotation consistency feature value; S502. Obtain the dynamic-static synergy factor based on the annotation consistency feature value and the annotation dynamic evolution feature value; S503. Obtain the annotation quality assessment score based on the annotation consistency feature value, the first weight value, the dynamic-static synergy factor, and the annotation dynamic evolution feature value, wherein the calculation formula is:

[0029] in, This indicates the quality assessment score. This indicates the consistency feature value of the annotation. This represents the first weight value. This indicates the annotation of dynamically evolving feature values. This represents the dynamic-static synergy factor.

[0030] As described in steps S501-S503 above, the present invention first obtains a corresponding first weight value based on the annotation consistency feature value. This first weight value represents the importance of the consistency feature in quality assessment and can be adjusted according to domain knowledge. Then, a dynamic-static synergy factor is obtained based on the annotation consistency feature value and the annotation dynamic evolution feature value. The calculation of this factor aims to capture the interaction between static consistency and dynamic evolution. A specific implementation is based on the correlation or ratio relationship between the two feature values, for example, P(H) = f(y(Z), P(q)), where the function f can be designed as a function to measure the degree of deviation between the two features. Its physical meaning lies in its characterization of the "dynamic-static consistency" of the annotation data. When an annotation data point simultaneously has high consistency feature values ​​and high dynamic evolution feature values ​​(i.e., high activity and high volatility), this may indicate an inherent contradiction in its quality, and the dynamic-static synergy factor can penalize this. Conversely, if both are low or the changing trends are coordinated, it can be considered normal. Thus, the dynamic-static synergy factor represents the interaction between consistency and dynamic evolution; for example, high consistency but high dynamics may indicate instability. Finally, the annotation quality assessment score is obtained based on the annotation consistency feature value, the first weight value, the dynamic-static synergy factor, and the annotation dynamic evolution feature value. The formula integrates consistency and dynamic evolution features, and the dynamic-static synergy factor adjusts the contribution of dynamic features to generate a comprehensive quality score. Therefore, by introducing configurable weight coefficients and innovative dynamic-static synergy factors, a comprehensive quality assessment model that transcends simple indicator superposition is constructed. Compared to the fixed weight summation or independent threshold judgment methods commonly used in existing technologies, this invention not only allows for flexible adjustment of the assessment focus according to task characteristics, but more importantly, it captures the complex interaction between static quality and dynamic behavior through a synergy factor mechanism. This allows the final annotation quality assessment score to more profoundly reveal the overall quality status of the annotation data. This refined fusion strategy directly contributes to improving the discrimination ability of the assessment system in complex scenarios, especially in distinguishing between complex situations such as "static compliance but dynamic abnormality" or "dynamic stability but static defects," providing more accurate judgment criteria and effectively avoiding the risk of misjudgment and omission due to a single assessment dimension or isolated feature processing.

[0031] In one embodiment, step S6, which involves evaluating the quality of the standard annotation data based on the annotation quality assessment score to obtain the evaluation result, includes: S601. Determine whether the annotation quality assessment score is greater than a preset threshold. S602. If the quality assessment score of the annotation is not greater than the preset threshold, the standard annotation data is determined to be qualified annotation data. S603. If the annotation quality assessment score is greater than the preset threshold, the standard annotation data is determined to be unqualified annotation data.

[0032] As described in steps S601-S603 above, this invention transforms the annotation quality assessment score obtained through multi-dimensional feature depth extraction and fusion calculation into an assessment result with direct guiding significance. Compared to existing quality judgment methods that may rely on isolated indicators, lack systematic fusion, or have opaque decision-making processes, this step forms a complete closed loop with the aforementioned steps. It ensures that the quality assessment is not only multi-dimensional and in-depth, but also that the conclusions are clear and actionable. This design effectively avoids misjudgments (judging actually unstable data as qualified) or omissions (misjudging data with fluctuations but controllable overall quality as unqualified) caused by focusing only on static results or ignoring dynamic processes, providing a reliable quality assurance mechanism for the management and application of large-scale annotation data.

[0033] All numerical calculations in this scheme are performed after normalization to obtain standard values.

[0034] This application also provides a standard labeled data verification and quality assessment system based on multi-dimensional feature extraction, including: The first acquisition module is used to acquire the data to be labeled and to extract features from the data to be labeled based on multi-dimensional features to obtain the standard labeled data to be verified, wherein the standard labeled data includes structured labeled data and unstructured labeled data. The partitioning module is used to partition nodes according to the structured annotation data to obtain multiple standard annotation node data, and to obtain the flipping annotation frequency and annotation structure drift degree according to the multiple standard annotation node data, and to obtain the annotation consistency feature value according to the flipping annotation frequency and annotation structure drift degree; The second acquisition module is used to acquire an annotation behavior sequence based on the unstructured annotation data, and to acquire a verification and correction backtracking sequence and a context semantic association sequence based on the annotation behavior sequence. The third acquisition module is used to acquire the annotation evolution trajectory sequence based on the annotation consistency feature value and the context semantic association sequence, and to acquire the annotation dynamic evolution feature value based on the verification correction backtracking sequence and the annotation evolution trajectory sequence; The fourth acquisition module is used to obtain the annotation quality evaluation score based on the annotation consistency feature value and the annotation dynamic evolution feature value; The evaluation module is used to evaluate the quality of the standard annotation data based on the annotation quality evaluation score, and obtain the evaluation result.

[0035] In one embodiment, the second acquisition module includes: The first acquisition unit is used for the steps of dividing the structured annotation data into nodes to obtain multiple standard annotation node data, obtaining the annotation flipping frequency and annotation structure drift degree based on the multiple standard annotation node data, and obtaining annotation consistency feature values ​​based on the annotation flipping frequency and annotation structure drift degree, including: The second acquisition unit is used to acquire an annotation logic unit based on the structured annotation data, wherein the annotation logic unit includes a start annotation unit, an annotation attribute unit, and an annotation end unit; A partitioning unit is used to partition the structured annotation data according to the annotation logic unit to obtain multiple standard annotation node data; The third acquisition unit is used to acquire multiple timestamp sequences and multiple annotation content change records for each standard annotation node data, and to acquire the number of flips of multiple annotation labels based on the multiple annotation content change records; The fourth acquisition unit is used to acquire the number of annotation operations within a preset time window based on the multiple timestamp sequences; The fifth acquisition unit is used to acquire multiple first flip label frequencies based on the multiple number of labeling operations and the multiple number of flips of the label, and to acquire an average flip label frequency based on the multiple first flip label frequencies, and to use the average flip label frequency as the flip label frequency. The sixth acquisition unit is used to acquire the graph structure features of each standard labeled node data, and to acquire multiple graph structure feature difference values ​​within a preset adjacent time window based on the multiple graph structure features shown, and to acquire the average graph structure feature difference value of the multiple graph structure feature difference values. The calculation unit is used to calculate the difference ratio based on the average difference value of the map structure features and the preset difference value of the map structure features, and to use the difference ratio as the drift degree of the labeled structure. The seventh acquisition unit is used to acquire the connection coefficient in the annotation process based on the starting annotation unit, the annotation attribute unit and the annotation end unit; The fusion unit is used to obtain a label consistency feature value by weighted fusion based on the flip label frequency, the label structure drift degree and the connection coefficient.

[0036] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0037] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0038] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0039] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0040] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for verifying and evaluating the quality of standard labeled data based on multi-dimensional feature extraction, characterized in that, include: Obtain the data to be labeled, and extract features from the data to be labeled based on multi-dimensional features to obtain the standard labeled data to be verified, wherein the standard labeled data includes structured labeled data and unstructured labeled data; The structured annotation data is divided into nodes to obtain multiple standard annotation node data. The annotation flipping frequency and annotation structure drift are obtained based on the multiple standard annotation node data. The annotation consistency feature value is obtained based on the annotation flipping frequency and annotation structure drift. Annotation behavior sequence is obtained based on the unstructured annotation data, and verification and correction backtracking sequence and context semantic association sequence are obtained based on the annotation behavior sequence; The annotation evolution trajectory sequence is obtained based on the annotation consistency feature value and the context semantic association sequence, and the annotation dynamic evolution feature value is obtained based on the verification and correction backtracking sequence and the annotation evolution trajectory sequence; The annotation quality assessment score is obtained based on the annotation consistency feature value and the annotation dynamic evolution feature value; The standard annotation data is evaluated based on the annotation quality assessment score to obtain the evaluation result.

2. The method for standard labeled data verification and quality assessment based on multi-dimensional feature extraction according to claim 1, characterized in that, The steps of dividing the structured annotation data into nodes to obtain multiple standard annotation node data, obtaining the annotation flipping frequency and annotation structure drift based on the multiple standard annotation node data, and obtaining annotation consistency feature values ​​based on the annotation flipping frequency and annotation structure drift include: Annotation logic units are obtained based on the structured annotation data, wherein the annotation logic unit includes a start annotation unit, an annotation attribute unit, and an annotation end unit; The structured annotation data is divided according to the annotation logic unit to obtain multiple standard annotation node data; Obtain multiple timestamp sequences and multiple annotation content change records for each standard annotation node data, and obtain the number of flips for multiple annotation labels based on the multiple annotation content change records; The number of annotation operations within a preset time window is obtained based on the multiple timestamp sequences. Multiple first flip label frequencies are obtained based on the number of times the multiple labeling operations are performed and the number of times the multiple labeling tags are flipped, and an average flip label frequency is obtained based on the multiple first flip label frequencies, and the average flip label frequency is used as the flip label frequency. Obtain the graph structure features of each standard labeled node data, and obtain multiple graph structure feature difference values ​​within a preset adjacent time window based on multiple graph structure features, and obtain the average graph structure feature difference value of the multiple graph structure feature difference values; The difference ratio is calculated based on the average difference value of the structural features of the graph and the difference value of the preset structural features of the graph, and the difference ratio is used as the drift degree of the labeled structure. The connection coefficient in the annotation process is obtained based on the starting annotation unit, the annotation attribute unit, and the annotation ending unit; The annotation consistency feature value is obtained by weighted fusion based on the flipping annotation frequency, the annotation structure drift degree, and the connection coefficient.

3. The method for standard labeled data verification and quality assessment based on multi-dimensional feature extraction according to claim 1, characterized in that, The steps of obtaining the annotation behavior sequence based on the unstructured annotation data, and obtaining the verification and correction backtracking sequence and the context semantic association sequence based on the annotation behavior sequence, include: The storage path structure and annotation task association structure of the annotation data are obtained based on the unstructured annotation data, and the annotation organization topology is constructed based on the storage path structure and annotation task association structure. Obtain the corresponding operation logs based on the unstructured labeled data, and generate a labeling behavior sequence based on the operation logs, the labeling organization topology, and preset operation time points; Multiple annotation connection nodes are obtained based on the annotation behavior sequence; The number of verification and correction backtracking steps for each of the multiple labeled connection nodes is obtained, and a verification and correction backtracking sequence is generated based on the multiple verification and correction backtracking steps and a preset time window. Obtain corresponding node semantic data based on multiple labeled connection nodes, and perform semantic vectorization processing on the multiple node semantic data to obtain multiple node semantic vectors; The semantic vectors of the multiple nodes are associated according to the labeled time order to obtain the context semantic association sequence.

4. The method for standard labeled data verification and quality assessment based on multi-dimensional feature extraction according to claim 1, characterized in that, The steps of obtaining the annotation evolution trajectory sequence based on the annotation consistency feature value and the context semantic association sequence, and obtaining the annotation dynamic evolution feature value based on the verification and correction backtracking sequence and the annotation evolution trajectory sequence, include: Extract semantic change vectors based on the contextual semantic association sequence; Based on the temporal alignment algorithm, the historical sequence of the annotation consistency feature values ​​and the semantic change vector are used to generate an annotation evolution trajectory sequence, wherein the evolution trajectory includes version switching points and semantic leap degrees; The correction intensity distribution is obtained based on the verification and correction backtracking sequence, and the inter-version correction density is calculated based on the correction intensity distribution and the version switching point; The evolutionary activity is obtained by weighted fusion of the semantic transition degree and the modified density, and the evolutionary activity is subjected to stability analysis within a preset time to obtain multiple fluctuation characteristic values. The fluctuation characteristic coefficient is calculated based on multiple fluctuation characteristic values, and the fluctuation characteristic coefficient is used as the labeling dynamic evolution characteristic value.

5. The method for standard labeled data verification and quality assessment based on multi-dimensional feature extraction according to claim 1, characterized in that, The step of obtaining the annotation quality assessment score based on the annotation consistency feature value and the annotation dynamic evolution feature value includes: The first weight value is obtained based on the annotation consistency feature value; The dynamic-static synergy factor is obtained based on the annotation consistency feature value and the annotation dynamic evolution feature value; The annotation quality assessment score is obtained based on the annotation consistency feature value, the first weight value, the dynamic-static synergy factor, and the annotation dynamic evolution feature value, wherein the calculation formula is: in, This indicates the quality assessment score. This indicates the consistency feature value of the annotation. This represents the first weight value. This indicates the annotation of dynamically evolving feature values. This represents the dynamic-static synergy factor.

6. The method for standard labeled data verification and quality assessment based on multi-dimensional feature extraction according to claim 1, characterized in that, The step of evaluating the quality of the standard annotation data based on the annotation quality assessment score to obtain the evaluation result includes: Determine whether the annotation quality assessment score is greater than a preset threshold; If the quality assessment score of the annotation is not greater than the preset threshold, the standard annotation data is determined to be qualified annotation data; If the quality assessment score of the annotation is greater than the preset threshold, the standard annotation data is determined to be unqualified annotation data.

7. A standard labeled data verification and quality assessment system based on multi-dimensional feature extraction, characterized in that, include: The first acquisition module is used to acquire the data to be labeled and to extract features from the data to be labeled based on multi-dimensional features to obtain the standard labeled data to be verified, wherein the standard labeled data includes structured labeled data and unstructured labeled data. The partitioning module is used to partition nodes according to the structured annotation data to obtain multiple standard annotation node data, and to obtain the flipping annotation frequency and annotation structure drift degree according to the multiple standard annotation node data, and to obtain the annotation consistency feature value according to the flipping annotation frequency and annotation structure drift degree; The second acquisition module is used to acquire an annotation behavior sequence based on the unstructured annotation data, and to acquire a verification and correction backtracking sequence and a context semantic association sequence based on the annotation behavior sequence; The third acquisition module is used to acquire the annotation evolution trajectory sequence based on the annotation consistency feature value and the context semantic association sequence, and to acquire the annotation dynamic evolution feature value based on the verification correction backtracking sequence and the annotation evolution trajectory sequence; The fourth acquisition module is used to obtain the annotation quality evaluation score based on the annotation consistency feature value and the annotation dynamic evolution feature value; The evaluation module is used to evaluate the quality of the standard annotation data based on the annotation quality evaluation score, and obtain the evaluation result.

8. The standard labeled data verification and quality assessment system based on multi-dimensional feature extraction according to claim 7, characterized in that, The second acquisition module includes: The first acquisition unit is used for the steps of dividing the structured annotation data into nodes to obtain multiple standard annotation node data, obtaining the annotation flipping frequency and annotation structure drift degree based on the multiple standard annotation node data, and obtaining annotation consistency feature values ​​based on the annotation flipping frequency and annotation structure drift degree, including: The second acquisition unit is used to acquire an annotation logic unit based on the structured annotation data, wherein the annotation logic unit includes a start annotation unit, an annotation attribute unit, and an annotation end unit; A partitioning unit is used to partition the structured annotation data according to the annotation logic unit to obtain multiple standard annotation node data; The third acquisition unit is used to acquire multiple timestamp sequences and multiple annotation content change records for each standard annotation node data, and to acquire the number of flips of multiple annotation labels based on the multiple annotation content change records; The fourth acquisition unit is used to acquire the number of annotation operations within a preset time window based on the multiple timestamp sequences; The fifth acquisition unit is used to acquire multiple first flip label frequencies based on the multiple number of labeling operations and the multiple number of flips of the label, and to acquire an average flip label frequency based on the multiple first flip label frequencies, and to use the average flip label frequency as the flip label frequency. The sixth acquisition unit is used to acquire the graph structure features of each standard labeled node data, and to acquire multiple graph structure feature difference values ​​within a preset adjacent time window based on the multiple graph structure features shown, and to acquire the average graph structure feature difference value of the multiple graph structure feature difference values. The calculation unit is used to calculate the difference ratio based on the average difference value of the map structure features and the preset difference value of the map structure features, and to use the difference ratio as the drift degree of the labeled structure. The seventh acquisition unit is used to acquire the connection coefficient in the annotation process based on the starting annotation unit, the annotation attribute unit and the annotation end unit; The fusion unit is used to obtain a label consistency feature value by weighted fusion based on the flip label frequency, the label structure drift degree and the connection coefficient.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.