A prompt word driven data labeling method and system

By generating joint guidance vectors and fusing domain features with operational behavior vectors, and by using reinforcement learning to optimize the prompt word knowledge graph, the inefficiency and low accuracy of prompt word-driven data annotation in existing technologies are solved, achieving efficient and high-quality data annotation.

CN121683727BActive Publication Date: 2026-05-08HANGZHOU SUOYI NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU SUOYI NETWORK TECHNOLOGY CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing prompt-driven data annotation technologies struggle to achieve high accuracy and adaptability, and fail to effectively coordinate prompts with operational behaviors, resulting in low annotation efficiency and unstable quality.

Method used

By generating a joint guidance vector, fusing the domain features and operational behavior vectors of the annotation task, and using reinforcement learning to iteratively optimize the prompt word knowledge graph, annotation defects are identified and prompt word parameters are adjusted, forming a closed-loop feedback mechanism.

Benefits of technology

It improves the accuracy and efficiency of data annotation, ensures that prompt words meet the requirements of annotation tasks and adapt to actual operation processes, dynamically adapts to annotation quality requirements, reduces error rates, and optimizes the adaptability of prompt words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121683727B_ABST
    Figure CN121683727B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data processing, and provides a data labeling method and system based on prompt word driving. A joint guide vector is generated by fusing the domain features of the labeling task and the operation behavior vector of the labeling personnel, which greatly reduces the labeling errors caused by guide bias, improves the labeling efficiency and the preliminary labeling quality. The initial labeling result is analyzed to identify the labeling defect types and the domain distribution difference of the associated prompt words, which provides a clear target for subsequent knowledge graph parameter optimization. The labeling defect types and the prompt word domain distribution difference are taken as the state space, and a reward function is constructed in combination with the labeling accuracy and the domain adaptability. The parameters of the prompt word knowledge graph are iteratively corrected through reinforcement learning, and high-quality prompt words can be continuously output. After the iteration and stability of the prompt word knowledge graph, the fusion coefficients of the joint guide vector are updated based on the quality evaluation data feedback to ensure the collaborative adaptation of the joint guide vector and the optimized knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a data annotation method and system based on prompt words. Background Technology

[0002] With the popularization of artificial intelligence and big data technologies, data annotation, as a fundamental step in model training and algorithm optimization, has seen its technical system continuously evolve. Early data annotation relied on manual annotation, sentence by sentence and frame by frame, which was inefficient, highly influenced by subjective experience, and difficult to guarantee consistent annotation quality.

[0003] To address this issue, the industry has gradually developed semi-automatic methods such as rule-driven annotation and template-based annotation. These methods reduce manual intervention and improve annotation efficiency by pre-setting domain rules or annotation templates. In recent years, prompt-based annotation technology has become the mainstream development direction. Its core logic is to guide annotators or automated tools to complete annotation tasks through prompts, achieving an efficient process of prompt guidance and rapid annotation.

[0004] This technology breaks away from the dependence of traditional rule-driven methods on fixed scenarios. It can cover more annotation needs by expanding prompt words and is widely used in various data annotation scenarios such as text, images and audio, promoting the development of data annotation towards intelligence and flexibility.

[0005] Although prompt-driven annotation technology has made some progress, it still has many limitations in practical applications, making it difficult to meet the requirements of high-precision and high-adaptability annotation. Existing prompt-driven methods mostly rely on a single dimension of guidance and fail to achieve deep collaboration between prompt expertise and annotation operation behavior, resulting in the retrieved prompts deviating from the annotation task specifications or reducing annotation efficiency and process standardization. Existing technologies mostly perform overall statistics on defects generated during the annotation process, without establishing a precise correlation between defect types and prompts, making it impossible to locate the guidance deviation of the prompts themselves, resulting in a lack of clear targets for subsequent optimization.

[0006] Based on the shortcomings of the existing technology, the technical problem to be solved in this application is how to achieve the accuracy, efficiency and quality stability of data annotation through the coordinated guidance of prompts and operations. Summary of the Invention

[0007] The purpose of this application is to overcome the shortcomings of the prior art and provide a data annotation method and system based on prompt words.

[0008] To achieve the above objectives, this application adopts the following technical solution:

[0009] Firstly, a prompt word-driven data annotation method is provided, which includes: receiving annotation task requests, parsing annotation tasks to extract domain features of annotation tasks, extracting operation behavior vectors from the operation behavior of annotation personnel, and fusing domain features and operation behavior vectors to generate joint guidance vectors;

[0010] Based on the joint guiding vector, prompt words are retrieved from the prompt word knowledge graph. The data to be labeled and prompt words are processed to output the initial labeling results and confidence scores. The initial labeling results are analyzed to identify the differences in the domain distribution of labeling defect types and related prompt words.

[0011] Initial annotation results with confidence scores below the confidence threshold are selected. The reward function is determined based on annotation accuracy and domain adaptability. The domain distribution difference between annotation defect types and associated prompt words is used as the state space. The parameters of the prompt word knowledge graph are iteratively corrected through reinforcement learning.

[0012] When the continuous iterations tend to stabilize, the iterative optimization is terminated and the labeled dataset is output. Based on the quality evaluation data of the iterative process, the fusion coefficient of the joint guiding vector is updated.

[0013] Optionally, the domain features of the extraction and annotation task include:

[0014] Receive annotation task requests and perform hierarchical semantic decomposition of the text description of the annotation task requests, including basic semantics, task objectives, and data constraints, to obtain initial features;

[0015] Retrieve historical domain features from historical annotation tasks, and use cosine similarity to match the distance between initial features and historical domain features to filter historical annotation tasks of the same type with feature similarity not less than the similarity threshold.

[0016] Identify cross-domain overlapping features in the initial features, calculate the weights by weighting the annotation accuracy and feature similarity of similar historical annotation tasks, and retain the cross-domain overlapping features with high weights as core features;

[0017] Basic weights are assigned to core features, and the basic weights are dynamically adjusted based on the contribution of core features to the labeling accuracy in similar historical labeling tasks. After weighted fusion, the domain features of the labeling task are formed.

[0018] Optionally, extracting the operation behavior vector includes:

[0019] Acquire the basic, interactive, and decision-making behaviors of annotation personnel at different times in similar historical annotation tasks;

[0020] Based on the historical annotation accuracy of similar annotation tasks, the behavior weights of different operation behaviors are dynamically allocated according to the operation type.

[0021] A behavior feature matrix is ​​constructed based on behavior weights and operational behaviors. Operation patterns are clustered using density clustering algorithms, and abnormal operational behaviors are eliminated to extract the behavior features of different operation patterns.

[0022] By adjusting the weights of the encoding dimensions based on the domain features of the annotation task, the behavioral features are mixed-encoded and normalized to generate an operational behavior vector consistent with the domain feature dimensions.

[0023] Optionally, generating the joint guiding vector includes:

[0024] Retrieve the preset domain ontology knowledge base, extract the quantitative value of the feature importance of the current annotation task's domain, and assign domain adaptation weights to the domain features according to the proportion of the quantitative value;

[0025] Select a small sample of data to be labeled for the current labeling task, complete the small sample labeling based on the operation behavior vector, and determine the effective weight of the operation behavior vector based on the degree of fit between the labeling results and the standard labeling.

[0026] Using domain adaptation weights and effective behavior weights as fusion coefficients, we weight and fuse domain features and operational behavior vectors with consistent dimensions to generate a joint guidance vector.

[0027] Optionally, the search suggestions include:

[0028] The prompt word knowledge graph is preprocessed by classifying the prompt words according to their respective domains and quantifying them into domain category vectors, and then associating them with the operation behavior tags of the personnel labeled in the corresponding domains.

[0029] Extract the corresponding domain component from the joint guidance vector, perform cosine similarity coarse matching with the domain category vector of the prompt word knowledge graph, and filter out candidate prompt words belonging to the domain of the current annotation task;

[0030] Extract the corresponding operation components from the joint guidance vector, calculate the matching degree between the operation behavior label and the operation component in the candidate prompt words, retain the prompt words with a matching degree greater than the matching threshold, and obtain the usable candidate set;

[0031] Based on the fusion coefficient of the joint guiding vector, the prompt words in the available candidate set are weighted and sorted. The number of prompt words is preset based on the annotation complexity of the current annotation task, and the prompt words are output according to the sorting results.

[0032] Optionally, the output initial annotation results and confidence levels include:

[0033] The output prompts are classified according to the proportion of the fusion coefficient. Prompts with a fusion coefficient greater than the preset proportion are set as domain prompts and operation prompts, respectively.

[0034] Perform domain semantic annotation on the data to be annotated based on domain prompt words and domain features, and perform operation adaptation annotation based on operation prompt words and operation behavior vectors to generate domain annotation results and operation annotation results;

[0035] Calculate the degree of fit between the two types of annotation results. If the degree of fit is greater than the degree of fit threshold, merge them into the initial annotation result. Otherwise, call the preset annotation standard fragment to correct the divergent parts and generate the initial annotation result.

[0036] Based on the consistency of the generated results, the matching degree of the prompt words and the fit of the operation annotation results, the confidence of the prompt word adaptability and the operation association is calculated and then weighted and summed to output the confidence of the initial annotation results.

[0037] Optionally, the difference in domain distribution between the identified and labeled defect types and associated prompt words includes:

[0038] Based on the confidence level of the initial annotation results and the discrepancies between the two types of annotation results, the discrepancies with confidence levels less than the confidence threshold are filtered out, and the discrepancies are classified according to their sources to identify the types of annotation defects.

[0039] The annotation defect type is bound to the domain prompt words and operation prompt words that generate the initial annotation result, and the frequency of prompt words with different proportions of fusion coefficient triggering the corresponding annotation defect type is statistically analyzed;

[0040] Based on the subdivision of the domain to which the current annotation task belongs, calculate the defect induction rate of the same prompt word under different subdivisions based on frequency, and statistically analyze the distribution ratio of annotation defect types among different subdivisions to obtain the domain distribution differences of the associated prompt words corresponding to the annotation defect types.

[0041] Optionally, the parameters of the iteratively revised prompt word knowledge graph include:

[0042] Initial annotation results with confidence scores below the confidence threshold are selected, and a reward function is determined based on annotation accuracy and domain adaptability. The parameters of the reward function are then dynamically adjusted according to differences in domain distribution.

[0043] The differences in the domain distribution of labeled defect types and associated prompt words are quantitatively represented as a state space. Through reinforcement learning, the parameters of the prompt word knowledge graph are hierarchically corrected according to the different proportions of the frequency of the labeled defect type corresponding to the prompt word and the fusion coefficient.

[0044] After each round of correction, prompt words are retrieved again based on the corrected prompt word knowledge graph and labeled. The cumulative score of the reward function is calculated. If the cumulative score does not reach the convergence threshold, the state space is updated based on the new domain distribution difference until the iteration tends to stabilize.

[0045] Optionally, the fusion coefficients of the feedback update joint guiding vector include:

[0046] Based on the quality assessment data from the iterative process, the annotation accuracy and defect induction rate are extracted, and the fusion coefficients of the joint guiding vectors are correlated to form a correlation relationship.

[0047] Based on the correlation, the adjustment rules are determined in conjunction with the deviation direction of the quality assessment data. The fusion coefficients are updated based on the adjustment rules. The joint guidance vector is regenerated using a small sample of unlabeled data to complete the labeling with the search prompt words and monitor the labeling effect.

[0048] By comparing the annotation results with the quality benchmark when the iterations tend to stabilize, it is possible to determine whether the fusion coefficients of the joint guiding vector after the update should be solidified.

[0049] Secondly, a prompt-word-driven data annotation system is provided, which includes: a task parsing module, a data annotation module, an intelligent iteration module, and a quality feedback module;

[0050] The task parsing module is used to receive annotation task requests, parse the annotation tasks to extract the domain features of the annotation tasks, extract the operation behavior vectors from the operation behavior of the annotators, and fuse the domain features and operation behavior vectors to generate a joint guidance vector.

[0051] The data annotation module retrieves prompt words from the prompt word knowledge graph based on the joint guiding vector, processes the data to be annotated and the prompt words to output the initial annotation results and confidence level, and analyzes the initial annotation results to identify the differences in the domain distribution of annotation defect types and related prompt words;

[0052] The intelligent iteration module is used to filter initial annotation results with confidence scores below the confidence threshold. It determines the reward function based on annotation accuracy and domain adaptability, and uses the difference in domain distribution between annotation defect types and associated prompt words as the state space. It iteratively corrects the parameters of the prompt word knowledge graph through reinforcement learning.

[0053] The quality feedback module is used to terminate iterative optimization and output labeled datasets when continuous iterations tend to stabilize. Based on the quality evaluation data of the iterative process, it provides feedback to update the fusion parameters of the joint guiding vector.

[0054] Compared with existing technologies, the beneficial effects of this application are as follows: By integrating the domain features of the annotation task with the operational behavior vectors of the annotators to generate a joint guidance vector, the domain features ensure the professionalism and domain standardization of the prompt words, while the operational behavior vectors conform to the practical habits of annotation. The synergy between the two ensures that the retrieved prompt words not only meet the requirements of the annotation task but also adapt to the actual operation process, significantly reducing annotation errors caused by guidance deviations and improving annotation efficiency and initial annotation quality. Analyzing the initial annotation results identifies the differences in the domain distribution of annotation defect types and related prompt words, thereby clarifying the defect induction patterns of different prompt words in different domain dimensions, providing a clear target for subsequent knowledge graph parameter optimization, and improving the targeting and effectiveness of the optimization.

[0055] Using the differences between the labeled defect types and the domain distribution of prompt words as the state space, a reward function is constructed by combining labeling accuracy and domain adaptability. The parameters of the prompt word knowledge graph are iteratively corrected through reinforcement learning, so that reinforcement learning has clear environmental feedback and optimization guidance, can dynamically adapt to the defect distribution pattern and labeling quality requirements, and the corrected knowledge graph is more in line with the labeling task requirements, and can continuously output high-quality prompt words.

[0056] After the prompt word knowledge graph has been iterated and stabilized, the fusion coefficient of the joint guidance vector is updated based on the feedback of quality assessment data, forming a complete closed loop of guidance, annotation, defect analysis, optimization and feedback. This ensures that the joint guidance vector and the optimized knowledge graph are synergistically adapted, so that the annotation effect is continuously optimized and tends to be stable. At the same time, the closed loop system can reuse optimization experience and reduce the repeated optimization cost of similar annotation tasks.

[0057] This application is not limited to a specific field or data type. Through the generalized fusion logic of domain features and operational behavior vectors and the universal analysis method of defect distribution, it can flexibly adapt to various data annotation tasks such as text, images and audio, as well as the domain norms and operating habits of different industries. Compared with the dependence of existing technologies on specific scenarios, it has a wider range of applications and stronger engineering application value. Attached Figure Description

[0058] Figure 1 A flowchart illustrating a prompt-word-driven data annotation method provided in this application embodiment;

[0059] Figure 2 The following is a flowchart illustrating the logic for generating a joint guiding vector, as provided in an embodiment of this application.

[0060] Figure 3 This is a logical flowchart illustrating the difference in domain distribution between the identified and labeled defect types and associated prompts provided in an embodiment of this application. Detailed Implementation

[0061] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more pairs. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0062] Example 1

[0063] refer to Figure 1 This application provides a flowchart of a data annotation method based on prompt words, which includes:

[0064] S1. Receive annotation task requests, parse the annotation tasks to extract the domain features of the annotation tasks, and extract the operation behavior vectors from the operation behaviors of the annotators. Fuse the domain features and operation behavior vectors to generate a joint guidance vector.

[0065] Furthermore, extracting domain features for annotation tasks includes:

[0066] Receive annotation task requests and perform hierarchical semantic decomposition of the text description of the annotation task requests, including basic semantics, task objectives, and data constraints, to obtain initial features;

[0067] Retrieve historical domain features from historical annotation tasks, and use cosine similarity to match the distance between initial features and historical domain features to filter historical annotation tasks of the same type with feature similarity not less than the similarity threshold.

[0068] Identify cross-domain overlapping features in the initial features, calculate the weights by weighting the annotation accuracy and feature similarity of similar historical annotation tasks, and retain the cross-domain overlapping features with high weights as core features;

[0069] Basic weights are assigned to core features, and the basic weights are dynamically adjusted based on the contribution of core features to the labeling accuracy in similar historical labeling tasks. After weighted fusion, the domain features of the labeling task are formed.

[0070] In audit scenarios, annotation requests are often unstructured text, mixing audit requirements, data constraints, and redundant expressions. If this type of text is directly used for feature extraction, semantic confusion can lead to deviations in subsequent feature matching, making it impossible to accurately locate the direction of audit annotation. After receiving audit annotation requests, we use natural language processing technology to perform hierarchical semantic segmentation, fully aligning with the core requirements of financial audit annotation. First, we perform basic semantic level segmentation, completing text segmentation and part-of-speech tagging, and removing stop words. Text segmentation uses industry-standard natural language processing tools to ensure accurate segmentation of audit-related terms. Part-of-speech tagging is used to distinguish between nouns and verbs, helping to remove invalid information. Stop words refer to auxiliary words and conjunctions that do not contribute to actual semantics, as these words can interfere with the extraction of core features.

[0071] Then, the task objectives were broken down into hierarchical levels to identify core actions and target objects. The core actions are annotations, and the target objects are accounts receivable details in the audit scenario. The annotation requirements were clarified to distinguish between aging intervals and details related to bad debt provision. The boundaries of the task objectives were locked in through keyword matching to eliminate vague expressions and clarify the priority of each annotation content. Finally, the data constraint levels were broken down, and semantic matching combined with regular expressions was used to extract constraints. The data scope was clarified to be the general ledger and subsidiary ledger texts that had been preliminarily verified. The data format constraint was to process only formal electronic texts and exclude handwritten drafts and blurry scans.

[0072] These constraints are integrated with basic semantics and task objectives to form initial features, ensuring that the initial features cover the core requirements and constraint boundaries of the audit annotation task. This effectively solves the problems of semantic confusion and redundant information interference in audit texts, achieves accurate decomposition and structured representation of annotation requirements, provides standardized input for subsequent dimensional alignment with historical domain features, significantly reduces matching bias, and improves the efficiency of subsequent feature processing, ensuring that the extracted domain features fit the domain attributes of financial audit.

[0073] Historical annotation tasks involve massive amounts of audit-related data, covering different types of tasks such as financial statement audits and internal control audits. Blindly reusing all historical data would introduce interfering features from different types of audit tasks. These features are irrelevant to the current accounts receivable annotation task and would cause subsequent core feature extraction to deviate from the target. First, historical domain features of historical annotation tasks are retrieved. The features of various previously completed audit annotation tasks are transformed according to the hierarchical semantic decomposition logic to ensure that the dimensions of historical domain features are consistent with the dimensions of the current initial features. The feature transformation follows the same hierarchical decomposition rules to ensure the uniformity of the feature structure. Then, historical domain features are classified and stored according to audit subdomains for easy retrieval. Each historical domain feature is associated with corresponding annotation effect data, providing a basis for subsequent screening and verification.

[0074] Next, feature dimension alignment and normalization are performed. Since the descriptions of historical domain features may differ across different historical annotation tasks, terminology normalization must be completed first. Terminology normalization means unifying synonymous expressions into standard terms, such as unifying "accounts receivable" and "accounts receivable" into standard expressions for audit scenarios. The distance between the initial features and historical domain features is matched using cosine similarity. Cosine similarity is used to quantify the feature similarity between two features; the higher the feature similarity, the greater the similarity value. The similarity threshold is a preset benchmark value based on the matching effect of historical audit annotation tasks. This benchmark value can be adjusted according to the feature complexity of the current annotation task to ensure the accuracy of selecting similar tasks. Historical similar annotation tasks with feature similarity not less than the similarity threshold are selected, and non-similar annotation tasks with feature similarity less than the similarity threshold are removed.

[0075] Finally, a secondary verification is performed on the selected historical similar annotation tasks. The overlap rate of the task objectives is used to determine whether the selected historical similar annotation tasks are highly consistent with the current accounts receivable detail annotation task, and to exclude interfering tasks with high similarity but different objectives. Thus, by classifying and storing historical domain features and normalizing terms, the problems of inconsistent semantics and classification confusion of historical domain features are solved, the accuracy of feature similarity matching is improved, and the similarity threshold and secondary verification effectively exclude interference from dissimilar annotation tasks and biased tasks, ensuring that the selected historical annotation tasks are consistent with the current task attributes. At the same time, the scope of historical data processing is narrowed, the amount of computation for subsequent feature extraction is reduced, the efficiency of the overall process is improved, and high-quality data support is provided for core feature extraction.

[0076] The overlapping parts between the historical domain features of similar historical annotation tasks and the initial features of the current annotation task include both core features of accounts receivable annotation in the audit scenario and general redundant features. General redundant features do not contribute positively to the accuracy of audit annotation, and retaining them will interfere with subsequent weight adjustment and feature fusion. The overlapping parts between the initial features of the current annotation task and the features of similar historical tasks are identified, namely cross-domain overlapping features. These features are semantic units common to similar annotation tasks in the audit scenario. The identification process adopts a semantic matching method, combined with the semantics of the audit domain, to ensure the accurate identification of cross-domain overlapping features. After the identification is completed, the cross-domain overlapping features are deduplicated to avoid feature duplication that will cause deviations in subsequent weighted calculations.

[0077] Then, a two-dimensional weighted calculation is performed on each cross-domain overlapping feature. The purpose of the weighted calculation is to assign a weight to each cross-domain overlapping feature, which is used to determine the importance of the cross-domain overlapping feature. The first dimension is the annotation accuracy of similar historical annotation tasks. Annotation accuracy refers to the actual contribution of the cross-domain overlapping feature to the quality of the annotation results in similar historical annotation tasks. The contribution is judged by the accuracy and error rate of past annotations. If the annotation accuracy significantly improves and the error rate decreases when a certain cross-domain overlapping feature appears in past tasks, then the annotation accuracy weight corresponding to the cross-domain overlapping feature is higher. The second dimension is the feature similarity between the cross-domain overlapping feature and the current annotation task. Feature similarity is quantified by the degree of correlation between the cross-domain overlapping feature and the core objective of the current annotation task. The higher the degree of correlation, the higher the weight of the feature similarity.

[0078] In conjunction with data constraints, meaningless cross-domain overlapping features within the constraint range are downweighted until they are eliminated. After the weighted calculation is completed, a weight screening standard is set. The screening standard is dynamically adjusted based on the annotation experience in the audit field to ensure that core features are retained while redundant features are eliminated. Cross-domain overlapping features with weights greater than the weight threshold are retained as core features, which will serve as the core objects for subsequent domain feature fusion. This achieves accurate screening of cross-domain overlapping features, retaining core features that are both practical and relevant, eliminating redundant information, and avoiding confusion in subsequent weight adjustments due to redundancy of core features. The dual-dimensional weighted calculation combines historical annotation data with the requirements of the current annotation task to avoid the one-sidedness of single-dimensional screening, improve the quality of core features, and ensure that core features can accurately adapt to the accounts receivable annotation task in the audit scenario.

[0079] Different core features contribute differently to the accuracy of audit annotation. Fixed weights cannot adapt to the actual needs of audit tasks. Some core features have contributed highly in historical audit tasks, while the importance of others needs to be adjusted based on the specific circumstances of the current task. A base weight is assigned to each core feature. The base weight is an initial weight set based on the general importance of the core feature in audit annotation tasks. This base weight is not a fixed value and will be adjusted based on subsequent actual contributions. Then, data from similar historical annotation tasks is retrieved to analyze the actual contribution of each core feature to the annotation accuracy. If a core feature can significantly improve the accuracy of audit annotation and reduce the annotation error rate in past annotation tasks, the base weight of that core feature is increased. If a core feature is prone to causing ambiguity in multiple annotation tasks, leading to an increase in the annotation error rate, the base weight of that core feature is decreased.

[0080] After the basic weights are adjusted, they are normalized to avoid over-emphasis on any core feature and to ensure that the sum of the basic weights is a fixed value. Finally, the core features with the adjusted basic weights are weighted and fused. During the fusion process, the semantic information and basic weight ratio of each core feature are preserved, while maintaining the logical relationship between the core features. This facilitates the subsequent collaborative fusion with the operational behavior vector, ultimately forming domain features adapted to the accounts receivable labeling task in the audit scenario. This ensures that the weight allocation of the domain features aligns with the actual needs of the audit task, strengthens the guiding role of high-contribution core features, improves the adaptability of the domain features to the audit labeling task, avoids imbalance of basic weights through weight normalization, and ensures the synergy between core features through logical relationship fusion. The final generated domain features can accurately represent the core needs of the audit labeling task, providing high-quality input for the subsequent generation of joint guidance vectors.

[0081] Furthermore, the extraction of the operation behavior vector includes:

[0082] Acquire the basic, interactive, and decision-making behaviors of annotation personnel at different times in similar historical annotation tasks;

[0083] Based on the historical annotation accuracy of similar annotation tasks, the behavior weights of different operation behaviors are dynamically allocated according to the operation type.

[0084] A behavior feature matrix is ​​constructed based on behavior weights and operational behaviors. Operation patterns are clustered using density clustering algorithms, and abnormal operational behaviors are eliminated to extract the behavior features of different operation patterns.

[0085] By adjusting the weights of the encoding dimensions based on the domain features of the annotation task, the behavioral features are mixed-encoded and normalized to generate an operational behavior vector consistent with the domain feature dimensions.

[0086] The operational behavior of audit annotation tasks reflects the professional habits and efficient annotation logic of the annotators. Operations at different time points can fully reflect the entire annotation process. Extracting only a single operation or a fragment of operation cannot capture the behavioral patterns in the annotation process, and the behavioral features constructed later will be one-sided. In the operation log, the operation behavior data of similar historical annotation tasks are fully acquired. The acquisition process records each operation step in chronological order to ensure the chronological integrity of the operation behavior. Basic operations include basic interactive actions such as mouse clicks, text selection, cursor positioning, and content copying in the annotation process. The focus is on acquiring basic actions directly related to audit annotation, and irrelevant system operations are eliminated. Interactive operations cover the function call behavior of annotation tools, such as calling audit annotation templates, enabling the quick division function of aging intervals, triggering data verification tools, and linking general ledger and subsidiary ledger data. The timing and frequency of function calls are recorded simultaneously.

[0087] The decision-making process focuses on professional judgment behaviors in audit annotation, including determining the aging range of accounts receivable, determining the provision ratio for bad debts, marking suspicious data, and verifying the annotation results. It also records the preceding and subsequent adjustment actions corresponding to these decision-making behaviors. After acquisition, the actions are sorted and organized by timestamp, and duplicate records and invalid operations triggered by system anomalies are removed from the operation logs to form time-series data of the actions. This comprehensively captures the operational habits and professional decision-making logic of the audit annotation personnel, avoiding the one-sidedness and fragmentation of the action data. The structured time-series action data provides standardized input for subsequent action weight allocation and action pattern extraction, ensuring that subsequent processing is based on real and complete action data. At the same time, eliminating invalid operations reduces data redundancy and improves subsequent processing efficiency.

[0088] Different types of operational behaviors contribute significantly to the accuracy of audit annotation. Basic operations are mostly auxiliary actions, decision-making operations directly affect the accuracy of annotation results, and interactive operations affect annotation efficiency and error rate. If a fixed weight allocation method is used, the role of high-value operations cannot be reflected, which will lead to subsequent behavioral characteristics being interfered with by invalid operations. By retrieving annotation result data of similar historical annotation tasks and combining it with the time series data of operational behaviors, a correlation between operation type and historical annotation accuracy is established. The actual impact of operational behaviors on annotation accuracy in similar historical annotation tasks is analyzed. If the annotation result accuracy rate of a certain type of decision-making behavior is high and the review and modification rate is low, it indicates that the operation contributes highly to accuracy and is assigned a higher behavioral weight. Decision-making behaviors include the determination of aging based on detailed account reconciliation.

[0089] For interactive operations, if calling a certain tool function significantly reduces the annotation error rate, the behavior weight of that type of interactive operation is increased. Tool functions include automatic account aging verification. For basic operations, only actions directly related to annotation accuracy are retained and assigned behavior weights. Basic operations without direct impact are assigned low behavior weights or even eliminated, including arbitrary cursor movement. During the allocation of behavior weights, adjustments are made according to operation type to ensure that the differences in behavior weights within the same type of operation reflect actual contribution. The behavior weights for different types of operations are adapted to the core needs of audit annotation, forming dynamically adjusted behavior weights that are synchronously associated with corresponding operational behaviors. This achieves differentiated allocation of operational behavior weights, highlighting the core value of decision-making operations and efficient interactive operations, weakening the interference of ineffective basic operations. Weight allocation is based on historical annotation accuracy, ensuring that behavior weight settings reflect actual annotation effects, providing accurate weight basis for subsequent construction of the behavior feature matrix, and improving the effectiveness and relevance of behavior features.

[0090] Scattered operational behaviors and their corresponding weights cannot be directly used to generate operational behavior vectors. They need to be integrated into structured data through a behavioral feature matrix. Furthermore, the actions of the annotators exhibit repetitive patterns and abnormal movements. Directly using the raw operational data would result in redundant behavioral features and the inclusion of interfering information. A behavioral feature matrix is ​​constructed using the organized operational behaviors as row vectors and their corresponding weights as column vectors. During the construction of the row matrix, the operational behaviors are aligned dimensionally according to time sequence to ensure accurate association of the weights corresponding to each operational behavior. Additionally, attribute features of the operational behaviors are supplemented, such as operation duration and triggering conditions, to enrich the matrix dimensions. Then, a density clustering algorithm is used to cluster the operational behaviors in the behavioral feature matrix. During clustering, the similarity and weights of the operational behaviors are used as the core criteria to aggregate operational sequences with the same patterns, forming different operational modes.

[0091] For example, the sequential operations of "retrieving the aging template → selecting the corresponding details → determining the aging range → triggering the verification function" are aggregated into a type of efficient annotation pattern; actions that deviate from normal annotation patterns, such as repeated clicks, misselected text, and meaningless function calls, are identified as abnormal operations and eliminated; after clustering, for each type of operation pattern, the core operation behavior and corresponding behavior weights are extracted, the temporal logic and key nodes of the operation pattern are sorted out, and the behavioral features of the operation pattern are formed to ensure that the behavioral features cover all types of efficient annotation patterns; thus, the structured integration of operation behavior and behavior weights is achieved through the behavioral feature matrix, providing standardized input for cluster analysis. Density clustering effectively refines the operation pattern, eliminates interference from abnormal operations, avoids redundancy of behavioral features, and improves the representativeness and stability of behavioral features. The extracted operation patterns are in line with the efficient habits of audit annotators, providing high-quality behavioral features for subsequent coding processing.

[0092] Operational behavior features must be precisely matched with the domain features extracted earlier; otherwise, the generated operational behavior vectors will not meet the core requirements of the audit annotation task. Adjusting the weight of the encoding dimension can strengthen operational behavior features related to the domain features and weaken irrelevant features. Hybrid encoding and normalization can integrate multi-dimensional behavioral features, avoid dimensional imbalance, and ensure that the generated operational behavior vectors are consistent with the domain feature dimensions, meeting the requirements of subsequent joint guidance vector fusion. Retrieve domain features and sort out the relationship between domain features and operational behavior features to clarify the key points of adaptation. For example, the aging interval determination in the domain features corresponds to the aging determination decision operation in the operational mode, and the bad debt provision corresponds to the relevant annotation and verification operation. For these related dimensions, increase the weight of the encoding dimension of the corresponding operational behavior feature to strengthen the domain adaptability. For operational behavior features that are weakly related to the core audit requirements, appropriately decrease the weight of the encoding dimension.

[0093] The encoding process employs a hybrid encoding approach, integrating the temporal features, behavioral weights, and operational patterns of the operational behaviors. It retains the core logic and behavioral weights of the operational patterns while ensuring consistency between the encoding dimensions and the domain feature dimensions, achieving dimensional alignment between the two types of features. After encoding, normalization is performed to balance the weights of each encoding dimension, preventing overemphasis on any single dimension and eliminating representational differences across dimensions. Ultimately, an operational behavior vector adapted to the audit accounts receivable labeling task is generated. This vector encompasses both the efficient operational patterns of the labelers and precise adaptation to domain features. This ensures accurate adaptation between the operational behavior vector and domain features, guaranteeing that the operational behavior vector aligns with task requirements. Hybrid encoding and normalization eliminate dimensional imbalances and representational differences, improving the stability and universality of the operational behavior vector. This provides high-quality input for the subsequent fusion of joint guidance vectors, ensuring that the fused vector accurately guides keyword retrieval.

[0094] like Figure 2 As shown, specifically, generating the joint guiding vector includes:

[0095] Retrieve the preset domain ontology knowledge base, extract the quantitative value of the feature importance of the current annotation task's domain, and assign domain adaptation weights to the domain features according to the proportion of the quantitative value;

[0096] Select a small sample of data to be labeled for the current labeling task, complete the small sample labeling based on the operation behavior vector, and determine the effective weight of the operation behavior vector based on the degree of fit between the labeling results and the standard labeling.

[0097] Using domain adaptation weights and effective behavior weights as fusion coefficients, we weight and fuse domain features and operational behavior vectors with consistent dimensions to generate a joint guidance vector.

[0098] Domain features encompass multiple dimensions, each with varying guiding value for audit annotation tasks. Applying equal weights would weaken the guiding role of core domain features, causing the subsequent fusion of joint guidance vectors to deviate from audit annotation requirements. A pre-defined audit domain ontology knowledge base is retrieved. This knowledge base is built upon financial audit industry standards and historical annotation experience, containing feature concepts, hierarchical relationships, and importance judgment rules related to accounts receivable annotation, covering core features such as aging, bad debt provision, and suspicious data marking. A quantitative value of the feature importance of the current annotation task's domain is extracted from the domain ontology knowledge base, based on the feature's core importance and scope of influence in audit annotation.

[0099] Domain-adaptive weights are assigned to each domain feature based on its quantified value proportion, prioritizing the increase in the weight proportion of core features, such as features related to aging interval determination and bad debt provision ratio, which directly affect the accuracy of audit annotation results and are assigned higher domain-adaptive weights. Auxiliary domain features, such as annotation text format specifications, are assigned lower domain-adaptive weights. After weight allocation, fine-tuning is performed based on the specific scenario of the current accounts receivable annotation task to ensure that the weight proportions adapt to the task requirements, forming domain-adaptive weights and associating them with corresponding domain features. This relies on the domain ontology knowledge base to achieve precise allocation of domain-adaptive weights, strengthen the guiding role of core features, avoid feature imbalance caused by equal weights, and ensure that the weight allocation conforms to the professional rules of the audit domain, ensuring that domain features can accurately reflect the core needs of audit annotation, providing high-quality domain dimension input for subsequent joint guidance vector fusion, and improving the domain adaptability of the joint guidance vector.

[0100] Although the operational behavior vectors are constructed based on historical labeled data, their actual adaptability in the current labeling task needs to be verified. The details of different audit labeling tasks vary, and historical operational behaviors may not be fully adaptable to the current labeling task. The behavior weights determined solely by historical labeled data will have biases. A small sample of data to be labeled for the current accounts receivable labeling task should be selected. The sample should cover typical scenarios such as different aging intervals, the existence of bad debt risks, and whether the data is suspicious, to ensure that the sample is representative and can fully verify the adaptability of the operational behavior vectors. The small sample labeling is completed based on the operational behavior vectors. During the labeling process, the efficient operation modes and decision logic contained in the vectors are reused, such as the operation sequence for determining the aging of accounts receivable and the function call method for data verification. After the labeling is completed, the small sample labeling results are compared with the preset standard labeling results to analyze the labeling fit.

[0101] The labeling fit is reflected in the logical consistency, accuracy, and completeness of the labeling results and the standard results. The focus is on whether the labeling results guided by the operational behavior vectors conform to auditing professional rules. The effective weights of the operational behavior vectors are determined based on the labeling fit. If the labeling results guided by a certain type of operational behavior vector have a high fit with the standard results, it indicates strong adaptability and is assigned a higher effective weight. If the labeling fit is low, it indicates a deviation in the operational behavior vector, and the effective weight is reduced, and the operational mode that needs optimization is marked. This ultimately forms the effective weights of the operations that fit the current labeling task. By verifying the actual adaptability of the operational behavior vectors through small-sample labeling, deviations caused by historical labeling data are avoided, ensuring that the effective weights of the operations fit the current audit labeling task. The determination of effective weights based on the labeling fit between the results and the standard achieves dynamic calibration of the operational behavior vectors, strengthens the role of efficient operational modes, weakens the impact of poorly adaptable operational behaviors, and improves the practicality of the operational behavior vectors.

[0102] Domain features and operational behavior vectors provide guidance from two dimensions: domain attributes and operational habits, respectively. Using either vector alone has limitations; domain features alone lack operational guidance, while operational behavior vectors alone are prone to deviating from core domain requirements. Domain adaptation weights and effective behavior weights are used as fusion coefficients to weight and fuse domain features and operational behavior vectors with consistent dimensions. During fusion, the core features of both types of vectors are retained according to the fusion coefficients, taking into account both audit domain attributes and the operational habits of the annotators. For example, domain features for aging determination are prioritized for fusion with corresponding efficient operation mode features to strengthen the core guidance logic. Strict consistency of vector dimensions is ensured during fusion. Based on the processed domain features and operational behavior vectors, features from different dimensions are integrated accordingly to avoid fusion failure due to dimension misalignment.

[0103] After fusion, normalization is performed to eliminate representational differences between different dimensions of features, balance the weight ratio of the two types of vectors, and avoid over-emphasis on single-dimensional features. The resulting joint guidance vector embodies both the domain requirements for auditing accounts receivable annotation and the efficient operational logic of the annotators, providing precise guidance for subsequent keyword retrieval. This achieves deep fusion of domain features and operational behavior vectors, balancing domain expertise with operational efficiency, overcoming the limitations of single-vector guidance. Normalization ensures consistent vector dimensions and balanced weights, improving the stability and universality of the joint guidance vector. It can accurately guide knowledge graph retrieval of keywords suitable for the current annotation task, indirectly improving annotation efficiency and accuracy.

[0104] S2. Based on the joint guiding vector, retrieve prompt words from the prompt word knowledge graph, process the data to be labeled and prompt words to output the initial labeling results and confidence level, and analyze the initial labeling results to identify the differences in the domain distribution of labeling defect types and related prompt words.

[0105] Furthermore, search suggestions include:

[0106] The prompt word knowledge graph is preprocessed by classifying the prompt words according to their respective domains and quantifying them into domain category vectors, and then associating them with the operation behavior tags of the personnel labeled in the corresponding domains.

[0107] Extract the corresponding domain component from the joint guidance vector, perform cosine similarity coarse matching with the domain category vector of the prompt word knowledge graph, and filter out candidate prompt words belonging to the domain of the current annotation task;

[0108] Extract the corresponding operation components from the joint guidance vector, calculate the matching degree between the operation behavior label and the operation component in the candidate prompt words, retain the prompt words with a matching degree greater than the matching threshold, and obtain the usable candidate set;

[0109] Based on the fusion coefficient of the joint guiding vector, the prompt words in the available candidate set are weighted and sorted. The number of prompt words is preset based on the annotation complexity of the current annotation task, and the prompt words are output according to the sorting results.

[0110] The prompt words stored in the prompt word knowledge graph cover multiple fields and types of content, resulting in cross-domain interference, semantic ambiguity, and a lack of operational relevance. If used directly for retrieval, the prompt words will have poor adaptability to the current audit annotation task and will not be able to accurately guide the annotation operation. The preset prompt word knowledge graph is preprocessed throughout the entire process. First, the prompt words are classified and sorted according to field categories. Based on audit industry standards and task requirements, the prompt words are divided into audit sub-fields such as accounts receivable, accounts payable, and fixed assets. At the same time, cross-domain prompt words that are not related to auditing are removed to ensure that the field category division fits the current task scenario.

[0111] For the categorized accounts receivable domain prompts, natural language processing technology is used to quantify them into domain category vectors. The quantification process focuses on the semantic association between prompts and audit features, retaining core semantic information such as aging determination, bad debt provision, and suspicious data marking, to achieve a structured representation of the semantics of prompts. Then, the operation behavior tags of the personnel marked under the corresponding domain are associated. The operation behavior tags are based on the three types of operation behaviors extracted above, covering tags related to basic operations, interactive operations, and decision-making operations.

[0112] For example, the prompts for dividing the aging period are associated with the action tags "select details → determine the period → trigger verification," and the prompts for bad debt provision are associated with the action tags "retrieve provision rules → verify data → confirm ratio." This ensures that each prompt corresponds to a clear action guide. The final preprocessed prompt set has both domain attributes and action-related characteristics. This effectively eliminates cross-domain interference prompts, achieves accurate domain classification and semantic standardization of prompts, avoids retrieval bias caused by ambiguous prompts, and the association of action tags makes the prompts not only domain-guided but also adaptable to the operational habits of the labelers, improving the usability of subsequent retrieval prompts and providing high-quality input for accurate matching.

[0113] Even after preprocessing, there are still a large number of prompts. Direct fine-grained matching would increase computation and be inefficient. By using coarse matching of domain components, prompts that match the current task domain attributes can be quickly filtered out, narrowing the search scope and eliminating prompts with poor domain adaptability. This allows for precise matching at the subsequent operation level, focusing on the target and balancing search efficiency and accuracy. The corresponding domain components are extracted from the joint guiding vector. These domain components are the product of domain adaptability weights and domain features, and they centrally carry the core domain features of accounts receivable labeling, including semantic information such as aging range, bad debt provision, and suspicious data.

[0114] The domain component is coarsely matched with the domain category vector using cosine similarity. Similarity matching focuses on the semantic fit between the two domains, prioritizing the selection of prompts closely related to features such as aging determination and bad debt provision. After matching, filtering conditions are set to remove prompts with low similarity or irrelevant to accounts receivable labeling, retaining prompts that meet the domain requirements to form candidate prompts. This quickly narrows the scope of prompt search, significantly reducing the computational load of subsequent fine matching and improving search efficiency. By filtering based on semantic fit, it ensures that candidate prompts are highly compatible with the domain attributes of the current audit task, eliminating cross-domain and low-relevance prompts, and providing high-quality candidate objects for accurate matching at the subsequent operational level.

[0115] While candidate prompts possess domain adaptability, the adaptability of the operational behavior labels corresponding to different prompts to the operational behavior vector of the current task varies. Domain matching alone cannot guarantee that prompts can guide efficient annotation operations. The corresponding operational components are extracted from the joint guidance vector. These operational components are the product of operational behavior features and effective weights, integrating the efficient operational patterns and decision-making logic of audit annotators, covering the core features of three types of operational behaviors. For each prompt in the candidate prompts, its associated operational behavior label is extracted, and the matching degree between this operational behavior label and the operational component of the joint guidance vector is calculated. The matching degree judgment focuses on the consistency of the operational sequence and the fit of the operational logic, with a key focus on verifying whether the operational behavior label corresponding to the prompt is consistent with operational patterns such as aging determination and bad debt provision.

[0116] If the operation tag associated with the prompt word is "quickly select details → automatically match aging intervals → trigger data verification", and it is consistent with the efficient operation mode in the operation component, then the matching degree meets the requirements. If the operation tag associated with the prompt word deviates significantly from the current operation logic, or only contains invalid basic operations, then the matching degree is insufficient. Based on the adaptation requirements of audit annotation operations, a matching threshold is set, prompt words with a matching degree greater than the matching threshold are retained, and prompt words with poor adaptability are eliminated to form a usable candidate set, which has both domain adaptability and operation adaptability. This achieves accurate screening of prompt words from domain adaptability to operation adaptability, ensuring that usable candidate prompt words can not only meet the needs of accounts receivable audit, but also guide annotators to follow efficient operation habits, avoid the decrease in annotation efficiency caused by the disconnect between prompt words and operation logic, and enhance the practical guidance value of prompt words.

[0117] Within the available candidate set, different prompts offer varying guiding value for the annotation task. Some prompts correspond to core annotation steps, while others are only used for auxiliary operations. Furthermore, the complexity of different audit annotation tasks varies, necessitating dynamic adjustments to the required number and priority of prompts. Weighted sorting strengthens the priority of core prompts, and pre-setting the number based on complexity ensures that the output prompts align with the actual needs of the task, avoiding redundant prompts from interfering with the annotation process. The fusion coefficient of the joint guidance vector serves as the weighting basis, encompassing domain adaptation weight and behavioral effectiveness weight. The weighting process considers both the domain importance and operational adaptability of the prompts. Core prompts with high domain adaptation weights and high compatibility with efficient operation modes, such as those related to accurate determination of aging intervals and confirmation of bad debt provision ratios, are assigned higher weights, increasing their ranking priority. Auxiliary prompts, such as those related to annotation format specifications and data saving reminders, are assigned lower weights, reducing their ranking priority.

[0118] Available candidate prompts are sorted in descending order based on weighted scores to ensure that core prompts are prioritized. The number of prompts is determined based on the complexity of the current accounts receivable labeling task. If the task involves complex aging classification and multi-scenario bad debt determination, resulting in high labeling complexity, the number of output prompts is appropriately increased to cover core and auxiliary stages. If the task scenario is simple, involving only routine accounts receivable detail labeling, the number of output prompts is reduced, retaining only core prompts. Prompts are output based on the sorting results and the preset number of prompts to provide precise guidance for subsequent labeling tasks. This weighted sorting clarifies the priority of prompts, ensuring that core prompts guide the labeling of core stages first, improving labeling efficiency and accuracy. The number of output prompts is dynamically adjusted according to labeling complexity to avoid redundant prompts or missing core prompts, achieving personalized adaptation of prompt output and allowing prompts to accurately match audit labeling tasks of varying complexity.

[0119] Specifically, the output of the initial annotation results and confidence levels includes:

[0120] The output prompts are classified according to the proportion of the fusion coefficient. Prompts with a fusion coefficient greater than the preset proportion are set as domain prompts and operation prompts, respectively.

[0121] Perform domain semantic annotation on the data to be annotated based on domain prompt words and domain features, and perform operation adaptation annotation based on operation prompt words and operation behavior vectors to generate domain annotation results and operation annotation results;

[0122] Calculate the degree of fit between the two types of annotation results. If the degree of fit is greater than the degree of fit threshold, merge them into the initial annotation result. Otherwise, call the preset annotation standard fragment to correct the divergent parts and generate the initial annotation result.

[0123] Based on the consistency of the generated results, the matching degree of the prompt words and the fit of the operation annotation results, the confidence of the prompt word adaptability and the operation association is calculated and then weighted and summed to output the confidence of the initial annotation results.

[0124] Although the output prompts are filtered and sorted, they encompass both domain-oriented and operation-oriented attributes. Different prompts have different core functions for annotation. If they are not graded and distinguished, direct use for annotation will lead to confusion between domain semantic guidance and operation process guidance, and will not accurately adapt to the different dimensions of annotation needs. The fusion coefficient of the joint guidance vector is retrieved, and the ratio of domain adaptation weight to behavior effectiveness weight in the fusion coefficient is used as the grading basis. The preset ratio standard adapts to the audit annotation needs and is used to distinguish the functional focus of the prompts. The ratio of the fusion coefficient of each output prompt is calculated. Prompts with a domain adaptation weight ratio greater than the preset ratio are set as domain prompts. These prompts focus on the core semantics of accounts receivable audit, such as the definition of aging range, bad debt provision rules, and suspicious data judgment criteria. Their core function is to ensure the domain professionalism and semantic accuracy of the annotation results.

[0125] Prompt words with a proportion of effective weight exceeding a preset proportion are designated as operation prompt words. These prompt words correspond to efficient operation processes for annotators, such as quick selection of detailed text, automatic verification of aging results, and standardized saving of annotation content. Their core function is to guide annotation operations to conform to an efficient mode. After classification, the two types of prompt words undergo secondary verification to ensure no overlap or confusion. Prompt words with critical proportions and ambiguous functions are removed, resulting in two clearly defined types of prompt words. This achieves precise functional classification of prompt words, clarifies the annotation guidance of different prompt words, avoids conflicts between domain semantics and operation process guidance, and allows subsequent annotation stages to have different focuses. This ensures both the professional semantic accuracy of audit annotations and conforms to the operating habits of annotators, improving the orderliness and efficiency of the annotation process.

[0126] Single-dimensional annotation cannot simultaneously ensure both audit domain professionalism and operational efficiency. Annotation based solely on domain-specific keywords, while guaranteeing semantic accuracy, lacks operational process guidance and is inefficient. Annotation based solely on operational keywords, while adhering to efficient processes, is prone to deviating from domain semantic norms, leading to annotation errors. This approach performs domain semantic annotation on the accounts receivable details to be annotated, based on domain-specific keywords and domain characteristics. During the annotation process, domain-specific keywords serve as semantic guidance, and the core dimensions of accounts receivable domain characteristics are combined to verify the detailed data one by one, determining the aging range, whether bad debt provisions have been made, and whether the data has been marked as suspicious. This ensures that the annotation results comply with audit industry standards and task requirements, generating domain-specific annotation results that emphasize semantic accuracy and professional compliance.

[0127] Simultaneously, based on operation prompts and operation behavior vectors, operation-adaptive annotations are performed. Guided by the operation prompts and following the efficient operation patterns implied by the operation behavior vectors, operations such as selecting detailed data, entering annotation content, and triggering verification are completed. This ensures that each step conforms to the habitual workflow of the annotators, generating operation annotation results that highlight the standardization of operations and the efficiency of the process. The two types of annotations are performed in parallel, recording the annotation process and results separately to ensure data integrity and traceability. The dual-dimensional annotation achieves a complementary integration of audit professional semantics and efficient operation processes, avoiding the limitations of a single annotation dimension while ensuring that the annotation results conform to industry standards and adapt to practical habits. The generation of the two types of annotation results provides a clear object for subsequent consistency verification and provides dual guarantees for improving the quality of the final annotation results.

[0128] Discrepancies may exist between domain annotation results and operational annotation results, possibly stemming from deviations between semantic guidance and operational guidance. Direct merging could lead to contradictory annotation results, affecting accuracy. The degree of fit between domain and operational annotation results is calculated, focusing on logical consistency, content completeness, and semantic accuracy. Key checks include consistency in core content such as aging interval determination, bad debt provision accrual, and suspicious data marking. Simultaneously, it verifies whether operational annotation results conform to domain semantic specifications and whether they fit operational process requirements. A fit threshold is set based on the consistency requirements of audit annotations. If the fit between the two types of annotation results exceeds the threshold, it indicates good consistency, and the two types of annotation results are directly merged, retaining the core consistent content to form the initial annotation result.

[0129] If the degree of fit is not greater than the degree of fit threshold, it indicates a discrepancy. A pre-defined standard annotation fragment is retrieved. This standard fragment is constructed based on audit industry standards and historical annotation experience, covering the core rules and discrepancy handling solutions for accounts receivable annotation. The discrepancies are analyzed one by one against the pre-defined standard fragment. Combining semantic guidance from domain prompts and process guidance from operational prompts, the discrepancies are corrected. Priority is given to ensuring that the corrected annotation results comply with audit professional standards, while also considering operational adaptability. After correction, an initial annotation result is generated, simultaneously recording the discrepancy points and the basis for correction. Thus, by judging the degree of fit, the two types of annotation results are accurately integrated and discrepancies are corrected, avoiding the inclusion of contradictory content in the final annotation result. This ensures the accuracy and consistency of the initial annotation result. The pre-defined standard fragment provides a professional basis for discrepancy correction, avoiding deviations caused by subjective corrections, while also considering professional standards and operational adaptability, making the initial annotation result both compliant and practically feasible.

[0130] Although the initial annotation results have been fused or corrected, their reliability still needs to be quantified to provide a basis for subsequent defect analysis and iterative correction. Confidence scores are generated based on three types of indicators: result consistency, prompt word suitability, and operation association. Result consistency is determined by the degree of fit between two types of annotation results; the higher the fit, the better the result consistency index, reflecting the internal stability of the initial annotation results. Prompt word suitability is determined by the matching degree between prompt words and the annotation task, focusing on verifying the semantic suitability of domain prompt words and the procedural suitability of operation prompt words; the stronger the suitability, the better the prompt word suitability index. Operation association is determined by the degree of fit between the operation annotation results and the operation behavior vector, reflecting the fit between the annotation operation and the efficient mode.

[0131] The final confidence score is generated by weighted summation of three types of indicators. The weight allocation aligns with audit annotation requirements, prioritizing result consistency and prompt word suitability to ensure that the confidence score truly reflects the credibility of the initial annotation results. After weighted calculation, the initial annotation results and their corresponding confidence scores are output, synchronously linked to the evaluation criteria of various indicators to ensure traceability of the confidence score. This achieves a quantitative assessment of the credibility of the initial annotation results, providing clear judgment criteria for subsequent processes. The confidence score comprehensively considers multiple dimensions of indicators, avoiding the one-sidedness of single indicator evaluation, and can fully reflect the stability, suitability, and relevance of the annotation results. At the same time, it provides precise direction for subsequent screening of low-confidence results, prompt word optimization, and operational process adjustments.

[0132] like Figure 3 As shown, the differences in the domain distribution of identified defect types and associated prompts include:

[0133] Based on the confidence level of the initial annotation results and the discrepancies between the two types of annotation results, the discrepancies with confidence levels less than the confidence threshold are filtered out, and the discrepancies are classified according to their sources to identify the types of annotation defects.

[0134] The annotation defect type is bound to the domain prompt words and operation prompt words that generate the initial annotation result, and the frequency of prompt words with different proportions of fusion coefficient triggering the corresponding annotation defect type is statistically analyzed;

[0135] Based on the subdivision of the domain to which the current annotation task belongs, the defect induction rate of the same prompt word under different subdivisions is calculated based on the frequency, and the distribution ratio of annotation defect types among different subdivisions is statistically analyzed to obtain the domain distribution differences of the associated prompt words corresponding to the annotation defect types.

[0136] In the initial annotation results, the consistency and accuracy of the high-confidence portions have been verified and do not require in-depth analysis. However, low-confidence results are often accompanied by annotation discrepancies, which are high-risk areas for defects. If the low-confidence discrepancies are not focused on, the defect analysis will be generalized, making it impossible to accurately pinpoint the root cause of the problem. The confidence levels of the initial annotation results and the discrepancy records of the two types of annotation results are retrieved. The set confidence threshold is adapted to the reliability requirements of audit annotation. The discrepancies with confidence levels lower than the confidence threshold are selected as the core objects of defect analysis, and high-confidence undisputed content is removed to reduce redundancy. Then, the discrepancies are classified according to their sources. Based on the audit annotation scenario, three core sources of discrepancies and their corresponding annotation defect types are identified. The first is domain semantic discrepancy, which stems from semantic guidance deviations of domain prompt words or misunderstandings of domain features. The corresponding annotation defect type is semantic deviation defect, such as confusion in the aging interval determination criteria and misunderstanding of bad debt provision rules.

[0137] Second, there are operational process discrepancies, stemming from improper guidance in operational prompts or mismatches in operational behavior vectors. These are labeled as non-standard operational defects, such as disordered operational order, omissions in verification function calls, and data association errors. Third, there are mixed discrepancies, involving both semantic and operational issues. These are labeled as mixed defects, such as operational process misalignment due to semantic deviations and semantic labeling errors caused by non-standard operations. After classification, the discrepancy characteristics and manifestations of each labeled defect type are labeled to form a structured labeled defect type, ensuring clear and traceable defect identification. This accurately identifies high-risk defect areas, avoids inefficient generalization of defect analysis, clarifies the essence and manifestation of defects by source, eliminates confusion of defect types, and provides standardized input for subsequent binding with prompts and frequency statistics, ensuring that defect analysis focuses on core issues and improves the overall relevance and efficiency of the analysis.

[0138] Defects are often related to biases in prompt words. The types of labeled defects caused by prompt words with different fusion coefficient ratios vary. If the defect type is not bound to the prompt word or the frequency of occurrence is not statistically analyzed, it is impossible to determine the cause of the defect and the probability of the prompt word causing the defect. Subsequent analysis of domain distribution differences will lack data support and it will be impossible to accurately locate high-risk prompt words. The identified defect types are bound one by one to the domain prompt words and operation prompt words used to generate the initial labeling results. During the binding process, the core prompt words that induce the defect are accurately matched by combining the divergence characteristics of the defect type and the guiding role of the prompt words. For example, the semantic deviation defect of the aging interval is bound to the corresponding domain prompt word, the disordered operation sequence defect is bound to the corresponding operation prompt word, and the mixed defect is bound to both types of prompt words at the same time.

[0139] The prompts are categorized according to the proportion of their fusion coefficients, distinguishing between three types: those with a high proportion of domain-adaptation weight, those with a high proportion of behavioral effectiveness weight, and those with a balance between the two. The frequency of each type of prompt triggering its corresponding defect type is then statistically analyzed. During the statistical analysis, the proportion of the fusion coefficient, the labeled defect type, and the corresponding divergent content of the prompts are recorded simultaneously. Defects caused by subjective errors of the labelers (i.e., those not triggered by the prompts) are eliminated to ensure that the frequency data reflects the true correlation between the prompts and defects, forming a correlation data between the proportion of prompts, fusion coefficients, labeled defect types, and frequencies. This establishes a precise correlation between prompts and labeled defect types, clarifies the defect-inducing patterns guided by the weight proportions of different fusion coefficients, and identifies prompts that frequently induce defects as optimization priorities. The correlation data provides data support for subsequent calculations of defect induction rates and analysis of domain distribution differences, preventing distribution difference analysis from being detached from actual causes and improving the practicality of the analysis results.

[0140] The sensitivity of defects varies across the domain subdivisions of audit annotation tasks. The probability of the same prompt word triggering a defect differs across different subdivisions. Simply counting the overall trigger frequency cannot reveal the distribution pattern of defects within the domain, nor can it accurately pinpoint high-risk subdivisions and their corresponding prompt words. This analysis breaks down the scope of the current accounts receivable annotation task into subdivisions. Core subdivisions include aging interval determination, bad debt provision ratio confirmation, suspicious data marking, and data correlation verification. Each subdivision corresponds to a core aspect of audit annotation. Based on correlated data, under each subdivision, the defect trigger rate of the same prompt word across different subdivisions is calculated by combining the frequency of defects triggered by the prompt word with the total annotation volume for that subdivision. The defect trigger rate reflects the relative probability of a prompt word triggering a defect in a specific dimension.

[0141] Simultaneously, the distribution ratio of various labeled defect types across different sub-dimensions is statistically analyzed to clarify the main defect types and their proportions under each sub-dimension. For example, the aging interval dimension is dominated by semantic deviation defects, while the data association verification dimension is dominated by operational non-standard defects. By combining the defect induction rate and distribution ratio, the domain distribution differences of associated prompts corresponding to labeled defect types are identified, clarifying high-risk sub-dimensions, corresponding high-induction-rate prompts, and core defect types. The distribution characteristics and causes of differences in each dimension are recorded simultaneously. This reveals the distribution patterns of defects and prompts in each sub-dimension of the audit field, accurately locates high-risk dimensions and high-induction-rate prompts, avoids blind optimization work, and clearly presents the defect characteristics and causes of different sub-dimensions. This provides targeted directions for subsequent iterations to correct knowledge graph parameters and adjust fusion coefficients, while also providing data support for optimizing the audit labeling process and improving the domain ontology knowledge base.

[0142] S3. Filter the initial annotation results with confidence scores less than the confidence threshold, determine the reward function based on annotation accuracy and domain adaptability, use the difference in domain distribution between annotation defect types and associated prompt words as the state space, and iteratively correct the parameters of the prompt word knowledge graph through reinforcement learning.

[0143] Specifically, the parameters for iteratively refining the prompt word knowledge graph include:

[0144] Initial annotation results with confidence scores below the confidence threshold are selected, and a reward function is determined based on annotation accuracy and domain adaptability. The parameters of the reward function are then dynamically adjusted according to differences in domain distribution.

[0145] The differences in domain distribution between labeled defect types and associated prompt words are quantitatively represented as a state space. Through reinforcement learning, the parameters of the prompt word knowledge graph are hierarchically corrected according to the different proportions of the frequency of the labeled defect type corresponding to the prompt word and the fusion coefficient.

[0146] After each round of correction, prompt words are retrieved again based on the corrected prompt word knowledge graph and labeled. The cumulative score of the reward function is calculated. If the cumulative score does not reach the convergence threshold, the state space is updated based on the new domain distribution difference until the iteration tends to stabilize.

[0147] In the initial annotation results, the low-confidence portion corresponds to high defect risk and is the core object of iterative optimization. If the reward function only focuses on accuracy, it is easy to deviate from audit standards. If it only focuses on domain adaptability, it will ignore operational defects. At the same time, the domain distribution differences have clearly defined the distribution patterns and causes of defects in each subdivision. If the reward function parameters are fixed, it is impossible to optimize high-risk dimensions in a targeted manner. The initial annotation results are retrieved, and the initial annotation results with confidence levels lower than the preset confidence threshold are selected. These results are often accompanied by defects such as semantic deviation and operational non-standardization, and correspond to clear domain distribution characteristics. Based on the annotation accuracy of similar historical tasks, the reward function is constructed in combination with the audit domain adaptability requirements of the current task. The annotation accuracy is based on the logical consistency with the standard results, while the domain adaptability is based on whether it complies with the accounts receivable audit standards and whether it is suitable for core links such as aging determination and bad debt provision.

[0148] Then, the parameters of the reward function are adjusted according to the differences in domain distribution. If the proportion of semantic defects in the aging interval judgment dimension is high, the reward weight related to semantic accuracy guidance in that dimension is increased. If the operational defects in the data association verification dimension are concentrated, the corresponding reward parameters are strengthened to adapt the operation process, guiding the optimization focus to high-risk dimensions. During the adjustment process, the fusion coefficient ratio of the associated prompt words is simultaneously adjusted to ensure that the parameter adjustment matches the functional positioning of the prompt words, forming a reward function that adapts to the current defect distribution. This accurately identifies high-risk optimization targets, avoiding the waste of optimization resources on high-confidence, high-quality results. The reward function parameters are dynamically adjusted in combination with the differences in domain distribution, ensuring that the optimization direction conforms to audit professional standards and can specifically solve frequently occurring defects, eliminating the blind optimization caused by fixed parameters, and providing a scientific reward guide for subsequent reinforcement learning.

[0149] The difference in domain distribution is the core manifestation of the correlation between labeled defect types and prompt words. It needs to be quantitatively represented and transformed into a state space that reinforcement learning can handle. Otherwise, it cannot provide clear environmental feedback for the correction of knowledge graph parameters. At the same time, prompt words with different fusion coefficient ratios cause different labeled defect types and frequencies. If the parameters are uniformly corrected globally, it will weaken the optimization effect of core prompt words and lead to an imbalance in knowledge graph parameters. The difference in domain distribution between labeled defect types and associated prompt words is quantitatively represented. The quantification process focuses on the defect characteristics of each audit subdivision, the frequency of prompt word induction, and the proportional correlation of fusion coefficients. It does not involve specific numerical calculations, but only transforms it into a state space that reinforcement learning can recognize through semantic structuring, clearly presenting the correspondence between subdivisions, labeled defect types, prompt words, and induction patterns.

[0150] The state space is input into reinforcement learning, and combined with the constructed reward function, hierarchical correction is performed according to the ratio of the frequency of the corresponding defect to the fusion coefficient. For prompt words with a high domain adaptation weight ratio, the semantic guidance parameters are corrected in a key way to optimize the accuracy of core semantics such as aging determination and bad debt provision, and reduce semantic bias defects. For operation prompt words with a high effective weight ratio, the operation process guidance parameters are adjusted to standardize the operation association logic such as data selection and verification triggering, and reduce operation non-standard defects. For prompt words with a balanced fusion coefficient, the semantics and operation guidance parameters are verified simultaneously to take into account the adaptability of both.

[0151] During the correction process, auditing standards and historical optimization experience were referenced to avoid over-correction of reward function parameters, which could lead to new adaptation problems. The parameter adjustment logic and basis for each type of prompt word were recorded simultaneously. The quantified state space provides clear environmental feedback for reinforcement learning, ensuring that the correction of prompt word knowledge graph parameters is based on real defect patterns. Layered correction achieves accurate optimization of prompt word knowledge graph parameters, avoiding parameter imbalance caused by global uniform correction. At the same time, the correction process is deeply bound to the functional positioning and defect-inducing characteristics of prompt words, improving the adaptability of prompt word knowledge graph parameters.

[0152] A single parameter correction may result in over-adjustment or incomplete optimization. If applied directly to full annotation, it can lead to fluctuations in annotation performance. Based on the corrected prompt word knowledge graph, the entire process of prompt word retrieval, dual-dimensional annotation, and result fusion is re-executed. The annotation object remains the current accounts receivable details data to ensure consistency between the verification scenario and the original task. The cumulative score of the reward function is calculated simultaneously. The level of the cumulative score reflects the adaptation effect of the corrected prompt word knowledge graph, such as improved accuracy of core association annotation, domain adaptability, and reduced defect rate. A convergence threshold is set based on the optimization requirements of audit annotation to determine whether the iteration has reached a stable effect. If the cumulative score reaches the convergence threshold, it indicates that the parameter correction is effective, the prompt word knowledge graph adapts to the current task, and the iteration terminates.

[0153] If the convergence threshold is not reached, it indicates that there is still room for optimization. The new domain distribution differences generated in this verification are retrieved, the state space of reinforcement learning is updated, and the correction and verification process is repeated until the cumulative score reaches the target and the iteration tends to be stable. After stabilization, the parameters of the current prompt word knowledge graph are solidified, and the parameter adjustment trajectory and defect optimization effect during the iteration process are recorded simultaneously to provide a reference for subsequent similar audit tasks. This forms a closed-loop optimization of correction, verification, and update, effectively avoiding the problem of deviation in a single correction, ensuring that the prompt word knowledge graph parameters continuously adapt to the audit annotation requirements, and the parameters after iterative convergence can significantly reduce the defect induction rate, improve the semantic accuracy and operational adaptability of prompt words, and at the same time provide standardized and high-quality knowledge graph support for subsequent similar tasks, reducing the cost of repeated optimization.

[0154] S4. When the continuous iterations tend to stabilize, terminate the iterative optimization and output the labeled dataset. Based on the quality evaluation data of the iterative process, update the fusion coefficient of the joint guiding vector.

[0155] Specifically, the fusion coefficients for updating the joint guiding vector include:

[0156] When the continuous iterations tend to stabilize, the iterative optimization is terminated and the labeled dataset is output. Based on the quality assessment data of the iterative process, the labeling accuracy and defect induction rate are extracted, and the fusion coefficients of the joint guiding vectors are respectively associated to form a correlation.

[0157] Based on the correlation, the adjustment rules are determined in conjunction with the deviation direction of the quality assessment data. The fusion coefficients are updated based on the adjustment rules. The joint guidance vector is regenerated using a small sample of unlabeled data to complete the labeling with the search prompt words and monitor the labeling effect.

[0158] By comparing the annotation results with the quality benchmark when the iterations tend to stabilize, it is possible to determine whether the fusion coefficients of the joint guiding vector after the update should be solidified.

[0159] The stabilization of the prompt word knowledge graph during iteration only indicates that the graph parameters are adapted to the current annotation requirements. However, the fusion coefficient of the joint guiding vector is still an intermediate value in the iteration process and may not have reached the optimal adaptation state. If this fusion coefficient is directly used, it will lead to insufficient synergy between the joint guiding vector and the stabilized knowledge graph. When the cumulative score of the prompt word knowledge graph reaches the convergence threshold and tends to stabilize after multiple iterations, the iteration optimization is terminated and the current annotation dataset is output. At the same time, the quality assessment data of the entire iteration process is retrieved. Quality indicators including annotation accuracy and defect induction rate are extracted from the quality assessment data. Annotation accuracy reflects the logical consistency and audit compliance of the annotation results after iteration with the standard results. Defect induction rate reflects the overall situation of semantic deviation and non-standard operation caused by various prompt words.

[0160] The quality indicators are correlated with the fusion coefficients of the joint guidance vector to establish a relationship. The accuracy of annotation is primarily correlated with the effective weight of the behavior, as this weight directly affects the guidance effect of the operation prompts, thus determining the standardization of the annotation process and the accuracy of the results. The defect incidence rate is primarily correlated with the domain adaptation weight, as this weight dominates the semantic guidance accuracy of the domain prompts and is directly related to the high incidence of defects in the audit subdivisions, forming a correlation between the quality indicators and the fusion coefficients. This accurately captures the inherent correlation between the fusion coefficients and annotation quality, avoiding blindly adjusting the coefficients without considering actual results. The structured correlation provides a clear basis for subsequent adjustment rules, ensuring that fusion coefficient optimization focuses on core issues. Relying on data from the iterative process guarantees the reliability of the correlation, laying a solid foundation for fusion coefficient adjustment.

[0161] Different deviation directions of quality indicators correspond to different fusion coefficient adaptation issues. If a uniform adjustment rule is adopted, it cannot solve various problems. For example, low annotation accuracy is due to insufficient effective weight of behavior, and high defect induction rate is due to imbalance of domain adaptation weight. Directly applying the updated fusion coefficient to the full annotation carries risks, as it will cause a decline in the quality of the full annotation due to the deviation of the fusion coefficient. Based on the correlation, targeted adjustment rules are formulated in combination with the deviation direction of quality assessment data. If the deviation in annotation accuracy is due to non-standard operation process, it indicates that the effective weight of behavior is not adapted enough. The adjustment rule focuses on optimizing the effective weight of behavior, strengthening the fit with efficient operation mode, and fine-tuning the domain adaptation weight to help improve the semantic guidance accuracy. If the deviation in defect induction rate is concentrated in the audit sub-dimension, it indicates that the domain adaptation weight is not targeted enough. The adjustment rule focuses on increasing the proportion of domain adaptation weight in the core dimension, weakening the weight of auxiliary dimensions, and reducing semantic deviation defects.

[0162] The fusion coefficients are updated according to the adjustment rules to ensure that the weight allocation takes into account both the professionalism of the audit field and the adaptability of operations. Small sample data covering typical scenarios are selected for annotation, including complex aging classification and bad debt judgment in multiple scenarios. The joint guidance vector is regenerated based on the updated fusion coefficients. The annotation is completed by searching the stabilized prompt word knowledge graph. The semantic guidance accuracy, operational process adaptability and defect occurrence are monitored during the annotation process to form the annotation effect. The adjustment rules are aligned with the essence of quality indicator deviations to achieve precise optimization of the fusion coefficients. This avoids the imbalance of fusion coefficients caused by a one-size-fits-all adjustment. Small sample verification greatly reduces the risk of full-scale annotation, can quickly discover adaptation problems after coefficient updates, improve optimization efficiency, and ensure that the updated fusion coefficients can effectively improve the synergy between the joint guidance vector and the stabilized prompt word knowledge graph.

[0163] The effectiveness of small-sample validation after updating the fusion coefficients needs to be compared with the quality benchmark when the iteration is stable to determine whether the optimization is effective. If the updated annotation effect is not better than the quality benchmark, it indicates that there is a deviation in the adjustment of the fusion coefficients and it needs to be re-optimized. If the annotation effect meets the standard, solidifying the fusion coefficients can ensure that the optimal parameters can be reused for subsequent full annotation and similar tasks, avoiding repeated adjustments and forming a complete optimization loop. The quality benchmark when the iteration tends to be stable is retrieved. This quality benchmark includes core indicators such as annotation accuracy, defect induction rate and audit domain adaptability, which are the core basis for judging the optimization effect of the fusion coefficients.

[0164] The annotation results validated with a small sample are compared one by one with the quality benchmark. The focus is on verifying whether the annotation accuracy has improved and whether the defect induction rate has decreased. It is also verified whether the annotation results guided by the updated fusion coefficient meet the accounts receivable audit standards and whether they are suitable for the needs of core links such as aging determination and bad debt provision. If the comparison results show that the annotation effect is better than or equal to the quality benchmark and the defects of each audit sub-dimension are effectively controlled, it indicates that the updated fusion coefficient has better adaptability. The parameters of the fusion coefficient and the stable prompt word knowledge graph are solidified to form a parameter combination for subsequent full-scale annotation and similar audit tasks.

[0165] If the comparison results do not meet the quality benchmark requirements, it indicates that there is room for optimization in adjusting the fusion coefficients. The process involves returning to the previous verification and reconstructing the relationships based on the current quality data to update the adjustment rules and fusion coefficients. This verification process is repeated until the desired effect is achieved. By comparing with the quality benchmark, the effectiveness of the updated fusion coefficients is ensured, preventing the solidification of inferior fusion coefficients that could lead to a decline in annotation quality. The solidified fusion coefficients, combined with the parameters of the prompt word knowledge graph, form a synergistic combination. This ensures the high-quality progress of the current annotation task and provides standardized support for subsequent similar audit annotation tasks, reducing the cost of repetitive optimization and improving the standardization and efficiency of the annotation process, thereby generating a high-quality annotation dataset. Simultaneously, it provides reusable parameter templates for similar accounts receivable annotation tasks, reducing optimization time at the task initiation stage and improving overall work efficiency.

[0166] Example 2

[0167] This application provides a prompt-driven data annotation system, which includes a task parsing module, a data annotation module, an intelligent iteration module, and a quality feedback module.

[0168] The task parsing module is used to receive annotation task requests, parse the annotation tasks to extract the domain features of the annotation tasks, extract the operation behavior vectors from the operation behavior of the annotators, and fuse the domain features and operation behavior vectors to generate a joint guidance vector.

[0169] The data annotation module retrieves prompt words from the prompt word knowledge graph based on the joint guiding vector, processes the data to be annotated and the prompt words to output the initial annotation results and confidence level, and analyzes the initial annotation results to identify the differences in the domain distribution of annotation defect types and related prompt words;

[0170] The intelligent iteration module is used to filter initial annotation results with confidence scores below the confidence threshold. It determines the reward function based on annotation accuracy and domain adaptability, and uses the difference in domain distribution between annotation defect types and associated prompt words as the state space. It iteratively corrects the parameters of the prompt word knowledge graph through reinforcement learning.

[0171] The quality feedback module is used to terminate iterative optimization and output a labeled dataset when the continuous iterations tend to stabilize. Based on the quality evaluation data of the iterative process, it provides feedback to update the fusion parameters of the joint guiding vector.

[0172] The technical principle of the system in this application embodiment is similar to the method described above in this application embodiment. Therefore, the system implementation is the same as the method implementation, and the detailed technical means will not be repeated.

Claims

1. A data annotation method based on prompt words, characterized in that, include: Receive annotation task requests, parse the annotation tasks to extract domain features of the annotation tasks, extract the operation behavior vectors from the operation behavior of the annotators, and fuse the domain features and operation behavior vectors to generate a joint guidance vector; Based on the joint guiding vector, prompt words are retrieved from the prompt word knowledge graph. The data to be labeled and prompt words are processed to output the initial labeling results and confidence scores. The initial labeling results are analyzed to identify the differences in the domain distribution of labeling defect types and related prompt words. The differences in the domain distribution of the identified defect types and associated prompts include: Based on the confidence level of the initial annotation results and the discrepancies between the two types of annotation results, the discrepancies with confidence levels less than the confidence threshold are filtered out, and the discrepancies are classified according to their sources to identify the types of annotation defects. The annotation defect type is bound to the domain prompt words and operation prompt words that generate the initial annotation result, and the frequency of prompt words with different proportions of fusion coefficient triggering the corresponding annotation defect type is statistically analyzed; Based on the subdivision of the domain to which the current annotation task belongs, the defect induction rate of the same prompt word under different subdivisions is calculated based on the frequency, and the distribution ratio of annotation defect types among different subdivisions is statistically analyzed to obtain the domain distribution differences of the associated prompt words corresponding to the annotation defect types. Initial annotation results with confidence scores below the confidence threshold are selected. The reward function is determined based on annotation accuracy and domain adaptability. The domain distribution difference between annotation defect types and associated prompt words is used as the state space. The parameters of the prompt word knowledge graph are iteratively corrected through reinforcement learning. When the continuous iterations tend to stabilize, the iterative optimization is terminated and the labeled dataset is output. Based on the quality evaluation data of the iterative process, the fusion coefficient of the joint guiding vector is updated.

2. The data annotation method based on prompt words as described in claim 1, characterized in that, The domain features of the extraction and annotation task include: Receive annotation task requests and perform hierarchical semantic decomposition of the text description of the annotation task requests, including basic semantics, task objectives, and data constraints, to obtain initial features; Retrieve historical domain features from historical annotation tasks, and use cosine similarity to match the distance between initial features and historical domain features to filter historical annotation tasks of the same type with feature similarity not less than the similarity threshold. Identify cross-domain overlapping features in the initial features, calculate the weights by weighting the annotation accuracy and feature similarity of similar historical annotation tasks, and retain the cross-domain overlapping features with high weights as core features; Basic weights are assigned to core features, and the basic weights are dynamically adjusted based on the contribution of core features to the labeling accuracy in similar historical labeling tasks. After weighted fusion, the domain features of the labeling task are formed. Cross-domain overlap features represent the overlapping part of the initial features of the current annotation task with the historical domain features of similar historical annotation tasks. Corresponding weight proportions are assigned to annotation accuracy and feature similarity. The weight of cross-domain overlap features is obtained by weighting annotation accuracy and feature similarity with the corresponding weight proportions, so as to determine the importance of cross-domain overlap features.

3. The data annotation method based on prompt words as described in claim 2, characterized in that, Extracting the operation behavior vector includes: Acquire the basic, interactive, and decision-making behaviors of annotation personnel at different times in similar historical annotation tasks; Based on the historical annotation accuracy of similar annotation tasks, the behavior weights of different operation behaviors are dynamically allocated according to the operation type. A behavior feature matrix is ​​constructed based on behavior weights and operational behaviors. Operation patterns are clustered using density clustering algorithms, and abnormal operational behaviors are eliminated to extract the behavior features of different operation patterns. The coding dimension weights are adjusted based on the domain features of the annotation task, and the behavioral features are mixed-encoded and normalized to generate an operation behavior vector consistent with the domain feature dimensions. The coding dimension weights are adjusted according to the correlation between the domain features and behavioral features of the annotation task.

4. The data annotation method based on prompt words as described in claim 3, characterized in that, The generation of the joint guiding vector includes: Retrieve the preset domain ontology knowledge base, extract the quantitative value of the feature importance of the current annotation task's domain, and assign domain adaptation weights to the domain features according to the proportion of the quantitative value; Select a small sample of data to be labeled for the current labeling task, complete the small sample labeling based on the operation behavior vector, and determine the effective weight of the operation behavior vector based on the degree of fit between the labeling results and the standard labeling. Using domain adaptation weights and effective behavior weights as fusion coefficients, we weight and fuse domain features and operational behavior vectors with consistent dimensions to generate a joint guidance vector.

5. The data annotation method based on prompt words as described in claim 4, characterized in that, The search suggestions include: The prompt word knowledge graph is preprocessed by classifying the prompt words according to their respective domains and quantifying them into domain category vectors, and then associating them with the operation behavior tags of the personnel labeled in the corresponding domains. Extract the corresponding domain component from the joint guidance vector, perform cosine similarity coarse matching with the domain category vector of the prompt word knowledge graph, and filter out candidate prompt words belonging to the domain of the current annotation task; Extract the corresponding operation components from the joint guidance vector, calculate the matching degree between the operation behavior label and the operation component in the candidate prompt words, retain the prompt words with a matching degree greater than the matching threshold, and obtain the usable candidate set; Based on the fusion coefficient of the joint guiding vector, the prompt words in the available candidate set are sorted in descending order by weight. The number of prompt words is preset based on the annotation complexity of the current annotation task, and the prompt words are output according to the sorting result. The fusion coefficient of the joint guiding vector includes the domain adaptation weight and the behavior effectiveness weight.

6. The data annotation method based on prompt words as described in claim 5, characterized in that, The initial annotation results and confidence levels of the output include: The output prompts are classified according to the proportion of the fusion coefficient. Prompts with a fusion coefficient greater than the preset proportion are set as domain prompts and operation prompts, respectively. Perform domain semantic annotation on the data to be annotated based on domain prompt words and domain features, and perform operation adaptation annotation based on operation prompt words and operation behavior vectors to generate domain annotation results and operation annotation results; Calculate the degree of fit between the two types of annotation results. If the degree of fit is greater than the degree of fit threshold, merge them into the initial annotation result. Otherwise, call the preset annotation standard fragment to correct the divergent parts and generate the initial annotation result. Based on the consistency of the generated results, the matching degree of the prompt words and the fit of the operation annotation results, the confidence of the prompt word adaptability and the operation association is calculated and then the confidence of the initial annotation result is output after weighted summation.

7. The data annotation method based on prompt words as described in claim 6, characterized in that, The parameters of the iteratively corrected prompt word knowledge graph include: Initial annotation results with confidence scores below the confidence threshold are selected, and a reward function is determined based on annotation accuracy and domain adaptability. The parameters of the reward function are then dynamically adjusted according to differences in domain distribution. The differences in the domain distribution of labeled defect types and associated prompt words are quantitatively represented as a state space. Through reinforcement learning, the parameters of the prompt word knowledge graph are hierarchically corrected according to the different proportions of the frequency of the labeled defect type corresponding to the prompt word and the fusion coefficient. After each round of correction, prompt words are retrieved again based on the corrected prompt word knowledge graph and labeled. The cumulative score of the reward function is calculated. If the cumulative score does not reach the convergence threshold, the state space is updated based on the new domain distribution difference until the iteration tends to stabilize.

8. The data annotation method based on prompt words as described in claim 7, characterized in that, The fusion coefficients of the feedback update joint guiding vector include: Based on the quality assessment data from the iterative process, the annotation accuracy and defect induction rate are extracted, and the fusion coefficients of the joint guiding vectors are correlated to form a correlation relationship. Based on the correlation, the adjustment rules are determined in conjunction with the deviation direction of the quality assessment data. The fusion coefficients are updated based on the adjustment rules. The joint guidance vector is regenerated using a small sample of unlabeled data to complete the labeling with the search prompt words and monitor the labeling effect. By comparing the annotation results with the quality benchmark when the iterations tend to stabilize, it is possible to determine whether the fusion coefficients of the joint guiding vector after the update should be solidified.

9. A cue-driven data annotation system, used to implement the cue-driven data annotation method according to any one of claims 1-8, characterized in that, include: The module includes a task parsing module, a data annotation module, an intelligent iteration module, and a quality feedback module. The task parsing module is used to receive annotation task requests, parse the annotation tasks to extract the domain features of the annotation tasks, extract the operation behavior vectors from the operation behavior of the annotators, and fuse the domain features and operation behavior vectors to generate a joint guidance vector. The data annotation module retrieves prompt words from the prompt word knowledge graph based on the joint guiding vector, processes the data to be annotated and the prompt words to output the initial annotation results and confidence level, and analyzes the initial annotation results to identify the differences in the domain distribution of annotation defect types and related prompt words; The intelligent iteration module is used to filter initial annotation results with confidence scores below the confidence threshold. It determines the reward function based on annotation accuracy and domain adaptability, and uses the difference in domain distribution between annotation defect types and associated prompt words as the state space. It iteratively corrects the parameters of the prompt word knowledge graph through reinforcement learning. The quality feedback module is used to terminate iterative optimization and output a labeled dataset when the continuous iterations tend to stabilize. Based on the quality evaluation data of the iterative process, it provides feedback to update the fusion parameters of the joint guiding vector.

Citation Information

Patent Citations

  • Artificial intelligence data annotation and cue word automatic construction engine system

    CN120950973A

  • Space-time knowledge graph question and answer method, system and device and storage medium

    CN121235050A