Visual annotation method for text and image data
By introducing a text segmentation model based on sentence-level sequence annotation and a cross-domain low-sample data pre-labeling model based on Prompt templates in the annotation method, the existing annotation methods cannot meet the problem that multimodal knowledge graph construction and lack of user experience are solved, and efficient and accurate knowledge graph entity and relationship recognition is achieved, which is suitable for multimodal large model pre-labeling in the military field.
Patent Information
- Application Number
- CN202411734190.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-29
AI Technical Summary
The existing annotation methods cannot meet the application scenarios of multimodal knowledge graph construction, lack user experience, low intelligent annotation efficiency, cannot be directly applied to the knowledge graph information extraction model, and lack the pre-labeling ability of multimodal large models in the military field.
It provides a visual annotation method for text and image data, adopts a cross-domain low-sample data pre-notation model based on sentence-level sequence annotation, a cross-domain low-sample data pre-notation model based on sentence-level sequence annotation, realizes intelligent pre-notation and collaborative annotation, and improves the efficiency and quality of data annotation.
It improves the accuracy of knowledge graph entity and relationship recognition, improves the efficiency and quality of data labeling, realizes cross-domain intelligent data labeling, and can be directly applied to knowledge graph information extraction models, especially for multimodal large model pre-labeling in the military field.
Smart Images

Figure CN120045707A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text and image annotation corpus in the military field, and specifically to a visual annotation method for text and image data, and further to functions such as classification management, visual data annotation, intelligent pre-annotation, and annotation result statistics. The present invention can provide accurate learning samples for tasks such as data labeling, knowledge graph construction, event extraction, and target recognition in the field of artificial intelligence, and can be used for training, testing, and verification and evaluation of the performance of deep learning models in the field of text and image data processing. Background Art
[0002] With the rapid development of artificial intelligence technology, data labeling, as an important data processing method, is gradually being widely used. Data labeling is to annotate or label data manually or semi-automatically to help algorithm models better understand and process data. Data labeling has always been an important step in training supervised learning algorithms because it provides the correct output so that machine learning algorithms can gradually adjust and improve their prediction accuracy. The significance and main role of data labeling are mainly reflected in the following aspects:
[0003] (1) Provide a training data set. Data labeling is the basic work of building a training data set. When training an artificial intelligence model, a large amount of labeled data is required as training samples. Through data labeling, the correct label can be added to each sample in the data set, thus providing a reference for the model to learn.
[0004] (2) Improve model performance. Data labeling can improve the performance of machine learning models. By adding labels to samples in a dataset, the model can better understand and learn the characteristics of the data during training. Labeled data can help the model identify key information and gradually adjust and improve the prediction accuracy and overall performance of the model.
[0005] (3) Promote the development of intelligent applications. Data annotation is an important part of developing intelligent applications. Whether it is image recognition, speech recognition or natural language processing, a large amount of annotated data is required to train the model. Through data annotation, sufficient training data can be provided for intelligent applications, thereby improving the recognition and understanding capabilities of the application and meeting user needs.
[0006] Data labeling has important practical significance in the field of artificial intelligence. Through data labeling, we can provide training data sets, improve model performance, promote algorithm research and development, support the development of intelligent applications, and promote industry development and innovation.
[0007] The existing annotation methods mainly have the following two problems:
[0008] (1) Existing annotation methods often focus on covering all data types and annotation tasks, or only provide annotation functions for specified data types in a specific field. Not only can they not meet the application scenarios of building multimodal knowledge graphs, but data annotation in the field of knowledge graph construction requires annotating entity types, entity relationship types, and entity attributes in the text based on the knowledge graph metamodel.
[0009] (2) The existing data standardization methods lack user experience, which seriously affects the efficiency and quality of data annotation. The data annotation results are often displayed in the form of tables, etc., and the annotation process cannot be visualized on the original text by combining background color, labels and lines.
[0010] (3) The intelligent labeling of existing labeling methods often requires a large amount of training data to provide the recognition performance and accuracy of the model. The domain characteristics are relatively strong, and the cross-domain recognition efficiency is low and the effect is poor.
[0011] (4) It is not bound to the domain ontology, and the annotation result data cannot be directly applied to the training of the knowledge graph information extraction model. Additional data format, annotation system and other conversions are required.
[0012] (5) Existing annotation methods lack the ability to pre-annotate based on large multimodal models in the military field. Annotating entities, attributes, relationships, events, and other elements such as military equipment, institutions, countries, regulations, and geographic information requires a lot of manpower and time costs. Summary of the invention
[0013] The purpose of the technical solution of the present invention is to improve the accuracy of knowledge graph entity and relationship recognition, improve data standard efficiency, and intelligently label cross-domain (military field) data.
[0014] In order to solve the above technical problems, the technical solution of the present invention provides a visual annotation method for text and image data, comprising the following steps:
[0015] Obtain the corpus, determine whether the corpus is an image, and if so, convert it into text. Proofread the text-based corpus, remove tags, and extract corpus description information.
[0016] A text segmentation model based on sentence-level sequence annotation is used to divide the corpus description information into sentences and add special tags to obtain a character sequence. The embedding vector obtained through the character embedding layer is summed with the position vector and the segment vector to obtain the final character vector. The final character vector is mapped to each sentence through the BERT encoder. The K character output vectors corresponding to each sentence are average pooled to obtain the final sentence vector. Each sentence encoding is mapped through the output layer and the softmax layer to classify whether each sentence is a paragraph boundary to achieve paragraph analysis.
[0017] According to the pre-established terminology database and domain vocabulary database, the sentences after paragraph analysis are segmented and denoised, and the segmented words are weighted according to their importance in the sentence. The weighted words are converted into hash values, and the hash values are used to generate weighted digital strings according to the weighted weights. Each weighted digital string is accumulated to obtain a sequence digital string and the dimensionality reduction process is performed to obtain the final hash value. Multiple final hash values are compared to determine the differences between sentences, so as to realize duplicate data processing;
[0018] According to entity extraction, attribute extraction, and event extraction, a prompt template combining answer-type prompts and task-type prompts is constructed. According to entity extraction, corresponding entity categories and entities, attribute extraction, corresponding entity categories, entities, and relationships, and event extraction, corresponding event types, events, argument roles, and arguments are constructed to establish a schema template.
[0019] The Prompt-based cross-domain few-sample data pre-annotation model includes an input layer, an encoding layer, and a UniLM decoding layer. The schema template is concatenated as the input of the input layer. The position encoding part in the encoding layer uses RoPE rotation position encoding. The UniLM decoding layer uses Seq2Seq attention mask mask mode to input bidirectional modeling and output unidirectional modeling to achieve conditional generation.
[0020] The input sequence is divided into a source sequence, a target sequence, and a separator between the source sequence and the target sequence. A random mask matrix is applied to the target sequence. The unmasked tokens in the source sequence and the target sequence are selected to predict the masked tokens for information reorganization tasks.
[0021] Input some sentences into the Prompt-based cross-domain few-sample data pre-annotation model, and output the pre-annotation results according to the task type;
[0022] The quantity, quality, and completion of the labeled data are counted according to different tasks, and multiple pre-labeling results are combined to display them in the form of knowledge graphs, tables, and charts.
[0023] Preferably, the initial context representation formula in the UniLM decoding layer is as follows:
[0024]
[0025] In the formula, the input sequence is X = {x 0 ,x 1 ,…,x l}, x l Indicates the length of the sequence.
[0026] Preferably, the final representation of the initial context representation by stacking n layers is as follows:
[0027]
[0028] Preferably, when inputting some sentences into the Prompt-based cross-domain few-sample data pre-annotation model, a label smoothing loss function with the following formula is used:
[0029]
[0030] In the formula, y i represents the i-th element of the true label vector, p i is the probability value of the i-th element predicted by the model, ε is a smoothing coefficient less than 1, and k is the number of categories.
[0031] The technical solution of the present invention proposes a visual annotation method for text and image data. By annotating the entity types, entity relationship types and entity attributes in the text and adding visual annotation of image data, the accuracy and performance of the information extraction model constructed by the knowledge graph are improved, and the image data is further annotated, solving the problem of not being able to meet the construction application scenarios of the multimodal knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Principles of data annotation technology;
[0033] Figure 2 Text segmentation model for sentence-level sequence annotation;
[0034] Figure 3 De-duplication detection process for collected data;
[0035] Figure 4 Enter the template for Prompt;
[0036] Figure 5 A cross-domain few-sample data pre-labeling model based on Prompt;
[0037] Figure 6 It is the Mask strategy of the UniLM model;
[0038] Figure 7 Annotate the system architecture for data;
[0039] Figure 8 Principles of data annotation technology;
[0040] Fig. 9 Provides a visual annotation interface for text data. DETAILED DESCRIPTION
[0041] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall within the scope limited by the appended claims of the application equally.
[0042] The principle of a visual annotation method for text and image data provided by an embodiment of the present invention is as follows: Figure 1 As shown in the figure, on the basis of the traditional data labeling process, the data labeling processing and text data pre-labeling stages are added, which greatly improves the data labeling efficiency and makes the data labeling method have strong cross-domain intelligent data labeling performance. The principles of the data labeling process are introduced below, and the implementation principles of data labeling processing and text data pre-labeling are introduced in detail.
[0043] (1) Annotation data collection
[0044] Annotated data collection mainly provides annotated corpus management for tasks such as annotated corpus entities and entity relationship extraction, including data maintenance and user service functions. Data maintenance mainly includes corpus entry, proofreading, storage, modification, deletion, tag removal and corpus description information management.
[0045] (2) Annotation data processing
[0046] Corpus annotation processing mainly realizes the automatic processing of corpus, mainly including paragraph analysis and duplicate data processing. Paragraph analysis mainly realizes the recognition and calculation of the content structure of corpus, so that the corpus can obtain a comparative display effect in the annotation interface; duplicate data processing uses the improved duplicate checking algorithm model to quickly compare the collected corpus with the corpus stored in the system and obtain results quickly.
[0047] 1) Text segmentation model based on sentence-level sequence labeling
[0048] Document segmentation is defined as the automatic prediction of the boundaries of a document (paragraph or chapter). Existing document segmentation work has mainly focused on written text, including two main methods: unsupervised and supervised. Document segmentation is a task that strongly relies on the information of long text chapters. Sentence-by-sentence classification models are prone to model performance obstacles when using semantic information of long texts, while hierarchical models have problems such as large amount of computation and slow reasoning speed. To this end, a text segmentation model based on sentence-level sequence annotation is proposed for data such as spoken ASR manuscripts. The model framework is shown in the figure below. Figure 2 shown.
[0049] ① The model models document segmentation as a sentence-level sequence labeling task. First, the input document is segmented into sentences. Each sentence is segmented by a tokenizer and a special tag is added.
[0050] ② The character sequence after word segmentation obtains the embedding vector through the character embedding layer, and the final character vector is obtained by element-wise summing with the position vector and segment vector;
[0051] ③Then the character vector is input into the BERT encoder, and the character vector output by the encoder is mapped to each sentence. The K character output vectors corresponding to each sentence are averaged and pooled to obtain the final sentence vector. Finally, each sentence encoding is mapped through the output layer and the softmax layer to classify whether each sentence is a paragraph boundary.
[0052] 2) Duplicate data processing
[0053] In the data collection process, in addition to text segmentation, the focus needs to be on data duplication checking to improve the quality of collected data. The specific data duplication detection process is as follows: Figure 3 shown.
[0054] ① Word segmentation. Word segmentation needs to be able to identify domain terms, concepts and other professional knowledge. This method improves the recognition effect of domain vocabulary and the accuracy of word segmentation by establishing a terminology database and a domain vocabulary database and writing them into the word segmenter. After denoising based on the word segmentation results, weighting is performed according to the importance of the vocabulary in the sentence. The more important the vocabulary is in the sentence, the greater the weight.
[0055] ② Hash. Use the hash algorithm to convert each word into a hash value, converting the vocabulary into numbers that can be run by computers, realizing similarity calculation and improving similarity calculation performance;
[0056] ③Weighting. Based on the hash generated in the previous step, a weighted digital string needs to be formed according to the weight of the word.
[0057] ④Merge, add up the sequence values calculated for each word to get a sequence number string.
[0058] ⑥ Dimensionality reduction: perform dimensionality reduction on the calculation result of the previous step. If each bit is greater than 0, it is recorded as 1, and if it is less than 0, it is recorded as 0 to form the final hash value.
[0059] Through the above process, the problem of calculating the similarity of text content can be converted into a numerical calculation problem, and the differences between text contents can be judged by comparing the differences in binary digital strings.
[0060] (3) Pre-annotation of text data
[0061] On the basis of the manual labeling function, an intelligent pre-labeling function is provided to quickly complete data labeling, improve data labeling efficiency, and reduce the cost of manual data labeling. Data pre-labeling can instantly label entities, attributes, relationships and other information in the text according to the documents uploaded by users. It is necessary to solve problems such as cold start, few samples, and cross-domain. There is an urgent need for a general pre-labeling model to complete the pre-labeling of basic text information without training or with a small amount of sample training. To this end, a cross-domain few-sample data pre-labeling technology based on Prompt is proposed, and the Prompt template that combines the answer-type Prompt and the task-type Prompt is used for optimization. That is, task prompts are added at the beginning of the text, and a uniformly constructed schema template is added at the end of the text. For example Figure 4 shown.
[0062] In the prompt template, INS represents the task type, such as entity extraction, attribute extraction, event extraction, etc.; s-type is a pre-constructed schema type, such as fighter, ship, personnel, country, etc.; ele_n[un_n] represents an element. Among them, ele_n is the text description of the element, and [un_n] is the sentinel token, which represents the element. For example, in the entity recognition task, it can represent entity categories and entities; in the relationship extraction task, it can represent entity categories, entities and relationships; in the event extraction task, it represents event type, event, argument role and argument respectively. Different elements correspond to different sentinel tokens, which can distinguish different elements.
[0063] The Prompt-based cross-domain few-sample data pre-labeling model is mainly composed of an input layer, an encoding layer, and a UniLM decoding layer. The input layer concatenates the task prompt, the original text, and the uniformly constructed schema template as input. The encoding layer uses three basic input representations, among which the position encoding part uses RoPE rotation position encoding; the UniLM layer uses the Seq2Seq attention mask, that is, bidirectional input modeling and unidirectional output modeling to achieve conditional generation. The overall structure is as follows: Figure 5 shown.
[0064] The role of the input layer is to process the original text into a form that the model can understand, which is a very critical step in the information extraction task. In the model, first, a unified schema is built for all training data, that is, the entities, attributes and relationships in all subtasks of information extraction are uniformly defined and constrained, and the elements to be extracted are structured. Then, the task description of information extraction, the preset schema template and the uploaded document are spliced to generate an input sequence to guide the model to generate the required structured text.
[0065] The encoding layer uses the Transformer architecture as the basic structure and uses three input representations, namely word embedding, segment embedding, and position embedding. Word embedding is a method of mapping discrete word symbols into continuous real number vectors; segment embedding is a method of representing the relationship between different text paragraphs; position embedding is a method of representing the relative or absolute position information of words in a sequence, which can help the model capture the order relationship and contextual information between words. The model uses RoPE rotation position encoding as position embedding, which encodes the absolute position in a rotation matrix and adds explicit relative position dependency to the self-attention formula, so that the model can support longer text sequences and achieve improved results.
[0066] The decoding layer uses the UniLM model based on the Decoder-Only Transformer architecture. By modifying different attention masks, different language model models such as bidirectional, unidirectional, and seq2seq can be modeled. The UniLM model consists of a multi-layer Transformer encoder. Suppose the input sequence is X = {x 0 ,x 1 ,…,x l}, x l Indicates the length of the sequence, and inputs X into the first layer Transformer to obtain the initial context representation as follows:
[0067]
[0068] After N layers of Transformer stacking:
[0069] H n =Transformer n (H n ),n∈[1,N]
[0070] Get the final representation:
[0071]
[0072] UniLM uses different mask matrices to determine which tokens the model pays attention to in the next operation. Therefore, under the same set of parameters, UniLM can change the operation mode of the model directly by changing the mask method. The UniLM model can generally use three different mask strategies to adapt to different tasks, as follows: Figure 6 shown.
[0073] Since the sequence-to-sequence prediction strategy divides the input sequence into two parts, the source sequence and the target sequence, and adds a separator between the two parts, the source sequence is not masked, and the target sequence randomly masks some tokens. The model can use the unmasked tokens in the source and target sequences to predict the masked tokens, which is more suitable for information compilation tasks. Therefore, the model chooses the sequence-to-sequence masking method for modeling.
[0074] A small amount of training data is input into the prompt template-based extraction model for training, using the label smoothing loss function, the formula is as follows:
[0075]
[0076] Among them, y i represents the i-th element of the true label vector, p i is the probability value of the i-th element predicted by the model, ε is a smoothing coefficient less than 1, and k is the number of categories. Finally, the pre-labeling result is output according to the task type.
[0077] (4) Annotation result statistics
[0078] Data annotation result statistics can count the quantity, quality, and completion of the annotated data according to different tasks, and also support the quantity, quality, and completion of different labels for the same task. It provides annotated corpus visualization function, which can display the annotation results in the form of knowledge graphs, tables, and charts.
[0079] The data annotation system architecture is as follows Figure 7 shown.
[0080] Among them, (1) annotated corpus management
[0081] Corpus management is used to manage corpora for tasks such as classification, entity relationship extraction, and label extraction. It includes data maintenance, automatic corpus processing, and user service functions. Data maintenance mainly includes functions such as corpus entry, proofreading, storage, modification, deletion, format conversion, merging, tag removal, and corpus description information management; automatic corpus processing mainly realizes the automated processing of corpora, mainly including word segmentation, annotation, text segmentation, merging, corpus alignment, tag processing, etc.; user service functions in corpus management mainly include query, retrieval, number of entries, sharing, downloading, etc.
[0082] (2) Labeling task management
[0083] Labeling task management includes functions such as task creation, editing, deletion, query and task assignment, and supports centralized labeling and multi-person collaborative labeling. Labeling task information includes task name, task description, task type, recommended model, collaborative labeling personnel, and assigned labeling quantity. Labeling task types include entity labeling, semantic relationship labeling, one-way text labeling, multiple text labeling, and sentiment labeling. Data labeling tasks support multi-person collaborative labeling, and can view labeling progress and review status, and can review collaborative labeling results; administrators can assign labeling tasks to different users or teams, specify the scope and requirements of labeling, and the task assignment function ensures that labeling work is carried out in an orderly manner; provide labeling task progress tracking function, administrators can view the completion status of labeling tasks in real time. Progress tracking helps managers promptly discover and solve problems that arise during the labeling process.
[0084] (3) Dataset Management
[0085] Dataset management provides unified management of annotated corpora, supports querying, filtering and exporting of annotation results, and users can choose from a variety of export formats, such as CSV, JSON, XML, etc. The system provides an export wizard to help users quickly export annotation results. The system automatically saves the annotation results and records the annotation process of each annotated data, including the annotating user, time, etc., to ensure the traceability of the annotation process. It supports the construction of annotated datasets by brushing the annotated data records for different algorithm model training scenarios.
[0086] (4) Labeling quality inspection
[0087] The annotation quality check supports auditors to manually check the annotation results to ensure the annotation quality. After the annotator submits the annotation data, the auditor is supported to manually check the annotation results according to the set sampling ratio, and can review the annotation quality item by item, support the re-marking of unqualified annotation results, and provide a review progress display function. The administrator role can review the data quality of ordinary users. Provide an audit status query function to display the number of audited and unaudited data. Provide an annotation corpus review function, display the annotated text paragraphs, and can review them item by item and in batches. It can display the proportion of repeatedly annotated text paragraphs and the consistency statistics of repeated annotation results.
[0088] (5) Intelligent pre-labeling
[0089] Based on the manual labeling function, we provide intelligent pre-labeling function to quickly complete data labeling, improve data labeling efficiency, and reduce the cost of manual data labeling. Intelligent pre-labeling refers to the process of using the existing algorithms in the system to generate labeling results based on the labels and data learning training in the current labeling stage. When selecting the pre-labeling algorithm model, you need to pay attention to matching the algorithm type with the label type of the data set.
[0090] (6) Data labeling results statistics
[0091] Data annotation result statistics can count the quantity, quality, and completion of the annotated data according to different tasks, and also support the quantity, quality, and completion of different labels for the same task. It provides annotated corpus visualization function, which can display the annotation results in the form of knowledge graphs, tables, and charts.
[0092] The embodiment of the present invention performs intelligent pre-annotation and collaborative annotation for weapon equipment text and image data in the military field, mainly including the steps of ontology management, annotation corpus management, annotation task management, annotation quality management and annotation result statistics. The process is as follows: Figure 8 shown.
[0093] (1) Ontology management stage
[0094] Before data annotation, users need to use ontology management to create, import, and modify the military weapons and equipment ontology, including entities, relationships, attributes and other elements, to form a weapons and equipment ontology with complete concepts, clear hierarchies, and obvious differences, which is used to guide data annotation content.
[0095] (2) Annotated corpus management stage
[0096] After the ontology is determined, users use the annotated corpus to manage the text and image data such as military news, white papers, blue books, regulations, etc. to be annotated that are connected to relational databases, graph databases, and ontology file systems. Then, they use paragraph analysis and deduplication functions to process duplicate data, realize the recognition and processing of paragraph structure, punctuation and other contents of the corpus, and improve the text display effect and annotation efficiency.
[0097] (3) Labeling task management
[0098] After data processing, the administrator uses the annotation task management function to create an annotation task and choose whether to enable intelligent pre-annotation. This system has a built-in self-developed multimodal large model in the military field. No training is required. Only prompt settings can be used to perform intelligent pre-annotation of entities, relationships, attributes and other elements on the military corpus to be annotated, thereby improving annotation efficiency. After accessing the multimodal large model in the military field, the administrator assigns the annotation task to team members for system annotation, and the administrator can view the annotation progress in real time.
[0099] (4) Labeling quality management
[0100] After the data labeling is completed, the administrator uses the labeling quality management function to review the labeling result data. The labeled data that passes the review forms the training, test, and evaluation data sets, and the labeled data that fails the review is returned to the team members for modification or re-labeling.
[0101] (5) Annotation result statistics
[0102] After the labeling task is completed, administrators and users can view the labeling statistics, including task dimension statistics, label dimension statistics, and visualized label corpus, to intuitively understand the completion status of the labeling task and the labeling results.
[0103] The beneficial effects of the embodiments of the present invention are as follows:
[0104] (1) The embodiment of the present invention is an important component of data analysis platforms such as multimodal knowledge graph platforms, deep learning algorithm platforms, data science platforms, and knowledge middle platforms. It can access data in professional data analysis platforms and complete the annotation of original corpus data. Data annotation is the basis for deep learning algorithm training. By providing high-quality and large amounts of training data for text, image and other unstructured data processing and analysis algorithm models, the accuracy and performance of data processing algorithm models are improved, and the data analysis platform's ability to compile and process professional data, knowledge, data analysis, and data intelligence applications is enhanced.
[0105] (2) The embodiment of the present invention supports traditional manual data labeling and provides a data pre-labeling function. It is possible to select an existing algorithm model in the data labeling system to automatically label the corpus, and after the labeling is completed, it is manually modified and improved to finally complete the labeling of the data corpus. At the same time, the data labeling system supports allocating data labeling work to different users in the form of tasks, and completing the data labeling work through multi-person collaboration. Compared with traditional data labeling methods, data pre-labeling and collaborative data labeling can greatly improve data labeling efficiency and save data labeling costs. At the same time, the embodiment of the present invention provides functions such as data labeling review, data set management, and data statistical analysis, which further improves the functions of the data labeling system. The drag-and-drop visual data labeling method greatly improves the UI interactive experience of data labeling, and displays the data labeling content and progress in real time and intuitively, thereby improving data labeling efficiency and quality.
[0106] Implementation Example 1: Data Annotation Based on Knowledge Graph Metamodel
[0107] (1) Modeling is performed under the metamodel management function module of the knowledge graph platform to design the entity types, relationship types, and entity attributes of the knowledge graph;
[0108] (2) In the entity and relationship annotation module of the data annotation system, it is possible to access and display entity types and relationship types, and distinguish them by different colors;
[0109] (3) Select the leaf node of the classification tree on the left side of the page. The annotated corpus is mounted on the leaf node, and the non-leaf nodes are the classification of the corpus. You can also enter a new annotated corpus and mark the words in the corpus by clicking the entity type label with the mouse. The annotation system adds a background color to the marked words and adds a type label;
[0110] (4) Click the relationship type label with the mouse, first click a marked entity, drag the mouse, and the annotation system will display an arc with an arrow, and then connect it to another marked entity to form a completed connection. The starting position of the connection is the head entity, the end position of the arrow is the tail entity, and the name in the middle of the connection is the relationship between the head and tail entities. The annotation system can automatically calculate the height according to the annotation of entities and relationships, and rearrange the paragraphs in the annotated corpus according to the calculation results to achieve the best display effect;
[0111] (5) In the entity attribute annotation interface, select the leaf node of the corpus classification tree on the left. The detailed content of the corpus and the annotated entities will be displayed on the right side of the page. Select the text corresponding to the attribute value with the mouse, and the annotation system will add a background color and label the attribute value. Select a labeled entity and drag the mouse to display an arc with an arrow. As the mouse is dragged, the arc is connected to a labeled attribute value to form a completed connection. The starting position of the connection is the entity, the end position of the connection arrow is the attribute value, and the name in the middle of the connection is the attribute type name. The annotation system can also automatically calculate the display height according to the annotation of the entity attributes, and rearrange the paragraphs of the annotated corpus to achieve the best display effect.
[0112] Implementation example 2: Image annotation
[0113] (1) In the image annotation module, the left side of the page contains the image annotation corpus, including the name, description, and thumbnail of the image corpus. The top of the page contains the image annotation label, the bottom contains the image content, and the right side contains the image annotation toolbar;
[0114] (2) Click on the image annotation label (for example: wheel, track, antenna, etc.) and complete the annotation on the image by dragging a frame. After the annotation is completed, the annotation system can color and mark the annotated area on the image and add a corresponding label;
[0115] (3) Right-click on the marked area to delete the marked content, including color, area, and label;
[0116] (4) The annotation toolbar on the right side of the page provides functions such as image zooming in, zooming out, restoring, and deleting, which facilitates the annotation of image corpus.
[0117] The embodiments of the present invention not only provide manual and automatic annotation functions, but also focus on improving the interactive effect of the data annotation UI. Manual data annotation can be completed by dragging and dropping, and the visual display effect of data annotation results can be improved by color, connection, label, etc., thereby improving the efficiency and experience of the data annotation process. At the same time, the key links of data annotation are further improved by the annotation data quality detection, annotation result statistical analysis, and data set management module, and the data annotation process is optimized and detected, thereby improving the quality and annotation efficiency of the annotation data. The text data annotation function interface is performed by using the method and system of the embodiments of the present invention. Fig. 9 .
[0118] The embodiment of the present invention creatively integrates the metamodel of the knowledge graph into the annotation system to provide data classification labels for data annotation. This ensures that the annotation result data set is highly consistent with the knowledge graph metamodel, and can be directly used for knowledge graph construction tasks such as knowledge graph entity extraction, relationship extraction, and entity attribute extraction, thereby improving the accuracy and performance of the information extraction model constructed by the knowledge graph.
[0119] The embodiment of the present invention improves the data annotation component based on the d3 visualization library, and can distinguish the annotated data content by color according to different data labels, intuitively display the relationship between data labels through connecting lines, and display the relationship category name on the connecting line. It supports direct deletion of annotated phrases and relationships by right-clicking the mouse. The improved annotation control makes the data annotation process more intuitive and friendly, and indirectly improves the data annotation efficiency.
[0120] The embodiment of the present invention not only provides a variety of manual data labeling modes, but also supports artificial intelligence labeling based on pre-trained models to reduce the cost of manual labeling. It proposes a cross-domain small sample data pre-labeling technology based on Prompt, and optimizes the Prompt template that combines answer-type Prompt and task-type Prompt. It can obtain better label classification and recognition effects in different fields, so that the data labeling system has strong cross-domain intelligent data labeling performance.
Claims
1. A visual annotation method for text and image data, characterized in that: The following steps are involved: Obtain the corpus, determine whether the corpus is an image, and if so, convert it into text. Proofread the text-based corpus, remove tags, and extract corpus description information. A text segmentation model based on sentence-level sequence annotation is used to divide the corpus description information into sentences and add special tags to obtain a character sequence. The embedding vector obtained through the character embedding layer is summed with the position vector and the segment vector to obtain the final character vector. The final character vector is mapped to each sentence through the BERT encoder. The K character output vectors corresponding to each sentence are average pooled to obtain the final sentence vector. Each sentence encoding is mapped through the output layer and the softmax layer to classify whether each sentence is a paragraph boundary to achieve paragraph analysis. According to the pre-established terminology database and domain vocabulary database, the sentences after paragraph analysis are segmented and denoised, and the segmented words are weighted according to their importance in the sentence. The weighted words are converted into hash values, and the hash values are used to generate weighted digital strings according to the weighted weights. Each weighted digital string is accumulated to obtain a sequence digital string and the dimensionality reduction process is performed to obtain the final hash value. Multiple final hash values are compared to determine the differences between sentences, so as to realize duplicate data processing; According to entity extraction, attribute extraction, and event extraction, a prompt template combining answer-type prompts and task-type prompts is constructed. According to entity extraction, corresponding entity categories and entities, attribute extraction, corresponding entity categories, entities, and relationships, and event extraction, corresponding event types, events, argument roles, and arguments are constructed to establish a schema template. The Prompt-based cross-domain few-sample data pre-annotation model includes an input layer, an encoding layer, and a UniLM decoding layer. The schema template is concatenated as the input of the input layer. The position encoding part in the encoding layer uses RoPE rotation position encoding. The UniLM decoding layer uses Seq2Seq attention mask mask mode to input bidirectional modeling and output unidirectional modeling to achieve conditional generation. The input sequence is divided into a source sequence, a target sequence, and a separator between the source sequence and the target sequence. A random mask matrix is applied to the target sequence. The unmasked tokens in the source sequence and the target sequence are selected to predict the masked tokens for information reorganization tasks. Input some sentences into the Prompt-based cross-domain few-sample data pre-annotation model, and output the pre-annotation results according to the task type; The quantity, quality, and completion of the labeled data are counted according to different tasks, and multiple pre-labeling results are combined to display them in the form of knowledge graphs, tables, and charts.
2. A visual annotation method for text and image data according to claim 1, characterized in that: The initial context representation formula in the UniLM decoding layer is as follows: In the formula, the input sequence is X = {x0, x1, ..., x l }, x l Indicates the length of the sequence.
3. A visual annotation method for text and image data as claimed in claim 2, characterized in that: The final representation of the initial context representation after stacking n layers is as follows:
4. A visual annotation method for text and image data according to claim 1, characterized in that: When some sentences are input into the Prompt-based cross-domain few-sample data pre-annotation model, the label smoothing loss function with the following formula is used: In the formula, y i represents the i-th element of the true label vector, p i is the probability value of the i-th element predicted by the model, ε is a smoothing coefficient less than 1, and k is the number of categories.
Citation Information
Patent Citations
Judicial case knowledge graph construction method of dependency syntactic analysis relation extraction model
CN110597999A
Knowledge graph construction method and system for enclosed switchgear
CN112883197A
Judicial document index extraction method based on few-sample comparative learning
CN115878777A
Supervised Summarization and Structuring of Unstructured Documents
US20240012842A1
Multimodal few-shot learning with frozen language models
US20240282094A1
Cited By
Digital marketing method and system, storage medium and computer
CN120278748A
Offshore wind power corpus processing method and device
CN121074911A