A multimodal annotation method based on deep learning
Through deep learning technology and multimodal annotation method, the problems of low efficiency and poor accuracy of multimodal data are solved, and efficient and accurate annotation of images and text data are achieved, which is suitable for text annotation, image annotation and knowledge graph construction.
Patent Information
- Application Number
- CN202411157253.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-08-22
AI Technical Summary
The existing multimodal data annotation methods have problems with low labeling efficiency and poor accuracy, especially when dealing with complex correlations between different modal data, traditional methods require a lot of manpower and are prone to subjective errors, making it difficult to achieve efficient and accurate labeling.
The multimodal annotation method based on deep learning is adopted, and the image and text data are jointly processed by the multimodal annotation module, model training module, intelligent annotation module and annotation echo module using the deep learning model, feature extraction and fusion is used using the YOLO algorithm and word embedding model, and multimodal loss function trained with the weighted fusion to realize automated annotation and real-time echo.
It improves the efficiency and accuracy of multimodal data annotation, and can efficiently and accurately label images and text data. It is suitable for text annotation, image annotation and knowledge graph construction and provides an efficient and reliable multimodal data annotation solution.
Smart Images

Figure CN119202934B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a system and method for multimodal data annotation, which uses a target detection algorithm of a deep learning model to speed up the efficiency of annotators in annotating data to be annotated. Background Art
[0002] Multimodal data annotation is a key research area in computer vision and data science. It involves assigning semantic labels, metadata, or annotations to data containing multiple types of information, such as text, images, audio, and video. These annotations enable computer systems to understand and process this information. The goal of multimodal data annotation is to associate data with semantic information, such as categorical labels, descriptive text, and sentiment polarity, to facilitate computer systems' understanding and analysis of the data. This annotation is typically performed by human annotators, who add appropriate labels and annotations based on the data's content and characteristics. However, manual annotation is associated with high costs, time consumption, difficulty in data integration, and challenges in consistency and standardization. The importance of multimodal data annotation lies in improving machine understanding capabilities. When computer systems can understand data from different modalities, they can better process and analyze it, thereby improving performance in various applications.
[0003] With the rapid development of digital technology, the acquisition and application of multimodal data has become commonplace in real life and across various industries. This multimodal data typically includes information in different forms, such as images and text, which are interrelated. The comprehensive utilization of this information can help achieve more accurate and comprehensive data analysis, identification, and application. However, the labeling of multimodal data faces many challenges. Traditional labeling methods typically require a large amount of manpower, are prone to subjective errors during the labeling process, and are difficult to guarantee in terms of labeling efficiency and accuracy. Furthermore, the complex correlations between different modal data also increase the complexity of the labeling task.
[0004] Existing technologies often rely on manual annotation. Even when automated annotation tools are utilized to some extent, problems such as low annotation efficiency and poor accuracy still exist. Therefore, there is an urgent need for a method and system that can efficiently and accurately annotate multimodal data. Providing an automated, efficient, and accurate multimodal data annotation system and method is an urgent task to meet the growing demand for multimodal data annotation and promote the application and development of multimodal data in various fields.
[0005] Patent publication number CN113535949A discloses a multimodal joint event detection method based on images and sentences. To identify events from both images and sentences, existing unimodal datasets can be used to train image and text event classifiers. Furthermore, an image-sentence matching module is trained using existing image-title pairs to identify images and sentences with the highest semantic similarity in multimodal articles, thereby obtaining feature representations of image entities and words in the common space. Finally, the model is tested using a small amount of multimodal annotated data, and a shared event classifier is used to obtain the events and their types described in the images and sentences, respectively.
[0006] Patent publication number CN115311512A discloses a data annotation method, apparatus, device, and storage medium. The dataset to be annotated is input into a preset automatic annotation model to obtain a first dataset; the annotation data of the same frame in the first dataset are fused and completed to obtain a second dataset; redundantly annotated data in the second dataset is determined, and the second dataset is cross-validated and integrated based on the redundantly annotated data to obtain a completed dataset. Point cloud and image data are acquired through a collector, multimodal annotation is used to cross-validate and enrich semantic information, and the multimodal annotations are fused.
[0007] Patent publication number CN115937738A discloses a training method, apparatus, device, and storage medium for a video annotation model. The method includes: acquiring video data and extracting keyframes from the video data; performing feature extraction on the frames to obtain feature data for the frames in different modalities; constructing subgraphs corresponding to the different modalities based on the feature data in different modalities; performing aggregation operations on the subgraphs corresponding to the different modalities using a graph neural network to obtain a target graph; obtaining predicted annotation results of the graph neural network for video annotation of the keyframes based on the target graph; and training the graph neural network until convergence based on the predicted annotation results and the actual annotation results of the keyframes to obtain a video annotation model.
[0008] Looking at existing multimodal annotation tools, the development of intelligent models has been relatively slow. Most tools can only annotate single files, such as text, images, or videos. The file processing process is complex, the annotation format is simple, and they cannot add information about the relationships between annotation tags. Annotation efficiency is also a key evaluation metric for a tool. Existing tools have limited development potential for intelligent annotation. Incorporating deep learning models can significantly automate annotation, improve accuracy, enable multimodal fusion, and enhance real-time performance and adaptability.
[0009] Different file contents, such as text and images, require completely different annotation requirements, and annotation models need to be specifically trained for these different contents. To address this issue, the present invention provides a multimodal annotation method based on deep learning, which can directly perform multimodal information annotation on various forms of data, train and optimize the model, and display the annotated information in real time. Summary of the Invention
[0010] To address the above issues, the present invention provides a multimodal annotation method based on deep learning. This method uses deep learning technology to efficiently and accurately annotate data in multiple modalities, including images and text.
[0011] The technical solution of the present invention is:
[0012] A multimodal annotation method based on deep learning, comprising the following steps:
[0013] 1) Use the multimodal annotation module to annotate data content:
[0014] 11) Labeling task definition: clearly define the specific content and standards of the labeling task, determine the data type and labeling system to be labeled, and ensure that the labelers understand the requirements of the labeling task;
[0015] 12) Selection of annotation format: Choose an appropriate multimodal annotation format to ensure that the annotation format can meet the requirements of the annotation task.
[0016] 13) Integration of annotation results: Integrate the annotation results into a unified data set to ensure that the organization and format of the data meet the requirements of subsequent tasks.
[0017] 2) Use deep learning models to learn the marked content:
[0018] 21) Collect annotated datasets, including input data and corresponding annotation information. Ensuring the quality of the dataset and the accuracy of the annotations is crucial for training deep learning models;
[0019] 22) Select an appropriate deep learning model and determine its structure and parameters based on the characteristics of the public dataset arXiv Dataset and the complexity of the annotation task. The arXiv Dataset is a database of 1.7 million articles, including relevant features such as article title, author, category, abstract, and PDF full text. The data is stored in JSON format.
[0020] 23) Use the prepared dataset to train the deep learning model. During the training process, the weights and parameters of the model are adjusted through the backpropagation algorithm and optimizer, so that the model gradually learns the characteristics and patterns of the annotation content. In order to improve the model's ability to learn the annotation content when processing multimodal data, the loss function is improved so that it can better capture the relationship and characteristics between each modality. The present invention uses a weighted fusion multimodal loss function:
[0021] L=α·L text +β·L img +γ·L fusion
[0022] Among them, L text is the loss for text modality, L img is the loss of image data, L fusion is the loss of modal fusion, and α, β, and γ are different weight coefficients.
[0023] 24) Use the validation dataset to validate the trained model, evaluate the model’s performance and generalization ability, and tune the model based on the validation results to improve the model’s performance;
[0024] 25) Apply the trained deep learning model to new unlabeled data to perform inference and prediction. The model automatically labels or classifies the input data and generates corresponding output results.
[0025] Furthermore, for the annotation of multimodal data forms, different annotation methods are provided, such as text, image, audio, etc., which are annotated using different annotation methods, and annotated datasets are constructed by manual annotation.
[0026] Furthermore, it is necessary to select a suitable annotation tool and use this system as an annotation tool to perform data annotation. This system provides annotation methods for different data, such as image types, which can perform target recognition and classification, key point annotation, image description, region annotation, and attribute annotation.
[0027] Furthermore, the deep learning module trains, predicts and evaluates the annotated data, and performs incremental training on the deep learning model based on the updated data obtained by the annotators modifying the information identified by the intelligent annotation.
[0028] Furthermore, the labeled data is collected and stored in a designated database for use in the next step of model training. The storage database uses a non-relational database.
[0029] Furthermore, the YOLO algorithm is used on the collected dataset to pre-train the deep learning model, train and tune the model parameters, and the YOLO algorithm is combined with the manually annotated labels to train the model to detect objects in the image and identify their categories and attributes.
[0030] Furthermore, the trained deep learning model is integrated with YOLO and the annotation model to create a multimodal framework that combines the YOLO object detection and annotation models. This framework combines YOLO and the intelligent annotation model into a single multimodal annotation framework, which can be used to perform multimodal annotation on new, unannotated files. This framework uses the YOLO model to detect objects in images, extract features for each detected object, and predict the category and attributes of each object using the text annotation model. The framework uses the YOLO algorithm for object detection and recognition in the image modality, with the recognition results transmitted to the database in JSON format, and the annotation model displays them on the front end. In the text modality, the deep learning model directly extracts and annotates entities, also saving the data in JSON format.
[0031] A server, characterized in that it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the above method.
[0032] A computer-readable storage medium stores a computer program thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0033] The present invention is divided into four modules: multimodal annotation module, model training module, intelligent annotation module and annotation echo module. Figure 1 A module demonstrating multimodal annotation methods.
[0034] The multimodal annotation module uses various annotation methods to annotate content, including text, tables, and images, locate their location and information, and store the annotated information. The model training module trains models based on annotated data from the database, learning and optimizing model parameters. The intelligent annotation module applies a trained model to automatically annotate unannotated documents and stores the annotation results in the database. The annotation display module displays annotation results at the corresponding locations in the document based on the model output.
[0035] The four modules of the present invention are described in detail below.
[0036] 1. Multimodal Annotation Module
[0037] The multimodal annotation module needs to complete the annotation of various data formats in the document and distinguish different data types. First, data collection is carried out and the parts of the document content that need to be annotated are described in detail.
[0038] Define the annotation scheme and determine the type and format of annotation. Text annotation types include: classification, which assigns text to predefined categories. Entity recognition: Identify specific types of entities in the text, such as names, places, or organizations. Relationship extraction: Identify the relationships between entities in the text. Then use the annotation module to annotate the text. For images and other forms of data, the annotation scheme is similar, but the form of annotation needs to be changed accordingly. Images, as a visual medium containing digital data, cannot be directly expressed intuitively using text. We stipulate that the multimodal annotation module uses coordinate selection for annotation and can be associated with relationships.
[0039] Categorical annotation is a text data annotation technique that categorizes text data into different categories or labels and assigns a label to each category. To assign a corresponding category label to each piece of text, the annotator must categorize the text according to a predefined classification system. Saving categorical annotations is typically relatively simple, requiring only the text content and labels to be stored, with additional information added based on the actual scenario.
[0040] Bounding box annotation is an image annotation technique that involves drawing rectangular boxes around objects in an image and assigning them a class label. It's one of the most commonly used annotation types in computer vision and is used to train object detection and recognition models. Use the selection tools in the annotation module to draw bounding boxes on the image and assign each box a class label indicating the object's category. Class labels can be predefined or custom. Save the annotations in your desired format.
[0041] A bounding box is usually represented as a quadruple (x, y, w, h).
[0042] in:
[0043] x and y are the coordinates of the upper left corner of the bounding box;
[0044] w and h are the width and height of the bounding box.
[0045] Polygon annotation is an image annotation technique that uses polygonal shapes to outline objects or regions in an image. It is more sophisticated than bounding box annotation and can more accurately represent the shape of an object. Use the polygon drawing tool in the annotation module to draw bounding boxes on the image. Each bounding box is assigned a class label indicating the class to which the object belongs. A polygon can be represented by a set of ordered vertices:
[0046] P=[(x1,y1),(x2,y2),...,(xn,yn)],
[0047] in:
[0048] (xi,yi) is the coordinate of the i-th vertex;
[0049] n is the number of vertices of the polygon.
[0050] As described above, after annotating the original data documents, the annotated data is stored in the database. The database is responsible for managing and processing data, including operations such as storage, retrieval, update, and deletion. The database automatically organizes and indexes data to achieve fast and efficient access. The annotated data is stored in the database using a non-relational database, and each data entry is stored in the database in JSON format. JSON data is in the key-value pair format, which is very convenient to store in a non-relational database. Furthermore, converting the data to JSON format when retrieving it makes it easier for downstream modules to use and analyze it, and it also significantly improves the training speed of the model.
[0051] Attachment Figure 2 Shows the architecture of the multimodal annotation module.
[0052] 2. Model Training Module
[0053] To handle multimodal inputs of image and text data, we can design a joint deep learning model consisting of two branches: an image processing branch and a text processing branch. These two branches are responsible for processing image and text data respectively, and their features are fused in subsequent layers to achieve joint multimodal processing.
[0054] 1. Image processing branch
[0055] Image feature extractor: This uses a convolutional neural network as the foundational model for the image processing branch, extracting features from image data. Pooling layer: This performs a pooling operation on the feature maps output by the convolutional layer, reducing feature dimensionality and improving model robustness. Fully connected layer: This connects the output of the pooling layer to one or more fully connected layers to learn higher-level image feature representations.
[0056] 2. Text processing branch
[0057] Word Embedding Layer: Represents text data as word embedding vectors, using either a pre-trained word embedding model or training your own. Recurrent Neural Network: Processes sequences of word embedding vectors using a recurrent neural network model to capture the semantic information of the text data. Pooling Layer or Global Average Pooling: Pools the output of the recurrent neural network to obtain a fixed-length representation of the text data.
[0058] 3. Multimodal Fusion
[0059] Feature fusion layer: Fusion of features from the image processing branch and the text processing branch using methods such as concatenation, addition, and weighted averaging. Fully connected layer: Inputs the fused features into the fully connected layer to learn the joint representation of multimodal features.
[0060] The model is trained and optimized using a training set with annotated image and text data, and parameter optimization is performed using a loss function and optimizer. The model is validated and fine-tuned using a validation set to prevent overfitting and improve generalization. This model can process both image and text data simultaneously and effectively fuse their features, enabling multimodal input processing and joint learning.
[0061] Attachment Figure 3 This is the overall architecture of the model training module.
[0062] 3. Intelligent Labeling Module
[0063] The intelligent annotation module intelligently annotates text, images, and other information within documents requiring annotation. In conjunction with the multimodal annotation module, it uses a trained intelligent model for automated annotation. The quality of intelligent annotation depends on the accuracy of the model. After annotation is complete, the results can be manually reviewed and modified.
[0064] Multimodal annotation is performed based on the model. The entity information and relationship information obtained by annotation are placed in the database using a unified format and displayed in the document. Annotation screening is performed based on pre-defined knowledge ontology to realize an automatic annotation tool that is convenient for users to use.
[0065] This module can automatically create annotation entities and relationship information, and can also manually add entity and relationship information that is missing after module processing. Annotation entities and relationship information that are not used after annotation is completed can be manually deleted.
[0066] Text annotation uses the Word2Vec model to vectorize the defined entity and relationship names. Based on the cosine similarity of the vectors, the entity and relationship categories marked in the intelligent annotation model are calculated for each name (including entity and relationship names). The selected entity and relationship names are then filtered for the intelligent annotation model output. The intelligent annotation module model outputs entity and relationship dictionaries. The categories in the dictionaries are used to filter the categories required for the annotation project. Annotations are generated and the annotations or descriptions of the text content are output to the user to assist in understanding the text content or performing related applications.
[0067] Multimedia information annotation involves different types of data, such as images, video, audio, and text, and utilizes common multimedia information conversion models. First, images are converted to text using an image-to-text conversion model. The image content is converted into a text description to assist in image content recognition. Next, a pre-trained convolutional neural network model is combined with a recurrent neural network. The image is fed into the network, converted into digital features, and its implicit features are extracted. The image features extracted by the convolutional neural network are used as the initial hidden state of the recurrent neural network, which then gradually generates a text sequence.
[0068] Attachment Figure 4 This is the overall architecture of the intelligent labeling module.
[0069] 4. Annotation Echo Module
[0070] The entity information and entity relationship information obtained from the model is stored in a unified format in the database and displayed on the document. It is then annotated and filtered according to predefined knowledge ontologies, making it easier for users to use automatic annotation tools. This module provides timely feedback and confirmation of annotation results during the annotation process, displaying the results in an intuitive way, such as drawing bounding boxes on the image and displaying the annotated portion in the text, so that users can clearly see the annotated content.
[0071] 1. The user establishes the annotation project ontology and relationships in the annotation tool, including the entity categories that need to be annotated in the annotation project and the relationship categories between entities.
[0072] 2. Use the Word2Vec model to vectorize user-defined entity names and relationship names, and calculate the entity and relationship categories marked in the intelligent labeling model corresponding to each name based on the cosine similarity of the vectors.
[0073] 3. Filter the output of the intelligent annotation model based on the entity and relationship names selected in step 2. The output of the intelligent annotation module model is an entity dictionary and a relationship dictionary. The categories required for the annotation project are filtered based on the category names in the dictionary.
[0074] 4. Position the selected entities and relationships based on the document’s text and text coordinate information dictionary, and locate them at the coordinates on the document.
[0075] 5. Create an intelligent annotation layer on the original document, build annotation boxes based on coordinates, and mark out entity categories and relationship categories.
[0076] Attachment Figure 5 This is the overall architecture of the annotation echo module.
[0077] Compared with the prior art, the present invention has the following positive effects:
[0078] The present invention realizes automatic labeling through deep learning technology. Compared with traditional labeling methods, the system and method of the present invention have the advantages of high labeling efficiency and high accuracy. It can be widely used in text labeling, image labeling, knowledge graph construction and other fields, providing an efficient and reliable solution for the labeling of multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 Block diagram of the multimodal annotation method.
[0080] Figure 2 This is the architecture diagram of the multimodal annotation module.
[0081] Figure 3 This is the overall architecture diagram of the model training module.
[0082] Figure 4 This is the overall architecture diagram of the intelligent labeling module.
[0083] Figure 5 This is the overall architecture diagram of the annotation echo module.
[0084] Figure 6 This is the overall architecture diagram of the general training model. DETAILED DESCRIPTION
[0085] The details of the present invention are described in detail below.
[0086] 1. Multimodal Data Conversion
[0087] 1. Text data conversion
[0088] Text vectorization uses the word embedding model Word2Vec to convert text into vector representations. As a deep learning-driven word embedding algorithm, Word2Vec successfully captures the semantic relationships and contextual information between words by mapping them into a low-dimensional real vector space, providing a powerful and flexible tool for subsequent text analysis tasks. Word2Vec primarily implements word embedding through two models: Continuous Bag of Words (CBOW) and Skip-Gram. CBOW: Predicts a word based on its contextual vocabulary, emphasizing the aggregated semantics of the word. Skip-Gram: Given a word, predicts its surrounding contextual vocabulary, focusing on the diffuse semantics of the word.
[0089] The computational method used by the model in text sequence uses the basic solution framework of statistical language model.
[0090] S=w1,w2,w3,...,w T
[0091] Its probability can be expressed as:
[0092]
[0093] That is, the joint probability of the sequence is converted into the product of a series of conditional probabilities. The problem becomes how to predict the conditional probabilities of these given words:
[0094] p(w t |w1,w2,...,w t-1 )
[0095] Due to its huge parameter space, such a primitive model wastes a lot of resources in practice. We chose a method that reduces the error and greatly reduces the waste of resources:
[0096] p(w t |w1,w2,...,w t-1 )≈p(w t |w t-n+1 ,...,w t-1 )
[0097] The input layer is the One-Hot encoded word vector of the context word, V is the number of words in the vocabulary, N is the dimension of the custom word vector, and C is the number of context words. Initialize a weight matrix W v×N , multiply the matrix by the left of all input One-Hot encoded word vectors to obtain vectors w1, w2, w3, w with a dimension of N c , N is set by the task. The resulting vectors w1, w2, w3, w c Add and average as the hidden layer vector h. Initialize another weight matrix W' N×V , multiply this matrix by the hidden layer vector h, and then process it through the activation function to obtain the V-dimensional vector y, where each element of y represents the probability distribution of each corresponding word.
[0098] 2. Image data conversion
[0099] The objects in the image not only have visual features, but also contain semantic information. Converting them into text data can provide additional contextual information for the object detection model, helping to more accurately identify and locate the object. Through cross-modal fusion, the model can learn the relationship between images and text, thereby improving the performance of object detection.
[0100] In the YOLO object detection algorithm, input layer data is normalized to map the pixel value range of the input image to a smaller range, typically between 0 and 1. Normalization ensures that the input data is within the same scale range, preventing large differences in pixel values between different images. This helps the network model better learn image features, improving model stability and convergence speed.
[0101] Calculate the mean and variance, normalize and calculate the mean and variance of the input tensor in a specific dimension. Assume that the shape of the input tensor X is [N, D] (N is the batch size, D is the feature dimension), calculate the mean μ and variance σ for each sample 2
[0102]
[0103]
[0104] Normalization: normalize each feature of each sample to make its mean 0 and variance 1
[0105]
[0106] Among them, ε is a small number used to prevent division by zero. Scaling and offset, scaling and offset operations are performed on the normalized features, which are completed by the learned parameters γ (scaling) and β (offset)
[0107] y i,d =γx′ i,d +β
[0108] γ and β are trainable parameters whose dimensions are the same as the dimensions of the input features.
[0109] 2. General Model Training
[0110] YOLO is an object detection algorithm that can process data from different modalities simultaneously. We have made some changes to YOLO to enable it to better detect multimodal data and obtain good model performance by training it on manually annotated datasets. Its model structure mainly includes the following components:
[0111] 1. Multi-channel input: Image data mainly consists of Mosaic image enhancement, adaptive anchor box calculation, and adaptive image scaling. Mosaic image enhancement is a commonly used data enhancement method in target detection. It can generate new training images by combining multiple different images. Randomly select four different images, reset the long side of the image, and keep the image ratio unchanged to generate a large image. Randomly generate a center point in the large image. Based on the center point of the large image, splice the four images together, perform M matrix transformation on the spliced image, crop based on the upper left corner of the large image, and transform the label coordinates of each independent image into label coordinates based on the large image. After the M matrix transformation, the final coordinates are obtained. Finally, the synthesized image and the corresponding label box are used for model training.
[0112] Mosaic data augmentation helps improve the robustness of the model because it requires the model to be able to handle overlaps and background information between different objects. This data augmentation technique is commonly used in object detection tasks in the YOLO series to improve the performance and generalization ability of the model.
[0113] To process text data, we use a pre-trained word embedding model or train our own to represent the text data as a sequence of word embedding vectors. We then use a recurrent neural network model to process these embedding vectors and capture the semantic information of the text data. For each word in each sentence, we generate a subword sequence, and the model predicts a named entity label for each subword. The tokenizer results are saved as a subword list, with a length equal to the number of words in the sentence. Entity recognition is then performed to extract the named entities and their categories from the sentence. The name, length, offset within the sentence, entity type, and original sentence of each extracted entity are stored in a dictionary.
[0114] 2. Backbone layer: The Backbone layer's image operations consist of a Focus structure and a CSP structure, and incorporate a CLIP (Constrastive Language-Image Pre-training) model structure to match the relationship between images and text.
[0115] The Focus structure in YOLO is a convolutional neural network layer used for feature extraction. It is used to compress and combine the information in the input feature map to extract higher-level feature representations. The Focus structure is a special convolution operation in YOLO. It is used as the first convolutional layer in the network to downsample the input feature map to reduce the amount of computation and parameters.
[0116] The Focus structure can divide the input feature into four sub-graphs and concatenate the four sub-graphs to obtain a smaller feature map. Assuming that the size of the input feature map is N×N×C, where N is the size of the feature map and C is the number of channels, the calculation process of the Focus structure is: separate the input feature map into four sub-graphs and obtain two sub-graphs of size The feature map is recorded as x and y; the horizontal and vertical convolution operations with a step size of 2 are performed on x and y respectively, and two sizes are obtained. The feature map is recorded as x' and y'; x' and y' are spliced together to obtain a size of The feature map is recorded as z; the horizontal and vertical convolution operations with a step size of 2 are performed on z to obtain a size of The feature map is the output of the Focus structure.
[0117] The Cross Stage Partial (CSP) architecture is a key component of YOLO. It effectively reduces network parameters and computational overhead while improving feature extraction efficiency. The core idea behind the CSP architecture is to split the input features into two parts: one part is processed by a small convolutional network, while the other part is directly processed by the next layer. The two feature maps are then concatenated and used as input to the Neck network in the next layer.
[0118] The CSP architecture involves splitting the input feature map into two parts, one for sub-network processing and the other for direct processing in the next layer. In the sub-network, a convolutional layer is used to compress the input feature map, followed by a series of convolution operations, and finally a convolutional layer for expansion. This allows for the extraction of relatively few high-level features. In the next layer, the Neck network, the feature map processed by the sub-network is concatenated with the directly processed feature map, followed by another series of convolution operations. This combines low-level, detailed features with high-level, abstract features, improving feature extraction efficiency.
[0119] The CLIP structure is added to match the relationship between image and text, and an automatic labeling method is introduced to generate region-text pairs instead of directly using image-text pairs for pre-training. Region detection is performed on the image to identify different objects and their locations in the image. Specifically, the labeling method includes three steps: (1) Noun phrase extraction: We first use the n-gram algorithm to extract noun phrases from the text; (2) Pseudo-labeling: A pre-trained open vocabulary detector is used. The vocabulary detector contains multiple categories and is mainly used for comparison results. Pseudo boxes are generated for a given noun phrase of each image, thereby providing rough region-text pairs. (3) Filtering: Use pre-trained CLIP to evaluate the relevance of image-text pairs and region-text pairs, and filter out pseudo annotations and images with low relevance. We further filter redundant bounding boxes by combining methods such as non-maximum suppression (NMS).
[0120] The Neck network aggregates feature maps of different scales through top-down and bottom-up approaches, creating a rich multi-scale feature pyramid. The goal is to ensure that features at each level possess both high semantic information and precise spatial location information, thereby improving the detection of small objects. This network is designed to handle multi-scale objects and improve the detection of both small and large objects. The network primarily performs a series of network layers that mix and combine image features and pass them to the output.
[0121] 3. Multi-channel output: Image data is primarily used in YOLO's output, which is a prediction box. Each prediction box consists of a confidence score, a class probability, and a bounding box position. YOLO's output can output information such as the location, size, and class of objects in the image, allowing for subsequent tasks such as object recognition and tracking.
[0122] YOLO uses the Intersection over Union (IoU) loss function for bounding box detection, which measures the difference between the predicted bounding box and the ground-truth bounding box. IoU loss is a variation of Intersection over Union (IoU), a metric used to measure the degree of overlap between the predicted and ground-truth bounding boxes. In object detection, IoU is often used to assess the overlap between the predicted and ground-truth bounding boxes to determine if the predicted box is correct.
[0123] IoU is expressed as the intersection-over-union ratio of the real box and the predicted box in target detection and is defined as follows:
[0124]
[0125] Where A and B represent the predicted bounding box and the ground-truth bounding box, respectively. This formula calculates the ratio of the intersection area of the two bounding boxes to the union area. When the ground-truth box and the predicted box completely overlap, the IoU is 1. Therefore, the bounding box regression loss function in object detection is defined as:
[0126] LossIoU=1-IoU
[0127] Specifically, for each predicted bounding box, we calculate its IoU value with all true bounding boxes, and then select the true bounding box with the largest IoU as its corresponding matching target, thereby calculating its IoU loss.
[0128] In the target detection task, an object may be detected by multiple prediction boxes. In order to avoid multiple detections of the same object, it is necessary to filter the repeated prediction boxes. This process is called non-maximum suppression (NMS). After outputting the results, YOLO will perform NMS on the overlapping target boxes to obtain the final detection results. NMS will compare the IoU values of the remaining prediction boxes except the retained prediction boxes with the prediction box with the highest confidence in the same category. If the IoU value of a prediction box and the selected prediction box exceeds the preset threshold, then the prediction box will be suppressed, that is, it will not be considered a valid detection result.
[0129] NMS essentially sorts all predicted boxes by confidence, from high to low. Starting with the most confident box, it then iterates through each prediction box, determining whether the IOU value between that prediction box and all subsequent prediction boxes exceeds a certain threshold. If the IOU value is greater than the threshold, the prediction box is removed from the candidate box list; otherwise, it is retained. It then iterates through the next prediction box, repeating the above steps until all prediction boxes have been iterated. Ultimately, the remaining prediction boxes are the result of NMS processing, meaning that each object corresponds to only one prediction box.
[0130] The text data is output probabilistically at the output end. After probability prediction in the network, the recognition result with the highest probability is selected for output. The prediction results are stored in the database in the form of a dictionary. The entity information with the highest probability predicted by the named entity model is linked to the corresponding entity number in the entity dictionary in the database.
[0131] Attachment Figure 6 The overall architecture of the general training model.
[0132] 3. Intelligent annotation display
[0133] The core function of intelligent annotation is to intelligently annotate various multimodal data within documents to improve information comprehension and utilization efficiency. The design concept of this module is derived from in-depth research and practice in multimodal data processing. It combines the multimodal annotation technology of the YOLO model to build an intelligent annotation system that can automatically identify entities in documents and analyze the relationships between entities.
[0134] First, the foundation is the various datasets stored in databases, which contain a large number of manually annotated documents and provide sufficient data support for model training. Different types of data are transformed in different ways within the datasets, which can then be input into the model for training and then used for prediction.
[0135] For text data, the BIO sequence annotation rule is used. This rule provides a clear logical framework for entity annotation. The first word of an entity is annotated as B (Begin), non-first words are annotated as I (Inside), and non-entity words are annotated as O (Outside). In this way, each training sentence is represented as a sequence consisting of a BX label, an IX label, and an O label, where X represents the entity category. This annotation method not only helps the model understand the entity structure in the text, but also provides a strong training foundation for subsequent named entity recognition.
[0136] For image data, the target bounding box is usually represented by a rectangular box that surrounds the target object and is used to identify the location and size of the target. It is represented by the coordinates of the upper left corner of the rectangle (x1, y1) and the coordinates of the lower right corner (x2, y2). In addition, the center coordinates of the bounding box rectangle (x c ,y c ) and width and height (w,h). By annotating the target bounding box, the model can learn how to accurately locate the target in the image and identify and classify it.
[0137] During the training phase, the YOLO model was used. This model performs exceptionally well when processing diverse data types, effectively leveraging information from large-scale data to improve performance on multimodal labeling tasks. The specific training process, including model architecture design, hyperparameter selection, and training strategy development, is described in detail in the implementation section of this manual.
[0138] After model training is complete, the intelligent annotation module can intelligently annotate new documents. The annotation results are stored in the database and displayed in the annotated document. This model's annotation capabilities improve the efficiency of annotating various types of data within a document. After annotation is complete, any discrepancies can be manually corrected, while also improving the accuracy of subsequent unannotated data. Overall, the design and implementation of the intelligent annotation module is based on a deep understanding of data, providing users with efficient and accurate document annotation services.
[0139] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Those skilled in the art may modify or make equivalent substitutions for the technical solutions of the present invention without departing from the spirit and scope of the present invention. The scope of protection of the present invention shall be based on the claims.
Claims
1. A multimodal annotation method based on deep learning, comprising the following steps: 1) Use the multimodal annotation module to annotate the multimodal data to obtain an annotated dataset; 2) using the YOLO algorithm to train a deep learning model using the dataset, and adjusting the weights and parameters of the deep learning model through a backpropagation algorithm and an optimizer so that the deep learning model gradually learns the characteristics and patterns of the annotated content; 3) For a multimodal data to be annotated, input it into the trained deep learning model for inference and prediction to generate the annotation results of the multimodal data to be annotated; The YOLO algorithm includes a multi-channel input, a backbone layer, and a multi-channel output; the backbone layer includes a Focus structure, a CSP structure, and a CLIP structure; and for each annotated multimodal data in the dataset, the training method is: 21) The multi-channel input terminal represents the text data in the multimodal data as a sequence of word embedding vectors, Using a recurrent neural network model to capture semantic information of the text data from a word embedding vector sequence and predicting a named entity label for each subword of each sentence in the text data; the multi-channel input end sequentially performs mosaic image enhancement, adaptive anchor box calculation, and adaptive image scaling on the image data in the multimodal data to generate a composite image and its label box, and then executes step 22); 22) Using the Focus structure to extract low-level features of the synthesized image and input them into the CSP structure; 23) The CSP structure divides the input low-level features into two parts, processes one part through a convolutional network to obtain high-level features, and then splices the high-level features with the other part of the low-level features and inputs them into the Neck network; the Neck network extracts the input spliced features and inputs them into the CLIP structure; 24) The CLIP structure generates region-text pairs based on the input features and the text data in the multimodal data. The method for generating region-text pairs is as follows: 241) extracting noun phrases from text data using an n-gram algorithm; 242) generating pseudo boxes for each given noun phrase using an open vocabulary detector to obtain rough region-text pairs; 243) evaluating the correlation between image-text pairs and region-text pairs, and filtering pseudo annotations and images with correlations below a set threshold to obtain predicted boxes for the multimodal data; 25) The multi-channel output end uses a non-maximum suppression method to process overlapping prediction boxes in the prediction boxes of the multimodal data to obtain a final detection result.
2. The method according to claim 1, characterized in that The method for annotating multimodal data using the multimodal annotation module is as follows: first, according to the specific content and standards of the set annotation task, determine the data type and label system that need to be annotated; then, select the multimodal annotation form according to the annotation task, annotate the multimodal data, and obtain the data set.
3. The method according to claim 1 or 2, characterized in that The loss function in the optimizer is LossIoU=1-IoU; IoU is the intersection-over-union ratio between the true box labeled in a multimodal data and the predicted box of the multimodal data.
4. The method according to claim 1 or 2, characterized in that The structure and parameters of the deep learning model are determined according to the characteristics of the multimodal data and the labeling task.
5. The method according to claim 1 or 2, characterized in that The multimodal data includes images and text.
6. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Multi-modal joint event detection method based on pictures and sentences
CN113535949A
Data annotation method and device, equipment and storage medium
CN115311512A
Video annotation model training method and device, equipment and storage medium
CN115937738A
Multi-modal model training method, device and equipment and readable storage medium
CN116561570A