An intelligent labeling method and device based on multi-modal collaboration and a storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN EXTREME VISION TECH CO LTD
- Filing Date
- 2026-06-23
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies suffer from insufficient accuracy and stability in multimodal data annotation, especially due to inconsistencies in candidate annotation results caused by differences in the analytical perspectives of different prompts.
By configuring multiple sets of differentiated prompts and performing multiple rounds of semantic parsing, a candidate label set is constructed and mapped to a preset semantic space to form a semantic action field, the target annotation result is determined, and video data is generated through the target sample set and cross-modal alignment data, and the annotation process is iteratively updated.
It improves the accuracy and stability of multimodal data annotation results, enhances the overall utilization efficiency of multimodal data, and forms a closed-loop annotation and data generation system.
Smart Images

Figure CN122433749A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to an intelligent annotation method, device and storage medium based on multimodal collaboration. Background Technology
[0002] With the development of artificial intelligence technology, multimodal data annotation has been widely used in scenarios such as model training, data management, and content generation. Currently, intelligent models are typically used to analyze multimodal data such as images, text, audio, or video to generate corresponding annotation results.
[0003] To improve annotation accuracy, some technologies use different prompts to analyze the data to be annotated for the same annotation task, thereby obtaining multiple candidate annotation results. However, because different prompts have different analytical perspectives, different candidate annotation results are often generated for the same data to be annotated. Existing technologies usually select the final annotation result from multiple candidate annotation results according to preset rules, but in practical applications, problems such as insufficient accuracy and stability of the annotation results still easily occur. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides an intelligent annotation method, apparatus, and storage medium based on multimodal collaboration.
[0005] The technical solution provided in this application is described below:
[0006] The first aspect of this application provides an intelligent annotation method based on multimodal collaboration, including: Configure multiple sets of differentiated prompts for the same annotation task, and perform multiple rounds of semantic parsing on the multimodal data to be annotated based on the multiple sets of differentiated prompts to obtain multiple candidate annotation results; Filter through multiple candidate annotation results to determine the candidate label set; The candidate label set is mapped to a preset semantic space, and a semantic action field is constructed based on the mapping results; The target annotation result is determined based on the semantic field. Based on the target annotation results and the corresponding multimodal data, construct a target sample set and cross-modal alignment data; Video data is generated based on the target sample set and cross-modal alignment data; The target sample set and the cross-modal alignment data are iteratively updated using the video data.
[0007] Optionally, the candidate label set is mapped to a preset semantic space, and a semantic action field is constructed based on the mapping result, including: The candidate tag set is mapped to a preset semantic space to obtain the semantic coordinates of each candidate tag in the preset semantic space; Based on the semantic coordinates, the credibility information of the candidate tags, and the semantic association between the candidate tags, the semantic influence strength of each candidate tag in the semantic space is determined. Each candidate label is used as a semantic action source, and a semantic action field is constructed based on the semantic coordinates and the semantic action intensity.
[0008] Optionally, based on the semantic coordinates, the credibility information of the candidate tags, and the semantic association between the candidate tags, the semantic influence strength of each candidate tag in the semantic space is determined, including: Obtain the credibility information of each candidate label; The semantic relationships between candidate tags are determined based on the distribution of candidate tags in the preset semantic space; The semantic influence strength of each candidate label is determined based on the credibility information and the semantic association.
[0009] Optionally, the semantic association between candidate tags is determined based on the distribution of candidate tags in the preset semantic space, including: Obtain the semantic coordinates of each candidate tag in the preset semantic space, and calculate the semantic distance between any two candidate tags based on the semantic coordinates; The semantic relevance between candidate tags is determined based on the semantic distance. The semantic association between candidate tags is established based on the semantic association degree.
[0010] Optionally, multiple sets of differentiated prompts can be configured for the same annotation task, and multiple rounds of semantic parsing can be performed on the multimodal data to be annotated based on the multiple sets of differentiated prompts to obtain multiple candidate annotation results, including: Based on the same annotation task, construct target description prompts, scene constraint prompts, and semantic reasoning prompts respectively; The multimodal data to be labeled is semantically parsed based on the target description prompts, the scene constraint prompts, and the semantic reasoning prompts to obtain multiple candidate labeling results.
[0011] Optionally, a target sample set and cross-modal alignment data are constructed based on the target annotation results and the corresponding multimodal data, including: Based on the target annotation results, target semantic entities are extracted from the multimodal data, and unified semantic anchors between different modalities are constructed using the target semantic entities; The semantic mapping relationship between different modalities is determined based on the unified semantic anchor point; The target sample set and cross-modal aligned data are constructed using the semantic mapping relationship.
[0012] Optionally, video data is generated based on the target sample set and cross-modal alignment data, including: Extract target semantic information from the target sample set; Establish association constraint relationships between different modes based on the cross-modal alignment data; Construct a semantic constraint chain based on the target semantic information and the associated constraint relationship; Video data is generated based on the semantic constraint chain.
[0013] Optionally, iteratively updating the target sample set and the cross-modal alignment data using the video data includes: Perform semantic analysis on the video data to obtain new candidate tags; The newly added candidate tags are mapped to the preset semantic space, and the semantic association between the newly added candidate tags and the candidate tag set is determined; The target sample set is supplemented based on the semantic association, and the cross-modal alignment data is updated using the supplemented target sample set.
[0014] Optionally, determining the target annotation result based on the semantic action field includes: The comprehensive influence intensity of each semantic region in the semantic field is obtained, and the semantic region with the largest comprehensive influence intensity is determined as the semantic convergence region. Target annotation results are generated based on the candidate labels in the semantic convergence region.
[0015] A second aspect of this application provides an intelligent annotation device based on multimodal collaboration, comprising: The parsing unit is used to configure multiple sets of differentiated prompt information according to the same annotation task, and to perform multiple rounds of semantic parsing on the multimodal data to be annotated according to the multiple sets of differentiated prompt information to obtain multiple candidate annotation results; The first determining unit is used to filter multiple candidate annotation results and determine the candidate label set; The first construction unit is used to map the candidate label set to a preset semantic space and construct a semantic action field based on the mapping result; The second determining unit is used to determine the target annotation result based on the semantic action field; The second construction unit is used to construct a target sample set and cross-modal alignment data based on the target annotation results and the corresponding multimodal data; The generation unit is used to generate video data based on the target sample set and cross-modal alignment data; An update unit is used to iteratively update the target sample set and the cross-modal alignment data using the video data.
[0016] A third aspect of this application provides an intelligent annotation device based on multimodal collaboration, comprising: Processor, memory, input / output units, and bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, which the processor invokes to perform the method as described in the first aspect and any one of the first aspects.
[0017] A fourth aspect of this application provides a computer-readable storage medium on which a program is stored, which, when executed on a computer, performs the method as described in the first aspect and any one of the first aspects.
[0018] As can be seen from the above technical solutions, this application has the following beneficial effects: This application constructs multiple sets of differentiated prompts for the same annotation task to obtain multiple candidate annotation results, and constructs a semantic action field based on the candidate label set to achieve semantic collaborative analysis and target annotation result generation among multiple candidate annotation results. Furthermore, it uses the target annotation results to construct a target sample set and cross-modal alignment data, then generates video data based on the target sample set and cross-modal alignment data, and then uses the video data to update the target sample set and cross-modal alignment data in reverse, thereby forming a collaborative linkage mechanism between annotation result generation, multimodal association establishment, data generation, and data feedback update.
[0019] This enables a mutually reinforcing data flow relationship, allowing the target annotation results not only to serve as annotation outputs but also to participate in subsequent sample construction, cross-modal association establishment, and data generation processes. The subsequently generated data, in turn, continuously feeds back into the preceding annotation process, achieving synchronous enhancement of annotation capabilities. This effectively improves the accuracy and stability of multimodal data annotation results and enhances the overall utilization efficiency of multimodal data. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of an embodiment of the intelligent annotation method based on multimodal collaboration in this application; Figure 2This is a schematic flowchart of a specific embodiment of step 103 in the intelligent annotation method based on multimodal collaboration in this application; Figure 3 This is a schematic flowchart of a specific embodiment of step 202 in the intelligent annotation method based on multimodal collaboration in this application; Figure 4 This is a schematic flowchart of a specific embodiment of step 101 in the intelligent annotation method based on multimodal collaboration of this application; Figure 5 This is a schematic flowchart of a specific embodiment of step 105 in the intelligent annotation method based on multimodal collaboration in this application; Figure 6 This is a schematic flowchart of a specific embodiment of step 106 in the intelligent annotation method based on multimodal collaboration in this application; Figure 7 This is a schematic flowchart of a specific embodiment of step 107 in the intelligent annotation method based on multimodal collaboration in this application; Figure 8 This is a schematic diagram of an embodiment of the intelligent annotation device based on multimodal collaboration according to this application; Figure 9 This is a schematic diagram of another embodiment of the intelligent annotation device based on multimodal collaboration of this application. Detailed Implementation
[0022] It should be noted that the multimodal data intelligent annotation, video data generation and data iterative optimization method disclosed in this embodiment is not limited to a single execution subject, and can be adapted to any type of data processing carrier such as terminal devices, servers, cloud computing platforms, artificial intelligence processing systems or distributed data processing clusters that have the capabilities of data processing, semantic parsing, model calculation and data iterative update.
[0023] In practical applications, the appropriate execution entity can be flexibly selected based on the volume of multimodal data processing, computational accuracy requirements, real-time requirements, and deployment scenarios. Lightweight, small-batch multimodal data annotation and iterative processing tasks can be completed independently by local intelligent terminals or embedded processing devices; large-scale, high-dimensional, massive multimodal data batch parsing, annotation optimization, video generation, and iterative processing tasks can be deployed on cloud servers, distributed computing clusters, or professional AI data processing platforms.
[0024] The core process, processing logic, and technical principles of this method are not limited by the hardware form, deployment method, device model, or computing power specifications of the executing entity. Any technical carrier that can implement the subsequent processes of multiple sets of differentiated prompt configurations, multi-round semantic parsing, candidate tag filtering, semantic action field construction, target result determination, sample set and cross-modal aligned data construction, video data generation, and data iterative update is within the scope of protection of this technology. The following section elaborates on this method in conjunction with specific processing steps.
[0025] To improve annotation accuracy, some technologies use different prompts to analyze the data to be annotated for the same annotation task, thereby obtaining multiple candidate annotation results. However, because different prompts have different analytical perspectives, different candidate annotation results are often generated for the same data to be annotated. Existing technologies usually select the final annotation result from multiple candidate annotation results according to preset rules, but in practical applications, problems such as insufficient accuracy and stability of the annotation results still easily occur.
[0026] Based on this, this application provides an intelligent annotation method, device and storage medium based on multimodal collaboration, which can effectively improve the accuracy and stability of multimodal data annotation results and enhance the overall utilization efficiency of multimodal data.
[0027] Please see Figure 1 This application discloses an intelligent annotation method based on multimodal collaboration, the method comprising: 101. Configure multiple sets of differentiated prompt information according to the same annotation task, and perform multiple rounds of semantic parsing on the multimodal data to be annotated according to the multiple sets of differentiated prompt information to obtain multiple candidate annotation results; 102. Filter the multiple candidate annotation results to determine the candidate label set; 103. Map the candidate label set to a preset semantic space, and construct a semantic action field based on the mapping result; 104. Determine the target annotation result based on the semantic field; 105. Construct a target sample set and cross-modal alignment data based on the target annotation results and the corresponding multimodal data; 106. Generate video data based on the target sample set and cross-modal alignment data; 107. Iteratively update the target sample set and the cross-modal alignment data using the video data.
[0028] In this embodiment, multiple sets of differentiated prompts are first configured according to the same annotation task, and multiple rounds of semantic parsing are performed on the multimodal data to be annotated based on the multiple sets of differentiated prompts to obtain multiple candidate annotation results. After obtaining multiple candidate annotation results, the multiple candidate annotation results are filtered to determine the candidate label set. Then, the candidate label set is mapped to a preset semantic space, and a semantic action field is constructed based on the mapping result. Next, the target annotation result is determined based on the semantic action field. Then, a target sample set and cross-modal alignment data are constructed based on the target annotation result and the corresponding multimodal data. Further, video data is generated based on the target sample set and cross-modal alignment data. Finally, the target sample set and cross-modal alignment data are iteratively updated through the video data.
[0029] In step 101, when performing the same multimodal data annotation task, multiple sets of differentiated prompts are first configured according to the task requirements. Each set of prompts is customized around the current annotation requirements to avoid homogenization of prompt content. After configuring multiple sets of differentiated prompts, each set of differentiated prompts is used sequentially to perform multiple rounds of semantic parsing on the original multimodal data to be annotated, such as images, text, audio, and video. Each set of prompts is broken down into multiple sub-prompts in a step-by-step questioning manner to complete rounds of parsing. Each round of semantic parsing will output a preliminary annotation. After parsing is completed, the output content corresponding to all rounds and all prompts is finally summarized to obtain a sufficient number of candidate annotation results. For example, for the annotation task of "scene object classification", three different types of prompts are set, focusing on object appearance description, object functional attributes, and scene association. Each type of prompt is further refined into multiple rounds of questions. After the model parses round by round, it will produce multiple different candidate annotation results.
[0030] In step 102, after obtaining all candidate annotation results, the result filtering is completed by combining voting and election mechanisms. Specifically, firstly, voting judgment rules are applied to the multi-round parsing results generated by the step-by-step prompt information of a single group. As long as any round output result within a single group is determined to be positively correlated and valid annotation content, the annotation information corresponding to that group is retained. Then, the results after filtering all groups are summarized, and a second filtering is carried out according to a preset election threshold. The frequency of occurrence of valid annotation content in all candidate annotation results is counted. When the frequency of occurrence of a certain annotation content exceeds the set threshold, it is included in the qualified range. Then, invalid annotation content that is duplicate, contradictory, or obviously erroneous is removed. Finally, a standardized candidate label set is formed. This process relies on an unsupervised filtering method, which can purify valid annotation labels from massive candidate results at low cost.
[0031] In step 103, after the candidate label set is determined, each candidate label in the set is mapped to a pre-constructed unified semantic space. This semantic space is built based on multimodal feature rules and can unify the expression forms of text labels, image features, and video features. During the label mapping process, the semantic features, association features, and modal association attributes corresponding to each candidate label are extracted, allowing discrete labels to form feature vectors in the same dimension. Then, based on the distribution of feature vectors obtained after all label mappings, the distance relationship between vectors, and the degree of association, a corresponding semantic action field is further constructed. The semantic action field will intuitively reflect the semantic weight, mutual association strength, and semantic coverage of each candidate label. The higher the semantic similarity of the labels, the more concentrated their positions are in the action field; the greater the semantic difference, the farther the spatial distribution distance, thus completing the transformation of labels from text form to spatial feature form.
[0032] In step 104, based on the constructed semantic action field, the final target annotation result is determined by combining the feature distribution within the field, semantic weights, and label associations. Specifically, within the semantic action field, the semantic region with the largest comprehensive action strength weight is preferentially selected as the semantic convergence region, and the target annotation result is generated based on the candidate labels in the semantic convergence region. After passing through the semantic action field, the rationality of the candidate labels can be judged from a global perspective. Compared with a single result judgment method, this method can further improve the accuracy of the annotation result, and finally output a unique target annotation result that meets the task requirements.
[0033] In step 105, after determining the target annotation results, these results are paired and integrated with the corresponding original multimodal data. Specifically, on the one hand, the paired and valid data are organized and collected to build a target sample set for model training. This sample set can be directly used as training material for single-modal small models, supplementing the positive samples required for intra-modal annotation. On the other hand, combined with multimodal feature extraction rules, cross-modal feature processing is performed on the multimodal data with target annotation results. Visual features are extracted from image data using a convolutional neural network, and text features are extracted from text data using a Transformer encoder. Then, feature fusion is completed through a multimodal Transformer. Similarity is calculated based on the fused features, and suitable datasets are selected to build cross-modal aligned data, thereby completing the process of data classification, feature processing, and alignment optimization.
[0034] In step 106, the video generation process is initiated using the completed target sample set and cross-modal aligned data processed by feature alignment as the basic data source. Specifically, the pre-trained dynamic module parameters are first fused into the basic diffusion model to form the original video generation model. Simultaneously, a dual-model architecture is built, with a mature diffusion model as the teacher model and a lightweight diffusion model as the student model. Model training is conducted using adversarial distillation combined with probabilistic flow distillation. The training process employs a distributed parallel training mode, distributing different models to different devices for computation. During the video generation stage, a specially designed motion module processes the dynamic motion information in the video's temporal dimension. Combined with the labeled information and cross-modal features from the data source, high-quality video data with complete image details and natural dynamic effects is quickly generated. This fully leverages the supporting role of labeled and aligned data, realizing the transformation of labeled data into video content.
[0035] In step 107, the recirculated video data serves as new multimodal data to be labeled, re-entering the processes of prompt parsing, candidate result selection, semantic processing, and target labeling determination. This generates new labeling results and expands the target sample set. Simultaneously, the new video data, combined with the corresponding labeling results, undergoes multimodal feature extraction, fusion, and similarity calculation again, continuously supplementing the volume and diversity of cross-modal alignment data. As video data continuously recirculates and repeatedly participates in the labeling and alignment process, the target sample set and cross-modal alignment data undergo continuous iterative optimization. The model labeling accuracy, cross-modal alignment effect, and video generation capability are also continuously improved, enabling the entire intelligent labeling and data generation system to operate in a closed loop.
[0036] Here are some practical application scenarios to illustrate this: First, for the unified task of labeling pedestrian behavior in street photography, several different sets of prompts were written. Then, based on the prompts, the images, text, and video frames were repeatedly analyzed and interpreted to obtain many preliminary labeling results, such as various candidate contents such as parks, roads, pedestrians, walking, cycling, indoors, and running.
[0037] Next, these preliminary results are filtered out according to established rules. Obviously incorrect content is removed. For example, if the scene is clearly outdoors, the "indoor" label is deleted; if the person is just walking normally, inappropriate labels such as "running" or "cycling" are removed. Finally, the effective candidate labels of road, pedestrian, walking, and outdoors are left.
[0038] Then, these labels are converted into feature vectors that the system can recognize, placed in a unified semantic space, and a semantic action field is built based on the proximity of the association between the labels. The priority is determined by combining the feature distribution within the field, and finally the formal annotation result is obtained: an outdoor road scene where people are walking normally.
[0039] After obtaining the final labels, the labels are bound together with the corresponding images, text, and original video materials to form a standard sample set that can be used for model training. At the same time, visual features of the images and semantic features of the text are extracted to complete cross-modal feature alignment and form a matching alignment dataset.
[0040] Based on the organized samples and aligned data, an optimized diffusion model is used, combined with a dedicated dynamic and motion processing module, to efficiently generate multiple new short videos of "pedestrians walking on the street" in batches, with smooth visuals and consistent content and annotation requirements.
[0041] Finally, the newly generated short videos are treated as new material and fed back into the initial annotation stage. The process of parsing, filtering, and aligning is repeated, continuously adding training samples and alignment data. This allows the annotation model and video generation model to be continuously optimized, resulting in fewer errors and more stable performance, forming a complete self-iterative closed loop. Through these operations, the accuracy and stability of multimodal data annotation results can be effectively improved, and the overall utilization efficiency of multimodal data can be enhanced.
[0042] Please refer to Figure 2 According to some embodiments of the present invention, in step 103, the candidate tag set is mapped to a preset semantic space, and a semantic action field is constructed based on the mapping result. Specifically, this may include, but is not limited to, the following: 201. Map the candidate tag set to a preset semantic space to obtain the semantic coordinates of each candidate tag in the preset semantic space; 202. Based on the semantic coordinates, the credibility information of the candidate tags, and the semantic association between the candidate tags, determine the semantic influence strength of each candidate tag in the semantic space; 203. Take each candidate label as a semantic action source, and construct a semantic action field based on the semantic coordinates and the semantic action intensity.
[0043] In this embodiment, after the candidate label set is screened and summarized, a unified preset semantic space pre-built and adapted for this multimodal annotation task is first retrieved. This semantic space is designed for multimodal features such as text, images, and videos, and can quantify and represent information such as label semantics, modal feature associations, and annotation confidence. When mapping candidate labels to the preset semantic space, a semantic encoding model is first used to vectorize the candidate labels to generate corresponding semantic feature vectors. Then, the semantic coordinates of the candidate labels are determined based on the position of the semantic feature vectors in the preset semantic space. The semantic space can be a vector space, an embedding space, or a knowledge graph space. Next, each valid candidate label in the candidate label set is sequentially semantically mapped. Combining the label semantic features and multimodal data association features retained from previous rounds of semantic parsing, voting, and election screening, and leveraging the feature representation capabilities of the multimodal large model, the unique semantic coordinates corresponding to each candidate label in the preset semantic space are calculated and determined. These coordinates accurately record the semantic position and feature dimension distribution of a single label. After the above operations, the originally discrete text labels can be transformed into quantifiable and computable point data within the semantic space.
[0044] After obtaining the semantic coordinates of all candidate labels, and combining the judgment criteria accumulated from the previous analysis of multiple sets of differentiated prompt words, multiple rounds of step-by-step questioning, and the dual screening mechanism, the credibility information corresponding to each candidate label is extracted. Credibility can be comprehensively quantified based on the number of positive correlation results in the election mechanism, the voting results of a single set of prompt words, and the degree of multimodal feature matching. Simultaneously, the spatial distance between labels is calculated based on their semantic coordinates. Specifically, the distance between two candidate labels can be calculated using this formula: ,in, Indicates semantic distance. This represents the coordinates of candidate label i in the k-th dimension of the semantic space. Let represent the coordinates of candidate label j in the k-th dimension of the semantic space, and n represent the dimension of the semantic space.
[0045] By combining modal fusion features extracted from text Transformer and image CNN, the semantic relationships between different candidate labels are derived, clarifying the semantic similarity, subordination, and opposition among the labels. Integrating the semantic coordinates of a single label, its own credibility value, and the semantic relationships between pairs of labels, the semantic influence strength of each candidate label in the current semantic space is calculated according to preset calculation rules. Influence strength directly reflects the discourse power, semantic influence, and potential probability of a single label becoming the final annotation result within the overall labeling system, thus completing a deep feature deduction from position coordinates to influence strength.
[0046] After clarifying the semantic coordinates and corresponding semantic influence strengths of all candidate tags, each candidate tag is treated as an independent semantic influence source. The semantic coordinates of each tag are used as the spatial landing point of the influence source, and the calculated semantic influence strength is used as the radiation capacity and influence weight of the influence source. Combining the established semantic relationships between tags, the influence propagation logic of the spatial field is simulated, and the overall arrangement and coupling of the tags are completed within the overall preset semantic space, thereby constructing a complete semantic influence field. This semantic influence field integrates the spatial location, credibility, mutual semantic influence, and influence radiation range of all candidate tags, and can fully present the overall semantic distribution of candidate tags.
[0047] Please refer to Figure 3 According to some embodiments of the present invention, in step 202, the semantic influence strength of each candidate label in the semantic space is determined based on the semantic coordinates, the credibility information of the candidate labels, and the semantic association between the candidate labels. Specifically, this may include, but is not limited to, the following: 301. Obtain the credibility information of each candidate label; 302. Determine the semantic relationships between candidate tags based on their distribution in the preset semantic space; 303. Determine the semantic influence strength of each candidate tag based on the credibility information and the semantic association.
[0048] In this embodiment, after organizing the candidate label set and generating the semantic coordinates of each label, the credibility information corresponding to all candidate labels in the set is collected one by one. This credibility information is comprehensively judged based on the full-process data of multimodal large model combined with step-by-step prompt words to carry out multi-round semantic parsing, voting mechanism and election mechanism. Specifically, on the one hand, the results of single-group step-by-step prompt words being judged as positively correlated in multi-round questioning are counted. On the other hand, the number and proportion of times the candidate label reaches the election threshold after all prompt word groups are counted by the election mechanism are calculated. At the same time, the credibility of each candidate label is quantitatively assigned by combining the multimodal fusion feature matching degree extracted by image modality CNN and text modality Transformer, as well as the cross-modal feature similarity calculation results, so as to form quantifiable and comparable credibility information. This credibility information objectively reflects the effectiveness and reliability level of different candidate labels after multi-level screening.
[0049] Next, based on the determined semantic coordinates of each candidate label and their overall distribution within the preset semantic space, the semantic relationships between each pair of candidate labels are analyzed and determined. Specifically, based on the semantic coordinates, Euclidean distance is used as a metric to calculate the spatial distance between different label points. The closer the spatial distance, the higher the semantic overlap and the stronger the association; the farther the spatial distance, the greater the semantic difference and the weaker the association. Simultaneously, combined with the feature information fused from the multimodal Transformer, different association types such as semantic similarity, hierarchical subordination, and semantic contradiction between labels are distinguished, comprehensively outlining the pairwise semantic relationships within the entire candidate label set and clarifying the inherent connections of mutual influence and constraint among all labels in the semantic space.
[0050] After obtaining the quantified credibility information of each candidate label and the complete semantic relationships between labels, the semantic influence strength of each candidate label is calculated by combining these two sets of data. The semantic influence strength can be determined by the following formula: F = α × C + β × R, where F represents the semantic influence strength, C represents the credibility of the candidate label, R represents the semantic relevance of the candidate label, and α and β represent weighting coefficients. During the calculation, the label's own credibility is used as the basic weight; the higher the credibility value, the stronger the label's basic influence. Dynamic adjustments are then made based on the semantic relationships between labels. For labels with highly similar semantics, the influence strength is adjusted collaboratively based on their distribution and mutual influence. For labels with semantic conflicts, weight balancing is performed based on the association features. The final calculation is completed by combining the feature weighting rules output by the modality alignment shared parameter matrix, ultimately determining the semantic influence strength of each candidate label in the semantic space. This semantic influence strength directly reflects the semantic influence and priority of a single label within the overall annotation system.
[0051] Please refer to Figure 4 According to some embodiments of the present invention, in step 101, multiple sets of differentiated prompt information are configured according to the same annotation task, and multiple rounds of semantic parsing are performed on the multimodal data to be annotated according to the multiple sets of differentiated prompt information to obtain multiple candidate annotation results. Specifically, this may include, but is not limited to, the following: 401. Construct target description prompts, scene constraint prompts, and semantic reasoning prompts for the same annotation task; 402. Perform semantic parsing on the multimodal data to be labeled based on the target description prompts, the scene constraint prompts, and the semantic reasoning prompts to obtain multiple candidate labeling results.
[0052] In this embodiment, for the same multimodal data intelligent annotation task currently being carried out, and considering the annotation requirements, data characteristics, and the overall technical requirements for subsequent sample selection and cross-modal alignment, three types of differentiated prompt information are constructed for classification: target description prompt information, scene constraint prompt information, and semantic reasoning prompt information. The target description prompt information defines and explains the core annotation object, annotation category, and basic attributes of this annotation task, clarifying the core annotation content that needs to be extracted from the multimodal data. The scene constraint prompt information limits the effective range of the annotation results based on the application scenario, environmental conditions, and boundary rules corresponding to the multimodal data, avoiding invalid outputs that are detached from the actual scenario. The semantic reasoning prompt information, relying on logical deduction and association analysis, guides the model to perform deep semantic inference by combining the association information of different modalities such as text and images within the data, and to uncover implicit annotation features.
[0053] Meanwhile, each type of prompt information is further broken down into multi-level sub-prompt content according to the step-by-step prompting approach, forming a hierarchical questioning logic. The three types of prompt information complement each other from the three dimensions of core content, scene boundaries, and in-depth reasoning, forming a complete and distinctive prompt information system, providing standardized input for multi-round semantic parsing.
[0054] Next, leveraging a multimodal large model, the constructed target description prompts, scene constraint prompts, and semantic reasoning prompts are sequentially invoked to perform multiple rounds of semantic parsing on the multimodal data to be labeled. During the parsing process, interactions and feature interpretations are initiated round by round according to the multi-level sub-prompt content decomposed within each type of prompt. The model combines different modal data such as images and text to respond to the guidance requirements of each type of prompt: identifying the core labeled subject based on the target description prompts, verifying whether the labeled content matches the corresponding scene rules according to the scene constraint prompts, and finally completing deep semantic association and logical judgment through semantic reasoning prompts.
[0055] The three types of prompts are parsed and output corresponding results independently. Each step-by-step parsing under a single type of prompt generates corresponding intermediate output content. After integrating the parsing output content of all categories and all rounds, multiple sets of candidate annotation results with different dimensions and interpretation angles can be obtained.
[0056] Please refer to Figure 5 According to some embodiments of the present invention, in step 105, a target sample set and cross-modal alignment data are constructed based on the target annotation results and the corresponding multimodal data. Specifically, this may include, but is not limited to, the following: 501. Extract target semantic entities from multimodal data based on the target annotation results, and construct unified semantic anchors between different modalities using the target semantic entities; 502. Determine the semantic mapping relationship between different modal data based on the unified semantic anchor point; 503. Construct the target sample set and cross-modal alignment data through the semantic mapping relationship.
[0057] In this embodiment, after obtaining the target annotation results, target semantic entities that represent the core content of the data and are strongly related to the annotation purpose are extracted from various multimodal raw data such as images and text. These semantic entities include keywords and semantic concepts in the text dimension, as well as semantic units in different modalities such as target objects and visual elements in the image dimension. Combined with the modal alignment shared parameter matrix built into the multimodal large model, unified semantic representation processing is performed on the extracted target semantic entities to eliminate semantic expression differences between different modalities. Target semantic entities scattered in single modalities such as text and images are associated and bound, thereby establishing a unified semantic anchor point that can connect all modalities and serve as a benchmark for localization. The unified semantic anchor point is used to represent semantic entities with the same semantic orientation in different modal data. This semantic anchor point serves as the semantic intersection benchmark for various modal data, connecting the determined target annotation results and the extracted semantic entities.
[0058] After constructing a unified semantic anchor, image modal features extracted by a convolutional neural network and text modal features output by a Transformer encoder are combined. Multi-dimensional feature fusion is achieved through a multi-modal Transformer. Then, similarity metrics such as Euclidean distance are used to calculate the matching degree and association distance of different modal data features relative to the unified semantic anchor. Based on the level of feature similarity and the correspondence rules between semantic entities and anchors, the one-to-one semantic mapping relationship between various modal data such as text, images, and newly added videos is identified and determined. The semantic transformation rules, corresponding logic, and weight allocation of different modal content are clarified, enabling each modal data to complete mutual interpretation and association matching based on the same semantic benchmark. This establishes cross-modal semantic association rules and opens up semantic communication channels between different modal data.
[0059] Based on the established complete semantic mapping relationship, the target annotation results, original multimodal data, and semantic mapping rules are integrated and collected. On the one hand, in accordance with the standard requirements of intramodal annotation, the multimodal data that has been bound to the target annotation results and completed semantic entity association is sorted and classified to form a standardized target sample set. The high-quality positive samples in this sample set can be directly used for the training iteration of single-modal small models, continuously improving the efficiency and accuracy of intramodal intelligent annotation. On the other hand, relying on the established semantic mapping relationship, the modal alignment shared parameter matrix, and the human feedback reinforcement learning optimization strategy, batch alignment processing is carried out on the multimodal data. The dataset is sorted according to feature similarity, and human feedback judgment and feature weighting optimization of highly similar data are prioritized, ultimately generating compliant and usable cross-modal aligned data.
[0060] Please refer to Figure 6 According to some embodiments of the present invention, in step 106, video data is generated based on the target sample set and cross-modal alignment data, which may specifically include, but is not limited to, the following: 601. Extract target semantic information from the target sample set; 602. Establish the association constraint relationship between different modes based on the cross-modal alignment data; 603. Construct a semantic constraint chain based on the target semantic information and the associated constraint relationship; 604. Generate video data based on the semantic constraint chain.
[0061] In this embodiment, after constructing the target sample set and cross-modal alignment data, target semantic information is first extracted from the target sample set. The target sample set stores filtered and verified multimodal positive samples, including paired images, text, and corresponding target annotation results. The sample content aligns with the target requirements of this annotation task. Leveraging the feature extraction capabilities of a multimodal large-scale model, semantic analysis is performed on each group of data in the sample set. For text samples, a Transformer encoder is used to decompose the sentence into keywords, semantic logic, entity information, and scene descriptions. For image samples, a convolutional neural network is used to extract visual semantics such as the main subject, scene elements, action states, and visual attributes. These are then combined with the determined target annotation results for semantic fusion. Ultimately, target semantic information that represents the core content of the data and the task's objective is obtained. For example, if the annotation task is outdoor animal behavior recognition, and the target sample set contains text data of "horses running on the grassland" and corresponding real-world images, then core target semantic information such as "the main subject is a horse, the scene is a grassland, and the behavior is running" will be extracted.
[0062] After extracting the target semantic information, association constraints are established between different modalities, such as images, text, and newly added video modalities, based on the processed cross-modal alignment data. Specifically, the cross-modal alignment data itself has already undergone multimodal feature fusion, similarity ranking, and human feedback optimization, clarifying the correspondence rules between different modalities in terms of content, features, and semantics. Based on this type of data, the binding rules, content correspondence boundaries, feature matching requirements, and logical constraints between modalities are sorted out, forming association constraints between modalities. Taking outdoor animal behavior recognition as an example, based on the alignment data, it can be seen that the text modality describing "horse running" must correspond to the image modality where the main subject is a horse and its posture is in a dynamic running state. Scene words and action words in the text need to be matched one by one with the visual elements and dynamic features in the image. At the same time, it is stipulated that there should be no semantic contradictions or content misalignments between modalities. For example, when the text is labeled "running", the image modality cannot match a static image of a horse standing. Such rules serve as the association constraints between different modalities to ensure that the data of each modality remains consistent in content and logic, and to carry out the extracted target semantic information, thus defining the modality matching rules for the semantic link.
[0063] After acquiring the target semantic information and the intermodal constraints, the two are combined to construct a complete semantic constraint chain. The semantic constraint chain describes the order of association and constraint rules between semantic elements during video generation. It can include scene constraints, subject constraints, behavioral constraints, temporal continuity constraints, or modal consistency constraints. Using the target semantic information extracted from the sample set as the main thread, the semantic constraint chain connects and arranges the dispersed semantic units according to scene logic, action sequence, and content association. Simultaneously, the cross-modal association constraints established in the previous step are embedded into each node of the entire chain, binding constraints to the modal representation, feature standards, and logical limitations corresponding to each segment of semantic content. For example, a link is constructed using "grassland-horse-running" as the core semantic thread. Modal constraints are superimposed at each node of the link: the "grassland" semantic node requires the text modality to contain corresponding scene vocabulary and the image modality to present grassland terrain features; the "horse" semantic node limits the subject of each modality to horses, excluding other animals; and the "running" semantic node constrains all modalities to reflect dynamic movement characteristics. Simultaneously, the entire link strictly adheres to the alignment rule of one-to-one correspondence between text and image modalities, preventing issues of disjointed modal content or semantic conflicts. This entire semantic constraint chain not only determines the complete semantic logic and narrative order that the video content needs to express but also standardizes the presentation of different modal content through associative constraints, becoming the core carrier connecting semantic information and the video generation model.
[0064] Based on the established semantic constraint chain, the diffusion model, optimized through dynamic module fusion and distillation, is invoked to generate video data. This involves a deep integration of dynamic modules, motion modules, and the basic diffusion model, while simultaneously utilizing teacher-student models trained with adversarial distillation and probabilistic flow distillation. Specifically, the complete semantic logic, temporal sequence, and modal constraints of each node in the semantic constraint chain are first analyzed. The motion module specifically addresses dynamic semantics such as "running" in the chain, processing dynamic information such as motion trajectory and action changes in the video's temporal dimension. The dynamic module also adapts to the overall video's visual style, scene transitions, and other global features. The diffusion model generates frames one by one according to the requirements of the semantic constraint chain, adhering to modal association constraints throughout the process. This ensures a high degree of semantic consistency between the video content and the original text and image semantics, preventing semantic deviation, subject misalignment, and inconsistent actions.
[0065] For example, the model generates continuous video footage based on the semantic constraint chain of "grassland, horse, running". The footage continuously presents the grassland scene, with the horse as the main subject and maintaining a running dynamic effect. The scene, subject, and action are consistent with the preceding semantics and constraint rules throughout the process. Finally, it outputs complete video data with excellent image quality, smooth dynamics, and accurate semantics. The generated video data will also flow back to the front-end annotation and alignment stage according to the established process, thereby achieving closed-loop iterative optimization throughout the entire process.
[0066] Please refer to Figure 7 According to some embodiments of the present invention, step 107, which iteratively updates the target sample set and the cross-modal alignment data using the video data, may specifically include, but is not limited to, the following: 701. Perform semantic analysis on the video data to obtain new candidate tags; 702. Map the newly added candidate tags to the preset semantic space, and determine the semantic association between the newly added candidate tags and the candidate tag set; 703. Supplement the target sample set according to the semantic association relationship, and update the cross-modal alignment data through the supplemented target sample set.
[0067] In this embodiment, after the video data is generated, semantic parsing is first performed on the returned video-type multimodal data. The video data includes image modal information such as frames, frame sequences, and dynamic content, as well as associated text information. Following the parsing logic of the entire method, multiple sets of differentiated step-by-step prompts are configured. Combined with a multimodal large model, multiple rounds of semantic interpretation are performed on the video's frame-by-frame content, overall semantics, and dynamic behavior. The core semantic features and content attributes of the video data are mined layer by layer according to a step-by-step questioning approach. Referring to the output rules of the aforementioned candidate annotation results, new candidate tags corresponding to the video content are extracted from the parsing results. These new candidate tags are generated based on the video's dynamic features and multimodal semantics, effectively supplementing the original tag system and providing a basic tag basis for subsequent sample and data updates.
[0068] After obtaining the new candidate tags corresponding to the video data, each new candidate tag is semantically mapped one by one, and its coordinates, semantic feature vector, and weight parameters in the semantic space are quantified. Combined with the existing candidate tag set stored in the semantic space, the semantic distance, overlap, and association strength between the new candidate tag and each tag in the candidate tag set are calculated using similarity measures such as Euclidean distance. This clarifies the semantic relationship between the two, distinguishes between different tag types with high semantic overlap, partial semantic association, and semantic independence, and thus defines the position and role of the new tag in the overall tag system, achieving association matching between new and old tags under a unified semantic framework.
[0069] After clarifying the semantic relationship between the newly added candidate tags and the existing candidate tag set, the target sample set is supplemented and optimized based on the association determination results. For newly added candidate tags with high semantic relevance and excellent content matching, they are bound to the corresponding video multimodal data, organized into new positive samples according to sample specifications, and incorporated into the original target sample set, expanding the data volume and content coverage of the sample set. For newly added candidate tags with independent semantics that can enrich tag dimensions, samples are also constructed and added to the sample set simultaneously.
[0070] After the sample set is supplemented and updated, iterative optimization of cross-modal alignment data is initiated based on the updated complete target sample set. The modal alignment shared parameter matrix within the multimodal large model is invoked, and video image modal features are extracted using a convolutional neural network and corresponding text modal features are extracted using a Transformer encoder. Cross-modal feature fusion is then completed via a multimodal Transformer. The dataset is then reordered based on feature similarity and optimized using feedback reinforcement learning. Finally, the overall update of cross-modal alignment data is completed, further enhancing the multimodal feature fusion and data alignment capabilities.
[0071] Please see Figure 8The second aspect of this application provides an intelligent annotation device based on multimodal collaboration, the device comprising: The parsing unit 801 is used to configure multiple sets of differentiated prompt information according to the same annotation task, and to perform multiple rounds of semantic parsing on the multimodal data to be annotated according to the multiple sets of differentiated prompt information to obtain multiple candidate annotation results; The first determining unit 802 is used to filter multiple candidate labeling results and determine a set of candidate labels; The first construction unit 803 is used to map the candidate label set to a preset semantic space and construct a semantic action field based on the mapping result; The second determining unit 804 is used to determine the target annotation result based on the semantic action field; The second construction unit 805 is used to construct a target sample set and cross-modal alignment data based on the target annotation results and the corresponding multimodal data; Generation unit 806 is used to generate video data based on the target sample set and cross-modal alignment data; The update unit 807 is used to iteratively update the target sample set and the cross-modal alignment data using the video data.
[0072] Please see Figure 9 This application also provides an intelligent annotation device based on multimodal collaboration, the device comprising: Processor 901, memory 902, input / output unit 903, bus 904; The processor 901 is connected to the memory 902, the input / output unit 903, and the bus 904; The memory 902 stores a program, and the processor 901 calls the program to execute any of the methods described above.
[0073] This application also relates to a computer-readable storage medium on which a program is stored, which, when run on a computer, causes the computer to perform any of the methods described above.
Claims
1. A smart annotation method based on multimodal collaboration, characterized in that, include: Configure multiple sets of differentiated prompts for the same annotation task, and perform multiple rounds of semantic parsing on the multimodal data to be annotated based on the multiple sets of differentiated prompts to obtain multiple candidate annotation results; Filter through multiple candidate annotation results to determine the candidate label set; The candidate label set is mapped to a preset semantic space, and a semantic action field is constructed based on the mapping results; The target annotation result is determined based on the semantic field. Based on the target annotation results and the corresponding multimodal data, construct a target sample set and cross-modal alignment data; Video data is generated based on the target sample set and cross-modal alignment data; The target sample set and the cross-modal alignment data are iteratively updated using the video data.
2. The intelligent annotation method based on multimodal collaboration according to claim 1, characterized in that, Mapping the candidate label set to a preset semantic space, and constructing a semantic action field based on the mapping results, including: The candidate tag set is mapped to a preset semantic space to obtain the semantic coordinates of each candidate tag in the preset semantic space; Based on the semantic coordinates, the credibility information of the candidate tags, and the semantic association between the candidate tags, the semantic influence strength of each candidate tag in the semantic space is determined. Each candidate label is used as a semantic action source, and a semantic action field is constructed based on the semantic coordinates and the semantic action intensity.
3. The intelligent annotation method based on multimodal collaboration according to claim 2, characterized in that, Based on the semantic coordinates, the credibility information of candidate tags, and the semantic relationships between candidate tags, the semantic influence strength of each candidate tag in the semantic space is determined, including: Obtain the credibility information of each candidate label; The semantic relationships between candidate tags are determined based on the distribution of candidate tags in the preset semantic space; The semantic influence strength of each candidate label is determined based on the credibility information and the semantic association.
4. The intelligent annotation method based on multimodal collaboration according to claim 3, characterized in that, Determining the semantic relationships between candidate tags based on their distribution in the preset semantic space includes: Obtain the semantic coordinates of each candidate tag in the preset semantic space, and calculate the semantic distance between any two candidate tags based on the semantic coordinates; The semantic relevance between candidate tags is determined based on the semantic distance. The semantic association between candidate tags is established based on the semantic association degree.
5. The intelligent annotation method based on multimodal collaboration according to claim 1, characterized in that, Multiple sets of differentiated prompts are configured for the same annotation task, and multiple rounds of semantic parsing are performed on the multimodal data to be annotated based on the multiple sets of differentiated prompts to obtain multiple candidate annotation results, including: Based on the same annotation task, construct target description prompts, scene constraint prompts, and semantic reasoning prompts respectively; The multimodal data to be labeled is semantically parsed based on the target description prompts, the scene constraint prompts, and the semantic reasoning prompts to obtain multiple candidate labeling results.
6. The intelligent annotation method based on multimodal collaboration according to claim 1, characterized in that, Based on the target annotation results and the corresponding multimodal data, a target sample set and cross-modal alignment data are constructed, including: Based on the target annotation results, target semantic entities are extracted from the multimodal data, and unified semantic anchors between different modalities are constructed using the target semantic entities; The semantic mapping relationship between different modalities is determined based on the unified semantic anchor point; The target sample set and cross-modal aligned data are constructed using the semantic mapping relationship.
7. The intelligent annotation method based on multimodal collaboration according to claim 1, characterized in that, Video data is generated based on the target sample set and cross-modal alignment data, including: Extract target semantic information from the target sample set; Establish association constraint relationships between different modes based on the cross-modal alignment data; Construct a semantic constraint chain based on the target semantic information and the associated constraint relationship; Video data is generated based on the semantic constraint chain.
8. The intelligent annotation method based on multimodal collaboration according to claim 1, characterized in that, Iteratively updating the target sample set and the cross-modal alignment data using the video data includes: Perform semantic analysis on the video data to obtain new candidate tags; The newly added candidate tags are mapped to the preset semantic space, and the semantic association between the newly added candidate tags and the candidate tag set is determined; The target sample set is supplemented based on the semantic association, and the cross-modal alignment data is updated using the supplemented target sample set.
9. The intelligent annotation method based on multimodal collaboration according to claim 1, characterized in that, Determining the target annotation result based on the semantic action field includes: The comprehensive influence intensity of each semantic region in the semantic field is obtained, and the semantic region with the largest comprehensive influence intensity is determined as the semantic convergence region. Target annotation results are generated based on the candidate labels in the semantic convergence region.
10. An intelligent annotation device based on multimodal collaboration, characterized in that, include: The parsing unit is used to configure multiple sets of differentiated prompt information according to the same annotation task, and to perform multiple rounds of semantic parsing on the multimodal data to be annotated according to the multiple sets of differentiated prompt information to obtain multiple candidate annotation results; The first determining unit is used to filter multiple candidate labeling results and determine the candidate label set; The first construction unit is used to map the candidate label set to a preset semantic space and construct a semantic action field based on the mapping result; The second determining unit is used to determine the target annotation result based on the semantic action field; The second construction unit is used to construct a target sample set and cross-modal alignment data based on the target annotation results and the corresponding multimodal data; The generation unit is used to generate video data based on the target sample set and cross-modal alignment data; An update unit is used to iteratively update the target sample set and the cross-modal alignment data using the video data.
11. An intelligent annotation device based on multimodal collaboration, characterized in that, The device includes: Processor, memory, input / output units, and bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, which the processor invokes to perform the method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains a program that, when executed on a computer, performs the method as described in any one of claims 1 to 9.