An image labeling intelligent analysis system and method based on multi-modal fusion

By distinguishing between stable and fluctuating labels, generating correction chains, differentiating anomaly types, and optimizing multimodal fusion strategies, the problems of high annotation error rate and resource waste in existing systems are solved, achieving efficient and accurate image annotation.

CN120852940BActive Publication Date: 2026-01-09北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511367590.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-09
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing image annotation systems suffer from problems such as poor adaptability to labeling scenarios, inefficient handling of abnormal events, weak ability to distinguish similar objects, lack of dynamic correction mechanisms, and crude multimodal fusion strategies in multimodal data annotation tasks, resulting in high annotation error rates, waste of resources, and increased computational costs.

Method used

By distinguishing between stable and fluctuating labels, a correction chain is generated, and mechanical anomalies are differentiated from analytical anomalies. A correction identification requirement array is constructed, a multimodal fusion strategy is optimized, and dynamic correction feature extraction is performed to improve annotation accuracy and efficiency.

Benefits of technology

It effectively reduces the annotation error rate, improves the ability to distinguish similar objects, reduces the consumption of computing resources, ensures the stability of annotation accuracy and efficiency, and is suitable for large-scale image annotation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852940B_ABST
    Figure CN120852940B_ABST
Patent Text Reader

Abstract

The application discloses a kind of image labeling intelligent analysis system and method based on multi-modal fusion, it is related to image labeling technical field, including label distinguishing module, correction chain analysis module, abnormal labeling event distinguishing module, correction identification demand array analysis module and early warning response module;Label distinguishing module is used to output stable label and fluctuation label based on the result of distinguishing;Correction chain analysis module is used to analyze the environmental characteristics in the effective image labeling event where fluctuation label is located, and generates the correction chain of corresponding fluctuation label;Abnormal labeling event distinguishing module is used to extract the abnormal labeling event stored in history, and the abnormal type of abnormal labeling event is mechanical abnormality and analysis abnormality;Correction identification demand array analysis module is used to construct correction identification demand array in image labeling analysis identification stage;Early warning response module carries out early warning response to the correction identification demand array corresponding to stable label and fluctuation label in real-time response image labeling analysis identification stage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image annotation, in particular to an image annotation intelligent analysis system and method based on multi-modal fusion. BACKGROUND

[0002] In the field of image annotation technology, the existing system faces multiple problems such as poor label scene adaptability, inefficient abnormal event processing, weak similar object distinguishing ability, lack of dynamic correction mechanism, and extensive multi-modal fusion strategy when dealing with multi-modal data annotation tasks. On the one hand, the existing system does not distinguish the stability of labels in different scenes, and uniformly uses a fixed strategy for annotation, which leads to annotation deviation of some labels due to environmental feature changes. Moreover, there is a lack of verification mechanism for different label types, which cannot perform scene-based secondary verification on labels susceptible to scene changes, further exacerbating errors. On the other hand, mechanical abnormalities and analysis abnormalities in system operation are not clearly distinguished and processed. Mechanical abnormalities are difficult to trigger maintenance prompts in time, and analysis abnormalities cannot accurately locate the error source, but only can be re-labeled generally, wasting resources. At the same time, when facing similar object annotation, the existing system mostly relies on single modal features, does not fully utilize the complementarity of multi-modal data, and uses an extensive "all-modal superposition" method for multi-modal fusion, which does not select efficient modal according to task requirements, increases computational cost, and reduces precision due to invalid modal interference. In addition, the feature extraction process is fixed, and there is a lack of dynamic correction system based on historical data, so that initial feature extraction deviation cannot be traced and corrected, making it difficult to meet the annotation requirements of high precision and high stability. SUMMARY

[0003] The present application aims to provide an image annotation intelligent analysis system and method based on multi-modal fusion to solve the problems in the prior art.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical scheme: an image annotation intelligent analysis method based on multi-modal fusion, the method comprising the following steps:

[0005] Step S100: collecting content labels for information annotation of each image data in historical stored image annotation events and feature columns required for generating content labels; the feature column refers to multi-source features extracted from each content in image data for matching labels; analyzing whether there are scene differences of each content label in effective image annotation events in effective image annotation events, and outputting stable labels and fluctuating labels based on the analysis results;

[0006] Step S200: analyzing the environmental features of the effective image annotation events where the fluctuating labels are located, and generating a correction chain corresponding to the fluctuating labels;

[0007] Step S300: extract the historically stored abnormal labeling events, distinguish the abnormal types of the abnormal labeling events as mechanical abnormalities and analysis abnormalities; when the mechanical abnormalities, the system prompts the model to respond to maintenance; when the analysis abnormalities, analyze the abnormal labeling events where the stable labels are located, and construct the correction recognition requirement array of each stable label in the image labeling analysis and recognition stage;

[0008] Step S400: when the analysis abnormalities, analyze the fluctuation labels in response to the trigger correction chain; when the analysis result is a label abnormality, extract the abnormality of the feature column; when the analysis result is a label normality, return to step S300 to output the correction recognition requirement array of each fluctuation label in the image labeling analysis and recognition stage;

[0009] Step S500: in real-time response to the image labeling analysis and recognition stage, pre-warning response is performed on the stable labels and the fluctuation labels that meet the corresponding correction recognition requirement array.

[0010] Further, step S100 includes the following specific steps:

[0011] Step S110: the effective image labeling event refers to the labeling event in which the content label of the image output by the system is consistent with the actual image content; the target event set of the content label is formed by searching all the effective image labeling events containing the search item in the historical record; the target feature column required for generating the corresponding search item of each effective image labeling event in the target event set is extracted as the target feature column; and the target feature column is bound with the effective image labeling event to form an association pair;

[0012] Step S120: analyze all the association pairs formed under each search item, and mark the association pairs with the same target feature column in the search combination as safe association pairs if the effective image labeling events of any two association pair records have a similarity less than a similarity threshold; if all the association pairs are marked as safe association pairs, output that there is no scene difference in the corresponding content label in the effective image labeling event; if there is an association pair that is not marked as a safe association pair, output that there is a scene difference in the corresponding content label in the effective image labeling event;

[0013] Step S130: output the content label with scene difference as a fluctuation label, and output the content label without scene difference as a stable label.

[0014] The purpose of distinguishing the labels is to verify and update the fluctuation labels after the extraction is completed when there is real-time feature scene extraction, so as to ensure the accuracy and precision of the fluctuation label extraction feature column, and to reduce the error rate of image labeling to a certain extent from the source.

[0015] Further, the correction chain corresponding to the fluctuation label is generated, including the following specific steps:

[0016] Step S210: Each fluctuation label is taken as a master node, each master node records each valid image label event in the target event set as an independent sub-node directly connected to the master node, and the content label recorded by each independent sub-node except the fluctuation label corresponding to the master node is extracted as a secondary node connected to the independent sub-node; the image data containing the master node, the independent sub-node and the secondary node form the environment characteristics corresponding to the master node;

[0017] Step S220: A rectangular coordinate system is established with the center point of the image data, the same physical dimension is drawn in all image data, the coordinate data of each secondary node located in the image data is marked, and the corresponding feature column is stored in the link of the corresponding independent sub-node and secondary node; all independent sub-nodes and secondary nodes of the same master node are traversed until the correction chain recording the coordinate data is generated.

[0018] Further, step S300 includes the following specific steps:

[0019] Step S310: Mechanical anomaly refers to that the system does not perform information labeling on the content label, and analysis anomaly refers to that the system labels the content label incorrectly; other content labels recorded in the abnormal labeling event where each stable label is located are collected, the content label corresponding to the same information as the incorrect information of the stable label is bound as an abnormal label, and an abnormal pair corresponding to the stable label is formed; if the information labeled incorrectly by the stable label is not the same as other content labels of the image, the content label with the maximum similarity to the feature column corresponding to the stable label is extracted to form an abnormal pair;

[0020] Step S320: All abnormal labeling events are traversed to generate a set of abnormal pairs corresponding to each stable label; each abnormal pair in the abnormal pair set is taken as a search target, and it is searched whether the same search target is recorded in the valid image labeling event; if not, the abnormal pair is output as a pre-warning response pair; if there is a record, the valid image labeling event and the abnormal labeling event recording the search target are extracted to form an event group to be analyzed;

[0021] Step S330: The feature modal type A1, A2 extracted in the image labeling analysis and recognition stage of the valid image labeling event and the abnormal labeling event in the event group to be analyzed is obtained, and the feature modal type refers to the modal data type corresponding to the feature extraction when the image labeling is used; the difference modal type a of the event group to be analyzed is calculated as A1∩A2;

[0022] When a=∅, the corresponding abnormal pair is marked as a pre-warning response pair;

[0023] When the number of difference modal types contained in a is one, the difference modal type is extracted as the correction recognition requirement array of the corresponding stable label when the abnormal pair in which the stable label is located coexists in the same image data; the correction recognition requirement array refers to image annotation analysis and recognition of the image in which the stable label is located, and the modal data type required to be extracted when the abnormal pair corresponding to the stable label exists, which is used to assist feature recognition to output corresponding annotation information;

[0024] When the number of difference modal types contained in a is greater than one, the number of times D1 that each difference modal type is output as valid in the historical annotation event containing the abnormal pair and the number of times D2 that each difference modal type is output as invalid are extracted, and the effective annotation participation index Q of each difference modal type is calculated using the formula: Q=D1 / (D1+D2). The larger the effective annotation participation index, the higher the influence degree of the difference modal type on improving the annotation accuracy in the historical annotation event containing the abnormal pair, and the greater the distinguishing effect. All the difference modal types contained in a are sorted in descending order according to the corresponding effective annotation participation index, and a requirement sequence is generated. The requirement sequence is used as the correction recognition requirement array of the corresponding stable label when the abnormal pair in which the stable label is located coexists in the same image data.

[0025] Further, step S400 includes the following specific steps:

[0026] The other content labels and corresponding coordinate data and feature columns of the image data in the abnormal annotation event in which each fluctuation label is located are extracted. The correction chain corresponding to the fluctuation label is extracted and stored, and the same correction chain as the link formed by the independent subnode and the secondary node in the correction chain is found for comparison. If the feature column recorded in the main node of the comparison correction chain is different from the feature column recorded in the fluctuation label, the analysis result is output as label abnormality, otherwise it is output as label normal. When the label is abnormal, it indicates that the generation of the corresponding abnormal annotation event is due to the annotation error caused by the difference in the feature column of the label itself.

[0027] Further, step S500 includes the following process:

[0028] When the real-time response image annotation analysis and recognition stage is reached, if the label record is a pre-warning response pair, an abnormal risk is directly pre-warned. If it is determined that the content to be annotated is a stable label, the content label that satisfies the historical record abnormal pair is extracted from the real-time image data in which the stable label is located, and the corresponding correction recognition requirement array is extracted based on the abnormal pair for pre-warning response. The abnormal pair indicates that the stable label coexists in the same image data and is prone to cause annotation errors. Therefore, when the real-time image data to be annotated contains a content label that can form an abnormal pair with the stable label, the correction recognition requirement array is extracted for further verification to improve the accuracy of image annotation. For example, if the correction recognition requirement array is a text modal data, the system can better assist annotation.

[0029] If it is determined that the required label is a fluctuation label, first, the fluctuation label and the correction chain are corrected, and when the pre-warning feature column extracts an exception, the feature column of the fluctuation label that is the same as the environment feature of the extraction and correction chain is replaced; and the content label of the abnormal pair obtained from the real-time image data where the fluctuation label after replacing the feature column is located is extracted, and the corresponding correction recognition requirement array is extracted based on the abnormal pair for pre-warning response.

[0030] An image label intelligent analysis system based on multi-modal fusion, the system comprises a label distinguishing module, a correction chain analysis module, an abnormal label event distinguishing module, a correction recognition requirement array analysis module and a pre-warning response module.

[0031] The label distinguishing module is used to collect the content labels of each image data completed information labeling in the historical stored image label events and the feature columns required for generating the content labels, extract the effective image label events to analyze whether there is scene difference of each content label in the effective image label events, and output stable labels and fluctuation labels based on the analysis results.

[0032] The correction chain analysis module is used to analyze the environment features of the effective image label events where the fluctuation labels are located, and generate the correction chain corresponding to the fluctuation labels.

[0033] The abnormal label event distinguishing module is used to extract the abnormal label events stored in history, and distinguish the abnormal types of the abnormal label events as mechanical abnormalities and analysis abnormalities.

[0034] The correction recognition requirement array analysis module is used to analyze the abnormal label events where the stable labels and the fluctuation labels are located when analyzing the abnormalities, and construct the correction recognition requirement array in the image label analysis and recognition stage.

[0035] The pre-warning response module is used to pre-warning response to the correction recognition requirement array corresponding to the stable labels and the fluctuation labels in the real-time response image label analysis and recognition stage.

[0036] Further, the correction chain analysis module comprises an environment feature generation unit and a chain data increasing unit.

[0037] The environment feature generation unit is used to take each fluctuation label as a master node, record each effective image label event in the target event set as an independent sub-node directly connected to the master node for each master node, extract other content labels recorded by each independent sub-node except the fluctuation label corresponding to the master node as a secondary node connected to the independent sub-node; and the image data containing the master node, the independent sub-node and the secondary node constitute the environment feature corresponding to the master node.

[0038] The chain data increasing unit is used for marking the coordinate data of each time node located at the image data and the corresponding feature column stored in the link of the corresponding independent sub-node and the time node;All independent sub-nodes and time nodes of the same main node are traversed until the correction chain of the coordinate data of each link record is stored.

[0039] Further, the correction recognition demand array analysis module comprises an abnormal pair extraction unit, a to-be-analyzed event group forming unit and a difference mode type analysis unit.

[0040] The abnormal pair extraction unit is used for extracting the content label of the abnormal pair when the similarity of the corresponding stable label feature column is the largest.

[0041] The to-be-analyzed event group forming unit is used for taking each abnormal pair in the abnormal pair set as a search target, searching whether the same search target is recorded in the effective image annotation event;If not, the abnormal pair is output as a pre-warning response pair;If there is a record, the effective image annotation event and the abnormal annotation event recording the search target are extracted to form a to-be-analyzed event group.

[0042] The difference mode type analysis unit is used for analyzing the difference mode type in the to-be-analyzed event group based on the feature mode type, and outputting the correction recognition demand array under different conditions.

[0043] Compared with the prior art, the beneficial effects of the present application are:

[0044] The present application effectively solves many pain points of the existing image annotation system by constructing a whole-process optimization scheme of "label classification-exception handling-dynamic correction-precise fusion". First, by analyzing the scene difference of the label in the effective image annotation event, it is divided into stable label and fluctuation label, and the correction chain containing environmental features, coordinate data and feature column is generated for the fluctuation label, realizing the scene verification update of the fluctuation label, and reducing the annotation error caused by scene change from the source.

[0045] The mechanical abnormality and the analysis abnormality are clearly distinguished, the system maintenance prompt is triggered when the mechanical abnormality to shorten the fault response time, the error source is accurately positioned by constructing the abnormal pair and the to-be-analyzed event group when the analysis abnormality, without the need for full-process re-annotation, and the fault handling efficiency is improved.

[0046] By calculating the effective annotation participation index of each mode, the correction recognition requirement array is constructed, the mode with high contribution degree is preferentially called in the similar object coexistence scene, the complementarity of multi-modal data is fully given, the similar object distinguishing ability is significantly improved, and the annotation error rate is reduced; at the same time, the correction analysis mechanism for fluctuating labels and the abnormality early warning response mechanism for stable labels are established, and a dynamic correction and early warning system is formed, so that the system can continuously learn and optimize from historical data, and ensure the long-term annotation precision stability; finally, the efficient mode is screened according to the effective annotation participation index, the blind calling of all modes is avoided, the annotation precision is ensured, the calculation resource consumption is reduced, the annotation efficiency is improved, and the actual application scene of large-scale image annotation is more suitable. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 It is a structural schematic diagram of the image annotation intelligent analysis method based on multi-modal fusion. DETAILED DESCRIPTION

[0048] All other embodiments obtained by those of ordinary skill in the art without creative labor on the basis of the embodiments in the present application belong to the scope of protection of the present application.

[0049] Embodiment: as shown in the figure, the present application provides an image annotation intelligent analysis method based on multi-modal fusion, the method comprising the following steps: Figure 1

[0050] Step S100: collect the content labels of each image data complete information annotation in the historical stored image annotation event and the feature columns required for generating the content labels; the feature column refers to the multi-source features extracted from each content in the image data for matching the labels; analyze whether the content labels in the effective image annotation event have scene differences, and output stable labels and fluctuating labels based on the analysis results;

[0051] Step S200: analyze the environmental features of the effective image annotation event where the fluctuating label is located, and generate a correction chain corresponding to the fluctuating label;

[0052] Step S300: extract the historical stored abnormal annotation event, and distinguish the abnormal types of the abnormal annotation event as mechanical abnormality and analysis abnormality; when the mechanical abnormality occurs, the system prompts the model to maintain; when the analysis abnormality occurs, analyze the abnormal annotation event where the stable label is located, and construct a correction recognition requirement array of each stable label in the image annotation analysis and recognition stage;

[0053] ​Step S400: When analyzing the abnormality, the fluctuation label response trigger is analyzed and compared; when the analysis result is label abnormality, the pre-warning feature column extraction is abnormal; when the analysis result is label normal, return to step S300 to output the correction identification requirement array of each fluctuation label in the image labeling analysis and identification stage;

[0054] Step S500: In real-time response to the image labeling analysis and identification stage, the correction identification requirement array corresponding to the stable label and the fluctuation label is pre-warned and responded.

[0055] Step S100 includes the following specific steps:

[0056] Step S110: The effective image labeling event refers to the content label output by the system identification corresponding to the actual image content; the content label is taken as a retrieval item, and all effective image labeling events containing the retrieval item in the historical record are searched to form a target event set of the content label, the target feature column required for generating the corresponding retrieval item of each effective image labeling event in the target event set is extracted as the target feature column; and the target feature column is bound with the effective image labeling event to form an association pair;

[0057] Step S120: Analyze all association pairs formed under each retrieval item, and take the effective image labeling events with image data similarity less than the similarity threshold as the search combination; mark the association pairs with the same target feature column in the search combination as safe association pairs; traverse all association pairs, if all association pairs are marked as safe association pairs, output that there is no scene difference in the corresponding content label in the effective image labeling event; if there is an association pair that is not marked as a safe association pair, output that there is scene difference in the corresponding content label in the effective image labeling event;

[0058] Step S130: Output the content label with scene difference as a fluctuation label, and output the content label without scene difference as a stable label.

[0059] The purpose of distinguishing the label is to verify and update the fluctuation label after the extraction is completed when there is real-time feature scene extraction, to ensure the accuracy and precision of the fluctuation label extraction feature column, and to reduce the error rate of image labeling to a certain extent from the source.

[0060] Generating a correction chain corresponding to the fluctuation label includes the following specific steps:

[0061] Step S210: Each fluctuation label is taken as a master node, each master node records each valid image label event in the target event set as an independent sub-node directly connected to the master node, and extracts other content labels recorded by each independent sub-node except the fluctuation label corresponding to the master node as a secondary node connected to the independent sub-node; image data containing the master node, independent sub-node and secondary node form the environment characteristics corresponding to the master node;

[0062] Step S220: A rectangular coordinate system is established with the center point of the image data, the same physical dimension is drawn in all image data, the coordinate data of each secondary node located in the image data is marked, and the corresponding feature column is stored in the link of the corresponding independent sub-node and secondary node; all independent sub-nodes and secondary nodes of the same master node are traversed until the correction chain recording the coordinate data of each link is generated.

[0063] As shown in the embodiment: the fluctuation label is "strawberry", and in image 1 and image 2, image 1 contains other content labels "cherry" and "tree"; image 2 contains other content labels "strawberry" and "table"; the following environment characteristics are formed:

[0064] "strawberry"→image 1→"cherry", "tree";

[0065] →image 2→"strawberry", "table";

[0066] The coordinate data and feature column are added to form the correction chain as follows:

[0067] "strawberry"→image 1→"cherry (coordinate 11, feature column 11)", "tree (coordinate 12, feature column 12)";

[0068] →image 2→"strawberry (coordinate 21, feature column 21)", "table (coordinate 22, feature column 22)".

[0069] Step S300 includes the following specific steps:

[0070] Step S310: The mechanical anomaly refers to that the system does not mark the information of the content label, and the analysis anomaly refers to that the system marks the information of the content label incorrectly; the other content labels recorded in the abnormal marking event of each stable label are collected, the content label corresponding to the same information as the incorrect marking information of the stable label is bound as the abnormal label, and an abnormal pair corresponding to the stable label is formed; if the information marked incorrectly by the stable label is not the same as the other content labels of the image, the content label with the maximum similarity to the feature column corresponding to the stable label is extracted to form an abnormal pair; as shown in the embodiment: if the stable label is "lychee", the image data further includes content labels "strawberry" and "table"; when the image that should be marked as lychee is marked as strawberry in the corresponding abnormal marking event, the abnormal pair of the stable label "lychee" at this time is "strawberry"; if the image that should be marked as lychee is marked as longan, and there is no longan image in the image data, the feature columns of the other content labels in the image data are analyzed for similarity, for example, the feature column of the strawberry includes "red", "conical", and "surface particles"; the feature column of the table includes "red-brown" and "square"; the one with the highest similarity is the strawberry; and the abnormal pair of the stable label "lychee" at this time is "strawberry";

[0071] Step S320: All abnormal marking events are traversed to generate an abnormal pair set corresponding to each stable label; each abnormal pair in the abnormal pair set is taken as a retrieval target, and it is determined whether the same retrieval target is recorded in the valid image marking event; if not, the abnormal pair is output as a pre-warning response pair; if yes, the valid image marking event and the abnormal marking event recording the retrieval target are extracted to form an event group to be analyzed;

[0072] Step S330: The feature modal types A1 and A2 extracted in the image marking analysis and identification stage of the valid image marking event and the abnormal marking event in the event group to be analyzed are obtained, and the feature modal type refers to the modal data type corresponding to the feature extraction in the image marking; the difference modal type a of the event group to be analyzed is calculated as A1∩A2;

[0073] When a=∅, the corresponding abnormal pair is marked as a pre-warning response pair;

[0074] When the number of difference modal types contained in a is one, the difference modal type is extracted as a correction identification requirement array of the corresponding stable label when the corresponding stable label and the abnormal pair are in the same image data; the correction identification requirement array refers to the modal data type required to be extracted when the stable label is in the image in the image marking analysis and identification, which is used to assist the feature identification to output the corresponding marking information;

[0075] When the number of difference modal types contained in a is greater than one, the number of times D1 and the number of times D2 of outputting correct labeling in the historical labeled events containing the abnormal pair of each difference modal type are extracted, and the effective labeling participation index Q of each difference modal type is calculated using the formula: Q=D1 / (D1+D2). The greater the effective labeling participation index, the higher the influence degree of the difference modal type on improving the labeling accuracy in the historical labeled events containing the abnormal pair, and the greater the distinguishing effect. All the difference modal types contained in a are sorted in descending order according to the corresponding effective labeling participation index, and a demand sequence is generated. The demand sequence is used as the correction recognition demand array of the corresponding stable label when the abnormal pair is in the same image data.

[0076] As shown in the embodiments: the difference modal types contained in a are text modal and audio modal, the text modal corresponds to extracting semantic features, and the audio modal corresponds to extracting audio features.

[0077] If the abnormal pair is litchi and strawberry, all image data containing litchi and strawberry labels and having text modal as a difference modal type are extracted, the number of times of labeled events and the number of times of labeled errors are recorded respectively, the effective labeling participation index is calculated, and all image data containing litchi and strawberry labels and having audio modal as a difference modal type are extracted, the number of times of labeled events and the number of times of labeled errors are recorded respectively, the effective labeling participation index is calculated, and the index sizes of the two difference modal types are compared, sorted, if the participation index corresponding to the text modal is greater, the text modal is preferentially selected to assist in correction recognition when the image data contains litchi and strawberry features for recognition and labeling, thereby improving the accuracy of similar object recognition.

[0078] Step S400 includes the following specific steps:

[0079] The other content labels and corresponding coordinate data and feature columns contained in the image data in the abnormal labeling events of each fluctuation label are extracted, the correction chain corresponding to the fluctuation label stored in the history is extracted, and the same correction chain as the link formed by the independent sub-nodes and the secondary nodes in the correction chain is searched and compared. If the feature column recorded in the main node of the comparison correction chain is different from the feature column recorded in the fluctuation label, the analysis result is output as a label abnormality, otherwise it is output as a label normal. When the label is abnormal, it can be explained that the generation of the corresponding abnormal labeling event is due to the labeling error caused by the difference in the extraction of the feature column of the label itself.

[0080] In the present application, the fluctuation label returns to step S300 to output the output of the correction identification demand array when it is determined that the label is normal. The correction needs to be distinguished and analyzed according to the correction chain corresponding to the fluctuation label. One correction chain corresponds to one type of fluctuation label. That is, in step S310, the abnormal labeling events in which each fluctuation label is located are distinguished according to all abnormal labeling events corresponding to the same type of fluctuation label. For example, the fluctuation label e has two correction chains, respectively recording the feature columns: feature 1, feature 2; and recording the feature columns: feature 2, feature 3; when analyzing the abnormal labeling events in which each fluctuation label is located, the abnormal labeling events 1, 2, and 3 corresponding to the correction chain 1 form an analysis set, and the abnormal labeling events 4 and 5 corresponding to the correction chain 2 form an analysis set, which is still analyzing the fluctuation label e.

[0081] Step S500 includes the following processes:

[0082] When the real-time response image labeling analysis identification phase, when the label record is a pre-warning response pair, directly pre-warning response abnormal risk; if it is determined that the content to be labeled is a stable label, the content label that meets the historical record abnormal pair in the real-time image data in which the stable label is located is extracted, and the corresponding correction identification demand array is extracted based on the abnormal pair for pre-warning response; the abnormal pair indicates that the problem that is easy to cause labeling error in the same image data with the stable label, so when the real-time image data to be labeled exists the content label that can form an abnormal pair with the stable label, the correction identification demand array is extracted for further verification, and the accuracy of image labeling is improved, such as the correction identification demand array being text modal data, assisting the system to better label;

[0083] If it is determined that the content to be labeled is a fluctuation label, the correction of the fluctuation label and the correction chain is performed first, the feature column of the fluctuation label is replaced when the abnormality in the pre-warning feature column is extracted, the feature column is the same as the environment feature of the correction chain; and the content label that meets the historical record abnormal pair under the same correction chain is extracted from the real-time image data in which the fluctuation label after the replacement of the feature column is located, and the corresponding correction identification demand array is extracted based on the abnormal pair for pre-warning response.

[0084] An image labeling intelligent analysis system based on multi-modal fusion, the system comprising a label distinguishing module, a correction chain analysis module, an abnormal labeling event distinguishing module, a correction identification demand array analysis module, and a pre-warning response module;

[0085] The label distinguishing module is used to collect the content labels of each image data in the historical storage image labeling events and the feature columns required for generating the content labels, to extract effective image labeling events to analyze whether each content label has scene difference in the effective image labeling events, and to output stable labels and fluctuation labels based on the analysis results;

[0086] The correction chain analysis module is configured to analyze the environmental features in the effective image annotation events where the fluctuation labels are located, and generate correction chains corresponding to the fluctuation labels.

[0087] The abnormal annotation event distinguishing module is configured to extract the historical stored abnormal annotation events, and distinguish the abnormal types of the abnormal annotation events as mechanical abnormalities and analysis abnormalities.

[0088] The correction recognition requirement array analysis module is configured to analyze the abnormal annotation events where the stable labels and the fluctuation labels are located when analyzing the analysis abnormalities, and construct the correction recognition requirement arrays in the image annotation analysis and recognition stage.

[0089] The early warning response module is configured to respond to the correction recognition requirement arrays corresponding to the stable labels and the fluctuation labels in the image annotation analysis and recognition stage.

[0090] The correction chain analysis module comprises an environmental feature generation unit and a chain data increasing unit.

[0091] The environmental feature generation unit is configured to take each fluctuation label as a main node, record each effective image annotation event in a target event set as an independent sub-node directly connected to the main node for each main node, extract other content labels recorded by each independent sub-node except the fluctuation label corresponding to the main node as a secondary node connected to the independent sub-node, and construct image data comprising the main node, the independent sub-node and the secondary node as the environmental feature corresponding to the main node.

[0092] The chain data increasing unit is configured to mark the coordinates of each secondary node in the image data and store the corresponding feature column in the link between the corresponding independent sub-node and the secondary node, and traverse all independent sub-nodes and secondary nodes of the same main node until a correction chain recording the coordinate data of each link is generated.

[0093] The correction recognition requirement array analysis module comprises an abnormal pair extraction unit, a to-be-analyzed event group construction unit and a difference modal type analysis unit.

[0094] The abnormal pair extraction unit is configured to extract the content labels having the maximum similarity with the feature column corresponding to the stable label as an abnormal pair.

[0095] The to-be-analyzed event group construction unit is configured to take each abnormal pair in the abnormal pair set as a search target, search whether the same search target is recorded in the effective image annotation event, output the abnormal pair as an early warning response pair if no record is found, and extract the effective image annotation event and the abnormal annotation event recording the search target to construct a to-be-analyzed event group if there is a record.

[0096] The difference modal type analysis unit is configured to analyze the difference modal type based on the feature modal type in the to-be-analyzed event group, and output the correction recognition requirement array in different cases.

[0097] It should be pointed out finally that the above only describes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent replacements to some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An intelligent image annotation and analysis method based on multimodal fusion, characterized in that: The method includes the following steps: Step S100: Collect the content tags and feature columns required for generating content tags from the image data of each image in the historical image annotation events; the feature columns refer to the multi-source features extracted from each content in the image data for matching tags; extract valid image annotation events to analyze whether there are scene differences in each content tag in the valid image annotation events, and output stable tags and fluctuating tags based on the analysis results; Step S200: Analyze the environmental features in the valid image annotation events where the wave label is located, and generate the corresponding correction chain for the wave label; Step S300: Extract historically stored abnormal annotation events, distinguish the abnormality types of abnormal annotation events as mechanical abnormalities and analysis abnormalities; when mechanical abnormalities occur, respond to the system prompts the model for maintenance; when analysis abnormalities occur, analyze the abnormal annotation events of stable labels, and construct an array of correction and recognition requirements for each stable label in the image annotation analysis and recognition stage. Step S400: When analyzing anomalies, compare and analyze the correction chain triggered by the fluctuation label response; when the analysis result is that the label is abnormal, warn of abnormal feature column extraction; when the analysis result is that the label is normal, return to step S300 to output the correction and recognition requirement array of each fluctuation label in the image annotation analysis and recognition stage. Step S500: In the real-time response image annotation analysis and recognition stage, an early warning response is given to the array of correction and recognition requirements that meet the requirements of stable labels and fluctuating labels.

2. The intelligent image annotation and analysis method based on multimodal fusion according to claim 1, characterized in that: Step S100 includes the following specific steps: Step S110: The effective image annotation event refers to the annotation event in which the content label of the output image identified by the system matches the actual image content; using the content label as the retrieval item, search for all effective image annotation events in the historical records that contain the retrieval item to form the target event set of the content label; extract the required feature columns corresponding to the retrieval item for each effective image annotation event in the target event set as the target feature columns; and bind the target feature columns with the effective image annotation events to form an association pair; Step S120: Analyze all associated pairs formed under each search term. Take any two associated pairs whose image data similarity is less than the similarity threshold as valid image annotation events as search combinations. Mark the associated pairs with the same target feature column in the search combinations as safe associated pairs. Traverse all associated pairs. If all associated pairs are marked as safe associated pairs, output that the corresponding content label has no scene difference in the valid image annotation events. If there are associated pairs that are not marked as safe associated pairs, output that the corresponding content label has scene difference in the valid image annotation events. Step S130: Output content tags with scene differences as fluctuating tags, and output content tags without scene differences as stable tags.

3. The image annotation intelligent analysis method based on multimodal fusion according to claim 2, characterized in that: The process of generating the corresponding correction chain for the fluctuation label includes the following specific steps: Step S210: Take each wave label as a master node. Each master node records each valid image label event in the target event set as an independent child node directly connected to the master node. Extract the other content labels recorded by each independent child node except for the wave label corresponding to the master node as the secondary nodes connected to the independent child nodes. The image data containing the master node, independent child nodes and secondary nodes constitute the environmental features of the corresponding master node. Step S220: Establish a rectangular coordinate system with the center point of the image data, draw the same physical dimensions in all image data, mark the coordinate data and corresponding feature columns of each secondary node in the image data it is located in, and store them in the link between the corresponding independent child node and the secondary node; traverse all independent child nodes and secondary nodes of the same primary node until a correction chain for storing coordinate data is generated for each link record.

4. The image annotation intelligent analysis method based on multimodal fusion according to claim 2, characterized in that: Step S300 includes the following specific steps: Step S310: The mechanical anomaly refers to the system's failure to label the content tags, and the analytical anomaly refers to the system's incorrect labeling of the content tags; collect other content tags recorded in the anomaly labeling events of each stable tag, bind the content tags corresponding to the incorrect labeling information of the stable tags as anomaly tags, and form an anomaly pair for the corresponding stable tags; If the information incorrectly labeled by the stable label is different from other content labels in the image, the content label with the highest similarity to the feature column of the corresponding stable label is extracted to form the abnormal pair. Step S320: Traverse all anomaly annotation events and generate an anomaly pair set corresponding to each stable label; using each anomaly pair in the anomaly pair set as the retrieval target, search whether the same retrieval target is recorded in the valid image annotation events; If no record is found, the anomaly pair is output as a warning response pair; if a record exists, the valid image annotation events and anomaly annotation events of the record retrieval target are extracted to form an event group to be analyzed. Step S330: Obtain the feature modality types A1 and A2 extracted during the image annotation analysis and recognition stage of the valid image annotation events and abnormal annotation events in the event group to be analyzed. The feature modality type refers to the modality data type used for image annotation during feature extraction; calculate the differential modality type a = A1 ∩ A2 of the event group to be analyzed. When a=∅, the corresponding anomaly pair is marked as an early warning response pair; When the number of differential modal types contained in a is one, the differential modal type is extracted as the correction and recognition requirement array when the corresponding stable label is in the same image data as the anomaly pair; the correction and recognition requirement array refers to the modal data type that needs to be extracted when performing image annotation analysis and recognition on the image where the stable label is located, and when the anomaly pair corresponding to the stable label exists, in order to assist feature recognition and output corresponding annotation information. When the number of differential modal types contained in 'a' is greater than one, extract the number of times each differential modal type was validly labeled (D1) and the number of times it was incorrectly labeled (D2) in the historical labeling events containing the anomaly pair. Calculate the effective labeling participation index Q for each differential modal type using the formula: Q = D1 / (D1 + D2). Sort all differential modal types contained in 'a' in descending order of their corresponding effective labeling participation indices to generate a demand sequence. Use this demand sequence as the correction and recognition demand array for the corresponding stable label when the anomaly pair is in the same image data.

5. The image annotation intelligent analysis method based on multimodal fusion according to claim 3, characterized in that: Step S400 includes the following specific steps: Extract other content tags, corresponding coordinate data, and feature columns from the image data of each fluctuation tag in the abnormal annotation event; extract the correction chain corresponding to the fluctuation tag in the historical storage; find the correction chain that forms the same link as the independent child nodes and secondary nodes in the correction chain for comparison; if the feature columns recorded by the main node in the comparison correction chain are different from the feature columns recorded by the fluctuation tag, the analysis result is output as the tag is abnormal, otherwise the output is the tag is normal.

6. The intelligent image annotation and analysis method based on multimodal fusion according to claim 5, characterized in that; Step S500 includes the following process: During the real-time response image annotation analysis and recognition stage, when the label is recorded as an early warning response pair, an early warning response risk is directly issued; if the content to be annotated is determined to be a stable label, the content label that meets the historical anomaly pair is extracted from the real-time image data where the stable label is located, and the corresponding correction and recognition requirement array is extracted based on the anomaly pair for early warning response. If the content to be labeled is determined to be a fluctuation label, the fluctuation label and the correction chain are first corrected. When the feature column extraction is abnormal, the feature column of the fluctuation label is replaced with the one that is the same as the environmental feature of the correction chain. The content label of the fluctuation label after the feature column is replaced is obtained from the real-time image data where the fluctuation label is located, and the abnormal pair under the same correction chain in the historical record is obtained. The corresponding correction recognition requirement array is extracted based on the abnormal pair for the early warning response.

7. An intelligent image annotation analysis system based on multimodal fusion, using the intelligent image annotation analysis method based on multimodal fusion as described in any one of claims 1-6, characterized in that: The system includes a label differentiation module, a correction chain analysis module, an abnormal labeling event differentiation module, a correction and identification requirement array analysis module, and an early warning response module. The label differentiation module is used to collect the content labels of each image data in the historical image labeling events and the feature columns required for the generation of content labels, extract the effective image labeling events, analyze whether there are scene differences in each content label in the effective image labeling events, and output stable labels and fluctuating labels based on the analysis results. The correction chain analysis module is used to analyze the environmental features in the valid image annotation events where the wave label is located, and generate the correction chain for the corresponding wave label. The anomaly labeling event differentiation module is used to extract historically stored anomaly labeling events and differentiate the anomaly types of the anomaly labeling events into mechanical anomalies and analytical anomalies. The correction and recognition requirement array analysis module is used to analyze the abnormal labeling events of stable and fluctuating labels when analyzing anomalies, and to construct the correction and recognition requirement array in the image labeling analysis and recognition stage. The early warning response module is used to provide early warning responses to the array of correction and recognition requirements corresponding to stable labels and fluctuating labels during the real-time response image annotation analysis and recognition stage.

8. The image annotation intelligent analysis system based on multimodal fusion according to claim 7, characterized in that: The correction chain analysis module includes an environmental feature generation unit and a chain data augmentation unit; The environmental feature generation unit is used to take each fluctuation label as a master node, and each master node records each valid image label event in the target event set as an independent child node directly connected to the master node. It extracts the other content labels recorded by each independent child node except for the fluctuation label corresponding to the master node as the secondary node connected to the independent child node. The image data containing the master node, independent child nodes and secondary nodes constitutes the environmental feature of the corresponding master node. The chain data addition unit is used to mark the coordinate data and corresponding feature columns of each secondary node in the image data it is located in, and store them in the link between the corresponding independent child node and the secondary node; traverse all independent child nodes and secondary nodes of the same primary node until a correction chain is generated to store the coordinate data of each link record.

9. The image annotation intelligent analysis system based on multimodal fusion according to claim 8, characterized in that: The correction identification requirement array analysis module includes an anomaly pair extraction unit, an event group composition unit to be analyzed unit, and a differential modality type analysis unit. The anomaly extraction unit is used to extract the content tag that has the highest similarity to the feature column of the corresponding stable tag to form an anomaly pair; The unit constituting the event group to be analyzed is used to search whether the same search target is recorded in the valid image annotation event, with each anomaly pair in the anomaly pair set as the search target. If no record is found, the anomaly pair is output as a warning response pair; if a record exists, the valid image annotation events and anomaly annotation events of the record retrieval target are extracted to form an event group to be analyzed. The differential modality type analysis unit is used to analyze the differential modality type output correction recognition requirement array under different conditions based on the feature modality type analysis in the event group to be analyzed.

Citation Information

Patent Citations

  • Query-driven large-scale human face data labeling method

    CN106228120A

  • Image sample label correction method and device, electronic equipment and storage medium

    CN117372754A