A multimodal data adaptive annotation method, system and smart terminal

By using an adaptive annotation method, which automatically generates annotation content using templates and human feedback, the problem of high cost and low efficiency of manual annotation is solved, and a fast, flexible and accurate annotation process is achieved.

CN120217198BActive Publication Date: 2025-10-28SICHUAN BAIJIA DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510279732.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-10-28
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Existing manual annotation methods require a large amount of human resources, are costly, and have low annotation efficiency.

Method used

By acquiring data information and annotation locations, determining the data type and environment, generating concise data, sending it to the human end, receiving feedback on the annotation method number, automatically generating annotation content using the template, and adjusting it in conjunction with the human annotation content to optimize the annotation model.

Benefits of technology

It improves the speed and efficiency of annotation, enhances the flexibility and accuracy of annotation, and reduces the workload of manual evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217198B_ABST
    Figure CN120217198B_ABST
Patent Text Reader

Abstract

This invention relates to a multimodal data adaptive annotation method, system, and intelligent terminal, belonging to the field of data annotation technology. It includes acquiring data information and annotation locations; determining the data type, data content, and data environment at the annotation location based on the annotation location; integrating the data type and data environment to form concise data and sending it to a preset human terminal; receiving the annotation method number fed back from the human terminal and searching for the corresponding annotation template from a preset template database; inputting the data content and data environment into a preset annotation model to obtain automatically filled content; and filling the annotation template with the automatically filled content to obtain the actual annotation content and performing annotation. This invention automatically generates annotations based on templates and actual conditions, eliminating the need for manual annotation and improving the speed and efficiency of annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data annotation technology, and in particular to a multimodal data adaptive annotation method, system and intelligent terminal. Background Technology

[0002] Multimodal data refers to data containing multiple different types of information, such as text, images, audio, and video. The core of multimodal data lies in enhancing the understanding and processing of information by fusing data from different modalities. For example, combining images and text can be used for image-text recognition, while combining audio and video can be used for video caption generation.

[0003] With the widespread adoption of devices such as the Internet of Things, social media, and surveillance cameras, the generation rate of multimedia data such as images and videos is growing exponentially. This massive amount of data requires efficient and accurate annotation for subsequent analysis and mining. Currently, the annotation method for images, videos, and text is usually done manually.

[0004] Regarding the aforementioned technologies, existing manual annotation methods require a large amount of human resources, are costly, and have low annotation efficiency. Summary of the Invention

[0005] To address the problems of existing manual annotation methods requiring significant human resources, incurring high costs, and exhibiting low annotation efficiency, this invention provides a multimodal data adaptive annotation method, system, and intelligent terminal.

[0006] Firstly, the present invention provides a multimodal data adaptive annotation method, which adopts the following technical solution:

[0007] An adaptive annotation method for multimodal data includes:

[0008] Obtain data information and label locations;

[0009] Determine the data type, data content, and data environment of the data information at the labeled location;

[0010] Based on data type and data environment, a short data set is generated and sent to the pre-set human terminal;

[0011] Receive the annotation method number from the human terminal and find the corresponding annotation template from the preset template database;

[0012] Input the data content and data environment into the preset annotation model to obtain automatically populated content;

[0013] The autofill content will fill the annotation template to obtain the actual annotation content and then annotate it.

[0014] By adopting the above technical solution, the annotation location is first determined, then a brief message is generated and sent to the human end. Upon receiving the annotation number from the human end, the template can be directly referenced. The template automatically generates annotations based on the actual situation, eliminating the need for manual annotation and improving the speed and efficiency of annotation.

[0015] Optional, also includes:

[0016] When receiving the annotation method number, receive the manually annotated content fed back by the human terminal;

[0017] Analyze manually annotated content to determine the annotation type;

[0018] Based on the annotation type, manually annotated content is replaced or added to the actual annotation content to update it to the actual annotation content.

[0019] By adopting the above technical solution, if the received content contains manually entered annotations, the corresponding annotations in the template can be replaced, thus improving the flexibility and accuracy of annotation.

[0020] Optionally, methods for replacing or adding manually labeled content to actual labeled content based on labeled type to form adjusted labeled content include:

[0021] When manually annotated content exists, the actual annotated content that has not yet been replaced or added is defined as the expected annotated content.

[0022] Determine the corresponding auto-fill content based on the expected annotation content and annotation type, and define the auto-fill content as the auto-annotation content.

[0023] When the automatically labeled content is missing, manually labeled content is added to the actual labeled content to create adjusted labeled content.

[0024] When automatically annotated content exists, determine the content deviation based on the manually annotated content and the automatically annotated content;

[0025] The corresponding allowable deviation range is retrieved from the preset error database based on the annotation type;

[0026] When the content deviation falls within the allowable deviation range, the manually annotated content is replaced with the automatically annotated content to form the adjusted annotation content;

[0027] When the content deviation does not fall within the allowable deviation range, a mandatory content prompt is generated based on the annotation type;

[0028] Send brief data and required field prompts to all other human terminals, and re-receive feedback on human annotations;

[0029] The most frequently identical manually annotated content among all other manually annotated entries is defined as new manually annotated content.

[0030] When new manually annotated content and automatically annotated content are the same, no adjustment content will be generated for the annotation;

[0031] When the new manually annotated content is the same as the manually annotated content, the manually annotated content will replace the automatically annotated content to form the adjusted annotation content.

[0032] By adopting the above technical solution, when the manual annotations and the annotations in the template are different, multiple people are needed to evaluate them in order to verify whether the annotation content is correct, thus improving the accuracy of the annotation content.

[0033] Optionally, methods for generating mandatory content prompts based on annotation type when the content deviation does not fall within the allowable deviation range include:

[0034] Retrieve historically labeled content stored in a preset historical repository based on data type, data content, and data environment;

[0035] The proportion of manually annotated content should be the same as that of historically annotated content.

[0036] The automatic equal ratio is determined based on the automatically annotated content and the historical annotated content;

[0037] The ratio deviation is determined based on both automatic and manual identical ratios;

[0038] When the ratio deviation exceeds the preset threshold for ignoring manual input, no adjustment will be made and no mandatory field prompt will be generated.

[0039] When the ratio deviation is less than the preset automatic ignore threshold, the manually annotated content will be replaced with the automatically annotated content to form the adjusted annotation content without forming a mandatory content prompt;

[0040] When the proportional deviation is greater than the automatic threshold but less than the manual threshold, a mandatory content prompt is generated based on the annotation type.

[0041] By adopting the above technical solution, based on historical data, if the results are basically the same as those obtained manually or automatically, it indicates that the other one may have caused a temporary problem. In this case, multiple evaluations are not required, reducing the workload caused by multiple evaluations and improving the efficiency of annotation.

[0042] Optionally, replacing automatically labeled content with manually labeled content also includes the following methods:

[0043] A training set is formed based on manually labeled content, data information, and data types;

[0044] The labeled model is trained based on the training set to obtain an optimized labeled model and then output it.

[0045] Optionally, it also includes a method for generating short data and sending it to a human, the method including:

[0046] Retrieve historical data;

[0047] Determine information similarity based on data information and historical data information;

[0048] When there is information similarity greater than a preset similarity threshold, the historical data information is defined as similar historical data information;

[0049] The similar historical annotation method number is determined based on similar historical data information and annotation location;

[0050] The similar historical annotation method number is received as the annotation method number, and no short data is sent to the human end;

[0051] When there is no information similarity greater than the similarity threshold, a short data set is generated and sent to the human end.

[0052] By adopting the above technical solution, when there are historical annotations that are very similar, the annotation can be set directly according to the historical number, and even the annotation number does not need to be marked by personnel, which further improves the speed and efficiency of annotation.

[0053] Optionally, methods for determining information similarity based on data information and historical data information include:

[0054] Based on the data type, determine the impact data in the data information and the historical impact data in the historical data information;

[0055] The corresponding data weight is retrieved from the preset weight database based on the data type.

[0056] The similarity of the baseline is determined by comparing the impact data with the corresponding historical impact data;

[0057] A weighted average is calculated based on all cardinal similarities and their corresponding data weights, and this average is output as the information similarity.

[0058] Optional, also includes:

[0059] Based on the data type, identify key data in the data information and historical key data in the historical data information;

[0060] When key data and historical key data are consistent, determine the impact data in the data information and the historical impact data in the historical data information based on the data type.

[0061] When key data and historical key data are inconsistent, the preset dissimilarity will be used as the information similarity for output.

[0062] By adopting the above technical solution, the similarity algorithm is based on the premise that the data must be consistent. If they are inconsistent, no calculation is required, which improves the rigor of similarity calculation.

[0063] Secondly, this invention provides a multimodal data adaptive annotation system, which adopts the following technical solution:

[0064] A multimodal data adaptive annotation system, comprising:

[0065] The acquisition module is used to acquire data information, annotation location, annotation method number, and manual annotation content;

[0066] A memory for storing the program of the control method for the multimodal data adaptive annotation method as described above;

[0067] The processor loads and executes programs from memory.

[0068] By adopting the above technical solution, the annotation location is first determined, then a brief message is generated and sent to the human end. Upon receiving the annotation number from the human end, the template can be directly referenced. The template automatically generates annotations based on the actual situation, eliminating the need for manual annotation and improving the speed and efficiency of annotation.

[0069] Thirdly, the present invention provides a smart terminal, which adopts the following technical solution:

[0070] A smart terminal includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed using any of the methods described above.

[0071] By adopting the above technical solution, the annotation location is first determined, then a brief message is generated and sent to the human end. Upon receiving the annotation number from the human end, the template can be directly referenced. The template automatically generates annotations based on the actual situation, eliminating the need for manual annotation and improving the speed and efficiency of annotation.

[0072] In summary, the present invention includes at least one of the following beneficial technical effects:

[0073] 1. The template automatically generates annotations based on the actual situation, eliminating the need for manual annotation and improving the speed and efficiency of annotation;

[0074] 2. If the received content contains manually entered annotations, the corresponding annotations in the template can be replaced, improving the flexibility and accuracy of the annotations;

[0075] 3. The system can be set directly according to the historical numbering, eliminating the need for personnel to annotate the numbers, which further improves the speed and efficiency of annotation. Attached Figure Description

[0076] Figure 1 This is a flowchart of a multimodal data adaptive annotation method in an embodiment of this application.

[0077] Figure 2 This is a flowchart illustrating a method for replacing or adding manually labeled content to actual labeled content based on the labeled type, as described in this application embodiment, to form an adjusted labeled content.

[0078] Figure 3 This is a flowchart of a method for generating mandatory content prompts based on annotation type when the content deviation does not fall within the allowable deviation range, as described in this application embodiment.

[0079] Figure 4 This is a flowchart illustrating the method for generating short data and sending it to a human in this application embodiment.

[0080] Figure 5 This is a flowchart of a method for determining information similarity based on data information and historical data information in an embodiment of this application.

[0081] Figure 6 This is a block diagram of a multimodal data adaptive annotation method in an embodiment of this application. Detailed Implementation

[0082] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0083] This invention discloses an adaptive annotation method for multimodal data. (Refer to...) Figure 1 An adaptive annotation method for multimodal data includes:

[0084] Step 100: Obtain data information and label locations.

[0085] The data information refers to all information about the data that requires manual annotation. This data includes various formats such as text, images, and videos. The annotation location refers to the position in the data information where annotations need to be made. Both the data information and the annotation location are obtained through manual input or input from the corresponding media.

[0086] Step 101: Determine the data type, data content, and data environment of the data information at the marked location.

[0087] Data type refers to the type of data information at the labeled location, including text, images, audio, and video. Data content refers to the content of the data. Data environment refers to the environment in which the data exists. This can be determined by built-in functions or tools in the programming language. Data content can be directly read or obtained through corresponding viewing tools, such as a debugger. Data environment refers to the storage, runtime, or usage environment of the data. This can be determined using code analysis tools, such as SonarQube and Checkmarx, to analyze data types and data flow in the code. Alternatively, data visualization tools, such as Tableau and Power BI, can be used to display and analyze the structure and content of the data.

[0088] Step 102: Based on the data type and data environment, integrate the data to form a short summary and send it to the preset human terminal.

[0089] Brief data refers to concise data information designed for easy viewing and analysis by users, eliminating the need for reading and understanding the entire data. It is presented using keywords. The data content is not included here to avoid requiring human analysis. The "human end" refers to the mobile or fixed device that the information can be received by the personnel who will be manually annotating it, such as a mobile phone.

[0090] Step 103: Receive the annotation method number from the human terminal and find the corresponding annotation template from the preset template database.

[0091] The annotation method number is a number assigned by the corresponding staff member after reviewing the brief data and considering the appropriate annotation method for the given location. The annotation template is the template corresponding to the annotation method number. The database stores the mapping relationship between annotation method numbers and annotation templates. The staff member responsible for this task sets up a series of templates, assigns a number to each template, and then records all the numbers and their corresponding templates. When the system receives a corresponding annotation method number, it automatically retrieves the corresponding annotation template from the database and outputs it.

[0092] Step 104: Input the data content and data environment into the preset annotation model to obtain automatically populated content.

[0093] The annotation model is a model trained using a deep neural network. This model is mainly used to generate corresponding annotations based on the data. The auto-fill content is generated based on the actual data content and data environment, producing partial content suitable for annotation output, which, combined with templates, results in complete content.

[0094] Step 105: Fill the annotation template with the autofill content to obtain the actual annotation content and then annotate it.

[0095] Step 106: When receiving the annotation method number, receive the manual annotation content fed back by the manual terminal.

[0096] The manually annotated content refers to content annotated manually. This content is entered here along with the annotation method number.

[0097] If there is no manually annotated content, proceed directly to step 105 and end the process.

[0098] Step 107: Analyze the manually annotated content to obtain the annotation type.

[0099] The annotation type refers to the type of manually annotated content. Here, the type does not refer to the format of the annotation, but rather to the corresponding annotation category.

[0100] Step 108: Replace or add manually annotated content to the actual annotation content based on the annotation type to update the annotation to the actual annotation content.

[0101] Reference Figure 2 Methods for replacing or adding manually labeled content to actual labeled content based on labeled type to create adjusted labeled content include:

[0102] Step 200: When there is manually annotated content, the actual annotated content that has not yet been replaced or added is defined as the expected annotated content.

[0103] Step 201: Determine the corresponding auto-fill content based on the expected annotation content and annotation type, and define the auto-fill content as the auto-annotation content.

[0104] Here, the automatic annotation content corresponding to the annotation type is obtained.

[0105] Step 202: When the automatically labeled content does not exist, add the manually labeled content to the actual labeled content to form the adjusted labeled content.

[0106] If the automatically labeled content does not exist, it means that the content that conflicts with the manually labeled content does not exist. In this case, the manually labeled content can be directly added to the actual labeled content to form the adjusted labeled content.

[0107] Step 203: When automatically annotated content exists, determine the content deviation based on manually annotated content and automatically annotated content.

[0108] The presence of automatically labeled content indicates that there are two instances of the same label type. Therefore, attention needs to be paid to the corresponding deviation. Content deviation refers to the difference between manually labeled content and automatically labeled content. This is calculated using a database lookup method. When the system receives the corresponding manually labeled content and automatically labeled content, it automatically retrieves the corresponding deviation from the database. If the content is a number, it can also be calculated using a formula.

[0109] Step 204: Find the corresponding allowable deviation range from the preset error database based on the annotation type.

[0110] The allowable deviation range is the range of deviations that can be tolerated for a given annotation type. This deviation range is set manually. The database stores the mapping relationship between annotation types and allowable deviation ranges. This is determined by experts in the field who experiment with different deviation values ​​for each annotation type to obtain a common-sense threshold, which is then recorded. When the system receives a corresponding annotation type, it automatically retrieves the corresponding allowable deviation range from the database and outputs it.

[0111] Step 205: When the content deviation falls within the allowable deviation range, replace the automatically labeled content with the manually labeled content to form the adjusted labeled content.

[0112] When the content deviation falls within the allowable deviation range, it means that although there is some deviation or no deviation between the manually annotated content and the automatically annotated content, it is within an understandable range. Therefore, the subjective content of the human annotation can be taken as the main factor, and the manually annotated content is replaced with the automatically annotated content to form the adjusted annotation content.

[0113] Step 206: When the content deviation does not fall within the allowable deviation range, generate a mandatory content prompt based on the annotation type.

[0114] The "Required Fields" prompt reminds the human operator to include this information in the feedback. The format here can be directly output as a textual expression of the annotation type; for example, if the annotation type is "Dimension," simply output "Dimension."

[0115] If the content deviation does not fall within the allowable deviation range, it indicates that there is an inconsistency between the manually annotated content and the automatically annotated content, and this is not a minor inconsistency. Therefore, it is necessary to determine whether the error is due to human error or error in the automatically filled-in system.

[0116] Step 207: Send the brief data and required field prompts to all other human terminals, and re-receive the feedback from human annotations.

[0117] The purpose of sending this message is to require all other human-based terminals to add the corresponding human annotation content for the annotation type.

[0118] Step 208: Filter out the human-annotated content with the most identical content from all other human-annotated content and define it as new human-annotated content.

[0119] The purpose of filtering is to determine what the correct content is.

[0120] Step 209: When the new manual annotation content and the automatic annotation content are the same, no adjustment annotation content is generated.

[0121] If the new manually annotated content and the automatically annotated content are the same, it means that the automatic annotation is correct and there is no need to replace it.

[0122] Step 210: When the new manual annotation content is the same as the manual annotation content, replace the automatic annotation content with the manual annotation content to form the adjusted annotation content.

[0123] If the new manually added annotation is the same as the original manually added annotation, it means that the original manually added annotation is identical. In this case, the annotation will be directly replaced to correct the error caused by the automatic filling.

[0124] Here, when the new manually labeled content differs from both the automatically labeled content and the manually labeled content, an alarm can be issued directly to remind staff to perform timely repairs.

[0125] Reference Figure 3 Methods for generating mandatory content prompts based on annotation type when the content deviation does not fall within the allowable deviation range include:

[0126] Step 300: Retrieve historical annotation content stored in the preset historical repository based on data type, data content, and data environment.

[0127] A historical repository is a database that stores annotations made during a historical process that have been verified as correct, along with related content. Historical annotations are annotations stored in the database that are essentially identical in data type, content, and environment. The historical repository stores the mapping relationship between data type, content, environment, and historical annotation content. It can be obtained by those skilled in the art by storing all data type, content, and environment when storing historical annotations.

[0128] Step 301: Determine the same proportion of manually annotated content and historically annotated content.

[0129] The "human-identical proportion" refers to the percentage of historical annotations that are identical to manually annotated content among all historical annotations. It is determined by first comparing the two sets of annotations, counting the number of identical historical annotations, and then dividing by the total number.

[0130] Step 302: Determine the automatic same ratio based on the automatically annotated content and the historical annotated content.

[0131] The automatic identical ratio refers to the proportion of historical annotations that are identical to the automatically annotated content among all historical annotations. This is determined by first comparing the two sets of annotations, counting the number of identical historical annotations, and then dividing by the total number.

[0132] Step 303: Determine the ratio deviation based on automatic and manual identical ratios.

[0133] The proportional deviation is the difference between the same proportional measurement performed manually and the same proportional measurement performed automatically. It is calculated by subtracting the two.

[0134] Step 304: When the ratio deviation is greater than the preset threshold for ignoring manual input, no adjustment will be made and no mandatory content prompt will be generated.

[0135] The threshold for ignoring manual annotations is set to a value exceeding which deviations in manually annotated content can be disregarded. When the proportional deviation exceeds this threshold, it indicates a large automatic similarity ratio, which can be ignored. This suggests that the automatically annotated content is more accurate, and no adjustments are needed. Furthermore, there's no need to generate mandatory content prompts, reducing the workload of sending information to other human endpoints and receiving feedback.

[0136] Step 305: When the ratio deviation is less than the preset automatic ignore threshold, replace the automatic annotation with the manual annotation to form the adjusted annotation and not to form the required content prompt.

[0137] If the value is less than the automatic threshold, it means that the deviation of the automatically labeled content can be ignored. When the ratio deviation is less than the automatic threshold, it means that the manual same ratio is large, and the automatic same ratio can be ignored. This indicates that the manual labeling content is more accurate, so the manual labeling content replaces the automatic labeling content to form adjusted labeling content without creating a mandatory content prompt.

[0138] Step 306: Form a training set based on manually labeled content, data information, and data types.

[0139] The training set is the collection of data used to train the labeled model. Using the training set, the algorithm can optimize parameters such as the model's weights and biases, enabling the model to perform well on the training data.

[0140] Step 307: Train the labeled model based on the training set to obtain an optimized labeled model and output it.

[0141] Step 308: When the proportional deviation is greater than the automatic threshold value but less than the manual threshold value, a mandatory content prompt is generated based on the annotation type.

[0142] If the ratio deviation is greater than the automatic threshold for ignoring but less than the manual threshold for ignoring, it means that no manually annotated or automatically annotated content can be ignored at this time, and the method in steps 200-210 needs to be followed.

[0143] Reference Figure 4 It also includes a method for whether to generate short data and send it to the human end, the method including:

[0144] Step 400: Retrieve historical data information.

[0145] Historical data information refers to data from historical processes. This data can be retrieved either from a new historical database or from the historical repository in step 300.

[0146] Step 401: Determine the information similarity based on the data information and historical data information.

[0147] Information similarity refers to the degree of similarity between currently received data and historical data. This can be achieved by breaking down all content of the data, comparing each element at different locations and categories, and then using a normalization method to convert it into similarity scores for the same category and unit. Further details can be provided in subsequent steps and will not be elaborated upon here.

[0148] Step 402: When there is information similarity greater than a preset similarity threshold, the historical data information is defined as similar historical data information.

[0149] The similarity threshold is a value that is set manually. If the similarity is exceeded, the similarity is very high and the content can be used as a reference.

[0150] Step 403: Determine the number of similar historical annotation methods based on similar historical data information and annotation location.

[0151] The similarity history annotation method number is the number used when similar historical data was annotated at the annotation location. The determination method can be either a search method, meaning all content is stored in the database, or it can include the annotation method number selected at the annotation location.

[0152] Step 404: The similar historical annotation method number is received as the annotation method number, and no short data is generated and sent to the human end.

[0153] When the similarity is high, the historical labeling method can be used as a reference for numbering, thus eliminating the need to send the data to the human end and reducing the workload of the human end.

[0154] Step 405: When there is no information similarity greater than the similarity threshold, generate short data and send it to the human end.

[0155] If it does not exist, it means that there is no reference in the historical process, and a brief data set still needs to be generated and sent to the human end.

[0156] Reference Figure 5 Methods for determining information similarity based on data information and historical data information include:

[0157] Step 500: Determine the key data in the data information and the historical key data in the historical data information based on the data type.

[0158] Key data refers to the most crucial data within the overall data set, as it determines the degree of similarity. Historical key data refers to the most critical data within the historical data set.

[0159] Step 501: When the key data and historical key data are consistent, determine the impact data in the data information and the historical impact data in the historical data information based on the data type.

[0160] Influencing data refers to data within the data information that has a certain impact on the annotation and affects the content of the annotation. Historical influencing data refers to historical data that has a certain impact on the annotation and affects the content of the annotation.

[0161] Here, both the impact data and historical impact data will affect the similarity between the two.

[0162] If the key data and historical key data are consistent, it means that the data that must be consistent is consistent, and the similarity can be determined directly based on other data.

[0163] Step 502: Find the corresponding data weight from the preset weight database based on the data type.

[0164] Data weight is the proportion of similarity corresponding to data types.

[0165] Step 503: Determine the base similarity by comparing the impact data with the corresponding historical impact data.

[0166] Cardinal similarity is the degree of similarity between the influencing data and the corresponding historical influencing data. It can be calculated by directly comparing the corresponding types of data, such as numerical comparisons.

[0167] Step 504: Calculate the weighted average based on all cardinal similarities and their corresponding data weights, and output it as the information similarity.

[0168] The weighted average is an important statistical measure that takes into account the relative importance of each data point. By appropriately allocating weights, the weighted average can more accurately reflect the overall characteristics of the data. In practical applications, the selection of weights needs to be based on the specific problem's context and objectives.

[0169] Step 505: When the key data and historical key data are inconsistent, output the preset dissimilarity as the information similarity.

[0170] Dissimilarity is the degree of complete dissimilarity. When the key data and historical key data are inconsistent, it means that the most crucial information is inconsistent. In this case, there is no need to consider other factors, and the dissimilarity can be output directly.

[0171] Based on the same inventive concept, embodiments of the present invention provide a multimodal data adaptive annotation system.

[0172] Reference Figure 6 A multimodal data adaptive annotation system, comprising:

[0173] The acquisition module is used to acquire data information, annotation location, annotation method number, and manual annotation content;

[0174] A memory for storing the program of the control method for the multimodal data adaptive annotation method as described above;

[0175] The processor loads and executes programs from memory.

[0176] Based on the same inventive concept, embodiments of the present invention provide a smart terminal, including a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as a multimodal data adaptive annotation method.

[0177] Those skilled in the art will clearly understand that for the sake of convenience and brevity, the division of the above-mentioned functional modules is only used as an example for illustration. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-mentioned systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0178] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A multimodal data adaptive annotation method, characterized in that, include: Obtain data information and label locations; Determine the data type, data content, and data environment of the data information at the labeled location; Based on data type and data environment, a short data set is generated and sent to the pre-set human terminal; Receive the annotation method number from the human terminal and find the corresponding annotation template from the preset template database; Input the data content and data environment into the preset annotation model to obtain automatically populated content; The autofill content is used to fill in the annotation template to obtain the actual annotation content and then annotate it; This also includes: When receiving the annotation method number, receive the manually annotated content fed back by the human terminal; Analyze manually annotated content to determine the annotation type; Based on the annotation type, replace or add manually annotated content to the actual annotation content to update it to the actual annotation content; The methods for replacing or adding manually annotated content to the actual annotated content based on the annotation type to form adjusted annotation content include: When manually annotated content exists, the actual annotated content that has not yet been replaced or added is defined as the expected annotated content. Determine the corresponding auto-fill content based on the expected annotation content and annotation type, and define the auto-fill content as the auto-annotation content. When the automatically labeled content is missing, manually labeled content is added to the actual labeled content to create adjusted labeled content. When automatically annotated content exists, determine the content deviation based on the manually annotated content and the automatically annotated content; The corresponding allowable deviation range is retrieved from the preset error database based on the annotation type; When the content deviation falls within the allowable deviation range, the manually annotated content is replaced with the automatically annotated content to form the adjusted annotation content; When the content deviation does not fall within the allowable deviation range, a mandatory content prompt is generated based on the annotation type; Send brief data and required field prompts to all other human terminals, and re-receive feedback on human annotations; The most frequently identical manually annotated content among all other manually annotated entries is defined as new manually annotated content. When new manually annotated content and automatically annotated content are the same, no adjustment content will be generated for the annotation; When the new manually annotated content is the same as the manually annotated content, the manually annotated content will replace the automatically annotated content to form the adjusted annotation content.

2. The multimodal data adaptive annotation method according to claim 1, characterized in that, Methods for generating mandatory content prompts based on annotation type when the content deviation does not fall within the allowable deviation range include: Retrieve historically labeled content stored in a preset historical repository based on data type, data content, and data environment; The proportion of manually annotated content should be the same as that of historically annotated content. The automatic equal ratio is determined based on the automatically annotated content and the historical annotated content; The ratio deviation is determined based on both automatic and manual identical ratios; When the ratio deviation exceeds the preset threshold for ignoring manual input, no adjustment will be made and no mandatory field prompt will be generated. When the ratio deviation is less than the preset automatic ignore threshold, the manually annotated content will be replaced with the automatically annotated content to form the adjusted annotation content without forming a mandatory content prompt; When the proportional deviation is greater than the automatic threshold but less than the manual threshold, a mandatory content prompt is generated based on the annotation type.

3. The multimodal data adaptive annotation method according to claim 2, characterized in that, After replacing automatically labeled content with manually labeled content, the following methods are also included: A training set is formed based on manually labeled content, data information, and data types; The labeled model is trained based on the training set to obtain an optimized labeled model and then output it.

4. The multimodal data adaptive annotation method according to claim 1, characterized in that, It also includes a method for whether to generate a short data entry and send it to a human, the method including: Retrieve historical data; Determine information similarity based on data information and historical data information; When there is information similarity greater than a preset similarity threshold, the historical data information is defined as similar historical data information; The similar historical annotation method number is determined based on similar historical data information and annotation location; The similar historical annotation method number is received as the annotation method number, and no short data is sent to the human end; When there is no information similarity greater than the similarity threshold, a short data set is generated and sent to the human end.

5. The multimodal data adaptive annotation method according to claim 4, characterized in that, Methods for determining information similarity based on data information and historical data information include: Based on the data type, determine the impact data in the data information and the historical impact data in the historical data information; The corresponding data weight is retrieved from the preset weight database based on the data type. The similarity of the baseline is determined by comparing the impact data with the corresponding historical impact data; A weighted average is calculated based on all cardinal similarities and their corresponding data weights, and this average is output as the information similarity.

6. The multimodal data adaptive annotation method according to claim 5, characterized in that, Also includes: Based on the data type, identify key data in the data information and historical key data in the historical data information; When key data and historical key data are consistent, determine the impact data in the data information and the historical impact data in the historical data information based on the data type. When key data and historical key data are inconsistent, the preset dissimilarity will be used as the information similarity for output.

7. A multimodal data adaptive annotation system, characterized in that, include: The acquisition module is used to acquire data information, annotation location, annotation method number, and manual annotation content; A memory for storing the program of the control method for the multimodal data adaptive annotation method as described in any one of claims 1 to 6; The processor loads and executes programs from memory.

8. A smart terminal, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent filling system and method based on large model

    CN119514513A