Video Entity Annotation Confidence Calculation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Single-modal entity annotation methods in video content analysis face accuracy bottlenecks due to incomplete information, and reliance on knowledge bases introduces errors, making verification processes uncertain.
Innovation Solution
A method and apparatus that recognize entities in videos, calculate confidence levels, and expand related entities using a knowledge base to improve annotation accuracy by integrating single-modal and multi-modal fusion, with confidence level calculations to infer and verify entity correctness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If single-modal entity annotation methods are used, then the processing is simple and fast, but the annotation accuracy cannot reach 100% due to incomplete information
Solution Approach 1:
The patent combines multiple single-modal annotation methods (face recognition, video fingerprint recognition, text recognition) into a multi-modal fusion system. By merging the results from different recognition modalities and verifying them against a knowledge base, the system achieves higher entity annotation accuracy while managing complexity through systematic integration of multiple sources.
2Measurement precision
If knowledge base verification is used to improve accuracy, then entity recognition accuracy can be improved, but errors are introduced when the knowledge base itself contains mistakes
Solution Approach 1:
The system implements a feedback mechanism where the knowledge base verification process continuously refines entity recognition results. When discrepancies are found between recognition results and knowledge base information, the system adjusts its confidence levels and can correct errors. The feedback loop allows the system to learn from knowledge base verification outcomes and improve subsequent recognition accuracy while accounting for knowledge base limitations.
3Measurement precision
If multiple recognition methods are fused, then the entity annotation accuracy is improved, but the system complexity and computational burden increase
Solution Approach 1:
The patent segments the multi-modal fusion process into distinct, manageable modules: face recognition module, video fingerprint recognition module, text recognition module, and knowledge base verification module. Each module operates independently to produce entity annotations and confidence levels, which are then integrated. This segmentation reduces system complexity by making each component independent and maintainable while still achieving high overall accuracy through their coordinated operation.
Data Source
AI summary
A method and an apparatus for outputting information are provided according to embodiments of the disclosure. The method includes: recognizing a target video, to recognize at least one entity and obtain a confidence degree of each entity, the entity including a main entity and related entities; matching the at least one entity with a pre-stored knowledge base to determine at least one candidate entity; obtaining at least one main entity by expanding the related entities of the at least one candidate entity based on the knowledge base, and obtaining a confidence degree of the obtained main entity; and calculating a confidence level of the obtained main entity based on the confidence degree of each of the related entities of the at least one candidate entity and the confidence degree of the obtained main entity, and outputting the confidence level of the obtained main entity.


