Multimedia Recognition Model Segmentation for Seesaw Effect
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multimedia recognition models face a 'seesaw phenomenon' where improvements in one recognition task lead to decreased performance in another, affecting media recognition accuracy and generalization capability.
Innovation Solution
The method involves obtaining sample multimedia data, predicting object types and attributes using an initial recognition model, correcting predictions based on deviations, and adjusting the model to create a target recognition model that improves accuracy and reduces prediction errors across tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the initial multimedia recognition model is trained to improve recognition effect of one recognition task, then the recognition effect of that task is improved, but the recognition effect of another recognition task decreases
Solution Approach 1:
The patent segments the recognition tasks into separate prediction branches within the multimedia recognition model. Each branch is dedicated to a specific recognition task (e.g., object type recognition, object attribute recognition), allowing independent optimization of each task without negatively impacting others. This segmentation resolves the seesaw phenomenon by enabling specialized processing for each task while maintaining overall model coherence.
2Measurement precision
If the initial multimedia recognition model is adjusted to reduce prediction errors, then prediction accuracy is improved, but model complexity increases
Solution Approach 1:
The patent introduces prediction deviation as an intermediary parameter that mediates between the initial model predictions and the final corrected predictions. By calculating and applying prediction deviation, the system achieves higher prediction accuracy without substantially increasing model complexity, as the deviation calculation is based on statistical properties rather than adding complex model structures.
Data Source
AI summary
Embodiments of this application disclose a multimedia data processing method performed by a computer device. The method includes: obtaining sample multimedia data, and a labeled object type and a labeled object attribute of a sample object in the sample multimedia data; predicting an object type and an object attribute of the sample object by applying the sample multimedia data to an initial multimedia recognition model; adjusting the initial multimedia recognition model based on the predicted object type, the predicted object attribute, the labeled object type, and the labeled object attribute, to obtain a target multimedia recognition model, the target multimedia recognition model being configured for recognizing a target object type and a target object attribute of an object in target multimedia data. According to this application, media recognition accuracy of a multimedia recognition model can be improved.


