The invention discloses an AI-based creative thinking auxiliary generation method and
system, relates to the technical field of
artificial intelligence, and solves the problem of inaccurate sound and picture matching in a traditional method by performing
timestamp alignment and
feature extraction on audio data and a visual
image frame and calculating the correlation between the audio data and the visual
image frame by using a cross-
modal attention mechanism. According to the method, user
eye movement track data is introduced, real attention points of a user are mapped into a visual sequence, and an optimized weight matrix is generated by constructing attention masks and fusing model attention weights, so that a generation result is more in line with
perception key points of the user. And meanwhile, a feedback mechanism is established based on the synchronization error
score, and when the sound and the picture are detected to be asynchronous, the visual frame
timestamp can be dynamically adjusted, so that the self-adaptive correction of the content is realized. On the whole, the method has remarkable advantages in the aspects of improving
modal alignment precision, enhancing
user perception consistency and optimizing generation result naturalness.