Dynamic material binding and editing system based on AI reasoning

Through the dynamic material binding and editing system based on AI reasoning, the problem of low efficiency of traditional material processing methods has been solved, and high-precision subtitle recognition, automatic storyboard calibration and efficient material pre-selection have been achieved, meeting the diverse needs of creators.

CN120805906APending Publication Date: 2025-10-17JIANGXI HIP HOP TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510964008.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional material processing methods are inefficient, making it difficult to accurately identify subtitle files and storyboards, and unable to efficiently screen out materials that meet creative needs. In addition, the editing function is not flexible enough to meet diverse needs.

Method used

A dynamic material binding and editing system based on AI reasoning is adopted, including a subtitle import and recognition module, a storyboard calibration module, a picture analysis and reasoning module, a material pre-selection and capture module, and a material editing module. Deep learning and reinforcement learning algorithms are used for subtitle recognition, storyboard calibration, material pre-selection and editing, and multimodal fusion technology is combined to improve material matching.

Benefits of technology

It improves the accuracy of subtitle recognition, enhances the automation and accuracy of storyboard calibration, improves the fit of material pre-selection and editing efficiency, reduces the amount of manual operation, and meets the diverse needs of creators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805906A_ABST
    Figure CN120805906A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of animation production, and discloses a dynamic material binding and editing system based on AI reasoning, and the system comprises a subtitle importing and recognition module which is used for importing an SRT subtitle file and automatically recognizing and extracting a subtitle text and corresponding time axis track information; the subtitle calibration module is used for carrying out subtitle calibration on the subtitle text; the picture analysis and reasoning module is used for carrying out picture description reasoning based on the split mirror calibration result and screening out a role, scene and article prop list; the material pre-selection and capture module is used for performing intelligent pre-selection on a material library according to the list and capturing materials; and the material editing module is used for dynamically adjusting and editing the grabbed materials. A high-precision subtitle recognition algorithm fusing lexical analysis and syntactic analysis technologies in natural language processing is adopted through the subtitle import and recognition module, and compared with a traditional recognition algorithm, the recognition accuracy is improved by 30% on complex sentence pattern and special character processing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of animation production, in particular to a dynamic material binding and editing system based on AI reasoning. BACKGROUND

[0002] In the field of film, animation and multimedia content creation, material acquisition and editing is a very key link. The traditional material processing method relies on a large number of manual operations, which is low in efficiency and prone to errors. For example, in processing subtitles, the traditional recognition algorithm has low accuracy in complex sentence processing and special character processing, and it is difficult to accurately identify SRT subtitle files of different formats and different language styles, which brings great difficulty to subsequent shot calibration and picture analysis, and cannot provide accurate data basis. In the aspect of shot calibration, if only relying on manual shot, not only a lot of time and energy are consumed, but also the shot standards of different personnel may be different, affecting the consistency of the work; and the existing automatic shot technology is not intelligent enough, and it is difficult to meet the shot demand of complex content. In the aspects of picture analysis and reasoning, material preselection and grabbing, and material editing, the traditional technology also faces many challenges, and it is difficult to efficiently and accurately filter out the role, scene, and prop list, cannot quickly obtain the material that meets the shot demand, and the material editing function is not rich and flexible enough to meet the increasingly diversified needs of creators. SUMMARY

[0003] (I) Technical problems solved In view of the deficiencies of the prior art, the present application provides a dynamic material binding and editing system based on AI reasoning, which solves the problem of "low efficiency of traditional method" in the above background art.

[0004] (II) Technical solutions To achieve the above purpose, the present application is implemented by the following technical solutions: a dynamic material binding and editing system based on AI reasoning, comprising: a subtitle import and recognition module for importing SRT subtitle files and automatically recognizing and extracting subtitle text and corresponding time axis track information; a shot calibration module for calibrating the subtitle text; a picture analysis and reasoning module for performing picture description reasoning based on the shot calibration result, and filtering out a role, scene, and prop list; a material preselection and grabbing module for intelligently preselecting and grabbing materials from a material library according to the list; a material editing module for dynamically adjusting and editing the grabbed materials; an export and synchronization module for exporting and synthesizing the edited materials into a draft, and backing up an engineering file.

[0005] Preferably, the split-screen calibration module comprises an intelligent split-screen calibration unit and a contrast split-screen calibration unit, the intelligent split-screen calibration unit adopts an intelligent split-screen module based on deep learning to automatically split-screen process the subtitle text, and the contrast split-screen calibration unit is used to upload text information of manual split-screen and compare and calibrate with the intelligent split-screen processing result.

[0006] Preferably, the intelligent split-screen module adopts an improved Transformer model to realize automatic split-screen processing through training on a large amount of simple animation split-screen data.

[0007] Preferably, the contrast split-screen calibration unit compares the manual split-screen with the software processing result through a difference comparison algorithm, and the difference comparison algorithm can automatically mark the split-screen difference part and provide error prompt and modification suggestion.

[0008] Preferably, the picture analysis and reasoning module uses a DeepSeek reasoning model and a picture reasoning module to perform picture description reasoning, the picture reasoning module introduces an attention mechanism optimization algorithm to accurately capture split-screen key information, and the picture analysis and reasoning module further screens out a role, scene, and prop list through a text mining algorithm.

[0009] Preferably, the material preselection and grabbing module comprises an AI preselection function unit and an image retrieval and grabbing unit, the AI preselection function unit adopts knowledge graph technology to intelligently preselect a material library according to the list based on semantic understanding, the image retrieval and grabbing unit adopts a hybrid retrieval algorithm combining local features and global features to analyze and grab preselected materials, in addition, the AI preselection function unit further introduces a multi-modal fusion technology to fuse and analyze text semantic information and visual feature information of the material, construct a multi-modal semantic space, calculate similarity in the multi-modal semantic space according to the role, scene, and prop list, and obtain preselected materials that are more suitable for split-screen requirements, the image retrieval and grabbing unit adopts a reinforcement learning algorithm to take the matching degree of the material and the split-screen as a reward signal, constantly optimize the retrieval strategy, and improve the accuracy of analyzing and grabbing preselected materials, and the material preselection and grabbing module further comprises a federated learning module, through federated learning with other terminal devices or systems, more sample data is obtained to optimize the preselection and grabbing model under the premise of protecting data privacy.

[0010] Preferably, the material editing module comprises a role binding unit, a scene and article processing unit and a comprehensive editing unit, the role binding unit is used for manually binding the role list screened by reasoning, supports selecting different pupil expressions for the role, and the expression can not be set for a specific role, the scene and article processing unit is used for operating the scene list and the article prop list screened by reasoning, calling an AI preselection function for preselection grabbing, and supporting user reselection, and the comprehensive editing unit is used for automatically arranging materials according to subtitles and element results, supporting attribute adjustment of the materials including but not limited to position, zoom, size and color, supporting custom preset, supporting dynamic adjustment of materials in a track, advanced image operation and parent-child element binding nesting.

[0011] Preferably, the track adjustment function of the comprehensive editing unit comprises dynamic adjustment of materials in a track, synchronous positioning to corresponding materials in a canvas by clicking a top-left mark of the materials in the track, and batch adjustment of properties of multiple selected materials in a track.

[0012] Preferably, the advanced image operation function of the comprehensive editing unit comprises manipulation deformation, perspective deformation and pen frame selection operation of any material in a track, and historical record saving of the operation-deformed content.

[0013] Preferably, the export and synchronization module adopts a special data conversion protocol to export the draft to the Moov video editing software for synchronous recognition, and adopts an incremental backup and version control technology to complete backup of the engineering file.

[0014] (Three) beneficial effects The application provides a dynamic material binding and editing system based on AI reasoning. (1) The dynamic material binding and editing system based on AI reasoning has the following beneficial effects:

[0015] (2) The dynamic material binding and editing system based on AI reasoning uses an improved Transformer model in the intelligent storyboard calibration unit of the storyboard calibration module during use, realizes automatic storyboard processing through training on a large amount of stick figure animation storyboard data, compares the manual storyboard with the software processing result through a difference comparison algorithm, automatically marks the storyboard difference part and provides error prompts and modification suggestions. This not only improves the automation degree of storyboard calibration, but also ensures the accuracy of the storyboard through manual comparison and calibration, meets the needs of different users for storyboard calibration, makes the storyboard calibration more in line with the creative intention, and provides a more reasonable storyboard basis for subsequent picture analysis and material processing.

[0016] (3) The dynamic material binding and editing system based on AI reasoning uses the AI preselection function unit of the material preselection and grabbing module to fuse and analyze the text semantic information and the visual feature information of the material based on semantic understanding, uses knowledge graph technology combined with multi-modal fusion technology, constructs a multi-modal semantic space, and the preselected material has an improved 40% fit compared with single modal preselection. In addition, the image retrieval and grabbing unit uses a hybrid retrieval algorithm combining local features and global features and a reinforcement learning algorithm to continuously optimize the retrieval strategy, takes the matching degree of the material and the storyboard as the reward signal, and improves the accuracy of the preselected material analysis and grabbing. This makes the material preselection more in line with the storyboard requirements, the grabbing process more accurate and efficient, and reduces the workload of manual material screening, providing higher quality material resources for material editing. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 The system module framework and the interactive flowchart of the present application are shown. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0019] Please refer to Figure 1 The present application provides a dynamic material binding and editing system based on AI reasoning, which includes a subtitle import and recognition module, a storyboard calibration module, a picture analysis and reasoning module, a material preselection and grabbing module, a material editing module, and an export and synchronization module. Specifically, The subtitle import and recognition module is used for importing an SRT subtitle file and automatically recognizing and extracting the subtitle text and corresponding timeline track information. A high-precision subtitle recognition algorithm is adopted. The algorithm combines morphological analysis and syntactic analysis techniques in natural language processing and can accurately recognize SRT subtitle files in different formats and different language styles. The specific calculation steps are as follows: The input SRT subtitle text is scanned by character stream, the words are recognized by a finite state automaton, and a sequence of lexical units is constructed. For example, for the sentence "The main character walks into a bright room", the morphological analysis will decompose it into "The main character", "walks into", "bright", and "room" lexical units, and assign each lexical unit a corresponding part-of-speech label. Based on the sequence of lexical units, a dependency syntax analysis algorithm is used to construct a syntax tree. Assuming that the set of lexical units is , the dependency syntax analysis calculates the dependency relationship score between words , constructs a dependency relationship matrix , wherein , the score calculation formula is: , wherein is a semantic similarity function, is a syntactic rule matching degree function, and are weight coefficients. By finding the maximum spanning tree algorithm, a syntax tree is constructed from the dependency relationship matrix , so that the grammatical structure of the sentence is clearly presented, and accurate analysis of complex sentence patterns is achieved. Compared with traditional recognition algorithms, the recognition accuracy is improved by 30% in complex sentence patterns and special character processing, providing accurate data basis for subsequent mirror calibration and picture analysis.

[0020] The mirror calibration module is used for mirror calibration of the subtitle text, including an intelligent mirror calibration unit and a contrast mirror calibration unit. The intelligent mirror calibration unit uses an intelligent mirror module based on deep learning to automatically process the subtitle text. The contrast mirror calibration unit is used to upload manually divided text information and compare it with the intelligent mirror processing result. The intelligent mirror module uses an improved Transformer model to automatically process the mirror by training a large amount of simple animation mirror data. In addition, the contrast mirror calibration unit compares the manual mirror with the software processing result by using a difference comparison algorithm. The difference comparison algorithm can automatically mark the mirror difference part and provide error prompts and modification suggestions. The picture analysis and reasoning module is used for picture description reasoning based on the shot calibration result, and screens out a role, a scene, and an article prop list. Specifically, picture description reasoning is performed by using a DeepSeek reasoning model and a picture reasoning module. The picture reasoning module introduces an attention mechanism optimization algorithm to accurately capture shot key information. The picture analysis and reasoning module also screens out a role, a scene, and an article prop list by using a text mining algorithm. When calculating attention weights, first, the word vector sequence of the shot text is encoded to obtain a hidden state . Then, the attention weights are calculated, wherein , , is a context vector, and is a function calculated by a multi-layer perception machine. The function highlights the weights of key plots and core elements, avoids secondary information interference, and makes the generated picture description more suitable for animation production requirements. Then, a role, a scene, and an article prop list are screened out by using a text mining algorithm. The algorithm combines named entity recognition and semantic correlation analysis technology. The named entity recognition uses a BiLSTM-CRF model to process the shot text and identify potential entities (such as role names, scene descriptions, and article prop names). Then, a semantic correlation analysis algorithm is used to calculate the semantic similarity between entities, wherein is a word vector representation of an entity . According to the semantic similarity, the correlation between entities is established, so that relevant elements are accurately extracted from the shot text, and the correlation between elements is established, thereby providing comprehensive data support for subsequent material preselection.

[0021] The material preselection and grabbing module is used for intelligent preselection of a material library according to the list and grabbing of the material, and specifically includes an AI preselection function unit and an image retrieval and grabbing unit. The AI preselection function unit preselects the material library intelligently based on semantic understanding and knowledge graph technology according to the list. The image retrieval and grabbing unit analyzes and grabs the preselected material by using a hybrid retrieval algorithm combining local features and global features. In addition, the AI preselection function unit also introduces a multi-modal fusion technology, fuses and analyzes the text semantic information and the visual feature information of the material, constructs a multi-modal semantic space, calculates the similarity in the multi-modal semantic space according to the role, scene, and prop list, and obtains preselected material that is more suitable for the shot requirement. The image retrieval and grabbing unit uses a reinforcement learning algorithm, takes the matching degree of the material and the shot as a reward signal, constantly optimizes the retrieval strategy, improves the accuracy of the analysis and grabbing of the preselected material, and is provided with a federal learning module. Through federal learning with other terminal devices or systems, more sample data is obtained to optimize the preselection and grabbing model under the premise of protecting data privacy. In the AI preselection function unit, first, the text semantic information is converted into a semantic vector through a word vector model The visual features (color histogram, shape descriptor, texture feature, etc.) of the material are extracted into a visual feature vector through a convolutional neural network . Then, a multi-modal semantic space is constructed, the semantic vector and the visual feature vector are mapped to the same space through linear transformation, and a fusion vector is obtained , wherein , is a transformation matrix, is a bias vector. According to the role, scene, and prop list, the cosine similarity is calculated in the multi-modal semantic space , and preselected material that is more suitable for the shot requirement is obtained. Compared with single-modal preselection, the fitting degree of the preselected material is improved by 40%; Further described, in the image retrieval and grabbing unit, the local features are extracted into key points and feature descriptors in the image by using a SIFT algorithm, and the global features are extracted into overall gradient direction histogram features of the image by using a HOG algorithm. The local features and the global features are fused to obtain a comprehensive feature vector . In the reinforcement learning algorithm, the state space is defined as the current material set and related shot information, the action space is different retrieval strategies (such as adjusting the keyword weight of retrieval, changing the feature matching threshold, etc.), and the reward function is , wherein is the relevance of the material and the shot, is the diversity of the material set, and is a weight coefficient. By constantly trying different retrieval strategies and updating the strategy according to the reward signal, the Q-learning algorithm is used to update the Q value table: wherein is a learning rate, is a discount factor, is a current state, is a current action, is a reward, is a next state, constantly optimizing the retrieval strategy to improve the accuracy of pre-selected material analysis and extraction.

[0022] The material editing module is used for dynamic adjustment and editing of the extracted material. Specifically, it includes a character binding unit, a scene and item processing unit, and a comprehensive editing unit. The character binding unit is used for manual binding of the character list after reasoning and screening, supports selecting different pupil expressions for characters, and can not set expressions for specific characters. The scene and item processing unit is used for operating the scene list and item prop list after reasoning and screening, calling the AI pre-selection function for pre-extraction, and supporting user re-selection. The comprehensive editing unit is used for automatically arranging materials according to subtitles and element results, supports attribute adjustment of materials including but not limited to position, scaling, size, and color adjustment, supports custom presets, supports dynamic adjustment of materials in the track, advanced image operations, and parent-child element binding and nesting. The track adjustment function of the comprehensive editing unit includes dynamic adjustment of materials in the track, synchronous positioning to the corresponding material in the canvas by clicking the upper left corner marker of the material in the track, and batch adjustment of properties of multiple selected materials in the track. In addition, the advanced image operation function of the comprehensive editing unit includes manipulating and deforming any material in the track, perspective deformation, and pen frame selection operation, and saving the history record of the deformed content; The export and synchronization module is used for exporting and synthesizing the edited material into a draft and backing up the engineering file. It uses a special data conversion protocol to export the draft to the Moov video editing software for synchronous recognition, and uses incremental backup and version control technology to complete the backup of the engineering file.

[0023] Although embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made therein without departing from the principles and spirit of the application, the scope of which is defined by the appended claims and their equivalents.

Claims

1. Dynamic material binding and editing system based on AI reasoning, characterized by: include: Subtitle import and recognition module, used to import SRT subtitle files and automatically identify and extract subtitle text and corresponding timeline track information; A storyboard calibration module, used for performing storyboard calibration on the subtitle text; The image analysis and reasoning module is used to perform image description reasoning based on the storyboard calibration results and filter out the list of characters, scenes, and items and props; A material pre-selection and grabbing module, used to intelligently pre-select and grab materials from the material library according to the list; Material editing module, used to dynamically adjust and edit captured materials; The export and synchronization module is used to export the edited materials into a synthesis draft and back up the project files.

2. The AI-based dynamic material binding and editing system according to claim 1, characterized in that: The storyboard calibration module includes an intelligent storyboard calibration unit and a comparative storyboard calibration unit. The intelligent storyboard calibration unit uses a deep learning-based intelligent storyboard module to automatically storyboard the subtitle text. The comparative storyboard calibration unit is used to upload the text information of the manual storyboard and compare and calibrate it with the intelligent storyboard processing results.

3. The AI-based dynamic material binding and editing system according to claim 2, characterized in that: The intelligent storyboard module adopts an improved Transformer model and realizes automatic storyboard processing by training a large amount of simple animation storyboard data.

4. The AI-based dynamic material binding and editing system according to claim 2, characterized in that: The comparison storyboard calibration unit compares the manual storyboard with the software processing result through a difference comparison algorithm. The difference comparison algorithm can automatically mark the difference part of the storyboard and provide error prompts and modification suggestions.

5. The AI-based dynamic material binding and editing system according to claim 1, characterized in that: The picture analysis and reasoning module uses the DeepSeek reasoning model and the picture reasoning module to perform picture description reasoning. The picture reasoning module introduces the attention mechanism optimization algorithm to accurately capture the key information of the storyboard. The picture analysis and reasoning module also uses the text mining algorithm to filter out the list of characters, scenes, and items and props.

6. The AI-based dynamic material binding and editing system according to claim 1, characterized in that: The material pre-selection and capture module includes an AI pre-selection functional unit and an image retrieval and capture unit. The AI ​​pre-selection functional unit is based on semantic understanding and adopts knowledge graph technology to intelligently pre-select the material library according to the list. The image retrieval and capture unit adopts a hybrid retrieval algorithm that combines local features and global features to analyze and capture the pre-selected materials. In addition, the AI ​​pre-selection functional unit also introduces multimodal fusion technology to fuse and analyze text semantic information with visual feature information of the material to construct a multimodal semantic space. Similarity calculation is performed in the multimodal semantic space based on the list of characters, scenes, and items and props to obtain pre-selected materials that are more in line with the storyboard requirements. The image retrieval and capture unit adopts a reinforcement learning algorithm and uses the matching degree between the material and the storyboard as a reward signal to continuously optimize the retrieval strategy and improve the accuracy of the analysis and capture of the pre-selected materials. The material pre-selection and capture module is also provided with a federated learning module. By performing federated learning with other terminal devices or systems, more sample data is obtained to optimize the pre-selection and capture model while protecting data privacy.

7. The AI-based dynamic material binding and editing system according to claim 1, characterized in that: The material editing module includes a role binding unit, a scene and item processing unit, and a comprehensive editing unit. The role binding unit is used to manually bind the role list after reasoning and screening, supports selecting different pupil expressions for the role, and does not need to set expressions for specific roles. The scene and item processing unit is used to operate the scene list and item prop list after reasoning and screening, calls the AI ​​pre-selection function for pre-selection and capture, and supports users to replace the selection again. The comprehensive editing unit is used to automatically arrange the material according to the subtitles and element results, supports the adjustment of the material attributes including but not limited to position, scaling, size, and color, supports custom presets, supports dynamic adjustment of materials within the track, advanced image operations, and parent-child element binding and nesting.

8. The AI-based dynamic material binding and editing system according to claim 7, characterized in that: The track adjustment function of the comprehensive editing unit includes supporting dynamic adjustment of the materials in the track. Clicking the mark in the upper left corner of the material in the track can synchronously locate the corresponding material in the canvas, and supporting batch adjustment of properties of multiple selected materials in the track.

9. The AI-based dynamic material binding and editing system according to claim 7, characterized in that: The advanced image manipulation functions of the comprehensive editing unit include manipulation and deformation, perspective deformation, and pen selection of any material in the track, and historical record preservation of the content after the operation and deformation.

10. The AI-based dynamic material binding and editing system according to claim 1, characterized in that: The export and synchronization module uses a dedicated data conversion protocol to export the draft to the Jianying software for synchronous recognition, and uses incremental backup and version control technology to complete the backup of the project files.