Audio book figure surround animation effect generation method, terminal equipment and storage medium

By layering the animation layer and combining audio timing and storyline characteristics, the problem of single animation effects of audio book is solved, and the dynamic matching of animation and phonological rhythm and plot development is achieved, which significantly improves the emotional fit between animation and content.

CN120047586AInactive Publication Date: 2025-05-27SHENZHEN MAIFENG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510519605.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The animation effects of existing audio books are relatively single, lacking dynamic interaction and immersion, and the accuracy of character extraction and cutout is limited, making it difficult to deal with complex backgrounds and details, resulting in unreal animation effects.

Method used

By layering the technical architecture of character animation layer, foreground animation layer and background animation layer, combining the audio timing and storyline characteristics in the audio book content information, the animation effect is achieved dynamic matching with the phonological rhythm and plot development.

Benefits of technology

It achieves dynamic matching of animation effects and phonological rhythm and plot development, breaks through the mechanized presentation method of the traditional single animation layer, and allows the character, foreground and background to change synergistically according to the internal space-time logic of the audio book, significantly improving the emotional compatibility between animation and content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047586A_ABST
    Figure CN120047586A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of data processing, and discloses an audio book figure surrounding animation effect generation method, terminal equipment and a storage medium. The audio book character surrounding animation effect generation method comprises the steps that a character animation layer and a foreground animation layer are generated according to content information of a target audio book, and the content information comprises story plots, audio data and an audio time sequence of the audio data; generating a background animation layer according to the foreground animation layer and the content information; and synthesizing the foreground animation layer, the character animation layer and the background animation layer to obtain an animation effect corresponding to the target audio book. According to the method, the background light and shadow change and the foreground can dynamically form relevance adjustment, a traditional mechanical presentation mode of a single animation layer is broken through, the role, the foreground and the background can cooperatively change according to the spatial-temporal logic in the audio book, and the emotional integrating degree of the animation and the content is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and particularly relates to a method for generating an animated effect of surround characters in an audiobook, a terminal device, and a storage medium. Background Art

[0002] With the popularization of digital content consumption, audiobooks, as a new media form, have been widely welcomed.

[0003] Currently, the animated effects of audiobooks are relatively single. For example, the character actions and background light and shadow changes cannot be matched in real time according to the voice emotion or plot conflict, resulting in a mechanical animation presentation. A new technical means is needed to solve the above technical problems. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method for generating an animated effect of surround characters in an audiobook, a terminal device, and a storage medium, which can solve the problems in the related art that the animated effects of audiobooks are relatively single, lack dynamic interaction and immersion, and have limited accuracy in character extraction and matte extraction, making it difficult to process complex backgrounds and details, resulting in unrealistic animated effects.

[0005] The first aspect of the present invention provides a method for generating an animated effect of surround characters in an audiobook, including: Generating a character animation layer and a foreground animation layer according to the content information of the target audiobook, where the content information includes the plot, audio data, and audio time sequence of the audio data; Generating a background animation layer according to the foreground animation layer and the content information; Combining the foreground animation layer, the character animation layer, and the background animation layer to obtain the animated effect corresponding to the target audiobook.

[0006] Optionally, in the first implementation manner of the first aspect of the present invention, after the step of combining the foreground animation layer, the character animation layer, and the background animation layer to obtain the animated effect corresponding to the target audiobook, the method further includes: Real-time detecting an interaction instruction initiated by a user; If the interaction instruction is received, then in response to the interaction instruction, adjusting the animated effect.

[0007] Optionally, in the second implementation manner of the first aspect of the present invention, the step of if the interaction instruction is received, then in response to the interaction instruction, adjusting the animated effect includes: If the interaction instruction is received, render parameters corresponding to the interaction instruction are matched, where the render parameters include first render sub-parameters corresponding to the character animation layer, second render sub-parameters corresponding to the foreground animation layer, and / or third render sub-parameters corresponding to the background animation layer; Adjust the animation effect according to the render parameters.

[0008] Optionally, in the third implementation manner of the first aspect of the present invention, the step of generating the character animation layer and the foreground animation layer according to the content information of the target audiobook includes: Obtain a static picture; Extract the character layer picture from the static picture; Generate the character animation layer according to the character layer picture and the content information, and generate the foreground animation layer according to the content information of the target audiobook.

[0009] Optionally, in the fourth implementation manner of the first aspect of the present invention, the step of generating the foreground animation layer according to the content information of the target audiobook includes: Generate a basic foreground animation layer according to the content information of the target audiobook, and analyze the emotional parameters in the audio time sequence; When it is determined that there is an intense emotion paragraph according to the emotional parameters, adjust the dynamic effect segment corresponding to the intense emotion paragraph in the basic foreground animation layer to obtain the foreground animation layer.

[0010] Optionally, in the fifth implementation manner of the first aspect of the present invention, the step of generating the basic foreground animation layer according to the content information of the target audiobook includes: Obtain the content information of the target audiobook; Perform a matching operation in a preset material library according to the plot in the content information to obtain the basic foreground animation layer.

[0011] Optionally, in the sixth implementation manner of the first aspect of the present invention, the step of extracting the character layer picture from the static picture includes: Extract the initial character layer picture from the static picture according to a character extraction large model; Perform texture mapping on the initial character layer picture according to OpenGL to obtain the character layer picture.

[0012] Optionally, in the seventh implementation manner of the first aspect of the present invention, the step of synthesizing the foreground animation layer, the character animation layer, and the background animation layer to obtain the animation effect corresponding to the target audiobook includes: Match the basic picture layer of the audiobook in a preset material library; The basic image layer, the background animation layer, the character animation layer, and the foreground animation layer are sequentially superimposed to obtain the animation effect corresponding to the target audio book.

[0013] In a second aspect, an embodiment of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for generating surround animation effects of audio book characters when executing the computer program.

[0014] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method for generating surround animation effects of audio book characters are implemented.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes the above-mentioned method for generating surround animation effects of audio book characters.

[0016] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows: through the technical architecture of hierarchically generating character animation layers, foreground animation layers and background animation layers, combined with the audio timing and storyline features in the audio book content information, the dynamic matching of animation effects with voice rhythm and plot development is achieved. By driving the generation of character animation layers and foreground animation layers based on audio timing data, the character actions and voice playback progress can be synchronized in real time; by regenerating the background animation layer based on the foreground animation layer and content information, the background light and shadow changes can be adjusted in a correlated manner with the foreground dynamics, breaking through the traditional mechanized presentation method of a single animation layer, so that the characters, foreground and background can change in coordination according to the inherent time and space logic of the audio book, significantly improving the emotional fit between the animation and the content. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0018] Figure 1 A schematic diagram of an embodiment of a method for generating a surround animation effect of an audio book character in an embodiment of the present invention; Figure 2 It is a schematic diagram of a specific embodiment after step S103 of the method for generating surround animation effects of audio book characters in an embodiment of the present invention; Figure 3It is a schematic diagram of a specific embodiment of step S101 of the method for generating an animated effect of surround characters in an audiobook according to an embodiment of the present invention; Figure 4 It is a schematic diagram of a specific embodiment of step S1013 of the method for generating an animated effect of surround characters in an audiobook according to an embodiment of the present invention; Figure 5 It is a schematic diagram of an embodiment of a terminal device according to an embodiment of the present invention. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.

[0020] It should be noted that the terms "including", "comprising" and "having" and any variations thereof in the description and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, terminal, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. In the terms in the claims, description and drawings of the present invention, relational terms such as "first" and "second" are only used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any such actual relationship or order between these entities / operations / objects.

[0021] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0022] With the popularization of digital content consumption, audiobooks, as a new media form, have been widely welcomed.

[0023] Currently, the animation effects of audiobooks are relatively single. For example, the character actions and background light and shadow changes cannot be matched in real time according to the voice emotion or plot conflict, resulting in a mechanical animation presentation. A new technical means is needed to solve the above technical problems.

[0024] In view of this, an embodiment of the present invention provides a method for generating an animated effect of surrounding characters in an audiobook, a terminal device, and a storage medium. Through a technical architecture that hierarchically generates a character animation layer, a foreground animation layer, and a background animation layer, and combines the audio timing and the plot characteristics in the content information of the audiobook, a dynamic matching between the animation effect and the speech rhythm and the plot development is achieved. Based on the audio timing data to drive the generation of the character animation layer and the foreground animation layer, the character actions can be synchronized with the speech playback progress in real time; by generating the background animation layer based on the foreground animation layer and the content information, the background light and shadow changes can be adjusted in association with the foreground dynamics, breaking through the mechanical presentation method of the traditional single animation layer, enabling the character, the foreground, and the background to change synergistically according to the internal time and space logic of the audiobook, and significantly improving the emotional fit between the animation and the content.

[0025] In order to illustrate the technical solution of the present invention, the following will be described through specific embodiments.

[0026] Figure 1 FIG. shows a schematic flowchart of the implementation of a method for generating an animated effect of surrounding characters in an audiobook provided by an embodiment of the present invention. This method can be applied to a terminal device. The terminal device can be a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, etc.

[0027] Specifically, the above method for generating an animated effect of surrounding characters in an audiobook may include the following steps S101 to S103.

[0028] Step S101, generate a character animation layer and a foreground animation layer according to the content information of the target audiobook, where the content information includes the plot, audio data, and the audio timing of the audio data.

[0029] In an embodiment of the present invention, the terminal device (the device for generating the animated effect of surrounding characters in an audiobook) analyzes the content information of the target audiobook, including its plot, original audio data, and audio timing (such as timestamps, paragraph markers, etc.).

[0030] Based on the plot, determine the character elements to be animated (such as character actions, expressions); at the same time, combine the time series characteristics of the audio data (such as speech rhythm, paragraph intervals) to design an appropriate animation frame sequence for the character, and generate an independent character animation layer.

[0031] Further, generate a foreground animation layer according to the scene description or preset rules in the content information, such as surface dynamic elements related to the plot such as dynamic light effects, text floating special effects, etc. This step realizes the precise association between the animation effect and the content of the audiobook by disassembling the content information into independent animation data of the character and the foreground.

[0032] Step S102, generating a background animation layer according to the foreground animation layer and content information.

[0033] In an embodiment of the present invention, based on the character and foreground animation layers, the dynamic characteristics of the foreground animation layer (such as changes in light effect intensity, special effect coverage) and the scene context in the audio book content information (such as the type of environment in which the story takes place, atmosphere description) are further analyzed to dynamically generate a background animation layer.

[0034] For example, if the foreground animation layer contains a high-frequency flickering light effect, the background animation layer can generate a corresponding dynamic light and shadow diffusion effect; if the content information indicates that the scene is rainy, the background animation layer generates a looping animation of raindrops falling or fog flowing. By linking the background animation layer with the foreground animation layer and the content information, the background dynamics are consistent with the overall scene logic, enhancing the visual hierarchy.

[0035] Step S103, synthesizing the foreground animation layer, the character animation layer and the background animation layer to obtain the animation effect corresponding to the target audio book.

[0036] In the embodiment of the present invention, through multi-layer rendering technology (such as OpenGL texture overlay), according to the preset overlay order (for example: background animation layer as the bottom layer, character animation layer in the middle, foreground animation layer on the top layer), the dynamic elements of each layer are merged into a unified animation effect. During the synthesis process, each frame of animation is synchronized and calibrated according to the audio timing to ensure that the character action, foreground special effects and background dynamics strictly match the audio playback progress.

[0037] For example, in an audio-tagged dialogue segment, the lip movements of the character animation layer are synchronized with the speech waveform, while the light and shadow changes of the background animation layer adjust the gradient speed according to the emotional intensity of the dialogue, ultimately outputting a surround animation effect that is highly synchronized with the audiobook content.

[0038] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows: through the technical architecture of hierarchically generating character animation layers, foreground animation layers and background animation layers, combined with the audio timing and storyline features in the audio book content information, the dynamic matching of animation effects with voice rhythm and plot development is achieved. By driving the generation of character animation layers and foreground animation layers based on audio timing data, the character actions and voice playback progress can be synchronized in real time; by regenerating the background animation layer based on the foreground animation layer and content information, the background light and shadow changes can be adjusted in a correlated manner with the foreground dynamics, breaking through the traditional mechanized presentation method of a single animation layer, so that the characters, foreground and background can change in coordination according to the inherent time and space logic of the audio book, significantly improving the emotional fit between the animation and the content.

[0039] The existing interaction methods for audiobook animations rely on fixed instructions, resulting in rigid user operations and the inability to express complex intentions. For example, users cannot directly adjust the dynamic intensity of animation elements through gestures or flexibly switch interaction modes according to scene requirements. Based on this, an alternative embodiment of the present invention is proposed.

[0040] Referring to Figure 2 , Figure 2 FIG. is a schematic diagram of a specific embodiment after step S103 of the method for generating an audiobook character surround animation effect in an embodiment of the present invention. The following specific implementation manners are further included after step S103.

[0041] Step S201, continuously detect the interaction instructions initiated by the user in real time.

[0042] In an embodiment of the present invention, after generating the animation effect of the target audiobook, the user operations are continuously monitored through a preset interaction interface (such as touch event monitoring, voice command recognition, or physical button response). The types of interaction instructions include but are not limited to switching animation styles, adjusting the intensity of dynamic elements (such as the strength of light effects), pausing / resuming animation playback, etc. The user input signals are captured through an event-driven mechanism, and the instruction types and operation parameters (such as sliding direction, click position, voice keywords) are parsed to provide data support for subsequent responses.

[0043] Optionally, the user interaction behaviors are captured in real time through a multimodal input interface (such as a camera, touch screen, gyroscope sensor), including but not limited to gesture actions (such as waving, pinching, sliding), voice commands, or changes in the tilt angle of the device. For example, when the user draws a "circle" gesture in the air, it indicates switching the background style, or adjusts the animation playback speed by pinching the fingers. The recognition of interaction instructions is based on a pre-trained AI model (such as a gesture recognition model, voice semantic parsing model), which converts unstructured inputs (such as gesture images, voice waveforms) into structured instruction types and parameters. The input signals are continuously monitored and invalid operations (such as accidental touches) are filtered to ensure the accuracy and real-time performance of instruction recognition.

[0044] Step S202, if an interaction instruction is received, then in response to the interaction instruction, adjust the animation effect.

[0045] In an embodiment of the present invention, when it is confirmed that a valid interaction instruction is received, the rendering logic of the current animation effect is dynamically modified according to the instruction type. For example, if the user triggers the "switch background style" instruction, a new background animation layer is loaded from the preset resource library to replace the original layer; if the user adjusts the "character action speed", the frame rate parameters of the character animation layer are recalculated and the rendering process is updated in real time. The adjustment process needs to be synchronized with the audio playback progress to avoid the problem of out-of-sync audio and video caused by interaction operations. The adjusted animation effect takes effect immediately through the rendering engine, and the user can immediately perceive the visual changes.

[0046] Optionally, after recognizing a valid interaction instruction, the user behavior is associated with a specific animation rendering parameter adjustment logic through an instruction mapping table. For example, a "swipe right" gesture is mapped to "switch to the next foreground special effect", and a "pinch with two fingers" is mapped to "reduce the character animation frame rate". The adjustment process is based on a hierarchical control mechanism: independent adjustable parameters (such as transparency, motion trajectory, particle special effect density) are defined for the character animation layer, foreground animation layer, and background animation layer respectively, enabling the user's instructions to act precisely on the target layer. For example, when the user magnifies a certain area through a gesture, the detail rendering of the character animation layer in that area is dynamically enhanced, while the complexity of the background animation layer is weakened, achieving focus-guided animation optimization.

[0047] In the embodiments of the present invention, by supporting natural interaction methods such as gestures, users are given more intuitive and flexible animation control capabilities. Through an AI model, unstructured instructions (such as gesture trajectories, speech semantics) are automatically parsed, and the user's intentions are dynamically mapped to the adjustment strategies of animation rendering parameters, getting rid of the limitations of traditional fixed instructions (such as button clicks).

[0048] The existing adjustment methods for audiobook animations are usually global modifications (such as overall acceleration / slowdown), and users cannot finely customize the visual effects according to their needs. For example, when a user hopes to enhance the expressive power of a character's actions, the traditional method may be forced to synchronously change the background dynamics, destroying the scene coordination. Based on this, the present invention proposes an optional embodiment.

[0049] Step S201 further includes the following specific implementation manners.

[0050] Step S2011, if the interaction instruction is received, match the rendering parameters corresponding to the interaction instruction, where the rendering parameters include the first rendering sub-parameters corresponding to the character animation layer, the second rendering sub-parameters corresponding to the foreground animation layer, and / or the third rendering sub-parameters corresponding to the background animation layer.

[0051] In the embodiments of the present invention, after detecting a user interaction instruction, the specific rendering parameters to be adjusted by the instruction are parsed through a preset instruction-parameter mapping table. The rendering parameters are split into three categories according to the animation layer. The first rendering sub-parameters corresponding to the character animation layer (such as character action frame rate, displacement trajectory), the second rendering sub-parameters corresponding to the foreground animation layer (such as light effect intensity, particle special effect density), and the third rendering sub-parameters corresponding to the background animation layer (such as dynamic blur coefficient, color gradient speed).

[0052] For example, when the user issues an instruction of "enhancing character actions", the "frame rate increase ratio" and "displacement amplitude coefficient" in the first rendering sub-parameters are matched; if the instruction is "reducing background complexity", the "motion blur level" and "layer transparency" in the third rendering sub-parameters are matched.

[0053] Step S2012, adjust the animation effect according to the rendering parameters.

[0054] In the embodiments of the present invention, according to the matched rendering parameters, independent adjustments are respectively performed on the character, foreground, and background animation layers: modifying the bone animation interpolation algorithm of the character animation layer to adapt to the new frame rate parameters; updating the particle emitter attributes (such as quantity, motion speed) of the foreground animation layer to change the light effect intensity; adjusting the shader parameters (such as blur radius, hue mixing weight) of the background animation layer to achieve complexity control. During the adjustment process, the timing synchronization of each layer is maintained to ensure that the playback progress of the multi-layer animation after parameter change is strictly aligned with the audio timing, avoiding visual tearing or out-of-sync audio and video caused by hierarchical adjustment.

[0055] In the embodiments of the present invention, through mapping user interaction instructions to independent adjustments of hierarchical rendering parameters, multi-dimensional fine control of the animation effect is achieved. The user can simultaneously trigger coordinated adjustments such as accelerating character actions, enhancing foreground light effects, and deepening background hues through a single interaction instruction (such as "enhancing dramatic effects") without having to operate layer by layer. The hierarchical parameter linkage mechanism not only ensures the flexibility of animation effect adjustment (such as only modifying foreground parameters without affecting other layers), but also supports global effect optimization across layers, solving the problems of coarse animation adjustment granularity and single effect in traditional technologies, and significantly improving the user's precise control ability over animation presentation.

[0056] The character extraction of traditional audiobook animations relies on manual matte painting or simple threshold segmentation, resulting in rough character edges and lost details (such as hair, semi-transparent clothing) when dealing with complex backgrounds, and static pictures are difficult to directly convert into dynamic effects. Based on this, an alternative embodiment of the present invention is proposed.

[0057] Refer to Figure 3 , Figure 3 is a schematic diagram of a specific embodiment of step S101 of the method for generating an audiobook character surround animation effect in the embodiments of the present invention. Step S101 further includes the following specific embodiments.

[0058] Step S1011, obtain a static picture.

[0059] In an embodiment of the present invention, static pictures matching the current storyline are obtained from the material library associated with the audiobook or user-uploaded resources. For example, if the target audiobook chapter describes "an adventure in the forest", then a static image of a character in a forest scene is selected. The sources of the static pictures include preset illustrations, third-party copyrighted materials, or user-defined content.

[0060] Step S1012: Extract the character layer picture from the static picture.

[0061] In an embodiment of the present invention, through a pre-trained large model for human extraction (such as segmentation models like DeepLabv3, U-Net, etc.), semantic segmentation is performed on the static picture to separate the character main area and generate an initial character layer picture. The model identifies features such as character edges and clothing details through a convolutional neural network, accurately removing complex backgrounds (such as overlapping trees, light and shadow interference), ensuring that the edges of the character layer are smooth and free of residual noise. For example, for a picture of a character standing in front of a bookshelf, the model can accurately segment the character outline and eliminate the interference of the bookshelf background.

[0062] Step S1013: Generate a character animation layer based on the character layer picture and content information, and generate a foreground animation layer according to the content information of the target audiobook.

[0063] In an embodiment of the present invention, based on the character layer picture and the content information of the audiobook (such as audio timings, emotion tags), a dynamic character animation layer is generated: According to the key time nodes in the audio timings, bone animations or deformation animations are added to the character layer picture. For example, in a dialogue paragraph, character lip-sync actions and limb swing animations are generated through keyframe interpolation. Combining with the scene description in the content information (such as "stormy night"), dynamic elements (such as raindrops falling, lightning special effects) are matched from a preset animation library to generate a foreground animation layer adapted to the storyline. For example, a particle effect layer of sword glows and flashes is generated in the "intense battle" plot.

[0064] In the embodiment of the present invention, by accurately extracting the character layer from the static picture and independently generating animations, the authenticity of the character animations and the resource utilization efficiency are significantly improved. Traditional methods rely on pre-rendered character animation materials, while this solution dynamically segments the characters in the static picture through an AI model and generates an adapted action sequence based on the audio timings, which not only reduces the animation production cost but also avoids problems such as edge jagging or background residue caused by traditional matte painting techniques. For example, in a complex background scene, the character layer can be independently rendered into a smooth lip-sync animation without redrawing the character materials, while the foreground animation layer dynamically generates special effects such as light effects / snow and rain according to the plot, enhancing the fit between the animation and the content.

[0065] The foreground animation effect of traditional audiobooks is fixed and cannot be dynamically adjusted according to the emotional changes of the audio (such as intense conversations and gentle narrations), resulting in a mechanical animation presentation. Based on this, an alternative embodiment of the present invention is proposed.

[0066] Referring to Figure 4 , Figure 4 FIG. is a schematic diagram of a specific embodiment of step S1013 of the method for generating an audiobook character surrounding animation effect in an embodiment of the present invention. Step S1013 further includes the following specific implementation manners.

[0067] Step S10131: Generate a basic foreground animation layer according to the content information of the target audiobook, and parse the emotional parameters in the audio time series.

[0068] In the embodiment of the present invention, basic foreground animation resources are matched based on the content information of the target audiobook. For example, an explosion spark special effect template is loaded according to the "war scene".

[0069] At the same time, emotional analysis is performed on the audio time series data to extract emotional parameters (such as intonation intensity, speech rate change, background music intensity), and the audio paragraphs are labeled with emotional labels such as "calm" and "intense" through a pre-trained AI model (such as an LSTM-based emotion classifier). For example, when a sudden increase in high-frequency energy and an increase in speech rate are detected in the audio waveform, the paragraph is labeled as an intense emotion paragraph.

[0070] Step S10132: When it is determined that there is an intense emotion paragraph according to the emotional parameters, adjust the corresponding dynamic effect segment of the intense emotion paragraph in the basic foreground animation layer to obtain the foreground animation layer.

[0071] In the embodiment of the present invention, when an intense emotion paragraph (such as a shouting segment in a battle scene) is recognized, the dynamic effect segment of the basic foreground animation layer corresponding to the paragraph is enhanced and adjusted. For example, the particle emission frequency is increased in the explosion spark template, and the spark diffusion range is expanded. The displacement speed of the dynamic elements is increased (such as the flying speed of sword lights and shadows), and the movement path is recalculated through the key frame interpolation algorithm. Additional special effects (such as screen vibration) are dynamically superimposed according to the emotional intensity value (such as the intensity score).

[0072] The adjusted dynamic effect segment is seamlessly integrated with the basic foreground animation layer to form the final foreground animation layer. For example, the basic light effect is used in the calm paragraph, and the high-density particle explosion effect is automatically triggered in the intense paragraph.

[0073] In the embodiments of the present invention, through a dynamic adjustment mechanism driven by emotional parameters, a high degree of synchronization between the foreground animation effect and the audio emotion is achieved. When a fierce emotional passage is detected, the dynamic expressiveness of the foreground animation is automatically enhanced (such as enhancing the light effect and accelerating the movement speed), so that the visual special effects can form a strong association with the emotional fluctuations of the voice (such as anger and tension). The fit between the animation and the plot and the user's immersive experience are significantly improved.

[0074] Most of the foreground animation resources of traditional audiobooks are fixed templates or manually designed paragraph by paragraph, resulting in long resource generation time and difficulty in adapting to diverse storylines. Based on this, the present invention proposes an alternative embodiment.

[0075] Step S10131 also includes the following specific implementation manners.

[0076] Step S101311, obtain the content information of the target audiobook.

[0077] In the embodiments of the present invention, complete structured content information is extracted from the metadata of the audiobook, including but not limited to: keywords of the storyline (such as "space adventure", "mystery case in the ancient castle"), scene description text (such as "the character runs in the storm"), semantic tags of the audio data (such as "suspense", "romance"), etc. The content information is subjected to entity recognition and scene classification through natural language processing technologies (such as the BERT model) to generate standardized feature vectors for subsequent material matching.

[0078] Step S101312, perform a matching operation in a preset material library according to the storyline in the content information to obtain a basic foreground animation layer.

[0079] In the embodiments of the present invention, the storyline feature vectors in the content information are matched with the animation templates in the preset material library, including animation resources classified by scene type (such as "forest", "city") and emotion tags (such as "tense", "warm"), such as particle special effects (rain and snow, flame), dynamic light effects (neon flashing, sunlight penetration), etc. Through similarity calculation (such as cosine similarity) or a rule engine (such as keyword matching), the basic animation template most suitable for the current storyline is selected. For example, if the storyline includes "spacecraft battle", special effect templates such as "laser beam" and "star particles" are matched and combined into a basic foreground animation layer.

[0080] Arrange the matched animation resources along the time axis and adjust the playback rhythm according to the paragraph division of the audio time sequence (such as chapter marks) to form a basic foreground animation layer synchronized with the content of the audiobook.

[0081] In the embodiment of the present invention, the generation efficiency and scene adaptability of the basic foreground animation layer are significantly improved through the intelligent matching mechanism of the preset material library. The optimal animation template is automatically selected based on the plot characteristics, which effectively avoids the inefficiency of traditional manual production of animation resources, and at the same time ensures that the dynamic elements are highly consistent with the plot logic.

[0082] Traditional character extraction technologies (such as color threshold segmentation and manual cutout) are prone to edge jaggedness and residual background noise in semi-transparent areas (such as background color mixed between hair strands) when processing complex backgrounds, resulting in hard edges and unnatural fusion of the generated animated characters with the background. Based on this, the present invention proposes an optional embodiment.

[0083] Step S1012 also includes the following specific implementation methods.

[0084] Step S10121, extracting the large model based on the character, extracting the initial character layer image in the static image.

[0085] In an embodiment of the present invention, the character extraction model refers to an image segmentation model based on deep learning (such as DeepLabv3, U-Net), which can accurately identify and segment the main area of ​​the character from the static picture by training a large amount of labeled data, and generate a character layer picture with a transparent background (Alpha channel).

[0086] The static image is input into the pre-trained character extraction model. The model analyzes the image features through a convolutional neural network and outputs a character mask. The mask marks the precise location of the character's pixels (such as hair and clothing edges) and is superimposed on the original image at the pixel level to generate an initial character layer image (retaining the main body of the character and making the background area transparent). For example, for a picture of a character standing in front of a complex bookshelf, the model can accurately segment the character's outline and eliminate the interference of the bookshelf texture.

[0087] Step S10122, texture mapping is performed on the initial character layer image according to OpenGL to obtain the character layer image.

[0088] In the embodiment of the present invention, texture mapping refers to a rendering technology in OpenGL, which fits a two-dimensional image (texture) to the surface of a three-dimensional model.

[0089] In this step, the segmented character layer image is converted into a texture resource suitable for animation rendering, including but not limited to adjusting the resolution, optimizing edge anti-aliasing and other processing. The initial character layer image is loaded as an OpenGL texture object, and it is bound to a preset animation mesh model (such as a 2D plane or a simple 3D model bound to the character's skeleton) through texture coordinate mapping. At the same time, through the blending and multisampling functions of OpenGL, the character's edges are anti-aliased to eliminate the jagged defects that may be caused by segmentation. For example, the transparent edges of the character's hair are feathered so that they blend naturally with the background when the subsequent animation is superimposed.

[0090] In the embodiment of the present invention, the accuracy and rendering quality of the character layer image are significantly improved by combining the character extraction large model with OpenGL texture mapping. The deep learning model solves the problem of residual character edges in complex backgrounds caused by traditional threshold segmentation, while OpenGL's texture optimization processing ensures that the segmented character layer has smooth edges and no pixel defects in animation rendering. For example, in scenes where the character and background colors are similar, the model can still accurately extract hair details and eliminate aliasing through texture mapping, so that the dynamic character animation and the background layer are superimposed to present a natural transition, which improves the authenticity and visual immersion of the animation.

[0091] Traditional audiobook animations are rendered using a single layer, which results in dynamic elements being mixed in the same layer and cannot be adjusted or reused independently. For example, modifying a character's action requires re-rendering the entire screen, which is inefficient and prone to visual errors. Based on this, the present invention proposes an optional embodiment.

[0092] Step S103 also includes the following specific implementation methods.

[0093] Step S1031, matching the basic image layer of the audio book in the preset material library.

[0094] In the implementation manner of the present invention, the basic image layer refers to a static background image used to carry dynamic animation layers (such as characters, foreground special effects). Its content is related to the audiobook scene, but it does not have dynamic effects itself. The preset material library stores a database of basic image resources, which are classified by scene tags (such as "forest" and "city") and color style (such as "dark" and "bright"), and support keyword retrieval and feature matching.

[0095] In this step, the content information of the audiobook is parsed, and a suitable basic image layer is selected from the preset material library through a semantic matching algorithm. For example, if the keyword "future city" is entered, a static background image containing high-rise buildings and aircraft is matched as the basic layer.

[0096] Step S1032: Sequentially stack the basic image layer, background animation layer, character animation layer, and foreground animation layer to obtain the animation effect corresponding to the target audiobook.

[0097] In an embodiment of the present invention, stacking refers to fusing multiple layers into a single image in sequence during graphics rendering. OpenGL controls the layer transparency and stacking effect (such as multiply, screen blend) through the blending mode.

[0098] Specifically, stack them in the order of "basic image layer, background animation layer, character animation layer, foreground animation layer". During the stacking process, the rendering order and transparency parameters of each layer are dynamically controlled by the content information. For example, in the "character invisibility" scenario, temporarily reduce the transparency of the character animation layer to make it semi-transparent and blend into the background.

[0099] In an embodiment of the present invention, through the hierarchical stacking rendering mechanism, efficient synthesis of the animation effect and optimization of the visual hierarchy are achieved. The basic image layer provides the static framework of the scene, and the background, character, and foreground animation layers sequentially stack dynamic elements, making each layer independently controllable. For example, in the "explosion scene", the basic layer maintains the building structure, the background layer renders the smoke diffusion, the character layer shows the running action, and the foreground layer stacks the flame special effect. After the four layers are stacked, a rich three-dimensional visual effect is formed. Significantly improves the flexibility and rendering efficiency of animation production.

[0100] As Figure 5 shown, it is a schematic diagram of a terminal device provided by an embodiment of the present invention. The terminal device 500 may include: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501, such as an audiobook character surround animation effect generation program. When the processor 501 executes the computer program 503, the steps in the above-mentioned various embodiments of the audiobook character surround animation effect generation are implemented.

[0101] The computer program can be divided into one or more modules / units. One or more modules / units are stored in the memory 502 and executed by the processor 501 to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0102] The terminal device may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art can understand that Figure 5 this is only an example of the terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.

[0103] The so-called processor 501 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0104] The memory 502 may be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device. The memory 502 may also be an external storage device of the terminal device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the terminal device. Further, the memory 502 may also include both the internal storage unit and the external storage device of the terminal device. The memory 502 is used to store computer programs and other programs and data required by the terminal device. The memory 502 may also be used to temporarily store data that has been output or is to be output.

[0105] It should be noted that for the sake of convenience and brevity of description, the structure of the above terminal device may also refer to the specific description of the structure in the method embodiments, which will not be elaborated here.

[0106] The embodiment of the present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method for generating an audiobook character surround animation effect can be implemented.

[0107] The embodiment of the present invention provides a computer program product, and when the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above method for generating an audiobook character surround animation effect when executed.

[0108] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0109] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0110] In the embodiments provided by the present invention, it should be understood that the disclosed terminal devices and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. Another point is that the couplings or direct couplings or communication connections shown or discussed between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0111] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0113] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0114] The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A method for generating a surround animation effect of an audio book character, characterized in that: include: Generate a character animation layer and a foreground animation layer according to content information of a target audio book, wherein the content information includes a storyline, audio data, and an audio timing of the audio data; Generate a background animation layer according to the foreground animation layer and the content information; The foreground animation layer, the character animation layer and the background animation layer are synthesized to obtain an animation effect corresponding to the target audio book.

2. The method for generating surround animation effects of audio book characters according to claim 1, characterized in that: After the step of synthesizing the foreground animation layer, the character animation layer and the background animation layer to obtain the animation effect corresponding to the target audio book, the method further includes: Real-time detection of interactive commands initiated by users; If the interaction instruction is received, the animation effect is adjusted in response to the interaction instruction.

3. The method for generating the surround animation effect of audio book characters as claimed in claim 2, characterized in that: If the interaction instruction is received, the step of adjusting the animation effect in response to the interaction instruction includes: If the interaction instruction is received, matching rendering parameters corresponding to the interaction instruction, the rendering parameters including a first rendering sub-parameter corresponding to the character animation layer, a second rendering sub-parameter corresponding to the foreground animation layer and / or a third rendering sub-parameter corresponding to the background animation layer; The animation effect is adjusted according to the rendering parameters.

4. The method for generating the surround animation effect of audio book characters according to claim 1, characterized in that: The step of generating the character animation layer and the foreground animation layer according to the content information of the target audio book comprises: Get a static image; Extracting a character layer image from the static image; The character animation layer is generated according to the character layer image and the content information, and the foreground animation layer is generated according to the content information of the target audio book.

5. The method for generating the surround animation effect of audio book characters according to claim 4, characterized in that: The step of generating the foreground animation layer according to the content information of the target audio book comprises: Generate a basic foreground animation layer according to the content information of the target audio book, and parse the emotion parameters in the audio time sequence; When it is determined according to the emotion parameter that there is an intense emotion segment, the dynamic effect segment corresponding to the intense emotion segment in the basic foreground animation layer is adjusted to obtain the foreground animation layer.

6. The method for generating the surround animation effect of audio book characters according to claim 5, characterized in that: The step of generating a basic foreground animation layer according to the content information of the target audio book comprises: Acquire the content information of the target audio book; According to the storyline in the content information, a matching operation is performed in a preset material library to obtain the basic foreground animation layer.

7. The method for generating the surround animation effect of audio book characters according to claim 4, characterized in that: The step of extracting the character layer image in the static image comprises: Extracting the initial character layer image from the static image based on the character extraction model; The initial character layer image is texture mapped according to OpenGL to obtain the character layer image.

8. The method for generating surround animation effects of audio book characters according to claim 1, characterized in that: The step of synthesizing the foreground animation layer, the character animation layer and the background animation layer to obtain the animation effect corresponding to the target audio book comprises: Matching the basic image layer of the audio book in the preset material library; The basic image layer, the background animation layer, the character animation layer, and the foreground animation layer are sequentially superimposed to obtain the animation effect corresponding to the target audio book.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for generating surround animation effects of audio book characters as described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating surround animation effects of audio book characters as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Animation production system and method, storage medium, and program product

    CN111598983A

  • Animation generation method and system, medium and electronic terminal

    CN113744369A

  • Audio animation methods and apparatus

    US20130332167A1