Dynamic cartoon rendering method, system and equipment based on music driving and medium

Through deep learning technology, music features are extracted and sentiment analysis is performed, and the rendering effect of comics image is dynamically adjusted, which solves the problem of static singleness of the existing comic reading experience, realizes music-driven dynamic rendering, and improves the fun and stickiness of the user's reading experience.

CN120147502APending Publication Date: 2025-06-13SHANGHAI BILIBILI TECH CO LTD
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202510324343.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-13

Smart Images

  • Figure CN120147502A_ABST
    Figure CN120147502A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a dynamic cartoon rendering method, system and device based on music driving and a medium, and the method comprises the steps: carrying out the feature extraction of to-be-matched music data through employing a convolutional neural network, and obtaining music features; performing sentiment analysis according to the music characteristics by using a long-short-term memory network to obtain a sentiment analysis result; and dynamically adjusting the rendering effect of the cartoon image according to the sentiment analysis result. According to the method and the device, the dynamic adjustment of the cartoon image rendering effect is driven in real time through the music data, the interaction between the music data and the cartoon is conveniently increased, immersive cross-medium experience is provided for a user, the cartoon reading experience is greatly enriched, and the interestingness of cartoon reading is effectively increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of Internet technologies, and particularly to a method, system, device, and medium for dynamic rendering of comics based on music driving. Background Art

[0002] Comics are an art form that depicts life, current events, etc. in a simple and exaggerated way, consisting of a group of images organized in the order of the story plot. Existing client-based comic reading is usually triggered according to the user's page-turning operation, etc., and simply displays the images to the user one by one in the original order. However, this comic reading experience is static, single, lacks interest, and reduces the user's reading stickiness. Summary of the Invention

[0003] In view of the above problems, the present application proposes a method, system, device, and medium for dynamic rendering of comics based on music driving to solve the following problems: the existing comic reading experience is static, single, and lacks interest.

[0004] According to one aspect of the embodiments of the present application, a method for dynamic rendering of comics based on music driving is provided, including:

[0005] Using a convolutional neural network to extract features from the music data to be paired to obtain music features;

[0006] Using a long short-term memory network to perform sentiment analysis based on the music features to obtain a sentiment analysis result;

[0007] According to the sentiment analysis result, dynamically adjust the rendering effect of the comic images.

[0008] Further, using a convolutional neural network to extract features from the music data to be paired to obtain music features further includes:

[0009] Obtain the music data to be paired;

[0010] Convert the music data into a spectrogram, and input the spectrogram into the convolutional neural network, and use the convolutional neural network to extract features from the spectrogram to obtain music features.

[0011] Further, using a long short-term memory network to perform sentiment analysis based on the music features to obtain a sentiment analysis result further includes:

[0012] Input the music features into the long short-term memory network, use the long short-term memory network to extract the temporal dependence relationship, analyze the music sentiment change, and obtain the output result of the long short-term memory network;

[0013] Map the output result to the sentiment category through a fully connected layer to obtain a sentiment analysis result.

[0014] Further, according to the sentiment analysis result, dynamically adjusting the rendering effect of the comic image further includes:

[0015] Obtain the comic image, and determine the target matching mode between the music data and the comic image;

[0016] According to the sentiment analysis result and the target matching mode, dynamically determine the rendering parameters corresponding to the comic image, and adjust the rendering effect of the comic image according to the rendering parameters.

[0017] Further, the sentiment analysis result is a sentiment vector;

[0018] According to the sentiment analysis result and the target matching mode, dynamically determining the rendering parameters corresponding to the comic image, and adjusting the rendering effect of the comic image according to the rendering parameters further includes:

[0019] Expand the sentiment vector to the same shape as the comic image as the style factor;

[0020] Perform arithmetic processing on the comic image and the style factor and then input them into the style conversion network. Use the style conversion network to dynamically determine the rendering parameters corresponding to the comic image according to the target matching mode and perform style conversion processing on the comic image according to the rendering parameters to obtain and display the stylized comic image.

[0021] Further, determining the target matching mode between the music data and the comic image further includes:

[0022] In response to the matching mode setting request executed by the user, set the matching mode corresponding to the matching mode setting request as the target matching mode between the music data and the comic image.

[0023] According to another aspect of the embodiments of the present application, a music-driven comic dynamic rendering system is provided, including:

[0024] A music feature extraction module, adapted to extract features from the music data to be paired by using a convolutional neural network to obtain music features;

[0025] A sentiment analysis module, adapted to perform sentiment analysis on the music features by using a long short-term memory network to obtain a sentiment analysis result;

[0026] A comic rendering module, adapted to dynamically adjust the rendering effect of the comic image according to the sentiment analysis result.

[0027] According to yet another aspect of the embodiments of the present application, a computing device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0028] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned music-driven comic dynamic rendering method.

[0029] According to another aspect of the embodiments of the present application, a computer storage medium is provided, in which at least one executable instruction is stored, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned music-driven comic dynamic rendering method.

[0030] According to still another aspect of the embodiments of the present application, a computer program product is provided, including at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned music-driven comic dynamic rendering method.

[0031] According to the technical solution provided by the embodiments of the present application, deep learning technology is introduced into the field of comic dynamic rendering. The music features of music data are extracted by using a convolutional neural network, and sentiment analysis is performed based on the music features by using a long short-term memory network, so that the sentiment analysis result can be obtained conveniently and accurately; the rendering effect of the comic image is dynamically adjusted according to the sentiment analysis result, realizing the dynamic adjustment of the rendering effect of the comic image driven by music data in real time, conveniently increasing the interaction between music data and comics, providing an immersive cross-media experience for users, greatly enriching the comic reading experience, effectively increasing the interest of comic reading, helping to attract users to read comics by matching different music data, and enhancing the user's reading interest and reading stickiness.

[0032] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and understandable, the following specifically illustrates the specific implementation manners of the embodiments of the present application. Brief Description of the Drawings

[0033] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the embodiments of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0034] Figure 1 A flowchart showing the process of a music-driven comic dynamic rendering method according to an embodiment of the present application is shown;

[0035] Figure 2 A flowchart showing the process of a music-driven comic dynamic rendering method according to another embodiment of the present application is shown;

[0036] Figure 3 The structural block diagram of a music-driven comic dynamic rendering system according to an embodiment of the present application is shown;

[0037] Figure 4 The structural schematic diagram of a computing device according to an embodiment of the present application is shown. Detailed implementation manners

[0038] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0039] First, the noun terms related to one or more embodiments of the present application are explained.

[0040] Convolutional Neural Networks (CNN): A type of feedforward neural network that contains convolutional calculations and has a deep structure, and is one of the representative algorithms of deep learning.

[0041] Long Short-Term Memory networks (LSTM): A special type of Recurrent Neural Network (RNN) designed to address the vanishing or exploding gradient problems encountered by traditional RNNs when processing long time series data. LSTM can effectively capture long-term dependencies in time series through a design called the "gating" mechanism, which makes them perform well in dealing with problems in multiple fields such as language modeling, speech recognition, and time series prediction.

[0042] Figure 1 The flow schematic diagram of a music-driven comic dynamic rendering method according to an embodiment of the present application is shown, as Figure 1 shown, the method includes the following steps:

[0043] Step S101, using a convolutional neural network to extract features from the music data to be paired, obtaining music features.

[0044] In the embodiments of the present application, the interaction between music data and comics is increased, and the dynamic adjustment of the rendering effect of comic images is driven in real time through music data, so as to provide users with an immersive cross-media experience, greatly enrich the comic reading experience, and increase the fun of comic reading.

[0045] In order to accurately match the rendering effect of comic images with the rhythm, emotion, etc. of the music data to be paired, the embodiments of the present application introduce deep learning technology into the field of comic dynamic rendering, and use deep learning algorithms to analyze and extract the music features of the music data. Specifically, the music data that needs to be paired with the comic images, that is, the music data to be paired, is obtained; then a convolutional neural network is used to extract features from the music data to be paired, so as to obtain music features, which may include features such as music rhythm, melody, and chords. Among them, the convolutional neural network is a deep learning network, which is particularly suitable for processing data with grid structures, such as images, videos, and audio signals, and can automatically extract features and perform classification or regression tasks.

[0046] Step S102, use a long short-term memory network to perform sentiment analysis based on the music features to obtain a sentiment analysis result.

[0047] Considering that the long short-term memory network can effectively capture temporal dependencies, in the embodiments of the present application, the long short-term memory network is used to analyze the sentiment tendency of the music data. Specifically, the music features obtained in step S101 can be used as the input of the long short-term memory network, and the long short-term memory network is used to analyze the music emotion changes according to the music features, so as to obtain a sentiment analysis result. The sentiment analysis result is used to reflect the sentiment category expressed by the music data, such as cheerful, sad, tense, etc.

[0048] Step S103, dynamically adjust the rendering effect of the comic image according to the sentiment analysis result.

[0049] After obtaining the sentiment analysis result, the rendering effect of the comic image can be adjusted in real time and dynamically according to the sentiment analysis result. For example, the picture color, texture, lines, dynamic effects, etc. of the comic image can be dynamically adjusted according to the sentiment analysis result, and special effects can also be added, which is not specifically limited here. For example, when the music rhythm speeds up, the action scenes in the comic image will become more intense; when the music melody becomes soothing, the picture color of the comic image will also become softer synchronously.

[0050] According to the method for dynamically rendering comics based on music drive provided by the embodiments of the present application, deep learning technology is introduced into the field of comic dynamic rendering. The music features of music data are extracted by using a convolutional neural network, and the long short-term memory network is used to perform sentiment analysis based on the music features, so that the sentiment analysis results can be obtained conveniently and accurately; according to the sentiment analysis results, the rendering effect of the comic image is dynamically adjusted, realizing the dynamic adjustment of the rendering effect of the comic image driven by music data in real time, conveniently increasing the interaction between music data and comics, providing users with an immersive cross-media experience, greatly enriching the comic reading experience, effectively increasing the interest of comic reading, helping to attract users to read comics by matching different music data, and enhancing users' reading interest and reading stickiness.

[0051] Figure 2 The flowchart of the method for dynamically rendering comics based on music drive according to another embodiment of the present application is shown, as Figure 2 shown, the method includes the following steps:

[0052] Step S201, obtain the music data to be matched.

[0053] In an alternative embodiment, a music library containing multiple music data may be pre-set in the client that provides the comic reading function, and the user can select the music data to be matched from the music library according to their own preferences.

[0054] In another alternative embodiment, the user may not select from the music library, but upload the music data to be matched by themselves. For example, the user can download their favorite music data from the network or the like, and then complete the upload of the music data through the client.

[0055] Step S202, convert the music data into a spectrogram, input the spectrogram into a convolutional neural network, and use the convolutional neural network to extract features from the spectrogram to obtain music features.

[0056] To facilitate the analysis of the music features of the music data, the music data can be first converted into a spectrogram by using a spectrogram conversion tool, and the spectrogram can visually represent the change of the frequency components of the music data over time. Specifically, the continuous music data can be segmented into multiple overlapping small segments (i.e., frames) by the spectrogram conversion tool. Usually, the length of each frame is 20 to 40 milliseconds, and there is 50% to 75% overlap between adjacent frames. Moreover, to reduce spectral leakage, each frame signal is multiplied by a window function (such as a Hann window, a Hamming window, etc.); then the Fourier transform is calculated for each frame to obtain the spectrum of that frame; then the spectra of all frames are arranged in chronological order and other processing to obtain the spectrogram.

[0057] The spectrogram may be a Mel Spectrogram, etc. The Mel Spectrogram is a spectrum representation adjusted based on the perceptual characteristics of the human auditory system. Compared with the traditional linear frequency scale spectrogram, the Mel Spectrogram uses the Mel scale, which is a scale based on the nonlinear perception of sound frequency by human hearing. The Mel scale increases the resolution of the low frequency band and reduces the resolution of the high frequency band, which is more in line with the human ear's perception characteristics of sounds of different frequencies.

[0058] The spectrogram is input into the convolutional neural network, and the music features such as rhythm and melody in the spectrogram are extracted through the convolutional layer in the convolutional neural network.

[0059] Step S203, using a long short-term memory network to perform sentiment analysis based on the music features to obtain a sentiment analysis result.

[0060] Among them, the emotional features of music data are modeled using a long short-term memory network, the emotional changes of music are analyzed, and they are mapped to the rendering style of comics. Specifically, the music features can be input into the long short-term memory network, the long short-term memory network is used to extract the temporal dependency, analyze the emotional changes of music, and obtain the output results of the long short-term memory network; then the output results are mapped to the emotional category through the fully connected layer to obtain the emotional analysis results.

[0061] Among them, multiple emotion categories can be pre-divided, such as cheerfulness, sadness, tension, etc., and the output results of the long short-term memory network are mapped to the pre-divided multiple emotion categories through the fully connected layer, so as to accurately determine the emotion category to which the music data belongs and obtain the emotion analysis result. The emotion analysis result can specifically be an emotion vector used to reflect the emotion category to which it belongs. In an embodiment of the present application, the emotion analysis result is obtained based on the analysis of music features, so the emotion analysis result already contains information about the music features, that is, the music features have been presented in the form of emotion analysis results at this stage, so the emotion analysis result can not only reflect its emotion category, but also the rhythm, melody, etc.

[0062] Step S204, acquiring a comic image, and determining a target matching pattern between the music data and the comic image.

[0063] The user selects the comic image he wants to read, and the comic image is obtained in step S204. In addition, in order to meet the diverse experience needs of users, multiple matching modes between music data and comic images are set, and a target matching mode needs to be determined from the multiple matching modes. Among them, the multiple matching modes may include: rhythm matching mode, melody matching mode, emotion matching mode and mixed matching mode.

[0064] Specifically, the rhythm matching mode refers to adjusting the frequency and amplitude of actions and dynamic effects in the comic images according to the beats or rhythm changes of the music (such as the speed of drum beats, etc.). The melody matching mode refers to controlling the movement, transitions, or local transformations of the elements in the comic images according to the ups and downs and pitch changes of the music melody. The emotion matching mode refers to adjusting the tone, brightness, and style of the comic images according to the emotions conveyed by the music (such as joy, sadness, tension, etc.). The hybrid matching mode refers to the matching mode designed by comprehensively considering various music features. According to the rhythm, melody, and emotion information, it adjusts the frequency and amplitude of actions and dynamic effects in the comic images, adjusts the movement, transitions, or local transformations of the elements in the images, and adjusts the tone, brightness, and style of the comic images, etc.

[0065] In an alternative embodiment, it supports users to customize the selection of the matching mode, effectively enhancing the personalized experience. The user can select a matching mode from multiple matching modes as the target matching mode between the music data and the comic images by executing a matching mode setting request. Specifically, in response to the matching mode setting request executed by the user, the matching mode corresponding to the matching mode setting request is set as the target matching mode between the music data and the comic images.

[0066] In another alternative embodiment, if the user does not make a selection of the matching mode, the system can default to one of the matching modes as the target matching mode. For example, the system can default to the hybrid matching mode as the target matching mode.

[0067] Step S205: Dynamically determine the rendering parameters corresponding to the comic images according to the emotion analysis result and the target matching mode, and adjust the rendering effect of the comic images according to the rendering parameters.

[0068] After obtaining the emotion analysis result and the target matching mode, the rendering parameters corresponding to the comic images can be dynamically determined according to the emotion analysis result and the target matching mode. Among them, the rendering parameters may include: color parameters, texture and line parameters, dynamic effect parameters, and special effect parameters, etc. Specifically, the color parameters refer to the parameters used to change the overall atmosphere of the comic images, and may include parameters such as hue, saturation, contrast, and brightness. The texture and line parameters may include parameters such as line thickness, brush stroke style, and texture details, which are used to adapt to the visual style requirements under different music emotions. The dynamic effect parameters refer to the parameters used to achieve dynamic rendering effects, and may include parameters such as action amplitude, movement speed, and transition effects (such as blur, deformation, etc.). The special effect parameters may include parameters such as special filters, light and shadow effects, and particle effects, which can be triggered at key music rhythm points or emotional turning points.

[0069] After determining the rendering parameters corresponding to the comic images, the comic images can be rendered according to the rendering parameters, thereby realizing the dynamic adjustment of the rendering effect of the comic images.

[0070] In practical applications, the dynamic adjustment of the rendering effect of comic images can be conveniently achieved through a style conversion network. Among them, the style conversion network is a convolutional network that has been pre-trained and can quickly perform style conversion on images. Specifically, the result of sentiment analysis can be a sentiment vector. Considering that the shape of the sentiment vector may be different from that of the comic image, for the convenience of processing, the sentiment vector can be extended to the same shape as the comic image as a style factor. Then, after performing arithmetic processing on the comic image and the style factor, they are input into the style conversion network. The style conversion network dynamically determines the rendering parameters corresponding to the comic image according to the target matching pattern and performs style conversion processing on the comic image according to the rendering parameters, obtaining and presenting the stylized comic image, thus realizing the adaptive adjustment of the comic rendering effect based on the sentiment analysis result of music data.

[0071] The following shows a specific implementation method for modularly realizing each process of music feature extraction, sentiment analysis, and comic rendering, and the description is as follows:

[0072] The music feature extraction module can use the Mel spectrogram and a convolutional neural network to extract music features such as rhythm and melody. Specifically, in the music feature extraction module: create a Mel spectrogram converter for converting music data into a Mel spectrogram; and define a convolutional neural network, which includes two convolutional layers (such as Conv2d), a ReLU activation function, and a max pooling layer (such as MaxPool2d) for extracting music features such as rhythm and melody from the Mel spectrogram.

[0073] The sentiment analysis module uses a long short-term memory network (i.e., LSTM) to model the sentiment features of music data and outputs a sentiment vector as the result of sentiment analysis. Specifically, in the sentiment analysis module: create an LSTM and a fully connected layer, input the music features as input into the LSTM. The LSTM extracts the temporal dependence relationship and analyzes the change of music sentiment to obtain an output sequence. Select the output of the last time step of the output sequence as the output result of the LSTM, and map the output result of the LSTM to the sentiment category through the fully connected layer, thus obtaining the sentiment vector, that is, the result of sentiment analysis.

[0074] The comic rendering module dynamically adjusts the style of comic images according to the emotion vector and realizes style conversion through a convolutional network. Specifically, in the comic rendering module: a style conversion network is defined to perform style conversion on comic images. The style conversion network may include two convolutional layers (such as Conv2d), ReLU activation function, and Sigmoid activation function. The image style is adjusted through the convolutional layer, and the Sigmoid activation function ensures that the output value is within the range of [0,1]; and the emotion vector is expanded to the same shape as the comic image as the style factor. For example, the shape of the comic image is (1,3,256,256), representing three RGB channels and a resolution of 256×256; the comic image is multiplied by the style factor and then input into the style conversion network, and the style conversion network performs style conversion processing to obtain the stylized comic image.

[0075] In the process of system integration, the music feature extraction module, emotion analysis module, and comic rendering module are initialized. Music features are extracted from the input music data, emotion analysis is performed based on the extracted music features, and then the style of the comic image is converted according to the emotion analysis result. During the example run, the required music data and comic images are created; instantiation is performed; the music data and comic images are input into the system to obtain the stylized comic image; the shape of the stylized comic image is printed, thus realizing the dynamic adjustment of the rendering effect of the comic image based on music drive.

[0076] According to the music-driven comic dynamic rendering method provided by the embodiments of the present application, the music data and comic images are dynamically combined. The music features of the music data are extracted by using a convolutional neural network, and emotion analysis is performed based on the music features by using a long short-term memory network. The comic dynamic rendering is driven by the emotion analysis result, realizing cross-media interaction and providing users with an immersive comic reading experience; moreover, it can perform real-time rendering, respond to music changes in real time, and adaptively and dynamically adjust the comic rendering effect; in addition, this solution can be well applied to application scenarios such as digital comic platforms and interactive entertainment, supporting users to independently select music data and comic images, and also supporting users to customize the matching mode between music data and comic images, enabling users to participate in comic creation, effectively enhancing the personalized experience and user participation, improving the interest of comic reading, helping to attract users to read comics by matching different music data, and enhancing the user's reading interest and reading stickiness.

[0077] Figure 3 The structural block diagram of a music-driven comic dynamic rendering system according to an embodiment of the present application is shown, as Figure 3 shown, the system includes: a music feature extraction module 310, an emotion analysis module 320, and a comic rendering module 330.

[0078] The music feature extraction module 310 is adapted to: extract features from the music data to be paired using a convolutional neural network to obtain music features;

[0079] The sentiment analysis module 320 is adapted to: perform sentiment analysis based on the music features using a long short-term memory network to obtain a sentiment analysis result.

[0080] The comic rendering module 330 is adapted to: dynamically adjust the rendering effect of the comic image according to the sentiment analysis result.

[0081] Optionally, the music feature extraction module 310 is further adapted to: obtain the music data to be paired; convert the music data into a spectrogram, input the spectrogram into the convolutional neural network, and use the convolutional neural network to extract features from the spectrogram to obtain music features.

[0082] Optionally, the sentiment analysis module 320 is further adapted to: input the music features into the long short-term memory network, use the long short-term memory network to extract temporal dependencies, analyze the changes in music sentiment, and obtain the output result of the long short-term memory network; map the output result to sentiment categories through a fully connected layer to obtain a sentiment analysis result.

[0083] Optionally, the comic rendering module 330 is further adapted to: obtain the comic image, determine the target matching pattern between the music data and the comic image; dynamically determine the rendering parameters corresponding to the comic image according to the sentiment analysis result and the target matching pattern, and adjust the rendering effect of the comic image according to the rendering parameters.

[0084] Optionally, the sentiment analysis result is a sentiment vector; the comic rendering module 330 is further adapted to: expand the sentiment vector to the same shape as the comic image as a style factor; perform arithmetic processing on the comic image and the style factor and then input them into the style conversion network, use the style conversion network to dynamically determine the rendering parameters corresponding to the comic image according to the target matching pattern and perform style conversion processing on the comic image according to the rendering parameters to obtain and display the stylized comic image.

[0085] Optionally, the system further includes: a user interaction module 340, adapted to, in response to a matching pattern setting request executed by the user, set the matching pattern corresponding to the matching pattern setting request as the target matching pattern between the music data and the comic image.

[0086] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated here.

[0087] According to the music-driven comic dynamic rendering system provided by the embodiments of the present application, music data is dynamically combined with comic images. The convolutional neural network is used to extract the music features of the music data, and the long short-term memory network is used to perform sentiment analysis based on the music features. The comic dynamic rendering is driven by the sentiment analysis results, realizing cross-media interaction and providing users with an immersive comic reading experience. Moreover, it can perform real-time rendering, respond to music changes in real time, and adaptively adjust the comic rendering effect dynamically. In addition, this solution can be well applied to application scenarios such as digital comic platforms and interactive entertainment, supporting users to select music data and comic images by themselves, and also supporting users to customize the matching mode between music data and comic images, enabling users to participate in comic creation, effectively enhancing the personalized experience and user participation, improving the interest of comic reading, helping to attract users to read comics by matching different music data, and enhancing the user's reading interest and reading stickiness.

[0088] The embodiments of the present application provide a non-volatile computer storage medium. The computer storage medium stores at least one executable instruction or computer program, and the executable instruction or computer program can enable the processor to perform the operations corresponding to the music-driven comic dynamic rendering method in any of the above method embodiments.

[0089] The embodiments of the present application provide a computer program product. The computer program product includes at least one executable instruction or computer program, and the executable instruction or computer program can enable the processor to perform the operations corresponding to the music-driven comic dynamic rendering method in any of the above method embodiments.

[0090] Figure 4 The structural schematic diagram of a computing device according to an embodiment of the present application is shown. The specific implementation of the computing device is not limited in the specific embodiments of the present application.

[0091] As Figure 4 shown, the computing device may include: a processor 402, a communications interface 404, a memory 406, and a communication bus 408.

[0092] Among them: The processor 402, the communications interface 404, and the memory 406 communicate with each other through the communication bus 408. The communications interface 404 is used to communicate with network elements of other devices such as clients or other servers. The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the embodiments of the music-driven comic dynamic rendering method for the computing device described above.

[0093] Specifically, the program 410 may include program code, which includes computer operation instructions.

[0094] The processor 402 may be a central processing unit (CPU), or a specific application integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0095] The memory 406 is used to store the program 410. The memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0096] The program 410 is specifically configured to cause the processor 402 to execute the music-driven comic dynamic rendering method in any of the above method embodiments. For the specific implementation of each step in the program 410, reference may be made to the corresponding steps and units in the above music-driven comic dynamic rendering embodiments, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated herein.

[0097] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. In addition, the embodiments of the present application are not directed to any specific programming language. It should be understood that the content of the embodiments of the present application described herein can be implemented using various programming languages, and the description of the specific language above is for disclosing the best mode of the embodiments of the present application.

[0098] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0099] Similarly, it should be understood that, for the purpose of streamlining the present disclosure and facilitating the understanding of one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed embodiments of the present application require more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the embodiments of the present application.

[0100] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0101] In addition, those skilled in the art will be able to understand that although some of the embodiments described herein include certain features included in other embodiments but not other features, the combination of the features of different embodiments means that it is within the scope of the embodiments of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0102] Each component embodiment of the embodiments of the present application may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) may be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present application. The embodiments of the present application may also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for performing part or all of the methods described herein. Such a program implementing the embodiments of the present application may be stored on a computer-readable medium, or may be in the form of one or more signals. Such signals may be downloaded from an Internet website, or provided on a carrier signal, or in any other form.

[0103] It should be noted that the above embodiments illustrate the embodiments of the present application rather than limit the embodiments of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The embodiments of the present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

Claims

1. A comic dynamic rendering method based on music driving, comprising: Use convolutional neural network to extract features of the music data to be matched and obtain music features; Using a long short-term memory network to perform sentiment analysis based on the music features, and obtaining a sentiment analysis result; According to the sentiment analysis result, the rendering effect of the comic image is dynamically adjusted.

2. According to the method of claim 1, the step of extracting features from the music data to be matched using a convolutional neural network to obtain music features further comprises: Obtaining music data to be matched; The music data is converted into a spectrogram, and the spectrogram is input into the convolutional neural network. The convolutional neural network is used to extract features from the spectrogram to obtain the music features.

3. According to the method of claim 1, the step of performing sentiment analysis based on the music features using a long short-term memory network to obtain a sentiment analysis result further comprises: Inputting the music features into the long short-term memory network, using the long short-term memory network to extract temporal dependencies, analyzing music emotion changes, and obtaining an output result of the long short-term memory network; The output result is mapped to the sentiment category through a fully connected layer to obtain the sentiment analysis result.

4. According to the method according to any one of claims 1 to 3, dynamically adjusting the rendering effect of the comic image according to the sentiment analysis result further comprises: Acquire the comic image, and determine a target matching pattern between the music data and the comic image; According to the sentiment analysis result and the target matching pattern, rendering parameters corresponding to the comic image are dynamically determined, and the rendering effect of the comic image is adjusted according to the rendering parameters.

5. The method according to claim 4, wherein the sentiment analysis result is a sentiment vector; The dynamically determining the rendering parameters corresponding to the comic image according to the sentiment analysis result and the target matching mode, and adjusting the rendering effect of the comic image according to the rendering parameters further comprises: Expanding the emotion vector into the same shape as the comic image as a style factor; The comic image and the style factor are processed by operation and then input into a style transfer network. The style transfer network is used to dynamically determine the rendering parameters corresponding to the comic image according to the target matching pattern and perform style transfer processing on the comic image according to the rendering parameters to obtain and display the stylized comic image.

6. The method according to claim 4 or 5, wherein the determining a target matching pattern between the music data and the comic image further comprises: In response to a matching mode setting request performed by a user, a matching mode corresponding to the matching mode setting request is set as a target matching mode between the music data and the comic image.

7. A music-driven comic dynamic rendering system, comprising: A music feature extraction module is suitable for extracting features from the music data to be matched using a convolutional neural network to obtain music features; A sentiment analysis module, adapted to perform sentiment analysis according to the music features using a long short-term memory network to obtain a sentiment analysis result; The comic rendering module is adapted to dynamically adjust the rendering effect of the comic image according to the emotion analysis result.

8. A computing device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the music-driven comic dynamic rendering method as described in any one of claims 1-6.

9. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and the executable instruction enables a processor to execute operations corresponding to the music-driven comic dynamic rendering method as described in any one of claims 1 to 6.

10. A computer program product, comprising at least one executable instruction, wherein the executable instruction enables a processor to execute operations corresponding to the music-driven comic dynamic rendering method according to any one of claims 1 to 6.