Music playing interface display method and device, equipment and storage medium
By displaying natural and physical effects of audio emotions in the music playback interface, the problem that background animation effects in the prior art cannot feedback song emotions is solved, and the intuitive display of audio emotions and the improvement of interface effects is achieved.
Patent Information
- Application Number
- CN202510513394.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-08
AI Technical Summary
The existing music player interface background animation effects cannot feedback the emotions of the song, resulting in disconnection between interface interaction and the emotional perception of the song.
The first background animation of the first audio is displayed in the music playback interface, and the audio emotions are visualized and displayed through the special effects of natural phenomena and the motion characteristics of natural particles, including the first natural effects and the first physical effects.
The background screen effect of the music playback interface is in line with the audio emotions, and users can intuitively feel the audio emotions, improving the interface display effect and emotional impact ability.
Smart Images

Figure CN120448012A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, equipment, and storage medium for displaying a music playback interface. Background Art
[0002] Current music players are provided with an interface interaction, which displays different images on the interface to show the rhythm and style of the music.
[0003] In the related art, background animation is displayed on the display interface of the music player. The background animation can dynamically change with the time of music playback. The background animation displays spectrum elements that dynamically change with the rhythm of the music. The user can independently select the geographical location of the music player, and the background animation can be displayed as a local weather image or a local city landscape image accordingly.
[0004] However, the background animation presented in the above method cannot reflect the emotions of the song itself, resulting in a disconnect between the interface interaction and the emotional perception of the song. Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for displaying a music playback interface. The technical solutions provided by the embodiments of the present application are as follows:
[0006] According to one aspect of an embodiment of the present application, a method for displaying a music playback interface is provided, the method comprising:
[0007] Display the music playback interface of the first audio;
[0008] During the playback of the first audio, a first background animation of the first audio is displayed in the music playback interface;
[0009] Among them, the first background animation effect includes a first natural special effect corresponding to the first audio, the first natural special effect includes a first physical special effect corresponding to the first audio, the first natural special effect is used to visualize the audio emotion contained in the first audio through natural phenomena, and the first physical special effect is used to indicate the first motion characteristics of natural particles in the first natural special effect.
[0010] According to one aspect of an embodiment of the present application, a device for displaying a music playback interface is provided, the device comprising:
[0011] An interface display module, configured to display a music playback interface of the first audio;
[0012] A motion effect display module, configured to display a first background motion effect of the first audio in the music playback interface during playback of the first audio;
[0013] Among them, the first background animation effect includes a first natural special effect corresponding to the first audio, the first natural special effect includes a first physical special effect corresponding to the first audio, the first natural special effect is used to visualize the audio emotion contained in the first audio through natural phenomena, and the first physical special effect is used to indicate the first motion characteristics of natural particles in the first natural special effect.
[0014] According to one aspect of an embodiment of the present application, a computer device is provided, which includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned method for displaying the music playback interface.
[0015] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned method for displaying the music playback interface.
[0016] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned method for displaying the music playback interface.
[0017] The technical solutions provided in the embodiments of the present application can bring the following beneficial effects:
[0018] By displaying the first background motion effect of the first audio in the music playback interface during the playback of the first audio, the first background motion effect is used to feedback the audio emotion contained in the first audio. Compared with the related technology in which the background motion effect is disconnected from the emotional perception of the audio, the technical solution provided by the present application displays the first natural special effect, based on the phenomenon characteristics of natural phenomena and the motion characteristics of natural particles in natural phenomena, to visually display the audio emotion contained in the first audio, so that the picture effect of the background picture of the music playback interface can be consistent with the audio emotion contained in the first audio, breaking through the traditional interface interaction method, allowing users to intuitively feel the audio emotion contained in the first audio from the music playback interface, improving the picture display effect of the music playback interface, enhancing the emotional influence ability of the music playback interface, and helping to improve the playback effect of the first audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic diagram of a computer system provided by one embodiment of the present application;
[0020] Figure 2 This is a flowchart of a method for displaying a music playback interface provided by one embodiment of the present application;
[0021] Figure 3This is a schematic diagram of a first background animation effect in a music playback interface provided by an embodiment of the present application;
[0022] Figure 4 This is a flowchart of a method for generating a first background dynamic effect provided by an embodiment of the present application;
[0023] Figure 5 This is a schematic diagram of a process for generating a first background animation effect provided by an embodiment of the present application;
[0024] Figure 6 This is a block diagram of a display device for a music playback interface provided by one embodiment of the present application;
[0025] Figure 7 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0027] Please refer to Figure 1 , which shows a schematic diagram of a computer system provided by an embodiment of the present application. The computer system may include: a terminal device 10 and a server 20.
[0028] There can be one or more terminal devices 10. The terminal device 10 can be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, game console, e-book reader, multimedia player, wearable device, intelligent voice interaction device, smart home appliance, vehicle terminal, aircraft, etc.
[0029] The client of the music player program can be installed in the terminal device 10, and the music player program provides the function of playing audio. Optionally, the music player program can be an application program that needs to be downloaded and installed, or it can be an application program that can be used immediately, and this application does not limit this.
[0030] The server 20 is used to provide background services for the client of the music player program installed and running in the terminal device 10. For example, the server 20 can be the background server of the above-mentioned music player program. The server 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but is not limited to these. Optionally, the server 20 provides background services for the music player programs in multiple terminal devices 10 at the same time. The terminal device 10 and the server 20 can communicate with each other through the network.
[0031] In an embodiment of the present application, a user triggers playback of a first audio track, and a music playback interface for the first audio track is displayed on a client of a music playback program. During playback of the first audio track, a first background motion effect for the first audio track is displayed on the music playback interface based on the audio emotion contained in the first audio track. The first background motion effect includes a first natural special effect corresponding to the first audio track, which is used to visually display the audio emotion contained in the first audio track through natural phenomena. The first natural special effect includes a first physical special effect corresponding to the first audio track, which is used to indicate a first motion characteristic of natural particles in the first natural special effect.
[0032] Please refer to Figure 2 , which shows a flow chart of a method for displaying a music playback interface provided by an embodiment of the present application. The execution subject of each step of the method may be a terminal device. The method may include at least one of the following steps 210 to 220:
[0033] Step 210: Display the music playing interface of the first audio.
[0034] The first audio can be any type of audio, and this application does not limit it. For example, the first audio includes but is not limited to song audio, pure music audio, opera audio, audio book audio, etc. The first audio can be audio stored locally or stored in network data.
[0035] When a user triggers a first audio track in a music player program, a music player interface for the first audio track is displayed on the display interface of the music player program client. The music player interface is an interface for playing the first audio track. The music player interface displays audio information and playback information for the first audio track. The audio information for the first audio track is used to introduce the first audio track. For example, if the first audio track is a song, the audio information for the first audio track includes the song title, the artist name, lyrics information, album name, cover image, and audio production staff of the first audio track. If the first audio track is pure music, the audio information for the first audio track includes the name of the pure music track, the instrument played by the first audio track, the performer of the first audio track, and the cover image. The playback information for the first audio track is used to indicate the playback status and playback progress of the first audio track. The playback progress of the first audio track can be reflected by a playback progress bar for the first audio track. If the first audio track includes a vocal track, for example, if the first audio track is a song, the playback progress of the first audio track can also be reflected by the lyrics information of the currently playing first audio track.
[0036] Step 220: During the playback of the first audio, a first background animation of the first audio is displayed in the music playback interface.
[0037] Among them, the first background animation effect includes the first natural special effect corresponding to the first audio, the first natural special effect includes the first physical special effect corresponding to the first audio, the first natural special effect is used to visualize the audio emotions contained in the first audio through natural phenomena, and the first physical special effect is used to indicate the first motion characteristics of natural particles in the first natural special effect.
[0038] The first background motion effect is a dynamic background effect displayed in the background screen layer of the music playback interface. During the playback of the first audio, the display content in the background screen layer of the music playback interface changes dynamically. Optionally, during the pause of the first audio, the display content in the background screen layer of the music playback interface can be dynamically changing or fixed. In other words, during the pause of the first audio, the first background motion effect of the first audio can be displayed in the music playback interface, or the first background image of the first audio can be displayed in the music playback interface. The first background image is a static background image of the first background motion effect when the first audio is paused.
[0039] The first background animation is generated based on the audio emotion contained in the first audio. The first background animation is used to reflect the audio emotion contained in the first audio, so that the background image of the music playback interface can match the audio emotion contained in the first audio. The specific generation process of the first background animation can be referred to in the following embodiment and is not described here.
[0040] The first background animation includes a first natural effect corresponding to the first audio. This natural effect is used to visualize the audio emotion contained in the first audio through natural phenomena. The audio emotions contained in the first audio include, but are not limited to, sadness, loneliness, joy, soothing, peaceful, and exciting emotions.
[0041] In some embodiments, the first natural special effect is a natural special effect in one of the following natural fields: weather morphology field, geological morphology field, astronomical morphology field, and biological morphology field.
[0042] The natural special effects of each natural field can be divided into multiple special effect types, and different special effect types are used to indicate different natural phenomena in the natural field. For example, the special effect types of the natural special effects in the weather morphology field include but are not limited to sunny weather, rainy weather, snowy weather, breezy weather, thunder and lightning weather, hurricane weather and other weather types; the special effect types of the natural special effects in the geological morphology field include but are not limited to volcanic eruptions, glacial movement, earthquake phenomena, wind erosion phenomena, water erosion phenomena and other types of geological movement; the special effect types of the natural special effects in the astronomical morphology field include but are not limited to astronomical phenomena such as the Big Bang, star track movement, solar eclipse phenomenon, lunar eclipse phenomenon, meteor shower, tidal phenomenon and other astronomical phenomena; the biological morphology field includes but is not limited to biological phenomena such as fish movement, jellyfish swimming, geese flying south, reindeer migration, falling leaves, and withering flowers.
[0043] The first natural special effect is a special effect type that matches the audio emotion contained in the first audio. For example, if the audio emotion contained in the first audio is sadness, then when the first natural special effect is a natural special effect in the field of weather morphology, the first natural special effect may be a rainy weather special effect corresponding to sadness; when the first natural special effect is a natural special effect in the field of geological morphology, the first natural special effect may be a geological special effect of glacier movement corresponding to sadness; when the first natural special effect is a natural special effect in the field of astronomical morphology, the first natural special effect may be an astronomical special effect of star track movement corresponding to sadness; when the first natural special effect is a natural special effect in the field of biological morphology, the first natural special effect may be a biological special effect of jellyfish swimming corresponding to sadness.
[0044] Each special effect type in each natural field can include multiple natural special effects. The different natural special effects under each special effect type are used to indicate the different motion characteristics of natural particles in the natural phenomenon under that special effect type, that is, the different natural special effects under each special effect type are used to indicate different physical special effects. For example, the rain type weather special effect can include natural special effects such as light rain weather effect, moderate rain weather effect, and heavy rain weather effect. The different natural special effects under the rain type weather special effect are used to indicate the different falling speeds of raindrops in the rain special effect. The falling speed of raindrops in the light rain weather effect is smaller than the falling speed of raindrops in the moderate rain weather effect, and the falling speed of raindrops in the moderate rain weather effect is smaller than the falling speed of raindrops in the heavy rain weather effect.
[0045] The first natural special effect includes a first physical special effect corresponding to the first audio, and the first physical special effect corresponding to the first audio is used to indicate the first motion characteristics of the natural particles in the first natural special effect. The natural particles in the natural special effect are natural particles existing in the natural phenomenon corresponding to the natural special effect, and the particle type of the natural particles in the first natural special effect is determined by the natural phenomenon corresponding to the first natural special effect. For example, if the first natural special effect is a rain special effect in the field of weather morphology, the natural particles in the first natural special effect may refer to raindrops; if the first natural special effect is a snow special effect in the field of weather morphology, the natural particles in the first natural special effect may refer to snowflakes; if the first natural special effect is a fog special effect in the field of weather morphology, the natural particles in the first natural special effect may refer to fog; if the first natural special effect is a lightning special effect in the field of weather morphology, the natural particles in the first natural special effect may refer to lightning.
[0046] The motion characteristics of natural particles can be one of the following motion characteristics: the motion speed of natural particles, the frequency of occurrence of natural particles, the appearance time of natural particles, and the concentration of natural particles. The motion characteristics of natural particles are related to the properties of natural particles. For example, if the natural particles in the first natural special effect are raindrops, the first motion characteristics of natural particles are used to indicate the falling speed of raindrops; if the natural particles in the first natural special effect are rainbows, the first motion characteristics of natural particles are used to indicate the appearance time of rainbows; if the natural particles in the first natural special effect are fog, the first motion characteristics of natural particles are used to indicate the concentration of fog; if the natural particles in the first natural special effect are lightning, the first motion characteristics of natural particles are used to indicate the appearance frequency of lightning.
[0047] The first motion characteristic of the natural particles indicated in the first physical special effect is determined by the concentration of the emotion contained in the first audio. For example, if the audio emotion contained in the first audio is stronger, the movement speed of the natural particles indicated by the first motion characteristic is faster, or the frequency of occurrence of the natural particles indicated by the first motion characteristic is higher, or the appearance time of the natural particles indicated by the first motion characteristic is longer, or the concentration of the natural particles indicated by the first motion characteristic is higher. Exemplarily, if the audio emotion contained in the first audio is sadness and the first natural special effect is a rain effect, the lower the concentration of sadness contained in the first audio, the slower the falling speed of the raindrops indicated by the first physical special effect, and the higher the concentration of sadness contained in the first audio, the faster the falling speed of the raindrops indicated by the first physical special effect.
[0048] By setting multiple natural fields for the first natural special effect, the first natural special effect is provided with the possibility of natural special effects in multiple natural fields, enriching the music playback interface and helping to improve the interface display effect of the music playback interface.
[0049] Figure 3 A schematic diagram of the first background dynamic effect in the music playback interface is shown. The first audio is a song audio. The music playback interface 30 displays the audio information of the first audio and the first background dynamic effect of the first audio. The audio information of the first audio includes the song title and singer name 34 of the first audio, as well as the lyrics information 33 of the first audio. The first background dynamic effect of the first audio includes the first natural special effect 31 corresponding to the first audio. Figure 3 The first natural special effect 31 is a rain special effect. The first natural special effect 31 includes a first physical special effect corresponding to the first audio. The first physical special effect is used to indicate a first motion feature of a natural particle 32 in the first natural special effect 31. The natural particle 32 is a raindrop.
[0050] The technical solution provided in the embodiment of the present application displays the first background motion effect of the first audio in the music playback interface during the playback of the first audio, and the first background motion effect is used to feedback the audio emotion contained in the first audio. Compared with the background motion effect in the related art that is disconnected from the emotional perception of the audio, the technical solution provided in the present application displays the first natural special effect, based on the phenomenon characteristics of natural phenomena and the motion characteristics of natural particles in natural phenomena, to visually display the audio emotion contained in the first audio, so that the picture effect of the background picture of the music playback interface can be consistent with the audio emotion contained in the first audio, breaking through the traditional interface interaction method, allowing users to intuitively feel the audio emotion contained in the first audio from the music playback interface, improving the picture display effect of the music playback interface, enhancing the emotional influence ability of the music playback interface, and helping to improve the playback effect of the first audio.
[0051] In some embodiments, in response to a sliding operation on the first natural special effect, natural particles in the first natural special effect are displayed in the first background animation effect to move in a sliding direction corresponding to the sliding operation.
[0052] The sliding operation for the first natural special effect is an operation performed by the user to slide the natural particles in the first natural special effect. For example, the sliding operation for the first natural special effect can be an operation in which the user presses a finger at a display position of the natural particles in the first natural special effect and then slides the display screen.
[0053] When a user performs a swipe operation on the first nature effect, the natural particles in the first nature effect will move in the direction of the swipe operation. In response to canceling the swipe operation on the first nature effect, the natural particles in the first nature effect displayed in the first background animation will move in the original swipe direction. In other words, when the user no longer presses the display position of the natural particles and swipes the display screen, the natural particles in the first nature effect will continue to move in the original direction.
[0054] For example, if the first natural special effect is a rain effect, and the sliding direction corresponding to the sliding operation performed by the user is from left to right, the raindrops in the rain effect will move from left to right following the sliding operation. After the user releases his finger, the raindrops in the rain effect will fall downward in the original direction of movement.
[0055] In some embodiments, in response to a click operation on the first natural special effect, a decelerated motion of natural particles in the first natural special effect is displayed in the first background animation effect.
[0056] The click operation for the first natural special effect is an operation performed by the user to click on the natural particles in the first natural special effect. For example, the click operation for the first natural special effect can be an operation in which the user's finger clicks on the display position of the natural particles in the first natural special effect.
[0057] After the user clicks on the first nature effect, the natural particles in the first nature effect will decelerate in their original direction of motion. In response to another click on the first nature effect, the natural particles in the first nature effect displayed in the first background animation will move at their original speed. That is, after the user clicks on the natural particles in the first nature effect again, the natural particles in the first nature effect will resume their original speed.
[0058] For example, if the first natural special effect is a rain special effect, after the user clicks on the raindrops in the rain special effect, the raindrops in the rain special effect slow down and fall. After the user clicks on the raindrops in the rain special effect again, the raindrops in the rain special effect fall at the original speed.
[0059] In some embodiments, in response to a long press operation on the first natural special effect, the accelerated motion of natural particles in the first natural special effect is displayed in the first background animation.
[0060] The long press operation for the first natural special effect is an operation performed by the user to long press the natural particles in the first natural special effect. For example, the click operation for the first natural special effect can be an operation in which the user long presses the display position of the natural particles in the first natural special effect.
[0061] When the user performs a long-press operation on the first nature effect, the natural particles in the first nature effect accelerate in their original direction of motion. In response to canceling the long-press operation on the first nature effect, the natural particles in the first nature effect displayed in the first background animation move at their original speed. That is, after the user releases the long-press operation on the natural particles in the first nature effect, the natural particles in the first nature effect return to their original speed.
[0062] For example, if the first natural special effect is a rain special effect, when the user long presses the raindrops in the rain special effect, the raindrops in the rain special effect will fall faster. After the user releases his finger, the raindrops in the rain special effect will fall at the original speed.
[0063] In some embodiments, in response to a double-click operation on the first natural special effect, the natural particles in the first natural special effect are displayed in the first background animation effect to stop moving.
[0064] The double-click operation for the first natural effect is an operation performed by the user to double-click the natural particles in the first natural effect. For example, the double-click operation for the first natural effect can be an operation in which the user's finger double-clicks the display position of the natural particles in the first natural effect.
[0065] After the user double-clicks the first nature effect, the natural particles in the first nature effect stop moving. In response to another double-click on the first nature effect, the natural particles in the first nature effect continue to move within the first background animation. That is, after the user double-clicks the natural particles in the first nature effect again, the natural particles in the first nature effect resume moving.
[0066] For example, if the first natural special effect is a rain special effect, after the user double-clicks the raindrops in the rain special effect, the raindrops in the rain special effect stop falling, and after the user double-clicks the raindrops in the rain special effect again, the raindrops in the rain special effect resume falling.
[0067] By providing an interface interaction method between the first background animation and the user, the user can participate in the movement changes of the first background animation during the playback of the first audio, enriching the screen display effect of the music playback interface, helping to enhance the user's interest in listening to music, and increasing the user's attention to the first background animation.
[0068] Next, the generation process of the first background dynamic effect is introduced. The generation process of the first background dynamic effect can be executed on the terminal device side or on the server side, and this application does not limit this.
[0069] Please refer to Figure 4 , which shows a flow chart of a method for generating a first background animation provided by an embodiment of the present application. The execution subject of each step of the method is a computer device, which can be a terminal device or a server. The method can include at least one of the following steps 410 to 440:
[0070] Step 410: Obtain feature information of the first audio according to the first audio.
[0071] The feature information of the first audio may include rhythm features of the first audio and audio features of the first audio, the rhythm features of the first audio are used to indicate the beat information of the first audio and the rhythm pattern of the first audio, and the audio features of the first audio are used to indicate the audio melody of the first audio, the audio style of the first audio and the timbre of the instrument in the first audio.
[0072] In step 420 , an emotional feature of the first audio is obtained based on the feature information of the first audio through a neural network model. The emotional feature of the first audio is used to indicate the audio emotion contained in the first audio.
[0073] The feature information of the first audio is input into the neural network model, and the neural network model outputs the emotional feature of the first audio, where the emotional feature of the first audio is a feature vector of the audio emotion contained in the first audio.
[0074] Optionally, the neural network model can be a model based on deep learning, and the neural network model can include a convolutional neural network, a recurrent neural network, a long short-term memory network and a gated recurrent unit. The specific process of generating emotional features through the neural network model is not introduced in this application.
[0075] Optionally, the neural network model can also be a multimodal large model. The multimodal large language model here can be any publicly available large language model, such as a natural language model based on a transformer structure obtained by training with a large amount of data. The large amount of data can reach a sample level of more than 100 million, and this application does not limit this.
[0076] In step 430 , a first natural special effect and a first physical special effect are obtained based on the feature information of the first audio and the emotional features of the first audio using a large multimodal model.
[0077] A first natural special effect is obtained by using a large multimodal model based on the characteristic information of the first audio and the emotional characteristics of the first audio. Optionally, a first physical special effect corresponding to the first natural special effect can be determined based on a mapping relationship between natural special effects and physical special effects. Alternatively, a first physical special effect can be obtained by using a large multimodal model based on the characteristic information of the first audio and the emotional characteristics of the first audio. The specific process for determining the first natural special effect and the first physical special effect can be referred to in the following embodiments and will not be described here.
[0078] Step 440: Render and generate a first background motion effect according to the first natural special effect and the first physical special effect.
[0079] The first background motion effect is rendered and generated according to the first natural special effect and the first physical special effect through the dynamic rendering engine.
[0080] By combining the characteristic information of the first audio and the emotional characteristics of the first audio to determine the first natural special effect and the first physical special effect, the fit between the first natural special effect and the first physical special effect and the audio emotion contained in the first audio can be improved, thereby improving the emotional expression ability of the first background animation, making it easier for users to accurately understand the audio emotion contained in the first audio when listening to the audio.
[0081] Figure 5 A schematic diagram of the generation process of the first background motion effect is shown. First, the first audio is input into the audio analysis module, and the audio analysis module extracts the rhythm characteristics of the first audio and the audio characteristics of the first audio. Secondly, the rhythm characteristics of the first audio and the audio characteristics of the first audio are input into the neural network model, and the neural network model extracts the emotional characteristics of the first audio. Then, the emotional characteristics of the first audio are input into the multimodal large model, and the multimodal large model determines the first natural special effect and the first physical special effect. Finally, the first natural special effect and the first physical special effect are input into the dynamic rendering engine, and the dynamic rendering engine renders to generate the first background motion effect.
[0082] In some embodiments, step 410 includes at least one of sub-steps 411 - 412 .
[0083] In sub-step 411 , rhythm extraction is performed on the first audio to obtain a rhythm feature of the first audio. The rhythm feature of the first audio is used to indicate beat information of the first audio and a rhythm pattern of the first audio.
[0084] The beat information of the first audio is used to indicate the number of beats contained in each minute of the first audio. If the number of beats contained in each minute of the first audio is the same, for example, the first audio can be a song audio, then beat extraction is performed directly on the first audio to obtain the beat information of the first audio. If the number of beats contained in each minute of the first audio is different, for example, the first audio can be an audio type such as opera audio, audiobook audio, etc., then a first sliding window is used to cut out segments of the first audio to obtain at least two audio segments corresponding to the first audio, beat extraction is performed on the at least two audio segments corresponding to the first audio respectively, and beat information corresponding to the at least two audio segments corresponding to the first audio is obtained, and the beat information corresponding to the at least two audio segments corresponding to the first audio is averaged to obtain the beat information of the first audio.
[0085] The rhythm pattern of the first audio is used to indicate the organization and presentation of the rhythm in the first audio. Different rhythm patterns have different patterns of rhythm strength and note length. The rhythm pattern of the first audio includes but is not limited to rhythm patterns such as 2 / 4, 3 / 4, and 4 / 4 time.
[0086] In sub-step 412, feature extraction is performed on the first audio to obtain audio features of the first audio. The audio features of the first audio are used to indicate the audio melody of the first audio, the audio style of the first audio, and the timbre of the instrument in the first audio.
[0087] The audio melody of the first audio is used to indicate the melody mode of the first audio. The melody mode of the first audio can be one of major, minor, and folk modes. Different melody modes have different scale structures, and each melody mode can be further subdivided into specific modes. For example, the major mode can be subdivided into C major, G major, D major, A major, and other modes. Different melody modes convey different audio emotions. For example, tunes composed in major keys often convey positive emotions such as joy, happiness, excitement, and solemnity, while tunes composed in minor keys are often used to express emotions such as sadness, sorrow, contemplation, and mystery.
[0088] The audio style of the first audio can be expressed by various features, such as spectrum features, rhythm features, timbre features, energy features, etc.
[0089] The timbre of the instrument in the first audio is related to the accompaniment instrument used in the first audio. Different instruments convey different audio emotions. For example, if the instrument in the first audio is a guzheng, the first audio often conveys a variety of emotions such as tranquility, remoteness, joy, and sorrow. If the instrument in the first audio is a saxophone, the first audio often conveys a variety of emotions such as romance, melancholy, and affection.
[0090] If the first audio includes a human voice track, the audio features of the first audio can also be used to indicate text information associated with the first audio. For example, if the first audio is a song, the audio features of the first audio can also be used to indicate the lyrics of the first audio. If the first audio is an opera, the audio features of the first audio can also be used to indicate the lyrics of the first audio. If the first audio is an audiobook, the audio features of the first audio can also be used to indicate the text information of the audiobook corresponding to the first audio.
[0091] By extracting rhythm features and audio features from the first audio, various features of the first audio are fully extracted, so that the feature information of the first audio can be represented more accurately, thereby obtaining the first natural special effects and the first physical special effects based on the feature information of the first audio, which helps to improve the generation accuracy of the first background special effects.
[0092] In some embodiments, the first audio may be input into a feature extraction model, and the feature extraction model may input audio features of the first audio.
[0093] In some embodiments, preprocessing is performed on the first audio to obtain the preprocessed first audio; a fast Fourier transform is performed on the preprocessed first audio to obtain a power spectrum of the preprocessed first audio; the power spectrum is input into a Mel filter group, and the logarithm of the output information of the Mel filter group is taken to obtain the logarithmic energy of the power spectrum; a discrete cosine transform is performed on the logarithmic energy of the power spectrum to obtain audio features of the first audio.
[0094] Preprocessing operations include pre-emphasis, framing, and windowing. First, pre-emphasis is performed on the first audio to obtain the pre-emphasized first audio. A high-pass filter is used to enhance the high-frequency portion of the first audio, compensating for the attenuation of the high-frequency component during transmission, making the spectrum of the first audio flatter. Pre-emphasis can usually be achieved using a first-order differential equation. Then, framing is performed on the pre-emphasized first audio to obtain the framed first audio. The continuous first audio is divided into overlapping short-time frames, each frame typically containing 20 to 40 milliseconds of audio data, with 10 to 25 milliseconds of overlap between frames. Finally, windowing is performed on the framed first audio to obtain the preprocessed first audio. A window function (such as a Hamming window) is applied to each frame of audio data in the framed first audio to reduce frame edge effects and make spectral analysis smoother.
[0095] A fast Fourier transform is performed on the preprocessed first audio to obtain a power spectrum of the preprocessed first audio. That is, a fast Fourier transform is performed on each frame of windowed audio data to convert the time domain signal into a frequency domain signal to obtain a spectrum of each frame of windowed audio data, and then the square of the modulus of the spectrum is calculated to obtain a power spectrum of each frame of windowed audio data.
[0096] A set of triangular filters can be designed with center frequencies evenly distributed on the Mel scale, covering the entire frequency range (typically 0-8000 Hz). The number of filter banks is typically 20-40. The power spectrum is fed into the Mel filter bank, and the energy output of each filter is calculated. The logarithm of the energy output of each filter is taken to obtain the logarithmic energy of the power spectrum. Logarithmic energy simulates the nonlinear perception of sound intensity by the human ear.
[0097] A discrete cosine transform is performed on the logarithmic energy of the power spectrum to obtain MFCC (Mel Frequency Cepstral Coefficients) coefficients, which are used to indicate the audio characteristics of the first audio. Typically, the first 12-13 coefficients are retained, and the rest are discarded.
[0098] Optionally, the MFCC coefficients may be normalized, for example, by subtracting the mean and dividing by the standard deviation, to improve the stability and comparability of the audio features.
[0099] By extracting MFCC coefficients, we can simulate the perceptual characteristics of the human ear that are more sensitive to low frequencies and have lower resolution for high frequencies, so that features related to human perception can be extracted more effectively. In addition, the MFCC coefficients are highly robust and can improve the stability and reliability of the audio features of the first audio, making it easier to subsequently obtain the first background motion effect based on the audio features of the first audio.
[0100] In some embodiments, step 430 includes at least one of sub-steps 431 - 433 .
[0101] In sub-step 431 , a first natural special effect is determined from at least two preset candidate natural special effects according to feature information of the first audio and emotional features of the first audio using a multimodal large model.
[0102] The feature information of the first audio, the emotional characteristics of the first audio, and the first prompt text are input into the multimodal large model, and the multimodal large model outputs a first natural special effect. The first prompt text is used to instruct the multimodal large model to determine the first natural special effect from at least two candidate natural special effects. The first prompt text includes text information of at least two candidate natural special effects. The candidate natural special effects are natural special effects pre-set by a technician and are not limited to these in this application.
[0103] Exemplarily, if at least two candidate natural special effects are natural special effects in the field of weather forms, and at least two candidate natural special effects include fog special effects, drizzle special effects, snowfall special effects, clear sky special effects, breeze special effects, rainbow special effects, thunderstorm special effects, hurricane special effects and aurora special effects, then the first prompt text may be: Please determine the natural special effects that match the feature information of the input first audio and the emotional characteristics of the first audio from the following candidate natural special effects based on the feature information of the input first audio and the emotional characteristics of the first audio. The candidate natural special effects include fog special effects, drizzle special effects, snowfall special effects, clear sky special effects, breeze special effects, rainbow special effects, thunderstorm special effects, hurricane special effects and aurora special effects.
[0104] In sub-step 432, when at least two candidate natural special effects include multiple natural special effects corresponding to at least one special effect type, a first physical special effect corresponding to the first natural special effect is obtained according to a mapping relationship between the candidate natural special effects and the candidate physical special effects. The multiple natural special effects corresponding to each special effect type are used to indicate different motion characteristics of natural particles in the special effect type.
[0105] At least two candidate natural special effects include multiple natural special effects corresponding to at least one special effect type, which means that at least one special effect type exists in the at least two candidate natural special effects, each special effect type corresponds to multiple natural special effects, and the multiple natural special effects corresponding to each special effect type are used to indicate different motion characteristics of natural particles in that special effect type. Exemplarily, the at least two candidate natural special effects are natural special effects in the field of weather morphology, and the at least two candidate natural special effects include multiple natural special effects corresponding to a rain special effect type, and the multiple natural special effects corresponding to the rain special effect type are light rain weather effect, moderate rain weather effect, and heavy rain weather effect, respectively.
[0106] When at least two candidate natural special effects include multiple natural special effects corresponding to at least one special effect type, it can be considered that at least two candidate natural special effects have refined the natural special effects of each special effect type more finely, then each natural special effect under each special effect type can more accurately refine the emotional concentration of the audio emotion contained in the first audio. For example, light rain weather special effects can be used to represent audio with a low concentration of sad emotions, moderate rain weather special effects can be used to represent audio with a medium concentration of sad emotions, and heavy rain weather special effects can be used to represent audio with a high concentration of sad emotions.
[0107] Therefore, when at least two candidate natural special effects are pre-set, and multiple natural special effects are set for each special effect type, the candidate physical special effects corresponding to each candidate natural special effect can be pre-set simultaneously. Then, a mapping relationship between the candidate natural special effects and the candidate physical special effects is pre-set, so that the first physical special effect corresponding to the first natural special effect can be determined based on the mapping relationship between the candidate natural special effects and the candidate physical special effects. For example, if the first natural special effect is a light rain weather special effect, the falling speed of the raindrops indicated by the first physical special effect is 3 mm / s; if the first natural special effect is a moderate rain weather special effect, the falling speed of the raindrops indicated by the first physical special effect is 10 mm / s; if the first natural special effect is a heavy rain weather special effect, the falling speed of the raindrops indicated by the first physical special effect is 20 mm / s.
[0108] By pre-setting the mapping relationship between candidate natural effects and candidate physical effects, the first physical effect corresponding to the first natural effect can be directly determined based on the first natural effect, simplifying the process of determining the first physical effect and improving the efficiency of generating the first physical effect. Furthermore, because the candidate natural effects are refined natural effects under each effect type, the correlation between the first physical effect and the audio emotion contained in the first audio is guaranteed, allowing the first physical effect to more accurately represent the audio emotion contained in the first audio.
[0109] In sub-step 433, when at least two candidate natural special effects include a natural special effect corresponding to at least two special effect types, the emotion level corresponding to the emotion feature of the first audio is obtained according to the feature information of the first audio and the emotion feature of the first audio through the multimodal large model. The emotion level is used to measure the emotion concentration contained in the first audio; and according to the emotion level corresponding to the emotion feature of the first audio, the first physical special effect corresponding to the emotion feature of the first audio is obtained.
[0110] The at least two candidate natural special effects include a natural special effect corresponding to at least two special effect types, respectively. This means that there are at least two special effect types in the at least two candidate natural special effects, each special effect type corresponds to a natural special effect, and the natural special effect corresponding to each special effect type is used to indicate the special effect type. Exemplarily, the at least two candidate natural special effects are natural special effects in the field of weather morphology, and the at least two candidate natural special effects include a natural special effect corresponding to a rain special effect type. The natural special effect corresponding to the rain special effect type is used to indicate a rain natural special effect.
[0111] In the case that at least two candidate natural special effects include a natural special effect corresponding to at least two special effect types respectively, it can be considered that the at least two candidate natural special effects only include the natural special effects corresponding to the special effect types, and the natural special effects of each special effect type are not refined. Then, a natural special effect under each special effect type can only roughly represent the audio emotion contained in the first audio, and cannot accurately represent the emotional concentration contained in the first audio. For example, the rain special effect type can only be used to represent the sad emotion in the first audio, and cannot accurately represent the emotional concentration of the sad emotion contained in the first audio.
[0112] Therefore, it is necessary to first input the feature information of the first audio, the emotional features of the first audio, and the second prompt text into the multimodal large model, which then outputs the emotion level corresponding to the emotional features of the first audio. The second prompt text is used to instruct the multimodal large model to determine the emotion level corresponding to the emotional features of the first audio from at least two candidate emotion levels corresponding to the emotional features of the first audio. The second prompt text contains textual information about the at least two candidate emotion levels corresponding to the emotional features of the first audio.
[0113] The emotion level is used to measure the concentration of emotions contained in the first audio. The higher the concentration of emotions contained in the first audio, the higher the emotion level corresponding to the emotion feature of the first audio. For example, if the emotion level is divided into 5 levels, and the emotion level corresponding to the emotion feature of the first audio is level 4, it means that the emotion concentration contained in the first audio is higher. According to the functional relationship between the emotion level corresponding to the emotion feature of the first audio and the motion feature, the first physical special effect corresponding to the emotion feature of the first audio is calculated. The functional relationship can be a linear functional relationship or a nonlinear functional relationship, and this application does not limit this. For example, if the emotion level corresponding to the emotion feature of the first audio is level 4, the first motion feature indicated by the first physical special effect can be 16mm / s.
[0114] By determining the emotional level corresponding to the emotional characteristics of the first audio without refining the candidate natural types, and accurately measuring the emotional concentration contained in the first audio through the emotional level, the first motion feature indicated by the first physical special effect obtained based on the emotional level can accurately reflect the emotional concentration contained in the first audio, thereby improving the accuracy of the first physical special effect, helping to enhance the emotional appeal of the first background motion effect, and thus enhancing the emotional perception between the audio and the user.
[0115] Table 1 below shows a brief mapping relationship between the beat information of the first audio and the emotion mapping, weather special effects, and physical special effects.
[0116] Table 1
[0117]
[0118] As can be seen from Table 1 above, the slower the beat information of the first audio, the sadder the corresponding audio emotion, and the mapped weather effects are also more soothing and peaceful. The faster the beat information of the first audio, the more exciting the corresponding audio emotion, and the mapped weather effects are also more intense and shocking.
[0119] Table 2 below shows an example analysis of background animation for a song audio.
[0120] Table 2
[0121]
[0122] As shown in Table 2 above, the BPM (Beats Per Minute) is slow. Slow songs often convey pensive, sentimental, gentle, or subdued emotions, suitable for expressing deeper emotions. The song is in the key of B minor, which inherently carries a melancholic, sad, and mysterious quality. In classical and pop music, B minor is often used to express deep sorrow or restrained emotions. The style of the song is lyrical pop, which emphasizes emotional expression and often revolves around themes such as love, separation, and memory. Combined with the slow tempo and minor key, the song's mood leans towards restrained sadness rather than intense pain. The instrument in the song is the flute, which has a deep, ethereal sound that often conveys loneliness, desolation, and an ancient aesthetic. The lyrics are lyrical and narrative, often conveying emotions through storytelling or imagery rather than direct expression. This resulted in the final background motion effect design: raindrops falling diagonally (matching the slow BPM); occasional wind-blown mist and rain waves (corresponding to the sound of the flute); and an orange-yellow halo of music in the background (corresponding to the lyrical and narrative atmosphere of the lyrics).
[0123] In some embodiments, step 220 includes sub-step 221 .
[0124] Sub-step 221, when the first audio is played to the first audio segment, the second background animation of the first audio is displayed in the music playing interface.
[0125] Among them, the second background animation includes a first natural special effect corresponding to the first audio, the first natural special effect includes a second physical special effect corresponding to the first audio clip, and the second physical special effect is used to indicate the second motion characteristics of the natural particles in the first natural special effect.
[0126] The first audio clip is the audio clip corresponding to the first time period in the first audio. The first time period can be any time period in the first audio, and this application does not limit the length of the first time period. The first audio can be divided into multiple first audio clips. When the first audio plays to different first audio clips, the second background animation corresponding to different first audio clips is displayed in the music playback interface.
[0127] The method for determining the first natural special effect in the second background motion effect can refer to the above embodiment and will not be introduced here. The second physical special effect in the second background motion effect is a physical special effect determined based on the feature information of the first audio clip and the emotional characteristics of the first audio clip. The second motion characteristics indicated by the second physical special effect are used to visually display the audio emotions included in the first audio clip. When the first audio is played to different first audio clips, the second physical special effect displayed in the music playback interface can be used to indicate the emotional concentration of the audio emotions included in different first audio clips.
[0128] For example, the length of the first time period can be set to 1 minute, and a first audio segment can be selected for the first audio every 1 minute. When the first audio segment is played to the next first audio segment every 1 minute, the second physical effect corresponding to the next first audio segment is displayed in the first natural effect on the music playback interface. For example, if the first natural effect is a rainy weather effect, and the second motion characteristic indicated by the second physical effect corresponding to the previous first audio segment may be 5 mm / s, the second motion characteristic indicated by the second physical effect corresponding to the next first audio segment may be 10 mm / s.
[0129] In some embodiments, the generation process of the second background animation is as follows: based on the first audio clip, characteristic information of the first audio clip is obtained; based on the characteristic information of the first audio clip, the emotional characteristics of the first audio clip are obtained through a neural network model, and the emotional characteristics of the first audio clip are used to indicate the audio emotions contained in the first audio clip; based on the characteristic information of the first audio clip and the emotional characteristics of the first audio clip, the emotional level corresponding to the emotional characteristics of the first audio clip is obtained through a multimodal large model; based on the emotional level corresponding to the emotional characteristics of the first audio clip, the second physical special effect corresponding to the emotional characteristics of the first audio clip is obtained; based on the first natural special effect and the second physical special effect, the second background animation is rendered and generated.
[0130] The characteristic information of the first audio segment includes a rhythm feature and an audio feature of the first audio segment. The rhythm feature of the first audio segment indicates the beat information and rhythm pattern of the first audio segment, and the audio feature indicates the audio melody, audio style, and timbre of the instruments in the first audio segment. If the number of beats per minute of the first audio segment is the same, for example, the first audio segment may be a song, beat extraction is performed on the first audio segment to obtain beat information for the first audio segment. If the number of beats per minute of the first audio segment is different, for example, the first audio segment may be an opera, an audiobook, or other audio type, a second sliding window is used to segment the first audio segment to obtain at least two audio segments corresponding to the first audio segment. Beat extraction is performed on each of the at least two audio segments corresponding to the first audio segment to obtain beat information corresponding to each of the at least two audio segments corresponding to the first audio segment. The beat information corresponding to the at least two audio segments corresponding to the first audio segment is averaged to obtain beat information for the first audio segment. Performing rhythm extraction on the first audio segment to obtain a rhythm pattern of the first audio segment. Performing feature extraction on the first audio segment to obtain an audio feature of the first audio segment.
[0131] By determining the second physical special effect corresponding to the first audio clip, the second motion feature indicated by the second physical special effect can accurately display the audio emotion included in the first audio clip. Therefore, when the first audio is played to different time periods, the second motion feature indicated by the second physical special effect in the first background motion effect can be consistent with the audio emotion included in the first audio clip currently playing, which helps to enhance the emotional appeal of the first background motion effect, thereby enhancing the emotional perception between the audio and the user.
[0132] In some embodiments, the above method further includes step 230 .
[0133] Step 230: In response to the operation of inputting the first text for the first audio, during the playback of the first audio, a third background animation of the first audio is displayed in the music playback interface.
[0134] Among them, the third background animation includes the first natural special effect corresponding to the first audio, the first natural special effect includes the first physical special effect corresponding to the first audio and the first auxiliary special effect corresponding to the first text, the first auxiliary special effect is used to visualize the audio emotions contained in the first text through natural phenomena, and the first text is used to describe the auditory experience of the first audio.
[0135] The operation of inputting the first text for the first audio is an operation performed by the user to set the first text for the first audio, and the first text is used to describe the user's auditory experience of the first audio. Optionally, the first text can be pre-set before the first audio is played, or it can be added during the playback of the first audio, and this application is not limited to this. Based on the first text, a first auxiliary special effect corresponding to the first text can be generated, and the first auxiliary special effect is used to be added to the first natural special effect to enhance the expressiveness of the first natural special effect. The first auxiliary special effect is an auxiliary natural phenomenon in the natural phenomenon corresponding to the first natural special effect. For example, if the first natural special effect is a rainy natural special effect, the first auxiliary special effect can be a lightning special effect, a dark matter particle flow special effect, a breeze special effect and other auxiliary special effects. If the first natural special effect is a sunny natural special effect, the first auxiliary special effect can be a halo diffraction special effect, a light special effect, a light spot special effect and other auxiliary special effects.
[0136] In some embodiments, the generation process of the third background animation is: based on the first text, obtaining the text features of the first text; determining the first auxiliary special effect from at least two pre-set candidate auxiliary special effects based on the feature information of the first audio, the emotional features of the first audio and the text features of the first text through a multimodal large model; rendering and generating the third background animation based on the first natural special effect, the first physical special effect and the first auxiliary special effect.
[0137] Perform feature extraction on the first text to obtain text features of the first text. Input the feature information of the first audio, the text features of the first text and the third prompt text into the multimodal large model, and the multimodal large model outputs the first auxiliary special effect. The third prompt text is used to instruct the multimodal large model to determine the first auxiliary special effect from at least two candidate auxiliary special effects, and the third prompt text contains text information of at least two candidate auxiliary special effects. The candidate auxiliary special effects are natural special effects pre-set by technical personnel, and this application does not limit this.
[0138] The at least two candidate auxiliary special effects include at least two auxiliary special effects corresponding to at least two candidate natural special effects respectively. Each candidate natural special effect is pre-set with at least two auxiliary special effects. The multimodal large model first determines the first natural special effect from the at least two pre-set candidate natural special effects based on the feature information of the first audio and the emotional features of the first audio, and then filters out at least two auxiliary special effects corresponding to the first natural special effect from the at least two candidate auxiliary special effects. The multimodal large model then determines the first auxiliary special effect from the at least two auxiliary special effects corresponding to the first natural special effect based on the feature information of the first audio and the text features of the first text.
[0139] After obtaining the first natural special effect, the first physical special effect and the first auxiliary special effect, a third background dynamic effect is generated by rendering through a dynamic rendering engine.
[0140] By adding a first auxiliary effect corresponding to the first text to the first natural effect, users can participate in the design of the first background effect, enhancing the interaction between the first background effect and the user interface and increasing the user's interest in the audio playback. Furthermore, the first auxiliary effect allows the first background effect to accurately reflect the user's auditory experience of the first audio, enhancing the dynamic display effect of the first background effect and making it more personalized to the user's listening needs.
[0141] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0142] Please refer to Figure 6 , which shows a block diagram of a display device for a music playback interface provided by an embodiment of the present application. The device has the function of implementing the display method of the above-mentioned music playback interface, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device described above, or it can be set in a computer device. Figure 6 As shown, the device 600 may include: an interface display module 610 and a motion effect display module 620.
[0143] The interface display module 610 is used to display the music playback interface of the first audio.
[0144] The motion effect display module 620 is used to display the first background motion effect of the first audio in the music playback interface during the playback of the first audio;
[0145] Among them, the first background animation effect includes a first natural special effect corresponding to the first audio, the first natural special effect includes a first physical special effect corresponding to the first audio, the first natural special effect is used to visualize the audio emotion contained in the first audio through natural phenomena, and the first physical special effect is used to indicate the first motion characteristics of natural particles in the first natural special effect.
[0146] In some embodiments, the motion effect display module 620 is used to:
[0147] When the first audio is played to the first audio segment, a second background animation of the first audio is displayed in the music playing interface;
[0148] Among them, the second background animation includes a first natural special effect corresponding to the first audio, the first natural special effect includes a second physical special effect corresponding to the first audio clip, and the second physical special effect is used to indicate the second motion characteristics of the natural particles in the first natural special effect.
[0149] In some embodiments, the motion effect display module 620 is used to:
[0150] In response to an operation of inputting a first text into the first audio, during playback of the first audio, displaying a third background animation of the first audio in the music playback interface;
[0151] Among them, the third background animation includes the first natural special effect corresponding to the first audio, the first natural special effect includes the first physical special effect corresponding to the first audio and the first auxiliary special effect corresponding to the first text, the first auxiliary special effect is used to visualize the audio emotions contained in the first text through natural phenomena, and the first text is used to describe the auditory experience of the first audio.
[0152] In some embodiments, the apparatus 600 further includes an interactive display module, wherein the interactive display module is configured to:
[0153] In response to a sliding operation on the first natural special effect, displaying, in the first background motion effect, natural particles in the first natural special effect moving in a sliding direction corresponding to the sliding operation;
[0154] or,
[0155] In response to a click operation on the first natural special effect, displaying the decelerated motion of the natural particles in the first natural special effect in the first background motion effect;
[0156] or,
[0157] In response to a long press operation on the first natural special effect, displaying accelerated motion of natural particles in the first natural special effect in the first background motion effect;
[0158] or,
[0159] In response to a double-click operation on the first natural special effect, the natural particles in the first natural special effect are displayed in the first background motion effect to stop moving.
[0160] In some embodiments, the apparatus 600 further includes a motion effect generation module, wherein the motion effect generation module is configured to:
[0161] obtaining feature information of the first audio according to the first audio;
[0162] Obtaining, using a neural network model, an emotional feature of the first audio according to feature information of the first audio, where the emotional feature of the first audio is used to indicate an audio emotion contained in the first audio;
[0163] Obtaining the first natural special effect and the first physical special effect according to the feature information of the first audio and the emotional features of the first audio through a multimodal large model;
[0164] The first background motion effect is generated by rendering according to the first natural special effect and the first physical special effect.
[0165] In some embodiments, the feature information of the first audio includes a rhythm feature of the first audio and an audio feature of the first audio; and the motion effect generation module is configured to:
[0166] Performing rhythm extraction on the first audio to obtain a rhythm feature of the first audio, where the rhythm feature of the first audio is used to indicate beat information of the first audio and a rhythm pattern of the first audio;
[0167] Feature extraction is performed on the first audio to obtain audio features of the first audio, where the audio features of the first audio are used to indicate an audio melody of the first audio, an audio style of the first audio, and a timbre of an instrument in the first audio.
[0168] In some embodiments, the motion effect generation module is used to:
[0169] performing preprocessing on the first audio to obtain a preprocessed first audio;
[0170] Performing a fast Fourier transform on the preprocessed first audio to obtain a power spectrum of the preprocessed first audio;
[0171] Inputting the power spectrum into a Mel filter bank, taking the logarithm of output information of the Mel filter bank, and obtaining the logarithmic energy of the power spectrum;
[0172] Perform discrete cosine transform on the logarithmic energy of the power spectrum to obtain audio features of the first audio.
[0173] In some embodiments, the motion effect generation module is used to:
[0174] Determining, by the multimodal large model, the first natural special effect from at least two preset candidate natural special effects based on the feature information of the first audio and the emotional features of the first audio;
[0175] In the case where the at least two candidate natural special effects include multiple natural special effects corresponding to at least one special effect type, the first physical special effect corresponding to the first natural special effect is obtained according to the mapping relationship between the candidate natural special effects and the candidate physical special effects, and the multiple natural special effects corresponding to each special effect type are used to indicate different motion characteristics of natural particles in the special effect type.
[0176] In some embodiments, the motion effect generation module is used to:
[0177] When the at least two candidate natural special effects include one natural special effect corresponding to at least two special effect types, obtaining, by the multimodal large model, an emotion level corresponding to the emotion feature of the first audio according to feature information of the first audio and the emotion feature of the first audio, the emotion level being used to measure the concentration of emotion contained in the first audio;
[0178] The first physical effect corresponding to the emotional feature of the first audio is obtained according to the emotional level corresponding to the emotional feature of the first audio.
[0179] In some embodiments, the first natural special effect is a natural special effect in one of the following natural fields:
[0180] Weather morphology field;
[0181] the field of geomorphology;
[0182] the field of astronomical morphology;
[0183] Biomorphic field.
[0184] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0185] Please refer to Figure 7 , which shows a block diagram of a computer device 700 provided in one embodiment of the present application. The computer device 700 can be any electronic device with data calculation, processing, and storage functions. The computer device 700 can be used to implement the display method of the music playback interface provided in the above embodiment.
[0186] Typically, the computer device 700 includes a processor 701 and a memory 702 .
[0187] The processor 701 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0188] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 702 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned method for displaying the music playback interface.
[0189] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation on the computer device 700, and the computer device 700 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.
[0190] In an exemplary embodiment, a computer-readable storage medium is also provided, storing a computer program that, when executed by a processor of a computer device, implements the above-described method for displaying a music playback interface. Optionally, the computer-readable storage medium may be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, or an optical data storage device.
[0191] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the above-described method for displaying a music playback interface.
[0192] It should be noted that this application can display a prompt interface, pop-up window or output voice prompt information before collecting the user's relevant data and during the process of collecting the user's relevant data. The prompt interface, pop-up window or voice prompt information is used to remind the user that its relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are terminated, that is, the user's relevant data is not obtained. In other words, all user data collected by this application are processed strictly in accordance with the requirements of relevant national laws and regulations. The informed consent or separate consent of the personal information subject is obtained only when the user agrees and authorizes it to collect the data. Subsequent data use and processing are carried out within the scope of authorization of laws and regulations and the personal information subject, and the collection, use and processing of relevant user data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0193] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.
[0194] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for displaying a music playing interface, characterized in that: The method comprises: Display the music playback interface of the first audio; During the playback of the first audio, a first background animation of the first audio is displayed in the music playback interface; Among them, the first background animation effect includes a first natural special effect corresponding to the first audio, the first natural special effect includes a first physical special effect corresponding to the first audio, the first natural special effect is used to visualize the audio emotion contained in the first audio through natural phenomena, and the first physical special effect is used to indicate the first motion characteristics of natural particles in the first natural special effect.
2. The method according to claim 1, characterized in that During the playback of the first audio, displaying a first background animation effect of the first audio in the music playback interface includes: When the first audio is played to the first audio segment, a second background animation of the first audio is displayed in the music playing interface; Among them, the second background animation includes a first natural special effect corresponding to the first audio, the first natural special effect includes a second physical special effect corresponding to the first audio clip, and the second physical special effect is used to indicate the second motion characteristics of the natural particles in the first natural special effect.
3. The method according to claim 1 or 2, characterized in that The method further comprises: In response to an operation of inputting a first text into the first audio, during playback of the first audio, displaying a third background animation of the first audio in the music playback interface; Among them, the third background animation includes the first natural special effect corresponding to the first audio, the first natural special effect includes the first physical special effect corresponding to the first audio and the first auxiliary special effect corresponding to the first text, the first auxiliary special effect is used to visualize the audio emotions contained in the first text through natural phenomena, and the first text is used to describe the auditory experience of the first audio.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: In response to a sliding operation on the first natural special effect, displaying, in the first background motion effect, natural particles in the first natural special effect moving in a sliding direction corresponding to the sliding operation; or, In response to a click operation on the first natural special effect, displaying the decelerated motion of the natural particles in the first natural special effect in the first background motion effect; or, In response to a long press operation on the first natural special effect, displaying accelerated motion of natural particles in the first natural special effect in the first background motion effect; or, In response to a double-click operation on the first natural special effect, the natural particles in the first natural special effect are displayed in the first background motion effect to stop moving.
5. The method according to claim 1, wherein The method further comprises: obtaining feature information of the first audio according to the first audio; Obtaining, using a neural network model, an emotional feature of the first audio according to feature information of the first audio, where the emotional feature of the first audio is used to indicate an audio emotion contained in the first audio; Obtaining the first natural special effect and the first physical special effect according to the feature information of the first audio and the emotional features of the first audio through a multimodal large model; The first background motion effect is generated by rendering according to the first natural special effect and the first physical special effect.
6. The method according to claim 5, characterized in that The feature information of the first audio includes a rhythm feature of the first audio and an audio feature of the first audio; The obtaining, according to the first audio, feature information of the first audio includes: Performing rhythm extraction on the first audio to obtain a rhythm feature of the first audio, where the rhythm feature of the first audio is used to indicate beat information of the first audio and a rhythm pattern of the first audio; Feature extraction is performed on the first audio to obtain audio features of the first audio, where the audio features of the first audio are used to indicate an audio melody of the first audio, an audio style of the first audio, and a timbre of an instrument in the first audio.
7. The method according to claim 6, characterized in that The performing feature extraction on the first audio to obtain audio features of the first audio includes: performing preprocessing on the first audio to obtain a preprocessed first audio; Performing a fast Fourier transform on the preprocessed first audio to obtain a power spectrum of the preprocessed first audio; Inputting the power spectrum into a Mel filter bank, taking the logarithm of output information of the Mel filter bank, and obtaining the logarithmic energy of the power spectrum; Perform discrete cosine transform on the logarithmic energy of the power spectrum to obtain audio features of the first audio.
8. The method according to any one of claims 5 to 7, characterized in that The obtaining of the first natural special effect and the first physical special effect according to the feature information of the first audio and the emotional feature of the first audio by the multimodal large model includes: Determining, by the multimodal large model, the first natural special effect from at least two preset candidate natural special effects based on the feature information of the first audio and the emotional features of the first audio; In the case where the at least two candidate natural special effects include multiple natural special effects corresponding to at least one special effect type, the first physical special effect corresponding to the first natural special effect is obtained according to the mapping relationship between the candidate natural special effects and the candidate physical special effects, and the multiple natural special effects corresponding to each special effect type are used to indicate different motion characteristics of natural particles in the special effect type.
9. The method according to claim 8, characterized in that The method further comprises: When the at least two candidate natural special effects include one natural special effect corresponding to at least two special effect types, obtaining, by the multimodal large model, an emotion level corresponding to the emotion feature of the first audio according to feature information of the first audio and the emotion feature of the first audio, the emotion level being used to measure the concentration of emotion contained in the first audio; The first physical effect corresponding to the emotional feature of the first audio is obtained according to the emotional level corresponding to the emotional feature of the first audio.
10. The method according to claim 8 or 9, characterized in that The first natural special effect is one of the following natural special effects: Weather morphology field; the field of geomorphology; the field of astronomical morphology; Biomorphic field.
11. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method for displaying a music playback interface as described in any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the method for displaying a music playback interface as described in any one of claims 1 to 10.
13. A computer program product, characterized in that The computer program product includes a computer program, which is loaded and executed by a processor to implement the method for displaying a music playback interface as described in any one of claims 1 to 10.