Music driving fluid generation method and system based on multi-dimensional semantic mapping

By using multidimensional semantic mapping technology, music features and emotion vectors are mapped to fluid colors and physical parameters, which solves the problems of parameter dispersion and insufficient control in existing music visualization, achieves a unified rendering effect in multi-style scenes, and improves the synchronicity and scalability of music visualization.

CN121922147APending Publication Date: 2026-04-24EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EAST CHINA NORMAL UNIV
Filing Date
2026-01-23
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing music visualization technologies lack a unified space for emotion and color parameters, making it difficult to reuse them across different style scenes. Furthermore, the lack of a layered control mechanism for physical parameters results in animation effects that lack diversity and scalability.

Method used

By employing a multidimensional semantic mapping method, the audio features of music, such as spectrum, beat, structure, and energy, are collaboratively mapped with Valence-Arousal emotion vectors and palette information to fluid colors, physical parameters, and camera control parameters, thereby constructing a unified emotion and color parameter space and realizing a reusable rendering framework for various fluid styles and scenes.

Benefits of technology

It achieves differentiated fluid rendering under different musical emotions, improves audio-visual synchronization and scene reuse capabilities, and significantly enhances the scalability and unified parameter control of music visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121922147A_ABST
    Figure CN121922147A_ABST
Patent Text Reader

Abstract

The invention discloses a music driving fluid generation method and system based on multi-dimensional semantic mapping, and belongs to the technical field of computer graphics and audio signal processing. The method comprises the following steps that music input by a user is received, characteristics such as rhythm, frequency spectrum and energy are obtained through audio analysis, and emotion vectors and color information changing along with time are obtained through a music emotion recognition and color palette generation module; in a unified parameter space, the audio, emotion and color semantics are mapped into fluid colors, fluid simulation parameters and camera and environment control parameters, and fluid models such as shallow water waves or height fields are driven to generate multi-style fluid scenes changing along with music. Compared with an existing music visualization mode which only depends on frequency band energy to drive simple and special effects, the method can achieve consistent expression of music emotion, rhythm, fluid form and audio-visual atmosphere in multiple scenes, and can be applied to scenes such as music visualization, immersive interactive art display and emotion relaxation assistance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer graphics and audio signal processing technology, specifically relating to a method and system for generating music-driven fluid based on multidimensional semantic mapping. Background Technology

[0002] Existing music visualization technologies can map audio features such as spectrum, beat, or volume into simple geometric changes, particle effects, or flowing backgrounds. Mainstream music players also offer the ability to automatically generate dynamic visuals based on music. However, most of these visualization solutions rely on empirical rules, directly driving several coloring or displacement parameters based on fixed album art colors with a small amount of frequency energy or overall volume. The mapping relationships are scattered and strongly coupled with specific scenes, making it difficult to reuse them across different style scenarios.

[0003] In the field of fluid-based music visualization, existing work mainly relies on spectral analysis to map volume or rhythm to local perturbation intensity or ripple quantity. It lacks an integrated mapping framework that integrates musical features and emotional semantics with fluid color, fluid physical parameters, and camera and environmental control. It also lacks a hierarchical control mechanism for physical parameters, making it difficult to maintain a consistent emotional expression and audio-visual structure synchronization in scene rendering. In addition, existing work lacks research on the expression of the intrinsic emotions and rich rhythmic changes in music, resulting in insufficient differentiation in the animation effects of music with different emotions and insufficient scalability.

[0004] To address the aforementioned issues, this invention proposes a music-driven fluid generation method and system based on multidimensional semantic mapping. Summary of the Invention

[0005] The purpose of this invention is to provide a music-driven fluid generation method and system based on multi-dimensional semantic mapping to solve the problems of scattered audio feature mapping rules, strong coupling with specific scenes, lack of unified parameter space, and insufficient granularity of fluid parameter control in existing music visualization systems. In a unified emotion and color parameter space, this invention collaboratively maps the audio features of music, such as spectrum, beat, structure, and energy, as well as Valence-Arousal emotion vectors and palette information, into fluid color, fluid physical parameters, and camera and environmental control parameters, thereby realizing a reusable, emotion-differentiated music-driven fluid rendering framework under various fluid styles and scenes.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A music-driven fluid generation method based on multidimensional semantic mapping includes the following steps: S1. Music Analysis and Feature Extraction: Obtain the music input selected by the user, build a multi-resolution audio analysis pipeline locally, perform spectrum, tempo, structure and energy analysis on the music, and extract audio feature sequences that change over time; S2, Emotion Recognition: Based on the audio feature sequence obtained in S1, the emotion recognition model is called to analyze it and obtain the valence-arousal (V,A) emotion vector corresponding to each time slice, as well as the emotion category label for the whole song; at the same time, the recognition results are sent to the color palette generation model to generate the corresponding multi-channel color vector or color palette. S3. Emotion Mapping: In a unified emotion parameter space, the valence-arousal (V,A) emotion vector obtained in S2 is mapped to the parameters of the fluid's primary color in the HSL color space (or equivalent color space). The brightness is kept within a preset range, and the hue and saturation are obtained from the valence-arousal (V,A) emotion vector through a monotonic mapping function, which is used to determine the primary color of the current fluid material. S4. Energy Calculation and Visual Allocation: In the CIELCh perceptual uniform color space, based on the color palette obtained in S2, the chromaticity of each color and its CIEDE2000 color difference with neutral gray of the same brightness are calculated, thereby obtaining the visual perception energy of each color; according to the magnitude of the visual perception energy, the colors are mapped to the fluid body, ambient lighting and volumetric fog effect, so that high perception energy colors correspond to high saliency visual elements, and the brightness and saturation of each element are adjusted in real time through a preset energy modulation function; S5. Fluid Simulation and Parameter Mapping: Construct a fluid simulation system based on Shallow Water Equations and establish a three-level mapping mechanism from musical characteristics to fluid physical parameters (micro-meta-macro), specifically including: Microscopic mapping: Normalize the beats per minute (BPM) in the audio feature sequence to obtain the wave velocity parameters in the shallow water equation or equivalent fluid model; route the energy of multiple preset frequency bands to spatially separated disturbance sources or interactive bodies, inject disturbances into the fluid height field or velocity field, and update the fluid morphology as the spectral energy changes. Mesoscopic mapping: The source term of the fluid equation is topologically modified according to the valence-arousal (V,A) emotion vector; when the emotional semantics are high arousal, the source term is set as an upward jet source to generate a diffuse wave; when the emotional semantics are low valence, the source term is set as a downward sink source to generate an inward wave. Macroscopic mapping: mapping the overall energy envelope of music to the global viscosity coefficient and wave velocity parameters of a fluid system; S6. Camera Motion and Environment Control: Based on audio features and emotion vectors, select a predefined camera motion mode in the camera control state machine, and control the time evolution of environmental parameters such as camera path, angle of view, focal length, volumetric light, fog effect, exposure, and color temperature, so that the camera motion and environmental changes are synchronized with the energy and paragraph structure of the music. S7. Real-time Solving and Rendering Output: Input the color parameters, fluid physical parameters, fluid mesh state, camera and environment parameters obtained from S3 to S6 into the fluid rendering framework, and perform fluid solving and scene rendering through the graphics processing unit (GPU). Finally, input a fluid animation sequence aligned with the music timeline.

[0007] Preferably, the emotion recognition model and the color palette generation model described in S2 are invoked in parallel through a unified reasoning process, so that the main fluid color is directly determined by the valence-arousal (V,A) emotion vector; the environment and special effects colors are jointly driven by the color palette analysis results and the visual perception energy allocation, thereby simultaneously obtaining the main fluid color, environment color and focus color in a single music analysis, so as to achieve consistent emotion expression across color, fluid and camera subsystems.

[0008] Preferably, the formula for the mapping operation in S3 is expressed as follows:

[0009]

[0010]

[0011] in, Hue and H Indicates hue; Saturation and S Indicates saturation; V Indicates valence; A Indicates wakefulness; F It is a linear mapping function; H wrap , H end 、H max , H min The values ​​of each point representing the target hue; V min , V max , V mid Indicates the range of values ​​for source valence;x , x 1, x 2 represents the input value and the binary reference mapping source term, respectively; y 1, y 2 represents the binary reference mapping target item.

[0012] Preferably, the method for acquiring visual perception energy in S4 specifically includes: Use the CIELCh color space to obtain the chroma of each color. C *( C i For each color C i Calculate its neutral gray, which is the same as the perceived brightness (L). G i The CIEDE2000 color difference between them is then used to define the visual perception energy based on weights, as shown in the specific formula. as follows:

[0013] in, Represents visual perception energy; C i Representing each color; C *( C i () indicates the chroma of each color; Gi indicates neutral gray; The CIEDE2000 color difference represents the difference between neutral grays with the same perceived brightness for each color. w c This represents the weight value of each color chromaticity in visual perception energy; w d The weight of the CIEDE2000 color difference between neutral grays with the same perceived brightness in visual perception energy; High-energy colors are assigned to the fluid subject and foreground highlights, while low-energy colors are assigned to the background and sky elements. The brightness and saturation of these elements are constrained using an energy-based modulation function, ensuring that the overall image maintains visual contrast and a unified emotional tone across different music clips. The energy modulation function is defined as follows:

[0014] in, A ( e () indicates the assigned base color HSL value; M ={ M energy , M beat},M energy This represents the spectral energy of music. M beat Indicates the number of beats in the music; The saturation and brightness of each element are modulated in real time using an energy modulation function to satisfy:

[0015]

[0016]

[0017] in, Indicates the brightness of the HSL color of the highlighted element; The saturation of the HSL color representing the highlighted element. Indicates the saturation of the HSL colors for the background and sky elements; This is the sensitivity parameter.

[0018] Preferably, the construction of the fluid simulation system based on the Shallow Water Equations described in S5 further includes the following: We employ shallow water equations or equivalent two-dimensional height field fluid models, and iteratively update the height and velocity fields using fragment shaders and computation shaders on the GPU. The resulting wave velocity parameters include: converting the beat count (BPM) into shallow water wave propagation speed and converting the pitch into the fluid viscosity coefficient; The resulting frequency bands include low-frequency, mid-frequency, and high-frequency bands, which are used to drive: the generation and propagation of large-scale waves, the formation of mid-scale ripples and vortices, and small-scale high-frequency details or particle spray effects, respectively. By varying the location and intensity of the disturbance source, a fluid response to music is achieved through scene rendering.

[0019] Preferably, the camera control state machine in S6 includes at least the following states: Orbital cruise mode: The camera moves in an orbit around the center of the scene at a constant or gradually changing angular velocity; Diving motion mode: The camera rapidly approaches the fluid surface along a preset curve, and the depth of field changes are superimposed; Slow track mode: The camera moves or pans at a low speed, suitable for low-energy music clips; Stable stillness: The camera's field of view remains basically still, with only slight shaking or breathing-like zooming; The state machine performs state transitions based on music energy, beat phase, and segment labels, and performs smooth interpolation of camera trajectory and environmental parameters during switching to reduce visual abrupt changes and ensure synchronization with the music timeline.

[0020] This invention further protects a music-driven fluid generation system based on multidimensional semantic mapping, comprising: Music Analysis and Feature Extraction Module: Used to acquire music input selected by the user, build a multi-resolution audio analysis pipeline locally, perform spectrum, tempo, structure and energy analysis on the music, and extract audio feature sequences that change over time; Emotion Recognition Module: Based on the obtained audio feature sequence, the emotion recognition model is called to analyze it and obtain the valence-arousal (V,A) emotion vector corresponding to each time slice, as well as the emotion category label for the whole song; at the same time, the recognition results are sent to the color palette generation model to generate the corresponding multi-channel color vector or color palette. Emotion Mapping Module: In a unified emotion parameter space, the obtained valence-arousal (V,A) emotion vector is mapped to the parameters of the fluid's primary color in the HSL color space (or equivalent color space). The brightness is kept within a preset range, and the hue and saturation are obtained from the valence-arousal (V,A) emotion vector through a monotonic mapping function to determine the primary color of the current fluid material. Energy Calculation and Visual Allocation Module: In the CIELCh perceptual uniform color space, based on the obtained color palette, the chromaticity of each color and its CIEDE2000 color difference with neutral gray of the same brightness are calculated, thereby obtaining the visual perception energy of each color; according to the magnitude of the visual perception energy, the colors are mapped to the fluid body, ambient lighting and volumetric fog effect, so that high perception energy colors correspond to high saliency visual elements, and the brightness and saturation of each element are adjusted in real time through a preset energy modulation function; Fluid Simulation and Parameter Mapping Module: Constructs a fluid simulation system based on Shallow Water Equations and establishes a three-level mapping mechanism from musical characteristics to fluid physical parameters (micro-meso-macro). Specifically, this includes: Microscopic mapping: Normalize the beats per minute (BPM) in the audio feature sequence to obtain the wave velocity parameters in the shallow water equation or equivalent fluid model; route the energy of multiple preset frequency bands to spatially separated disturbance sources or interactive bodies, inject disturbances into the fluid height field or velocity field, and update the fluid morphology as the spectral energy changes. Mesoscopic mapping: The source term of the fluid equation is topologically modified according to the valence-arousal (V,A) emotion vector; when the emotional semantics are high arousal, the source term is set as an upward jet source to generate a diffuse wave; when the emotional semantics are low valence, the source term is set as a downward sink source to generate an inward wave. Macroscopic mapping: mapping the overall energy envelope of music to the global viscosity coefficient and wave velocity parameters of a fluid system; Camera motion and environment control module: Based on audio features and emotion vectors, it selects a predefined camera motion mode in the camera control state machine and controls the time evolution of environmental parameters such as camera path, angle of view, focal length, volumetric light, fog effect, exposure, and color temperature, so that the camera motion and environmental changes are synchronized with the energy and paragraph structure of the music. Real-time solving and rendering output module: Input the obtained color parameters, fluid physics parameters, fluid mesh state, camera and environment parameters into the fluid rendering framework, and perform fluid solving and scene rendering through the graphics processing unit (GPU), and finally input a fluid animation sequence aligned with the music timeline.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a music-driven fluid generation method and system based on multidimensional semantic mapping, solving the problems of traditional music visualization that rely only on a small amount of frequency band energy, lack a unified parameter space, and lack fine fluid control. Specifically, by constructing a multidimensional semantic mapping framework that integrates spectrum, beat, structure, energy, Valence-Arousal emotion vectors, and palette information, it unifies the modeling of color control, shallow water wave height field fluid physical parameters, and camera and environmental parameters. This enables multi-scale fluid morphology control from microscopic perturbations to macroscopic wave velocity and viscosity coefficients. Furthermore, it performs perception-based color energy allocation in the CIELCh color space using CIEDE2000 difference, and introduces a music structure-aware camera state machine for environmental linkage control, significantly improving audio-visual synchronization and scene reuse capabilities. This framework is adaptable to various fluid rendering styles and can be widely applied in music visualization, immersive interactive art displays, emotional relaxation, and digital content creation, demonstrating high practical value and promising application prospects. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings involved in the embodiments are now briefly described. Obviously, the drawings in the following description are merely illustrative of some embodiments of the present invention. For those skilled in the art, other forms of drawings can be constructed based on these drawings without creative effort.

[0023] Figure 1 This is an overall flowchart of a music-driven fluid generation method based on multidimensional semantic mapping proposed in this invention; Figure 2 This is a schematic diagram of the color mapping process mentioned in the embodiments of the present invention; Figure 3 This is a schematic diagram of the fluid parameter mapping process mentioned in the embodiments of the present invention; Figure 4 This is a schematic diagram of the output fluid animation sequence mentioned in the embodiments of the present invention; Figure 5-6 This is a schematic diagram illustrating an example of the usage process mentioned in the embodiments of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Please see Figure 1 This invention proposes a music-driven fluid generation method based on multidimensional semantic mapping, comprising the following steps: Step a: The system acquires the music input selected by the user, establishes a multi-resolution audio analysis pipeline locally, performs spectrum, tempo, structure and energy analysis on the music, and obtains audio feature sequences that change over time. Step b: Based on the audio feature sequence in step a, the emotion recognition model is called to analyze it, obtain the valence-arousal (V,A) emotion vector for each time slice and the emotion category label for the whole song, and send it to the color palette generation model at the same time to obtain the corresponding multi-channel color vector or color palette. Step c: Please see Figure 2 In a unified emotion parameter space, the (V,A) emotion vector from step b is mapped to the HSL color or equivalent color space parameters of the fluid's primary color. The brightness dimension remains within a preset range, and hue and saturation are obtained from V and A through a monotonic mapping function to determine the primary color tone of the current fluid material. The specific mapping formula is calculated as follows:

[0026]

[0027]

[0028] in, V min =-3.74; V mid =3.00; V max =3.53; H max =267°; H min =0°; H wrap =360°; H end =339°; A min =-3.4; A max =2.79; S min =0%; S max =100%; Step d: Please see Figure 3 In the CIELCh perceptual uniform color space, based on the chromaticity of each color in the palette obtained in step b and the CIEDE2000 color difference between each color and neutral gray of the same brightness, the visual perception energy of each color is calculated. Based on the magnitude of the visual perception energy, the colors are mapped to the fluid body, ambient lighting and volumetric fog effect, so that high perception energy colors correspond to highly significant visual elements, and the brightness and saturation of each element are adjusted in real time through a preset energy modulation function. Step e: A fluid simulation system based on the Shallow Water Equations is constructed, establishing a three-level mapping mechanism from musical characteristics to fluid physical parameters at the micro, meso, and macro levels. 1) Microscopic mapping: Normalize and map the beats per minute (BPM) in the audio features obtained in step a to obtain the wave velocity parameters in the shallow water equation or equivalent fluid model; route the energy of multiple preset frequency bands to spatially separated disturbance sources or interactive bodies, inject disturbances into the fluid height field or velocity field, so that the fluid morphology is updated with the change of spectral energy. 2) Mesoscopic mapping: Based on the emotional semantic coordinates in step b, the source term of the fluid equation is topologically modified; when the emotional semantic is high arousal, the source term is set to an upward jet source to generate a diffuse wave; when the emotional semantic is low valence, the source term is set to a downward sink source to generate an inward wave. 3) Macroscopic mapping: Map the overall energy envelope of the music in step a to the global viscosity coefficient and wave velocity parameters of the fluid system; Step f: Based on the audio features and emotion vectors in steps a and b, a predefined camera motion mode is selected in the camera control state machine, and the time evolution of camera path, viewpoint, focal length, and environmental parameters such as volumetric light, fog effect, exposure, and color temperature is controlled to keep the camera motion and environmental changes synchronized with the energy and paragraph structure of the music. Step g: The color parameters, fluid physics parameters, fluid mesh state, and camera and environment parameters obtained in steps c–f are input into the fluid rendering framework. The GPU performs fluid solving and scene rendering, outputting a fluid animation sequence aligned with the music timeline (e.g., Figure 4 (As shown).

[0029] The music-driven fluid generation method based on multidimensional semantic mapping proposed in this invention will be described below with reference to specific examples and related figures.

[0030] Example 1: A music-driven fluid generation method based on multidimensional semantic mapping includes the following steps: Step 1: The user inputs music 'q', and a corresponding playback and analysis session is established; Step 2: Input q into the audio analysis module to obtain audio features such as beat BPM(q), spectral energy S(q,t), overall energy envelope E(q,t), and segment label L(q,t), collectively referred to as F(q,t); Step 3: Upload q to the emotion recognition model and the color palette model to obtain the Valence–Arousal emotion vector Emo(q,t) that changes over time and the color palette sequence Pal(q,t) that matches the music style, respectively. Step 4: In the unified parameter space, calculate the parameter vector Param(t) = {C_fluid(t), C_env(t), P_phys(t), P_cam(t), P_env(t)} for the current time slice based on F(q,t), Emo(q,t), and Pal(q,t). Where C_fluid is the fluid color parameter, C_env is the ambient color parameter, P_phys is the fluid physical parameter (such as wave velocity, damping, disturbance source configuration, etc.), P_cam is the camera status parameter, and P_env is the ambient lighting parameter; Step 5: Based on P_phys(t), inject the spectral energy and rhythm information into the shallow water wave equation or the height field fluid model, update the height field and velocity field, and obtain the fluid state State_fluid(t) that changes with time. Step 6: Based on the color, camera and environment parameters in Param(t) and State_fluid(t), render in the fluid scene to obtain the image frame Image(t) corresponding to time t, and output it synchronously with the playback of music q; Step 7: Feed back Param(t) and part of State_fluid(t) to the user interface. The user interacts to obtain the parameter increment ΔParam(t). In the next time slice, update the unified parameter vector: Param(t+Δt) ← Param(t+Δt) ⊕ ΔParam(t) to affect the subsequent fluid simulation and rendering results.

[0031] Example 2: Based on Example 1, but with a difference, taking calm and soothing music as an example, specifically including the following: Step 1: The user inputs a calm and soothing piano piece, q1. The system establishes a playback and analysis session for q1 (e.g., ...). Figure 5 (as shown in A); Step 2: Input q1 into the audio analysis module to obtain F(q1,t), where BPM(q1)≈60–80, the spectral energy is mainly in the mid-low frequency range, the overall energy envelope is smooth, and the number of segments is relatively small and long (e.g., Figure 5 (as shown in B) Step 3: Upload q1 to the emotion recognition model and the color palette model to obtain Emo(q1,t) with high Valence and low Arousal, and a color palette Pal(q1,t) dominated by soft cool colors and light colors (e.g., Figure 5 (as shown in B) Step 4: Calculate Param1(t) = {C_fluid1(t), C_env1(t), P_phys1(t), P_cam1(t), P_env1(t)} based on F(q1,t), Emo(q1,t), and Pal(q1,t). Here, C_fluid1 and C_env1 represent soft, neutral colors; P_phys1 sets a lower wave velocity, higher damping, and lower disturbance intensity; P_cam1 represents a slow orbit or stable viewing angle; and P_env1 represents low-contrast, gently lit environmental parameters (e.g., ...). Figure 5 (as shown in C) Step 5: Based on P_phys1(t), inject the spectral energy and rhythm information into the shallow water wave equation or the height field fluid model to update the height field and velocity field, obtaining the slowly fluctuating and easily calmed fluid state State_fluid1(t) (e.g., Figure 5 (as shown in D); Step Six: Based on the color, camera and environment parameters in Param1(t) and State_fluid1(t), render the image frame Image1(t) in the fluid scene to obtain a water surface or fluid image with soft colors and smooth motion, and output it synchronously with the playback of music q1 (e.g., Figure 5 (as shown in E) Step 7: Feed back Param1(t) and part of State_fluid1(t) to the user interface. Users can slightly adjust the ripple amplitude, color saturation, etc. to obtain the parameter increment ΔParam1(t), which is used to update Param1(t+Δt) for subsequent time slices, so as to achieve personalized adjustment while maintaining the overall style of "calm and soothing".

[0032] Example 3: Based on Example 1, but with a difference, taking music with an exciting and stimulating style as an example, it specifically includes the following: Step 1: Input an exciting dance music track q2 to establish a corresponding playback and analysis session (e.g., ...). Figure 6 (as shown in A); Step 2: Input q2 into the audio analysis module to obtain F(q2,t), where BPM(q2)≈130–150, the spectral energy is significant in the mid-to-high frequencies, the overall energy envelope fluctuates significantly, and the segment structure is complex and highly contrasting (e.g., Figure 6 (as shown in B) Step 3: Upload q2 to the emotion recognition model and the color palette model to obtain Emo(q2,t) with moderately high Valence and significantly high Arousal, as well as the color palette Pal(q2,t) with warm tones, high saturation, and strong contrast (e.g., Figure 6 (as shown in B) Step 4: Calculate Param2(t) = {C_fluid2(t), C_env2(t), P_phys2(t), P_cam2(t), P_env2(t)} based on F(q2,t), Emo(q2,t), and Pal(q2,t), where C_fluid2 and C_env2 represent highly saturated contrasting colors, P_phys2 represents high wave speed, low damping, and high-frequency, high-intensity disturbances, P_cam2 represents rapid orbital movements or drastic changes in perspective, and P_env2 represents high-contrast environmental parameters with drastic changes in lighting. Figure 5 C); Step 5: Based on P_phys2(t), inject strong beats and high-frequency energy into the shallow water wave equation or height field fluid model to frequently trigger disturbance sources and enhance wave propagation speed, resulting in a large-amplitude, rapidly propagating, and difficult-to-quell fluid state, State_fluid2(t). (e.g.) Figure 6 As shown in D), the fluid state changes the corresponding source terms as the musical mood progresses, resulting in various topological modifications (such as...). Figure 6 (as shown in D); Step Six: Render image frame Image2(t) based on Param2(t) and State_fluid2(t), generating a vivid, high-contrast, and dramatically fluctuating fluid image. The camera moves rapidly with the beat, ambient light and shadow are significantly enhanced at climaxes, and the output is synchronized with the playback of music q2 (e.g., Figure 6 (as shown in E) Step 7: Feed back Param2(t) and part of State_fluid2(t) to the user interface. Users can reduce the perturbation density or lower the color contrast to obtain the parameter increment ΔParam2(t) based on their perception. This is used to update Param2(t+Δt) for subsequent time slices, controlling the intensity of visual stimulation while maintaining the overall "excitement" style.

[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating music-driven fluid based on multidimensional semantic mapping, characterized in that, Includes the following steps: S1. Music Analysis and Feature Extraction: Obtain the music input selected by the user, build a multi-resolution audio analysis pipeline locally, perform spectrum, tempo, structure and energy analysis on the music, and extract audio feature sequences that change over time; S2, Emotion Recognition: Based on the audio feature sequence obtained in S1, the emotion recognition model is called to analyze it, and the valence-arousal emotion vector corresponding to each time slice and the emotion category label of the whole song are obtained; at the same time, the recognition results are sent to the color palette generation model to generate the corresponding multi-channel color vector or color palette. S3, Emotion Mapping: In a unified emotion parameter space, the valence-arousal emotion vector obtained in S2 is mapped to the parameters of the fluid's primary color in the HSL color space. The brightness is kept within a preset range, and the hue and saturation are obtained from the valence-arousal emotion vector through a monotonic mapping function to determine the primary color of the current fluid material. S4. Energy Calculation and Visual Allocation: In the CIELCh perceptual uniform color space, based on the color palette obtained in S2, the chromaticity of each color and its CIEDE2000 color difference with neutral gray of the same brightness are calculated, thereby obtaining the visual perception energy of each color; according to the magnitude of the visual perception energy, the colors are mapped to the fluid body, ambient lighting and volumetric fog effect, so that high perception energy colors correspond to high saliency visual elements, and the brightness and saturation of each element are adjusted in real time through a preset energy modulation function; S5. Fluid Simulation and Parameter Mapping: Construct a fluid simulation system based on shallow water wave equations and establish a three-level mapping mechanism from musical characteristics to fluid physical parameters (micro-meso-macro), specifically including: Microscopic mapping: Normalize the beats per minute (BPM) in the audio feature sequence to obtain the wave velocity parameters in the shallow water equation or equivalent fluid model; route the energy of multiple preset frequency bands to spatially separated disturbance sources or interactive bodies, inject disturbances into the fluid height field or velocity field, and update the fluid morphology with the change of spectral energy. Mesoscopic mapping: The source terms of the fluid equation are topologically modified according to the valence-arousal emotion vector; when the emotional semantics are high arousal, the source terms are set as upward jet sources to generate diffusion waves; when the emotional semantics are low valence, the source terms are set as downward convergent sources to generate inward waves. Macroscopic mapping: mapping the overall energy envelope of music to the global viscosity coefficient and wave velocity parameters of a fluid system; S6. Camera Motion and Environment Control: Based on audio characteristics and emotion vectors, select a predefined camera motion mode in the camera control state machine, and control the time evolution of environmental parameters such as camera path, angle of view, focal length, volumetric light, fog effect, exposure, and color temperature, so that the camera motion and environmental changes are synchronized with the energy and paragraph structure of the music. S7. Real-time Solving and Rendering Output: Input the color parameters, fluid physical parameters, fluid mesh state, camera and environment parameters obtained from S3 to S6 into the fluid rendering framework, and perform fluid solving and scene rendering through the graphics processing unit (GPU). Finally, input a fluid animation sequence aligned with the music timeline.

2. The music-driven fluid generation method based on multidimensional semantic mapping according to claim 1, characterized in that, The emotion recognition model and the color palette generation model described in S2 are invoked in parallel through a unified reasoning process, so that the main color of the fluid is directly determined by the valence-arousal emotion vector. The environmental and special effects colors are driven by both the results of palette analysis and the allocation of visual perception energy, thereby simultaneously obtaining the fluid primary color, environmental color, and focal color in a single music analysis to achieve consistent emotional expression across color, fluid, and camera subsystems.

3. The music-driven fluid generation method based on multidimensional semantic mapping according to claim 1, characterized in that, The formula for the mapping operation described in S3 is as follows: in, Hue and H Indicates hue; Saturation and S Indicates saturation; V Indicates valence; A Indicates wakefulness; F It is a linear mapping function; H wrap , H end 、H max , H min The values ​​of each point representing the target hue; V min , V max , V mid Indicates the range of values ​​for source valence; x , x 1, x 2 represents the input value and the binary reference mapping source term, respectively; y 1, y 2 represents the binary reference mapping target item.

4. The music-driven fluid generation method based on multidimensional semantic mapping according to claim 1, characterized in that, The method for acquiring visual perception energy described in S4 specifically includes: Use the CIELCh color space to obtain the chromaticity of each color. C *( C i For each color C i Calculate its neutral gray, which is the same as the perceived brightness. G i The CIEDE2000 color difference between them is then used to define the visual perception energy based on weights, as shown in the specific formula. as follows: in, Represents visual perception energy; C i Representing each color; C *( C i () indicates the chroma of each color; Gi indicates neutral gray; The CIEDE2000 color difference represents the difference between neutral grays with the same perceived brightness for each color. w c This represents the weight value of each color chromaticity in visual perception energy; w d The weight of the CIEDE2000 color difference between neutral grays with the same perceived brightness in visual perception energy; High-energy colors are assigned to the fluid subject and foreground highlights, while low-energy colors are assigned to the background and sky elements. The brightness and saturation of these elements are constrained using an energy-based modulation function, ensuring that the overall image maintains visual contrast and a unified emotional tone across different music clips. The energy modulation function is defined as follows: in, A ( e () indicates the assigned base color HSL value; M ={ M energy , M beat }, M energy This represents the spectral energy of music. M beat Indicates the number of beats in the music; The saturation and brightness of each element are modulated in real time using an energy modulation function to satisfy: in, Indicates the brightness of the HSL color of the highlighted element; The saturation of the HSL color representing the highlighted element. Indicates the saturation of the HSL colors for the background and sky elements; This is the sensitivity parameter.

5. The music-driven fluid generation method based on multidimensional semantic mapping according to claim 1, characterized in that, The construction of a fluid simulation system based on shallow water wave equations described in S5 further includes the following: We employ shallow water equations or equivalent two-dimensional height field fluid models, and iteratively update the height and velocity fields using fragment shaders and computation shaders on the GPU. The resulting wave velocity parameters include: converting the beat count (BPM) into shallow water wave propagation speed and converting the pitch into the fluid viscosity coefficient; The resulting frequency bands include low-frequency, mid-frequency, and high-frequency bands, which are used to drive: the generation and propagation of large-scale waves, the formation of mid-scale ripples and vortices, and small-scale high-frequency details or particle spray effects, respectively. By varying the location and intensity of the disturbance source, a fluid response to music is achieved through scene rendering.

6. The music-driven fluid generation method based on multidimensional semantic mapping according to claim 1, characterized in that, The camera control state machine described in S6 includes at least the following states: Orbital cruise mode: The camera moves in an orbit around the center of the scene at a constant or gradually changing angular velocity; Diving motion mode: The camera rapidly approaches the fluid surface along a preset curve, and the depth of field changes are superimposed; Slow track mode: The camera moves or pans at a low speed, suitable for low-energy music clips; Stable stillness: The camera's field of view remains basically still, with only slight shaking or breathing-like zooming; The state machine performs state transitions based on music energy, beat phase, and segment labels, and performs smooth interpolation of camera trajectory and environmental parameters during switching to reduce visual abrupt changes and ensure synchronization with the music timeline.

7. A music-driven fluid generation system based on multidimensional semantic mapping, applying the method of any one of claims 1-6, characterized in that, include: Music Analysis and Feature Extraction Module: Used to acquire music input selected by the user, build a multi-resolution audio analysis pipeline locally, perform spectrum, tempo, structure and energy analysis on the music, and extract audio feature sequences that change over time; Emotion Recognition Module: Based on the obtained audio feature sequence, the emotion recognition model is called to analyze it, and the valence-arousal emotion vector corresponding to each time slice and the emotion category label of the whole song are obtained. At the same time, the recognition results are sent to the color palette generation model to generate the corresponding multi-channel color vector or color palette. Emotion Mapping Module: In a unified emotion parameter space, the obtained valence-arousal emotion vector is mapped to the parameters of the fluid's primary color in the HSL color space, where the brightness is kept within a preset range, and the hue and saturation are obtained from the valence-arousal emotion vector through a monotonic mapping function to determine the primary color of the current fluid material; Energy Calculation and Visual Allocation Module: In the CIELCh perceptual uniform color space, based on the obtained color palette, the chromaticity of each color and its CIEDE2000 color difference with neutral gray of the same brightness are calculated, thereby obtaining the visual perception energy of each color; according to the magnitude of the visual perception energy, the colors are mapped to the fluid body, ambient lighting and volumetric fog effect, so that high perception energy colors correspond to high saliency visual elements, and the brightness and saturation of each element are adjusted in real time through a preset energy modulation function; Fluid Simulation and Parameter Mapping Module: Constructs a fluid simulation system based on shallow water wave equations and establishes a three-level mapping mechanism from musical characteristics to fluid physical parameters (micro-meso-macro), specifically including: Microscopic mapping: Normalize the beats per minute (BPM) in the audio feature sequence to obtain the wave velocity parameters in the shallow water equation or equivalent fluid model; route the energy of multiple preset frequency bands to spatially separated disturbance sources or interactive bodies, inject disturbances into the fluid height field or velocity field, and update the fluid morphology with the change of spectral energy. Mesoscopic mapping: The source terms of the fluid equation are topologically modified according to the valence-arousal emotion vector; when the emotional semantics are high arousal, the source terms are set as upward jet sources to generate diffusion waves; when the emotional semantics are low valence, the source terms are set as downward convergent sources to generate inward waves. Macroscopic mapping: mapping the overall energy envelope of music to the global viscosity coefficient and wave velocity parameters of a fluid system; Camera motion and environment control module: Based on audio features and emotion vectors, it selects a predefined camera motion mode in the camera control state machine and controls the time evolution of environmental parameters such as camera path, angle of view, focal length, volumetric light, fog effect, exposure, and color temperature, so that the camera motion and environmental changes are synchronized with the energy and paragraph structure of the music. Real-time solving and rendering output module: Input the obtained color parameters, fluid physics parameters, fluid mesh state, camera and environment parameters into the fluid rendering framework, and perform fluid solving and scene rendering through the graphics processing unit (GPU), and finally input a fluid animation sequence aligned with the music timeline.