Visual content generation system and method based on music feature recognition

The sensor and audio data acquisition device obtains physiological and sound feature data in real time, and uses the smart chip to generate light and video content that matches the music and user status, solving the problems of weak correlation between visual content and music and independent control, realizing audio-visual linkage, and improving the adjustment of physical and mental effects and safety.

CN120378708APending Publication Date: 2025-07-25SHANGHAI ARTSBANG CULTURE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510431962.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing visual content generation method has weak correlation with the target music, and cannot be adjusted in real time based on physiological data, and the lighting and video control are independent, which poses safety risks.

Method used

The sensor and audio data acquisition device obtains physiological and sound feature data in real time, and uses intelligent chips to comprehensively analyze it to generate light and video content that matches the music and user status, achieving audio-visual linkage.

Benefits of technology

Visual background content is generated based on the client's physiological data, and the heart rate and breathing frequency are adjusted too fast, so that the client is calm and relaxed, improving the adjustment effect and safety of the visual content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378708A_ABST
    Figure CN120378708A_ABST
Patent Text Reader

Abstract

The invention provides a visual content generation system and method based on music feature recognition. The system comprises a sensor, an audio data acquisition device, an intelligent chip, a lighting device and a projection device. The sensor is connected with the intelligent chip to obtain physiological data such as heart rate and respiratory rate of a user in real time; the audio data acquisition device is connected with the intelligent chip, and obtains sound feature data of each piece of target music in real time; the intelligent chip comprehensively analyzes the music data and the physiological data, generates light and video content matched with the current music and the user state, and outputs the light and video content through the light device and the projection device. According to the method, intervention can be performed according to the real body state of the visitor, the visual background content is generated based on the real-time physiological data of the visitor, different solutions are matched according to different body states reflected by the physiological data of the visitor, and the too fast heart rate and respiratory rate are adjusted through visual traction, so that the visitor is calm and relaxed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target music. Specifically, it relates to a visual content generation system and method based on music feature recognition. Background Art

[0002] Currently, the existing visual content generation methods mainly still rely on traditional forms, and there is no specific visual content generation form for target music. The traditional visual content generation methods have the following bottlenecks:

[0003] The relevance between visual content and target music is weak. Currently, the visual content used to match target music on the market mostly takes forms such as "mandala" and "natural scenery", and the relevance of these contents to the music itself to be presented is weak, and there may be infringement risks. The weak relevance to music is further reflected in that there is no systematic relevance between the elements in visual content such as color, shape, light and shade, movement and stillness and the music elements (such as pitch, timbre, loudness, rhythm, melody, etc.), and the role of adjusting the body and mind brought by hearing cannot be further strengthened through vision.

[0004] Currently, the existing traditional audio-visual intervention content still stays in the category of simple receptive listening and viewing, without involving real-time visual content adjustment and cannot be generated based on real-time physiological data.

[0005] Traditional visual content creation generally comes from the ideas of the producer himself and has nothing to do with the specific function of adjusting the body and mind. In actual use, especially in clinical treatment, if patients with emotional or mental disorders are allowed to watch inappropriate visual content, it is entirely possible to cause serious consequences.

[0006] Currently, the existing visual content for adjusting the body and mind cannot be uniformly controlled. Lights and videos generally use two independent systems for control, and videos and lights cannot be controlled in the same system. Summary of the Invention

[0007] An embodiment of the present application provides a visual content generation system based on music feature recognition, including: a sensor, an audio data acquisition device, an intelligent chip, a lighting device, and a projection device; the sensor is connected to the intelligent chip and is used to obtain physiological data such as the heart rate and breathing rate of the user in real time; the audio data acquisition device is connected to the intelligent chip and is used to obtain sound feature data such as the pitch, loudness, and timbre of each target music in real time; the intelligent chip comprehensively analyzes the music data and physiological data, generates lighting and video content that matches the current music and the user's state, and outputs it through the lighting device and the projection device.

[0008] Wherein, the intelligent chip includes: a data coding set processing module, a lighting generation module, and a video generation module; the data coding set processing module is connected to the video generation module.

[0009] Among them, the sensor is used to collect real-time physiological data, the real-time volume and vibration intensity of the user, and perform data preprocessing, convert it into the required data format, and input it into the data coding set processing module.

[0010] Among them, the audio data collection device is used to collect sound feature data such as the pitch, loudness, and timbre of each target music, convert it into the required data format, and input it into the data coding set processing module.

[0011] Among them, the data coding set processing module includes a music coding unit and a physiological data coding unit. The data coding set processing module receives the data, performs music coding processing and physiological data coding processing, and sends the processed data to the video generation module to generate a video stream and play it through a projection device.

[0012] Among them, the light generation module is used to generate a light signal according to the touch input and the light mapping rule. The light signal matches the video content accordingly and is played through the light device.

[0013] In a second aspect, the present application provides a method for generating visual content based on music feature recognition, including: the sensor obtains the physiological data of the user in real time; the audio data collection device obtains the sound feature data of each target music in real time; the intelligent chip comprehensively analyzes the sound feature data and the physiological data, generates lights and video content that match the current music and the user's state, and outputs them through the light device and the projection device.

[0014] Among them, it includes: the sensor collects real-time physiological data, the real-time volume and vibration intensity of the user, and performs data preprocessing, converts it into the required data format, and inputs it into the data coding set processing module.

[0015] Among them, it includes: the audio data collection device collects the sound feature data of each target music, converts it into the required data format, and inputs it into the data coding set processing module, and the sound feature data includes: pitch, loudness, and timbre.

[0016] Among them, it includes: the data coding set processing module receives the data, performs music coding processing and physiological data coding processing, and sends the processed data to the video generation module to generate a video stream and play it through the projection device.

[0017] The visual content generation system and method based on music feature recognition in the embodiments of the present application have the following beneficial effects:

[0018] The present invention can intervene according to the real physical state of the visitor. The visual background content is generated based on the real-time physiological data of the visitor, and different solutions are matched according to the different physical states reflected by the physiological data of the visitor. Through visual traction, the too-fast heart rate and breathing frequency are adjusted to calm and relax the visitor. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 FIG. is a schematic structural diagram of a visual content generation system based on music feature recognition according to an embodiment of the present application;

[0020] Figure 2 FIG. is a schematic diagram of a visual content generation process according to an embodiment of the present application;

[0021] Figure 3 FIG. is a schematic structural diagram of a data encoding set processing module in a visual content generation system based on music feature recognition according to an embodiment of the present application;

[0022] Figure 4 FIG. is a flowchart of a method for generating visual content based on music feature recognition. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The present application will be further described below with reference to the drawings and embodiments.

[0024] In the following description, the terms "first" and "second" are only for the purpose of description and cannot be construed as indicating or implying relative importance. The following description provides multiple embodiments of the present invention, and different embodiments can be replaced or combined. Therefore, the present application can also be considered to include all possible combinations of the same and / or different embodiments described. Thus, if one embodiment includes features A, B, and C, and another embodiment includes features B and D, then the present application should also be considered to include embodiments containing one or more all other possible combinations of features A, B, C, and D, even though such an embodiment may not be explicitly described in the following content.

[0025] Embodiment 1

[0026] The main visual content generated based on the target music in the present application is linked to the key elements of the target music. Elements such as the hue, lightness, shape, and shape change of the main visual content are associated with the loudness, music speed, and attack speed of the music with the function of adjusting the body and mind, and the corresponding mechanism is set, and the corresponding function of adjusting the body and mind is realized through the matching of the light spectrum and the visual content.

[0027] Such as Figure 1-2As shown in the figure, a visual content generation system based on music feature recognition in this application includes: a sensor 10, an audio data acquisition device 20, an intelligent chip 30, a lighting device 40, and a projection device 50; the sensor 10 is connected to the intelligent chip 30 and is used to obtain physiological data such as the user's heart rate and breathing rate in real time; the audio data acquisition device 20 is connected to the intelligent chip 30 and is used to obtain sound feature data such as the pitch, loudness, and timbre of each target music in real time; the intelligent chip 30 comprehensively analyzes the music data and physiological data, generates lighting and video content that matches the current music and the user's state, and outputs it through the lighting device 40 and the projection device 50. The intelligent chip 30 includes: a data encoding set processing module 301, a lighting generation module 302, and a video generation module 303; the data encoding set processing module 301 is connected to the video generation module 303.

[0028] The sensor 10 collects the user's real-time physiological data (such as heart rate and breathing rate) and the user's real-time volume and vibration intensity preferences, preprocesses the data, converts it into the data format required for generation, and inputs it into the data encoding model 30; the audio data acquisition device 20 collects the sound feature data of a complete target music piece, and the sound feature data includes pitch, loudness, timbre, speed, attack speed, etc. The collected data is preprocessed, converted into the data format required for generation, and input into the data encoding model 30;

[0029] As Figure 3 shown in the figure, the data encoding set processing module 301 includes a music encoding unit and a physiological data encoding unit. The data encoding set processing module 301 receives the converted physiological data and sound feature data, and performs music encoding processing and physiological data encoding processing respectively;

[0030] The music encoding unit comprehensively collects the sound feature data of the input music piece, including but not limited to key parameters such as pitch, loudness, timbre, speed, and attack speed. These feature data are divided into discrete signal sequences s_1, s_2, s_3,..., s_n according to time slices. Among them, the composition of each discrete signal vector s_i includes:

[0031] · Pitch: Describes the frequency characteristics of the music in the current time slice.

[0032] · Loudness: Measures the volume intensity in this time slice.

[0033] · Timbre: Characterizes the unique attributes of the sound through a set of spectral features (such as MFCC, spectral centroid, spectral flatness, etc.).

[0034] · Tempo: The rhythm information of the music, dynamically estimated in combination with the timing characteristics.

[0035] · Attack Speed: Characterizes the dynamic change rate at the start of a sound.

[0036] · Other supplementary features: Such as zero-crossing rate, beat intensity, harmonic spectrum, etc., provide a more fine-grained description of the music.

[0037] After preprocessing the above feature data, normalization, noise reduction, and completion are performed to ensure data consistency and model robustness. Subsequently, this time series data is input into an encoder-decoder network based on Attention and Transformer.

[0038] The encoder-decoder network adopts an autoencoder architecture, consisting of an encoder and a decoder.

[0039] The encoder part includes:

[0040] · Input layer: Receives time series signals s_1, s_2, \dots, s_n, and each signal vector is transformed into a high-dimensional space representation through an embedding layer.

[0041] · Multi-Head Attention: Captures the global dependencies between different time slices through parallel computation of multiple attention heads.

[0042] · Feedforward Neural Network (FFN): Used for non-linear feature mapping to enhance feature expression ability.

[0043] · Layer Normalization and Residual Connections: Improve the stability and convergence speed of the model.

[0044] The output of the encoder is a high-dimensional vector representing the compressed feature embedding of the music.

[0045] The decoder part includes:

[0046] · Decoding input: Takes the feature embedding output by the encoder as input, combines it with the positional encoding of the time step to capture the sequential information in the time series.

[0047] · Cross Attention: During the decoding process, combines the output of the encoder with the decoder's own context information to generate a more accurate sequence representation.

[0048] · Output layer: Decodes the reconstructed signal corresponding to the input time slice through a fully connected layer and an activation function.

[0049] The embedded representation generated by the encoder is a high-dimensional feature representation of the music piece, which can describe the overall style, dynamic changes, and local details of the music. This embedding can be widely used in the following scenarios:

[0050] 1. Music similarity analysis: Recommends and classifies music pieces by comparing the similarity of the embedded vectors.

[0051] 2. Generative model training: Uses the embedding as a conditional input to generate music segments with a similar style.

[0052] 3. Retrieval system: Implements efficient music search and indexing based on the embedded vectors.

[0053] 4. Sentiment analysis: Studies the emotional impact of music on listeners by mapping the embedding into an emotional space.

[0054] This architecture combines the cutting-edge technologies of sequence modeling, can effectively capture the temporal characteristics and complex structures of music, and provides an efficient and accurate solution for music processing.

[0055] During the operation of the system, the physiological data encoding unit collects the user's physiological data and interaction preference data in real time, including but not limited to the following:

[0056] Physiological data:

[0057] · Heart Rate (HR): Uses wearable devices (such as heart rate belts or smart watches) to record the user's heart rate data in real time.

[0058] · Breathing Rate (BR): Collects the user's breathing rate through a chest sensor or a wearable device to characterize the user's current breathing rhythm.

[0059] · Galvanic Skin Response (GSR): Used to reflect the user's emotional arousal level.

[0060] · Blood Oxygen Saturation (SpO2): Monitors the oxygen content in the user's blood through an optical sensor.

[0061] · Body Temperature: Used to comprehensively analyze the user's health status and mood.

[0062] User preference data:

[0063] · Real-time volume preference: Dynamically changes according to the user's operation records of volume adjustment, capturing the user's volume needs in the current context.

[0064] · Vibration intensity preference: Real-time obtains the user's preference by adjusting the intensity settings of the vibration device, reflecting the user's need for tactile stimulation.

[0065] · Interaction context information: Records the user's current usage scenarios (such as resting, exercising, working) and the ambient sound interference situation.

[0066] The above data is divided into discrete signal sequences by time slices, and each signal vector includes the following:

[0067] 1. Physiological data: Physiological characteristics such as heart rate, respiratory rate, skin electrical activity, blood oxygen saturation, body temperature, etc.

[0068] 2. User preferences: Volume and vibration intensity settings at the current moment.

[0069] 3. Environmental variables: Optional features such as ambient sound intensity, room temperature, etc.

[0070] After these signals are preprocessed (such as normalization, noise reduction, and missing value filling), they are input into a Transformer-based encoder-decoder model for feature extraction and embedding generation.

[0071] Each signal is a multi-modal feature vector, including the following parts:

[0072] · Physiological signal embedding: Maps physiological data such as heart rate and respiratory rate to a unified high-dimensional feature space.

[0073] · User preference embedding: Vector representation of volume and vibration intensity preferences.

[0074] · Context feature embedding: Includes environmental variables and usage context encoding.

[0075] Encoder part:

[0076] 1. Embedding layer: Each modal feature is mapped to a high-dimensional vector representation through a dedicated embedding layer, and the sequential relationship of the time series is captured through positional encoding.

[0077] 2. Multi-modal fusion: Uses the multi-head attention mechanism to model the interaction of different modal features, capturing the correlation between physiological signals and user preferences.

[0078] 3. Temporal modeling: Extracts the temporal characteristics and global dependencies of the signal sequence by stacking multiple layers of Transformer encoders.

[0079] Decoder part:

[0080] 1. Decoding Input: The embedded features output by the encoder are combined with the positional encoding as the input.

[0081] 2. Multi-layer Decoding: Use the multi-head attention mechanism and the feed-forward network to decode the encoded features, generating prediction results or optimization suggestions for real-time feedback.

[0082] 3. Output Layer: Generate the embedded representation of the multimodal signal for further analysis or real-time interactive control.

[0083] The embedded vector generated by the encoder can be used as the multimodal feature expression of the user's state and preferences, and is applied to the following scenarios:

[0084] 1. Real-time Music Adjustment: Dynamically adjust the rhythm, volume, and sound effect style of the music according to the user's physiological state and preferences.

[0085] 2. Vibration and Tactile Feedback Optimization: Achieve personalized vibration intensity regulation based on the user's physiological signals to enhance the immersive experience.

[0086] 3. Emotional State Analysis: Use the embedded vector to infer the user's emotional state and provide contextual audio and tactile feedback.

[0087] 4. Health Monitoring and Intervention: Combine the embedded representation to design targeted health reminders or relaxation programs.

[0088] This architecture realizes the accurate perception and dynamic adaptation of the user's state by collecting and modeling the user's physiological data and preferences in real time, and is applicable to fields such as personalized audio regulation, health management, and immersive experience.

[0089] The data processed by the music encoding unit and the physiological data encoding unit is sent to the video generation module 303 to generate a video stream, generate video frames one by one, and play them through the projection device 50.

[0090] The lighting generation module 302 is used to generate lighting signals according to the touch input and the lighting mapping rules. The lighting signals are matched with the video content and played through the lighting device 40.

[0091] The visual content generation system based on music feature recognition in this application includes lighting signal output and video signal output. The lighting spectrum and the visual content are matched accordingly. In terms of color, yellow usually represents pleasure, and blue represents calm; for another example, in terms of shape, sharp represents tension, and smooth represents peace.

[0092] 1. Lighting Signal Output

[0093] 1.1 Spectrum Selection: Select a specific wavelength spectrum according to the function of adjusting the body and mind.

[0094] 1.2 Lighting signal generation: Generate lighting signals according to touch inputs and lighting mapping rules.

[0095] 2. Video stream output

[0096] 2.1 Data acquisition:

[0097] a. Completely acquire audio data.

[0098] b. Dynamically acquire physiological data in real time.

[0099] 2.2 Data processing:

[0100] a. Input the acquired data into their respective encoding units for processing.

[0101] b. Generate a set of video frames using the processed data.

[0102] 2.3 Video generation rules:

[0103] a. Color of the main object:

[0104] (1) Hue: Convert the selected spectrum in the lighting rules into the corresponding hue, and set the saturation to 100%.

[0105] (2) Lightness:

[0106] 1. When the average loudness is below 30 dB, the lightness is 50%.

[0107] 2. When the average loudness is between 30 - 50 dB, the lightness is 70%.

[0108] 3. When the average loudness is above 50 dB, the lightness is 100%.

[0109] b. Shape of the main object:

[0110] (1) Shape generation: Mutually map the average attack speed of the sound head in the timbre and the roundness of the polygon. As a fully enclosed abstract graph, the polygon can carry rich imagery and symbolic meanings. In the visual presentation of the target music, the polygon can convey the connotation and emotion of the music through continuous changes and expansions. The slower the average attack speed of the sound head, the rounder the polygon. The formula is as follows:

[0111]

[0112] R is the roundness (unit: degree), and t is the average attack speed of the sound head (unit: millisecond).

[0113] (2) Shape change: The speeds at which the saturation and lightness of the shape become 0 are both one-tenth of the music speed.

[0114] (3) Shape generation frequency and motion change speed: Both are one-tenth of the music speed.

[0115] c. Background content:

[0116] (1) Color: Corresponding the hue of the main object to the hue of the background content.

[0117] (2) Hue change:

[0118] 1. When the current heart rate ≤ 70 beats per minute or the respiratory rate ≤ 15 breaths per minute, the hue remains unchanged.

[0119] 2. When the current heart rate is between 70 - 90 beats per minute or the respiratory rate is between 15 - 19 breaths per minute, the hue changes 12 times per minute within the range of plus or minus 15 degrees.

[0120] 3. When the current heart rate ≥ 90 beats per minute or the respiratory rate ≥ 19 breaths per minute, the hue uniformly transitions from 15 times per minute to 12 times per minute within the range of plus or minus 15 degrees.

[0121] 3. Brightness:

[0122] 3.1. When the current heart rate ≤ 70 beats per minute or the respiratory rate ≤ 15 breaths per minute, the brightness is 0%.

[0123] 3.2 When the current heart rate is between 70 - 90 beats per minute or the respiratory rate is between 15 - 19 breaths per minute, the brightness uniformly transitions from 40% to 0%.

[0124] 3.3 When the current heart rate ≥ 90 beats per minute or the respiratory rate ≥ 19 breaths per minute, the brightness uniformly transitions from 60% to 0%.

[0125] 4. Saturation:

[0126] 4.1 When the current heart rate ≤ 70 beats per minute or the respiratory rate ≤ 15 breaths per minute, the saturation is 0%.

[0127] 4.2 When the current heart rate is between 70 - 90 beats per minute or the respiratory rate is between 15 - 19 breaths per minute, the saturation uniformly transitions from 40% to 0%.

[0128] 4.3 When the current heart rate ≥ 90 beats per minute or the respiratory rate ≥ 19 breaths per minute, the saturation uniformly transitions from 60% to 0%.

[0129] In this application, by analyzing the sound characteristics of the target music and the physiological data of the user, visual content that matches it is automatically generated, and the light spectrum matches the visual content accordingly, achieving visual-auditory dual-sensory linkage to enhance the effect of adjusting the body and mind.

[0130] In this application, the visual background content is generated based on the real-time physiological data of the visitor, can be intervened according to the real physical state of the visitor, and different solutions are matched according to different physical states reflected by the physiological data of the visitor. Through visual traction, the too-fast heart rate and breathing frequency are adjusted to make the visitor calm and relaxed.

[0131] As Figure 4 shown, this application also provides a method for generating visual content based on music feature recognition, including: S101, a sensor obtains the physiological data of the user in real time; S103, an audio data acquisition device obtains the sound feature data of each target music in real time; S105, an intelligent chip comprehensively analyzes the sound feature data and the physiological data, generates lighting and video content that matches the current music and the user's state, and outputs it through a lighting device and a projection device.

[0132] In some embodiments, the method for generating visual content based on music feature recognition in this application includes: the sensor collects real-time physiological data, the real-time volume and vibration intensity of the user, and performs data preprocessing, converts it into the required data format, and inputs it into the data coding set processing module.

[0133] In some embodiments, the method for generating visual content based on music feature recognition in this application includes: an audio data acquisition device collects the sound feature data of each target music, converts it into the required data format, and inputs it into the data coding set processing module, and the sound feature data includes: pitch, loudness and timbre.

[0134] In some embodiments, the method for generating visual content based on music feature recognition in this application includes: the data coding set processing module receives the data, performs music coding processing and physiological data coding processing, and sends the processed data to the video generation module to generate a video stream and play it through a projection device.

[0135] In this application, the embodiments of the method for generating visual content based on music feature recognition are basically similar to the embodiments of the system for generating visual content based on music feature recognition. For the related parts, please refer to the introduction of the embodiments of the system for generating visual content based on music feature recognition.

[0136] The above introduction is only the preferred embodiments of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A visual content generation system based on music feature recognition, characterized in that, Including: A sensor, an audio data acquisition device, an intelligent chip, a lighting device, and a projection device; The sensor is connected to the intelligent chip and is used to obtain the physiological data of the user in real time; The audio data acquisition device is connected to the intelligent chip and is used to obtain the sound feature data of each target music in real time; The intelligent chip comprehensively analyzes the sound feature data and the physiological data, generates lighting and video content that matches the current music and the user's state, and outputs it through the lighting device and the projection device.

2. The visual content generation system based on music feature recognition according to claim 1, wherein The intelligent chip includes: a data coding set processing module, a lighting generation module, and a video generation module; the data coding set processing module is connected to the video generation module.

3. The visual content generation system based on music feature recognition according to any one of claims 1-2, wherein The sensor is used to collect real-time physiological data, the real-time volume and vibration intensity of the user, and perform data preprocessing, convert it into the required data format, and input it into the data coding set processing module.

4. The visual content generation system based on music feature recognition according to any one of claims 1-2, characterized in that The audio data acquisition device is used to collect the sound feature data of each target music, convert it into the required data format, and input it into the data coding set processing module. The sound feature data includes: pitch, loudness, and timbre.

5. The visual content generation system based on music feature recognition according to any one of claims 1-2, characterized in that, The data coding set processing module includes a music coding unit and a physiological data coding unit. The data coding set processing module receives the data, performs music coding processing and physiological data coding processing, and sends the processed data to the video generation module to generate a video stream and play it through the projection device.

6. The visual content generation system based on music feature recognition according to any one of claims 1-2, wherein The lighting generation module is used to generate a lighting signal according to the touch input and the lighting mapping rule. The lighting signal matches the video content accordingly and is played through the lighting device.

7. A visual content generation method based on music feature recognition, characterized in that, Including: The sensor obtains the physiological data of the user in real time; The audio data acquisition device obtains the sound feature data of each target music in real time; The intelligent chip comprehensively analyzes the sound feature data and the physiological data, generates lighting and video content that matches the current music and the user's state, and outputs it through the lighting device and the projection device.

8. The method for generating visual content based on music feature recognition according to claim 7, wherein Including: The sensor collects real-time physiological data, the real-time volume and vibration intensity of the user, and performs data preprocessing, convert it into the required data format, and input it into the data coding set processing module.

9. The method for generating visual content based on music feature recognition according to claim 8, wherein Including: The audio data acquisition device collects the sound feature data of each target music, convert it into the required data format, and input it into the data coding set processing module. The sound feature data includes: pitch, loudness, and timbre.

10. The visual content generation method based on music feature recognition according to claim 9, wherein Including: The data coding set processing module receives the data, performs music coding processing and physiological data coding processing, and sends the processed data to the video generation module to generate a video stream and play it through the projection device.