Music signal feature extraction and light dynamic mapping control method and system
By employing technologies such as deep learning and spectrum analysis, we have achieved accurate identification of music signal characteristics and dynamic mapping control of lighting, solving the problem of poor synchronization between lighting and music in existing technologies and enhancing the audiovisual experience and live interactivity.
Patent Information
- Application Number
- CN202511083043.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for music signal feature extraction and dynamic lighting mapping control suffer from low beat recognition accuracy, limited style and emotion recognition range, incomplete drum beat recognition, and poor segmentation and special event detection capabilities. They also fail to effectively support color scheme switching for popular songs, resulting in poor synchronization between lighting performance and music.
By employing technologies such as deep learning, convolutional neural networks, spectrum analysis, and clustering algorithms, and through music beat recognition, style and emotion recognition, drum beat feature analysis, music segmentation, and special event detection, light driving files are generated to achieve dynamic mapping control between music signal features and lights.
It achieves accurate recognition of key features such as musical beat, style, and emotion, can identify drum details such as bass drum and snare drum, automatically segments music, detects special events, improves the synchronization of lighting and voice, enhances visual expressiveness and audience immersion, and improves the quality of audiovisual experience.
Smart Images

Figure CN120895055A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of music signal processing and light control, specifically to a music signal feature extraction and light dynamic mapping control method. BACKGROUND
[0002] In the prior art, the synchronization interaction between music and light mainly relies on preset programs or manual adjustment, which has significant limitations in processing music signal feature extraction and mapping control. Specifically, the current music signal feature extraction is often not accurate enough, especially for the recognition of key features such as music tempo, style, and emotional changes. Traditional technical means are difficult to capture subtle and complex music elements, resulting in poor synchronization effect between light performance and music, and failing to fully exhibit the rhythm and emotional fluctuations of music.
[0003] The prior art has many deficiencies in music signal feature extraction and light dynamic mapping control, including but not limited to low tempo recognition accuracy, limited style and emotion recognition range, imperfect drum point recognition, poor segmentation and special event detection capability, and lack of effective support for popular song color system switching, all of which urgently need a more advanced and accurate technical solution to break through. SUMMARY
[0004] The present application aims to solve the above problems of the prior art. A music signal feature extraction and light dynamic mapping control method and system are proposed. The technical solution of the present application is as follows: A music signal feature extraction and light dynamic mapping control method, comprising the following steps: Step 1, model-based music tempo recognition, including the steps of identifying measures, tempo points, measure numbers, time signatures, and calculating tempo; Step 2, music style and emotion recognition, including the steps of classifying song styles and identifying emotional tendencies, the music style and emotion recognition being based on the results of the music tempo recognition; Step 3, drum feature analysis and recognition, including the identification of bass drum and military drum hitting points, the drum feature analysis being based on the results of the music tempo recognition; Step 4, music segmentation recognition, based on the results of the music style recognition and the drum feature analysis; Step 5, special event detection, including the identification of vocal high notes, accompaniment silence, special drum points at the beginning, and fading at the end, the special event detection being based on the results of the music tempo recognition, the music style recognition, the drum feature analysis, and the music segmentation recognition; Step 6, popular song color system switching recognition and execution, the color system switching recognition being based on the results of the music style recognition and the emotion recognition. Step 7, the cloud generates an output corresponding sound feature driving file according to the input audio data using an AI model, and generates control instructions according to the file and the material segment file made by the lighting designer; the generation step is based on the results of the beat recognition, the music style and emotion recognition, the drum point feature analysis, the music segment recognition, the special event detection and the color system switching recognition and execution, to realize dynamic mapping control of music signal features and light.
[0005] A system using any of the methods, comprising: a feature recognition module, comprising: a beat recognizer for identifying measures, beat time points, measure numbers, time signatures and calculating tempos; a style recognizer for classifying song styles and identifying emotional tendencies; a drum point analyzer for identifying bass drum and snare drum hitting points; a segment recognizer for identifying music segments based on the results of the music style recognition; and a special event detector for detecting high-pitched vocals, accompaniment silences, special drum points at the beginning and fading at the end; a light control unit for real-time control of light equipment according to the output of the feature recognition module; a storage unit for storing the music feature data and the light driving file; the light driving file is a sound feature driving file generated by the cloud according to the input audio data using an AI model, and control instructions are generated according to the file and the material segment file made by the lighting designer; the generation step is based on the results of the beat recognition, the music style and emotion recognition, the drum point feature analysis, the music segment recognition, the special event detection and the color system switching recognition and execution, to realize dynamic mapping control of music signal features and light.
[0006] a central processing unit (CPU) for coordinating system workflow, processing the output of the feature recognition module and generating light control signals, the processing of the central processing unit is based on the results of the beat recognition, the music style recognition, the drum point analysis, the segment recognition and the special event detection.
[0007] The advantages and benefits of the present application are as follows: The application realizes accurate recognition of music signal characteristics and dynamic mapping control of light through various advanced technologies such as deep learning, convolutional neural networks, spectral analysis, clustering algorithms, etc. It can not only accurately capture key features such as the rhythm, style, and mood of music, but also identify drum details such as bass drums and military drums, automatically segment music, detect special music events such as high-pitched vocals and accompaniment silence, and switch color systems to enhance visual performance. The real-time data processing capability of the central processing unit and the cache mechanism of the storage unit together ensure the instantaneity and smoothness of the system response, while the multi-channel output capability of the light control unit provides a solid foundation for the diversity and hierarchy of light effects. The addition of the vocal separation algorithm makes the synchronization of light and vocals more precise, greatly improving the emotional impact of the performance and the audience's sense of immersion. Overall, the scheme significantly enhances the interactivity of music and light, improves the quality of audio-visual experience, and is suitable for various live performances, stage performances, and entertainment venues, effectively stimulating audience emotions and enhancing the visual impact and artistic expression of performances. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 is a preferred embodiment of the application music signal feature extraction and dynamic mapping control method flow chart provided by the application; Figure 2 is a preferred embodiment of the application music signal feature extraction and dynamic mapping control system schematic diagram provided by the application. DETAILED DESCRIPTION
[0009] The technical solutions in the embodiments of the application will be described in detail below with reference to the drawings of the embodiments of the application. The described embodiments are only a part of the embodiments of the application.
[0010] The technical solution of the application to solve the above technical problems is: As shown in Figure 1 A music signal feature extraction and dynamic mapping control method, comprising the following steps: Step 1, model-based music beat recognition, including the steps of identifying measures, beat time points, measure numbers, time signatures, and calculating tempo; Step 2, music style and emotion recognition, including the steps of classifying song styles and identifying emotional tendencies, the music style and emotion recognition being based on the results of the music beat recognition; Step 3, drum feature analysis and recognition, including the identification of bass drum and military drum hitting points, the drum feature analysis being based on the results of the music beat recognition; Step 4, music segmentation recognition, based on the results of the music style recognition and the drum feature analysis; Step 5, special event detection, including high-pitched voice, accompaniment silence, opening special drum and ending fade-out recognition, based on the results of the music beat recognition, music style recognition, drum feature analysis and music segmentation recognition; Step 6, popular song color system switching recognition and execution, based on the results of the music style recognition and emotion recognition; Step 7, generation of light driving file, based on the results of the beat recognition, music style and emotion recognition, drum feature analysis, music segmentation recognition, special event detection and color system switching recognition and execution, to realize the dynamic mapping control of music signal features and light.
[0011] The specific implementation process of the present application is as follows: 1. Cloud feature recognition: the music signal is transmitted to the cloud server through the network, and the sound feature driving file is generated after the deep learning model processing, which contains the information of music beat, style, emotion, drum feature, music segmentation and special event, etc., and is stored in the cloud.
[0012] 2. Real-time control of hardware box: - In the intelligent light mode, the hardware box downloads the sound feature driving file from the cloud, combines the "material segment file" and configuration file pre-arranged by the lighting designer, generates a DMX512 standard data stream of 30 frames per second, and realizes accurate display of real-time light effect through precise synchronization interface.
[0013] - In the scene light mode, the hardware box replicates the working principle of the traditional light system, uses the sound card to sample audio and analyzes it in real time, detects specific drum points, triggers event material playback according to the detection results, and plays fixed scene materials in a loop, switches modes according to user needs to respond to sudden changes in the live environment.
[0014] 3. Lighting designer material arrangement: the lighting designer uses a dedicated arrangement tool to arrange the recorded "DMX512 material segments" according to the music classification understanding and creative arrangement, and outputs the "material file" that can match the sound feature driving file generated by the cloud.
[0015] 4. Sound feature driving file: as a new way to describe music signals, this file can be manually and finely made by lighting designers to achieve extreme richness and accuracy, or it can be generated in batches through AI algorithms to significantly reduce costs. These two generation methods can be flexibly selected according to actual needs.
[0016] 5. Light show automation and intelligence: The introduction of the "sound feature driven" file enables the "light control" system to cover a type of song with the "material file" of the lighting designer, rather than arranging each song individually, greatly improving work efficiency, and also providing a solid foundation for the automation and intelligence of the light show.
[0017] The technical solution of the present application not only improves the audio-visual experience of live performances, but also optimizes the work process, reduces costs, and opens up new possibilities for the interaction of music and light.
[0018] Further, the step 1 of model-based music beat recognition specifically includes the following steps: 1.1 Volume normalization: using peak normalization or root mean square (RMS) energy normalization technology, the energy of the audio is adjusted to a preset range; 1.2 Bar and beat time point recognition: extract Mel frequency cepstral coefficients, spectral contrast, and zero-crossing rate from the original audio; use long short-term memory network (LSTM) to automatically recognize the repetition pattern, intensity change, and frequency distribution in the music through learning a large amount of music data, so as to accurately detect the start and end time of each bar and each beat point within the bar; 1.3 Number of bars and time signature determination: after identifying each bar, count the number of bars contained in the entire music; determine the basic time signature of the music according to the rhythm of the beats within the bar; 1.4 Tempo calculation: first, record the time point of the first beat in each bar; then, select multiple adjacent bars and calculate the sum of the number of beats between them; finally, divide this total number by the time difference to obtain the BPM value; 1.5 Integration and optimization: integrate the bar, beat time point, number of bars, time signature, and tempo information into a data structure; for errors that may occur during the identification process, such as misidentification caused by irregular noise, set a threshold or use a filter to exclude these abnormal values.
[0019] Further, step 2: music style and emotion recognition specifically includes the following steps: 2.1 Feature extraction: based on the music beat recognition results of the first stage, extract more detailed music features, including speed features, spectral features, rhythm patterns, melody features, and harmonic structures; 2.2 Style classification: use the trained classification model to classify music into the following styles using SVM support vector machine, decision tree, random forest, or deep learning model: DJ music: usually with high BPM, pop music; subdivide pop music into: fast songs, slow songs, and medium-speed songs: 2.3 Emotion Tendency Recognition: The emotion recognition model is mainly used for popular music, especially the refined subcategories. The model identifies the emotional state conveyed by the music by analyzing the movement of the melody line, the color of the harmony, the density and speed of the rhythm, and other comprehensive characteristics, including but not limited to: happiness, sadness, passion, tranquility.
[0020] Further, the step 3: Drum Feature Analysis and Recognition specifically includes: 3.1 Preprocessing: Based on the results of beat recognition, first perform audio signal preprocessing, including noise removal, frequency band segmentation, and volume normalization, to ensure that the drum signal is clear and separated from other audio tracks; 3.2 Drum Detection: Use short-time Fourier transform (STFT) to convert the audio signal into a frequency domain representation, which facilitates the identification of drum features; find energy mutation points in the low frequency band, and by setting an energy threshold and a detection window, the potential drum position can be located in the time domain; 3.3 Bass Drum and Snare Drum Classification: Use support vector machines (SVM) to classify the detected drum: extract features related to bass drums and snare drums, including spectral shape, zero-crossing rate, instantaneous frequency, and duration; use labeled bass drum and snare drum samples to train the classification model and optimize the model parameters; input the detected drum features into the trained model to predict whether each drum is a bass drum or a snare drum; 3.4 Precise Drum Hit Positioning: Based on the preliminary detection results of the drum, further refine the position of the hit point; use peak detection algorithms to accurately measure the hit time point of the drum; 3.5 Data Fusion: Combine the detected drum classification results with the beat recognition timestamps to create a list with drum type and precise hit time points, providing accurate timing and action instructions for real-time response of light control.
[0021] Further, the step 4, based on the results of music style recognition and drum feature analysis, performs music segmentation recognition, specifically including: 4.1 Establishment of Music Segmentation Principles: According to the results of music style recognition, determine the segmentation principles corresponding to different music styles; 4.2 Feature Selection and Extraction: Based on music style and drum features, select and extract key features for segmentation recognition; including: music energy profile: continuously monitor the energy changes of the music, identify energy peaks or valleys as potential segmentation points; rhythm density: quantify the density of drum hits or backbeats within each measure to distinguish between the introduction, chorus, and interlude sections of a song; and harmonic progression: analyze the rules of harmonic transitions; 4.3 Automatic segmentation of music using Hidden Markov Models (HMM), clustering analysis, and Variational Autoencoders (VAE), including: Hidden Markov Models (HMM): Train HMM models to learn the characteristic patterns of different musical sections, then use the models to segment the entire song; Clustering analysis: Cluster the time-series features, with features within the same cluster representing similar musical sections. Identify the segmentation points using cluster centers or boundary points; Variational Autoencoders (VAE): Use VAE to learn the latent representation of music, and identify the transition points of musical sections by changes in distance in the latent space; 4.4 Refinement and optimization: Refine and optimize the preliminary identified segments; 4.5 Result output: Output the optimized music segmentation results as a list of timestamps, each corresponding to the start point of a musical section. These results will be used to guide the dynamic changes of the lighting effects, ensuring that the lighting is closely synchronized with each part of the music, enhancing the consistency and depth of the audio-visual experience.
[0022] Further, the step 5: special event detection specifically includes: 5.1 Signal preprocessing: Preprocess the music signal, including denoising, spectral analysis, and smoothing filtering; 5.2 Vocal high pitch detection: Use signal processing techniques to identify peaks in the vocal frequency range, and use vocal recognition algorithms to confirm high pitch parts; 5.3 Accompaniment silence detection: Use energy detection and zero-crossing rate calculation to monitor sections of the music signal where energy suddenly decreases or is close to zero for a long time, to determine the silence period of the accompaniment; 5.4 Opening special drum point recognition: Analyze the characteristics of the first few seconds of the music to identify whether there are significant drum patterns or unusually high energy peaks; 5.5 End fade-out detection: Monitor the end of the music signal to find the trend of gradually decreasing energy.
[0023] Further, the step 6: popular song color system switching identification and execution specifically includes: 6.1 Color system selection: Based on the results of music style and emotion recognition, establish a mapping relationship between color systems and music style and emotion, including: fast-paced and cheerful songs tend to bright and saturated colors, while slow songs or sad melodies tend to soft and low-saturation colors; 6.2 Switching strategy: Design a color system switching strategy based on music segmentation and emotional turning points, and switch the color system at significant emotional changes or important melody nodes; 6.3 Execution control: Integrate the color system switching identification results into the lighting control system to ensure smooth and natural color system transitions at the right music positions.
[0024] Further, the step 7: light drive file generation technology extension, specifically includes: 7.1 Merge the results of beat recognition, music style recognition, emotion recognition, drum feature analysis, music segmentation recognition, special event detection and color system switching recognition to form a unified data set; 7.2 Time synchronization encoding: In the drive file, assign a timestamp to each event including beat, paragraph switching, color system change, to ensure that all light changes can be accurately executed according to the timeline of the music; 7.3 Dynamic parameter configuration: Set dynamic parameters for different types of light effects, including intensity, color, and flashing frequency. These parameters are adjusted according to the real-time changes of music characteristics to achieve the best visual effect.
[0025] Preferably, as shown in Figure 2 A system using any of the methods described above, comprising: A feature recognition module, including: a beat recognizer for identifying measures, beat time points, measure numbers, time signatures, and calculating tempo; a style recognizer for classifying song styles and identifying emotional tendencies; a drum analyzer for identifying bass drum and snare drum hitting points; a segmentation recognizer for identifying music segmentation based on the results of the music style recognition; and a special event detector for detecting high-pitched vocals, accompaniment silence, special drum at the beginning and fade-out at the end; A light control unit for real-time control of light equipment according to the output of the feature recognition module; A storage unit for storing the music feature data and the light drive file; A central processing unit (CPU) for coordinating system workflow, processing the output of the feature recognition module and generating light control signals, the processing of the central processing unit is based on the results of the beat recognition, the music style recognition, the drum analysis, the segmentation recognition and the special event detection.
[0026] Preferably, one aspect of the present application provides a music signal feature extraction and light dynamic mapping control method, comprising the steps of identifying measures, beat time points, measure quantities, time signatures, and calculating tempo. In terms of technology, the present embodiment adopts advanced audio analysis technology, pre-processes audio signals through a deep learning model, extracts timing features, and then identifies measure boundaries and beat positions in music to accurately calculate tempo. In principle, beat recognition is based on the periodicity and repeatability of music signals, and beat points are determined by analyzing the frequency spectrum and energy distribution of audio signals. In terms of effects, the technology in the present embodiment enables the light to accurately follow the rhythm changes of the music, enhancing the audio-visual synchronization and live atmosphere. In other embodiments, the generalization ability of beat recognition can be improved by improving the training data set of the model and increasing the samples of different styles of music to solve the problem of inaccurate beat recognition in complex music environments.
[0027] Further, in the music beat recognition step based on the model, a deep learning model is used for beat detection based on timing feature analysis to improve the accuracy of measure and beat time point recognition and achieve more accurate tempo calculation. The present embodiment technically utilizes the powerful feature extraction capability of deep neural networks, and through learning a large number of music samples, the model can capture subtle changes in music signals, thereby more accurately recognizing beats. In principle, the deep learning model simulates the learning process of the human brain through the connection of multiple neurons, abstracts high-level features of music signals step by step, and finally realizes accurate positioning of beats. In terms of effects, the accuracy of beat recognition is improved, the light control is more delicate, and the audience's immersion is enhanced. In other embodiments, attention mechanisms or recurrent neural networks (RNN) and other technologies can be introduced to further optimize the model structure and improve the processing capability of long sequence music signals, solving the stability problem of beat recognition in long time music playing.
[0028] Further, the music style and emotion recognition is based on the results of the music beat recognition, including the steps of classifying song style and identifying emotional tendency. In terms of technology, this embodiment uses convolutional neural network (CNN) and sentiment analysis algorithm to achieve accurate recognition of music style and emotion by analyzing the spectral features and lyrics of the music. In principle, CNN can extract local features from the music spectrum, and the sentiment analysis algorithm can understand the emotional tone of the song by combining the lyrics content. The two work together to comprehensively capture the style and emotion information of the music. In terms of effect, it enables the lights to adjust the form of expression according to the changes of music style and emotion, improving the artistic and ornamental performance. In other embodiments, more complex models such as long short-term memory network (LSTM) can be used to solve the limitations of single feature recognition and improve the accuracy of style and emotion recognition by fusing multi-dimensional information such as music speed, pitch and chord.
[0029] Further, the drum feature analysis and recognition are based on the results of the music beat recognition, especially including the identification of bass drum and military drum hitting points. In terms of technology, this embodiment uses spectral analysis technology and machine learning algorithm based on time series analysis to accurately identify the position and type of drum points by analyzing the frequency domain and time domain features of the audio signal. In principle, drum point recognition utilizes the unique pattern of drum sound in the frequency spectrum and the regularity of drum points in the time series. By filtering and matching these features through algorithms, the precise positioning of bass drum and military drum hitting points is achieved. In terms of effect, it enables the lights to respond instantly when the drum points appear, enhancing the rhythm and live interactivity. In other embodiments, more professional audio processing libraries can be used to solve the problem of complex drum set recognition by adding recognition of other percussion instruments such as cymbals and congas, thereby improving overall expressiveness.
[0030] Further, the music segmentation recognition is based on the results of music style recognition and drum feature analysis. In terms of technology, this embodiment uses clustering algorithm and deep learning technology to automatically divide songs into different sections by comprehensively analyzing the melody, rhythm and drum feature of the music. In principle, music segmentation recognition relies on the internal logic of music structure, and identifies the key points of section transition by analyzing the trend of music elements. In terms of effect, it ensures the natural transition of light effects between different music sections, improving music expressiveness and visual coherence. In other embodiments, music theory knowledge such as harmony analysis and form structure recognition can be introduced to further refine the segmentation rules and solve the problem of inaccurate segmentation in complex music structure.
[0031] Further, special event detection is based on the results of music beat recognition, music style recognition, drum feature analysis, and music segmentation recognition, including the identification of high-pitched vocals, accompaniment silence, special drum at the beginning, and fade-out at the end. Technically, this embodiment utilizes acoustic models and deep learning techniques to identify the timing of special events through spectral and intensity analysis of audio signals. In principle, special event detection relies on the abnormal characteristics of music signals, such as spectral spikes of high-pitched vocals and energy drops of accompaniment silence, to achieve accurate detection of special events by setting thresholds and pattern matching. In terms of effects, it enhances the expressiveness of lighting at musical climaxes and turning points, and improves the sensory experience of the audience. In other embodiments, more diverse event detection can be achieved by introducing multi-modal analysis, such as combining video analysis, to solve the problem of event recognition diversity in multiple scenarios.
[0032] Further, popular song color scheme switching recognition is based on the results of music style recognition and emotion recognition. Technically, this embodiment uses color psychology principles and color space mapping techniques to design color scheme switching schemes based on emotional climax points of songs. In principle, color psychology studies the psychological impact of different colors on people, and by associating music emotions with colors, it designs light color changes that match the emotions of the music. In terms of effects, it makes light colors fluctuate with music emotions, improving the visual and auditory enjoyment of the audience. In other embodiments, user preference learning can be introduced to dynamically adjust color scheme switching strategies based on audience feedback, addressing the technical challenges of individualized needs differentiation.
[0033] Further, the generation of light driving files is based on the results of beat recognition, music style and emotion recognition, drum analysis, music segmentation recognition, special event detection, and color scheme switching recognition to achieve dynamic mapping control of music signal features and light. Technically, this embodiment uses custom algorithms and time synchronization techniques to match music features with pre-set light effects, ensuring real-time synchronization of light control commands with music signal features. In principle, the generation of light driving files is a multi-factor decision-making process that needs to consider the speed, style, emotion, drum, section, and special events of music to calculate the best light control strategy through algorithms. In terms of effects, it achieves seamless integration of music and light, improving the accuracy of visual and auditory synchronization and artistic expression. In other embodiments, real-time rendering technology can be introduced to dynamically adjust light effects based on live environment and audience feedback, addressing the adaptability of light performance in dynamic environments.
[0034] Further, the custom algorithm is based on multi-dimensional information such as the spectrum, amplitude, rhythm, etc. of the music to comprehensively determine the parameters such as brightness, color, and flashing frequency of the light. In terms of technology, the embodiment maps various characteristics of the music signal to the parameters of the light effect through algorithm design, achieving precise light control. In terms of principle, the custom algorithm converts music signal characteristics into light control parameters through a mathematical model, such as mapping the brightness through the spectral peak and the flashing frequency through the rhythm intensity. In terms of effect, the light effect is more delicate and diversified, enhancing the artistic nature of the live performance and the audience's sense of participation. In other embodiments, machine learning technology can be introduced to allow the algorithm to learn and optimize itself, solving the problem of light effect adaptability under different music styles. Further, the music signal feature extraction and light dynamic mapping control system includes modules such as a beat recognizer, a style recognizer, a drum point analyzer, a segment recognizer, a special event detector, and a light control unit. In terms of technology, the embodiment constructs an integrated control system, with each module coordinated by a central processing unit (CPU) to achieve efficient extraction of music signal characteristics and intelligent control of light. In terms of principle, the system architecture design follows the modular principle, with each module focusing on a specific music feature recognition task, and through the unified scheduling of the CPU, achieving fast information transmission and processing. In terms of effect, the system's response speed and control accuracy are improved, making the light performance more smooth and lively. In other embodiments, edge computing technology can be introduced to offload some computing tasks to the light device end, reducing the burden on the central processor and solving the problem of delay and bandwidth limitation in large-scale light control systems.
[0035] Further, the light control unit has multi-channel output capability and can control different types of light devices such as RGB LED lights, stage laser lights, and dynamic projection lights, etc. In terms of technology, the light control unit of the embodiment uses advanced multi-channel output technology to control multiple types of light devices simultaneously, achieving complex light effect combinations. In terms of principle, the multi-channel output technology is based on digital signal processing (DSP) and wireless communication protocols, allowing independent control of each channel's light device to achieve precise light effect synchronization. In terms of effect, it provides diversified and hierarchical light effects, enhancing the flexibility and aesthetic nature of light performance. In other embodiments, virtual reality (VR) or augmented reality (AR) technology can be introduced to combine light effects with virtual scenes, creating a more immersive audio-visual experience and solving the problem of single and lack of innovation in traditional light effects.
[0036] Further, the special event detector includes a vocal separation algorithm that extracts the vocal signal from the accompaniment using deep learning techniques, and identifies the high and low pitch of the vocals and their duration. Technically, the vocal separation algorithm of the present embodiment uses a deep learning framework such as U-Net or Wavenet to process the audio signal, effectively separating the vocals and accompaniment. In principle, the vocal separation algorithm learns the spectral and temporal features of vocals and accompaniment to establish a model to distinguish the two signals, achieving accurate extraction of the vocal signal. In terms of effects, it enables the lights to respond to the high and low pitch changes of the vocals, enhancing the emotional expression and visual impact of the performance. In other embodiments, voice recognition technology can be introduced to identify specific words or tones in the vocals, enabling more precise synchronization of lights with the vocals, and solving the problem of light control in complex vocal changes.
[0037] The technical solution of the present application involves the working process and the use process. The working process is as follows: first, the music signal is input into the system, and the beats are identified by the beat identifier, and then the music style and emotion are analyzed by the style identifier, and then the drum point analyzer identifies the bottom drum and military drum hitting points, and the segment identifier identifies the segments according to the music style, and the special event detector detects the vocal high pitch, accompaniment silence, special drum points at the beginning and end, and other special events. On this basis, the central processing unit (CPU) integrates the outputs of all feature recognition modules, matches the music features with the light preset effects through a self-defined algorithm, and finally adjusts the light equipment in real time according to the control signal generated by the CPU, realizing the synchronization and interaction of music and light. In the use process, the user only needs to connect the music signal to the system, and the system automatically completes the extraction of music features and the control of light effects, without manual intervention, greatly simplifying the operation process and improving the efficiency and watchability of live performances.
[0038] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions.
[0039] It should be further noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or other elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0040] The above examples are to be understood only as illustrative of the present application and not restrictive of the scope of the present application. After reading the foregoing disclosure, many modifications and variations of the present application will be apparent to the skilled person. Such variations and modifications are to be considered as falling within the scope of the present application as defined by the claims.
Claims
1. A method for music signal feature extraction and dynamic lighting mapping control, characterized in that, Includes the following steps: Step 1, model-based music beat recognition, includes: identifying measures, beat time points, number of measures, time signature, and calculating tempo; Step 2, music style and emotion identification, includes: classifying song styles and identifying emotional tendencies, the music style and emotion identification being based on the results of music beat identification; Step 3, drum beat feature analysis and identification, including the identification of bass drum and snare drum strike points, the drum beat feature analysis is based on the results of the music beat identification; Step 4, music segmentation and recognition, is performed based on the results of the music style recognition and the drum beat feature analysis; Step 5, special event detection, including the identification of high vocal pitch, silent accompaniment, special drum beats at the beginning, and fade-out at the end. The special event detection is based on the results of music beat recognition, music style recognition, drum beat feature analysis, and music segmentation recognition. Step 6: Popular song color scheme switching recognition and execution, wherein the color scheme switching recognition is based on the results of the music style recognition and the emotion recognition; Step 7: The cloud uses an AI model to generate a corresponding sound feature driver file based on the input audio data, and generates control instructions based on the file and the material clip file created by the lighting engineer. The generation step is based on the results of the beat recognition, music style and emotion recognition, drum beat feature analysis, music segment recognition, special event detection and color system switching recognition and execution, so as to realize the dynamic mapping control of music signal features and lighting.
2. The method for music signal feature extraction and dynamic lighting mapping control according to claim 1, characterized in that, Step 1, model-based music beat recognition, specifically includes the following steps: 1.1 Volume Normalization: Using peak normalization or root mean square (RMS) energy normalization techniques, the audio energy is adjusted to a preset range; 1.2 Measure and beat time point recognition: Mel frequency cepstral coefficients, spectral contrast, and zero-crossing rate are extracted from the original audio. Using a Long Short-Term Memory (LSTM) network, through learning from a large amount of music data, the model automatically recognizes repetition patterns, intensity changes, and frequency distribution in the music, thereby accurately detecting the start and end times of each measure and each beat point within the measure. 1.3 Determining the Number of Measures and Time Signature: After identifying each measure, count the number of measures in the entire piece; determine the basic time signature of the music based on the rhythmic pattern within each measure; 1.4 Beat Calculation: First, record the time of the first beat in each measure; then, select multiple adjacent measures and calculate the sum of the number of beats between them; finally, divide this total by the time difference to obtain the BPM value; 1.5 Integration and Optimization: Integrate measure, beat time, measure number, beat number, and beat speed information into a single data structure; for errors that may occur during the recognition process, such as misidentification caused by irregular noise, these outliers should be excluded by setting thresholds or using filters.
3. The method for music signal feature extraction and dynamic lighting mapping control according to claim 1, characterized in that, Step 2: Music Style and Emotion Recognition specifically includes the following steps: 2.1 Feature Extraction: Based on the results of the first stage of music beat recognition, more detailed music features are extracted, including tempo features, spectrum features, rhythm patterns, melody features, and harmonic structure; 2.2 Style Classification: Using a trained classification model, employing SVM (Support Vector Machine), decision trees, random forests, or deep learning models, music is classified into the following styles: DJ music: typically with a high BPM, popular music; popular music is further subdivided into: fast songs, slow songs, and mid-tempo songs. 2.3 Emotional Tendency Recognition: The emotion recognition model is mainly used for popular music, especially for refined subcategories. The model identifies the emotional state conveyed by music by analyzing comprehensive features such as the direction of the melodic line, the color of harmony, and the density and tempo of the rhythm, including but not limited to: happiness, sadness, passion, and tranquility.
4. The method for music signal feature extraction and dynamic lighting mapping control according to claim 1, characterized in that, Step 3: Drumbeat Feature Analysis and Recognition specifically includes: 3.1 Preprocessing: Based on the results of beat recognition, the audio signal is first preprocessed, including noise removal, frequency band segmentation and volume normalization, to ensure that the drum signal is clear and separated from other audio tracks; 3.2 Drumbeat Detection: The audio signal is converted into a frequency domain representation using Short Time Fourier Transform (STFT) to facilitate the identification of drumbeat features; energy abrupt changes are found in the low-frequency band, and potential drumbeat locations can be located in the time domain by setting energy thresholds and detection windows; 3.3 Bass Drum and Snare Drum Classification: Support Vector Machine (SVM) is used to classify the detected drum beats: Features related to the bass drum and snare drum are extracted, including spectral shape, zero-crossing rate, instantaneous frequency, and duration; the classification model is trained using labeled bass drum and snare drum samples, and the model parameters are optimized; the detected drum beat features are input into the trained model to predict whether each drum beat is a bass drum or a snare drum. 3.4 Precise Positioning of Impact Points: Based on the preliminary detection results of the drumbeats, the position of the impact points is further refined; a peak detection algorithm is used to accurately measure the timing of the drumbeats. 3.5 Data Fusion: The detected drum beat classification results are combined with the timestamps of the beat recognition to create a list with drum beat type and precise hit time, providing accurate time and motion instructions for the real-time response of the lighting control.
5. The method for music signal feature extraction and dynamic lighting mapping control according to claim 1, characterized in that, Step 4, which involves music segmentation based on the results of the music style recognition and the drum beat feature analysis, specifically includes: 4.1 Establishment of Music Segmentation Principles: Based on the results of music style identification, determine the segmentation principles corresponding to different music styles; 4.2 Feature Selection and Extraction: Based on musical style and drum beat characteristics, key features for segmentation identification are selected and extracted, including: Musical Energy Profile: Continuously monitor the energy changes of the music and identify energy peaks or troughs as potential segmentation points; Rhythmic Density: Quantify the density of drum beats or downbeats within each measure to distinguish between the introduction, chorus, and interlude of the song; Harmonic Progression: Analyze the rules of harmonic transitions. 4.3 Hidden Markov Model (HMM), cluster analysis, and variational autoencoder (VAE) are used to achieve automatic music segmentation, including: Hidden Markov Model (HMM): By training the HMM model, the feature patterns of different music segments are learned, and then the model is used to divide the entire song into segments; Cluster analysis: Temporal features are clustered, and features within the same group represent similar segments of the music. The segmentation position is identified by the cluster center or boundary point; Variational autoencoder (VAE): The VAE is used to learn the latent representation of the music, and the transition points of music segments are identified by the changes in distance in the latent space. 4.4 Refinement and Optimization: Refine and optimize the initially identified segments; 4.5 Output Results: The optimized music segmentation results will be output as a list of timestamps, with each timestamp corresponding to the start point of a music segment. These results will be used to guide the dynamic changes of subsequent lighting effects, ensuring that the lighting and music are closely synchronized, and improving the consistency and depth of the audiovisual experience.
6. The method for music signal feature extraction and dynamic lighting mapping control according to claim 1, characterized in that, Step 5: Special event detection specifically includes: 5.1 Signal Preprocessing: Preprocessing the music signal, including noise reduction, spectrum analysis, and smoothing filtering; 5.2 High-pitched voice detection: Signal processing technology is used to identify peak values within the frequency range of human voices, and high-pitched parts are confirmed by combining them with human voice recognition algorithms; 5.3 Accompaniment silence detection: Using energy detection and zero crossover rate calculation, monitor sections in the music signal where the energy suddenly decreases or remains close to zero for a long time to determine the silence period of the accompaniment; 5.4 Identification of Special Drum Beats at the Beginning: By analyzing the characteristics of the first few seconds of the music, identify whether there are significant drum beat patterns or abnormally high energy peaks; 5.5 End-of-signal fade-out detection: Monitor the end of the music signal to find the trend of gradually decreasing energy.
7. The method for music signal feature extraction and dynamic lighting mapping control according to claim 1, characterized in that, Step 6: Popular Song Color System Switching Recognition and Execution specifically includes: 6.1 Color Scheme Selection: Based on the results of music style and emotion recognition, establish a mapping relationship between color scheme and music style and emotion, including: fast-paced and cheerful songs tend to have bright and saturated colors, while slow songs or sad melodies tend to have soft and low-saturation colors. 6.2 Switching Strategy: Design a color scheme switching strategy based on music segmentation and emotional turning points, and switch colors when the mood of the song changes significantly or at important melodic nodes; 6.3 Execution Control: Integrate color scheme switching recognition results into the lighting control system to ensure a smooth and natural color scheme transition at the appropriate music position.
8. The method for music signal feature extraction and dynamic lighting mapping control according to claim 1, characterized in that, Step 7: Extension of lighting driver file generation technology, specifically includes: 7.1 The results of beat recognition, music style recognition, emotion recognition, drum beat feature analysis, music segmentation recognition, special event detection, and color scheme switching recognition are merged to form a unified dataset; 7.2 Time Synchronization Encoding: In the driver file, a timestamp is assigned to each event, including beats, paragraph transitions, and color changes, to ensure that all lighting changes are executed precisely according to the music timeline; 7.3 Dynamic Parameter Configuration: Set dynamic parameters for different types of lighting effects, including intensity, color, and flashing frequency. These parameters are adjusted in real time according to the changes in music characteristics in order to achieve the best visual effect.
9. A system employing the method according to any one of claims 1-8, characterized in that, include: The feature recognition module includes: a beat recognizer for recognizing measures, beat timings, number of measures, time signature, and calculating tempo; a style recognizer for classifying song styles and recognizing emotional tendencies; a drum beat analyzer for recognizing bass drum and snare drum strike points; a segmentation recognizer for recognizing music segments based on the results of the music style recognition; and a special event detector for detecting high vocal pitches, silent accompaniment, special drum beats at the beginning, and fade-out at the end. A lighting control unit is used to control lighting equipment in real time based on the output of the feature recognition module; The storage unit is used to store the music feature data and the lighting drive file in the cloud; the cloud uses an AI model to generate the corresponding sound feature drive file based on the input audio data. The central processing unit (CPU) is used to coordinate the system workflow, process the output of the feature recognition module, and generate lighting control signals. The processing of the CPU is based on the results of the beat recognition, the music style recognition, the drum beat analysis, the segmentation recognition, and the special event detection.