LED lamp strip intelligent control method and system based on sound change

Through microphone array beamforming and DAGAN noise compensation combined with LSTM-TCN model optimization, the environmental noise interference and user intention identification problems of traditional voice control systems are solved, and intelligent lighting control with high accuracy and low power consumption is achieved.

CN120512801AInactive Publication Date: 2025-08-19ZHONGSHAN LANDE ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510948093.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-08-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional voice-controlled lighting control systems are susceptible to environmental noise interference, and the sound source positioning deviation leads to a high false trigger rate, making it difficult to capture the dynamic timing mode of user behavior. The optimization of existing models ignores the balance between real-time, energy efficiency and personalized needs, and is difficult to meet the deployment requirements of low-power embedded devices.

Method used

Microphone array beamforming algorithm is used to identify the sound source direction, combine the LSTM-TCN neural network and the MOPSO multi-objective particle swarm optimization algorithm for feature fusion and model optimization, and use DAGAN deep adversarial generation network for noise compensation, realizing user intention recognition and intelligent lighting control.

Benefits of technology

Improve sound source positioning accuracy in low signal-to-noise ratio environments, reduce error triggering rate, optimize real-time response time, support personalized scene switching, reduce power consumption and extend device life cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120512801A_ABST
    Figure CN120512801A_ABST
Patent Text Reader

Abstract

The invention relates to the field of LED lamp strip intelligent control, in particular to an LED lamp strip intelligent control method and system based on sound changes. The method comprises the following steps: acquiring environment sound signal data in an indoor microphone array, identifying a sound source direction in the environment sound signal data through a beam forming algorithm, extracting multi-dimensional feature data in the sound source signal source data, carrying out multi-modal feature fusion on the multi-dimensional feature data, and carrying out multi-modal feature fusion on the multi-modal feature data. The method comprises the steps of establishing an LSTM-TCN user behavior recognition model by utilizing a neural network architecture, performing three-layer optimization on network hyper-parameters, MFCC feature weights and lighting effect decisions of the model by utilizing a multi-target particle swarm optimization algorithm, performing noise compensation on multi-dimensional vector data by utilizing a DAGAN, and inputting the multi-dimensional vector data into a target user behavior recognition model. And obtaining the real-time intention of the user. The intention recognition accuracy is improved, and the false triggering rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent control of LED light strips, and in particular to an intelligent control method and system for LED light strips based on sound changes. Background Art

[0002] With the rapid development of smart home technology, lighting control systems based on voice interaction have become a research hotspot. Traditional voice control solutions often rely on keyword wake-up or fixed threshold triggering, which has significant limitations. First, a single microphone is susceptible to interference from ambient noise. Especially in complex home scenarios, interference sources such as multiple people talking can lead to deviations in sound source localization and high false trigger rates. Second, traditional acoustic feature extraction methods are static and single, making it difficult to capture the dynamic temporal patterns of user behavior, resulting in insufficient intent recognition accuracy. Finally, existing model optimization often focuses on a single metric, ignoring the balance between real-time performance, energy efficiency, and personalization, making it difficult to meet the deployment requirements of low-power embedded devices. Summary of the Invention

[0003] The purpose of the present invention is to solve the above problems and to design an intelligent control method and system for LED light strips based on sound changes.

[0004] The technical solution of the present invention to achieve the above object is that, in an intelligent control method of an LED light strip based on sound changes, the intelligent control method of an LED light strip comprises the following steps:

[0005] Acquire ambient sound signal data from an indoor microphone array, identify the direction of a sound source in the ambient sound signal data using a beamforming algorithm, and obtain sound source signal source data;

[0006] Extracting multi-dimensional feature data from the sound source signal source data, and performing multimodal feature fusion on the multi-dimensional feature data to obtain multi-dimensional vector data;

[0007] An LSTM-TCN user behavior recognition model was established using the LSTM-TCN neural network architecture. The MOPSO multi-objective particle swarm optimization algorithm was used to perform three-level optimization on the model's network hyperparameters, MFCC feature weights, and lighting effect decisions, resulting in the target LSTM-TCN user behavior recognition model.

[0008] The multi-dimensional vector data is noise-compensated using the DAGAN deep generative adversarial network and then input into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention;

[0009] Intelligent control of the LED light strip is performed based on the real-time intention of the user, including at least dynamic light mapping and light scene adaptation.

[0010] Furthermore, in the above-mentioned intelligent control method for LED light strips based on sound changes, the step of obtaining ambient sound signal data from an indoor microphone array, identifying the direction of the sound source in the ambient sound signal data through a beamforming algorithm, and obtaining the sound source signal source data includes:

[0011] Acquiring ambient sound signal data from an indoor microphone array, wherein the microphone array is at least arranged on a ceiling, a wall, and a floor;

[0012] Using a Hamming window to perform frame processing on the ambient sound signal data, wherein the frame length is 20 ms and the frame shift is 10 ms; performing a short-time Fourier transform on the frame-processed data to obtain first ambient sound signal data;

[0013] Calculate the time difference between the sound source in the first ambient sound signal data reaching different microphones based on the GCC-PHAT generalized cross-correlation-phase transform algorithm, and solve the sound source direction by the least squares method in combination with the uniform circular array to obtain the second ambient sound signal data;

[0014] The MVDR minimum variance distortionless response beamforming algorithm is used to enhance the signal of the sound source direction in the second environmental sound signal data, suppress interference from other directions, and obtain the source data of the sound source signal.

[0015] Furthermore, in the above-mentioned intelligent control method for LED light strips based on sound changes, the multi-dimensional feature data from the sound source signal source data is extracted, and multi-modal feature fusion is performed on the multi-dimensional feature data to obtain multi-dimensional vector data, including:

[0016] Extracting multi-dimensional feature data from the sound source signal source data, including at least time domain features, frequency domain features and dynamic features;

[0017] The time domain features include at least short-time energy, zero-crossing rate, amplitude envelope dynamic range and autocorrelation function peak ratio; the frequency domain features include at least Mel frequency cepstral coefficients, spectrum centroid and spectrum flux; the dynamic features include at least the first-order difference of the dynamic change rate of the time domain features and the frequency domain features and the second-order difference obtained by performing a differential calculation on the feature data after the first-order difference;

[0018] A unified feature vector is constructed for the time domain features, frequency domain features and dynamic features in the multidimensional feature data, the mean variance is normalized for each feature dimension, the multidimensional feature vector is organized by frame, the features of the current frame and the two frames before and after are taken, the statistics within the window are calculated, and the multidimensional vector data is obtained.

[0019] Furthermore, in the above-mentioned intelligent control method for LED light strips based on sound changes, the LSTM-TCN neural network architecture is used to establish an LSTM-TCN user behavior recognition model, and the MOPSO multi-objective particle swarm optimization algorithm is used to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and light effect decisions to obtain the target LSTM-TCN user behavior recognition model, including:

[0020] Based on the LSTM fusion time series modeling and TCN local feature extraction capabilities, spatiotemporal joint modeling is performed. The network structure of the model includes at least an input layer, an LSTM module, a TCN module, and a classification head.

[0021] The input layer is used to receive historical multi-dimensional vector data in the system, and its time window length is 10 frames;

[0022] The LSTM module includes a bidirectional LSTM layer and an output layer. The bidirectional LSTM layer is used to capture long-term dependencies:

[0023]

[0024] Among them, the number of hidden units in the LSTM module The optimization is performed using the MOPSO multi-objective particle swarm optimization algorithm with a range of (32,256); Represents the hidden state at the previous moment, Represents the input features at the current moment, Represents the output of the bidirectional LSTM layer;

[0025] The output layer of the LSTM module is used to concatenate the time step outputs into , represents the time step, represents a bidirectional LSTM;

[0026] The dilated convolution kernel size of the TCN module , expansion factor ,in is the number of layers, depth The residual connection and weight normalization calculation formula of the TCN module are as follows:

[0027]

[0028] The output layer of the TCN module: ,in is the number of channels, Represented as a convolution operation, represents the activation function;

[0029] The classification head includes global average pooling + fully connected layer, and the calculation formula is as follows:

[0030]

[0031] in, represents global average pooling, represents the weight of the fully connected layer, Indicates bias, Indicates calculating classification probability output.

[0032] Furthermore, in the above-mentioned LED light strip intelligent control method based on sound changes, the MOPSO multi-objective particle swarm optimization algorithm includes:

[0033] Establish the optimization framework of MOPSO multi-objective particle swarm optimization algorithm:

[0034]

[0035] in, Represents the hyperparameters of the target LSTM-TCN user behavior recognition model; Represents MFCC feature weight; Indicates lighting effect decision; represents the number of hidden units, Indicates the size of the convolution kernel of the dilated convolution stack; represents the depth of the dilation factor, Represents the learning rate; MFCC feature weight ,and ; (0.1, 0.2) represents the light effect parameter of intensity gain; A lighting parameter that indicates the ratio of hot to cold color temperature.

[0036] Furthermore, in the above-mentioned intelligent control method for LED light strips based on sound changes, the MOPSO multi-objective particle swarm optimization algorithm is used to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and light effect decisions to obtain the target LSTM-TCN user behavior recognition model, further comprising:

[0037] Initialize the particle swarm of the MOPSO multi-objective particle swarm optimization algorithm, set the swarm size to 50, and randomly generate particles that meet the constraints;

[0038] The calculation formula of the particle swarm velocity update rule is as follows:

[0039]

[0040] in, Represents the inertia weight, and decreases linearly, and represents the acceleration factor, and Belonging to the set (0,1) represents a random number; Indicates the In the iteration, Particles in Dimensional speed; Indicates the In the iteration, Particles in Dimensional speed; Indicates the Particles in The optimal historical position of individuals in the dimension; Indicates that the global optimal particle is The position of the dimension guides the search direction of the particle swarm; Indicates the In the iteration, Particles in The current position of the dimension;

[0041] The particle swarm is constrained. If the sum of the weights is not equal to 1, it is rescaled and the out-of-range parameters of the particle swarm are reflected back to the feasible domain by mirroring.

[0042] The particle swarm is maintained at the Pareto frontier. The top 20% of non-inferior particles are retained for the next iteration after each iteration. When the number of iterations is equal to 150, the optimal particle swarm is output and replaced as the hyperparameters of the model to obtain the target LSTM-TCN user behavior recognition model.

[0043] Furthermore, in the above-mentioned intelligent control method for LED light strips based on sound changes, the DAGAN deep adversarial generative network is used to perform noise compensation on the multi-dimensional vector data and then inputs it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention, including:

[0044] Obtain multi-dimensional vector data and design a DAGAN deep adversarial generative network architecture, which includes at least a generator and a discriminator. The generator is used to map noisy feature vectors to a clean feature space, and the discriminator is used to distinguish between the restored features output by the generator and the true clean features.

[0045] A noise baseline is established using feature statistics of non-speech segments. When it is detected that the features of the multi-dimensional vector data deviate from the baseline by more than a threshold, a DAGAN deep adversarial generative network is triggered to compensate.

[0046] Compensate for the high-frequency MFCC coefficients most affected by noise, and dynamically assign restoration weights through the attention layer within the generator. The discriminator outputs a restoration credibility score. When the score is lower than 0.7, the current frame is discarded and historical data interpolation is used.

[0047] The denoised features output by DAGAN are organized into sequences according to time windows, and the compensation confidence is added as an additional input channel. The sequence is then fed into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention.

[0048] To achieve the above-mentioned purpose, the technical solution of the present invention is as follows: further, in an intelligent control system for LED light strips based on sound changes, the intelligent control system for LED light strips comprises:

[0049] An environmental data acquisition module is used to obtain environmental sound signal data from an indoor microphone array, identify the direction of the sound source in the environmental sound signal data through a beamforming algorithm, and obtain sound source signal source data;

[0050] A data feature fusion module is used to extract multi-dimensional feature data from the source data of the sound source signal, perform multi-modal feature fusion on the multi-dimensional feature data, and obtain multi-dimensional vector data;

[0051] The recognition model building module is used to build an LSTM-TCN user behavior recognition model using the LSTM-TCN neural network architecture. The MOPSO multi-objective particle swarm optimization algorithm is used to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and lighting effect decisions to obtain the target LSTM-TCN user behavior recognition model.

[0052] The user intention recognition module is used to use the DAGAN deep adversarial generative network to compensate for the noise of the multi-dimensional vector data and then input it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention;

[0053] The light strip intelligent control module is used to intelligently control the LED light strip based on the user's real-time intention, including at least dynamic light mapping and light scene adaptation.

[0054] Furthermore, in the above-mentioned LED light strip intelligent control system based on sound changes, the environmental data acquisition module includes the following submodules:

[0055] An acquisition submodule is used to acquire ambient sound signal data from an indoor microphone array, wherein the microphone array is at least arranged on the ceiling, the wall and the ground;

[0056] A framing processing submodule is configured to perform framing processing on the ambient sound signal data using a Hamming window, wherein the frame length is 20 ms and the frame shift is 10 ms; and perform a short-time Fourier transform on the frame-processed data to obtain first ambient sound signal data;

[0057] A time difference calculation submodule is used to calculate the time difference between the sound source in the first ambient sound signal data reaching different microphones based on the GCC-PHAT generalized cross-correlation-phase transform algorithm, and solve the sound source direction by the least squares method in combination with the uniform circular array to obtain the second ambient sound signal data;

[0058] The data acquisition submodule is used to use the MVDR minimum variance distortionless response beamforming algorithm to enhance the signal of the sound source direction in the second environmental sound signal data, suppress interference from other directions, and obtain the source data of the sound source signal.

[0059] Furthermore, in the above-mentioned LED light strip intelligent control system based on sound changes, it is characterized in that the user intention recognition module has the following submodules:

[0060] The acquisition submodule is used to obtain multi-dimensional vector data and design the DAGAN deep adversarial generative network architecture, which includes at least a generator and a discriminator. The generator is used to map the noisy feature vector to a clean feature space, and the discriminator is used to distinguish the repaired features output by the generator from the real clean features.

[0061] A detection submodule is configured to establish a noise baseline using feature statistics of non-speech segments, and trigger a DAGAN deep adversarial generative network to compensate when it is detected that the features of the multi-dimensional vector data deviate from the baseline by more than a threshold.

[0062] The compensation submodule is used to compensate for the high-frequency MFCC coefficients most affected by noise. The restoration weights are dynamically allocated through the attention layer within the generator. The discriminator outputs a restoration credibility score. When the score is lower than 0.7, the current frame is discarded and the historical data interpolation is used.

[0063] A submodule is obtained, which is used to organize the denoised features output by DAGAN into sequences according to time windows, add compensation confidence as an additional input channel, and input it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention.

[0064] Its beneficial effects include: 1. Enhanced anti-interference capabilities: Through dual-stage processing of microphone array beamforming and DAGAN noise compensation, the sound source localization error is ≤3° in extreme environments with a signal-to-noise ratio as low as -5dB, improving intent recognition accuracy and reducing false trigger rates. 2. Real-time response optimization: The MOPSO algorithm is used to jointly optimize the hyperparameters and feature weights of the LSTM-TCN model, compressing inference latency. Combined with pipeline parallel computing, the overall system response time is shortened to meet real-time interaction requirements. 3. Enhanced personalized adaptation: The introduction of a light effect decision layer and a user preference learning mechanism improves user satisfaction by dynamically adjusting color temperature and intensity parameters and supports smooth switching between customized scenes. 4. Improved resource efficiency: The model parameter count is reduced by 50% (1.8M → 0.9M), supports 8-bit quantization deployment, reduces memory usage by 60%, and can run on ARM Cortex-M7-class chips with power consumption of ≤100mW. 5. Adaptive robustness: DAGAN's online incremental learning and noise environment perception module enables the system to adaptively converge within 24 hours after the emergence of new interference, extending its lifecycle. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Various other advantages and benefits will become apparent to those skilled in the art by reading the following detailed description of the preferred embodiment.The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.

[0066] Figure 1 This is a schematic diagram of a first embodiment of an intelligent control method for LED light strips based on sound changes in an embodiment of the present invention;

[0067] Figure 2 Schematic diagram of a second embodiment of an intelligent control method for LED light strips based on sound changes in an embodiment of the present invention;

[0068] Figure 3 Schematic diagram of a third embodiment of an intelligent control method for LED light strips based on sound changes in an embodiment of the present invention;

[0069] Figure 4 This is a schematic diagram of a first embodiment of an intelligent control system for LED light strips based on sound changes in an embodiment of the present invention. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0071] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a", "an", "" and "" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0072] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 As shown, the LED light strip intelligent control method based on sound changes includes the following steps:

[0073] Step 101: Acquire ambient sound signal data from an indoor microphone array, identify the direction of the sound source in the ambient sound signal data using a beamforming algorithm, and obtain sound source signal source data;

[0074] Specifically, in this embodiment, ambient sound signal data from an indoor microphone array is obtained, and the microphone array is at least arranged on the ceiling, the wall and the ground;

[0075] Using a Hamming window to perform frame processing on the ambient sound signal data, wherein the frame length is 20 ms and the frame shift is 10 ms; performing a short-time Fourier transform on the frame-processed data to obtain first ambient sound signal data;

[0076] The time difference between the sound source arriving at different microphones in the first ambient sound signal data is calculated based on the GCC-PHAT generalized cross-correlation-phase transform algorithm, and the direction of the sound source is solved by the least squares method in combination with the uniform circular array to obtain the second ambient sound signal data;

[0077] The MVDR minimum variance distortionless response beamforming algorithm is used to enhance the signal of the sound source direction in the second ambient sound signal data, suppress interference from other directions, and obtain the source data of the sound source signal.

[0078] Step 102: extracting multi-dimensional feature data from the sound source signal source data, performing multi-modal feature fusion on the multi-dimensional feature data, and obtaining multi-dimensional vector data;

[0079] Specifically, in this embodiment, multi-dimensional feature data is extracted from the sound source signal source data, including at least time domain features, frequency domain features and dynamic features;

[0080] The time domain features include at least short-time energy, zero-crossing rate, amplitude envelope dynamic range and autocorrelation function peak ratio; the frequency domain features include at least Mel frequency cepstral coefficients, spectrum centroid and spectrum flux; the dynamic features include at least the first-order difference of the dynamic change rate of the time domain features and the frequency domain features and the second-order difference obtained by performing a further difference calculation on the feature data after the first-order difference;

[0081] A unified feature vector is constructed for the time domain features, frequency domain features and dynamic features in the multi-dimensional feature data. The mean variance of each feature dimension is normalized, and the multi-dimensional feature vector is organized by frame. The features of the current frame and the two frames before and after are taken, and the statistics within the window are calculated to obtain multi-dimensional vector data.

[0082] Step 103: Establish an LSTM-TCN user behavior recognition model using the LSTM-TCN neural network architecture, and use the MOPSO multi-objective particle swarm optimization algorithm to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and light effect decision to obtain a target LSTM-TCN user behavior recognition model.

[0083] Specifically, in this embodiment, spatiotemporal joint modeling is performed based on the LSTM fusion time series modeling and TCN local feature extraction capabilities. The network structure of the model includes at least an input layer, an LSTM module, a TCN module, and a classification head.

[0084] The input layer is used to receive historical multi-dimensional vector data in the system, and its time window length is 10 frames;

[0085] The LSTM module includes a bidirectional LSTM layer and an output layer. The bidirectional LSTM layer is used to capture long-term dependencies:

[0086] ;

[0087] Among them, the number of hidden units in the LSTM module The optimization is performed using the MOPSO multi-objective particle swarm optimization algorithm with a range of (32,256); Represents the hidden state at the previous moment, Represents the input features at the current moment, Represents the output of the bidirectional LSTM layer;

[0088] The output layer of the LSTM module is used to concatenate the time step outputs into , represents the time step, represents a bidirectional LSTM;

[0089] The size of the dilated convolution kernel in the TCN module , expansion factor ,in is the number of layers, depth ; The residual connection and weight normalization calculation formula of the TCN module are as follows:

[0090] ;

[0091] The output layer of the TCN module: ,in is the number of channels, Represented as a convolution operation, represents the activation function;

[0092] The classification head includes global average pooling + fully connected layer, and the calculation formula is as follows:

[0093] ;

[0094] in, represents global average pooling, represents the weight of the fully connected layer, Indicates bias, Indicates calculating classification probability output.

[0095] Establish the optimization framework of MOPSO multi-objective particle swarm optimization algorithm:

[0096] ;

[0097] in, Represents the hyperparameters of the target LSTM-TCN user behavior recognition model; Represents MFCC feature weight; Indicates lighting effect decision; represents the number of hidden units, Indicates the size of the convolution kernel of the dilated convolution stack; represents the depth of the dilation factor, Represents the learning rate; MFCC feature weight ,and ; (0.1, 0.2) represents the light effect parameter of intensity gain; A lighting parameter that indicates the ratio of hot to cold color temperature.

[0098] Initialize the particle swarm of the MOPSO multi-objective particle swarm optimization algorithm, set the swarm size to 50, and randomly generate particles that meet the constraints;

[0099] The calculation formula of the particle swarm velocity update rule is as follows:

[0100] ;

[0101] in, Represents the inertia weight, and decreases linearly, and represents the acceleration factor, and Belonging to the set (0,1) represents a random number; Indicates the In the iteration, Particles in Dimensional speed; Indicates the In the iteration, Particles in Dimensional speed; Indicates the Particles in The optimal historical position of individuals in the dimension; Indicates that the global optimal particle is The position of the dimension guides the search direction of the particle swarm; Indicates the In the iteration, Particles in The current position of the dimension;

[0102] The particle swarm is constrained. If the sum of the weights is not equal to 1, it is rescaled and the out-of-range parameters of the particle swarm are reflected back to the feasible domain by mirroring.

[0103] The particle swarm is maintained at the Pareto frontier. The top 20% of non-inferior particles are retained for each iteration to enter the next iteration. When the number of iterations is equal to 150, the optimal particle swarm is output and replaced as the hyperparameter of the model to obtain the target LSTM-TCN user behavior recognition model.

[0104] Step 104: Use the DAGAN deep generative adversarial network to perform noise compensation on the multi-dimensional vector data and then input it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention;

[0105] Specifically, in this embodiment, multi-dimensional vector data is obtained and a DAGAN deep adversarial generative network architecture is designed, which includes at least a generator and a discriminator. The generator is used to map the noisy feature vector to a clean feature space, and the discriminator is used to distinguish the repaired features output by the generator from the real clean features.

[0106] The feature statistics of non-speech segments are used to establish a noise baseline. When the features of the multi-dimensional vector data deviate from the baseline by more than a threshold, the DAGAN deep adversarial generative network is triggered to compensate.

[0107] Compensate for the high-frequency MFCC coefficients most affected by noise, and dynamically assign restoration weights through the attention layer within the generator. The discriminator outputs a restoration credibility score. When the score is lower than 0.7, the current frame is discarded and historical data interpolation is used.

[0108] The denoised features output by DAGAN are organized into sequences according to time windows, and the compensation confidence is added as an additional input channel. The sequence is then fed into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention.

[0109] Step 105 : Intelligently control the LED light strip based on the user's real-time intention, including at least dynamic light mapping and light scene adaptation.

[0110] In this embodiment,

[0111] Dynamic Light Mapping

[0112] Create a sound feature-light parameter mapping table:

[0113] Pitch → Adjust the color temperature of the light strip (high pitch maps to cool white light, low pitch maps to warm yellow light).

[0114] Sound intensity → Controls the brightness level (soft commands reduce brightness, loud commands increase brightness).

[0115] Speed ​​of the rhythm → trigger dynamic effects (fast rhythm activates "Party Flash Mode", slow rhythm activates "Breathing Gradient Mode").

[0116] Step 4.2: Scene Adaptive Control

[0117] Preset scene mode library:

[0118] Home mode: Automatically illuminate the path light strip based on the location of footsteps.

[0119] Entertainment mode: Synchronize the light strip color changes with the music rhythm (extract the music beat characteristics to drive the light).

[0120] Security mode: Abnormal sounds (such as glass breaking) trigger the light strip to flash at high frequency to alarm.

[0121] Based on ambient light sensor data, dynamically adjust lighting parameters to avoid overexposure or glare.

[0122] Its beneficial effects include: 1. Enhanced anti-interference capabilities: Through dual-stage processing of microphone array beamforming and DAGAN noise compensation, the sound source localization error is ≤3° in extreme environments with a signal-to-noise ratio as low as -5dB, improving intent recognition accuracy and reducing false trigger rates. 2. Real-time response optimization: The MOPSO algorithm is used to jointly optimize the hyperparameters and feature weights of the LSTM-TCN model, compressing inference latency. Combined with pipeline parallel computing, the overall system response time is shortened to meet real-time interaction requirements. 3. Enhanced personalized adaptation: The introduction of a light effect decision layer and a user preference learning mechanism improves user satisfaction by dynamically adjusting color temperature and intensity parameters and supports smooth switching between customized scenes. 4. Improved resource efficiency: The model parameter count is reduced by 50% (1.8M → 0.9M), supports 8-bit quantization deployment, reduces memory usage by 60%, and can run on ARM Cortex-M7-class chips with power consumption of ≤100mW. 5. Adaptive robustness: DAGAN's online incremental learning and noise environment perception module enables the system to adaptively converge within 24 hours after the emergence of new interference, extending its lifecycle.

[0123] In this embodiment, please refer to Figure 2 In a second embodiment of the intelligent control method for LED light strips based on sound changes according to an embodiment of the present invention, ambient sound signal data from an indoor microphone array is obtained, and the direction of the sound source in the ambient sound signal data is identified by a beamforming algorithm to obtain source data of the sound source signal, including the following steps:

[0124] Step 201: Acquire ambient sound signal data from an indoor microphone array, where the microphone array is at least arranged on a ceiling, a wall, and a floor;

[0125] Step 202: Frame the ambient sound signal data using a Hamming window, where the frame length is 20 ms and the frame shift is 10 ms; perform short-time Fourier transform on the frame-processed data to obtain first ambient sound signal data;

[0126] Step 203: Calculate the time difference between the sound source in the first ambient sound signal data reaching different microphones based on the GCC-PHAT generalized cross-correlation-phase transform algorithm, and solve the sound source direction using the least squares method in combination with the uniform circular array to obtain the second ambient sound signal data;

[0127] Step 204: Use the MVDR minimum variance distortionless response beamforming algorithm to enhance the signal of the sound source direction in the second ambient sound signal data, suppress interference from other directions, and obtain source data of the sound source signal.

[0128] Its beneficial effect is enhanced anti-interference capability: through the dual-stage processing of microphone array beamforming and DAGAN noise compensation, in extreme environments with a signal-to-noise ratio as low as -5dB, the sound source positioning error is ≤3°, improving the accuracy of intent recognition and reducing the false trigger rate.

[0129] In this embodiment, please refer to Figure 3 In a third embodiment of the sound-based intelligent control method and system for LED light strips according to the present invention, temperature control of a power supply control system of smart clothing based on multi-step temperature prediction results includes the following steps:

[0130] Step 301: Obtain multi-dimensional vector data and design a DAGAN deep adversarial generative network architecture, which includes at least a generator and a discriminator. The generator is used to map noisy feature vectors to a clean feature space, and the discriminator is used to distinguish between the restored features output by the generator and the true clean features.

[0131] Step 302: Use the feature statistics of the non-speech segment to establish a noise baseline. When it is detected that the features of the multi-dimensional vector data deviate from the baseline by more than a threshold, trigger the DAGAN deep adversarial generative network to compensate.

[0132] Step 303: Compensate for the high-frequency MFCC coefficients most affected by noise, and dynamically assign restoration weights through the attention layer within the generator. The discriminator outputs a restoration credibility score. When the score is lower than 0.7, the current frame is discarded and historical data interpolation is used.

[0133] Step 304: Organize the denoised features output by DAGAN into a sequence according to the time window, add the compensation confidence as an additional input channel, and input it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention.

[0134] The above describes the LED light strip intelligent control method based on sound changes provided by the embodiment of the present invention. The following describes the LED light strip intelligent control system based on sound changes according to the embodiment of the present invention. Figure 4 In one embodiment of the present invention, an intelligent control system for LED light strips includes:

[0135] The environmental data acquisition module is used to obtain the environmental sound signal data from the indoor microphone array, identify the sound source direction in the environmental sound signal data through the beamforming algorithm, and obtain the source data of the sound source signal;

[0136] The data feature fusion module is used to extract multi-dimensional feature data from the source data of the sound source signal, perform multi-modal feature fusion on the multi-dimensional feature data, and obtain multi-dimensional vector data;

[0137] The recognition model building module is used to build an LSTM-TCN user behavior recognition model using the LSTM-TCN neural network architecture. The MOPSO multi-objective particle swarm optimization algorithm is used to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and lighting effect decisions to obtain the target LSTM-TCN user behavior recognition model.

[0138] The user intent recognition module uses the DAGAN deep generative adversarial network to compensate for noise in multi-dimensional vector data and then inputs it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intent;

[0139] The light strip intelligent control module is used to intelligently control the LED light strip based on the user's real-time intention, including at least dynamic light mapping and light scene adaptation.

[0140] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above-described embodiments. The above-described embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention claimed.

Claims

1. An intelligent control method for LED light strips based on sound changes, characterized in that: The LED light strip intelligent control method comprises the following steps: Acquire ambient sound signal data from an indoor microphone array, identify the direction of a sound source in the ambient sound signal data using a beamforming algorithm, and obtain sound source signal source data; Extracting multi-dimensional feature data from the sound source signal source data, and performing multimodal feature fusion on the multi-dimensional feature data to obtain multi-dimensional vector data; An LSTM-TCN user behavior recognition model was established using the LSTM-TCN neural network architecture. The MOPSO multi-objective particle swarm optimization algorithm was used to perform three-level optimization on the model's network hyperparameters, MFCC feature weights, and lighting effect decisions, resulting in the target LSTM-TCN user behavior recognition model. The multi-dimensional vector data is noise-compensated using the DAGAN deep generative adversarial network and then input into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention; Intelligent control of the LED light strip is performed based on the real-time intention of the user, including at least dynamic light mapping and light scene adaptation.

2. The LED light strip intelligent control method based on sound changes according to claim 1, characterized in that: The step of obtaining ambient sound signal data from an indoor microphone array, identifying a sound source direction in the ambient sound signal data by a beamforming algorithm, and obtaining sound source signal source data includes: Acquiring ambient sound signal data from an indoor microphone array, wherein the microphone array is at least arranged on a ceiling, a wall, and a floor; Using a Hamming window to perform frame processing on the ambient sound signal data, and performing short-time Fourier transform on the frame-processed data to obtain first ambient sound signal data; Calculating the time difference between the sound source in the first ambient sound signal data reaching different microphones based on the GCC-PHAT generalized cross-correlation-phase transform algorithm to obtain the second ambient sound signal data; The MVDR minimum variance distortionless response beamforming algorithm is used to enhance the sound source direction in the second ambient sound signal data to obtain sound source signal source data.

3. The LED light strip intelligent control method based on sound changes according to claim 1, characterized in that: The extracting multi-dimensional feature data from the sound source signal source data and performing multi-modal feature fusion on the multi-dimensional feature data to obtain multi-dimensional vector data includes: Extracting multi-dimensional feature data from the sound source signal source data, including at least time domain features, frequency domain features and dynamic features; A unified feature vector is constructed for the time domain features, frequency domain features and dynamic features in the multi-dimensional feature data, mean variance normalization is performed on each feature dimension, and the multi-dimensional feature vector is organized by frame to obtain multi-dimensional vector data.

4. The LED light strip intelligent control method based on sound changes according to claim 1, characterized in that: The LSTM-TCN user behavior recognition model is established using the LSTM-TCN neural network architecture. The MOPSO multi-objective particle swarm optimization algorithm is used to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and light effect decisions to obtain the target LSTM-TCN user behavior recognition model, including: Based on the LSTM fusion time series modeling and TCN local feature extraction capabilities, spatiotemporal joint modeling is performed. The network structure of the model includes at least an input layer, an LSTM module, a TCN module, and a classification head. The input layer is used to receive historical multi-dimensional vector data in the system, and its time window length is 10 frames; The LSTM module includes a bidirectional LSTM layer and an output layer, and the bidirectional LSTM layer is used to capture long-term dependencies.

5. The LED light strip intelligent control method based on sound changes according to claim 4, characterized in that: The MOPSO multi-objective particle swarm optimization algorithm includes: Establish the optimization framework of MOPSO multi-objective particle swarm optimization algorithm: ; in, Represents the hyperparameters of the target LSTM-TCN user behavior recognition model; Represents MFCC feature weight; Indicates lighting effect decision; represents the number of hidden units, Indicates the size of the convolution kernel of the dilated convolution stack; represents the depth of the dilation factor, Represents the learning rate; MFCC feature weight ,and ; (0.1, 0.2) represents the light effect parameter of intensity gain; A lighting parameter that indicates the ratio of hot to cold color temperature.

6. The LED light strip intelligent control method based on sound changes according to claim 4, characterized in that: The MOPSO multi-objective particle swarm optimization algorithm is used to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and light effect decisions to obtain the target LSTM-TCN user behavior recognition model, and further includes: Initialize the particle swarm of the MOPSO multi-objective particle swarm optimization algorithm, set the swarm size to 50, and randomly generate particles that meet the constraints; The particle swarm is constrained. If the sum of the weights is not equal to 1, it is rescaled and the out-of-range parameters of the particle swarm are reflected back to the feasible domain by mirroring. The particle swarm is maintained at the Pareto frontier. The top 20% of non-inferior particles are retained for the next iteration after each iteration. When the number of iterations is equal to 150, the optimal particle swarm is output and replaced as the hyperparameters of the model to obtain the target LSTM-TCN user behavior recognition model.

7. The LED light strip intelligent control method based on sound changes according to claim 1, characterized in that: The DAGAN deep generative adversarial network is used to perform noise compensation on the multi-dimensional vector data and then input it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention, including: Obtain multi-dimensional vector data and design a DAGAN deep adversarial generative network architecture, which includes at least a generator and a discriminator. The generator is used to map noisy feature vectors to a clean feature space, and the discriminator is used to distinguish between the restored features output by the generator and the true clean features. A noise baseline is established using feature statistics of non-speech segments. When it is detected that the features of the multi-dimensional vector data deviate from the baseline by more than a threshold, a DAGAN deep adversarial generative network is triggered to compensate. Compensate for the high-frequency MFCC coefficients most affected by noise, and dynamically assign restoration weights through the attention layer within the generator. The discriminator outputs a restoration credibility score. When the score is lower than 0.7, the current frame is discarded and historical data interpolation is used. The denoised features output by DAGAN are organized into sequences according to time windows, and the compensation confidence is added as an additional input channel. The sequence is then fed into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention.

8. The LED light strip intelligent control system based on sound changes is characterized by: The LED light strip intelligent control system includes the following modules: An environmental data acquisition module is used to obtain environmental sound signal data from an indoor microphone array, identify the direction of the sound source in the environmental sound signal data through a beamforming algorithm, and obtain sound source signal source data; A data feature fusion module is used to extract multi-dimensional feature data from the source data of the sound source signal, perform multi-modal feature fusion on the multi-dimensional feature data, and obtain multi-dimensional vector data; The recognition model building module is used to build an LSTM-TCN user behavior recognition model using the LSTM-TCN neural network architecture. The MOPSO multi-objective particle swarm optimization algorithm is used to perform three-layer optimization on the model's network hyperparameters, MFCC feature weights, and lighting effect decisions to obtain the target LSTM-TCN user behavior recognition model. The user intention recognition module is used to use the DAGAN deep adversarial generative network to compensate for the noise of the multi-dimensional vector data and then input it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention; The light strip intelligent control module is used to intelligently control the LED light strip based on the user's real-time intention, including at least dynamic light mapping and light scene adaptation.

9. The LED light strip intelligent control system based on sound changes according to claim 8, characterized in that: The environmental data acquisition module includes the following submodules: An acquisition submodule is used to acquire ambient sound signal data from an indoor microphone array, wherein the microphone array is at least arranged on the ceiling, the wall and the ground; A frame processing submodule is used to perform frame processing on the ambient sound signal data using a Hamming window, and perform short-time Fourier transform on the frame-processed data to obtain first ambient sound signal data; a time difference calculation submodule, configured to calculate the time difference between the sound source in the first ambient sound signal data reaching different microphones based on a GCC-PHAT generalized cross-correlation-phase transform algorithm to obtain second ambient sound signal data; The data acquisition submodule is used to enhance the sound source direction in the second environmental sound signal data by using the MVDR minimum variance distortionless response beamforming algorithm to obtain the source data of the sound source signal.

10. The LED light strip intelligent control system based on sound changes according to claim 8, characterized in that: The user intention recognition module has the following submodules: The acquisition submodule is used to obtain multi-dimensional vector data and design the DAGAN deep adversarial generative network architecture, which includes at least a generator and a discriminator. The generator is used to map the noisy feature vector to a clean feature space, and the discriminator is used to distinguish the repaired features output by the generator from the real clean features. A detection submodule is configured to establish a noise baseline using feature statistics of non-speech segments, and trigger a DAGAN deep adversarial generative network to compensate when it is detected that the features of the multi-dimensional vector data deviate from the baseline by more than a threshold. The compensation submodule is used to compensate for the high-frequency MFCC coefficients that are most affected by noise, and dynamically allocates the repair weights through the attention layer inside the generator; The discriminator outputs a restoration credibility score. When the score is lower than 0.7, the current frame is discarded and the historical data interpolation is used; A submodule is obtained, which is used to organize the denoised features output by DAGAN into sequences according to time windows, add compensation confidence as an additional input channel, and input it into the target LSTM-TCN user behavior recognition model to obtain the user's real-time intention.