Method, system, and storage medium for offline voice control of light freeform composition patterns
By using an offline voice control system that utilizes voice activity detection and local recognition technology, combined with preset and custom modes to drive LED light panels, the flexibility and personalization issues of existing low-power lighting control systems are solved, enabling low-power, high-response complex pattern combinations and personalized displays.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HAIBEN ELECTRONIC TECH CO LTD
- Filing Date
- 2026-06-09
- Publication Date
- 2026-07-21
AI Technical Summary
Existing lighting control systems struggle to achieve flexible and complex pattern combinations in low-power standby mode, and lack support for user-customized patterns, resulting in users being unable to obtain a rich and varied lighting visual experience in environments without internet access or battery-powered scenarios.
The system employs an offline voice control method, continuously monitoring ambient sound signals through a voice activity detection module. This activates the offline voice recognition module for local recognition, combining a preset scene mode library with user-defined modes to drive the LED dot matrix light panel to display corresponding light patterns or effects. It supports offline recognition in both Chinese and English and multiple lighting modes.
It enables flexible control of complex pattern combinations and personalized custom patterns under low power conditions, improves the user interaction experience, avoids network dependence and privacy leakage risks, and has a response latency of less than 200ms.
Smart Images

Figure CN122438232A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent lighting control technology, specifically to a method, system, and storage medium for offline voice-controlled free combination of light patterns. Background Technology
[0002] Existing lighting control technologies typically rely on cloud servers for voice command parsing and processing. Users issue voice commands via smart speakers or networked devices; the signals are transmitted over the network to the cloud for recognition, and then control commands are sent to the lighting terminal. In addition, some existing technologies also provide offline voice control solutions based on local microcontrollers. These solutions use pre-set wake words and a simple command dictionary to perform basic operations such as switching on / off and changing the color of local LED strips or bulbs. These solutions achieve voice interaction functionality to a certain extent and establish basic hardware connections and control logic.
[0003] Existing lighting control solutions mainly have the following problems:
[0004] First, voice control relies on cloud-based recognition. Most mainstream voice-controlled lighting solutions on the market currently depend on cloud-based voice recognition services (such as connecting a smart speaker to a cloud server for voice recognition). These solutions require a continuous network connection. When the network is interrupted or unstable, the voice control function completely fails. Furthermore, user voice data needs to be uploaded to a cloud server for processing, posing a risk of privacy leaks. In addition, the response latency of cloud recognition is typically over 500 milliseconds, resulting in a poor user experience.
[0005] Secondly, some offline voice solutions have limited functionality. While some existing products offer offline voice control, such as the Chinese utility model patent CN216414636U which discloses an intelligent offline voice-controlled RGBW light strip controller, it only supports simple on / off and color switching operations. It cannot display complex light pattern combinations, lacks scene mode presets, and cannot freely combine light patterns according to user needs. Furthermore, the voice wake-up modules in existing offline voice solutions have high power consumption, typically exceeding 50mW during continuous operation, making them unsuitable for battery-powered portable lighting devices.
[0006] It is evident that in existing technologies, lighting control systems often struggle to achieve flexible and complex pattern combination control in low-power standby mode, and lack effective support for user-customized patterns, resulting in users being unable to obtain a rich and varied lighting visual experience in environments without internet access or battery-powered scenarios. Summary of the Invention
[0007] This invention provides a method, system, and storage medium for offline voice-controlled light patterns that can freely combine patterns. It can solve the technical problems of existing light control systems that are difficult to achieve flexible and complex pattern combination control in low-power standby mode and lack support for user-customized patterns.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A first aspect of the present invention provides a method for offline voice-controlled free combination of light patterns, comprising the following steps:
[0010] In standby mode, the voice activity detection module continuously monitors ambient sound signals;
[0011] When the voice activity detection module detects a human voice signal that meets the preset conditions, the offline voice recognition module is activated.
[0012] The offline speech recognition module recognizes the collected speech data to obtain speech commands; the speech commands include wake word commands or control command commands.
[0013] If the command is identified as a control command, the control command is parsed and the corresponding lighting control strategy is matched from the preset scene mode library.
[0014] The LED dot matrix light panel is driven to display corresponding light patterns or effects according to the lighting control strategy.
[0015] The scene mode library stores at least one festival mode, at least one party mode, and a user-defined mode; the user-defined mode supports the generation of dot matrix graphic data uploaded through an external terminal.
[0016] In an optional embodiment, in step S1, the voice activity detection module employs a two-level wake-up mechanism, specifically including:
[0017] Level 1 detection: The energy value of ambient sound is monitored through an ultra-low power energy detection circuit. When the energy value exceeds the first threshold, the second level detection is triggered.
[0018] The second level of detection: By using a human voice duration verification algorithm or spectral feature analysis, it is determined whether the sound signal is a human voice. If so, a wake-up signal is generated to activate the offline speech recognition module.
[0019] In an optional embodiment, in step S3, the offline speech recognition module supports offline recognition of both Chinese and English.
[0020] The offline speech recognition module has a pre-built Chinese wake-up word library, an English wake-up word library, a Chinese command word library, and an English command word library;
[0021] The offline speech recognition module is configured to match speech commands from the corresponding dictionary based on the user-defined language pattern or automatic recognition mode.
[0022] In an optional embodiment, in step S4, the holiday mode includes at least one of the following: birthday mode, Valentine's Day mode, Children's Day mode, and Halloween mode.
[0023] Each festival mode corresponds to a specific combination of light colors, flashing frequency, and dot matrix pattern;
[0024] Step S5 specifically includes: according to the matched holiday mode, controlling each LED bead on the LED dot matrix light board to light up or turn off according to a preset time sequence, so as to dynamically display the corresponding dot matrix pattern.
[0025] In an optional embodiment, the method further includes a user-defined pattern construction step, specifically comprising:
[0026] Receive image data from an external terminal;
[0027] The image data is downsampled and converted into a pixel matrix that matches the resolution of the LED dot matrix light panel;
[0028] The pixel matrix is converted into lighting control data and stored in the scene mode library as a user-defined mode;
[0029] The downsampling process includes color quantization and dithering algorithms to adapt to the color display capabilities of the LED panel.
[0030] In an optional embodiment, the control command instructions further include basic control instructions, which include switch control, color control, brightness adjustment, and mode switching.
[0031] When a color control command is detected, the overall color tone of the LED dot matrix display is adjusted.
[0032] When a brightness adjustment command is detected, the PWM duty cycle of the LED dot matrix light board is adjusted to change the brightness;
[0033] When a mode switching command is detected, different sub-patterns or effects are cyclically switched under the currently selected mode type.
[0034] In an optional embodiment, in step S5, if the currently matched mode is a party mode, the method further includes:
[0035] Real-time acquisition of ambient audio signals via audio input interface;
[0036] Extract the rhythmic features or frequency components of the environmental audio signal;
[0037] The flashing frequency, brightness variation amplitude, or color change speed of the LED dot matrix light panel are dynamically adjusted according to the rhythm characteristics or frequency components to achieve synchronized changes in lighting effects with the rhythm of the music.
[0038] In one alternative embodiment, the method is run on a low-power microcontroller unit (MCU).
[0039] In the standby state of step S1, only the voice activity detection module works, and the overall power consumption of the system is less than 16mW;
[0040] After the offline speech recognition module is activated in step S2, the system enters normal working state with a response delay of less than 200ms.
[0041] A second aspect of the present invention provides a system for offline voice-controlled free combination of light patterns, comprising:
[0042] The voice acquisition module is used to collect ambient sound signals;
[0043] A voice activity detection module, connected to the voice acquisition module, is used to continuously monitor ambient sound signals in standby mode and generate a wake-up signal when a human voice signal that meets preset conditions is detected.
[0044] An offline speech recognition module, connected to the speech activity detection module, is activated upon receiving the wake-up signal and performs offline speech data recognition to obtain speech commands.
[0045] The main control module is connected to the offline voice recognition module and the LED driver module respectively, and is used to parse the voice commands, retrieve the corresponding lighting control strategy from the memory, and generate control signals;
[0046] The LED driver module is used to drive the LED dot matrix light board to display corresponding light patterns according to the control signal;
[0047] A memory for storing a scene mode library, the scene mode library including at least one holiday mode, at least one party mode and user-defined modes;
[0048] The user-defined mode supports the generation of dot matrix graphic data uploaded by an external terminal and received via a wireless communication module.
[0049] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for offline voice-controlled free combination of light patterns as described in any one of claims 1 to 8.
[0050] The solution provided by this invention continuously monitors the system in standby mode via a voice activity detection module, activating the offline voice recognition module only when human voice is detected. During normal monitoring, the voice recognition circuit does not need to operate at full speed, significantly reducing standby power consumption. Offline voice recognition is performed locally, eliminating the need for network transmission and avoiding latency and privacy risks. After parsing, control commands are matched with lighting control strategies from a scene mode library. This library includes holiday modes, party modes, and user-defined modes that support uploading dot matrix graphic data from external terminals, allowing for flexible access to diverse pattern resources. Finally, the matched strategy drives the LED dot matrix light panel to display the corresponding lighting pattern or effect, achieving a complete closed loop from voice input to complex dynamic pattern output. This solves the problem of existing technologies being unable to simultaneously handle complex pattern combination control and personalized customization needs under low power conditions, improving the adaptability and interactive experience of intelligent lighting systems in different scenarios. Attached Figure Description
[0051] Figure 1 A flowchart of a method for offline voice-controlled free combination of light patterns provided by the present invention;
[0052] Figure 2 The flowchart illustrates the two-level wake-up mechanism for voice activity detection in the offline voice-controlled light free combination pattern method provided by the present invention.
[0053] Figure 3 A flowchart illustrating the construction of a user-defined pattern in the offline voice-controlled light free combination pattern method provided by the present invention;
[0054] Figure 4 The flowchart shows the party mode in the offline voice-controlled light free combination pattern method provided by the present invention. Detailed Implementation
[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] Please refer to Figure 1 for a flowchart of a method for offline voice-controlled free combination of light patterns provided in an embodiment of the present invention. The method includes the following steps:
[0057] Step S1: In standby mode, the ambient sound signal is continuously monitored through the voice activity detection module;
[0058] The Voice Activity Detection (VAD) module is a processing unit used to determine the presence of human voice activity in an audio stream. Acting as a front-end monitoring sentinel for the system, it remains operational even when the system is in low-power standby mode. This module acquires analog sound signals from the environment and converts them into digital signals, calculating the signal's energy value, zero-crossing rate, or spectral characteristics in real time. The monitored object is continuous ambient sound signals, which can originate from a built-in microphone array or a single microphone. This module operates with a low sampling frequency and computational complexity, maintaining overall system power consumption at a low level (e.g., below 16mW) in battery-powered scenarios.
[0059] Step S2: When the voice activity detection module detects a human voice signal that meets the preset conditions, the offline voice recognition module is activated;
[0060] A voice signal meeting preset conditions refers to a signal whose energy value exceeds a specific threshold and whose duration or spectral characteristics conform to a human voice statistical model after analysis by the voice activity detection module. Preset conditions are set according to the actual application scenario, such as setting an energy threshold to exclude slight background noise, or setting a minimum duration to exclude brief impulse noise. The offline voice recognition module is a voice decoding engine embedded in the local microcontroller. Once activated, it performs semantic analysis on the voice data. The activation mechanism is triggered by a hardware interrupt signal or software flag generated by the voice activity detection module. Only when valid human voice is confirmed to have been captured will the system wake up the offline voice recognition module from sleep mode, switching it from low-power mode to full-speed operating mode.
[0061] For example, when there is continuous conversation or clear shouting in the environment, the voice activity detection module detects that the signal energy is consistently above -40dBFS for 50ms, and the spectral center is located in the 300Hz-3400Hz human voice frequency band. This is then determined to be a human voice signal that meets preset conditions, and a wake-up command is sent to activate the offline voice recognition module. If it is only wind noise or low-frequency noise from running appliances, the conditions are not met, and the offline voice recognition module remains off. This two-stage linkage mechanism achieves a balance between power consumption and response speed, ensuring that user commands are not missed while minimizing energy consumption.
[0062] This step enables the system to dynamically switch from listening mode to recognition mode. The execution entity is the system's power management unit or main control MCU, which performs power-on or clock-enable operations based on the judgment results of the preceding modules. This improves the system's energy efficiency, enabling long-term standby based on battery power, while ensuring that high computing resources are only consumed when real human voices are present.
[0063] Step S3: The collected voice data is recognized by the offline speech recognition module to obtain voice commands; voice commands include wake word commands or control command commands.
[0064] The offline speech recognition module is a computing unit that can perform acoustic model matching and language model decoding locally without an internet connection. Voice commands are control requests with specific semantics issued by the user in natural language. Wake-up words are specific word sequences used to further confirm the user's intent and fully activate system control (such as "Xiao Guang Xiao Guang" or "Hello Light"); control commands are specific operation instructions issued by the user after wake-up or directly (such as "Turn on birthday mode," "Turn red," etc.). Voice commands originate from digital speech data streams after analog-to-digital conversion. The offline speech recognition module has a pre-built bilingual dictionary containing Chinese and English, which can map the input speech waveform into corresponding text strings or command codes according to the user-defined language pattern or automatic recognition mode.
[0065] For example, when a user says "Turn on Party Mode", the offline speech recognition module matches the speech segment with a preset Chinese command dictionary and outputs the corresponding control command code CMD_VALENTINE; if the user says "Turn on Party Mode" in English, the module matches the English dictionary and outputs the corresponding English command code.
[0066] The core of this step lies in utilizing local computing power to achieve semantic understanding, without relying on cloud servers. The execution method involves algorithmic processes such as feature extraction, acoustic scoring, and decoding search. The resulting structured speech command data is directly used as input for subsequent logical parsing, avoiding the experience degradation caused by network latency and instability, while simultaneously ensuring the privacy and security of user voice data.
[0067] Step S4: If the control command is identified, the control command is parsed and the corresponding lighting control strategy is matched from the preset scene mode library;
[0068] The scene mode library is a data structure stored in non-volatile memory (such as Flash) used to map voice commands to lighting execution parameters. The lighting control strategy is a set of parameters defining the color, brightness, blink frequency, and time sequence of each LED on the LED matrix light board. The scene mode library stores at least one holiday mode, at least one party mode, and a user-defined mode. Holiday modes are preset fixed patterns and dynamic effects for specific holidays; party modes include dynamic algorithms that change with the music rhythm; user-defined modes support the generation of dot matrix graphic data uploaded from external terminals (such as smartphone apps), allowing users to personalize the displayed content. The parsing process compares the voice command code with the index key-value pair in the scene mode library; if a match is found, the corresponding lighting control strategy data is retrieved and loaded into video memory or a buffer.
[0069] Step S5: Drive the LED dot matrix light panel to display the corresponding light patterns or effects according to the lighting control strategy;
[0070] An LED dot matrix light panel is a display panel composed of multiple light-emitting diodes arranged in a matrix. Its resolution can be configured according to actual needs (such as 8x8, 16x16, etc.). The lighting control strategy contains all the information required to drive the LED dot matrix light panel, including the RGB color values of each LED, the PWM duty cycle, and the refresh time sequence. The driving process involves the main control module sending control signals to the LED driver chip through GPIO ports or dedicated communication interfaces (such as SPI, I2C) based on the loaded strategy data, controlling each LED to light up, turn off, or change color according to preset logic. For holiday mode, this manifests as a gradual change of static patterns or a specific sequence of flashing; for party mode, it manifests as a rapid color jump; and for user-defined mode, it accurately reproduces the pixel matrix graphics uploaded from an external terminal.
[0071] For example, when a birthday pattern is matched, the lighting control strategy instructs the central area of the LED dot matrix light panel to display yellow pixels in the shape of a cake, while the surrounding area flashes in a rainbow color cycle; when a user-uploaded custom heart pattern is matched, the system strictly follows the pixel coordinates in the uploaded data, controlling the corresponding LED beads to display red, while the rest are turned off.
[0072] This step is the final execution stage of the technical solution, transforming the digital control strategy into visual light and shadow effects. The execution is carried out by an LED driver module in conjunction with the main control MCU. Through precise timing control and color reproduction, the system can present various complex light patterns and dynamic effects, providing intuitive feedback to user voice commands and enhancing the immersive experience of human-computer interaction.
[0073] This invention, through the synergy of the aforementioned technical features, constructs a complete offline voice-controlled lighting combination scheme. The voice activity detection module and the offline voice recognition module work in a hierarchical manner. In standby mode, the system maintains low-power listening only; upon detecting a human voice signal, the recognition engine is activated. This mechanism reduces the high power consumption problem of traditional constantly-on solutions and eliminates dependence on network connections, ensuring real-time response and data privacy. Furthermore, the parsing module is tightly coupled with the scene mode library, allowing flexible invocation of preset holiday modes, party modes, or user-defined modes. The user-defined mode supports uploading dot matrix graphic data from external terminals. Combined with the driving capabilities of the LED dot matrix light board, it breaks through the limitations of fixed patterns in traditional lighting control, achieving true free combination. This hardware and software collaborative architecture provides personalized lighting display effects while maintaining low power consumption and high responsiveness, solving the problems of limited functionality and poor scalability in existing technologies.
[0074] In one embodiment, such as Figure 2 The diagram shows a flowchart of a method for voice activity detection using a two-level wake-up mechanism provided by an embodiment of the present invention. The method further includes further limitations on the specific implementation of the voice activity detection module in step S1.
[0075] Step S11: First-level detection: The energy value of ambient sound is monitored by an ultra-low power energy detection circuit. When the energy value exceeds the first threshold, the second-level detection is triggered.
[0076] The first-level detection utilizes hardware circuitry or low-power algorithms to acquire and quantify the amplitude or power of ambient sound waves in real time. The ultra-low-power energy detection circuit specifically includes a microphone preamplifier, a rectifier filter circuit, and a voltage comparator, with its operating current controlled at the microamp level. This circuit acts as the system's gatekeeper, continuously coarsely screening ambient sound signals. Its core function is to quickly eliminate background noise below a preset energy level, such as wind noise, distant vehicle sounds, or background noise from electronic devices. The first threshold is a dynamic or fixed value set based on the average ambient noise level in the actual application scenario. For example, in a quiet bedroom environment, the first threshold can be set to a voltage value corresponding to a 40 dB sound pressure level; while in a noisy living room environment, this threshold can be automatically adjusted to a voltage value corresponding to a 60 dB sound pressure level. When the detected ambient sound energy value instantaneously or repeatedly exceeds this first threshold, it indicates the possible presence of valid human voice input. At this time, the circuit outputs a high-level signal to trigger the second-level detection. Through this preliminary judgment based on energy values, the system does not need to activate complex digital signal processing units, maintaining extremely low standby power consumption during most silent or low-noise periods.
[0077] Step S12: Second-level detection: Determine whether the sound signal is a human voice by using a human voice duration verification algorithm or spectral feature analysis. If so, generate a wake-up signal to activate the offline speech recognition module.
[0078] The second-level detection is a refined recognition process initiated after the first-level detection is triggered. It addresses the issue of false wake-ups caused by interference from sudden noises (such as clapping or door slamming) when relying solely on energy detection. The voice duration verification algorithm counts the duration of a sound signal exceeding the energy threshold. If this duration falls within the typical time window of human speech (e.g., 200ms to 2s), it is identified as a potential human voice; if the duration is too short (e.g., a sudden impact sound) or too long (e.g., continuous wind noise), it is filtered out. Spectral feature analysis performs a Fast Fourier Transform on the acquired sound signal to extract frequency domain features, focusing on the distribution of signal energy in the main human voice frequency band from 300Hz to 3400Hz and the variation of the fundamental frequency. For example, if a sound signal is detected that not only exceeds the first threshold in energy but also has a duration of approximately 800ms, and spectral analysis shows obvious formant features at 500Hz, 1000Hz, and 2000Hz, consistent with the spectral distribution of vowel pronunciation, then the signal is identified as a human voice. Once a human voice is detected, the system immediately generates a wake-up signal and sends it to the power management pin or interrupt pin of the offline speech recognition module, switching it from sleep mode to full-speed operation. Through the combination of ultra-low power energy detection in the first stage and high-precision feature analysis in the second stage, the system's sensitivity to weak human voices is ensured while preventing invalid wake-ups caused by non-human noise, thus reducing additional power consumption due to false wake-ups.
[0079] This invention achieves a balance between power consumption control and recognition accuracy through the synergy of the aforementioned two-stage wake-up mechanism. The ultra-low power energy detection circuit, as the first line of defense, filters out most meaningless background noise at a minimal energy cost, ensuring the system consumes only a small amount of power during long standby periods. Building upon this, a human voice duration verification algorithm or spectral feature analysis serves as the second line of defense, utilizing the unique time and frequency domain attributes of human voice for secondary confirmation, eliminating interference signals with high energy but not human voice (such as the sound of a heavy object falling or the start / stop sound of an air conditioner). This progressively layered filtering logic ensures that the high-power offline voice recognition module is only activated when a valid command is generated with a very high probability, avoiding energy waste and response delays caused by frequent start / stop cycles. Ultimately, this mechanism maintains the overall standby power consumption of the system at a low level (e.g., below 16mW), while simultaneously improving the robustness of voice control, ensuring that users can stably control the free combination of light patterns via voice commands in various complex acoustic environments.
[0080] In another optional embodiment, the method further includes extended configuration of speech recognition language patterns to support seamless interaction in multilingual environments.
[0081] The offline speech recognition module supports bilingual offline recognition in Chinese and English.
[0082] The offline speech recognition module integrates a software or hardware decoding engine inside a low-power microcontroller unit (MCU). Its function is to directly complete the conversion from acoustic wave signals to text instructions locally without connecting to an external network server. This module supports bilingual offline recognition in Chinese and English and integrates acoustic models and language models for two languages, namely Chinese Mandarin and standard English. By loading bilingual acoustic feature extraction algorithms, this module can automatically adapt to the pronunciation characteristics of different languages. For example, when processing Chinese, it focuses on the extraction of tone features, and when processing English, it focuses on the extraction of phoneme liaison and stress features. For example, users can say "Turn on the birthday mode" in Chinese or "Turn on BirthdayMode" in English, and the system can complete the parsing locally.
[0083] The offline speech recognition module internally pre-sets a Chinese wake-up word library, an English wake-up word library, a Chinese command word library, and an English command word library.
[0084] The word library is a structured data set stored in a non-volatile memory (such as Flash) and is used to define specific words recognizable by the system and their corresponding acoustic templates. The Chinese wake-up word library and the English wake-up word library store the acoustic fingerprints of keywords such as "Xiaoguang Xiaoguang" and "Hello Light" for activating the system respectively. The Chinese command word library and the English command word library store specific control instructions such as "Red", "Brighter", etc. and their variants. These word libraries are pre-written before factory shipment or through firmware upgrade. The generation process includes collecting a large number of audio samples with multiple accents and noises, and after training, quantifying them into a compact feature vector matrix to adapt to the storage limitations of embedded devices. For example, the Chinese command word library may include "Valentine's mode" and its synonym "Romantic mode", and the English command word library stores "Valentine'sMode" and "Romantic Mode" correspondingly. By storing wake-up words and command words independently by language, the system can cover instructions in two languages while keeping the memory occupancy low.
[0085] The offline speech recognition module is configured to be able to match voice instructions from the corresponding word library according to the language mode set by the user or the automatic recognition mode.
[0086] Language mode refers to the target language strategy adopted by the system during voice matching. This includes user-specified single language mode (Chinese only or English only) and automatic recognition mode determined by the system's intelligent judgment. User-defined language modes are typically determined through physical button combinations, APP configuration, or specific voice setting commands. Once set, the system will lock and load the corresponding single-language dictionary to reduce computational load. Automatic recognition mode involves the system detecting voice activity and then simultaneously or rapidly polling both Chinese and English acoustic models for scoring. The system automatically determines the language based on the highest confidence score and locks the matching range for subsequent command words. For example, in automatic recognition mode, if the user first says the English wake word "HelloLight," the system automatically switches the context to English and then only searches the English command dictionary for commands such as "PartyMode"; the reverse is also true. This flexible configuration allows the system to meet the rapid response needs of users accustomed to a fixed language while also adapting to free interaction scenarios in mixed language environments.
[0087] This invention constructs an offline multilingual voice control system through the synergy of the aforementioned technical features. By combining built-in Chinese and English bilingual recognition capabilities with pre-built dictionaries for each language, the system achieves real-time response to users speaking various languages without relying on cloud computing power. A dynamic switching mechanism between user-defined mode and automatic recognition mode allows the system to adjust resource allocation according to actual usage scenarios: focusing on single-language matching to improve speed and accuracy in environments with a defined language, and ensuring recognition success rate through dual-path verification in uncertain environments. This design reduces standby and operating power consumption while protecting user voice privacy data from local storage, enriching the international application scenarios of lighting control systems.
[0088] In one embodiment, the method further includes a more detailed execution process of steps S4 and S5 of the above embodiment.
[0089] In step S4, the holiday mode includes at least one of the following: birthday mode, Valentine's Day mode, Children's Day mode, and Halloween mode; each holiday mode corresponds to a specific combination of light colors, flashing frequency, and dot matrix pattern.
[0090] Holiday modes are collections of lighting configurations pre-stored in the scene mode library, associated with specific traditional holidays or anniversaries. Birthday mode is typically associated with the "Happy Birthday" voice command, and its corresponding light color combination can be set to a multi-colored alternation (e.g., red, yellow, blue, green cycle), with a flashing frequency set to a medium-speed, cheerful rhythm (e.g., 2Hz), and the default dot matrix graphic is a birthday cake or gift box shape. Valentine's Day mode is associated with the "Valentine's Day" command, with a color combination limited to a pink and red gradient, a flashing frequency simulating a heartbeat rhythm (e.g., a variable frequency flashing pattern that starts fast and then slows down), and a dot matrix graphic in the shape of a heart. Children's Day mode is associated with the "Children's Day" command, using a high-saturation rainbow color sequence, and the dot matrix graphic is a smiley face or balloon shape. Halloween mode is associated with the "Halloween" command, with a color combination primarily in orange and purple, and the dot matrix graphic is a pumpkin lantern or bat shape. These light color combinations, flashing frequencies, and dot matrix patterns are loaded from memory into the running memory during the initialization phase by the main control module, allowing for rapid recall upon recognition of corresponding voice commands. For example, when a user says "Activate birthday mode," the system immediately locks the birthday mode parameter set, where the color combination determines the RGB component output values of the LED beads, the flashing frequency determines the time interval for state transitions, and the dot matrix pattern defines which coordinate positions of the LED beads are lit. By binding holiday semantics with specific visual parameters in this way, the lighting equipment can create a specific holiday atmosphere without requiring users to manually adjust complex colors and dynamic effects.
[0091] Step S5 specifically includes: according to the matched holiday mode, controlling each LED bead on the LED dot matrix light board to light up or turn off according to a preset time sequence, so as to dynamically display the corresponding dot matrix pattern;
[0092] The preset time sequence is a collection of frames that transforms static dot matrix graphics into dynamic animation effects. Each frame contains the on / off state and color information of all LED beads at that moment. The control process is as follows: the main control module reads the frame data queue for the currently matched holiday mode and sends control signals to the LED driver module sequentially according to the set refresh rate (e.g., 30fps or dynamically adjusted based on the flashing frequency). For each time slot, the LED driver module drives the corresponding LED beads to light up or turn off based on the received signal, utilizing the persistence of vision to create a continuous dynamic image. In specific execution, if the current mode is birthday mode, the system not only displays a static cake graphic but also controls the LED beads in the candle section to simulate a flickering flame effect according to the preset time sequence; if it is Valentine's Day mode, it controls the LED beads around the heart-shaped graphic to exhibit a breathing light-like brightness change. This mechanism of controlling the on / off state of LED beads according to a time sequence, in close coordination with the aforementioned holiday mode parameters, transforms static graphic data into dynamic visual effects. For example, in Halloween mode, the eyes of the pumpkin lantern can repeatedly switch between on and off in a 0.5-second sequence to simulate a blinking effect. This precise timing control enhances the vividness and appeal of the light display.
[0093] This invention constructs a complete mechanism for creating a festive atmosphere through the synergy of the aforementioned technical features. By pre-setting multiple representative holiday modes (such as birthdays, Valentine's Day, Children's Day, and Halloween) and binding each mode with specific color combinations, flashing frequencies, and dot matrix graphics, the system can automatically retrieve and execute the lighting logic upon receiving voice commands. Using a preset time sequence to control the on / off state of each LED on the LED dot matrix light board, static graphics are transformed into dynamic animation effects, solving the problem of monotonous lighting patterns and lack of dynamic variation in existing technologies. This fully automated processing, from voice recognition to pattern matching to dynamic graphic rendering, avoids the tedious operation of manual programming or complex settings by users.
[0094] In another optional embodiment, as shown in Figure 3, a flowchart of constructing a user-defined mode is provided by an embodiment of the present invention. The method further includes a step of constructing the user-defined mode, specifically including:
[0095] Step S10: Receive image data from an external terminal;
[0096] External terminals are smart devices, such as smartphones, tablets, or laptops, that establish a communication connection with the lighting control system. Image data is the raw image file transmitted from the external terminal to this system via a wireless communication module (such as Bluetooth BLE or Wi-Fi), and its format includes JPEG, PNG, or BMP. This step obtains the personalized visual material that the user expects to display, serving as the source input for generating custom lighting patterns. The user can select photos from their local album through an application (APP) installed on the external terminal, or directly use the camera to capture live footage, and send the selected image data to the lighting control system through the application programming interface. The received image data typically has a pixel size and rich color depth far exceeding the physical resolution of the LED dot matrix light panel, and cannot be directly used to drive LED beads; it requires subsequent processing to adapt to the hardware display capabilities.
[0097] Step S20: Downsample the image data to convert it into a pixel matrix that matches the resolution of the LED dot matrix light panel;
[0098] Downsampling is the process of compressing high-resolution original image data spatially to a target resolution, which strictly corresponds to the physical arrangement specifications of the LED dot matrix light panel, such as 8x8, 16x16, or 32x32. This process involves not only reducing the image size but also color quantization and dithering algorithms to adapt to the limited color display capabilities of the LED light panel. Color quantization maps millions of colors in the original image to the limited color gamut that LED chips can display (such as a specific color set generated by mixing RGB colors), determining the closest usable color value through clustering algorithms or lookup tables. Dithering, building upon color quantization, uses error diffusion techniques (such as Floyd-Steinberg dithering) to distribute quantization errors to adjacent pixels, simulating richer midtones and gradient effects to the human eye, avoiding obvious color banding or blocky distortion. This downsampling process can transform arbitrarily complex natural images into discrete pixel matrices suitable for low-resolution LED array displays, preserving the main outlines and lighting features of the original image while ensuring data format compatibility with the hardware driver layer.
[0099] Step S30: Convert the pixel matrix into lighting control data and store it in the scene mode library as a user-defined mode;
[0100] Lighting control data is the underlying instruction set that directly drives the LED driver module, containing the brightness value (PWM duty cycle) and color channel value of each LED. The conversion process maps each pixel in the generated pixel matrix to a corresponding register configuration parameter or serial data stream. Storing it in the scene mode library means binding the generated lighting control data with a unique mode identifier (ID) and writing it into non-volatile memory (such as Flash), making it a persistent resource that the system can call. This step allows user-defined patterns to be invoked at any time by voice commands, just like preset holiday or party modes. For example, when a user uploads a photo of a birthday cake, the system processes it to generate corresponding dot matrix data and saves it as the "My Birthday Cake" mode. Subsequently, when the user says "Open My Birthday Cake" or a similar custom wake word, the system can retrieve the data from the scene mode library and drive the light panel display. This achieves automated conversion and solidification from general image data to a dedicated lighting control protocol.
[0101] This invention solves the problem in existing technologies where users find it difficult to easily convert personalized images into lighting patterns by receiving image data from external terminals, performing downsampling processing including color quantization and dithering algorithms, and generating and storing lighting control data. The data access from external terminals provides a rich source of content, color quantization in the downsampling process ensures accurate color reproduction within the hardware color gamut, and the dithering algorithm compensates for the visual loss caused by low resolution and low color depth. Together, these two methods enhance the detail and realism of the final displayed pattern. The generated lighting control data is integrated into a scene mode library, allowing custom modes to be parsed and executed by the offline speech recognition module on an equal footing with system preset modes. This achieves a high degree of personalization in lighting displays while ensuring privacy and low-latency response.
[0102] In another optional embodiment, the control command instructions further include basic control instructions, which include switch control, color control, brightness adjustment instructions, and mode switching instructions.
[0103] Basic control commands are direct operational instructions used to adjust the basic state of lighting equipment, originating from specific keywords or phrases input by the user via voice. Switch control commands trigger the power-on / off logic of the LED dot matrix light panel; color control commands contain predefined color value mappings, such as "red," "blue," etc.; brightness adjustment commands correspond to requests to increase or decrease light intensity; and mode switching commands are used to poll for different sub-effects in the current scene. These commands, together with scene mode commands, constitute a complete voice control set. For example, when a user says "Turn the lights to a warm yellow," the system recognizes it as a color control command and extracts the corresponding color temperature parameter; when the user says "Brighter," the system recognizes it as a brightness adjustment command.
[0104] When a color control command is detected, the overall color tone of the LED dot matrix display is adjusted.
[0105] Adjusting the overall color tone involves the main control module identifying color keywords, searching a preset color mapping table, obtaining the corresponding red (R), green (G), and blue (B) primary color component values, and converting these component values into drive signals to be sent to the LED driver module. This process specifically includes: parsing the color semantics in the voice command; if the command is "purple," then retrieving the RGB coordinate values of purple from the repository (e.g., R=128, G=0, B=128); the main control module converts these values into PWM signal duty cycle configurations, controlling the three sub-light-emitting units of each pixel on the LED dot matrix board. For example, upon receiving the command "switch to ocean blue," the system increases the duty cycle of the green and blue channels of the full-screen LED beads and decreases the duty cycle of the red channel.
[0106] When a brightness adjustment command is detected, the PWM duty cycle of the LED dot matrix light board is adjusted to change the brightness;
[0107] Adjusting the PWM duty cycle controls the average current of the LED chips by changing the duty cycle ratio of the pulse width modulation signal, thus linearly changing the luminous brightness. The system maintains a current brightness level variable. When a "brighten" command is received, this variable is increased by a preset step size (e.g., 10%), and the PWM output waveforms of each color channel are recalculated. When a "dim" command is received, this variable is decreased. If the current brightness has reached its maximum or minimum value, the current state is maintained or an audible alert is issued. For example, if the current PWM duty cycle is 50%, and the user issues a "dim" command, the system will reduce the duty cycle to 40%. At this time, the overall luminous flux of the LED matrix light panel decreases, but the chromaticity coordinates remain unchanged. Utilizing PWM dimming technology avoids the color shift problems that may occur with analog dimming.
[0108] When a mode switching command is detected, different sub-patterns or effects are cyclically switched under the currently selected mode type.
[0109] The cyclic switching of different sub-patterns or effects involves traversing multiple pre-stored sub-item sequences within the current major scene mode (such as holiday mode or party mode) without altering the overall scene mode. A linked list of indexes is maintained in memory for each mode type, containing multiple specific animation frame sequences or static graphic data. When a mode switching command (such as "next" or "change") is detected, the main control module increments the current index pointer by one. If the pointer exceeds the end of the linked list, it is reset to zero, and new sub-pattern data is loaded into the display memory, driving the LED dot matrix light panel to refresh the display. For example, in the currently selected birthday mode, the light panel displays a static image of a birthday cake upon receiving the first switching command; upon receiving the command again, it automatically switches to a flashing candle effect; and on the third command, it switches to a floating colored balloon effect. This fine-grained switching mechanism within the same context allows users to quickly preview and select their preferred lighting effect through simple voice interaction.
[0110] This invention, through the refinement and execution of the aforementioned basic control commands, enables control over the color, brightness, and internal effects of lighting in an offline state. The combination of color control commands and the RGB mapping mechanism allows users to define the ambient atmosphere using natural language; brightness adjustment commands, combined with PWM duty cycle adjustment technology, provide a smooth brightness change experience while ensuring color consistency; and the collaborative work of mode switching commands and the sub-pattern index list solves the cumbersome problem of needing to exit and re-enter mode switching in traditional solutions. These technologies work together to construct a high-response, low-power offline voice-controlled lighting system.
[0111] In another alternative embodiment, such as Figure 4 The diagram shows a flowchart illustrating how lighting effects change synchronously with the rhythm of music, according to an embodiment of the present invention. The method further includes:
[0112] Step S51: Acquire ambient audio signals in real time through the audio input interface;
[0113] The audio input interface is the physical channel through which the system connects to external sound sensors or microphones, used to convert analog sound wave signals into digital audio data streams. This step is executed by an audio acquisition process running on a low-power microcontroller unit, which continuously acquires raw audio samples from the environment after party mode is activated. The system digitizes the analog signal input from the microphone by configuring an analog-to-digital converter (ADC) at a preset sampling rate (e.g., 16kHz or 44.1kHz), generating discrete time-series data. For example, when a user plays background music in a party setting, the audio input interface acquires sound waveforms at a frequency of 16,000 times per second, forming continuous PCM data frames.
[0114] Step S52: Extract rhythmic features or frequency components of the ambient audio signal;
[0115] Rhythmic features are parameters in audio signals that characterize beat intensity, periodic pulses, or changes in energy envelope, while frequency components represent the energy distribution of the audio signal across different frequency bands. This step is a signal processing procedure based on the previously acquired digital audio data, separating the key control factors driving light changes from complex mixed sounds. The system uses a Fast Fourier Transform (FFT) algorithm to convert the time-domain audio signal to the frequency domain, analyzing the energy of low-frequency drum beats (e.g., 20Hz-250Hz) as rhythmic features, or analyzing the spectral centroid of the entire frequency band as frequency components. For example, when the ambient audio contains strong bass drum beats, the algorithm detects a sudden increase in energy peaks near 60Hz and marks it as a rhythmic pulse; when the music is in a high-pitched melody, the proportion of energy in the high-frequency band increases, which is identified as a specific frequency component feature. Through this feature extraction method, abstract sound signals can be quantified into calculable numerical indicators.
[0116] Step S53: Dynamically adjust the flashing frequency, brightness variation amplitude, or color change speed of the LED dot matrix light panel according to the rhythm characteristics or frequency components to achieve synchronized changes in lighting effects with the music rhythm.
[0117] Dynamic adjustment establishes a mapping relationship between audio feature parameters and light driving parameters, and updates the light status in real time based on changes in audio features. This step is executed by using extracted rhythm features or frequency components as input variables and calculating the corresponding PWM duty cycle, color index increment, or refresh interval through a predefined mapping function. When a strong rhythmic feature (such as a heavy bass drum beat) is detected, the system triggers high-brightness flashing or rapid color switching on the LED dot matrix light panel; when abundant high-frequency components are detected, the speed of color changes is accelerated or the amplitude of brightness changes is increased. For example, if the extracted rhythmic feature shows a beat interval of 0.5 seconds, the LED light panel is controlled to alternate between bright and dark at a frequency of 2Hz; if frequency component analysis shows extremely strong low-frequency energy, the brightness change amplitude is set to the maximum value, creating a strong burst of light. The LED dot matrix light panel no longer plays static patterns according to a fixed script, but responds in real time with the fluctuations of the live music, providing dynamic interactivity for party modes.
[0118] This invention achieves synchronization between lighting effects and ambient music through the synergy of the aforementioned technical features. An audio input interface acquires ambient audio signals in real time, providing the system with a sensory interface for the external sound field. Signal processing algorithms extract rhythmic features or frequency components, transforming unstructured sound information into structured control commands. Based on this, these audio features are directly mapped to the flashing frequency, brightness variation, or color change speed of the LED dot matrix light panel, constructing a closed-loop feedback mechanism from hearing to seeing. This synergistic cooperation enables the lighting system to capture rhythmic changes and melodic fluctuations in music and instantly transform them into visual light and shadow dynamics. In party mode, users can enjoy a light show that changes with the music without manual intervention, solving the problems of single lighting modes and inability to respond to real-time ambient sound in existing technologies.
[0119] In one embodiment, the method runs on a low-power microcontroller unit (MCU);
[0120] The low-power microcontroller unit (MCU) is the core computing and control chip that carries the offline voice-controlled light free combination pattern method of this invention. The MCU includes a central processing unit core, on-chip memory, a timer module, and various peripheral interfaces.
[0121] In the standby state of step S1, only the voice activity detection module works, and the overall system power consumption is less than 16mW;
[0122] The standby state is a silent monitoring phase when the system has completed initialization but has not received a valid wake-up signal. In this state, the overall system power consumption is below 16mW, achieved through a hardware and software collaborative energy-saving strategy. The MCU puts all high-power peripherals, except for the Voice Activity Detection (VAD) module, such as the offline speech recognition engine, LED dot matrix light board driver circuit, and audio power amplifier, into sleep or power-off states. The VAD module operates in ultra-low power mode, periodically sampling and initially judging the energy value or specific spectral characteristics of ambient sound signals without performing complex waveform decoding or semantic analysis. For example, the VAD module can be configured to wake up the MCU core every 10ms to perform microsecond-level energy integration calculations. If no sound signal exceeding a preset threshold is detected, it immediately re-enters sleep mode. Through this intermittent operation and partial wake-up mechanism, the system standby current can be controlled within a few milliamps (assuming a supply voltage of 3.3V-5V), and the overall power consumption is stably maintained below 16mW. This low-power characteristic makes the method of this invention suitable for battery-powered portable lighting devices or decorative lighting scenarios.
[0123] After activating the offline speech recognition module in step S2, the system enters normal working state with a response delay of less than 200ms.
[0124] Normal operation occurs when the voice activity detection module confirms the presence of a human voice signal and generates a wake-up signal, triggering the offline voice recognition module to run at full speed. A response latency of less than 200ms is the time interval between the end of the user's voice command and the start of the corresponding lighting action on the LED dot matrix light panel. This low latency characteristic is primarily due to the offline processing architecture and the efficient scheduling of the MCU: since there is no need to upload audio data to a cloud server for recognition, network transmission delays and cloud queuing time are eliminated; simultaneously, upon receiving the wake-up signal, the MCU immediately loads the pre-set Chinese and English dictionary model into the high-speed cache in full-frequency operation mode, directly extracting and matching features from the locally collected audio data. For example, when the user says "Open birthday mode," the MCU completes the digitization and feature vectorization of the voice signal within approximately 50ms, completes the matching with the local command dictionary within approximately 80ms, and uses the remaining time for command parsing and PWM signal generation, with the total time controlled within 200ms. This millisecond-level response speed provides users with an instant interactive experience, avoiding the operational lag caused by network fluctuations in traditional cloud solutions.
[0125] This invention achieves a balance between high performance and low power consumption. By combining a phased module start-stop strategy, the system retains only the most basic auditory perception capabilities in standby mode, reducing power consumption to below 16mW. Upon detecting a valid human voice, it quickly switches to full-function operation, utilizing local computing power to complete the entire closed loop from voice acquisition to lighting execution within 200ms. This operating mechanism overcomes the shortcomings of existing offline solutions, such as limited functionality or slow response, and also avoids the network dependence and privacy risks of cloud-based solutions.
[0126] This invention also provides a system for offline voice-controlled free combination of light patterns. The system includes: a voice acquisition module for acquiring ambient sound signals; a voice activity detection module connected to the voice acquisition module for continuously monitoring ambient sound signals in standby mode and generating a wake-up signal when a human voice signal meeting preset conditions is detected; an offline voice recognition module connected to the voice activity detection module for being activated upon receiving the wake-up signal and performing offline recognition of the voice data to obtain voice commands; a main control module connected to both the offline voice recognition module and the LED driver module for parsing voice commands, retrieving the corresponding light control strategy from the memory, and generating control signals; an LED driver module for driving the LED dot matrix light panel to display corresponding light patterns according to the control signals; and a memory for storing a scene mode library, including at least one holiday mode, at least one party mode, and a user-defined mode; wherein the user-defined mode supports the generation of dot matrix graphic data uploaded by an external terminal received via a wireless communication module.
[0127] The voice acquisition module is a hardware component used to convert acoustic signals into electrical signals, which can be an electret microphone or other types of acoustic-electric conversion devices. This module is connected to the voice activity detection module and serves as the front-end input interface of the system. It is responsible for capturing the sound information in the surrounding environment in real time and transmitting analog or digital audio streams to the subsequent processing unit. In practical applications, parameters such as the sensitivity and signal-to-noise ratio of this module can be set according to actual requirements.
[0128] The voice activity detection module is a logic circuit or processing unit used to determine whether the input sound signal contains valid human voices. It is connected to the voice acquisition module and is in the standby monitoring state of the system. Its core function is to perform activity determination, that is, to distinguish background noise from valid human voice activities. In the system linkage relationship, this module continuously receives signals from the voice acquisition module. Once a human voice signal that meets the preset conditions is detected, it generates a wake-up signal and sends it to the offline speech recognition module, triggering the system to enter the normal working state from the low-power standby state. This cooperation relationship optimizes the system power consumption and avoids the high energy consumption caused by the continuous operation of the offline speech recognition module when there is no valid speech. The specific implementation method of this module can be a hardware-based energy detection circuit, a software-based short-time energy zero-crossing rate analysis algorithm, or a two-stage wake-up mechanism.
[0129] The offline speech recognition module is built-in with an acoustic model and a language model, and can decode and match speech data without relying on a network connection. It is connected to the voice activity detection module and is only activated after receiving the wake-up signal. Its function is to convert continuous speech waveform data into specific text instructions or command codes. In the overall technical solution, this module starts after receiving the wake-up signal, deeply recognizes the speech data collected by the voice acquisition module and preliminarily screened, and outputs speech instructions including wake-up word instructions or control command instructions. Its cooperation with the voice activity detection module forms a "pre-detection + precise recognition" link, which not only ensures the timeliness of the response but also reduces the average power consumption. This module internally pre-sets a Chinese wake-up word library, an English wake-up word library, a Chinese command word library, and an English command word library, supporting bilingual offline recognition in Chinese and English. The specific recognition algorithm can adopt the hidden Markov model, deep neural network, etc.
[0130] The main control module is the central processing unit of the system, such as a microcontroller unit (MCU), digital signal processor (DSP), or embedded processor. It connects to the offline speech recognition module and the LED driver module, coordinating and scheduling the resources of the entire system. In system interaction, the main control module receives voice commands from the offline speech recognition module, performs semantic parsing to determine the user's intent; then it accesses the scene pattern library in memory to match the corresponding lighting control strategy; finally, it generates specific control signals based on the matched strategy and sends them to the LED driver module. This module is also responsible for managing the system's working state switching, such as controlling the relevant modules to return to standby mode after completing an instruction execution. The specific model, clock speed, and core architecture of the main control module can be selected based on processing complexity and power consumption requirements.
[0131] The LED driver module can be a circuit interface or driver chip used to drive the LED dot matrix light panel to emit light. This LED driver module is connected to the main control module and directly to the LED dot matrix light panel. Its function is to convert the logic control signals generated by the main control module into current or voltage signals that can drive the LED beads to light up, turn off, or adjust their brightness. In the overall scheme, the LED driver module controls the display state of each LED bead on the LED dot matrix light panel according to the control signals issued by the main control module, thereby presenting corresponding light patterns or dynamic effects. Its cooperation with the main control module realizes the final conversion from digital instructions to physical light effects. This LED driver module can integrate constant current driving function and PWM dimming function, and supports multiple color combinations of RGB or RGBW. Its specific driving method can be static driving or dynamic scanning driving; this embodiment of the invention does not impose any special limitations on this.
[0132] The memory is a non-volatile storage medium, such as Flash memory or EEPROM (Electrically Erasable Programmable Read-Only Memory), used to store program code, configuration data, and scene mode data. It is connected to the main control module (or via a bus) and stores the scene mode library. Its function is to provide persistent data storage for the system, ensuring that scene configurations are not lost after power failure. During system synchronization, the memory responds to read requests from the main control module, providing lighting control strategy data (such as dot matrix data, color sequences, and flashing frequencies) corresponding to holiday modes, party modes, and user-defined modes. The data structure stored in the scene mode library can be a lookup table, a binary file, or other formats.
[0133] The scene mode library can store at least one holiday mode, at least one party mode, and user-defined modes. Holiday modes can be preset lighting display schemes that match the atmosphere of a specific holiday, such as birthday mode, Valentine's Day mode, Children's Day mode, Halloween mode, etc., each mode corresponding to specific light color combinations, flashing frequencies, and dot matrix patterns. Party modes can be schemes that dynamically adjust lighting effects according to music rhythm or ambient sound changes. User-defined modes are personalized light patterns that users can define and upload through external means. These modes together constitute the core content resources of the system, enabling the system to meet diverse lighting needs.
[0134] The user-defined mode supports the generation of dot matrix graphic data uploaded from external terminals, meaning the system has scalability and personalization capabilities. Specifically, users can design or take pictures using external terminals such as mobile apps and tablets, transmit them to the system via a wireless communication module, and the system processes and stores them as user-defined patterns. This mechanism allows light patterns to no longer be limited to factory presets but can be updated at any time according to user preferences.
[0135] The system operation process of this invention is as follows: After the system is powered on, it enters standby mode. At this time, only the voice acquisition module and the voice activity detection module work, continuously monitoring the ambient sound. When the voice activity detection module detects a human voice signal that meets the preset conditions, it generates a wake-up signal to activate the offline voice recognition module. The offline voice recognition module starts and recognizes the acquired voice data to obtain specific voice commands and sends them to the main control module. The main control module parses the command and matches the corresponding lighting control strategy from the scene mode library in the memory. The main control module generates a corresponding control signal and sends it to the LED driver module. The LED driver module drives the LED dot matrix light panel to display the light pattern or dynamic effect matching the command according to the control signal. If a user-defined mode is involved, the external terminal can upload the dot matrix graphic data to the memory through the wireless communication module to update the scene mode library for subsequent calls.
[0136] Through the above technical solutions, this invention constructs a complete system architecture including voice acquisition, activity detection, offline recognition, main control scheduling, and LED driving, which can support the fully automated control of the entire process from sound input to light pattern output from the hardware level; the cascaded cooperation of the voice activity detection module and the offline voice recognition module achieves a balance between low power standby and fast response; the memory stores a scene mode library including festival, party, and user-defined modes, and supports receiving external data through the wireless communication module, providing rich and diverse and customizable lighting display effects.
[0137] In another aspect, the present invention also provides a computer-readable storage medium, which may be the computer-readable storage medium included in the apparatus described above; or it may be a standalone computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, which are used by one or more processors to execute the offline voice-controlled light free combination pattern method described in the present invention.
[0138] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for offline voice-controlled free combination of light patterns, characterized in that, Includes the following steps: Step S1: In standby mode, the ambient sound signal is continuously monitored through the voice activity detection module; Step S2: When the voice activity detection module detects a human voice signal that meets the preset conditions, the offline voice recognition module is activated; Step S3: The offline speech recognition module recognizes the collected speech data to obtain speech commands; the speech commands include wake word commands or control command commands. Step S4: If the control command is identified, the control command is parsed and the corresponding lighting control strategy is matched from the preset scene mode library; Step S5: Drive the LED dot matrix light panel to display the corresponding light pattern according to the lighting control strategy; The scene mode library stores at least one festival mode, at least one party mode, and a user-defined mode; the user-defined mode supports the generation of dot matrix graphic data uploaded through an external terminal.
2. The method according to claim 1, characterized in that, In step S1, the voice activity detection module employs a two-level wake-up mechanism, specifically including: Level 1 detection: The energy value of ambient sound is monitored through an ultra-low power energy detection circuit. When the energy value exceeds the first threshold, the second level detection is triggered. The second level of detection: By using a human voice duration verification algorithm or spectral feature analysis, it is determined whether the sound signal is a human voice. If so, a wake-up signal is generated to activate the offline speech recognition module.
3. The method according to claim 1, characterized in that, In step S3, the offline speech recognition module supports offline recognition of both Chinese and English. The offline speech recognition module has a pre-built Chinese wake-up word library, an English wake-up word library, a Chinese command word library, and an English command word library; The offline speech recognition module is configured to match speech commands from the corresponding dictionary based on the user-defined language pattern or automatic recognition mode.
4. The method according to claim 1, characterized in that, In step S4, the holiday mode includes at least one of the following: birthday mode, Valentine's Day mode, Children's Day mode, and Halloween mode. Each festival mode corresponds to a specific combination of light colors, flashing frequency, and dot matrix pattern; Step S5 specifically includes: according to the matched holiday mode, controlling each LED bead on the LED dot matrix light board to light up or turn off according to a preset time sequence, so as to dynamically display the corresponding dot matrix pattern.
5. The method according to claim 1, characterized in that, The method also includes a step for constructing user-defined patterns, specifically including: Step S10: Receive image data from an external terminal; Step S20: Downsample the image data to convert it into a pixel matrix that matches the resolution of the LED dot matrix light panel; Step S30: Convert the pixel matrix into lighting control data and store it in the scene mode library as a user-defined mode; The downsampling process includes color quantization and dithering algorithms to adapt to the color display capabilities of the LED light panel.
6. The method according to claim 1, characterized in that, The control command instructions also include basic control instructions, which include switch control, color control, brightness adjustment and mode switching. When a color control command is detected, adjust the overall color tone of the LED dot matrix display. When a brightness adjustment command is detected, the PWM duty cycle of the LED dot matrix light board is adjusted to change the brightness. When a mode switching command is detected, different sub-patterns or effects are cyclically switched under the currently selected mode type.
7. The method according to claim 1, characterized in that, In step S5, if the currently matched mode is the party mode, the method further includes: Step S51: Acquire ambient audio signals in real time through the audio input interface; Step S52: Extract the rhythmic features or frequency components of the ambient audio signal; Step S53: Dynamically adjust the flashing frequency, brightness change amplitude, or color change speed of the LED dot matrix light panel according to the rhythm characteristics or frequency components to realize that the lighting effect changes synchronously with the music rhythm.
8. The method according to any one of claims 1 to 7, characterized in that, The method is run on a low-power microcontroller unit (MCU). In the standby state of step S1, only the voice activity detection module is configured to work; After activating the offline speech recognition module in step S2, the configuration system enters normal working state.
9. A system for offline voice-controlled free combination of light patterns, characterized in that, include: The voice acquisition module is used to collect ambient sound signals; A voice activity detection module, connected to the voice acquisition module, is used to continuously monitor ambient sound signals in standby mode and generate a wake-up signal when a human voice signal that meets preset conditions is detected. An offline speech recognition module, connected to the speech activity detection module, is activated upon receiving the wake-up signal and performs offline speech data recognition to obtain speech commands. The main control module is connected to the offline voice recognition module and the LED driver module respectively, and is used to parse the voice commands, retrieve the corresponding lighting control strategy from the memory, and generate control signals; The LED driver module is used to drive the LED dot matrix light board to display corresponding light patterns according to the control signal. A memory for storing a scene mode library, the scene mode library including at least one holiday mode, at least one party mode and user-defined modes; The user-defined mode supports the generation of dot matrix graphic data uploaded by an external terminal and received via a wireless communication module.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for offline voice-controlled free combination of light patterns as described in any one of claims 1 to 8.