A voice recognition-based light intelligent adjustment method and device

CN122555026APending Publication Date: 2026-08-11SHENZHEN HAIBEN ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]有鉴于此,本申请实施例提供了一种基于语音识别的灯光智能调节方法及装置,旨在解决现有技术中存在的噪声环境识别效果差、无法兼容两线制灯光架构的问题

Benefits of technology

[0010] The beneficial effects of this application embodiment compared with the prior art are as follows: This application adopts an offline voice processing mode, which effectively reduces the standby power consumption of the device, greatly improves the voice recognition accuracy in complex noise environments, is compatible with two-wire lighting systems, does not require uploading voice data to the cloud, avoids the risk of privacy leakage, and will not affect the normal operation of the device due to network interruption, thereby comprehensively improving the stability, functionality, compatibility and user experience of lighting voice control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122555026A_ABST
    Figure CN122555026A_ABST
Patent Text Reader

Abstract

This application provides a method and device for intelligent lighting adjustment based on speech recognition, applicable to the field of data processing technology. The method includes: generating speech recognition information based on environmental sound monitoring information, a human voice recognition model, and an accent differentiation and recognition model; generating speech command word recognition information based on the speech recognition information and a speech command word recognition model; and generating lighting adjustment control command information based on the speech command word recognition information and the mapping relationship between the command words and lighting adjustment control commands, for lighting adjustment control processing. This application adopts an offline speech processing mode, effectively reducing device standby power consumption, significantly improving speech recognition accuracy in complex noise environments, and is compatible with two-wire lighting systems, comprehensively improving the stability, functionality, compatibility, and user experience of lighting voice control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to a method and device for intelligent lighting adjustment based on voice recognition. Background Technology

[0002] With the rapid development of the smart home industry, voice control has become the mainstream application in the field of smart lighting. Offline voice lighting control technology has also been gradually promoted and applied, and is widely used in various lighting products such as LED light strips, decorative lights, and ambient lights.

[0003] There are currently offline voice-controlled lighting solutions that can realize basic on / off and color switching functions for lights. Some solutions are equipped with a voice activity detection module to complete voice wake-up, and are mainly adapted to control three-wire and four-wire lighting systems.

[0004] However, existing offline voice lighting control solutions have limited lighting effects, low voice recognition accuracy in complex noisy environments, and cannot be adapted to two-wire lighting systems with only a power supply line. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method and apparatus for intelligent lighting adjustment based on voice recognition, which aims to solve the problems of poor noise environment recognition and incompatibility with two-wire lighting architecture in the prior art.

[0006] The first aspect of this application provides a method for intelligent lighting adjustment based on voice recognition, comprising: Obtain environmental sound monitoring information; Based on the environmental sound monitoring information, the preset human voice recognition model, and the preset accent differentiation and recognition model, speech recognition information is generated; Based on the speech recognition information and multiple preset speech command word recognition models, multiple speech command word recognition information is generated; Based on the recognition information of multiple voice command words and the mapping relationship between multiple preset command words and lighting adjustment control instructions, lighting adjustment control instruction information is generated so as to perform lighting adjustment control processing through the lighting adjustment control instruction information.

[0007] A second aspect of this application provides a voice recognition-based intelligent lighting adjustment device, comprising: An environmental sound monitoring information acquisition module is used to acquire environmental sound monitoring information; The speech recognition information generation module is used to generate speech recognition information based on the environmental sound monitoring information, the preset human voice recognition model, and the preset accent differentiation and recognition model. The voice command word recognition information generation module is used to generate multiple voice command word recognition information based on the voice recognition information and multiple preset voice command word recognition models; The lighting adjustment control instruction information generation module is used to generate lighting adjustment control instruction information based on the recognition information of multiple voice command words and the mapping relationship between multiple preset command words and lighting adjustment control instructions, so as to perform lighting adjustment control processing through the lighting adjustment control instruction information.

[0008] A third aspect of this application provides a terminal device, the terminal device including a memory and a processor, the memory storing a computer program executable on the processor, and the processor executing the computer program to implement the steps of the voice recognition-based intelligent lighting adjustment method described in the first aspect above.

[0009] A fourth aspect of this application provides a computer-readable storage medium, comprising: storing a computer program, wherein when executed by a processor, the computer program implements the steps of the voice recognition-based intelligent lighting adjustment method described in the first aspect above.

[0010] The beneficial effects of this application embodiment compared with the prior art are as follows: This application adopts an offline voice processing mode, which effectively reduces the standby power consumption of the device, greatly improves the voice recognition accuracy in complex noise environments, is compatible with two-wire lighting systems, does not require uploading voice data to the cloud, avoids the risk of privacy leakage, and will not affect the normal operation of the device due to network interruption, thereby comprehensively improving the stability, functionality, compatibility and user experience of lighting voice control. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram illustrating the implementation process of the intelligent lighting adjustment method based on voice recognition provided in Embodiment 1 of this application; Figure 2 This is a schematic diagram illustrating the implementation process of the intelligent lighting adjustment method based on voice recognition provided in Embodiment 2 of this application; Figure 3 This is a schematic diagram illustrating the implementation process of the intelligent lighting adjustment method based on voice recognition provided in Embodiment 3 of this application; Figure 4This is a schematic diagram illustrating the implementation process of the intelligent lighting adjustment method based on voice recognition provided in Embodiment 4 of this application; Figure 5 This is a schematic diagram illustrating the implementation process of the intelligent lighting adjustment method based on voice recognition provided in Embodiment 5 of this application; Figure 6 This is a schematic diagram illustrating the implementation process of the intelligent lighting adjustment method based on voice recognition provided in Embodiment Six of this application; Figure 7 This is a schematic diagram of the structure of the intelligent lighting adjustment device based on voice recognition provided in an embodiment of this application; Figure 8 This is a schematic diagram of the terminal device provided in the embodiments of this application. Detailed Implementation

[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0014] To illustrate the technical solution described in this application, specific embodiments are provided below.

[0015] Figure 1 The flowchart illustrating the implementation of the intelligent lighting adjustment method based on voice recognition provided in Embodiment 1 of this application is shown below in detail: Step S101: Obtain environmental sound monitoring information.

[0016] In this embodiment, the environmental sound monitoring information can be the full-frequency environmental audio signal collected by the device in real time, including various environmental noises and human voice signals. It can be continuously collected and acquired through the voice acquisition module, and can completely record the sound waveform and sound intensity data at different times in the space, providing the original audio data source for subsequent human voice detection, noise reduction processing and speech recognition.

[0017] Step S102: Generate speech recognition information based on the environmental sound monitoring information, the preset human voice recognition model, and the preset accent differentiation and recognition model.

[0018] In this embodiment, the preset voice recognition model can be manually preset and can be constructed using an architecture combining VAD pre-detection and a lightweight RNNoise neural network. The preset voice energy threshold value can be the corresponding trigger detection standard value, and the preset voice double verification duration can be a first preset duration and a second preset duration, respectively. The preset accent discrimination recognition model can also be manually preset and can be constructed using a multi-task multi-accent deep neural network model. It can rely on Mel-frequency cepstral coefficients to complete feature extraction. The preset difference between British and American accent probabilities can be a fixed standard value. Then, the preset voice recognition model can be used to perform voice activity detection and AI noise reduction processing on environmental sound monitoring information, filtering out interference sound sources such as TV sound, air conditioner sound, and fan noise. The processed pure voice is then input into the preset accent discrimination recognition model to complete the language and accent determination of Mandarin Chinese, British English, and American English. Then, the voice validity result, language type result, and accent type result are integrated to generate speech recognition information containing complete audio features, language information, and accent information.

[0019] Step S103: Generate multiple voice command word recognition information based on the voice recognition information and multiple preset voice command word recognition models.

[0020] In this embodiment, multiple preset voice command word recognition models can be manually preset. Different voice command word recognition models are trained specifically for Chinese command words, British English command words, and American English command words. The main body of the model can be built using a deep neural network. The preset command word matching confidence threshold can be set to the standard recognition threshold. The preset command word classification labels include basic control command labels, dynamic special effect command labels, timing control command labels, and combined special effect command labels. Then, the audio features, language and accent features in the voice recognition information are input into the matching preset voice command word recognition models, and then keyword retrieval, command classification and semantic parsing are completed in sequence. Finally, the recognition results of different dimensions are summarized to generate multiple voice command word recognition information corresponding to various types of lighting control command content.

[0021] Step S104: Based on the recognition information of the multiple voice command words and the mapping relationship between the multiple preset command words and the lighting adjustment control instructions, generate lighting adjustment control instruction information, so as to perform lighting adjustment control processing through the lighting adjustment control instruction information.

[0022] In this embodiment, the mapping relationship between multiple preset command words and lighting adjustment control instructions can be pre-set manually. This mapping relationship can include the correspondence rules between all voice command words and lighting action instructions, including turning lights on and off, color switching, brightness adjustment, and all control logic such as flowing water effects, meteor effects, flashing effects, breathing effects, various holiday combination effects, and timed switching. The pre-set dynamic effect operation parameters include the lighting sequence, light spot movement speed, brightness change curve, and light alternation frequency. The pre-set two-wire signal encoding parameters include the duty cycle values ​​of the start pulse, data pulse, and end pulse. Then, the recognition information of multiple voice command words is prioritized and its validity is verified. Then, the pre-set mapping relationship between command words and lighting adjustment control instructions is combined to complete instruction matching and parameter assignment. Then, the control instruction is encoded and encapsulated according to the two-wire transmission rules, thereby generating lighting adjustment control instruction information that includes light on / off, color, brightness, dynamic effects, combination scenes, timing functions, and two-wire transmission encoded content. Then, the lighting adjustment control instruction information can drive the LED driver module and the two-wire communication module to complete the corresponding lighting adjustment control processing.

[0023] The intelligent lighting adjustment method based on voice recognition provided in this application adopts an offline voice processing mode, which effectively reduces the standby power consumption of the device, significantly improves the voice recognition accuracy in complex noise environments, is compatible with two-wire lighting systems, does not require uploading voice data to the cloud, avoids the risk of privacy leakage, and will not affect the normal operation of the device due to network interruption, thereby comprehensively improving the stability, functionality, compatibility and user experience of lighting voice control.

[0024] Figure 2 The flowchart illustrating the implementation of the voice recognition-based intelligent lighting adjustment method provided in Embodiment 2 of this application is shown. The difference between this method and Embodiment 1 is that step S102 specifically includes: Step S201: Generate human voice recognition information based on the environmental sound monitoring information and the preset human voice recognition model.

[0025] In this embodiment, the preset human voice recognition model can be manually preset and can be constructed using a VAD pre-detection architecture. The manually preset human voice energy threshold value can be the corresponding trigger detection standard value. The manually preset human voice dual verification duration can be a first preset duration and a second preset duration, respectively. Then, the preset human voice recognition model can be used to detect voice activity on environmental sound monitoring information, distinguish between environmental noise and valid human voice signals, and then determine whether there is human voice in the current audio. Then, the human voice detection results and original audio feature data are sorted together to generate human voice recognition information containing human voice determination results and original audio features.

[0026] Step S202: Based on the preset human voice noise reduction processing model, noise reduction processing is performed according to the human voice recognition information to generate noise-reduced human voice recognition information.

[0027] In this embodiment, the preset human voice noise reduction processing model can be pre-set by the user and can be constructed using a lightweight RNNoise neural network. It can suppress common environmental noise sources such as television sound, air conditioner sound, and fan noise. The pre-set noise reduction gain parameter can be a fixed standard value, and the pre-set frequency domain feature extraction parameter can be a uniformly configured value. Then, the audio frequency domain features in the human voice recognition information are extracted. Subsequently, the speech gain curve is calculated through the preset human voice noise reduction processing model, and various environmental noises are filtered out. Then, the pure human voice audio data is retained, thereby generating noise-reduced human voice recognition information with interference signals removed.

[0028] Step S203: Generate speech recognition information based on the noise-reduced human voice recognition information and the preset accent differentiation and recognition model.

[0029] In this embodiment, the preset accent discrimination and recognition model can be pre-set by humans and can be constructed using a multi-task, multi-accent deep neural network model. Feature extraction can be completed using Mel-frequency cepstral coefficients. The pre-set probability judgment difference between British and American accents can be a fixed standard value. Then, human voice feature data is extracted from the noise-reduced human voice recognition information. The feature data is then input into the preset accent discrimination and recognition model to complete the language and accent determination of Mandarin Chinese, British English, and American English. Finally, all data such as human voice status, language type, and accent type are summarized to generate speech recognition information containing complete audio features, language information, and accent information.

[0030] The intelligent lighting adjustment method based on voice recognition provided in this application effectively reduces device standby power consumption and enhances noise filtering through offline voice processing mode. It continuously improves the accuracy of voice recognition in complex noise environments, is compatible with two-wire lighting systems, and eliminates the need to upload voice data to the cloud, effectively avoiding the risk of privacy leakage. The device operation is not affected by network status, thereby optimizing the operational stability and recognition accuracy of voice control for lighting.

[0031] Figure 3 The flowchart illustrating the implementation of the intelligent lighting adjustment method based on voice recognition provided in Embodiment 3 of this application is shown. Its difference from Embodiment 2 described above lies in: The preset voice recognition model includes a preset voice recognition trigger command generation sub-model and a preset voice recognition generation sub-model; Step S201 specifically includes: Step S301: Generate a sub-model based on the environmental sound monitoring information and the preset human voice recognition trigger command, and generate human voice recognition trigger command information.

[0032] In this embodiment, the preset human voice recognition trigger instruction generation sub-model can be preset by humans and can be constructed using a VAD network model. The preset human voice energy threshold value can be the corresponding trigger detection standard value, and the preset human voice duration threshold value can be 200 milliseconds. Then, the ambient sound monitoring information is used to perform real-time sound energy detection, and then it is determined whether the audio signal reaches the human voice trigger standard. Then, the corresponding trigger judgment result is output, thereby generating human voice recognition trigger instruction information for starting the subsequent recognition process.

[0033] Step S302: Generate human voice recognition information based on the environmental sound monitoring information, human voice recognition trigger instruction information, and preset human voice recognition generation sub-model.

[0034] In this embodiment, the preset voice recognition generation sub-model can be manually preset and can use a VAD network model to complete the secondary verification. The manually preset secondary detection duration threshold can be a fixed standard value, which is used to cooperate with the preceding module to complete the dual verification. Then, based on the voice recognition trigger instruction information, the formal voice recognition process is confirmed to start. Subsequently, the preset voice recognition generation sub-model is called to perform secondary voice activity detection on the environmental sound monitoring information to eliminate false triggering caused by instantaneous noise. Then, the dual detection results are integrated with the original audio features to generate voice recognition information containing voice judgment results and original audio features.

[0035] The intelligent lighting adjustment method based on voice recognition provided in this application significantly reduces the probability of false triggering caused by environmental noise through a two-level detection mechanism, ensuring the accuracy of voice recognition. It is used to stably adapt to two-wire lighting systems, does not rely on cloud networks or transmit voice data throughout the process, and takes into account the ability to prevent accidental touches, recognition accuracy, safety of use and device compatibility, thereby improving the overall effect of lighting voice control.

[0036] Figure 4 The flowchart illustrating the implementation of the voice recognition-based intelligent lighting adjustment method provided in Embodiment 4 of this application is shown. Its difference from Embodiment 1 above lies in: The preset voice command word recognition model includes a preset voice accent matching sub-model and a preset voice command word recognition sub-model; wherein, the preset voice accent matching sub-model and the preset voice command word recognition sub-model correspond one-to-one. Step S103 specifically includes: Step S401: Based on the speech recognition information and multiple preset speech accent matching sub-models, obtain the matched speech accent matching sub-model.

[0037] In this embodiment, multiple preset speech accent matching sub-models can be preset by humans. The main body of the model can be constructed using a multi-task multi-accent deep neural network model. The preset accent matching confidence threshold can be a fixed standard value. Then, language features and accent features in the speech recognition information are extracted. The feature data are then input into each preset speech accent matching sub-model to complete the matching operation. Finally, the matching results output by each model are compared to obtain the matched speech accent matching sub-model.

[0038] Step S402: Based on the matched speech accent matching sub-model, perform matching processing on multiple preset speech command word recognition sub-models to obtain the matched speech command word recognition sub-model.

[0039] In this embodiment, multiple preset voice command word recognition sub-models can be preset by humans. The main body of the model can be built using a deep neural network. Different models are adapted to the corresponding command words in Chinese, British English, and American English respectively. The preset model association rules are based on a unified configuration standard. Then, according to the language and accent type of the matched voice accent matching sub-model, the corresponding voice command word recognition sub-model is retrieved. Then, the model filtering and binding operation is completed to obtain the matched voice command word recognition sub-model.

[0040] Step S403: Generate multiple voice command word recognition information based on the speech recognition information and the matched speech command word recognition sub-model.

[0041] In this embodiment, the audio features, language features, and accent features in the speech recognition information are input into the matching speech command word recognition sub-model, and then keyword retrieval, instruction classification, and semantic parsing are carried out. Finally, all parsing results are integrated to generate multiple speech command word recognition information.

[0042] The intelligent lighting adjustment method based on speech recognition provided in this application improves the recognition accuracy for different languages ​​and accents, thereby effectively improving the pertinence and reliability of speech recognition and control.

[0043] Figure 5 The flowchart illustrating the implementation of the speech recognition-based intelligent lighting adjustment method provided in Embodiment 5 of this application is shown. Its difference from Embodiment 4 described above lies in: Multiple preset speech accent matching sub-models include a preset first speech accent matching sub-model and a preset second speech accent matching sub-model; Step S401 specifically includes: Step S501: Calculate the first speech accent matching probability information based on the speech recognition information and the preset first speech accent matching sub-model.

[0044] In this embodiment, the preset first speech accent matching sub-model can be preset by humans and can be constructed using a multi-task multi-accent deep neural network model. It is mainly used to recognize British English accents. The preset accent feature extraction parameters can be fixed configuration values. Then, human voice feature data in speech recognition information is extracted, and the feature data is input into the preset first speech accent matching sub-model for probability calculation. Then, the corresponding matching value is output to calculate the first speech accent matching probability information.

[0045] Step S502: Calculate the second speech accent matching probability information based on the speech recognition information and the preset second speech accent matching sub-model.

[0046] In this embodiment, the preset second speech accent matching sub-model can be pre-set manually and can be constructed using a multi-task multi-accent deep neural network model. It is mainly used to recognize American English accents. The pre-set accent feature extraction parameters are consistent with the first speech accent matching sub-model. Then, the human voice feature data in the speech recognition information is input into the preset second speech accent matching sub-model to complete the probability calculation. The matching results output by the model are then statistically analyzed, and the corresponding numerical content is organized to calculate the second speech accent matching probability information.

[0047] Step S503: Calculate the difference between the first and second speech accent matching probability information to obtain speech accent matching probability difference information.

[0048] In this embodiment, the first speech accent matching probability information and the second speech accent matching probability information are extracted and the difference is calculated. The result data of the calculation is recorded to obtain the speech accent matching probability difference information.

[0049] Step S504: Based on the speech accent matching probability difference information, the preset first speech accent matching sub-model, and the preset second speech accent matching sub-model, obtain the matched speech accent matching sub-model.

[0050] In this embodiment, the preset accent judgment difference standard can be a fixed value. Then, the voice accent matching probability difference information is compared with the preset judgment standard. The accent type corresponding to the current voice is determined by combining the difference size. Then, the matching model is selected to obtain the matched voice accent matching sub-model.

[0051] The intelligent lighting adjustment method based on speech recognition provided in this application effectively distinguishes multiple speech accents, improves the accuracy of speech recognition for multiple accents, and thus comprehensively optimizes the lighting voice control experience in multi-accent scenarios.

[0052] Figure 6 The flowchart illustrating the implementation of the speech recognition-based intelligent lighting adjustment method provided in Embodiment Six of this application is shown. Its difference from Embodiment One described above lies in: The voice command word recognition information includes basic lighting function adjustment command word recognition information and lighting effect function adjustment command word recognition information; The mapping relationship between multiple preset command words and lighting adjustment control commands includes the mapping relationship between preset basic function command words and lighting adjustment control commands, as well as the mapping relationship between preset special effects function command words and lighting adjustment control commands. The lighting adjustment control command information includes basic lighting function adjustment control command information and lighting special effects function adjustment control command information; Step S104 specifically includes: Step S601: Generate lighting basic function adjustment control instruction information based on the recognition information of the multiple lighting basic function adjustment command words and the preset mapping relationship between basic function command words and lighting adjustment control instructions.

[0053] In this embodiment, the mapping relationship between preset basic function command words and lighting adjustment control instructions can be preset by humans. It can include command words and control instructions corresponding to basic operations such as turning on the light, turning off the light, color switching, and brightness adjustment. The preset execution priority of basic instructions can be a fixed sorting standard. Then, the validity of multiple lighting basic function adjustment command word recognition information is verified. Then, the instruction matching is completed by combining the preset mapping relationship between basic function command words and lighting adjustment control instructions. Finally, the instructions are sorted and organized according to the execution priority to generate lighting basic function adjustment control instruction information.

[0054] Step S602: Based on the recognition information of the multiple lighting effect function adjustment command words and the preset mapping relationship between the effect function command words and the lighting adjustment control instructions, generate lighting effect function adjustment control instruction information.

[0055] In this embodiment, the mapping relationship between the preset special effects function command words and the lighting adjustment control instructions can be preset by humans. It includes command words and control instructions corresponding to flowing water effects, meteor effects, flashing effects, breathing effects, and various festival combination effects and timed switches. The preset special effects operation parameters include the lighting sequence, light spot movement speed, brightness change curve, and light alternation frequency. Then, the instruction content in the recognition information of multiple lighting special effects function adjustment command words is analyzed. Then, the parameter assignment and instruction matching are completed by combining the preset special effects function command words and the lighting adjustment control instructions. Finally, the instruction encapsulation is completed according to the two-wire signal encoding rules to generate lighting special effects function adjustment control instruction information.

[0056] The voice recognition-based intelligent lighting adjustment method provided in this application effectively reduces the probability of command confusion, ensures the stability of basic lighting operations and dynamic effects operation, thereby enriching the functional dimensions of voice lighting control and improving the operational reliability of intelligent lighting adjustment.

[0057] Corresponding to the method in the above embodiments, Figure 7 The diagram shows a structural block diagram of a voice recognition-based intelligent lighting adjustment device provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiments of this application are shown. Figure 7 The example of a voice recognition-based intelligent lighting adjustment device can be the execution subject of the voice recognition-based intelligent lighting adjustment method provided in the aforementioned embodiment 1.

[0058] Reference Figure 7 The voice recognition-based intelligent lighting control device includes: The environmental sound monitoring information acquisition module 710 is used to acquire environmental sound monitoring information; The speech recognition information generation module 720 is used to generate speech recognition information based on the environmental sound monitoring information, the preset human voice recognition model, and the preset accent differentiation and recognition model. The voice command word recognition information generation module 730 is used to generate multiple voice command word recognition information based on the voice recognition information and multiple preset voice command word recognition models; The lighting adjustment control instruction information generation module 740 is used to generate lighting adjustment control instruction information based on the recognition information of multiple voice command words and the mapping relationship between multiple preset command words and lighting adjustment control instructions, so as to perform lighting adjustment control processing through the lighting adjustment control instruction information.

[0059] The process by which each module in the voice recognition-based intelligent lighting control device provided in this application implements its respective function can be found in the foregoing. Figure 1 The description of Embodiment 1 shown will not be repeated here.

[0060] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0061] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0062] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0063] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0064] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. It should also be understood that although the terms "first," "second," etc., are used in the text to describe various elements in some embodiments of this application, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first table may be named a second table, and similarly, a second table may be named a first table, without departing from the scope of the various described embodiments. Both the first table and the second table are tables, but they are not the same table.

[0065] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0066] The voice recognition-based intelligent lighting adjustment method provided in this application can be applied to terminal devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality / virtual reality devices, laptops, super mobile personal computers, netbooks, and personal digital assistants. This application does not impose any restrictions on the specific type of terminal device.

[0067] For example, the terminal device may be a station in a WLAN, a cellular phone, a cordless phone, a session initiation protocol phone, a wireless local loop station, a personal digital processing device, a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, an in-vehicle device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a set-top box, a user premises equipment, and / or other devices for communication over a wireless system, as well as next-generation communication systems, such as mobile terminals in 5G networks or mobile terminals in future evolved public terrestrial mobile networks, etc.

[0068] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. For example... Figure 8 As shown, the terminal device 8 in this embodiment includes: at least one processor 80 ( Figure 8 Only one is shown in the image), and a memory 81 is stored in which a computer program 82 that can run on the processor 80 is stored. When the processor 80 executes the computer program 82, it implements the steps in the various embodiments of the intelligent lighting adjustment method based on voice recognition described above, for example... Figure 1 Steps S101 to S104 are shown. Alternatively, when the processor 80 executes the computer program 82, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 7 The functions of modules 710 to 740 are shown.

[0069] The terminal device 8 can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art will understand that... Figure 8 This is merely an example of terminal device 8 and does not constitute a limitation on terminal device 8. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may also include input transmission devices, network access devices, buses, etc.

[0070] The processor 80 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0071] In some embodiments, the memory 81 may be an internal storage unit of the terminal device 8, such as a hard disk or memory of the terminal device 8. The memory 81 may also be an external storage device of the terminal device 8, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., equipped on the terminal device 8. Furthermore, the memory 81 may include both internal and external storage units of the terminal device 8. The memory 81 is used to store operating systems, applications, bootloaders, data, and other programs, such as the program code of computer programs. The memory 81 can also be used to temporarily store data that has been sent or will be sent.

[0072] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0073] This application also provides a terminal device, which includes at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor. When the processor executes the computer program, it causes the terminal device to implement the steps in any of the above method embodiments.

[0074] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0075] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0076] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0077] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0078] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0079] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0080] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for intelligent adjustment of light based on voice recognition, characterized in that, include: Obtain environmental sound monitoring information; Based on the environmental sound monitoring information, the preset human voice recognition model, and the preset accent differentiation and recognition model, speech recognition information is generated; Based on the speech recognition information and multiple preset speech command word recognition models, multiple speech command word recognition information is generated; Based on the recognition information of multiple voice command words and the mapping relationship between multiple preset command words and lighting adjustment control instructions, lighting adjustment control instruction information is generated so as to perform lighting adjustment control processing through the lighting adjustment control instruction information.

2. The voice recognition based intelligent light adjustment method as claimed in claim 1, wherein, The step of generating speech recognition information based on the environmental sound monitoring information, the preset human voice recognition model, and the preset accent differentiation and recognition model specifically includes: Based on the environmental sound monitoring information and the preset human voice recognition model, human voice recognition information is generated; Based on a preset human voice noise reduction processing model, noise reduction processing is performed according to the human voice recognition information to generate noise-reduced human voice recognition information. Based on the noise-reduced human voice recognition information and the preset accent differentiation and recognition model, speech recognition information is generated.

3. The intelligent lighting adjustment method based on voice recognition as described in claim 2, characterized in that, The preset voice recognition model includes a preset voice recognition trigger command generation sub-model and a preset voice recognition generation sub-model; The step of generating human voice recognition information based on the environmental sound monitoring information and the preset human voice recognition model specifically includes: Based on the environmental sound monitoring information and the preset human voice recognition trigger command, a sub-model is generated to generate human voice recognition trigger command information; Based on the environmental sound monitoring information, the human voice recognition trigger instruction information, and the preset human voice recognition generation sub-model, human voice recognition information is generated.

4. The intelligent lighting adjustment method based on voice recognition as described in claim 1, characterized in that, The preset voice command word recognition model includes a preset voice accent matching sub-model and a preset voice command word recognition sub-model; wherein, the preset voice accent matching sub-model and the preset voice command word recognition sub-model correspond one-to-one. The step of generating multiple voice command word recognition information based on the speech recognition information and multiple preset voice command word recognition models specifically includes: Based on the speech recognition information and multiple preset speech accent matching sub-models, a matched speech accent matching sub-model is obtained; Based on the matched speech accent matching sub-model, multiple preset speech command word recognition sub-models are matched to obtain the matched speech command word recognition sub-model. Based on the speech recognition information and the matched speech command word recognition sub-model, multiple speech command word recognition information is generated.

5. The intelligent lighting adjustment method based on voice recognition as described in claim 4, characterized in that, Multiple preset speech accent matching sub-models include a preset first speech accent matching sub-model and a preset second speech accent matching sub-model; The step of obtaining the matched speech accent matching sub-model based on the speech recognition information and multiple preset speech accent matching sub-models specifically includes: Based on the speech recognition information and the preset first speech accent matching sub-model, the first speech accent matching probability information is calculated. Based on the speech recognition information and the preset second speech accent matching sub-model, the second speech accent matching probability information is calculated. The difference between the first and second speech accent matching probability information is calculated to obtain speech accent matching probability difference information. Based on the speech accent matching probability difference information, the preset first speech accent matching sub-model, and the preset second speech accent matching sub-model, the matched speech accent matching sub-model is obtained.

6. The intelligent lighting adjustment method based on voice recognition as described in claim 1, characterized in that, The voice command word recognition information includes command word recognition information for adjusting basic lighting functions and command word recognition information for adjusting lighting effects functions.

7. The intelligent lighting adjustment method based on voice recognition as described in claim 6, characterized in that, The mapping relationship between multiple preset command words and lighting adjustment control commands includes the mapping relationship between preset basic function command words and lighting adjustment control commands, as well as the mapping relationship between preset special effects function command words and lighting adjustment control commands. The lighting adjustment control command information includes basic lighting function adjustment control command information and lighting special effects function adjustment control command information; The step of generating lighting adjustment control instruction information based on the recognition information of multiple voice command words and the mapping relationship between multiple preset command words and lighting adjustment control instructions, and performing lighting adjustment control processing through the lighting adjustment control instruction information, specifically includes: Based on the recognition information of the multiple basic lighting function adjustment command words and the preset mapping relationship between basic function command words and lighting adjustment control instructions, basic lighting function adjustment control instruction information is generated. Based on the recognition information of the multiple lighting effect function adjustment command words and the preset mapping relationship between the effect function command words and the lighting adjustment control instructions, lighting effect function adjustment control instruction information is generated.

8. A voice recognition based intelligent light adjustment device, characterized in that, include: An environmental sound monitoring information acquisition module is used to acquire environmental sound monitoring information; The speech recognition information generation module is used to generate speech recognition information based on the environmental sound monitoring information, the preset human voice recognition model, and the preset accent differentiation and recognition model. The voice command word recognition information generation module is used to generate multiple voice command word recognition information based on the voice recognition information and multiple preset voice command word recognition models; The lighting adjustment control instruction information generation module is used to generate lighting adjustment control instruction information based on the recognition information of multiple voice command words and the mapping relationship between multiple preset command words and lighting adjustment control instructions, so as to perform lighting adjustment control processing through the lighting adjustment control instruction information.

9. A terminal device, comprising: The terminal device includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.