Illuminating lamp control method for predicting user emotion based on voice and related device
By collecting user voice and environmental data, and using an emotion recognition model to automatically adjust the color temperature and brightness of lighting, the problem of users needing professional knowledge to adjust the color temperature and brightness of smart lights is solved, thus achieving intelligent emotional adaptation of lighting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EARDA TECH CO LTD
- Filing Date
- 2026-03-11
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, when users adjust the color temperature and brightness of smart lights via voice, they need to have a high level of dimming knowledge, and the level of intelligence is low, making it difficult to adjust the lights to the ideal color temperature and brightness.
The system collects user voice signals and environmental data through a voice acquisition module, combines user attribute information, uses a user emotion recognition model to predict the user's emotion score, and automatically adjusts the color temperature and brightness of the lighting based on the emotion score.
It enables the adjustment of light color temperature and brightness to match the user's mood without user intervention, thus improving the intelligence level of lighting control.
Smart Images

Figure CN121865482A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart home technology, and in particular to a lighting control method and related device based on voice prediction of user emotions. Background Technology
[0002] With the development of smart home technology, smart lighting fixtures are widely used in people's lives. Users can control smart lighting fixtures to turn them on and off, as well as adjust their color temperature and brightness, via voice commands.
[0003] In existing technologies, when users adjust the color temperature and brightness of smart lights via voice, their voice needs to contain color temperature and brightness values, or semantics indicating the adjustment trend of color temperature and brightness, such as "brighter" or "warmer." On the one hand, this requires users to have a high level of dimming knowledge. When users have a low level of mastery of dimming knowledge related to color temperature and brightness, it is difficult to adjust the lights to the ideal color temperature and brightness. Users need to repeatedly adjust the color temperature and brightness of the lights through voice commands, or even give up adjusting the lights altogether. On the other hand, adjusting the color temperature and brightness of smart lights directly through user voice commands has a low level of intelligence. Summary of the Invention
[0004] This invention provides a lighting control method and related device based on voice prediction of user emotions, which adjusts the color temperature and brightness of the light according to the user's emotions, so that the color temperature and brightness of the light are adapted to the user's emotions, thereby improving the intelligence of lighting control.
[0005] In a first aspect, the present invention provides a lighting control method based on voice prediction of user emotions, comprising: After the lights are turned on and the smart mode is entered, the voice signal of the user to be identified is collected through the voice acquisition module. Collect environmental data of the environment where the lighting fixture is located and obtain attribute information of the user to be identified; The voice signal, the environmental data, and the attribute information are input into the user emotion recognition model to obtain the emotion score of the user to be identified. The target color temperature and target brightness of the lighting lamp are determined based on the emotion score. The color temperature of the lighting lamp is adjusted to the target color temperature, and the brightness of the lighting lamp is adjusted to the target brightness.
[0006] Secondly, the present invention provides a lighting control device based on voice prediction of user emotions, comprising: The voice signal acquisition module is used to acquire the voice signal of the user to be identified after the lighting is turned on and the smart mode is entered. An environmental data and attribute information acquisition module is used to collect environmental data of the environment where the lighting fixture is located and to obtain the attribute information of the user to be identified. The emotion score prediction module is used to input the voice signal, the environmental data, and the attribute information into the user emotion recognition model to obtain the emotion score of the user to be identified. A target color temperature and brightness determination module is used to determine the target color temperature and target brightness of the lighting lamp based on the emotion score; The lighting adjustment module is used to adjust the color temperature of the lighting lamp to the target color temperature and the brightness of the lighting lamp to the target brightness.
[0007] Thirdly, the present invention provides a lighting lamp, the lighting lamp comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the lighting control method based on voice prediction of user emotions as described in the first aspect of the present invention.
[0008] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a processor to execute and implement the lighting control method based on voice prediction of user emotions as described in the first aspect of the present invention.
[0009] This invention, after the lighting is turned on and enters intelligent mode, identifies the user's emotion score through the user's voice signal, environmental data, and user attribute information. Based on the user's emotion score, it further determines the target color temperature and target brightness, adjusting the lighting's color temperature and brightness to the target levels. This achieves the goal of controlling the lighting's color temperature and brightness by recognizing the user's emotions through their voice, attribute information, and the surrounding environment. This ensures the lighting's color temperature and brightness match the user's emotions without requiring the user to convey color temperature and brightness values via voice or intervention, thus improving the intelligence level of lighting control.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a lighting control method based on voice prediction of user emotions provided in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the user emotion recognition model in an embodiment of the present invention; Figure 3 This is a schematic diagram of a lighting control device based on voice prediction of user emotions provided in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the structure of the lighting lamp provided in Embodiment 3 of the present invention. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0014] Example 1 Figure 1 This is a flowchart of a lighting control method based on voice prediction of user emotions provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where the color temperature and brightness of the lighting are adjusted based on user emotions. This method can be executed by a lighting control device based on voice prediction of user emotions. This lighting control device can be implemented in hardware and / or software and can be configured within the lighting fixture. Figure 1 As shown, the lighting control method based on voice prediction of user emotions includes: S101. After the lighting is turned on and the smart mode is entered, the voice signal of the user to be identified is collected through the voice acquisition module.
[0015] In this embodiment, the lighting can be a smart lighting fixture for home or commercial use that is controlled by voice. The lighting fixture includes a fixture body and a voice acquisition module. The smart mode can refer to the lighting fixture being in a voice-controlled mode. After the lighting fixture is turned on and enters the smart mode, the voice acquisition module acquires an initial voice signal segment of a preset duration according to a preset cycle. The background noise signal and a mixed voice signal including at least two users to be identified are separated from the initial voice signal. The mixed voice signal is then separated to obtain the voice signal of each user to be identified. The user to be identified can be a user who controls the lighting fixture, or is within the range of the lighting fixture, or is in the space where the lighting fixture is located.
[0016] Specifically, users can turn on the lights and switch to smart mode by using voice commands (such as "turn on the lights") or manual operation (such as pressing the smart mode button on the lights). The lights will automatically enter smart mode after being turned on. After the smart mode is activated, the voice acquisition module (which can be a microphone array containing more than 3 omnidirectional microphones, with a sampling frequency of 16kHz and a sampling accuracy of 16bit) will be automatically activated and enter the voice acquisition state. The preset cycle of the voice acquisition module is set to 5 seconds (which can be adjusted according to the actual scenario, ranging from 3 to 10 seconds). The duration of the initial voice signal segment acquired each time is 2 seconds. That is, every 5 seconds, 2 seconds of initial voice signal is acquired to avoid excessive power consumption and excessive data volume caused by continuous acquisition, while ensuring that the user's effective voice can be captured.
[0017] The initial voice signal is processed by separating it, removing background noise signals and retaining the mixed voice signal containing at least two users to be identified. For example, a Conv-TasNet-based sound source separation model can be used to perform spectral analysis on the initial voice signal to identify and separate non-human background noise signals (including TV playback audio, short video audio, environmental noise, etc., with the identification threshold set to: non-human voice signals accounting for ≥30% are judged as background noise signals). The remaining signal is the mixed voice signal containing at least two users to be identified (such as the voice signal of parents talking in the living room in a family scene).
[0018] The mixed speech signal is further separated to obtain the speech signal of each user to be identified. For example, DPRNN (Dual-Path Recurrent Neural Network) is used. When the mixed speech signal is input, DPRNN performs clustering and separation based on the speech features of different users (such as fundamental frequency, speech rate, timbre, etc.), and outputs multiple independent single-user speech signals. After separation, each single-user speech signal is denoised (for example, by using Wiener filtering algorithm) to remove residual weak background noise and obtain a clean speech signal of the user to be identified, which is used for subsequent emotion recognition steps.
[0019] S102. Collect environmental data of the environment where the lighting is located and obtain attribute information of the user to be identified.
[0020] Environmental data can include the brightness, temperature, and current time of the environment in which the lighting is located. Attribute information can include various user attributes. In this embodiment, an environmental sensing module can be integrated into the lighting. This environmental sensing module can include a temperature sensor, a brightness sensor, and a timing module. The environmental sensing module is activated synchronously with the voice acquisition module and collects environmental data synchronously according to the same preset cycle as the voice acquisition, ensuring the temporal correlation between environmental data and voice signals and improving the accuracy of subsequent emotion recognition.
[0021] Taking a living room light in a home as an example, a temperature sensor can collect the ambient temperature of the living room in real time. The collected temperature data is converted into a digital signal and transmitted to the light's processor for storage. A brightness sensor (such as a photoresistor) can collect the combined brightness of the ambient natural light and other light sources in the living room in real time, avoiding interference from ambient brightness on the user's emotional judgment and subsequent lighting adjustment. The current system time is obtained through a timing module and divided into preset time periods (such as 7:00-12:00 for the morning period, 12:00-18:00 for the afternoon period, 18:00-23:00 for the evening period, and 23:00-7:00 for the night period). The time period information is converted into corresponding codes (such as 1 for the morning period and 2 for the afternoon period) to facilitate subsequent emotion recognition model recognition.
[0022] In addition, user attribute information can be obtained through pre-storage and real-time matching, balancing convenience and security. For example, when a user uses the smart mode for the first time, they can enter their personal attribute information through the associated mobile app. The attribute information can include gender (male / female), age (divided by age group: under 18, 18-30, 31-50, over 50), and occupation (divided by industry: student, company employee, freelancer, retiree, etc.). The entered attribute information is stored in the local storage module of the lighting fixture or the associated cloud server after being de-identified and encrypted, and is associated with the user's voice characteristics (such as timbre and speech rate). A user attribute information database is formed. Of course, attribute information can also be obtained through real-time matching. For example, after the single-user speech signal is separated in step S101, the user's speech features (such as fundamental frequency and Mel frequency cepstral coefficients) are extracted and matched with the speech features in the user attribute information database. When the matching confidence is ≥85%, the user's pre-stored gender, age, and occupation information are retrieved. If the matching fails (such as first-time use or changes in speech features), a prompt instruction is issued through the voice interaction module (such as "Please enter your gender and age information"). The user provides attribute information through voice feedback, which is collected and stored in real time for subsequent emotion recognition.
[0023] S103. Input the voice signal, environmental data and attribute information into the user emotion recognition model to obtain the emotion score of the user to be recognized.
[0024] In this embodiment, the user emotion recognition model can be a neural network module used to predict the user's emotion score. This user emotion recognition model can be trained through the following steps: S1. Construct a user emotion recognition model, which includes a user feature extraction sub-model and a user pair matching sub-model.
[0025] The user emotion recognition model in this embodiment can be a hybrid deep learning model, which includes a cascaded user feature extraction sub-model and a user pair matching sub-model. The user feature extraction sub-model can employ a multimodal feature fusion network. For speech signals, acoustic features (such as MFCC and spectral features) can be extracted using a CNN+LSTM network. For attribute information (gender, age, occupation) and environmental data samples (temperature, brightness, time of day), they can be converted into dense vectors through an embedding layer. Finally, all features are concatenated and fused through a fully connected layer, outputting a 256-dimensional user feature.
[0026] The user pair matching sub-model may include a similarity calculation layer, an emotion score lookup layer, and an emotion feature library. The similarity calculation layer is used to calculate the distance between any two user features as the similarity, and the two users with the highest similarity are used to generate user pairs. The emotion score lookup layer is used to find the emotion scores of users in the user pair in the emotion feature library.
[0027] S2. Obtain the training dataset, which includes user pairs, the emotion scores corresponding to the user pairs, and the voice signal samples, attribute information samples, and environmental data samples of each user pair.
[0028] User pairs can consist of two user samples with similar or identical emotion scores. Each user pair corresponds to one emotion score, and the voice signal samples, attribute information samples, and environmental data samples of the two user samples in the user pair can be different. For example, user sample A and user sample B form a user pair with an emotion score of 3 (the emotion score ranges from 0 to 10, where 0 represents extreme sadness and 10 represents extreme happiness). User sample A and user sample B each have their own voice signal samples, attribute information samples, and environmental data samples. Specifically, voice signals, attribute information, and environmental data of users with different emotions in real life can be collected, and user pairs can be manually formed and assigned emotion scores to obtain a training dataset.
[0029] The training dataset in this embodiment can cover user data of different age groups, genders, occupations and various environmental scenarios to improve the generalization ability of the user emotion recognition model.
[0030] S3. Input the speech signal samples, attribute information samples, and environmental data samples of multiple user samples in the training dataset into the user emotion recognition model.
[0031] In one embodiment, batches of user samples (e.g., Batch Size set to 32) from the training dataset are input into the constructed user emotion recognition model. For each batch of input data, it is simultaneously fed into the user feature extraction sub-model to ensure parallel processing efficiency.
[0032] S4. In the user feature extraction sub-model, features are extracted from the speech signal sample, attribute information sample and environmental data sample of each user sample to obtain the user features of each user sample.
[0033] Specifically, in the user feature extraction sub-model, multi-dimensional data of each input user sample is extracted in parallel. Speech signal samples can be pre-emphasized, framed, and windowed before being input into the feature extraction network to obtain deep speech features. Attribute information samples and environmental data samples can be converted into numerical features through one-hot encoding or embedding layer processing. After the model normalizes the above features, it generates 256-dimensional (or 512-dimensional, etc.) user features for each user sample through the feature fusion layer.
[0034] S5. In the user-pair matching sub-model, calculate the feature similarity between the user features of each user sample and the user features of other user samples.
[0035] In the user pair matching sub-model, based on all user features output in step S4, the cosine similarity, L1 distance, or L2 distance between any two user features is calculated to obtain the feature similarity. The closer the feature similarity is to 1, the higher the feature matching degree of the two user samples, and the more likely they are to form a user pair with close interaction.
[0036] S6. Generate predicted user pairs using the two user samples with the highest feature similarity.
[0037] Based on the similarity calculated in step S5, the Hungarian algorithm or greedy matching algorithm can be used to select the two user features with the highest similarity to pair them to generate predicted user pairs. It should be noted that if a batch contains N user samples (N is an even number), then N / 2 predicted user pairs are generated. If there are an odd number of user samples, then the single user sample with the lowest similarity is removed.
[0038] S7. Calculate the loss rate using user-pair samples and predicted user-pairs.
[0039] In one embodiment, a hybrid loss function can be constructed by combining a contrastive loss function with a mean squared error loss (MSE loss). This hybrid loss function calculates the loss rate between the predicted results and the actual samples. The contrastive loss penalizes mismatches between predicted and actual user pairs, bringing positive pairs (actual user pairs) closer in feature distance and negative pairs further apart. The mean squared error loss calculates the difference between the predicted sentiment score of a user pair (mapped via feature similarity) and the actual sentiment score (the sentiment score of the user pair). In one example, the total loss rate can be calculated using the following hybrid loss function: L = L contrast +λ×L MSE ; Where λ is the balance coefficient, which can be 0.5; ; N is the total number of user pairs in a batch of input data, y i The value of the i-th predicted user pair is determined when the i-th predicted user pair is the same as the user pair sample (positive sample, correct match). i =1, when the i-th predicted user pair is not the same as the user pair sample (negative sample, incorrect match). i =0,d i Let be the Euclidean distance between the user features of the two users in the i-th predicted user pair, and let margin be the boundary value (a preset constant used to control the feature distance boundary between positive and negative samples; a value of 2.0 is recommended, but can be adjusted according to training accuracy to ensure that the distance between positive samples is less than margin and the distance between negative samples is greater than margin). max(margin - d) i ,0) is used when the negative sample (y) i The feature distance d (=0) i If d is less than the margin, calculate the penalty term; if d i If the margin is greater than or equal to 0, no penalty is required.
[0040] For the correctly predicted user pair (y) i =1), by comparing the loss function L contrast This enables the model to be minimized. This makes the user characteristics of two users more similar, which is helpful for user pairs (y) with inaccurate predictions. i =0), the model can minimize This makes the distance between the user characteristics of two users greater than the margin, thereby improving the accuracy of the matching.
[0041] In one embodiment, the mean square error loss L MSE The calculation method is as follows: ; N is the total number of user pairs in a batch of input data, s i Let be the true sentiment score of the i-th user, which is the sentiment score of the user to which the i-th user belongs to the sample. Let represent the predicted sentiment score of the user pair to which the i-th user belongs. The sentiment score of the predicted user pair can be obtained by mapping the feature similarity of the predicted user pair. The mapping logic is to convert the feature similarity of the predicted user pair (0-1)×10 into a score of 0-10.
[0042] Through mean square error loss L MSE The average of the squared differences of all predicted user pairs is used to obtain the overall mood score prediction bias. By minimizing this bias, the model makes the predicted mood scores more closely match the actual mood scores, providing a precise basis for the target color temperature and brightness of subsequent lighting adjustments.
[0043] S8. Determine whether the conditions for stopping training are met.
[0044] In one embodiment, the training stop condition may be that the number of training epochs reaches a preset maximum number of epochs (100 epochs in this embodiment). In another embodiment, the training stop condition may be that the loss rate L decreases by less than a preset threshold for 5 consecutive epochs. If the training stop condition is met, then S9 is executed; otherwise, S10 is executed.
[0045] S9. Determine that the user emotion recognition model has completed training, and store the user features and emotion scores of each user in the sample into the emotion feature database.
[0046] When the training stop condition is met, the model training is terminated, the current model parameters are determined to be the optimal parameters, and the trained user emotion recognition model is obtained. At the same time, the user features of each user sample in all user pairs in the training dataset are associated and stored with their corresponding emotion scores to build an emotion feature library. This feature library is used for fast matching and emotion score query in subsequent practical applications.
[0047] S10. Adjust the model parameters of the user emotion recognition model according to the loss rate, and return to S3.
[0048] If the training stop condition is not met, the Adam optimizer (with a learning rate of 0.001) is used to backpropagate and update all trainable parameters of the user emotion recognition model (including the weights of the feature extraction sub-model and the user's weights on the matching sub-model) based on the total loss rate calculated in step S7, and then the process returns to S3 to continue training the model.
[0049] After the user emotion recognition model is trained and deployed, the emotion feature library includes user features and emotion scores from multiple user samples. For each user to be identified, ambient temperature, ambient brightness, current time period, the user's voice signal, and gender, age, and occupation are input into the user emotion recognition model. The user feature extraction sub-model extracts user features based on ambient temperature, ambient brightness, current time period, the user's voice signal, gender, age, and occupation. In the user pair matching sub-model, the feature similarity between the user to be identified and the user features of each user sample in the emotion feature library is calculated. The user sample with the highest feature similarity is used to generate the emotion recognition model. The system identifies user pairs and uses feature similarity as the confidence level for each pair. When the confidence level is greater than or equal to a preset confidence threshold, the emotion score of the user sample in the user pair is determined as the emotion score of the user to be identified. When the confidence level is less than the preset confidence threshold, the emotion recognition of the user to be identified is determined to have failed. The system then sends the user to be identified, along with ambient temperature, ambient brightness, current time period, the user's voice signal, gender, age, and occupation, to a human recognition terminal. The human recognition terminal user determines the emotion score based on the reviewer's operation. After receiving the emotion score returned by the human recognition terminal, the user features and emotion score of the user to be identified are stored in the emotion feature database.
[0050] Specifically, the environmental data (ambient temperature, ambient brightness, current time period) collected and processed in step S102, the attribute information of the user to be identified (gender, age, occupation), and the voice signal of the user to be identified separated in step S101 can be uniformly converted into a standardized format that the model can recognize. Among them, ambient temperature and ambient brightness are converted into numerical data (retaining 1 decimal place), the current time period uses the preset encoding (morning=1, afternoon=2, evening=3, night=4), gender (male=1, female=0), age (under 18=1, 18-30=2, 31-50=3, over 50=4), and occupation (student=1, company employee=2, freelancer=3, retiree=4) are converted into coded data. The voice signal is processed by pre-emphasis, framing, and windowing to extract MFCC acoustic features (128 dimensions). All data are normalized to the [0,1] interval to avoid the difference in units affecting the model's recognition accuracy.
[0051] The multi-dimensional data of each user to be identified is input into the user feature extraction sub-model. The user feature extraction sub-model extracts user features and outputs them to the user pair matching sub-model. The user pair matching sub-model calls the emotion feature library and uses the cosine similarity algorithm to calculate the feature similarity between the user features of the user to be identified and the user features of each user sample in the emotion feature library. The user sample with the highest feature similarity is selected and paired with the user to be identified to generate a user pair. The maximum feature similarity is directly used as the confidence score of the user pair. When the confidence score is greater than the confidence score threshold, the emotion of the user to be identified is determined to be successful. The emotion score of the user sample in the generated user pair is directly determined as the emotion score of the user to be identified. For example, if the emotion score of the user sample is 7.5 and the confidence score between the user to be identified and the user sample is 0.88, then the emotion score of the user to be identified is determined to be 7.5. This emotion score is used to determine the target color temperature and brightness of the lighting lamp.
[0052] In another embodiment, if the confidence level is less than the confidence level threshold, the emotion recognition of the user to be identified is deemed to have failed. The reasons for the failure may include significant differences between the characteristics of the user to be identified and those of all sample users in the emotion feature database, abnormal user voice characteristics, sudden changes in environmental data, etc., requiring a manual recognition completion process. The processor of the lighting lamp packages the relevant data of the user to be identified and sends it to a preset manual recognition terminal (such as a dedicated review computer or an APP on the mobile terminal of the user to be identified) via a wireless communication module (such as WiFi or Bluetooth). The packaged data includes: standardized multi-dimensional input data of the user to be identified (ambient temperature, ambient brightness, current time period, voice signal, gender, age, occupation), the 256-dimensional feature vector of the user to be identified, the confidence level value, and the recognition failure indicator. After receiving the data, the manual recognition terminal automatically displays the relevant information of the user to be identified (the voice signal can be played through the terminal, and the environmental data and attribute information are displayed in the form of a form). The reviewer, in conjunction with professional psychological standards, manually evaluates the emotion score of the user to be identified (the score range is 0-10, consistent with the emotion score dimensions of the sample users) based on the user's voice tone, attribute information, and environmental scene, and confirms the submission through the terminal operation.
[0053] The human recognition terminal feeds back the emotion score determined by the reviewer to the lighting processor via the wireless communication module. After receiving the emotion score, the processor associates the 256-dimensional feature vector of the user to be identified with the human-assessed emotion score and stores it in the emotion feature database, realizing the dynamic updating of the emotion feature database. The newly added data can be called up during subsequent training to improve the model's recognition accuracy for similar users and reduce the frequency of subsequent human recognition.
[0054] In this embodiment, whether the emotion recognition is successful or not, the emotion score of the user to be identified is determined. If there are multiple users to be identified, the emotion score of each user is determined one by one, so as to provide an accurate basis for determining the target color temperature and brightness of the lighting lamp based on the emotion score.
[0055] S104. Determine the target color temperature and target brightness of the lighting fixtures based on the emotion score.
[0056] In one embodiment, after completing the emotion recognition of the user to be identified, it can be determined whether the number of users to be identified is greater than or equal to 2. If not, the target color temperature and target brightness corresponding to the emotion score are found in the preset emotion and light parameter lookup table. If so, the voice features of the voice signal of each user to be identified are calculated, the emotion weight of the user to be identified is calculated based on the voice features, the comprehensive emotion score is calculated using the emotion weight and emotion score of each user to be identified, and the target color temperature and target brightness corresponding to the comprehensive emotion score are found in the preset emotion and light parameter lookup table.
[0057] This embodiment can pre-build a table of emotions and lighting parameters. This table can be pre-built based on the psychological principle of emotion and lighting adaptation and a large amount of user test data, and stored in the local storage unit of the lighting fixture. The table of lighting parameters can be flexibly adjusted according to the actual application scenario (home, office, etc.). In one example, the table of lighting parameters is as follows (using an emotion score range of 0-10, color temperature in K, and brightness in lux as an example): It should be noted that in the above comparison table, the mood score corresponds to the lighting parameters. If the mood score is a specific value (such as 7.5), the corresponding target color temperature and target brightness can be calculated using linear interpolation. For example, if 7.5 is in (5, 8], the color temperature = 3500 + (7.5-5) × (4500-3500) / (8-5) = 4333K. The brightness is calculated similarly to ensure adjustment accuracy.
[0058] After identifying the emotion scores of all users to be identified, the number of users N to be identified can be counted. It is determined whether N is greater than or equal to 2. When the number of users N to be identified is less than 2, it is determined that it belongs to a single-person scene. The emotion score of the user to be identified is directly extracted, and the corresponding range is found in the above-mentioned preset emotion and lighting parameter comparison table to determine the target color temperature and target brightness. For example, if the emotion score of the user to be identified is 7.2, which belongs to the interval (5, 8], the target color temperature = 3500 + (7.2-5) × (4500-3500) / (8-5) = 4233K, and the target brightness = 300 + (7.2-5) × (300-400) / (8-5) = 373 lux, which is suitable for a pleasant and relaxed emotional state.
[0059] If the number of users to be identified, N, is greater than or equal to 2, it is determined to be a multi-user scenario. A comprehensive emotion score needs to be calculated. The speech signal of each user to be identified after separation in step S101 can be retrieved, and the speech feature parameters can be extracted using the MFCC algorithm. These speech feature parameters may include parameters such as fundamental frequency, speech rate, Mel-frequency cepstral coefficients, and amplitude of the speech signal. The speech activity of each user is calculated using the speech feature parameters. In one embodiment, speech activity = (user speech duration / total speech duration) × (user speech energy / average speech energy), where user speech duration is the effective speech duration of the user within the acquisition period (excluding silent segments), total speech duration is the sum of the effective speech durations of all users to be identified, and speech energy is calculated through the amplitude of the speech signal. After obtaining the speech activity of each user to be identified, normalization processing is performed, and the normalized speech activity is determined as the emotion weight of the user to be identified.
[0060] For example, if there are two users to be identified, user A's voice duration is 1.2 seconds and voice energy is 0.8, and user B's voice duration is 0.8 seconds and voice energy is 0.6, then the total voice duration can be determined to be 2.0 seconds and the average voice energy is 0.7. Then, user A's voice activity = (1.2 / 2.0) × (0.8 / 0.7) ≈ 0.6857, and user B's voice activity = (0.8 / 2.0) × (0.6 / 0.7) ≈ 0.3429. After normalization, user A's emotion weight ≈ 0.667, and user B's emotion weight ≈ 0.333.
[0061] This embodiment can use a weighted summation algorithm to calculate the overall emotion score, combining the emotion weight w of each user. i And mood score s i The overall emotion score S is calculated using the following formula: ; Continuing with the example above, user A's emotion score is 8.0 and user B's emotion score is 6.0. Therefore, the overall emotion score S = (0.667 × 8.0) + (0.333 × 6.0) ≈ 7.334 points.
[0062] After calculating the overall emotion score, the same logic as in the single-person scene (range matching or linear interpolation) is used to find the corresponding target color temperature and target brightness to ensure that the lighting adjustment is adapted to the overall emotions of multiple people.
[0063] S105. Adjust the color temperature of the lighting lamp to the target color temperature and adjust the brightness of the lighting lamp to the target brightness.
[0064] In an optional embodiment, before executing S105, the color temperature difference can be calculated by measuring the difference between the target color temperature and the current color temperature, and the brightness difference can be calculated by measuring the difference between the target brightness and the current brightness. It is then determined whether the absolute value of the color temperature difference is greater than a preset color temperature value. If so, the step of adjusting the color temperature of the lighting lamp to the target color temperature is executed; otherwise, adjusting the color temperature of the lighting lamp is prohibited. Similarly, it is determined whether the absolute value of the brightness difference is greater than a preset brightness value. If so, the step of adjusting the brightness of the lighting lamp to the target brightness is executed; otherwise, adjusting the brightness of the lighting lamp is prohibited. This avoids frequent adjustments to the color temperature and brightness of the lighting lamp when the current brightness and color temperature are close to the target brightness and color temperature, reducing the frequency of light adjustment and extending the lifespan of the lighting lamp.
[0065] In this embodiment, when adjusting color temperature and brightness, it can first determine whether the target color temperature is greater than the upper limit color temperature or less than the lower limit color temperature. If so, if the target color temperature is greater than the upper limit color temperature, the color temperature of the lighting lamp is adjusted to the upper limit color temperature; if the target color temperature is less than the lower limit color temperature, the color temperature of the lighting lamp is adjusted to the lower limit color temperature. If not, the color temperature of the lighting lamp is smoothly adjusted to the target color temperature step by step according to the preset color temperature adjustment step. It also determines whether the target brightness is greater than the upper limit brightness or less than the lower limit brightness. If so, if the target brightness is greater than the upper limit brightness, the brightness of the lighting lamp is adjusted to the upper limit brightness; if the target brightness is less than the lower limit brightness, the brightness of the lighting lamp is adjusted to the lower limit brightness. If not, the brightness of the lighting lamp is smoothly adjusted to the target brightness step by step according to the preset brightness adjustment step. On the one hand, this avoids overshoot caused by the color temperature and brightness of the lighting lamp exceeding the upper or lower limit values. On the other hand, smoothly adjusting the color temperature and brightness step by step according to the preset adjustment step can avoid sudden changes in the color temperature and brightness of the lighting lamp, reducing the impact of lighting effects on user experience.
[0066] This invention, after the lighting is turned on and enters intelligent mode, identifies the user's emotion score through the user's voice signal, environmental data, and user attribute information. Based on the user's emotion score, it further determines the target color temperature and target brightness, adjusting the lighting's color temperature and brightness to the target levels. This achieves the goal of controlling the lighting's color temperature and brightness by recognizing the user's emotions through their voice, attribute information, and the surrounding environment. This ensures the lighting's color temperature and brightness match the user's emotions without requiring the user to convey color temperature and brightness values via voice or intervention, thus improving the intelligence level of lighting control.
[0067] Furthermore, the user emotion recognition model in this embodiment samples the voice signals, environmental data, attribute information, and emotions of real users to establish an emotion feature library. When recognizing user emotions, the features of the user to be recognized are matched with the features in the emotion feature library to generate user pairs, thereby determining the emotion score of the user to be recognized. This realizes the determination of the emotion score of the user to be recognized by referring to the user features of existing real users. The recognized emotion score is closer to the real scene, thus improving the accuracy of the recognized emotion score.
[0068] Example 2 Figure 3 This is a schematic diagram of a lighting control device based on voice prediction of user emotions, provided in Embodiment 2 of the present invention. Figure 3 As shown, the lighting control device based on voice prediction of user emotions includes: The voice signal acquisition module 301 is used to acquire the voice signal of the user to be identified after the lighting is turned on and the smart mode is entered. The environmental data and attribute information acquisition module 302 is used to collect environmental data of the environment where the lighting fixture is located and to obtain attribute information of the user to be identified. The emotion score prediction module 303 is used to input voice signals, environmental data and attribute information into the user emotion recognition model to obtain the emotion score of the user to be identified; The target color temperature and brightness determination module 304 is used to determine the target color temperature and target brightness of the lighting lamp based on the emotion score. The lighting adjustment module 305 is used to adjust the color temperature of the lighting lamp to the target color temperature and the brightness of the lighting lamp to the target brightness.
[0069] The lighting control device based on voice prediction of user emotions provided in the embodiments of the present invention can execute the lighting control method based on voice prediction of user emotions provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0070] Example 3 Figure 4 A schematic diagram of the structure of a lighting lamp 40 that can be used to implement an embodiment of the present invention is shown.
[0071] like Figure 4As shown, the lighting lamp 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 and a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded into the RAM 43 from storage unit 47. The RAM 43 can also store various programs and data required for the operation of the lighting lamp 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0072] Multiple components in the lighting lamp 40 are connected to the I / O interface 45, including: an input unit 46, such as a voice acquisition module, an environmental sensor, etc.; a storage unit 47, such as ROM, RAM, etc.; and a communication unit 48, such as a wireless communication transceiver. The communication unit 48 allows the lighting lamp 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0073] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as a lighting control method based on voice prediction of user emotions.
[0074] In some embodiments, the lighting control method based on voice prediction of user emotions can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 47. In some embodiments, part or all of the computer program can be loaded and / or installed on the lighting lamp 40 via ROM 42 and / or communication unit 48. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the lighting control method based on voice prediction of user emotions described above can be performed. Alternatively, in other embodiments, processor 41 can be configured to perform the lighting control method based on voice prediction of user emotions by any other suitable means (e.g., by means of firmware).
[0075] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0076] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0077] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0078] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0079] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A lighting control method based on voice prediction of user emotions, characterized in that, include: After the lights are turned on and the smart mode is entered, the voice signal of the user to be identified is collected through the voice acquisition module. Collect environmental data of the environment where the lighting fixture is located and obtain attribute information of the user to be identified; The voice signal, the environmental data, and the attribute information are input into the user emotion recognition model to obtain the emotion score of the user to be identified. The target color temperature and target brightness of the lighting lamp are determined based on the emotion score. The color temperature of the lighting lamp is adjusted to the target color temperature, and the brightness of the lighting lamp is adjusted to the target brightness.
2. The method according to claim 1, characterized in that, After the lights are turned on and the system enters smart mode, the voice signal of the user to be identified is collected through the voice acquisition module, including: After the lights are turned on and the smart mode is entered, the voice acquisition module collects initial voice signal segments of a preset duration according to a preset cycle. Separate the background noise signal and the mixed speech signal including at least two users to be identified from the initial speech signal; The mixed speech signal is separated to obtain the speech signal of each user to be identified.
3. The method according to claim 1, characterized in that, The user emotion recognition model is trained through the following steps: Construct a user emotion recognition model, which includes a user feature extraction sub-model and a user pair matching sub-model; Obtain a training dataset, which includes user pairs, emotion scores corresponding to the user pairs, and voice signal samples, attribute information samples, and environmental data samples for each user sample in the user pairs. The speech signal samples, attribute information samples, and environmental data samples of multiple user samples in the training dataset are input into the user emotion recognition model; The user feature extraction sub-model extracts features from the speech signal sample, attribute information sample, and environmental data sample of each user sample to obtain the user features of each user sample. The user-pair matching sub-model calculates the feature similarity between the user features of each user sample and the user features of other user samples. The two user samples with the highest feature similarity are used to generate predicted user pairs; The loss rate is calculated using the user pair samples and the predicted user pairs; Determine whether the conditions for stopping training are met; If so, determine that the user emotion recognition model has completed training, and store the user features and emotion scores of each user in each user sample in the emotion feature database; If not, adjust the model parameters of the user emotion recognition model according to the loss rate, and return to the step of inputting the speech signal samples, attribute information samples, and environmental data samples of multiple user samples in the training dataset into the user emotion recognition model.
4. The method according to any one of claims 1-3, characterized in that, The user emotion recognition model includes a user feature extraction sub-model and a user pair matching sub-model. The user pair matching sub-model includes an emotion feature library, which includes user features and emotion scores of multiple sample users. The environmental data includes environmental temperature, environmental brightness, and the current time period. The attribute information includes the user's gender, age, and occupation. The voice signal, the environmental data, and the attribute information are input into a user emotion recognition model to obtain the emotion score of the user to be identified, including: For each user to be identified, the ambient temperature, ambient brightness, current time period, the user's voice signal, the user's gender, age, and occupation are input into the user emotion recognition model; The user feature extraction sub-model extracts the user features of the user to be identified based on the ambient temperature, ambient brightness, current time period, the voice signal of the user to be identified, the user's gender, age, and occupation. In the user pair matching sub-model, the feature similarity between the user features of the user to be identified and the user features of each sample user in the emotion feature database is calculated. User pairs are generated by using the sample user with the highest feature similarity and the user to be identified, and the feature similarity is used as the confidence level of the user pair; When the confidence level is greater than or equal to a preset confidence threshold, the emotion score of the sample user in the user pair is determined as the emotion score of the user to be identified. When the confidence level is less than a preset confidence threshold, it is determined that the emotion recognition of the user to be identified has failed. The user to be identified is sent to the human recognition terminal along with ambient temperature, ambient brightness, current time period, the user's voice signal, gender, age, and occupation. The human recognition terminal user determines the emotion score based on the operation of the reviewer. After receiving the emotion score returned by the human recognition terminal, the user characteristics of the user to be identified and the emotion score are stored in the emotion feature database.
5. The method according to any one of claims 1-3, characterized in that, Determining the target color temperature and target brightness of the lighting lamp based on the emotion score includes: Determine if the number of users to be identified is greater than or equal to 2; If not, find the target color temperature and target brightness corresponding to the emotion score in the preset emotion and lighting parameter comparison table; If so, calculate the speech features of the speech signal for each user to be identified; The emotion weight of the user to be identified is calculated based on the voice features. A comprehensive emotion score is calculated using the emotion weight and emotion score of each user to be identified. Find the target color temperature and target brightness corresponding to the comprehensive emotion score in the preset emotion and lighting parameter comparison table.
6. The method according to any one of claims 1-3, characterized in that, Before adjusting the color temperature of the lighting lamp to the target color temperature and the brightness of the lighting lamp to the target brightness, the method further includes: The color temperature difference is obtained by calculating the difference between the target color temperature and the current color temperature, and the brightness difference is obtained by calculating the difference between the target brightness and the current brightness; Determine whether the absolute value of the color temperature difference is greater than a preset color temperature value; If so, proceed with the step of adjusting the color temperature of the lighting lamp to the target color temperature; If not, adjusting the color temperature of the lighting fixture is prohibited; Determine whether the absolute value of the brightness difference is greater than a preset brightness value; If so, proceed with the step of adjusting the brightness of the lighting lamp to the target brightness; If not, adjusting the brightness of the lighting is prohibited.
7. The method according to any one of claims 1-3, characterized in that, Adjusting the color temperature of the lighting lamp to the target color temperature and adjusting the brightness of the lighting lamp to the target brightness includes: Determine whether the target color temperature is greater than the upper limit color temperature or less than the lower limit color temperature; If so, when the target color temperature is greater than the upper limit color temperature, adjust the color temperature of the lighting lamp to the upper limit color temperature; when the target color temperature is less than the lower limit color temperature, adjust the color temperature of the lighting lamp to the lower limit color temperature. If not, adjust the color temperature of the lighting lamp step by step to the target color temperature smoothly according to the preset color temperature adjustment step size; Determine whether the target brightness is greater than the upper limit brightness or less than the lower limit brightness; If so, when the target brightness is greater than the upper limit brightness, adjust the brightness of the lighting lamp to the upper limit brightness; when the target brightness is less than the lower limit brightness, adjust the brightness of the lighting lamp to the lower limit brightness. If not, adjust the brightness of the lighting lamp step by step smoothly to the target brightness according to the preset brightness adjustment step size.
8. A lighting control device based on voice prediction of user emotions, characterized in that, include: The voice signal acquisition module is used to acquire the voice signal of the user to be identified after the lighting is turned on and the smart mode is entered. An environmental data and attribute information acquisition module is used to collect environmental data of the environment where the lighting fixture is located and to obtain the attribute information of the user to be identified. The emotion score prediction module is used to input the voice signal, the environmental data, and the attribute information into the user emotion recognition model to obtain the emotion score of the user to be identified. A target color temperature and brightness determination module is used to determine the target color temperature and target brightness of the lighting lamp based on the emotion score; The lighting adjustment module is used to adjust the color temperature of the lighting lamp to the target color temperature and the brightness of the lighting lamp to the target brightness.
9. A lighting lamp, characterized in that, The lighting fixture includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the lighting control method based on voice prediction of user emotions according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the lighting control method based on voice prediction of user emotions as described in any one of claims 1-7.