Atmosphere adjusting method and device of intelligent terminal, terminal and storage medium
By combining visual and radio frequency identification technologies to acquire physiological signals and facial images, the system can identify the user's emotional state, solving the problem of lack of emotion perception in existing technologies. This enables personalized atmosphere adjustment for smart terminals and improves the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-13
AI Technical Summary
Existing methods for adjusting the atmosphere of smart terminals only consider the user's voice and posture, lacking the perception of the user's emotional state and failing to meet personalized needs.
By combining image-based visual recognition and millimeter-wave radio frequency identification technology, the system can identify the user's emotional state and adjust the scene atmosphere of the smart terminal by acquiring the user's physiological signals and facial images.
It enables the smart terminal's atmosphere to be adjusted according to the user's emotional state, meeting the user's personalized needs and improving the user experience.
Smart Images

Figure CN121665075A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of equipment control technology. More specifically, this application relates to an atmosphere adjustment method, apparatus, terminal, and storage medium for a smart terminal. Background Technology
[0002] Traditional methods for adjusting the atmosphere of smart terminals (such as smart TVs) involve acquiring the voice information of all users within a preset scene, including decibel levels. The preset scene includes a smart terminal, which acts as an atmosphere adjustment device. If at least one user's decibel level is greater than or equal to a preset decibel level, the method acquires the voice information and posture of all users. Based on this information, the operating strategy of the atmosphere adjustment device is determined, and the device is then triggered to operate according to the strategy. This method only considers the user's voice and posture—two simple human-computer interaction methods—and lacks awareness of the user's emotional state. Therefore, existing technology cannot adjust the atmosphere of a smart terminal scene based on the user's emotional state, failing to meet personalized user needs.
[0003] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, terminal, and storage medium for adjusting the atmosphere of a smart terminal, which can adjust the scene atmosphere of the smart terminal according to the user's emotional state, meet the user's personalized needs, and improve the user experience. This application is mainly achieved through the following technical solutions: A first aspect of this application provides an atmosphere adjustment method for a smart terminal, comprising: Acquire the user's physiological signals and facial images, and obtain a target signal quality score based on the physiological signals and a target image quality score based on the facial images; The target signal quality score and the target image quality score are subjected to emotion probability vector recognition processing to obtain the target emotion probability vector. Based on the target emotion probability vector, a target atmosphere strategy is selected from the first preset rule; Adjust the scene atmosphere of the smart terminal based on the target atmosphere strategy.
[0005] According to one embodiment of this application, the physiological signals are chest vibration signals caused by the user's heartbeat and abdominal undulation signals caused by the user's breathing.
[0006] According to one embodiment of this application, the step of obtaining a target signal quality score based on the physiological signal includes: According to the second preset rule, one of the signal-to-noise ratios is selected as the target signal-to-noise ratio from the signal-to-noise ratio of the chest cavity vibration signal and the signal-to-noise ratio of the abdominal undulation signal; The target signal-to-noise ratio is processed by a first preset algorithm to calculate the signal quality score, thereby obtaining the target signal quality score.
[0007] According to one embodiment of this application, the step of obtaining a target image quality score based on the face image includes: The second preset algorithm is used to calculate the brightness uniformity score of the face image to obtain the target brightness uniformity score. The third preset algorithm is used to calculate the face occlusion rate of the face image to obtain the target face occlusion rate. The Laplacian variance algorithm is used to calculate the sharpness score of the face image to obtain the target sharpness score; The fourth preset algorithm is used to calculate the image quality score of the target image by processing the target brightness uniformity score, the target face occlusion rate and the target sharpness score.
[0008] According to one embodiment of this application, the step of performing emotion probability vector recognition processing on the target signal quality score and the target image quality score to obtain the target emotion probability vector includes: Obtain the target facial emotion probability vector and the target physiological emotion probability vector of the user; The fifth preset algorithm is used to perform weight calculation on the quality score of the target signal to obtain the physiological weight; The sixth preset algorithm is used to perform weight calculation on the quality score of the target image to obtain the face weight; The seventh preset algorithm is used to perform emotion probability vector recognition processing on the target facial emotion probability vector, the target physiological emotion probability vector, the physiological weight and the facial weight to obtain the target emotion probability vector.
[0009] According to one embodiment of this application, the steps of obtaining the target facial emotion probability vector and the target physiological emotion probability vector of the user include: A convolutional neural network is used to perform facial emotion probability vector recognition processing on the facial image to obtain the user's target facial emotion probability vector; The physiological signals are processed by feature vector extraction to obtain physiological emotion feature vectors; A deep learning model is used to perform physiological emotion probability vector prediction processing on the physiological emotion feature vector to obtain the target physiological emotion probability vector.
[0010] According to one embodiment of this application, the step of adjusting the scene atmosphere of a smart terminal based on the target atmosphere strategy includes: The display screen and audio of the smart terminal are adjusted based on the target atmosphere strategy.
[0011] A second aspect of this application provides an atmosphere adjustment device for a smart terminal, comprising: A multimodal perception module is used to acquire the user's physiological signals and facial images, and to obtain a target signal quality score based on the physiological signals and a target image quality score based on the facial images. The emotion fusion algorithm module is used to perform emotion probability vector recognition processing on the target signal quality score and the target image quality score to obtain the target emotion probability vector. The emotion content database module is used to select a target atmosphere strategy based on the target emotion probability vector in a first preset rule; The atmosphere rendering module is used to adjust the scene atmosphere of the smart terminal based on the target atmosphere strategy.
[0012] A third aspect of this application provides a smart terminal, including a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the steps of the atmosphere adjustment method of the smart terminal provided in the first aspect of this application.
[0013] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the atmosphere adjustment method for a smart terminal provided in the first aspect of this application.
[0014] The beneficial effects of the embodiments of this application include: This application embodiment acquires the user's physiological signals and facial images, and obtains a target signal quality score based on the physiological signals and a target image quality score based on the facial images. It then performs emotion probability vector recognition processing on the target signal quality scores and the target image quality scores to obtain a target emotion probability vector. Based on the target emotion probability vector, it selects a target atmosphere strategy from a first preset rule. Finally, it adjusts the scene atmosphere of the smart terminal based on the target atmosphere strategy. Compared to existing technologies that only consider the user's voice and posture—two simple human-computer interaction methods—this application embodiment combines multimodal information such as the user's physiological signals and facial images to identify the target emotion probability vector, thereby achieving the purpose of perceiving the user's emotional state. This allows for the adjustment of the smart terminal's scene atmosphere according to the user's emotional state, meeting the user's personalized needs and improving the user experience. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 The flowcharts for the atmosphere adjustment method of the smart terminal of this application are shown in some embodiments; Figure 2 This is a schematic block diagram of the atmosphere adjustment device for the smart terminal of this application in some embodiments; Figure 3 The following is a flowchart of the atmosphere adjustment device for the smart terminal of this application in some other embodiments; Figure 4 This is a block diagram illustrating the principle of the smart terminal in some embodiments of this application. Detailed Implementation
[0017] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0018] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0019] The terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0020] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product, or apparatus.
[0021] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items.
[0022] Traditional methods for adjusting the atmosphere of smart terminals (such as smart TVs) involve acquiring the voice information of all users within a preset scene, including decibel levels. The preset scene includes a smart terminal, which acts as an atmosphere adjustment device. If at least one user's decibel level is greater than or equal to a preset decibel level, the method acquires the voice information and posture of all users. Based on this information, the operating strategy of the atmosphere adjustment device is determined, and the device is then triggered to operate according to the strategy. This method only considers the user's voice and posture—two simple human-computer interaction methods—and lacks awareness of the user's emotional state. Therefore, existing technology cannot adjust the atmosphere of a smart terminal scene based on the user's emotional state, failing to meet personalized user needs.
[0023] To solve the aforementioned technical problems, the inventors made a keen discovery: image-based visual recognition and millimeter-wave-based radio frequency identification can be combined and applied to the atmosphere control of smart terminals.
[0024] Firstly, image-based visual recognition methods directly analyze overt emotional expressions by capturing visual features such as facial expressions, subtle muscle movements, and facial contours. Their advantage lies in their intuitive information and rich features, enabling the recognition of numerous subtle facial expression units. However, this method has the following drawbacks: First, it is highly sensitive to the environment; changes in lighting or occlusion (such as masks or glasses) can severely affect its performance. Second, it is susceptible to subjective manipulation; individuals can consciously control their facial expressions to conceal their true emotions, leading to distorted recognition results.
[0025] Meanwhile, millimeter-wave radio frequency identification (RFID) infers a user's internal emotional state by analyzing microscopic physiological signals (such as heartbeat, breathing rhythm, and skin tremors) sensed by radar wave reflections. Its advantages lie in its strong environmental robustness, being unaffected by lighting conditions and able to penetrate slight obstructions; at the same time, it operates in a non-image-based manner, making it more privacy-friendly.
[0026] In summary, image-based visual recognition and millimeter-wave radio frequency identification (RFID) offer significant complementarity in emotion recognition. The visual modality excels at capturing rich, "external" facial expressions, while the RFID modality is adept at perceiving "internal" physiological responses that are difficult to fake. They effectively compensate for each other's weaknesses. Therefore, deeply fusing these two modalities to identify a user's emotional state and adjust the atmosphere of a smart terminal environment can meet users' personalized needs.
[0027] The specific embodiments of this application will be further described below with reference to the accompanying drawings.
[0028] refer to Figure 1 The diagram shows a flowchart of an atmosphere adjustment method for a smart terminal provided in the first aspect of an embodiment of this application. The atmosphere adjustment method for the smart terminal is applied to a smart terminal, which may be a television. In other embodiments, the smart terminal may be other types of terminals, which can be specifically configured by those skilled in the art according to actual needs. Figure 1 The atmosphere adjustment method of the smart terminal includes: S1. Acquire the user's physiological signals and facial images, and obtain a target signal quality score based on the physiological signals and a target image quality score based on the facial images.
[0029] This application embodiment obtains the physiological signals by scanning the user's body with millimeter-wave radar. The step of obtaining the physiological signals by scanning the user's body with millimeter-wave radar can be implemented using existing technology.
[0030] The millimeter-wave radar operates at a frequency of 60 GHz, with a transmit power less than or equal to 10 dBm (compliant with FCC / CESAR standards), and a range resolution less than or equal to 5 cm. In other embodiments, the operating frequency, transmit power, and range resolution can be set to other values, which can be determined by those skilled in the art according to actual needs.
[0031] The millimeter-wave radar can penetrate the user's clothing to detect subtle movements in the user's chest cavity.
[0032] This application embodiment captures facial images using a camera. The camera is a color camera with a resolution greater than 1080p and a frame rate greater than 30fps, and the camera supports low-brightness enhancement.
[0033] Furthermore, the physiological signals are chest vibration signals caused by the user's heartbeat and abdominal rise and fall signals caused by the user's breathing. The chest vibration signals and the abdominal rise and fall signals are acquired simultaneously.
[0034] The frequency range of chest vibration caused by the user's heartbeat is 0.8–2.0 Hz, and the frequency range of abdominal undulation caused by the user's breathing is 0.1–0.5 Hz.
[0035] Further, the step of obtaining the target signal quality score based on the physiological signal includes: selecting one of the signal-to-noise ratios from the signal-to-noise ratio of the chest cavity vibration signal and the signal-to-noise ratio of the abdominal undulation signal as the target signal-to-noise ratio according to the second preset rule; and performing signal quality score calculation processing on the target signal-to-noise ratio using the first preset algorithm to obtain the target signal quality score.
[0036] The second preset rule selects the smaller signal-to-noise ratio from the signal-to-noise ratio of the chest cavity vibration signal and the signal-to-noise ratio of the abdominal undulation signal. In other embodiments, the second preset rule can be other rules, which can be set by those skilled in the art according to actual needs.
[0037] Furthermore, the calculation formula for selecting one of the signal-to-noise ratios from the chest cavity vibration signal and the abdominal undulation signal as the target signal-to-noise ratio according to the second preset rule is as follows: ; in, It is the target signal-to-noise ratio; It is a function that takes the minimum value; It is the signal-to-noise ratio of the thoracic vibration signal; It is the signal-to-noise ratio of the abdominal undulation signal.
[0038] Furthermore, the first preset algorithm is used to calculate the signal quality score of the target signal-to-noise ratio, and the calculation formula for obtaining the target signal quality score is as follows: ; in, It is the target signal quality score. The target signal-to-noise ratio is measured in decibels. The unit of 15 in the calculation formula is also decibel. In the calculation formula, "other" represents a linear mapping.
[0039] Further, the step of obtaining the target image quality score based on the face image includes: using a second preset algorithm to calculate the brightness uniformity score of the face image to obtain a target brightness uniformity score; using a third preset algorithm to calculate the face occlusion rate of the face image to obtain a target face occlusion rate; using a Laplacian variance algorithm to calculate the sharpness score of the face image to obtain a target sharpness score; and using a fourth preset algorithm to calculate the image quality score of the target brightness uniformity score, the target face occlusion rate, and the target sharpness score to obtain the target image quality score.
[0040] Furthermore, the second preset algorithm is used to calculate the brightness uniformity score of the facial image. The calculation formula for obtaining the target brightness uniformity score is as follows: ; ; ; in, It is the average brightness of the facial image; It is the total number of pixels in the face image; It is the first in the face image grayscale value of each pixel; It is the average brightness deviation of the facial image; It is the target brightness uniformity score; It is a coefficient used to prevent division by zero in the denominator. In some implementations, .
[0041] Furthermore, the third preset algorithm is used to calculate the face occlusion rate of the face image, and the calculation formula for obtaining the target face occlusion rate is as follows: ; in, It is the target face occlusion rate; It is the th out of 468 3D points Confidence level of a 3D point; It is an empirical threshold. .
[0042] It should be understood that the 468 3D points were obtained by recognizing the facial image using a facial landmark detector.
[0043] The facial landmark detector can be BlazeFace. BlazeFace is a lightweight algorithm for facial landmark detection in the MediaPipe framework. MediaPipe is a cross-platform framework used to build machine learning pipelines for processing time-series data such as video and audio.
[0044] Furthermore, the Laplacian variance algorithm is used to calculate the sharpness score of the face image. The step of obtaining the target sharpness score can be implemented by existing technology and will not be described in detail in this application.
[0045] Furthermore, the fourth preset algorithm is used to calculate the image quality score of the target brightness uniformity score, the target face occlusion rate, and the target sharpness score. The calculation formula for obtaining the target image quality score is as follows: ; in, It is the quality score of the target image; It is the weight of the target brightness uniformity score; It is the weight of the target face occlusion rate; It is the weight of the target sharpness score; This is the target sharpness score.
[0046] For example, It is 0.3. It is 0.5. It is 0.2. , , A higher value indicates that the facial image is more reliable.
[0047] S2. Perform emotion probability vector recognition processing on the target signal quality score and the target image quality score to obtain the target emotion probability vector.
[0048] Further, step S2 includes: acquiring the user's target facial emotion probability vector and target physiological emotion probability vector; using a fifth preset algorithm to perform weight calculation processing on the target signal quality score to obtain physiological weight; using a sixth preset algorithm to perform weight calculation processing on the target image quality score to obtain facial weight; and using a seventh preset algorithm to perform emotion probability vector recognition processing on the target facial emotion probability vector, the target physiological emotion probability vector, the physiological weight, and the facial weight to obtain the target emotion probability vector.
[0049] Further, the steps of obtaining the user's target facial emotion probability vector and target physiological emotion probability vector include: using a convolutional neural network to perform facial emotion probability vector recognition processing on the facial image to obtain the user's target facial emotion probability vector; performing feature vector extraction processing on the physiological signal to obtain a physiological emotion feature vector; and using a deep learning model to perform physiological emotion probability vector prediction processing on the physiological emotion feature vector to obtain the target physiological emotion probability vector.
[0050] The convolutional neural network mentioned is either MobileNet-V3 (also known as the third generation of mobile networks) or EfficientNet-B0 (the efficient network B0). MobileNet-V3 is a lightweight convolutional neural network model proposed by Google in 2019, optimized for mobile devices (i.e., smart terminals) and edge computing devices, aiming to improve computational efficiency and model performance. EfficientNet-B0 is the basic model in the EfficientNet series, proposed by Google in 2019, aiming to reduce the number of parameters and computational cost while maintaining high accuracy through network structure optimization.
[0051] The convolutional neural network is a trained neural network capable of recognizing four basic emotions: joy, anger, sorrow, and calmness.
[0052] The convolutional neural network runs on a local chip in the smart terminal.
[0053] For example, the target facial emotion probability vector can be 0.75 for joy, 0.05 for anger, 0.10 for sorrow, and 0.12 for calmness.
[0054] Further, the step of extracting feature vectors from the physiological signals to obtain physiological emotion feature vectors includes: extracting features from the chest vibration signal to obtain heart rate variability; extracting features from the abdominal undulation signal to obtain respiratory rate variability, wherein the heart rate variability and the respiratory rate variability constitute the physiological emotion feature vector.
[0055] The heart rate variability includes the root mean square of successive differences (RMSSD) of the differences between adjacent RR intervals, the standard deviation of normal-to-normal RR intervals (SDNN), low-frequency features, and / or high-frequency features.
[0056] The step of extracting features from the chest cavity vibration signal to obtain heart rate variability features can be achieved using existing technologies.
[0057] The step of extracting features from the abdominal undulation signal to obtain respiratory rate variability can be achieved using existing technologies.
[0058] The deep learning model is a trained LSTM-Attention model, which is a deep learning model that integrates a long short-term memory network and an attention mechanism.
[0059] The training process of the trained LSTM-Attention model can be achieved using existing technologies.
[0060] Furthermore, the fifth preset algorithm is used to perform weight calculation processing on the target signal quality score, and the calculation formula for obtaining the physiological weight is as follows: ; in, It is the physiological weight; It is the base of the natural logarithm; It is the temperature coefficient of Softmax. The larger the value, the smoother the Softmax transition. . The value can be dynamically adjusted.
[0061] Furthermore, the sixth preset algorithm is used to perform weight calculation processing on the quality score of the target image to obtain the facial weight. The calculation formula for this step is as follows: ; in, This refers to the facial weights.
[0062] Furthermore, the seventh preset algorithm is used to perform emotion probability vector recognition processing on the target facial emotion probability vector, the target physiological emotion probability vector, the physiological weight, and the facial weight. The calculation formula for obtaining the target emotion probability vector is as follows: ; in, It is the target emotion probability vector; It is the probability vector of the target facial emotion; It is the probability vector of the target physiological emotion.
[0063] The step of using the seventh preset algorithm to perform emotion probability vector recognition processing on the target facial emotion probability vector, the target physiological emotion probability vector, the physiological weight, and the facial weight to obtain the target emotion probability vector can be understood as using the Softmax weight mapping step.
[0064] It should also be understood that when the lighting is extremely poor (such as when the lights are turned off at night), Approaching 0, If the signal quality of the millimeter-wave radar is poor, the radar weight will be automatically increased to ensure the reliability of identification. Approaching 0, If the visual weight is automatically reduced, the embodiments of this application will automatically increase the visual weight to ensure the robustness of the method. Therefore, the embodiments of this application can solve the problem of single-camera recognition failure in low light, occlusion, and side-facing situations; furthermore, the embodiments of this application can achieve true "seamless interaction" without requiring user-wearable devices.
[0065] S3. Select a target atmosphere strategy from the first preset rules based on the target emotion probability vector.
[0066] The first preset rule contains multiple atmosphere strategies, and one of these atmosphere strategies has a one-to-one correspondence with the target emotion probability vector. Each atmosphere strategy includes a visual strategy, an auditory strategy, and an interaction strategy.
[0067] For example, when the target emotion probability vector is "joy" (representing a happy emotion), the visual strategy in the target atmosphere strategy is that the wallpaper screensaver of the smart terminal is in a dynamic bright style, the ambient light has a rainbow gradient, and the brightness gradually increases from medium to bright; the auditory strategy in the target atmosphere strategy is to play upbeat music at a medium volume; and the interactive strategy in the target atmosphere strategy is for an AI virtual avatar to pop up, smile, wave, and reply "I'm in a great mood today." This example corresponds to a home application scenario.
[0068] For example, when the target emotion probability vector is anger (representing the emotion of being angry), the visual strategy in the target atmosphere strategy is that the wallpaper and screensaver of the smart terminal are static dark red solid color, the picture is depressing and heavy, the ambient light is always red and the brightness is reduced; the auditory strategy in the target atmosphere strategy is noise reduction, dialogue enhancement, and volume reduction, etc.; the interactive strategy in the target atmosphere strategy is that an AI virtual image pops up, makes a soothing head-patting gesture, and replies "Take a deep breath first, I'll help you calm down." This example corresponds to a home application scenario.
[0069] For example, when the target emotion probability vector is sadness (representing grief), the visual strategy in the target atmosphere strategy is that the wallpaper and screensaver of the smart terminal are static, low-saturation, warm-toned, with a warmer color temperature, and the ambient lighting is warm orange and soft with reduced brightness; the auditory strategy in the target atmosphere strategy is to play soothing music or piano music at a medium volume; the interactive strategy in the target atmosphere strategy is for an AI virtual avatar to pop up, smile, wave, and reply, "You seem a little down. Would you like to hear a lighthearted story? I can keep you company." This example corresponds to a home application scenario.
[0070] For example, when the target emotion probability vector is calm, the visual strategy in the target atmosphere strategy is that the wallpaper and screensaver of the smart terminal are set to default or user-defined, the ambient light is a constant warm white, and the brightness is default; the auditory strategy in the target atmosphere strategy is user-defined, and the volume is medium; the interaction strategy in the target atmosphere strategy is that an AI virtual avatar pops up, smiles and waves, and replies "Just relax, call me anytime if you need anything." This example corresponds to a home application scenario.
[0071] In some implementations, the first preset rule follows the "emotion-scene-atmosphere strategy" triple mapping rule.
[0072] S4. Adjust the scene atmosphere of the smart terminal based on the target atmosphere strategy.
[0073] Furthermore, the step of adjusting the scene atmosphere of the smart terminal based on the target atmosphere strategy includes: adjusting the display screen and audio of the smart terminal based on the target atmosphere strategy.
[0074] Through the above implementation methods, this application embodiment combines multimodal information such as the user's physiological signals and facial images to identify the target emotion probability vector, so as to achieve the purpose of perceiving the user's emotional state. This allows the smart terminal to adjust the scene atmosphere according to the user's emotional state, meet the user's personalized needs, and improve the user experience.
[0075] After sensing the user's emotional state, the embodiments of this application can adaptively adjust the scene atmosphere of the smart terminal according to the user's personal preferences and real-time feedback, thereby avoiding the rendering effect of the scene atmosphere being stiff and formulaic, and further meeting the user's personalized needs.
[0076] The embodiments of this application employ two modalities of information: the user's physiological signals and facial images, which can reduce the misidentification rate of single-modal emotions.
[0077] The embodiments of this application can achieve personalized, contextualized, and adaptive optimization of the home entertainment environment.
[0078] When the smart terminal is a smart TV, the smart TV comes pre-installed with rendering scheme packages (i.e., atmosphere strategies) for different emotions (joy, anger, sorrow, and default) in different scenarios. See the following example for details: Example 1: When a user turns on the smart TV in the living room, the smart TV automatically recognizes it as a "home scene." Then, the camera captures the user's depressed expression, and millimeter-wave radar detects that the user is breathing slowly. The "depressed expression" and "slow breathing" are combined to determine a "sad mood." The smart TV then immediately invokes and executes a sad atmosphere strategy for the "home scene," which includes the following: 1) The screen becomes softer, the theme wallpaper and screensaver are changed to static low-saturation warm tones, the color temperature becomes warmer, the backlight is automatically reduced, and an AI avatar appears in the lower right corner of the screen and replies, "You seem a little down. Would you like to hear a lighthearted story? I can keep you company." 2) The ambient light strip of the smart TV breathes gently, displaying a warm, dark orange hue; the volume of the smart TV is reduced, and the background music is changed to soft piano music, etc.
[0079] 3) The smart TV continues to observe the user. If the user sits quietly without operating, it is considered positive feedback. The brightness and volume are maintained and then slightly reduced to guide the user to calm down. If the user gets up or says "change the set", the smart TV switches to another rendering scheme package under the mood of the scene, and records the preference at the same time.
[0080] Example 2: When a user connects a microphone to the smart TV and switches to karaoke mode, the smart TV automatically recognizes it as a "KTV scene." Then, it uses a camera to capture the user's smiling face and millimeter-wave radar to measure a rapid increase in the user's heart rate. The "smiling face" and "rapidly increased heart rate" are combined to determine a "happy and excited mood." Next, the smart TV immediately invokes and executes the happy atmosphere strategy under the "KTV scene," which includes the following: 1) The screen saturation is instantly increased, the theme wallpaper and screensaver are changed to dynamic neon, the brightness is turned up to the maximum, and the AI avatar changes into neon clothes and replies "Everyone wave with me"; 2) The ambient light strip of the smart TV will simultaneously enter a rainbow flashing mode, jumping to the rhythm of the music; the volume of the smart TV will be turned up, and rock music will be recommended; 3) The smart TV continuously observes the user. If the user continues to sing, wave, and jump, it is considered positive feedback. The TV maintains this feedback and slightly increases the speed of the light strip changes to enhance the atmosphere. If the user puts down the microphone or says "slow down," the TV switches to another rendering scheme package for the mood of that scene and records the preference. The new scheme will be recommended first when the same mood is present next time.
[0081] Example 3: When a user turns on the smart TV in a low-light environment in the living room, the smart TV automatically recognizes it as a "home scene"; due to the lights being off, the camera... The heart rate decreases, but the millimeter-wave radar remains unaffected. Based on the fusion method, the proportion of emotion recognition by the camera decreases, and the millimeter-wave radar measures a normal user heart rate. The "camera emotion recognition" and "normal user heart rate" are fused and determined to represent a "calm" emotion. Then, the smart TV immediately invokes and executes the calming strategy for the "home scenario," which includes the following: 1) The wallpaper and screensaver of the smart TV are either default or user-defined, the ambient light is a constant warm white with default brightness, and the AI avatar changes into neon clothes and replies "Relax, call me anytime if you need anything"; 2) The smart TV continuously observes the user, and if the user's emotions change, it calls up the corresponding atmosphere rendering package.
[0082] Based on the above examples, the embodiments of this application are also applicable to low-light scenes or occlusion scenes.
[0083] refer to Figure 2 The diagram shown is a schematic block diagram of an atmosphere adjustment device for a smart terminal provided in the second aspect of an embodiment of this application. Figure 2 In the above, the atmosphere adjustment device 100 of the smart terminal includes a multimodal perception module 101, an emotion fusion algorithm module 102, an emotion content database module 103, and an atmosphere rendering module 104.
[0084] The multimodal perception module 101 is used to acquire the user's physiological signals and facial images (see reference). Figure 3 The steps of "acquiring facial images" and "acquiring physiological signals" are shown, and a target signal quality score is obtained based on the physiological signals, and a target image quality score is obtained based on the facial images.
[0085] The emotion fusion algorithm module 102 is used to perform emotion probability vector recognition processing on the target signal quality score and the target image quality score to obtain a target emotion probability vector (see reference). Figure 3 The "target emotion probability vector" step in the process.
[0086] The emotion content database module 103 is used to select a target atmosphere strategy based on the target emotion probability vector from a first preset rule (see reference). Figure 3 (The "atmosphere strategy" steps).
[0087] The atmosphere rendering module 104 is used to adjust the scene atmosphere of the smart terminal based on the target atmosphere strategy.
[0088] The target atmosphere strategy includes visual, auditory, and interactive strategies, as referenced. Figure 3 As shown, the visual strategy of this application embodiment is used to adjust the wallpaper / screensaver / light strip / color tone, etc. of the smart terminal; the auditory strategy is used to adjust the volume / sound effects / music, etc. of the smart terminal; and the interaction strategy is used to adjust the virtual image rendering of the smart terminal.
[0089] In some implementations, the emotion fusion algorithm module 102 includes a facial expression recognition submodule, a physiological emotion recognition submodule, and a decision-level fusion unit.
[0090] The facial expression recognition submodule is used to perform facial emotion probability vector recognition processing on the facial image using a convolutional neural network to obtain the user's target facial emotion probability vector.
[0091] The physiological emotion recognition submodule is used to obtain the target facial emotion probability vector and the target physiological emotion probability vector of the user.
[0092] The physiological emotion recognition submodule is further used to perform feature vector extraction processing on the physiological signal to obtain a physiological emotion feature vector; and to perform physiological emotion probability vector prediction processing on the physiological emotion feature vector using a deep learning model to obtain the target physiological emotion probability vector.
[0093] The decision-level fusion unit is used to perform weight calculation processing on the target signal quality score using a fifth preset algorithm to obtain physiological weight; to perform weight calculation processing on the target image quality score using a sixth preset algorithm to obtain facial weight; and to perform emotion probability vector recognition processing on the target facial emotion probability vector, the target physiological emotion probability vector, the physiological weight, and the facial weight using a seventh preset algorithm to obtain the target emotion probability vector.
[0094] In some implementations, the emotion content database module 103 includes a user feedback learning submodule, which is used to collect explicit and implicit feedback in real time. (See reference...) Figure 3 The steps of "explicit feedback and implicit feedback" in the process.
[0095] The explicit feedback includes a return command, a like command, a close command, a switch command, a dislike command, and a like command. The return command is fed back by the return button on the remote control of the smart terminal, the like command is fed back by the like button on the remote control, the close command is fed back by the close button on the remote control, and the switch command, the dislike command, and the like command are voice commands output by the user.
[0096] The implicit feedback includes both positive and negative feedback.
[0097] The positive feedback includes the user continuously watching the video played on the smart terminal for more than or equal to 30 seconds, or the user's emotion changing from anger to calm.
[0098] Each rendering scheme package (i.e., atmosphere strategy) has an initial weight of 1.0. If positive feedback is received, the initial weight is increased by 0.1; if negative feedback is received, the initial weight is decreased by 0.1. This process can be referenced. Figure 3 The "Optimize Atmosphere Strategy Weights" step.
[0099] The application of explicit and implicit feedback can break through the traditional "fixed atmosphere" mode and realize personalized atmosphere rendering and continuous updating of atmosphere strategy for smart TVs.
[0100] In some implementations, the atmosphere rendering module 104 includes a visual rendering submodule, an auditory rendering submodule, and an interactive rendering submodule.
[0101] The visual rendering submodule is used to control the screen display of the smart TV and the external ambient light strip. More specifically, it controls the smart TV to switch wallpapers, screen savers and / or UI theme colors, and controls the external ambient light strip via HDMI or Wi-Fi to achieve synchronized color changes.
[0102] The auditory rendering submodule is used to control the stereo speakers to play audio. This submodule can call music APIs from QQ Music, NetEase Cloud Music, etc., to automatically match mood-themed playlists; it also supports the generation of white noise / natural sounds (such as rain, ocean waves, and campfire) for mood-soothing purposes.
[0103] The stereo speakers support Dolby Audio.
[0104] The interactive rendering submodule is used to control the display of a 3D AI virtual human on the smart TV screen and play the audio of the 3D AI virtual human. The appearance of the 3D AI virtual human is optional, and its appearance can change according to emotions. The appearance of the 3D AI virtual human can be a cartoon animal, a realistic style, a minimalist line drawing, or a user-uploaded image.
[0105] The 3D AI virtual human's audio supports multiple languages, such as Chinese (Mandarin / Cantonese), English, and Japanese. The speech rate of the 3D AI virtual human's audio is 200–300 wpm (wpm is words per minute), and this speech rate is adjustable.
[0106] refer to Figure 4 The diagram shown is a schematic block diagram of a smart terminal provided in the third aspect of an embodiment of this application. Figure 4 In this embodiment, the smart terminal 200 includes a processor 201 and a memory 202. The memory 202 is used to store computer programs, and the processor 201 is used to call and run the computer programs stored in the memory 202 to execute the steps of the smart terminal atmosphere adjustment method provided in the first aspect of the present application.
[0107] Those skilled in the art will understand that Figure 4The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the smart terminal to which the present invention is applied. A specific smart terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0108] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the atmosphere adjustment method for a smart terminal provided in the first aspect of this application.
[0109] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0110] The technical features of the above embodiments can be combined without changing the basic principles of this application. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0111] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the patent protection scope of this application should be determined by the appended claims.
Claims
1. An atmosphere adjustment method for a smart terminal, characterized in that, include: Acquire the user's physiological signals and facial images, and obtain a target signal quality score based on the physiological signals and a target image quality score based on the facial images; The target signal quality score and the target image quality score are subjected to emotion probability vector recognition processing to obtain the target emotion probability vector. Based on the target emotion probability vector, a target atmosphere strategy is selected from the first preset rule; Adjust the scene atmosphere of the smart terminal based on the target atmosphere strategy.
2. The atmosphere adjustment method for a smart terminal according to claim 1, characterized in that, The physiological signals are chest vibration signals caused by the user's heartbeat and abdominal undulation signals caused by the user's breathing.
3. The atmosphere adjustment method for a smart terminal according to claim 2, characterized in that, The steps for obtaining the target signal quality score based on the physiological signal include: According to the second preset rule, one of the signal-to-noise ratios is selected as the target signal-to-noise ratio from the signal-to-noise ratio of the chest cavity vibration signal and the signal-to-noise ratio of the abdominal undulation signal; The target signal-to-noise ratio is processed by a first preset algorithm to calculate the signal quality score, thereby obtaining the target signal quality score.
4. The atmosphere adjustment method for a smart terminal according to claim 1, characterized in that, The steps for obtaining the target image quality score based on the face image include: The second preset algorithm is used to calculate the brightness uniformity score of the face image to obtain the target brightness uniformity score. The third preset algorithm is used to calculate the face occlusion rate of the face image to obtain the target face occlusion rate. The Laplacian variance algorithm is used to calculate the sharpness score of the face image to obtain the target sharpness score; The fourth preset algorithm is used to calculate the image quality score of the target image by processing the target brightness uniformity score, the target face occlusion rate and the target sharpness score.
5. The atmosphere adjustment method for a smart terminal according to claim 1, characterized in that, The steps of performing emotion probability vector recognition processing on the target signal quality score and the target image quality score to obtain the target emotion probability vector include: Obtain the target facial emotion probability vector and the target physiological emotion probability vector of the user; The fifth preset algorithm is used to perform weight calculation on the quality score of the target signal to obtain the physiological weight; The sixth preset algorithm is used to perform weight calculation on the quality score of the target image to obtain the face weight; The seventh preset algorithm is used to perform emotion probability vector recognition processing on the target facial emotion probability vector, the target physiological emotion probability vector, the physiological weight and the facial weight to obtain the target emotion probability vector.
6. The atmosphere adjustment method for a smart terminal according to claim 5, characterized in that, The steps for obtaining the target facial emotion probability vector and the target physiological emotion probability vector of the user include: A convolutional neural network is used to perform facial emotion probability vector recognition processing on the facial image to obtain the user's target facial emotion probability vector; The physiological signals are processed by feature vector extraction to obtain physiological emotion feature vectors; A deep learning model is used to perform physiological emotion probability vector prediction processing on the physiological emotion feature vector to obtain the target physiological emotion probability vector.
7. The atmosphere adjustment method for a smart terminal according to claim 1, characterized in that, The steps for adjusting the scene atmosphere of the smart terminal based on the target atmosphere strategy include: The display screen and audio of the smart terminal are adjusted based on the target atmosphere strategy.
8. An atmosphere control device for a smart terminal, characterized in that, include: A multimodal perception module is used to acquire the user's physiological signals and facial images, and to obtain a target signal quality score based on the physiological signals and a target image quality score based on the facial images. The emotion fusion algorithm module is used to perform emotion probability vector recognition processing on the target signal quality score and the target image quality score to obtain the target emotion probability vector. The emotion content database module is used to select a target atmosphere strategy based on the target emotion probability vector in a first preset rule; The atmosphere rendering module is used to adjust the scene atmosphere of the smart terminal based on the target atmosphere strategy.
9. A smart terminal, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to perform the steps of the atmosphere adjustment method of the smart terminal according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the steps of the atmosphere adjustment method of the smart terminal according to any one of claims 1 to 7.