Data processing method, data processing device and electronic equipment
By analyzing audio-visual data features and real-time emotion recognition, and dynamically matching lighting effects and color schemes, the problem of monotonous lighting effects in existing technologies is solved, achieving synchronization of lighting with audio and video and emotion matching, thereby improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies cannot match different types of data when processing audio or video data, resulting in limited lighting effect modes and affecting user experience.
By analyzing the audio and video features of the audio-visual data and combining them with the emotional attributes of the target audience, the system dynamically matches the target lighting effects and color scheme to achieve synchronization between lighting and audio-visual content and emotional resonance.
It enhances the user experience by using real-time emotion recognition and audio-visual feature fusion to ensure that lighting is synchronized with audio and video, and that colors match the user's emotions, thus achieving personalized interaction.
Smart Images

Figure CN121815011A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data processing method, a data processing apparatus, and an electronic device. Background Technology
[0002] Currently, electronic devices typically analyze the characteristics of real-time audio or video signals during processing, mapping these characteristics to dynamic changes in lighting. This causes the electronic device's lighting to synchronize with the style of the audio or video. However, mapping the parsed audio or video signals to fixed display rules during processing cannot match different types of audio or video data. The resulting lighting effects are limited to a single mode, negatively impacting the user experience. Summary of the Invention
[0003] This application provides a data processing method and an electronic device.
[0004] On one hand, embodiments of this application provide a data processing method, including: Based on the first audio-visual data, a target lighting effect mode matching the first audio-visual data is determined; the first audio-visual data includes the audio features and / or video features of the currently playing content. Based on the target emotion attribute associated with the target object, a target color scheme mode that matches the target emotion attribute is determined; the target emotion attribute is obtained based on the image data and / or audio data corresponding to the target object. The target lighting effect mode and the target color scheme mode are merged to obtain the corresponding display information.
[0005] Optionally, the method further includes: Perform facial expression recognition on the image data corresponding to the target object to obtain a first emotion probability distribution; Perform voice emotion analysis on the audio data corresponding to the target object to obtain a second emotion probability distribution; Based on the audio features and / or video features corresponding to the first audio-visual data, a third emotion probability distribution is obtained; The target emotion attribute is determined by weighted fusion of at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution. Wherein, the weight of the first emotion probability distribution is greater than the weight of the second emotion probability distribution, and the weight of the second emotion probability distribution is greater than the weight of the third emotion probability distribution.
[0006] Optionally, the method further includes: The first emotion probability distribution and the second emotion probability distribution are weighted and fused to obtain the first emotion attribute; The probability distribution of the third emotion is weighted and fused to obtain the second emotion attribute; If the emotion category represented by the first emotion attribute conflicts with the emotion category represented by the second emotion attribute, the first emotion attribute is determined as the target emotion attribute.
[0007] Optionally, the step of weightedly fusing at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution to determine the target emotion attribute includes: Determine the first probability value, second probability value, and third probability value of the preset emotion category in the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution; Based on the weights corresponding to at least two of the first probability value, the second probability value, and the third probability value, the target probability value of the emotion category is obtained. The emotion category corresponding to the maximum value among the target probability values is determined as the target emotion attribute.
[0008] Optionally, determining the target lighting effect mode that matches the first audio-visual data based on the first audio-visual data includes: The audio data contained in the first audio-visual data is parsed to determine the audio features corresponding to the audio data contained in the first audio-visual data; the audio features include at least one of the following: volume, frequency distribution, rhythm complexity, and audio dynamic range; The video data contained in the first audio-visual data is parsed, and the corresponding video features are extracted. The video features include at least one of the following: brightness, hue, scene switching frequency, and scene category. Based on the audio features corresponding to the audio data contained in the first audio-visual data, and / or the video features corresponding to the video data contained in the first audio-visual data, the corresponding target lighting effect mode is determined from the preset lighting effect mode library.
[0009] Optionally, determining the target color scheme pattern that matches the target emotional attribute based on the target object's associated target emotional attribute includes: Based on a preset mapping relationship between emotion categories and color schemes, the color scheme corresponding to the emotion category represented by the target emotion attribute is determined as the target color scheme; wherein, the color scheme includes multiple preset color values with different visual functions.
[0010] Optionally, the operations of facial expression recognition on the image data corresponding to the target object and voice emotion analysis on the audio data corresponding to the target object can be executed synchronously and in parallel on different processor cores or threads.
[0011] Optionally, the method further includes: Based on the audio dynamic range of the audio data included in the first audio-visual data, the sampling rate for extracting the audio features corresponding to the first audio-visual data is adjusted; wherein, the audio dynamic range includes the volume range of the audio data included in the first audio-visual data.
[0012] On the other hand, embodiments of this application also provide a data processing apparatus, including: The acquisition module is used to determine a target lighting effect mode that matches the first audio-visual data based on the first audio-visual data; the first audio-visual data includes the audio features and / or video features of the currently playing content; The determination module is used to determine a target color scheme that matches the target emotional attribute based on the target emotional attribute associated with the target object; the target emotional attribute is obtained based on the image data and / or audio data corresponding to the target object. The processing module is used to merge the target lighting effect mode and the target color scheme mode to obtain the corresponding display information.
[0013] This application also provides another electronic device, including: Cameras are used to capture emotional information about the target subject; A microphone is used to acquire the voice information of the target object; A display for displaying target display information, wherein the target display information is obtained by fusing at least the emotional information of the target object, the voice information of the target object, and the audio and / or video features of the currently playing content included in the first audio-visual data; Memory, used to store executable programs; A processor is configured to execute the executable program to perform the following steps: Based on the first audio-visual data, a target lighting effect mode matching the first audio-visual data is determined; the first audio-visual data includes the audio features and / or video features of the currently playing content. Based on the target emotion attribute associated with the target object, a target color scheme mode that matches the target emotion attribute is determined; the target emotion attribute is obtained based on the image data and / or audio data corresponding to the target object. The target lighting effect mode and the target color scheme mode are merged to obtain the corresponding display information. Attached Figure Description
[0014] Figure 1 This is a flowchart of a data processing method according to an embodiment of this application; Figure 2 This is another flowchart of the data processing method according to an embodiment of this application; Figure 3 This is another flowchart of the data processing method according to an embodiment of this application; Figure 4 Examples of embodiments of this application Figure 2 A flowchart of one embodiment of step S700; Figure 5 Examples of embodiments of this application Figure 1 A flowchart of one embodiment of step S100; Figure 6 This is a schematic diagram of the structure of the data processing apparatus according to an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0015] Various embodiments and features of this application are described herein with reference to the accompanying drawings.
[0016] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.
[0017] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0018] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.
[0019] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application.
[0020] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.
[0021] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this application, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.
[0022] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.
[0023] Figure 1 A flowchart illustrating a data processing method according to an embodiment of this application is shown. An embodiment of this application provides a data processing method, such as... Figure 1 As shown, it includes: S100, based on the first audio-visual data, determine the target lighting effect mode that matches the first audio-visual data; the first audio-visual data includes the audio features and / or video features of the currently playing content; The data processing method of this embodiment is applied to an electronic device, which can be a smart terminal with a camera, microphone, screen, and model. When the electronic device plays audio and video content, the audio and video content being played can be obtained, and the corresponding first audio and video data can be obtained. The first audio and video data can be the multimedia content source currently being played by the electronic device, specifically including the audio features and / or video features of the currently played content. If the first audio and video data contains audio content, the audio data can be parsed to extract audio features, such as rhythm features (e.g., beat, rhythm intensity), spectral features (e.g., low-frequency, mid-frequency, and high-frequency energy distribution), dynamic features (e.g., volume transients, loudness range), and timbre features (e.g., the spectral feature vector of the sound wave). If the first audio and video data contains video content, the video data can be parsed to extract video features; for example, by analyzing the video frame sequence, visual rhythm (e.g., shot switching rate, image motion vector intensity), dominant color tone, scene brightness, and image content classification information can be extracted. If the first audio and video data contains both audio and video content, the aforementioned audio features and video features are extracted.
[0024] Audio and / or video features extracted from the first audio-visual data can be used for lighting effect control. Based on the extracted audio and / or video features, a matching target lighting effect mode can be selected from a predefined lighting effect mode library or dynamically generated. The extracted audio and / or video features are input into a preset mapping model of audio-visual features and lighting effects, or a model trained through machine learning. This model can output one or more matching target lighting effect modes. A lighting effect mode can be a set of rules describing the dynamic behavior of lighting, defining the rhythm, form, and dynamic parameters of lighting changes. For example, a strong rhythm mode where lighting changes follow the detected beat, mainly using instantaneous flashing or rapid switching; a soft gradient mode where lighting changes are smoothly associated with melodic lines or visual movement, using slow brightness fade-in / fade-out or color flow effects; and a scene enhancement mode that triggers preset lighting animations based on the video scene type.
[0025] S200, based on the target emotion attribute associated with the target object, determine a target color scheme mode that matches the target emotion attribute; the target emotion attribute is obtained based on the image data and / or audio data corresponding to the target object; In this embodiment, when determining the target lighting effect mode that matches the first audio-visual data, the target object in the environment can also be perceived through its sensors; for example, the target object can be a user watching a video or listening to audio. Specifically, the camera of the electronic device can be used to acquire an image or video stream containing the face of the target object. Using face detection and expression recognition algorithms, facial expression features are analyzed to obtain the image emotion probability distribution; for example, using a lightweight model of a convolutional neural network to analyze the user's facial expression features, the image emotion probability distribution is {"pleasant": 0.7, "calm": 0.2, "surprised": 0.1}. Audio data of the target object is collected through a microphone, and voice emotion recognition analysis is performed on the audio data to obtain the voice emotion probability distribution; for example, analyzing the user's acoustic features, such as pitch, speech rate, energy, and spectral features, the voice emotion probability distribution is {"excited": 0.6, "neutral": 0.3, "tired": 0.1}.
[0026] The target emotion attribute can be obtained based on the analysis results of one or more types of data, including image data and audio data corresponding to the target object. If only the image data corresponding to the target object is analyzed, the image emotion probability distribution can be determined as the target emotion attribute; if only the audio data corresponding to the target object is analyzed, the speech emotion probability distribution can be determined as the target emotion attribute. If the target emotion attribute is obtained based on both image data and audio data corresponding to the target object, the image emotion probability distribution and the speech emotion probability distribution can be weighted and fused, and the emotion category with the highest weighted probability can be determined as the target emotion attribute.
[0027] Based on the defined target emotional attribute, the corresponding target color scheme can be searched and determined in a pre-stored emotion-color mapping database. The color scheme can be the color tone, saturation tendency, and brightness range of the lighting. For example, a joyful / excited emotion can be mapped to warm colors with high saturation and medium-high brightness (such as orange or bright yellow) or a contrasting combination of multiple colors; a calm / relaxed emotion can be mapped to cool colors with medium-low saturation and medium-low brightness (such as light blue or pale cyan) or a soft monochrome gradient; a melancholic / sad emotion can be mapped to cool colors with low saturation and low brightness or neutral colors (such as dark blue or gray).
[0028] S300, the target lighting effect mode and the target color scheme mode are merged to obtain the corresponding display information.
[0029] In this embodiment, after determining the target lighting effect mode that matches the first audio-visual data and the target color scheme mode that matches the target emotional attribute, the target lighting effect mode and the target color scheme mode are merged to obtain the corresponding display information. Specifically, while maintaining the rhythm, variation form, and dynamic parameters defined by the target lighting effect mode, the color values used in the light changes are replaced or adjusted to conform to the color tone defined by the target color scheme mode. The display information can be control instructions parsed by the light driver, including parameters such as color value, brightness value, change time point, and transition curve. The light driver module controls devices such as RGB LED light strips and smart lights to display based on the display information. For example, if the target lighting effect mode is a strong rhythm mode, it is characterized by an instantaneous switch between two colors every 0.5 seconds; the target color scheme mode is a cheerful color scheme, with orange as the main color and bright yellow as the auxiliary color; the fused effect is that the light switches instantaneously between orange and bright yellow at a frequency of once every 0.5 seconds. The data processing method of this application analyzes the audio and video features of audio-visual content, performs style classification using a lightweight model, and intelligently matches the corresponding target lighting effect mode. It integrates multimodal emotion recognition of real-time user facial expressions and voice, performs weighted fusion based on the principle of prioritizing real-time physiological reactions of users, and determines the target emotional attributes and corresponding target color scheme modes that reflect the true mood. The target lighting effect mode and target color scheme mode are then fused, so that the lighting is synchronized with the audio-visual content in terms of rhythm and resonates with the user in terms of color and emotion, thereby improving the user experience. Through optimization strategies such as parallel computing and adaptive dynamic sampling rate, the method ensures real-time fast response and achieves efficient management of computing resources, realizing personalized linkage between lighting, audio-visual style, and emotion.
[0030] In one embodiment of this application, such as Figure 2 As shown, the method further includes: S400, Perform facial expression recognition on the image data corresponding to the target object to obtain a first emotion probability distribution; In this embodiment, after capturing an image of the target object's face using a camera, facial expression recognition is performed on the image data corresponding to the target object to obtain a first emotion probability distribution. Specifically, a face detection algorithm can be used to locate and crop the facial region of the target object to obtain a preprocessed facial image; a lightweight convolutional neural network model set in the electronic device is used to analyze the preprocessed facial image to obtain the first emotion probability distribution. For example, as shown in Table 1, a camera is used to capture a user's facial image; a lightweight expression recognition model is used to analyze facial key points, muscle movements, and other features to output a first emotion probability distribution that conforms to Ekman's six basic emotion models (happiness, sadness, anger, surprise, fear, disgust) or a more simplified emotional state (excitement, peace, melancholy, joy). For example, {happiness: 0.9, pleasure: 0.1}.
[0031]
[0032] Table 1: Emotion extraction using facial expression recognition technology, output probability based on Ekman's six basic emotion models. S500, Perform voice emotion analysis on the audio data corresponding to the target object to obtain a second emotion probability distribution; In this embodiment, after acquiring the audio data of the target object through a microphone, voice emotion analysis is performed on the audio data of the target object to obtain a second emotion probability distribution. Specifically, the user's voice signal is collected through a microphone, and the tone, speech rate, volume, and spectral characteristics of the voice signal are analyzed. A lightweight voice emotion classifier is used to extract the emotion features represented by the voice signal to obtain the second emotion probability distribution.
[0033] S600, based on the audio features and / or video features corresponding to the first audio-visual data, obtain the third emotion probability distribution; In this embodiment, audio features and / or video features corresponding to the first audio-visual data can be obtained based on the multimedia content source currently being played by the electronic device, and a third emotion probability distribution can be determined. Specifically, the audio data contained in the first audio-visual data is parsed to extract audio features; for example, rhythm features (such as beat, rhythm intensity), spectral features (such as low-frequency, mid-frequency, and high-frequency energy distribution), dynamic features (such as volume transients, loudness range), and timbre features (such as the spectral feature vector of the sound wave) can be extracted; the audio features are then parsed using an audio emotion recognition model to obtain the emotion probability corresponding to the audio features. The video data contained in the first audio-visual data is parsed to extract video features; for example, by analyzing the video frame sequence, visual rhythm (such as shot switching rate, image motion vector intensity), dominant color tone, scene brightness, and image content classification information can be extracted; the video features are then parsed using a lightweight image classification model to obtain the emotion probability corresponding to the video features.
[0034] If the first audio-visual data contains only audio features, the emotion probability corresponding to the audio features is determined as the third emotion probability distribution; if the first audio-visual data contains only video features, the emotion probability corresponding to the video features is determined as the third emotion probability distribution; if the first audio-visual data contains both audio and video features, the emotion probabilities corresponding to the audio features and the emotion probabilities corresponding to the video features are weighted and fused to obtain the third emotion probability distribution.
[0035] For example, as shown in Table 2, by analyzing the typical emotional associations of music genres, the corresponding probability distributions are obtained: {excitement: 0.6, anger: 0.3, vitality: 0.1}.
[0036]
[0037] Table 2: Probability distribution of output emotion based on typical emotion associations of audio type S700, at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution are weighted and fused to determine the target emotion attribute; wherein, the weight of the first emotion probability distribution is greater than the weight of the second emotion probability distribution, and the weight of the second emotion probability distribution is greater than the weight of the third emotion probability distribution.
[0038] In this embodiment, after obtaining the first emotion probability distribution corresponding to the image data of the target object, the second emotion probability distribution corresponding to the audio data of the target object, and the third emotion probability distribution corresponding to the audio features and / or video features of the first audio-visual data, at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution are weighted and fused to determine the target emotion attribute. Since facial expressions and voices are real-time physiological responses of users, they reflect the user's current mood more directly than the type of audio or video played (which may be background or preference). Therefore, the first emotion probability distribution corresponding to the image data of the target object is assigned the highest weight; the second emotion probability distribution corresponding to the audio data of the target object is assigned a lower weight than the first emotion probability distribution but a higher weight than the third emotion probability distribution; the third emotion probability distribution corresponding to the audio features and / or video features of the first audio-visual data is assigned a lower weight than the second emotion probability distribution. The first, second, and third emotion probability distributions are then superimposed according to their weights, and the distribution with the highest probability is taken as the target emotion attribute.
[0039] Specifically, the weight of the first emotion probability distribution can be set to 40%, the weight of the second emotion probability distribution to 35%, and the weight of the third emotion probability distribution to 25%. For example, when a user plays music using an electronic device, the first emotion probability distribution corresponding to the user's facial expression is happiness (0.9), which, after weighting, is 0.9 × 40% = 0.36; the second emotion probability distribution corresponding to the user's voice is excitement (0.7), which, after weighting, is 0.7 × 35% = 0.245; the third emotion probability distribution corresponding to the music being played is excitement (0.6), which, after weighting, is 0.6 × 25% = 0.15; the highest probability distribution after weighting is the first emotion probability distribution, which is happiness (0.36).
[0040] In one embodiment of this application, such as Figure 3 As shown, the method further includes: S800, the first emotion probability distribution and the second emotion probability distribution are weighted and fused to obtain the first emotion attribute; In this embodiment, a weighted fusion of the first and second emotion probability distributions is performed to obtain the first emotion attribute. Specifically, the first and second emotion probability distributions represent the two most direct real-time physiological signal sources for the user: facial expressions and speech. The weighted fusion of the first and second emotion probability distributions reflects the assessment of the immediacy and directness of the information source; typically, facial expression weight (e.g., 40%) is set higher than speech weight (e.g., 35%), because subtle changes in facial muscles are considered the most direct and difficult-to-disguise expression of emotion. The probability value of each emotion category in the first and second emotion probability distributions is multiplied by its corresponding weight and then summed to obtain a weighted composite probability distribution. The emotion category with the highest probability value can be determined as the first emotion attribute. For example, during a family movie viewing, the camera detected a user's upturned lips and contracted eye muscles, outputting an expression probability distribution of {happiness: 0.90, calmness: 0.10}. Simultaneously, the microphone captured the user's light humming, with a calm tone and stable rhythm, resulting in a speech analysis probability distribution of {pleasant: 0.80, neutral: 0.20}. After mapping pleasure to happiness and unifying the label, the weighted value was calculated: Happiness probability = 0.90 × 40% + 0.80 × 35% = 0.36 + 0.28 = 0.64. This value is higher than other emotions, so the primary emotional attribute can be determined as happiness.
[0041] S900, weighted fusion of the third emotion probability distribution to obtain the second emotion attribute; In this embodiment, the third emotion probability distribution is derived from feature analysis of the currently playing audio and / or video content. A weighted fusion of the third emotion probability distribution yields the second emotion attribute. Specifically, the emotion category with the highest probability value in the third emotion probability distribution can be directly selected as the second emotion attribute. The second emotion attribute represents the dominant emotional atmosphere that the audio-visual content attempts to create or is typically associated with. For example, suppose a clip from a classic tragedy film is being played, with somber and soothing classical music and a dark color palette. Through analysis of the audio and video content, the resulting third emotion probability distribution could be {sadness: 0.75, calm: 0.20, neutral: 0.05}; the second emotion attribute can be determined as sadness.
[0042] S1000, if the emotion category represented by the first emotion attribute conflicts with the emotion category represented by the second emotion attribute, the first emotion attribute is determined as the target emotion attribute.
[0043] In this embodiment, a first emotional attribute is compared with a second emotional attribute. If the emotional categories represented by the first and second emotional attributes are consistent or do not constitute a fundamental conflict, the attribute can be directly adopted or smoothed out. When a significant conflict is detected between the first and second emotional attributes, the first emotional attribute is preferentially identified as the target emotional attribute. This rule is designed based on user experience considerations. The user's facial expressions and voice are real-time physiological signal outputs, which reflect the user's current emotional experience better than audio-visual content as external environmental stimuli.
[0044] For example, the currently playing music might be a sad folk song, with a melancholic emotional tone. However, the user might be chatting with a friend, their facial expression identified as happy, and their tone of voice as cheerful. If the playing music's emotional tone (sadness) conflicts with the user's real-time emotional tone (happiness), the music's guidance will be ignored, and the result of the fusion of the user's facial expression and voice will take precedence. Specifically, the primary emotional attribute (happiness derived from the fusion of facial expression and voice) will be directly used as the final target emotional attribute. This ensures that subsequent lighting color schemes will match the warm tones associated with a happy mood, rather than following the cool tones of sad music, thus ensuring that the lighting environment matches the user's true emotional state.
[0045] In one embodiment of this application, such as Figure 4 As shown, the step of weightedly fusing at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution to determine the target emotion attribute includes: S710, determine the first probability value, second probability value and third probability value of the preset emotion category in the first emotion probability distribution, the second emotion probability distribution and the third emotion probability distribution; S720, based on the weights corresponding to at least two of the first probability value, the second probability value, and the third probability value, the target probability value of the emotion category is obtained; S730, the emotion category corresponding to the maximum value in the target probability values is determined as the target emotion attribute.
[0046] In this embodiment, firstly, raw probability values are extracted from the first, second, and third emotion probability distributions for the preset emotion categories. Then, according to predefined weighting coefficients that reflect the importance of different data sources, a weighted comprehensive score, i.e., the target probability value, is calculated for each emotion category. Finally, by comparing these target probability values, the emotion category corresponding to the maximum value among the target probability values is determined as the target emotion attribute.
[0047] Specifically, a predefined set of emotion categories for decision-making is used, such as {happiness, excitement, calmness, sadness, neutral}; this set covers the main emotional dimensions required for the target lighting color scheme's response. A first emotion probability distribution (facial expression), a second emotion probability distribution (voice), and a third emotion probability distribution (audio-visual content) are received and mapped to the emotion category set. For each emotion category in the set, its corresponding probability value is retrieved from the three mapped probability distributions; a first probability value is obtained from the first emotion probability distribution, a second probability value from the second, and a third probability value from the third. The first probability value (facial expression) has a weight of 40%, the highest, reflecting the high reliability of facial expressions as the most direct physiological signal. The second probability value (voice) has a weight of 35%, the second highest, reflecting that voice emotions are also a user's colloquial and immediate response. The third probability value (audio-visual content) has a weight of 25%, the lowest, reflecting that the emotions of external content may not be a direct reflection of the user's current mood. Based on the weights corresponding to at least two of the first, second, and third probability values, the target probability value of the emotion category can be obtained. These target probability values are compared, and the emotion category corresponding to the target probability value with the largest value is determined as the final target emotion attribute.
[0048] For example, a user is watching a fast-paced sci-fi action movie (audio-visual content), but has just completed a relaxing activity, and their actual mood is relatively pleasant and relaxed. The user is smiling, and the facial recognition model outputs the following first emotion probability distribution: {Happy: 0.85, Calm: 0.1, Neutral: 0.05}. When the user chats with someone else in a relaxed tone, the speech analysis model outputs the following second emotion probability distribution: {Happy: 0.70, Calm: 0.25, Neutral: 0.05}. The movie's fast-paced editing, explosive sound effects, and rousing background music, determined by the content analysis model, output the following third emotion probability distribution: {Excited: 0.80, Nervous: 0.15, Neutral: 0.05}. The target probability value for the emotion of happiness is 0.85×40% + 0.70×35% + 0.00×25% = 0.34 + 0.245 + 0 = 0.585; the target probability value for the emotion of excitement is 0.00×40% + 0.00×35% + 0.95×25% = 0 + 0 + 0.2375 = 0.2375; the target probability values for other emotions, such as calmness, are much lower than 0.585. Comparing the target probability values for all emotion categories, happiness has the highest value at 0.585; therefore, the final determined target emotion attribute is happiness.
[0049] In one embodiment of this application, such as Figure 5 As shown, determining the target lighting effect mode that matches the first audio-visual data based on the first audio-visual data includes: S110, parse the audio data contained in the first audio-visual data to determine the audio features corresponding to the audio data contained in the first audio-visual data; the audio features include at least one of the following: volume, frequency distribution, rhythm complexity, and audio dynamic range; In this embodiment, when the first audio-visual data includes audio data, the audio data contained in the first audio-visual data is parsed to determine the audio features corresponding to the audio data contained in the first audio-visual data. Specifically, the short-term root mean square energy of the audio data can be calculated as a real-time volume feature. This feature directly reflects the change in sound intensity and is the basis for triggering changes in light brightness. A fast Fourier transform is performed on the audio data to divide the spectrum into three frequency ranges: low frequency (20-250Hz), mid frequency (250-4kHz), and high frequency (4kHz-20kHz), and the energy proportion of each frequency range is determined to obtain the frequency distribution characteristics. For example, strong drum beats and bass will show a high proportion of low-frequency energy, while vocals and melodies are mostly concentrated in the mid-frequency range. The beat frequency (BPM) of the data is detected to analyze the regularity of the rhythm; the stability of the beat is measured by calculating the variance of the interval between consecutive beats. A small variance indicates a stable rhythm (such as electronic dance music), while a large variance indicates a complex and varied rhythm (such as free jazz or the cadenza of classical music). The difference between the peak volume and the average volume (or valley volume) of audio data is calculated as a dynamic range feature. Audio with a large dynamic range (such as classical symphony) has a range from subtle to grand, while audio with a small dynamic range and significant compression (such as some pop music) has a smooth change in loudness.
[0050] S120, the video data contained in the first audio-visual data is parsed and the corresponding video features are extracted. The video features include at least one of the following: brightness, hue, screen switching frequency and scene category. In this embodiment, when the first audio-visual data includes video data, the video data contained in the first audio-visual data is parsed to determine the video features corresponding to the video data contained in the first audio-visual data. Specifically, the average brightness value of the video frames contained in the video data can be calculated to obtain the brightness feature. This feature can be used to map the basic brightness level of the lights; dark scenes correspond to low-brightness ambient light, while bright scenes can appropriately increase the ambient light. The dominant color tone of the current frame in the video (such as warm orange-red and cool blue-cyan) is determined through a color histogram or dominant color extraction algorithm. This feature can provide a direct visual reference for the color selection of the lights. The timing of shot transitions in the video is detected, and the number of transitions per unit time is calculated, i.e., the frame transition frequency. This feature directly reflects the editing rhythm of the video; high-speed transitions can be mapped to rapid changes in lighting. A lightweight image classification neural network is used to analyze the video frames in real time, identify scene categories, and provide a key basis for selecting theme lighting effects.
[0051] S130, based on the audio features corresponding to the audio data contained in the first audio-visual data, and / or the video features corresponding to the video data contained in the first audio-visual data, determine the corresponding target lighting effect mode from the preset lighting effect mode library.
[0052] In this embodiment, a target lighting effect mode is determined from a preset lighting effect mode library based on the audio features corresponding to the audio data contained in the first audio-visual data and / or the video features corresponding to the video data. Specifically, the audio features and / or video features can be matched with the feature conditions of each mode in the lighting effect mode library, and the lighting effect mode with the highest matching degree is determined as the target lighting effect mode. If the first audio-visual data only contains audio data, the audio features can be matched with the feature conditions of each mode in the lighting effect mode library. For example, when analyzing an electronic dance music piece, the audio features are: stable rhythm, high proportion of low-frequency energy, high transient density, and medium dynamic range. A precise beat-synchronized flashing mode is matched in the lighting effect mode library. If the first audio-visual data only contains video data, the video features can be matched with the feature conditions of each mode in the lighting effect mode library. For example, when playing a fast-cut action movie, the video features are: extremely high scene switching frequency, intense fighting scene, cool color scheme, and strong brightness contrast. A fast-cut light and shadow scanning mode is matched in the lighting effect mode library. If the first audio-visual data includes both audio and video data, the audio and video features can be matched with the feature conditions of each mode in the lighting effect mode library. For example, as shown in Table 3, when playing a concert video, based on the audio features (exhilarating orchestral music, stable rhythm) and video features (brilliant stage lights, concert scene), the performance is characterized by a strong rhythm, prominent accents, and large volume fluctuations; the BMP is stable between 80 and 140 BPM; the low-frequency drum beats are dense; and there is a special frequency of electric guitar (3-6kHz high energy). The stage atmosphere enhancement mode can be matched in the lighting effect mode library. This mode follows the rhythm and intensity of the music, and in terms of color, it displays the brilliance of the stage lights. It has burst pulses (instant brightening during accents, followed by decay), and the dynamic range is characterized by a large brightness variation (rapid switching from 0-100%).
[0053]
[0054] Table 3: Music Genre Feature Extraction and Rhythm Pattern Mapping Methods In one embodiment of this application, determining a target color scheme that matches the target emotional attribute based on the target emotional attribute associated with the target object includes: Based on a preset mapping relationship between emotion categories and color schemes, the color scheme corresponding to the emotion category represented by the target emotion attribute is determined as the target color scheme; wherein, the color scheme includes multiple preset color values with different visual functions.
[0055] In this embodiment, an emotion-color mapping relationship database can be pre-set. Each emotion category in this database corresponds not to a color, but to a structured color scheme object. The color scheme object can contain an array or list of multiple preset color values (usually stored in HEX or HSV format), and each color value is explicitly assigned a specific visual function label, such as "primary color," "secondary color," "transition color," "embellishment color," or "balance color." Based on the determined target emotion attribute, and based on the preset mapping relationship between emotion categories and color schemes, the emotion-color mapping relationship database is queried, and the color scheme corresponding to the emotion category represented by the target emotion attribute is determined as the target color scheme.
[0056] For example, as shown in Table 4, a user is watching an animated film. The camera captures the user's continuous smiling face (facial expression recognition), and the microphone collects cheerful laughter and exclamations (voice emotion analysis). Through multimodal fusion, the target emotion attribute is determined to be "happiness." Using "happiness" as an index, the emotion-color mapping database is queried to obtain a structured color scheme pattern containing five colors: bright yellow, warm orange, light yellow, light pink, and off-white, along with their functional definitions. This pattern is then identified as the current target color scheme pattern. Based on the animated film's lively background music and bright visuals, the target lighting effect pattern is determined to be a cartoon rhythmic bouncing mode. Its characteristics include synchronized light changes with sound effects, elasticity, brief on / off cycles, and color jumps. The target color scheme mode and target lighting effect mode are merged to generate specific display information: the light uses a low-brightness off-white (#FFF8E1) as a constant background to create comfortable ambient light; when there is a light background music beat, the main body of the light fluctuates elastically between bright yellow (#FFD700) and light yellow (#FFF380) to simulate a cheerful pulse; when exaggerated sound effects or comical actions appear in the film, the light instantly switches to warm orange (#FFA500) and flashes brightly to enhance the comedic effect; when the plot reaches a joyful climax and the theme song plays, the light flashes rapidly while interspersing light pink (#FFD1DC) to push the emotion to its peak.
[0057]
[0058] Table 4: Based on color psychology, different color settings are matched to different user moods. This embodiment defines a structured color scheme with multiple colors, allowing colors for different functions to be flexibly adapted to different dynamic stages of the target lighting effect mode. This results in a final light show that has both a unified emotional tone and rich details and variations, enhancing the user experience.
[0059] In one embodiment of this application, the operations of facial expression recognition on the image data corresponding to the target object and voice emotion analysis on the audio data corresponding to the target object are executed synchronously and in parallel on different processor cores or threads.
[0060] In this embodiment, the two computationally intensive operations of facial expression recognition of target object image data and voice emotion analysis of audio data are deployed on different processor cores or independent execution threads of electronic devices for synchronous parallel execution, thereby reducing the overall system processing latency.
[0061] The electronic device in this embodiment can be a smart terminal with a multi-core CPU or heterogeneous computing unit (such as CPU+NPU). The device has a built-in high-definition camera and microphone that start simultaneously to capture video streams containing the user's face and ambient audio streams, respectively. After receiving the raw image data and audio data packets, a scheduling thread running on one CPU core can perform task assignment; send a frame of facial image data packet to the expression recognition processing thread; and send the corresponding audio data packet (e.g., a 500ms audio segment aligned with the image frame time) to the speech emotion analysis processing thread. The expression recognition thread is typically scheduled to execute on the device's neural processing unit (NPU) or a dedicated CPU core. This thread can perform the following operations: call a face detection model to locate and crop the facial region; perform normalization preprocessing (scaling, normalization) on the cropped image; input the preprocessed image into a preloaded lightweight convolutional neural network expression recognition model; and output a first emotion probability distribution for that frame. The speech emotion analysis thread is typically scheduled to execute on another independent CPU core or digital signal processor (DSP). This thread performs the following operations: preprocesses the audio data, including noise reduction, pre-emphasis, and frame segmentation; performs speech activity detection to separate valid user speech segments; extracts acoustic features from the valid speech segments; inputs the feature vectors into a preloaded lightweight speech emotion classifier; and outputs a second emotion probability distribution for the speech during that time period. Then, the first and second emotion probability distributions can be weighted and fused to obtain a first emotion attribute.
[0062] In one embodiment of this application, the method further includes: Based on the audio dynamic range of the audio data included in the first audio-visual data, the sampling rate for extracting the audio features corresponding to the first audio-visual data is adjusted; wherein, the audio dynamic range includes the volume range of the audio data included in the first audio-visual data.
[0063] In this embodiment, the sampling rate for audio feature extraction is dynamically adjusted based on the dynamic range of the audio content in the first audio-visual data. Specifically, the audio stream in the first audio-visual data is acquired or received at a higher default sampling rate (e.g., 48kHz), and dynamic range analysis is performed. The system calculates the dynamic range (DR) of the audio in real time. Dynamic range can be defined as the difference between the peak volume and the valley volume over a period of time (usually expressed in decibels, dB). One or more dynamic range thresholds are preset to distinguish the dynamic range of the audio. Low dynamic range typically refers to audio types with gentle volume fluctuations and weak contrasts, such as some folk songs, slow classical music, ambient music, and lyrical songs. For example, DR < 20 dB. High dynamic range typically refers to music types with dramatic volume fluctuations and a large number of transients and strong impact, such as rock, electronic dance music, symphonic climaxes, and movie action sound effects. For example, DR >= 20 dB (or a higher threshold such as 25 dB can be set). Based on the current DR value, perform a sampling rate adjustment: If DR < 20 dB, it can be identified as low dynamic range. Reduce the sampling rate of the audio processing to 22.05 kHz. This sampling rate still adequately covers the fundamental frequency and major harmonics of vocals and most instruments, meeting basic feature extraction requirements. Simultaneously, it reduces data processing volume by approximately 50%, significantly lowering CPU computational power consumption in subsequent Fourier transforms and feature calculations. If DR >= 20 dB, it can be identified as high dynamic range music. Maintain or switch the sampling rate to 48 kHz (or higher). A high sampling rate can more accurately capture rapid transient changes (such as drum beats and string picking) and high-frequency details in music, ensuring the extraction accuracy of key features such as rhythm detection and transient analysis, and avoiding feature loss or distortion due to insufficient sampling.
[0064] This application also provides a data processing apparatus corresponding to the data processing method. Since the principle of the data processing apparatus in this application is similar to the above-mentioned processing method, the implementation of the data processing apparatus can be referred to the implementation of the method, and the repeated parts will not be described again. Figure 6 This application provides a schematic diagram of the structure of a data processing apparatus according to an embodiment, which specifically includes: The acquisition module is used to determine a target lighting effect mode that matches the first audio-visual data based on the first audio-visual data; the first audio-visual data includes the audio features and / or video features of the currently playing content; The determination module is used to determine a target color scheme that matches the target emotional attribute based on the target emotional attribute associated with the target object; the target emotional attribute is obtained based on the image data and / or audio data corresponding to the target object. The processing module is used to merge the target lighting effect mode and the target color scheme mode to obtain the corresponding display information.
[0065] In one embodiment of this application, the determining module is further configured as follows: Perform facial expression recognition on the image data corresponding to the target object to obtain a first emotion probability distribution; Perform voice emotion analysis on the audio data corresponding to the target object to obtain a second emotion probability distribution; Based on the audio features and / or video features corresponding to the first audio-visual data, a third emotion probability distribution is obtained; The target emotion attribute is determined by weighted fusion of at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution. Wherein, the weight of the first emotion probability distribution is greater than the weight of the second emotion probability distribution, and the weight of the second emotion probability distribution is greater than the weight of the third emotion probability distribution.
[0066] In one embodiment of this application, the processing module is further configured as follows: The first emotion probability distribution and the second emotion probability distribution are weighted and fused to obtain the first emotion attribute; The probability distribution of the third emotion is weighted and fused to obtain the second emotion attribute; If the emotion category represented by the first emotion attribute conflicts with the emotion category represented by the second emotion attribute, the first emotion attribute is determined as the target emotion attribute.
[0067] In one embodiment of this application, the processing module is further configured as follows: Determine the first probability value, second probability value, and third probability value of the preset emotion category in the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution; Based on the weights corresponding to at least two of the first probability value, the second probability value, and the third probability value, the target probability value of the emotion category is obtained. The emotion category corresponding to the maximum value among the target probability values is determined as the target emotion attribute.
[0068] In one embodiment of this application, the determining module is further configured as follows: The audio data contained in the first audio-visual data is parsed to determine the audio features corresponding to the audio data contained in the first audio-visual data; the audio features include at least one of the following: volume, frequency distribution, rhythm complexity, and audio dynamic range; The video data contained in the first audio-visual data is parsed, and the corresponding video features are extracted. The video features include at least one of the following: brightness, hue, scene switching frequency, and scene category. Based on the audio features corresponding to the audio data contained in the first audio-visual data, and / or the video features corresponding to the video data contained in the first audio-visual data, the corresponding target lighting effect mode is determined from the preset lighting effect mode library.
[0069] In one embodiment of this application, the determining module is further configured as follows: Based on a preset mapping relationship between emotion categories and color schemes, the color scheme corresponding to the emotion category represented by the target emotion attribute is determined as the target color scheme; wherein, the color scheme includes multiple preset color values with different visual functions.
[0070] In one embodiment of this application, the acquisition module is further configured as follows: Based on the audio dynamic range of the audio data included in the first audio-visual data, the sampling rate for extracting the audio features corresponding to the first audio-visual data is adjusted; wherein, the audio dynamic range includes the volume range of the audio data included in the first audio-visual data.
[0071] Based on the same inventive concept, such as Figure 7 As shown, this embodiment also includes an electronic device, comprising: Cameras are used to capture emotional information about the target subject; A microphone is used to acquire the voice information of the target object; A display for displaying target display information, wherein the target display information is obtained by fusing at least the emotional information of the target object, the voice information of the target object, and the audio and / or video features of the currently playing content included in the first audio-visual data; Memory, used to store executable programs; A processor is configured to execute the executable program to perform the following steps: Based on the first audio-visual data, a target lighting effect mode matching the first audio-visual data is determined; the first audio-visual data includes the audio features and / or video features of the currently playing content. Based on the target emotion attribute associated with the target object, a target color scheme mode that matches the target emotion attribute is determined; the target emotion attribute is obtained based on the image data and / or audio data corresponding to the target object. The target lighting effect mode and the target color scheme mode are merged to obtain the corresponding display information.
[0072] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A data processing method, comprising: Based on the first audio-visual data, a target lighting effect mode matching the first audio-visual data is determined; the first audio-visual data includes the audio features and / or video features of the currently playing content. Based on the target emotion attribute associated with the target object, a target color scheme mode that matches the target emotion attribute is determined; the target emotion attribute is obtained based on the image data and / or audio data corresponding to the target object. The target lighting effect mode and the target color scheme mode are merged to obtain the corresponding display information.
2. The method according to claim 1, further comprising: Perform facial expression recognition on the image data corresponding to the target object to obtain a first emotion probability distribution; Perform voice emotion analysis on the audio data corresponding to the target object to obtain a second emotion probability distribution; Based on the audio features and / or video features corresponding to the first audio-visual data, a third emotion probability distribution is obtained; The target emotion attribute is determined by weighted fusion of at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution. Wherein, the weight of the first emotion probability distribution is greater than the weight of the second emotion probability distribution, and the weight of the second emotion probability distribution is greater than the weight of the third emotion probability distribution.
3. The method according to claim 2, further comprising: The first emotion probability distribution and the second emotion probability distribution are weighted and fused to obtain the first emotion attribute; The probability distribution of the third emotion is weighted and fused to obtain the second emotion attribute; If the emotion category represented by the first emotion attribute conflicts with the emotion category represented by the second emotion attribute, the first emotion attribute is determined as the target emotion attribute.
4. The method according to claim 2, wherein the weighted fusion of at least two of the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution to determine the target emotion attribute includes: Determine the first probability value, second probability value, and third probability value of the preset emotion category in the first emotion probability distribution, the second emotion probability distribution, and the third emotion probability distribution; Based on the weights corresponding to at least two of the first probability value, the second probability value, and the third probability value, the target probability value of the emotion category is obtained. The emotion category corresponding to the maximum value among the target probability values is determined as the target emotion attribute.
5. The method according to claim 1, wherein determining the target lighting effect mode matching the first audio-visual data based on the first audio-visual data includes: The audio data contained in the first audio-visual data is parsed to determine the audio features corresponding to the audio data contained in the first audio-visual data. The audio features include at least one of the following: volume, frequency distribution, rhythmic complexity, and audio dynamic range; The video data contained in the first audio-visual data is parsed, and the corresponding video features are extracted. The video features include at least one of the following: brightness, hue, scene switching frequency, and scene category. Based on the audio features corresponding to the audio data contained in the first audio-visual data, and / or the video features corresponding to the video data contained in the first audio-visual data, the corresponding target lighting effect mode is determined from the preset lighting effect mode library.
6. The method according to claim 1, wherein determining the target color scheme pattern matching the target emotional attribute based on the target emotional attribute associated with the target object includes: Based on a preset mapping relationship between emotion categories and color schemes, the color scheme corresponding to the emotion category represented by the target emotion attribute is determined as the target color scheme; wherein, the color scheme includes multiple preset color values with different visual functions.
7. The method according to claim 2, wherein, The operations of facial expression recognition on the image data corresponding to the target object and voice emotion analysis on the audio data corresponding to the target object are executed synchronously and in parallel on different processor cores or threads.
8. The method according to claim 1, further comprising: Based on the audio dynamic range of the audio data included in the first audio-visual data, the sampling rate for extracting the audio features corresponding to the first audio-visual data is adjusted; wherein, the audio dynamic range includes the volume range of the audio data included in the first audio-visual data.
9. A data processing apparatus, comprising: The acquisition module is used to determine a target lighting effect mode that matches the first audio-visual data based on the first audio-visual data; the first audio-visual data includes the audio features and / or video features of the currently playing content; The determination module is used to determine a target color scheme that matches the target emotional attribute based on the target emotional attribute associated with the target object; the target emotional attribute is obtained based on the image data and / or audio data corresponding to the target object. The processing module is used to merge the target lighting effect mode and the target color scheme mode to obtain the corresponding display information.
10. An electronic device, comprising: Cameras are used to capture emotional information about the target subject; A microphone is used to acquire the voice information of the target object; A display for displaying target display information, wherein the target display information is obtained by fusing at least the emotional information of the target object, the voice information of the target object, and the audio and / or video features of the currently playing content included in the first audio-visual data; Memory, used to store executable programs; A processor is configured to execute the executable program to perform the following steps: Based on the first audio-visual data, a target lighting effect mode matching the first audio-visual data is determined; the first audio-visual data includes the audio features and / or video features of the currently playing content. Based on the target emotion attribute associated with the target object, a target color scheme mode that matches the target emotion attribute is determined; the target emotion attribute is obtained based on the image data and / or audio data corresponding to the target object. The target lighting effect mode and the target color scheme mode are merged to obtain the corresponding display information.
Citation Information
Cited By
Light effect determination method, sound, atmosphere lamp device and robot
CN122138309A