Multi-modal fusion intelligent sensing switch control system
The intelligent sensor switch control system, which integrates multimodal fusion, enables accurate identification of user behavior and emotional state, dynamically adjusts the lighting environment, solves the problem that existing sensor switch systems cannot meet personalized needs, and improves the intelligence and comfort of lighting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHEN ZHEN TOADA ELECTRONICS CO LTD
- Filing Date
- 2025-07-22
- Publication Date
- 2026-05-19
AI Technical Summary
Existing home sensor switch systems rely solely on human movement for control, lacking the ability to understand and respond to complex home behavior scenarios, and thus failing to meet users' personalized comfort needs.
The intelligent sensor switch control system adopts multimodal fusion. It collects data in real time through the multimodal perception module, performs feature extraction and fusion through the dynamic extraction module, accurately identifies the user's behavior and emotional state through the state recognition module, generates a lighting parameter configuration scheme through the context decision module, and finally realizes intelligent lighting control through the control execution module.
It achieves accurate recognition of user behavior and emotional state, can dynamically adjust the lighting environment to meet personalized needs, improve the intelligence and comfort of lighting, and at the same time protect user privacy and enhance the system's adaptability and robustness.
Smart Images

Figure CN120568555B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology, and more specifically, to a multimodal fusion intelligent inductive switch control system. Background Technology
[0002] In the construction of modern smart homes, sensor switches, as one of the basic components of environmental control, are widely used in the automated management of subsystems such as lighting, air conditioning, and curtains. However, most existing household sensor switches are based on infrared or microwave sensors, which usually only realize switch action control based on human movement. They have limited functions and lack the ability to understand and respond to complex household behavior scenarios.
[0003] In real-world home environments, users' activity states and emotional needs vary significantly, such as a hurried morning commute, a relaxing evening rest, and social interactions during family gatherings. These state changes are often accompanied by multimodal characteristics such as user movement behavior, spatial dwelling patterns, interaction frequency, and voice communication. However, existing switch control systems struggle to comprehensively perceive this information and cannot dynamically adjust the lighting environment according to user scenarios, resulting in insufficiently intelligent lighting control that fails to meet personalized comfort needs.
[0004] In view of this, the present invention proposes a multimodal fusion intelligent inductive switch control system to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a multimodal fusion intelligent inductive switch control system, comprising:
[0006] The multimodal sensing module is used to acquire multimodal sensing data in real time.
[0007] The dynamic extraction module is used to apply a hierarchical spatiotemporal computing mechanism to extract and fuse features from multimodal sensing data, thereby obtaining a temporal representation vector sequence of user activity patterns.
[0008] The state recognition module is used to perform semantic parsing on the temporal representation vector sequence of user activity patterns to accurately identify the user's behavioral and emotional states.
[0009] The contextual decision-making module is used to dynamically generate lighting parameter configuration schemes based on a pre-built intelligent response decision engine, combined with behavioral state, emotional state, and multimodal perception data.
[0010] The control execution module is used to convert the lighting parameter configuration scheme into a lighting control command sequence and send it to the lighting control device via the KNX bus.
[0011] Furthermore, the multimodal perception data includes behavioral perception data and environmental perception data continuously collected at a time points; the behavioral perception data includes audio perception data and infrared perception data.
[0012] The steps to obtain the temporal representation vector sequence of user activity patterns include:
[0013] Step S1: Perform spatiotemporal alignment on the multimodal sensing data to form a spatiotemporal data matrix;
[0014] Step S2: Perform low-level feature extraction on the spatiotemporal data matrix to obtain behavioral feature vectors and audio feature vectors;
[0015] Step S3: Perform high-level fusion processing on the behavioral feature vector and the audio feature vector to generate a temporal representation vector sequence of user activity patterns.
[0016] Furthermore, based on all infrared sensing data, a trigger sequence is generated; the movement speed between every two adjacent infrared sensing data in the trigger sequence is calculated sequentially, and a movement speed vector is formed; different digital labels are set for different functional spaces in the home environment and marked as space labels; adjacent infrared sensing data in the trigger sequence with the same corresponding space label are taken as a trigger set; the dwell time of the space label corresponding to each trigger set is calculated and a dwell time vector is formed; the space labels corresponding to each trigger set are arranged in ascending order according to the trigger sequence to generate a movement trajectory; the movement speed vector, dwell time vector, and movement trajectory are combined into a behavior feature vector.
[0017] Furthermore, the audio perception data corresponding to each trigger set is acquired and marked as mobile audio data; real-time sound intensity and human voice frequency band energy are extracted from each mobile audio data, and sound intensity feature vector and human voice feature vector are constructed respectively; based on the sound intensity feature vector and human voice feature vector, the sound intensity fluctuation frequency and background noise index of each spatial label in the movement trajectory are calculated; the sound intensity feature vector, human voice feature vector, sound intensity fluctuation frequency and background noise index are integrated into an audio feature vector.
[0018] Furthermore, methods for accurately identifying users' behavioral and emotional states include:
[0019] Based on the temporal representation vector sequence of user activity patterns, the average entropy fluctuation is calculated, and the attribution degree of each behavioral state and emotional state is inferred. The attribution degree of each behavioral state and emotional state and the average entropy fluctuation are used as analysis data. The analysis data is input into the trained state recognition model to predict the corresponding state set. The state set includes a behavioral label and an emotional label. Based on the state set, the user's behavioral state and emotional state are identified.
[0020] Furthermore, methods for calculating average entropy fluctuations include:
[0021] Collect historical sequences, cluster the historical sequences to obtain b state category centers, and assign a corresponding state symbol to each state category center; calculate the matching difference between each multimodal feature vector and each state category center; based on the matching difference, obtain the matching category center of each multimodal feature vector from all state category centers, and map each multimodal feature vector to the corresponding state symbol to form a time-series symbol sequence;
[0022] Set a sliding window of length c, and divide the time series symbol sequence into d symbol windows according to the sliding window; calculate the symbol entropy of each symbol window and form a symbol entropy sequence; calculate the entropy fluctuation between every two adjacent symbol entropies in the symbol entropy sequence in turn, and average all entropy fluctuations to obtain the average entropy fluctuation.
[0023] Furthermore, methods for inferring the attribution degree of each behavioral state to an emotional state include:
[0024] For each dimension of the multimodal feature vector, multiple corresponding fuzzy sets are constructed. Each multimodal feature vector is converted into the degree of belonging to each corresponding fuzzy set through fuzzification technology. Fuzzy rules are defined. All fuzzified multimodal feature vectors are matched with the fuzzy rules, and fuzzy inference is performed using fuzzy inference methods to obtain fuzzy inference results. The fuzzy inference results include the degree of belonging to each behavioral state and emotional state.
[0025] Furthermore, the construction methods for the intelligent response decision engine include:
[0026] Define an oscillator set, where each oscillator corresponds to a behavioral or emotional state. Based on the identified behavioral and emotional states, retrieve the corresponding oscillator from the set. Using environmental perception data, correct the oscillator parameters for each oscillator, including angular frequency and initial phase. For all oscillators, construct an oscillator coupling network to further correct the corrected initial phase for each oscillator. Preset basic lighting parameters and dynamically optimize them based on the corrected oscillators to generate a lighting parameter configuration scheme.
[0027] Furthermore, the lighting parameters include brightness, hue, saturation, and the rhythm of change;
[0028] The dynamic optimization method for brightness is as follows: based on a preset brightness weighted set, the output value of each oscillator is weighted and summed, and then the brightness in the basic lighting parameters is added to obtain the dynamically optimized brightness; the dynamic optimization methods for hue and saturation are the same as those for brightness.
[0029] Furthermore, the dynamic optimization method for the changing rhythm is as follows: the absolute values of the corrected angular frequencies of all oscillators are averaged to obtain the dynamically optimized changing rhythm.
[0030] The technical effects and advantages of the multimodal fusion intelligent inductive switch control system of this invention are as follows:
[0031] Employing multimodal sensing methods such as sound and infrared to comprehensively perceive user behavior characteristics and environmental conditions overcomes the limitations of traditional single-sensor-based systems, enabling more accurate capture of complex user activity patterns in the home environment. By combining spatiotemporal feature extraction and semantic modeling, it achieves precise identification of user behavioral and emotional states, ensuring that intelligent lighting control accurately responds to environmental and user conditions. Based on an oscillator-coupled network, it constructs an intelligent response decision engine oriented towards user conditions, dynamically optimizing lighting parameters based on environmental perception data. This allows the lighting system to better adapt to and meet personalized needs in different usage scenarios. It achieves fully automatic intelligent control of home lighting, not only improving the intelligence and comfort of lighting but also effectively protecting user privacy and enhancing the system's adaptability and robustness, thereby meeting the needs of modern families for personalized and intelligent environments. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the multimodal fusion intelligent inductive switch control system of Embodiment 1 of the present invention. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] Example 1
[0035] Please see Figure 1 As shown, the multimodal fusion intelligent sensor switch control system of this embodiment includes a multimodal sensing module, a dynamic extraction module, a state recognition module, a situational decision-making module, and a control execution module; each module is connected by wired and / or wireless means to realize data transmission between modules.
[0036] The multimodal sensing module is used to collect multimodal sensing data in real time.
[0037] Multimodal perception data includes behavioral perception data and environmental perception data continuously collected at a time points, where a is an integer greater than 1;
[0038] Behavioral perception data includes audio perception data and infrared perception data. Audio perception data refers to acoustic signals collected by sound sensors, which are used to detect sound features such as human voices and conversations, thereby helping to determine the user's activity status and interaction. Infrared perception data refers to infrared radiation information collected by infrared sensors, i.e., the trigger time of infrared sensors, which are mainly used to detect the presence of human bodies, movement trajectories, and areas of stay, thereby helping to perceive the user's location and behavioral status.
[0039] Environmental sensing data includes light intensity and ambient temperature; light intensity is obtained through a light sensor, and ambient temperature is obtained through a temperature sensor.
[0040] It should be noted that the sound sensor, infrared sensor, light sensor, and temperature sensor are all installed in each functional space of the home environment, such as the bedroom, living room, and study, to achieve accurate perception and dynamic detection of multimodal sensing data in each functional area, and the sampling frequency of each sensor is set to be consistent; in particular, multiple infrared sensors are deployed in each functional space, which helps to infer the user's movement trajectory, speed changes, and other behavioral characteristics in the home environment by using the time sequence and spatial distribution information of the infrared sensor activation.
[0041] It should be understood that the sound sensor in this embodiment focuses on collecting acoustic signals to extract sound features (such as sound intensity, spectrum, etc.) for identifying the user's behavioral state, without recording or storing voice, thereby effectively protecting user privacy.
[0042] The dynamic extraction module is used to apply a hierarchical spatiotemporal computing mechanism to extract and fuse features from multimodal sensing data, thereby obtaining a temporal representation vector sequence of user activity patterns.
[0043] The steps to obtain the temporal representation vector sequence of user activity patterns include:
[0044] Step S1: Perform spatiotemporal alignment on the multimodal sensing data to form a spatiotemporal data matrix;
[0045] The timestamp corresponding to each data point in the multimodal sensing data is obtained through the built-in clocks in each sensor, and data with the same timestamp are grouped into a time-aligned data set. Different digital labels are set for different functional spaces and marked as spatial labels. According to the functional space of each sensor, each data point in the multimodal sensing data is labeled with a corresponding spatial label. Data with the same spatial label in each group of time-aligned data are grouped into a spatiotemporal aligned data set. A spatiotemporal data matrix is constructed based on the spatiotemporal aligned data set. In the spatiotemporal data matrix, the rows represent consecutive timestamps, the columns represent spatial labels of different functional spaces, and each element corresponds to a group of spatiotemporal aligned data sets.
[0046] Step S2: Perform low-level feature extraction on the spatiotemporal data matrix to obtain behavioral feature vectors and audio feature vectors;
[0047] All infrared sensing data are sorted from earliest to latest according to their corresponding timestamps to generate a trigger sequence. The movement speed between every two adjacent infrared sensing data in the trigger sequence is calculated and a movement speed vector is formed. Based on the spatial label corresponding to each infrared sensing data in the trigger sequence, adjacent infrared sensing data with the same spatial label are grouped into a trigger set. The time difference between the infrared sensing data with the latest timestamp and the infrared sensing data with the earliest timestamp in each trigger set is calculated to obtain the dwell time of the spatial label corresponding to each trigger set, and a dwell time vector is formed. The spatial labels corresponding to each trigger set are arranged in ascending order according to the trigger sequence to generate a movement trajectory. The movement speed vector, dwell time vector, and movement trajectory are combined into a behavior feature vector to reflect the user's movement rhythm and spatial state changes.
[0048] From multimodal sensing data, audio sensing data corresponding to the timestamp and spatial label of each trigger set is obtained and marked as mobile audio data. Real-time sound intensity and human voice frequency band energy are extracted from each mobile audio data, and sound intensity feature vector and human voice feature vector are constructed respectively. Based on the sound intensity feature vector and human voice feature vector, the sound intensity fluctuation frequency and background noise index corresponding to each spatial label in the movement trajectory are calculated. Real-time sound intensity represents the overall intensity of sound at an instant, and human voice frequency band energy represents the energy of the sound in the frequency range specific to human voice. The extraction method of real-time sound intensity and human voice frequency band energy is existing technology, and the specific process will not be elaborated here. The sound intensity feature vector, human voice feature vector, sound intensity fluctuation frequency and background noise index are integrated into an audio feature vector to reflect the acoustic activity in the home environment.
[0049] The method for calculating the moving speed between two adjacent infrared sensing data is as follows: Each infrared sensor is assigned a different digital tag and marked as an infrared tag; the corresponding sensor distance is obtained based on the infrared tags corresponding to two adjacent infrared sensing data; the distance between each pair of infrared sensors is obtained by a person skilled in the art through on-site measurement during the installation of the infrared sensors; the difference between the timestamps corresponding to two adjacent infrared sensing data is used as the time difference; the moving speed between two adjacent infrared sensing data is obtained based on the ratio of the sensor distance to the time difference.
[0050] The method for calculating the sound intensity fluctuation frequency is as follows: real-time sound intensities with the same corresponding trigger set are grouped into a sound intensity set; each real-time sound intensity in each sound intensity set is compared with a preset sound intensity threshold; real-time sound intensities with values greater than the sound intensity threshold are marked as active sound intensities, and real-time sound intensities with values less than or equal to the sound intensity threshold are marked as silent sound intensities; the number of switching times in each sound intensity set is counted, and the ratio of the number of switching times in each sound intensity set to the dwell time of the corresponding trigger set is taken as the sound intensity fluctuation frequency of each sound intensity set; wherein, the number of switching times is the total number of times active sound intensity switches to silent sound intensity, or silent sound intensity switches to active sound intensity; the sound intensity threshold is preset by those skilled in the art according to the actual situation.
[0051] The method for calculating the background noise index is as follows: subtract the human voice frequency band energy from the real-time sound intensity to obtain the noise energy; use the ratio of noise energy to real-time sound intensity as the background noise index; when the real-time sound intensity is 0, the background noise index is also 0.
[0052] Step S3: Perform high-level fusion processing on the behavioral feature vector and the audio feature vector to generate a temporal representation vector sequence of user activity patterns;
[0053] The behavioral feature vectors with the same timestamp are fused with the audio feature vectors to form a multimodal feature vector; all multimodal feature vectors are arranged according to their corresponding timestamps to generate a temporal representation vector sequence of user activity patterns.
[0054] The state recognition module is used to perform semantic parsing on the temporal representation vector sequence of user activity patterns to accurately identify the user's behavioral and emotional states.
[0055] Methods for accurately identifying a user's behavioral and emotional states include:
[0056] Historical sequences are collected, which are time-series representation vector sequences obtained from historical moments. These historical sequences are acquired through the database built into the switch control system. The K-Means clustering algorithm is used to cluster the historical sequences, obtaining b state category centers. Here, b is an integer greater than 1, and each state category center represents a user's basic state, such as low activity and low sound intensity, or high activity and high sound intensity. The K-Means clustering algorithm is an existing technology, and its specific process will not be elaborated upon here. A corresponding state symbol is assigned to each state category center; for example, low activity and low sound intensity is assigned the state symbol A, and high activity and high sound intensity is assigned the state symbol B, etc.
[0057] Calculate the Euclidean distance between each multimodal feature vector and each state category center, and label it as the matching difference degree; take the state category center with the smallest matching difference degree among all matching differences corresponding to each multimodal feature vector as the matching category center of the corresponding multimodal feature vector, and map each multimodal feature vector to the corresponding state symbol to form a temporal symbol sequence;
[0058] A sliding window with a length c is set, the specific value of c being preset by those skilled in the art based on actual conditions; according to the sliding window, the time sequence symbol is divided into d symbol windows; the occurrence frequency of each state symbol within each symbol window is counted, and the ratio of each occurrence frequency to the window length is taken as the occurrence frequency of each state symbol within each symbol window; based on the occurrence frequency, the symbol entropy (i.e., information entropy) of each symbol window is calculated, and a symbol entropy sequence is formed; the entropy fluctuation (i.e., the difference between each two adjacent symbol entropies) between every two adjacent symbol entropies in the symbol entropy sequence is calculated sequentially, and all entropy fluctuations are averaged to obtain the average entropy fluctuation.
[0059] For each dimension of the multimodal feature vector, multiple corresponding fuzzy sets are constructed. For example, the fuzzy sets corresponding to movement speed are fast, medium, and slow, while the fuzzy sets corresponding to sound intensity fluctuation frequency are high, medium, and low. Each multimodal feature vector is converted into the degree of belonging to its corresponding fuzzy set using fuzzification techniques. Fuzzification is the process of converting precise numerical values into the degree of belonging to a fuzzy set. Examples of fuzzification techniques include triangular and trapezoidal degree functions. Fuzzy rules are defined based on expert knowledge or relevant literature. All fuzzified multimodal feature vectors are matched with the fuzzy rules, and fuzzy inference methods (such as the Mamdani fuzzy inference model and the Sugeno fuzzy inference model) are used for fuzzy inference to obtain the fuzzy inference results. The fuzzy inference results include the degree of belonging to each behavioral state and emotional state (i.e., the degree of belonging to each behavioral state and the degree of belonging to each emotional state). Behavioral states include sitting quietly reading, watching TV, and family gatherings; emotional states include calm, anxiety, and excitement.
[0060] The attribution degree and average entropy fluctuation of each behavioral state and emotional state in the fuzzy inference results are used as analysis data. The analysis data is input into a trained state recognition model to predict the corresponding state set. The state set includes a behavioral label and an emotional label. The behavioral label is the numerical label corresponding to the behavioral state, and different behavioral states correspond to different behavioral labels. The emotional label is the numerical label corresponding to the emotional state, and different emotional states correspond to different emotional labels. Based on the state set, the user's behavioral state and emotional state are identified.
[0061] The state recognition model is a deep neural network model, which includes an input layer, hidden layers, and an output layer. Each hidden layer contains multiple neurons, and each neuron is connected to the neurons in the next layer. The connections contain weights that determine the importance and influence of the data transmitted in the neural network. An activation function is applied to each neuron between the hidden layer and the output layer. The activation function introduces non-linearity, allowing the network to learn more complex patterns and features. The deep neural network model is a current technology, and the specific training process will not be described in detail here.
[0062] It should be understood that the purpose of calculating average entropy fluctuations is to measure the stability and complexity of user state changes, thereby enhancing the accuracy of identifying behavioral and emotional states. By calculating the symbolic entropy under a sliding window and obtaining its average entropy fluctuations in the time series, the dynamic characteristics of user state transitions can be captured: if the average entropy fluctuations are small, it indicates that the user state is relatively stable, possibly corresponding to states such as sitting quietly reading or being calm; if the average entropy fluctuations are large, it indicates that the user state changes frequently, possibly corresponding to states such as family gatherings or anxiety. As a time-dynamic feature, average entropy fluctuations, together with the fuzzy inference results, constitute the input of the state recognition model, improving the state recognition model's ability to distinguish behavioral and emotional states and the depth of semantic parsing.
[0063] It should be noted that semantic parsing through the temporal representation vector sequence of user activity patterns can integrate the spatiotemporal dynamic features of multimodal perception data to accurately identify the user's behavioral and emotional states. Compared with traditional methods that rely on image processing, this embodiment avoids directly collecting and processing the user's visual images, effectively protecting user privacy and reducing the risk of sensitive information leakage. At the same time, it is not affected by environmental factors such as lighting and occlusion, and has stronger robustness and adaptability. By capturing continuous changes in behavior and emotions through temporal modeling, it improves the timeliness and accuracy of recognition, thus achieving more stable, comprehensive, and privacy-friendly user state perception in complex environments.
[0064] The contextual decision-making module is used to dynamically generate lighting parameter configuration schemes based on a pre-built intelligent response decision engine, combined with behavioral state, emotional state, and multimodal perception data.
[0065] The construction methods for intelligent response decision engines include:
[0066] Define a set of oscillators, where each oscillator corresponds to a behavioral or emotional state; the expression for an oscillator is: In the formula, O i (t) represents the output value of the i-th oscillator at time t, where time t is the running time elapsed since the start of the switching control system. A i ω represents the amplitude of the i-th oscillator.i This represents the angular frequency of the i-th oscillator. This represents the initial phase of the i-th oscillator; wherein, the amplitude, angular frequency, and initial phase of each oscillator are preset by those skilled in the art; the amplitude is set according to the intensity or subjective significance of the influence of the corresponding behavioral or emotional state on the lighting parameters, for example, excitement corresponds to a larger amplitude; the angular frequency is set according to the activity level of the corresponding behavioral or emotional state, the greater the activity level, the larger the corresponding angular frequency, for example, a family gathering corresponds to a higher angular frequency, while sitting and reading corresponds to a lower angular frequency; the initial phase is set according to the order in which the corresponding behavioral and emotional states occur in real life;
[0067] Based on the identified behavioral and emotional states, the corresponding oscillators are obtained from the oscillator set. It should be noted that the oscillators mentioned in the following process are all oscillators obtained based on the identified behavioral and emotional states.
[0068] Based on environmental perception data, the oscillator parameters of each oscillator are corrected. These parameters include angular frequency and initial phase. The method for correcting the oscillator parameters is as follows: a preset modulation set is established, comprising a frequency modulation set and a phase modulation set. The frequency modulation set includes the modulation coefficient of light intensity or ambient temperature on the corresponding angular frequency of each oscillator, and the phase modulation set includes the modulation coefficient of light intensity or ambient temperature on the corresponding initial phase of each oscillator. The modulation set reflects the degree of influence of light intensity and ambient temperature on behavioral or emotional states and is preset by those skilled in the art based on actual conditions. Based on the modulation set, the light intensity and ambient temperature are weighted and summed separately to calculate the parameter correction coefficient for each oscillator, which includes a frequency correction coefficient and a phase correction coefficient. The oscillator parameters of each oscillator are sequentially added to their corresponding parameter correction coefficients to complete the correction of the oscillator parameters for each oscillator.
[0069] For all oscillators, an oscillator coupling network is constructed to further correct the initial phase of each oscillator. The oscillator coupling network is constructed using the Kuramoto model to achieve phase coordination and dynamic interaction between oscillators. The Kuramoto model is an existing technology, and its specific expression will not be elaborated here. Basic lighting parameters are preset, and the basic lighting parameters are dynamically optimized based on the corrected oscillators to generate a lighting parameter configuration scheme.
[0070] The basic lighting parameters are preset by those skilled in the art based on the actual application scenario, human comfort standards, and user needs. The lighting parameters include brightness, hue, saturation, and change rhythm. Brightness represents the intensity of the light, indicating the brightness emitted by the light source. Hue represents the color of the light. Saturation represents the purity of the light color; a higher saturation indicates a more vibrant light color, while a lower saturation indicates a light color closer to gray. Change rhythm represents the frequency of light changes, such as light flashing or gradation, which affects the dynamic effect of the light.
[0071] The dynamic optimization method for brightness is as follows: based on a preset brightness weighted set, the output value of each oscillator is weighted and summed, and then the brightness in the basic lighting parameters is added to obtain the dynamically optimized brightness. The brightness weighted set includes the influence weight of each oscillator on the brightness, which is preset by those skilled in the art according to the influence intensity of the behavioral state or emotional state corresponding to each oscillator on the brightness. The dynamic optimization methods for hue and saturation are the same as those for brightness, except that the dynamic optimization method for hue uses a preset hue weighted set, and the dynamic optimization method for saturation uses a preset saturation weighted set.
[0072] The dynamic optimization method for the changing rhythm is as follows: the absolute values of the corrected angular frequencies of all oscillators are averaged to obtain the dynamically optimized changing rhythm.
[0073] The control execution module is used to convert the lighting parameter configuration scheme into a lighting control command sequence and send it to the lighting control device via the KNX bus.
[0074] The lighting control command sequence is a set of commands encoded according to the KNX protocol format that can be recognized and executed by lighting control devices to adjust the brightness, hue, saturation and rhythm of the lights.
[0075] KNX bus is an internationally standardized communication protocol and network technology for intelligent building control systems. It adopts a distributed architecture, allowing all connected devices to exchange data through a shared communication cable, thereby achieving interconnection and collaborative control between devices.
[0076] Lighting control equipment refers to lighting terminal equipment that integrates a KNX communication module in a home environment. It can receive and parse lighting control command sequences and adjust lighting parameters to achieve intelligent control effects.
[0077] This embodiment employs multimodal sensing methods, including sound and infrared, to comprehensively perceive user behavior characteristics and environmental conditions, overcoming the limitations of traditional single-sensor-based systems. It can more accurately capture complex activity patterns of users in their home environment. By combining spatiotemporal feature extraction and semantic modeling, it achieves precise identification of user behavior and emotional states, ensuring that intelligent lighting control accurately responds to environmental and user conditions. Based on an oscillator-coupled network, it constructs an intelligent response decision engine oriented towards user conditions, dynamically optimizing lighting parameters based on environmental perception data. This allows the lighting system to better adapt to and meet personalized needs in different usage scenarios. It achieves fully automatic intelligent control of home lighting, not only improving the intelligence and comfort of lighting but also effectively protecting user privacy and enhancing the system's adaptability and robustness, thereby meeting the needs of modern families for personalized and intelligent environments.
[0078] Example 2
[0079] This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable code that, when executed by the one or more processors, can perform the multimodal fusion intelligent sensor switch control system described above.
[0080] The methods or systems according to embodiments of this application can also be implemented using the architecture of the electronic device shown in this application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or hard disk, may store the multimodal fusion intelligent inductive switch control system provided in this application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in this application is merely exemplary; when implementing different devices, one or more components in the electronic device shown in this application may be omitted according to actual needs.
[0081] Example 3
[0082] One embodiment of this application discloses a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by a processor, a multimodal fusion intelligent inductive switch control system according to an embodiment of this application, as described with reference to the above figures, can be executed. The storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0083] Furthermore, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, this application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be executed by a processor to perform instructions corresponding to the method steps provided in this application, such as a multimodal fusion intelligent sensor switch control system. When this computer program is executed by a central processing unit (CPU), it performs the functions defined in the method of this application.
[0084] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0085] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0086] In the description of this invention, it should be understood that the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0087] In the description of this invention, unless otherwise stated, "a plurality of" means two or more.
[0088] In the description of this invention, "several" means one or more, and "a large number" means two or more.
[0089] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0090] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0091] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A multimodal fusion intelligent inductive switch control system, characterized in that, include: The multimodal sensing module is used to acquire multimodal sensing data in real time. The dynamic extraction module is used to apply a hierarchical spatiotemporal computing mechanism to extract and fuse features from multimodal sensing data, thereby obtaining a temporal representation vector sequence of user activity patterns. The steps to obtain the temporal representation vector sequence of user activity patterns include: Step S1: Perform spatiotemporal alignment on the multimodal sensing data to form a spatiotemporal data matrix; Step S2: Perform low-level feature extraction on the spatiotemporal data matrix to obtain behavioral feature vectors and audio feature vectors; Step S3: Perform high-level fusion processing on the behavioral feature vector and the audio feature vector to generate a temporal representation vector sequence of user activity patterns; The state recognition module is used to perform semantic parsing on the temporal representation vector sequence of user activity patterns to accurately identify the user's behavioral and emotional states. Methods for accurately identifying a user's behavioral and emotional states include: Based on the temporal representation vector sequence of user activity patterns, the average entropy fluctuation is calculated, and the attribution degree of each behavioral state and emotional state is inferred. The attribution degree of each behavioral state and emotional state and the average entropy fluctuation are used as analysis data. The analysis data is input into the trained state recognition model to predict the corresponding state set. The state set includes a behavioral label and an emotional label. Based on the state set, the user's behavioral state and emotional state are identified. Methods for calculating average entropy fluctuations include: Collect historical sequences, perform clustering on the historical sequences, and obtain... Each state category center is assigned a corresponding state symbol; the matching difference between each multimodal feature vector in the temporal representation vector sequence and each state category center is calculated; based on the matching difference, the matching category center of each multimodal feature vector is obtained from all state category centers, and each multimodal feature vector is mapped to the corresponding state symbol to form a temporal symbol sequence; Set window length The sliding window is used to divide the timing symbol sequence into... The system uses a symbol window to count the occurrences of each state symbol within each window and assigns the ratio of each occurrence to the window length as the frequency of each state symbol within that window. It then calculates the information entropy of each symbol window based on its frequency and labels it as the symbol entropy, forming a symbol entropy sequence. Finally, it calculates the difference between the entropies of every two adjacent symbols in the sequence to obtain entropy fluctuations. All entropy fluctuations are then averaged to obtain the average entropy fluctuation. The contextual decision-making module is used to dynamically generate lighting parameter configuration schemes based on a pre-built intelligent response decision engine, combined with behavioral state, emotional state, and multimodal perception data. The control execution module is used to convert the lighting parameter configuration scheme into a lighting control command sequence and send it to the lighting control device via the KNX bus.
2. The multimodal fusion intelligent inductive switch control system according to claim 1, characterized in that, Multimodal sensing data includes Behavioral and environmental perception data were continuously collected at various time points; the behavioral perception data included audio and infrared perception data.
3. The multimodal fusion intelligent inductive switch control system according to claim 2, characterized in that, Based on all infrared sensing data, a trigger sequence is generated; the movement speed between every two adjacent infrared sensing data in the trigger sequence is calculated sequentially, and a movement speed vector is formed; different digital labels are set for different functional spaces in the home environment and marked as space labels; adjacent infrared sensing data in the trigger sequence with the same corresponding space label are grouped into a trigger set; the dwell time of the corresponding space label for each trigger set is calculated and a dwell time vector is formed; the space labels corresponding to each trigger set are arranged in ascending order according to the trigger sequence to generate a movement trajectory; the movement speed vector, dwell time vector, and movement trajectory are combined into a behavior feature vector.
4. The multimodal fusion intelligent inductive switch control system according to claim 3, characterized in that, Acquire the audio perception data corresponding to each trigger set and mark it as mobile audio data; extract the real-time sound intensity and human voice frequency band energy from each mobile audio data, and construct the sound intensity feature vector and human voice feature vector respectively. Based on the sound intensity feature vector and the human voice feature vector, calculate the sound intensity fluctuation frequency and background noise index of each spatial tag in the movement trajectory; integrate the sound intensity feature vector, human voice feature vector, sound intensity fluctuation frequency and background noise index into an audio feature vector.
5. The multimodal fusion intelligent inductive switch control system according to claim 4, characterized in that, Methods for inferring the attribution of each behavioral state to an emotional state include: For each dimension of the multimodal feature vector, multiple corresponding fuzzy sets are constructed. Each multimodal feature vector is converted into the degree of belonging to each corresponding fuzzy set through fuzzification technology. Fuzzy rules are defined. All fuzzified multimodal feature vectors are matched with the fuzzy rules, and fuzzy inference is performed using fuzzy inference methods to obtain fuzzy inference results. The fuzzy inference results include the degree of belonging to each behavioral state and emotional state.
6. The multimodal fusion intelligent inductive switch control system according to claim 5, characterized in that, The construction methods for intelligent response decision engines include: Define an oscillator set, where each oscillator corresponds to a behavioral or emotional state. Based on the identified behavioral and emotional states, retrieve the corresponding oscillator from the set. Using environmental perception data, correct the oscillator parameters for each oscillator, including angular frequency and initial phase. For all oscillators, construct an oscillator coupling network to further correct the corrected initial phase for each oscillator. Preset basic lighting parameters and dynamically optimize them based on the corrected oscillators to generate a lighting parameter configuration scheme.
7. The multimodal fusion intelligent inductive switch control system according to claim 6, characterized in that, Lighting parameters include brightness, hue, saturation, and the rhythm of change; The dynamic optimization method for brightness is as follows: based on a preset brightness weighted set, the output value of each oscillator is weighted and summed, and then the brightness in the basic lighting parameters is added to obtain the dynamically optimized brightness. The dynamic optimization methods for hue and saturation are consistent with those for brightness.
8. The multimodal fusion intelligent inductive switch control system according to claim 7, characterized in that, The dynamic optimization method for the changing rhythm is as follows: the absolute values of the corrected angular frequencies of all oscillators are averaged to obtain the dynamically optimized changing rhythm.