Multi-mode interactive intelligent window based on AI and control method
By using a multimodal interactive smart window, combined with a speaker array, microphone array, and adaptive control algorithm, the problem of poor low-frequency noise isolation effect of traditional soundproof windows is solved, achieving efficient noise reduction, voice interaction, and 3D sound effects, meeting the personalized needs of modern high-quality indoor sound environment.
Patent Information
- Application Number
- CN202511372016.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-12-23
AI Technical Summary
Traditional soundproof windows have limited effectiveness in isolating low-frequency noise. Existing ANC systems are not suitable for large spaces or building structures. Smart windows have not yet integrated functions such as active noise reduction, human-computer interaction, 3D sound reconstruction, and ultrasonic insect repellent, making it difficult to meet the personalized needs of modern high-quality indoor sound environments.
It employs a frequency-division speaker array, microphone array, voice interaction unit, active noise reduction unit, infrared depth camera unit, intelligent controller and data processing module, combined with FxLMS adaptive control algorithm and deep learning, to realize a multimodal interactive intelligent window, integrating active noise reduction, voice interaction, 3D audio reconstruction and ultrasonic insect repellent functions.
It achieves local and global noise reduction, dynamic tracking noise reduction, supports voice interaction and 3D sound effects, has a significant insect repellent effect, and improves indoor acoustic comfort and personalized control.
Smart Images

Figure CN121191484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of architectural acoustics and smart home control technology, and in particular to an AI-based multimodal interactive smart window and control method. Background Technology
[0002] In today's urban living environment, external noise, such as traffic, construction, and human voices, significantly impacts indoor living and working comfort. Traditional soundproof windows primarily rely on materials to block sound waves, offering limited effectiveness against low-frequency noise and failing to meet the demands of a high-quality living environment. In recent years, active noise control (ANC) technology has been increasingly applied to enclosed spaces such as headphones, vehicles, and air conditioners. It reduces noise through sound wave cancellation, possessing unparalleled low-frequency suppression capabilities compared to traditional materials. However, existing ANC systems are generally unsuitable for integrated applications in large spaces or building structures. Furthermore, with the development of smart home technology, voice control, indoor sound field optimization, and intelligent environmental adjustment have become research hotspots. While curtain systems and music windows with voice assistants exist on the market, intelligent window devices integrating active noise cancellation, human-computer interaction, 3D sound reconstruction, ultrasonic insect repellent, and target tracking systems are still lacking. Therefore, there is an urgent need for a compact, functionally integrated, highly noise-reducing, and user-friendly intelligent window to meet the personalized needs of modern high-quality indoor acoustic environments. Summary of the Invention
[0003] The purpose of this invention is to provide an AI-based multimodal interactive smart window and control method to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides an AI-based multimodal interactive smart window, including a frequency-division speaker array embedded inside the window frame, and a microphone array disposed on the edge of the window frame. The microphone array includes an error microphone array and a reference microphone array, with the reference microphone array disposed on the outdoor side of the window and the error microphone array disposed on the indoor side of the window.
[0005] The window frame is also equipped with a voice interaction unit, an active noise cancellation unit, an intelligent controller, an infrared depth camera unit, a power module, and a data processing module.
[0006] Preferably, the noise reduction unit includes a signal acquisition module, a controller array, a frequency-division speaker array, an error estimation module, and an adaptive weight update module;
[0007] The controller array consists of multiple controllers, each of which is an FIR structure;
[0008] A crossover loudspeaker array includes multiple crossover loudspeakers, all of which are embedded in the internal wall structure of the window frame and arranged symmetrically.
[0009] The active noise reduction unit adopts a multi-channel adaptive filtering structure.
[0010] Preferably, the active noise cancellation unit runs the FxLMS adaptive control algorithm, which dynamically generates canceling sound waves based on the reference signal and the error signal, and emits them through a frequency-division loudspeaker array to form a reverse sound field acting on the target area indoors.
[0011] The adaptive weight update module performs online learning based on the FxLMS algorithm, and updates the weight parameters of the control filter according to the result of the filtered reference signal and the gradient value of the backpropagation error.
[0012] The adaptive weight update module is equipped with a dynamic delay compensation mechanism and filter stability constraints, and the active noise reduction unit is also equipped with a frequency selection mechanism.
[0013] Preferably, the infrared depth camera unit acquires the three-dimensional position information of people in the room and outputs it to the data processing module;
[0014] The data processing module uses deep learning to map 3D position information to acoustic path control parameters in real time, dynamically updates the target noise reduction area, and controls the sound field to automatically focus following the moving target.
[0015] Preferably, the infrared depth camera unit includes an indoor target recognition and 3D modeling system, which collects thermal imaging data and depth map information of the human body and extracts the three-dimensional coordinates, motion trajectory and posture information of the person in space.
[0016] The 3D modeling system processes image data in real time through a deep learning network, outputs the coordinates of joint nodes, the center position of the character and the direction of movement, matches the acoustic path database, dynamically selects crossover speakers and adjusts the output weight and phase difference of each crossover speaker.
[0017] Preferably, the voice interaction unit includes an audio synthesis module, a music playback module, and an insect repellent module, and the voice interaction unit embeds a speech recognition model based on the Transformer structure.
[0018] The audio synthesis module controls the crossover speaker array to output music, white noise, or ultrasonic signals of 20–65 kHz.
[0019] The insect repellent module is equipped with an insect repellent signal generator, a biological frequency adaptation database, an insect repellent status indicator, and a safety linkage mechanism. The insect repellent signal generator generates continuous waves, swept waves, frequency-modulated waves, or pulse-modulated waves of different frequencies within the range of 20kHz to 65kHz. The biological frequency adaptation database automatically matches common insects according to the current season and geographical region, and dynamically adjusts the insect repellent waveform and transmission power. The insect repellent status indicator and safety linkage mechanism automatically reduce the transmission power or pause the output when the window is open or people approach.
[0020] The music playback module includes a sound source storage unit, an audio synthesizer, an HRTF filter, and an echo space construction unit. The sound source storage unit stores music files, white noise, and a library of natural ambient sounds. The HRTF filter selects preset head-related transfer function data based on the user's position and orientation in the room to filter the stereo signal and construct a spatial virtual sound source. The audio synthesizer generates waveforms to synthesize natural sounds with adjustable parameters.
[0021] Preferably, the intelligent controller includes a main control MCU, a DSP, and a low-power wake-up controller, and the intelligent controller has a built-in resource scheduler.
[0022] Preferably, the window frame is made of aerospace-grade aluminum alloy or carbon fiber composite structure, and the window is a double-glazed vacuum glass structure.
[0023] The window frame has modular mounting slots and a reserved embedded channel. All wiring uses shielded cables, and there is an electrical isolation zone between the power supply and audio modules.
[0024] A control method for a multimodal interactive smart window based on AI includes the following steps:
[0025] S1. Acquire multi-source heterogeneous sensing information, including: a reference microphone array continuously collects external environmental noise signals, an error microphone array collects indoor sound field error signals, an infrared depth camera unit acquires the three-dimensional position, motion trajectory and posture information of indoor personnel, and a voice interaction unit receives and recognizes user voice commands.
[0026] S2. The data processing module performs spatiotemporal alignment, feature extraction and fusion processing on the multi-source heterogeneous sensing information in S1, and constructs an environmental perception model, which includes the sound field environment, user location and user intent.
[0027] S3, the intelligent controller makes distributed collaborative decisions based on the environmental perception model.
[0028] Preferably, the distributed collaborative decision-making in S3 is performed by the intelligent controller based on the environmental perception model, and the specific steps are as follows:
[0029] S31. The main control MCU parses the high-level semantics of the user's voice commands, switches the system working mode, and generates logic control commands.
[0030] S32 and DSP execute computationally intensive low-level acoustic signal processing algorithms. Based on the fused noise and error signals, they run the FxLMS algorithm to dynamically update the filter weights of the active noise control unit and generate anti-noise signals.
[0031] S33. Based on the personnel location information provided by the infrared depth camera module, call the acoustic path library, calculate the optimal output weight and phase difference of the frequency division loudspeaker array, perform dynamic sound field control and output, and realize dynamic focusing of the noise reduction area or sound field reconstruction of 3D audio.
[0032] S34. Generate specific ultrasonic insect repellent signals based on decision instructions.
[0033] Therefore, the present invention employs the above-mentioned AI-based multimodal interactive smart window and control method, which has the following beneficial effects:
[0034] (1) It has active noise reduction function, which can realize local noise reduction and global noise reduction. At the same time, combined with the infrared human body tracking system, it can realize dynamic tracking noise reduction.
[0035] (2) It has human-computer interaction function, which can realize dialogue with users, voice interaction, play music, news, stories, etc.;
[0036] (3) It can realize indoor 3D sound effects, form indoor 3D sound field, and enhance the indoor sound effect experience.
[0037] (4) It has an insect repellent effect. By emitting ultrasonic waves of different frequencies, it affects different types of insects and achieves the purpose of insect repellent.
[0038] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0039] Figure 1 This is a schematic diagram of the structure of an embodiment of the AI-based multimodal interactive smart window and control method of the present invention;
[0040] Figure 2 This is a schematic diagram of the active noise reduction unit control in an embodiment of an AI-based multimodal interactive smart window and control method according to the present invention.
[0041] Figure 3 This invention provides an embodiment of an AI-based multimodal interactive smart window and control method that utilizes a neural network model to establish an acoustic path library.
[0042] Figure 4This is an acoustic path matching flowchart of an embodiment of an AI-based multimodal interactive smart window and control method according to the present invention.
[0043] Figure 5 This is a flowchart illustrating an embodiment of the ultrasonic insect repellent method for a multimodal interactive smart window and control method based on AI according to the present invention.
[0044] Figure 6 This is a functional overview diagram of the intelligent controller in an embodiment of the AI-based multimodal interactive intelligent window and control method of the present invention;
[0045] Reference numerals: 1. Crossover speaker array; 2. Error microphone array; 3. Reference microphone array; 4. Voice interaction unit; 5. Active noise reduction unit; 6. Intelligent controller; 7. Infrared depth camera unit; 8. Power supply module; 9. Data processing module. Detailed Implementation
[0046] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0047] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0048] Example
[0049] Please see Figures 1-6 The present invention provides an AI-based multimodal interactive smart window, including a frequency-division speaker array 1 embedded inside the window frame, a microphone array set on the edge of the window frame, a voice interaction unit 4, an active noise reduction unit 5, a smart controller 6, an infrared depth camera unit 7, a power module 8, and a data processing module 9.
[0050] The crossover speaker array 1 is embedded in the internal wall structure of the window frame, arranged in a multi-channel symmetrical configuration. For example... Figure 1As shown, an embedded channel is reserved inside the frame of the smart window for the crossover speaker array 1 and the microphone array. Eight full-range, high-performance crossover speakers are evenly arranged inside the top, bottom, left, and right frames of the window. These crossover speakers have a frequency response capability of 50Hz–65kHz, meeting the combined requirements of active noise cancellation and ultrasonic output.
[0051] A reference microphone 3 is positioned on the outdoor edge of the window to collect ambient noise signals in real time; multiple error microphones 2 are positioned on the indoor edge to collect residual noise signals after active noise reduction. All microphones use MEMS miniature pickup units, which have high sensitivity and anti-interference capabilities.
[0052] The voice interaction unit 4 inside the room has a voice recognition function. The system has an embedded voice recognition model based on the Transformer structure. With the keyword wake-up module, it can recognize commands such as "noise reduction on" and "bedside area only".
[0053] The window frame is made of aerospace-grade aluminum alloy or carbon fiber composite structure, with modular mounting slots for embedding various functional components. The window glass is a double-glazed vacuum glass structure, providing basic passive sound insulation. Internal wiring uses shielded cables to prevent signal crosstalk; the crossover speaker openings are acoustically damped to prevent aesthetic damage and dust accumulation. The system features waterproof and anti-electrical interference designs, and all electroacoustic components have undergone EMC testing, complying with national building electronic product standards.
[0054] The window integrates multiple functions such as active noise control, voice interaction, 3D audio space reconstruction, ultrasonic insect repellent, human body tracking, and dynamic sound field control, forming a new type of intelligent building acoustic system. The design concept of this system is to transform traditional building windows from passive acoustic structures into controllable active sound field adjustment devices, realizing a leap in the function of building components from "noise isolation" to "noise elimination" and "environmental interaction".
[0055] The active noise cancellation unit 5 employs a multi-channel adaptive filtering algorithm, with the FxLMS algorithm as its core control strategy. It collects external ambient noise signals through reference microphones positioned on the outer edge of the window frame, and utilizes an array of crossover loudspeakers embedded inside the window frame to output interference signals with opposite phase and amplitude. These interference signals cancel out the original noise within the target area, significantly reducing the actual sound pressure level. The multi-channel adaptive FxLMS algorithm continuously adjusts the control filter weights based on the error signal, dynamically generating cancellation sound waves based on the reference and error signals. These waves are then emitted by the crossover loudspeaker array to cancel out the noise, forming a reverse sound field acting on the target area indoors, achieving effective low-frequency active noise cancellation control. The sampling frequency is set to 8–48 kHz, and the delay is controlled within 1 ms to achieve real-time noise reduction, adapting to low-frequency sources such as traffic noise and air conditioning noise. The specific process is as follows: Figure 2 As shown.
[0056] The FxLMS algorithm boasts strong real-time performance and adaptability, dynamically updating its control weights to accommodate changes in the indoor sound field. Error signals are collected by an array of error microphones deployed inside the window frame, and a feedback control algorithm is used to correct the filter weights, forming a closed-loop control system. Compared to traditional soundproof windows, this system offers significant advantages in the low-frequency range (<500Hz), making it particularly suitable for handling low-frequency noise sources such as traffic, air conditioning, and construction noise.
[0057] The active noise cancellation unit 5 includes a signal acquisition module, a controller array, a crossover speaker array 1, an error estimation module, and an adaptive weight update module. The controller array includes multiple controllers, each with an FIR structure. The crossover speaker array includes multiple sets of crossover speakers, all of which are embedded in the internal wall structure of the window frame and arranged symmetrically.
[0058] The active noise cancellation unit adopts a multi-channel adaptive filtering structure. Each channel contains several controllers and several cross-frequency speakers to form a cross-frequency speaker channel. Each cross-frequency speaker channel is equipped with a weighted controller to generate interference signals. The drive signal for each cross-frequency speaker is output by the controller.
[0059] The adaptive weight update module performs online learning based on the FxLMS algorithm. It updates the weight parameters of the control filter according to the result of the filtered reference signal and the gradient value of the backpropagation error. The module is equipped with a dynamic delay compensation mechanism and filter stability constraints to ensure that the control system can still converge and maintain stable output in a complex multi-path reflection environment.
[0060] The crossover speaker array 1 is driven by a power module and supports high-fidelity sound wave output with a sampling rate of over 65kHz; the active noise cancellation unit 5 is also equipped with a frequency selection mechanism, which allows users to specify the noise cancellation frequency band through voice commands or automatically identify the dominant noise frequency distribution and perform spectrum focusing.
[0061] To achieve spatial zoning noise reduction, the system constructs a virtual room sound field grid, supporting the division of a room into multiple acoustic sub-zones, such as office areas, sleeping areas, reading areas, and dining areas. After voice commands are recognized, the controller calls upon the corresponding crossover speaker and microphone combination for that sub-zone, applying active noise reduction only to that sub-domain, thereby reducing energy consumption and improving control precision. The system supports Mandarin Chinese, English, and other languages, and the voice recognition model can be upgraded via OTA.
[0062] The area-selective noise reduction and voice interaction mechanism allows users to personalize their noise reduction zones, such as applying noise reduction only to the "desk area," "sofa area," or "entire room." Users achieve this function through voice interaction. The voice recognition module collects user voice commands, and the system maps semantic commands to spatial control parameters. Using an internal spatial area division algorithm, the room is divided into multiple noise reduction sub-domains. Each sub-domain is associated with a specific crossover speaker-microphone combination, and the system activates the corresponding channel and controls it independently for that sub-domain. This approach achieves precise and energy-efficient noise reduction while allowing users to switch scenes as needed, enhancing the intelligent experience.
[0063] An infrared depth camera unit 7 is positioned on the right side of the window to collect real-time human motion trajectories and spatial posture information within the room. The image data is processed using YOLOv5, DepthNet, and Pose Estimation networks to extract human skeleton points and coordinate positions. Combined with a 3D acoustic modeling system for the room, an acoustic path library is established using neural networks. The specific process is as follows: Figure 3 As shown.
[0064] The infrared depth camera unit 7 is used to acquire the three-dimensional position information of people in the room and output it to the data processing module. This module uses deep learning to map the human body's spatial position to the acoustic path control parameters in real time, dynamically update the target noise reduction area, and control the sound field to automatically focus on the moving target.
[0065] The infrared depth camera unit 7 integrates an infrared thermal imaging camera and a depth camera, enabling real-time tracking of human movement within a room. Through an indoor target recognition and 3D modeling system, it collects thermal imaging data and depth map information of the human body. The image data is processed by a data processing module to identify the target person, extract key skeletal nodes, and calculate the person's 3D coordinates, movement trajectory, and posture information in space. The system processes the image data in real-time using a deep learning network, outputting joint node coordinates, the person's center position, and direction of movement. This data is combined with acoustic path data collected experimentally to build an acoustic path library through neural network learning, matching appropriate acoustic path databases. Based on the person's location and the room's acoustic path model, the system dynamically selects the crossover speaker channels and adjusts the output weights and phase differences of each crossover speaker, focusing sound interference on the person's location and dynamically constructing an acoustic focusing area. The controller automatically switches the crossover speaker array combination based on real-time location information, ensuring the active noise cancellation area dynamically adjusts within a 1m radius around the user's body, truly achieving intelligent acoustic adaptation with "person-to-sound" functionality for human body tracking control. This feature can be used in various scenarios such as mobile work, home fitness, and children's learning, significantly improving the individual auditory experience. The system has learning capabilities and can generate sound field response strategy templates based on the user's daily habits, enabling predictive switching.
[0066] The human body tracking control of the infrared depth camera unit 7 can also be linked with the voice interaction unit 4. Users can issue commands, and the system will automatically switch the target area according to the current location and area label.
[0067] The infrared depth camera module can be activated by voice. Based on the real-time position of the human body, the acoustic paths of the frequency-division loudspeakers and microphones are matched. The control system dynamically adjusts the output of the frequency-division loudspeaker array accordingly, focusing the noise reduction field within a 1-meter radius around the human body, achieving dynamic noise reduction. If the human body enters a new area, the system automatically switches the voice-controlled path after a 0.5-second delay to prevent discomfort caused by sudden changes in the sound field. The specific process is as follows... Figure 4 As shown.
[0068] The voice interaction unit 4 includes an audio synthesis module, a music playback module, and an insect repellent module, and the voice interaction unit embeds a speech recognition model based on the Transformer structure. The voice interaction unit 4 supports voice command switching between noise reduction mode, insect repellent mode, music playback mode, etc. The system controls the crossover speaker array to output music, white noise, or ultrasonic signals of 20-65kHz through the audio synthesis module.
[0069] The music playback module includes a local audio source storage unit, an audio synthesizer, an HRTF filter, and an echo space construction unit; among which, the audio source storage unit is used to store local or cloud-synchronized music files, white noise, and natural ambient sound material libraries.
[0070] The HRTF filter selects preset head-related transfer function data based on the user's position and orientation in the room to filter the stereo signal and construct a virtual sound source in space. The audio synthesizer also supports waveform generation to synthesize natural sounds with adjustable parameters, such as rain sounds, stream sounds, and insect chirps. After the synthesized signal is played through a frequency-division speaker array, it forms a superposition of real spatial propagation paths. Combined with room reverberation parameters, the sound field is adjusted to construct an immersive natural sound environment. The system supports voice command wake-up and scene linkage, which can start soft white noise playback and automatically switch the window noise reduction area to the bedside.
[0071] The 3D music playback function features 3D sound reconstruction and natural sound playback. To provide a high-quality auditory experience, this invention introduces a 3D spatial audio reconstruction module. The system has a built-in HRTF database and, based on the user's position and orientation indoors, simulates spatial awareness by controlling the output phase difference and sound pressure level of the crossover speaker array. Natural sounds such as waterfalls, insect chirps, and rain are generated by the system's embedded white noise generator or environmental sound library, and the immersive experience is enhanced through a reverberation algorithm. Users can select sound effects via voice commands. When the system is playing music, the active noise cancellation module enters a shared control mode, automatically reducing background noise to make the music sound clearer and smoother. Through the built-in HRTF database, combined with the user's spatial orientation and the position of the crossover speaker array, the system can preprocess the audio signal to reproduce the direction and distance of spatial sound sources. By simulating the time difference and intensity difference of sound waves at the ear, the user obtains an immersive auditory experience. In addition, the system supports playing white noise and natural soundscapes such as rain, ocean waves, and insect chirps to create a relaxing atmosphere, improve concentration, or aid in falling asleep. Sound effects can be activated via voice commands.
[0072] The crossover speaker 1 not only features noise control and music playback functions, but also utilizes a high-frequency response crossover speaker for insect repellent functionality. The insect repellent module achieves its function through the crossover speaker 1. This module includes an insect repellent signal generator, which, through the controller, can generate continuous waves, swept waves, frequency-modulated waves, or pulse-modulated waves of different frequencies within the 20kHz to 65kHz range. These waves interfere with the nervous system of insects, effectively repelling mosquitoes, moths, cockroaches, and other pests. The swept-frequency mode prevents pests from adapting to fixed frequencies, while the pulse mode enhances the energy impact effect. The system employs an adjustable swept-frequency method, automatically or according to user commands, changing the transmission frequency to adapt to different seasons and pest types. To ensure human safety, the output sound pressure level of the insect repellent system is controlled within a safe threshold, preventing auditory strain. This function is particularly suitable for summer nights, kitchens, bedrooms, and other similar environments, providing high comfort and safety.
[0073] The ultrasonic signal supports adaptive frequency switching to prevent insects from adapting, and also supports user voice commands or timed wake-up. The specific process is as follows: Figure 5As shown. The insect repellent signal can be controlled by a preset timer or by the user via voice commands. The system further includes a biological frequency adaptation database, which can automatically match the nerve-sensitive frequency bandwidth of common insects such as mosquitoes, cockroaches, and moths according to the current season and geographical region, and dynamically adjust the insect repellent waveform and transmission power. When the crossover speaker is working in insect repellent mode, its output sound pressure level is controlled between 90 and 110 dB and remains in the inaudible frequency band to prevent auditory interference to adults, infants, or pets. The system also has an insect repellent status indicator light and a safety linkage mechanism, which automatically reduces the transmission power or pauses the output when the window is open or people approach.
[0074] The intelligent controller 6 adopts a modular multi-processor architecture, including a main control MCU, a DSP, and a low-power wake-up controller. The MCU is responsible for speech recognition, human-computer interaction, and logic control; the DSP handles high-bandwidth tasks such as active noise cancellation, audio reconstruction, and ultrasonic output; and the low-power controller is used for standby monitoring and mode switching management. The core control module of the entire system consists of a multi-core MCU and a DSP. The MCU is responsible for coordinating and scheduling various modules, speech parsing, and system status monitoring; the DSP is used to perform high-bandwidth ANC operations and audio rendering. A detailed functional diagram is shown below. Figure 6 As shown.
[0075] The entire system is powered by an integrated embedded power module. This power module 8 supports three power supply methods: AC power input, solar-assisted charging, and lithium battery energy storage. It also features charging protection, power indication, and energy consumption statistics.
[0076] The control system operates based on an event-driven strategy, automatically switching system modes such as noise reduction mode, music playback mode, insect repellent mode, or standby mode based on multi-dimensional input signals such as voice commands, ambient light, personnel approach, and time schedules.
[0077] The system control circuit, window-type frequency-division speaker array, microphone array, and power lines are arranged in the wiring trough inside the window frame structure. The wiring adopts a shielded structure to prevent electromagnetic interference, and an electrical isolation zone is provided between the power supply and the audio module to improve system stability and safety. This claim significantly improves the stability, energy saving and intelligence of system operation through the integrated design of multi-source power supply, efficient scheduling and low-power intelligent control, which facilitates long-term operation of the product in complex building environments.
[0078] The control system features a multi-functional collaborative operation mechanism, supporting non-interference or timing-switching collaborative control strategies between modules. A built-in resource scheduler allocates processor resources and audio output bandwidth to prevent system response delays or acoustic interference when multiple functions are invoked simultaneously.
[0079] Intelligent voice-controlled windows comprehensively upgrade building windows in terms of sound insulation, active noise reduction, environmental interaction, and pest control, significantly improving the acoustic comfort and functionality of the living environment.
[0080] This invention also provides a control method for a multimodal interactive smart window based on AI, which adopts a multimodal fusion and distributed active control intelligent control method: the system continuously collects external environmental noise signals through a reference microphone array 3, collects indoor sound field error signals through an error microphone array 2, acquires the three-dimensional position, motion trajectory and posture information of indoor personnel through an infrared depth camera unit 7, and receives and recognizes user voice commands through a voice interaction unit 4; the data processing module 9 performs spatiotemporal alignment, feature extraction and fusion processing on the above multi-source heterogeneous sensing information to construct a unified environmental perception model that includes sound field environment, user position and user intent;
[0081] Distributed active control uses the intelligent controller 6 to make distributed collaborative decisions based on the environmental perception model: the main control MCU is responsible for parsing the high-level semantics of the user's voice commands and switching the system's working mode (noise reduction, music, insect repellent) accordingly, generating logic control commands; the DSP is responsible for executing the computationally intensive low-level acoustic signal processing algorithm, running the FxLMS algorithm based on the fused noise signal and error signal to dynamically update the filter weights of the active noise reduction unit 5 and generate anti-noise signals; based on the personnel position information provided by the infrared depth camera unit 7, the acoustic path library is called to calculate the optimal output weights and phase differences of each channel of the frequency division speaker array 1, realizing dynamic focusing of the noise reduction area or 3D audio sound field reconstruction; and specific ultrasonic insect repellent signals are generated according to the decision commands.
[0082] Dynamic sound field control and output are achieved through a crossover speaker array 1 as a shared actuator. Based on the acoustic signals calculated by the DSP, different functions are performed: the active noise control module outputs an anti-noise wave with the opposite phase to the external noise, creating a quiet zone in the user's area; the music playback module outputs an audio signal reconstructed by the spatial sound field based on the HRTF filter and the user's position, creating an immersive listening experience; the insect repellent module outputs a modulated ultrasonic signal of 20kHz-65kHz to repel pests, and the sound pressure level of the crossover speaker output is controlled between 90 and 110dB, and kept in the inaudible frequency range.
[0083] Intelligent resource scheduling and safety management utilize a built-in resource scheduler to uniformly schedule and manage the processor's computing resources and the output channels of the crossover speaker array. When multiple functions are requested simultaneously, the system switches between time sequences or allocates resources based on preset priorities or user commands to avoid interference between sound signals. The system continuously monitors the window status and the proximity of personnel, and automatically adjusts the output power or suspends output in abnormal situations through a safety linkage mechanism to ensure human and machine safety.
[0084] Therefore, this invention employs the aforementioned AI-based multimodal interactive smart window and control method, comprehensively realizing a multi-functional integrated smart window system encompassing active noise reduction, voice interaction, 3D music playback, ultrasonic insect repellent, and infrared human body tracking. Based on speech recognition and deep learning algorithms, it achieves regionalized noise reduction, contextual music rendering, and intelligent control of the indoor acoustic environment. This invention significantly improves the acoustic comfort and functionality of living environments and can be widely applied in buildings requiring high-quality acoustic environments, such as residences, office buildings, hospitals, and libraries, possessing broad market prospects.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multimodal interactive smart window based on AI, characterized in that: It includes a crossover speaker array, which is embedded inside the window frame. A microphone array is set on the edge of the window frame. The microphone array includes an error microphone array and a reference microphone array. The reference microphone array is located on the outside side of the window, and the error microphone array is located on the inside side of the window. The window frame is also equipped with a voice interaction unit, an active noise cancellation unit, an intelligent controller, an infrared depth camera unit, a power module, and a data processing module.
2. The AI-based multimodal interactive smart window according to claim 1, characterized in that: The active noise cancellation unit includes a signal acquisition module, a controller array, a frequency-division speaker array, an error estimation module, and an adaptive weight update module; The controller array consists of multiple controllers, each of which is an FIR structure; A crossover loudspeaker array includes multiple crossover loudspeakers, all of which are embedded in the internal wall structure of the window frame and arranged symmetrically. The active noise reduction unit adopts a multi-channel adaptive filtering structure.
3. The AI-based multimodal interactive smart window according to claim 2, characterized in that: The active noise cancellation unit runs the FxLMS adaptive control algorithm, which dynamically generates canceling sound waves based on the reference signal and error signal, and emits them through a crossover speaker array to form a reverse sound field acting on the target area indoors. The adaptive weight update module performs online learning based on the FxLMS algorithm, and updates the weight parameters of the control filter according to the result of the filtered reference signal and the gradient value of the backpropagation error. The adaptive weight update module is equipped with a dynamic delay compensation mechanism and filter stability constraints, and the active noise reduction unit is also equipped with a frequency selection mechanism.
4. The AI-based multimodal interactive smart window according to claim 1, characterized in that: The infrared depth camera unit acquires the three-dimensional position information of people in the room and outputs it to the data processing module; The data processing module uses deep learning to map 3D position information to acoustic path control parameters in real time, dynamically updates the target noise reduction area, and controls the sound field to automatically focus following the moving target.
5. The AI-based multimodal interactive smart window according to claim 4, characterized in that: The infrared depth camera unit includes an indoor target recognition and 3D modeling system. The indoor target recognition system collects thermal imaging data and depth map information of the human body and extracts the three-dimensional coordinates, motion trajectory and posture information of the person in space. The 3D modeling system processes image data in real time through a deep learning network, outputs the coordinates of joint nodes, the center position of the character and the direction of movement, matches the acoustic path database, dynamically selects crossover speakers and adjusts the output weight and phase difference of each crossover speaker.
6. The AI-based multimodal interactive smart window according to claim 1, characterized in that: The voice interaction unit includes an audio synthesis module, a music playback module, and an insect repellent module, and the voice interaction unit embeds a speech recognition model based on the Transformer structure. The audio synthesis module controls the crossover speaker array to output music, white noise, or ultrasonic signals of 20–65 kHz. The insect repellent module is equipped with an insect repellent signal generator, a biological frequency adaptation database, an insect repellent status indicator light, and a safety linkage mechanism. The insect repellent signal generator generates continuous waves, swept waves, frequency-modulated waves, or pulse-modulated waves of different frequencies in the range of 20kHz to 65kHz. The biological frequency adaptation database automatically matches common insects based on the current season and geographical region, and dynamically adjusts the insect repelling waveform and transmission power; the insect repelling status indicator light and safety linkage mechanism automatically reduce the transmission power or pause the output when the window is open or people approach. The music playback module includes a sound source storage unit, an audio synthesizer, an HRTF filter, and an echo space construction unit. The sound source storage unit stores music files, white noise, and a library of natural ambient sounds. The HRTF filter selects preset head-related transfer function data based on the user's position and orientation in the room to filter the stereo signal and construct a spatial virtual sound source. The audio synthesizer generates waveforms to synthesize natural sounds with adjustable parameters.
7. The AI-based multimodal interactive smart window according to claim 1, characterized in that: The intelligent controller includes a main control MCU, a DSP, and a low-power wake-up controller, and the intelligent controller has a built-in resource scheduler.
8. The AI-based multimodal interactive smart window according to claim 1, characterized in that: The window frame is made of aerospace-grade aluminum alloy or carbon fiber composite structure, and the window is a double-glazed vacuum glass structure. The window frame has modular mounting slots and a reserved embedded channel. All wiring uses shielded cables, and there is an electrical isolation zone between the power supply and audio modules.
9. A control method for an AI-based multimodal interactive smart window according to any one of claims 1-8, characterized in that, Includes the following steps: S1. Acquire multi-source heterogeneous sensing information, including: a reference microphone array continuously collects external environmental noise signals, an error microphone array collects indoor sound field error signals, an infrared depth camera unit acquires the three-dimensional position, motion trajectory and posture information of indoor personnel, and a voice interaction unit receives and recognizes user voice commands. S2. The data processing module performs spatiotemporal alignment, feature extraction and fusion processing on the multi-source heterogeneous sensing information in S1, and constructs an environmental perception model, which includes the sound field environment, user location and user intent. S3, the intelligent controller makes distributed collaborative decisions based on the environmental perception model.
10. The control method for a multimodal interactive smart window based on AI according to claim 9, characterized in that, The distributed collaborative decision-making in S3 is performed by the intelligent controller based on the environmental perception model, and the specific steps are as follows: S31. The main control MCU parses the high-level semantics of the user's voice commands, switches the system working mode, and generates logic control commands. S32 and DSP execute computationally intensive low-level acoustic signal processing algorithms. Based on the fused noise and error signals, they run the FxLMS algorithm to dynamically update the filter weights of the active noise control unit and generate anti-noise signals. S33. Based on the personnel location information provided by the infrared depth camera module, call the acoustic path library, calculate the optimal output weight and phase difference of the frequency division loudspeaker array, perform dynamic sound field control and output, and realize dynamic focusing of the noise reduction area or sound field reconstruction of 3D audio. S34. Generate specific ultrasonic insect repellent signals based on decision instructions.