A smart start / stop system for audio and video intercom speakers that integrates voice wake-up

CN122575354APending Publication Date: 2026-08-14SHENZHEN ZHILIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]目前,现有音视频对讲系统中的音频喇叭多采用持续待机或手动启停的控制方式,一方面,为确保对讲请求能够被及时响应,喇叭通常需要长期处于待机发声状态,这不仅会造成大量不必要的电能消耗,还会因长期持续工作导致喇叭音圈发热、绝缘层老化,显著缩短喇叭使用寿命,增加设备维护成本,另一方面,手动启停方式操作繁琐,无法根据实际对讲需求实现精准响应,在无人值守场景或双手忙碌状态下,用户难以快速控制喇叭启停,严重影响交互便捷性,为此,本发明提供一种融合语音唤醒的音视频对讲音频喇叭智能启停系统

Benefits of technology

1.本发明所述的一种融合语音唤醒的音视频对讲音频喇叭智能启停系统,通过意图感知层实现多维度对讲意图超前预判,结合相位差和能量比双因子信号分离算法,可精准隔离本地唤醒信号与远端对讲干扰信号,从根源规避音视频对讲过程中远端播放声音、环境嘈杂噪声引发的误唤醒问题,同时依托四级事件优先级调度机制,实现紧急事件优先响应、常规事件有序调度、无效事件抑制拦截,配合预控、跟随、恢复三段式喇叭时序控制,消除喇叭启停瞬时电流冲击与爆破杂音,大幅缩短对讲响应时延,保证各类场景下语音唤醒识别准确率与对讲通话清晰度,显著提升人机语音交互的智能化与流畅度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575354A_ABST
    Figure CN122575354A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of audio-visual intercom and voice wake-up technology. Specifically, it is an intelligent start-stop system for audio speakers in audio-visual intercom that integrates voice wake-up. It includes an intent perception layer module, which has a built-in intercom intent predictor, a two-way signal separator, a scene intent matcher, and an access control engine. As the front-end multi-source information perception and pre-decision entry point of the system, it is responsible for real-time acquisition of ambient sound field signals and audio-visual intercom link signals, and completes the isolation and separation of local voice wake-up signals and remote intercom signals. At the same time, it integrates voiceprint features, usage context, historical behavior, application scenarios, and user identity and permissions information to predict the user's intercom behavior intent in advance and output intent confidence, scene type, and access control results. This provides a reliable perception data source, intent decision basis, and scene adaptation parameters for the subsequent priority scheduling layer and state pre-control layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of audio and video intercom and voice wake-up technology, specifically an intelligent start-stop system for audio and video intercom speakers that integrates voice wake-up. Background Technology

[0002] With the acceleration of global digitalization, the widespread adoption of 5G networks, and the increasing maturity of artificial intelligence algorithms, smart audio devices are evolving from simple playback tools into all-around smart interactive terminals.

[0003] Audio and video intercom technology, as a convenient two-way communication method, has been widely used in smart homes, security monitoring, industrial communication, smart communities and other fields, becoming a key interactive carrier connecting the physical world and digital services. In these application scenarios, the audio speaker, as the core sound-generating component of the audio and video intercom system, directly determines the communication experience, device power consumption and service life.

[0004] Currently, most audio speakers in existing audio-visual intercom systems use continuous standby or manual start / stop control methods. On the one hand, to ensure that intercom requests can be responded to in a timely manner, the speakers usually need to be in a standby state for a long time. This not only causes a lot of unnecessary power consumption, but also causes the speaker voice coil to heat up and the insulation layer to age due to long-term continuous operation, significantly shortening the speaker's service life and increasing equipment maintenance costs. On the other hand, the manual start / stop method is cumbersome to operate and cannot achieve accurate response according to actual intercom needs. In unattended scenarios or when both hands are busy, users find it difficult to quickly control the speaker to start or stop, which seriously affects the convenience of interaction. Therefore, this invention provides an intelligent start / stop system for audio speakers in audio-visual intercom systems that integrates voice wake-up. Summary of the Invention

[0005] In order to overcome the shortcomings of the prior art, at least one technical problem raised in the background art is solved.

[0006] The technical solution adopted by this invention to solve its technical problem is: a smart start / stop system for audio and video intercom speakers that integrates voice wake-up, comprising: The intent perception layer module, through its built-in intercom intent predictor, two-way signal separator, scene intent matcher, and permission verification engine, and as the system's front-end multi-source information perception and pre-decision entry point, is responsible for real-time acquisition of environmental sound field signals and audio-visual intercom link signals, completing the isolation and separation of local voice wake-up signals and remote intercom signals. At the same time, it integrates voiceprint features, usage context, historical behavior, application scenarios, and user identity and permissions multi-dimensional information to predict the user's intercom behavior intent in advance and output intent confidence, scene type, and permission judgment results, providing a reliable perception data source, intent decision basis, and scene adaptation parameters for the subsequent priority scheduling layer and state pre-control layer. The priority scheduling layer module is configured with a multi-level wake-up event scheduling mechanism and a dynamic matching algorithm for intent and resources. It performs hierarchical resource allocation and action scheduling for wake-up events based on priority levels, and dynamically allocates system and speaker operating resources in combination with intent confidence, scene adaptability, and permission level coefficient. The state pre-control layer module adopts a three-segment speaker control architecture and dynamic power rail unit. After anticipating the intercom intention, it enters the pre-control stage in advance to preheat the speaker bias circuit. After the intercom is established, it dynamically adjusts the speaker output power and amplitude in real time according to the intercom voice characteristics. Before the intercom ends, it enters the recovery stage in advance to complete the power drop-off in an exponential manner. The speaker power supply voltage is dynamically adjusted through a multi-level switchable power module, and a power recovery circuit is configured to recover and utilize the speaker's idle back electromotive force. The speaker status monitoring and feedback module collects speaker voice coil temperature, impedance, and diaphragm vibration parameters in real time and transmits them back to the priority scheduling layer module and the status pre-control layer module. Combined with real-time parameters, it dynamically corrects the wake-up scheduling strategy and speaker start-stop control logic, realizing audio speaker seamless intelligent start-stop, anti-false wake-up, low-power operation, and overload self-protection, with intercom intent prediction, event hierarchical scheduling, and hardware status closed-loop perception.

[0007] Preferably, the intercom intent predictor in the intent perception layer module adopts a lightweight temporal attention model, builds a personalized intent model based on the user's historical intercom behavior, combines real-time sound field features to predict the user's intercom behavior in advance, and outputs the intent confidence level within the interval range.

[0008] Preferably, the bidirectional signal separator in the intent perception layer module adopts a phase difference and energy ratio dual-factor separation algorithm, and achieves accurate isolation and filtering of local wake-up signals and remote intercom signals by combining microphone array sound source phase locking with dynamic energy threshold determination.

[0009] Preferably, the scene intent matcher in the intent perception layer module has a built-in scene and intercom mode mapping library, which can automatically identify noisy application scenarios such as home, office, and outdoor. The permission verification engine adopts a voiceprint and permission dynamic binding mechanism to divide intercom access permissions into three levels: administrator, ordinary user, and visitor.

[0010] Preferably, the multi-level wake-up event scheduling mechanism of the priority scheduling layer module is specifically divided into four event levels: emergency, high priority, medium priority, and low priority. Differentiated system resource quotas, response times, and execution action logic are configured for different event levels.

[0011] Preferably, the intention and resource dynamic matching algorithm of the priority scheduling layer module uses intention confidence, scenario adaptability and permission level coefficient as correlation factors to calculate and allocate system running resources and loudspeaker working resource quotas in a coordinated manner.

[0012] Preferably, the three-stage speaker control architecture of the state pre-control layer module is a pre-control, follow-up, and recovery time-series closed-loop architecture. In the pre-control stage, the speaker bias circuit is activated in advance to preheat after the intercom intention is predicted to meet the standard. In the follow-up stage, the speaker operating condition is dynamically adjusted to adapt to the intercom voice characteristics. In the recovery stage, the speaker output power is gradually reduced by exponential decay before the intercom ends.

[0013] Preferably, the dynamic power rail unit of the state pre-control layer module is configured with a multi-level switchable power module, which supports adaptive switching of multiple power supply voltages to match the speaker power output requirements under different scenarios.

[0014] Preferably, the power recovery circuit built into the state pre-control layer module is used to collect the reverse electromotive force generated in the speaker's idle state and reuse the recovered electrical energy to the system's low-power standby power supply circuit.

[0015] Preferably, the speaker status monitoring and feedback module collects the speaker voice coil operating temperature, real-time impedance value, and diaphragm vibration amplitude parameters in real time, and transmits the collected parameters back in real time to form a closed-loop feedback, dynamically correcting the wake-up event scheduling strategy and the speaker start / stop and power control logic.

[0016] The beneficial effects of this invention are as follows: 1. The present invention discloses an intelligent start-stop system for audio speakers in audio-visual intercom that integrates voice wake-up. Through an intent perception layer, it achieves multi-dimensional advance prediction of intercom intent. Combined with a dual-factor signal separation algorithm based on phase difference and energy ratio, it can accurately isolate local wake-up signals from remote intercom interference signals, fundamentally avoiding false wake-up problems caused by remote sound playback and environmental noise during audio-visual intercom. Simultaneously, relying on a four-level event priority scheduling mechanism, it achieves priority response to emergency events, orderly scheduling of routine events, and suppression and interception of invalid events. Combined with a three-stage speaker timing control system (pre-control, follow-up, and recovery), it eliminates instantaneous current surges and popping noises during speaker start-up and shutdown, significantly shortening intercom response latency. This ensures accurate voice wake-up recognition and clear intercom calls in various scenarios, significantly improving the intelligence and fluency of human-computer voice interaction.

[0017] 2. The intelligent start-stop system for audio and video intercom speakers that integrates voice wake-up, as described in this invention, abandons the traditional crude mode of long-term standby and constant power operation of speakers. It uses an intention and resource dynamic matching algorithm to allocate system and speaker operating resources on demand. With the help of a multi-level dynamic power rail unit, the power supply voltage is adaptively switched according to the ambient noise, realizing intelligent adaptation to low-voltage and low-power operation in quiet scenarios and high-voltage and full-sound output in high-noise scenarios. At the same time, the built-in power recovery circuit can recover the back electromotive force generated during the switching of speaker operating conditions. After energy storage and voltage stabilization, it is reused in the system's low-power standby circuit, which greatly reduces the static standby power consumption and average operating power consumption of the whole machine, effectively extending the battery life of battery-powered audio and video intercom devices and meeting the requirements of low-power and energy-saving design. Attached Figure Description

[0018] The invention will now be further described with reference to the accompanying drawings.

[0019] Figure 1 This is a schematic diagram of the overall hierarchical closed-loop workflow in this invention; Figure 2 This is a flowchart of the three-stage horn intelligent start / stop control process of the state pre-control layer module in this invention. Detailed Implementation

[0020] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0021] like Figure 1-2 As shown in the embodiment of the present invention, an intelligent start / stop system for audio and video intercom speakers integrating voice wake-up includes: The intent perception layer module incorporates an intercom intent predictor, a two-way signal separator, a scene intent matcher, and an authorization verification engine. The intercom intent predictor uses a lightweight temporal attention model to predict user intercom intent based on voiceprint, context, and historical behavior data. The two-way signal separator uses a phase difference and energy ratio dual-factor separation algorithm to isolate local wake-up signals from remote intercom signals in real time. The scene intent matcher has a built-in scene and intercom mode mapping library to identify the current application scenario and match speaker operating parameters. The authorization verification engine verifies user intercom identity and authorization through dynamic binding of voiceprint and authorization. As the system's front-end multi-source information perception and pre-decision entry point, it is responsible for real-time acquisition of environmental sound field signals and audio / video intercom link signals, completing the isolation and separation of local voice wake-up signals and remote intercom signals. It also integrates voiceprint features, usage context, historical behavior, application scenarios, and user identity and authorization information to predict user intercom behavior intent in advance and output intent confidence, scene type, and authorization judgment results. This provides a reliable perception data source, intent decision basis, and scene adaptation parameters for the subsequent priority scheduling layer and state pre-control layer. The priority scheduling layer module is configured with a multi-level wake-up event scheduling mechanism and a dynamic matching algorithm for intent and resources. It performs hierarchical resource allocation and action scheduling for wake-up events based on priority levels, and dynamically allocates system and speaker operating resources in combination with intent confidence, scene adaptability, and permission level coefficient. The state pre-control layer module adopts a three-segment speaker control architecture and dynamic power rail unit. After anticipating the intercom intention, it enters the pre-control stage in advance to preheat the speaker bias circuit. After the intercom is established, it dynamically adjusts the speaker output power and amplitude in real time according to the intercom voice characteristics. Before the intercom ends, it enters the recovery stage in advance to complete the power drop-off in an exponential manner. The speaker power supply voltage is dynamically adjusted through a multi-level switchable power module, and a power recovery circuit is configured to recover and utilize the speaker's idle back electromotive force. The speaker status monitoring and feedback module collects core hardware parameters such as speaker voice coil operating temperature, AC equivalent impedance, diaphragm vibration amplitude and vibration frequency in real time with a fixed sampling period. After filtering, denoising and normalizing the collected raw sensor data, it is sent back to the priority scheduling layer module and the status pre-control layer module in a closed loop. The system dynamically iteratively corrects the wake-up event scheduling priority strategy, voice wake-up sensitivity threshold, speaker start-stop timing parameters and power control logic based on the speaker's real-time health status parameters. Ultimately, it achieves intelligent adaptive operation effects such as audio speaker imperceptible intelligent start-stop, strong resistance to environmental and remote voice false wake-up, ultra-low power consumption operation at all times, and self-protection against electrical overheating and impedance abnormality overload, based on the advance prediction of intercom intention, multi-event hierarchical scheduling, and closed-loop perception of speaker status at all times. The overall hardware architecture uses a RISC-V dual-core embedded main control chip as the core processing unit, paired with a six-channel array microphone audio acquisition unit, a WebRTC / SIP dual-protocol audio and video intercom communication unit, a Class D digital power amplifier driver unit, a multi-level switchable dynamic power rail unit, a speaker composite status sensor acquisition unit, a low-power coprocessor monitoring unit, and a power recovery energy storage unit; peripheral components include a storage unit, a clock unit, a level conversion unit, an overcurrent and overvoltage protection unit, and wired / wireless communication interfaces, adapting to various application scenarios such as smart homes, smart community buildings, industrial intercoms, park security, and indoor video doorbells.

[0022] like Figure 1-2 As shown, the intercom intent predictor in the intent perception layer module is independently deployed in the system's low-power coprocessor, without occupying the main core's computing resources, ensuring ultra-low power operation for normal monitoring; the predictor is equipped with a lightweight temporal attention AI model, which adopts a lightweight network structure design, prunes redundant convolutional layers and fully connected layers, adapts to the computing power limitations of embedded edge devices, and can complete real-time inference operations on the local end without cloud dependence. First, a user-specific behavioral feature database is pre-loaded and constructed using the system's local non-volatile storage unit. This database includes historical behavioral data such as the user's daily intercom initiation times, commonly used fixed wake-up word sequences, average duration of a single intercom session, intercom trigger frequency, common interaction scenarios, and human voice spectral characteristics. The model uses a continuous multi-frame temporal audio feature sequence as input data and automatically assigns feature weights to different audio frames through a temporal attention mechanism. It deeply mines the implicit correlation between environmental sound field signal features and the user's actual intercom behavior, quantifies and calculates the user's intercom intent confidence in real time, and sets a dedicated formula for calculating intercom intent confidence. In the formula: This represents the normalized confidence level of the intercom intent, with a value range of 0–100%. To collect the matching scores between human voices and the local voiceprint database in real time, This is the matching coefficient between the current time and the user's commonly used intercom time slots. The correlation between real-time sound field characteristics and historical intercom behavior characteristics; , , The preset fixed weighted correction coefficients are used, and the normalization constraints are strictly satisfied. =1, the coefficient can be locally calibrated and adaptively fine-tuned according to different application scenarios; The system has a preset three-level confidence threshold: when the calculated intent confidence is ≥80%, it is determined to be a highly reliable and valid intercom intent, and a pre-trigger signal is immediately output to the downstream scheduling module and pre-control module to start the speaker pre-control preparation process; when the confidence is between 50% and 80%, it is determined to be a suspicious intent, and only the voice recognition detection accuracy is improved, without starting the speaker hardware pre-control; when the confidence is below 50%, it is directly determined to be an invalid interference intent, and the system maintains the original low-power sleep listening state without triggering any subsequent hardware actions or link establishment actions. At the same time, the predictor has an offline self-learning iteration function, which can automatically record the user's new intercom behavior data every day, periodically update the intent feature library and model weights, and continuously improve the intent prediction accuracy.

[0023] like Figure 1-2 As shown, the bidirectional signal separator in the intent perception layer module adopts a phase difference and energy ratio dual-factor collaborative separation algorithm. It relies on a spatial sound field acquisition array formed by a six-channel array microphone arranged in an equilateral circular layout. By utilizing the propagation delay and phase difference characteristics of the signals from the same sound source received by different array elements, it achieves accurate spatial positioning and feature differentiation of local near-field human voice and far-field intercom sound source played by a remote speaker. When the splitter is working, it first performs synchronous sampling, filtering, noise reduction, and echo cancellation preprocessing on the raw audio signals from the six microphones. It then extracts the phase characteristics, time-domain energy characteristics, and spectral characteristics of each audio signal in real time. On one hand, by calculating the signal phase difference between different microphone array elements, it establishes a local human voice phase characteristic threshold range; signals deviating from this phase range are directly identified as remote intercom source signals. On the other hand, it synchronously calculates the ratio of the instantaneous local human voice audio energy to the instantaneous energy of the remote intercom playback audio, constructs a dynamic adaptive energy threshold judgment condition, and sets a dedicated calculation formula for signal isolation judgment. In the formula: This is the instantaneous energy ratio of local voice to remote intercom audio. The average energy of local human voice in short time frames acquired in real time; The average energy of short-time frames of audio transmitted from the audio / video intercom link to the remote playback of the local speaker; The algorithm employs a dual-condition joint decision mechanism: when the phase difference of the sound source matches the local near-field human voice characteristics, and the energy ratio satisfies... If the system determines that the signal is a valid local wake-up command, it will allow the user to proceed to the subsequent intent recognition and permission verification process; if the sound source phase deviates from the local human voice characteristics, or the energy ratio is lower... It directly identifies remote intercom interference signals, environmental noise signals, or echo signals, and performs truncation and filtering on such signals, and shields the wake-up trigger channel, eliminating the problem of local false wake-up caused by remote intercom sound and environmental noise from the bottom layer. At the same time, the splitter can adaptively and dynamically fine-tune the phase tolerance threshold and energy ratio judgment threshold according to the real-time environmental noise intensity, adapting to different working conditions such as quiet home, noisy outdoor, and strong industrial interference.

[0024] like Figure 1-2 As shown, the scene intent matcher in the intent perception layer module has a pre-built expandable scene and intercom mode mapping knowledge base. It has four standard scene templates built in by default: home, office, outdoor public place, and industrial site. Each template includes multi-dimensional scene fingerprint features such as environmental noise decibel range, sound field reverberation time, background spectrum distribution, and time period behavior characteristics. The matcher uses a real-time clustering algorithm to perform similarity matching on the multi-dimensional features of the collected environmental audio and automatically identify the application scene of the current device. After identifying the scene, the matcher calls the mapping knowledge base to automatically match the appropriate system operating parameters, including the voice wake-up sensitivity level, the initial default volume of the speaker, the low power sleep delay threshold, the anti-interference filtering strength, and the complete set of working parameters for intercom communication code rate. At the same time, the knowledge base supports users to add custom scene templates locally, and can manually enter special environmental features and configure exclusive operating parameters to adapt to personalized application needs. The permission verification engine adopts a voiceprint and dynamic permission binding architecture. It pre-supports the offline local input of exclusive voiceprint features for three types of users: administrators, ordinary resident users, and temporary visitors. An encrypted voiceprint feature library is established and stored in a local secure partition. The engine compares and matches the real-time collected wake-up voice features with the voiceprint library dimension by dimension. Once a match is successful, the corresponding identity permission level is automatically bound, and differentiated operation permissions are granted. Administrator privileges: Possess full permissions for emergency wake-up command response, system parameter configuration, manual adjustment of speaker working mode, forced start and stop of intercom links, and permission allocation and management; Regular user permissions: Only supports regular voice wake-up to initiate intercom, answer remote calls, and basic volume adjustment permissions; no permissions to modify system parameters. Temporary visitor privileges: Only allowed to passively answer intercom calls initiated by remote devices, prohibited from actively initiating intercom operations via voice wake-up, and intercom duration is subject to an automatic time limit mechanism; Meanwhile, the engine has a built-in blacklist and whitelist management mechanism that can record unfamiliar voiceprint characteristics on the blacklist and directly block their wake-up and intercom access requests. It supports temporary visitor time-limited dynamic authorization and can set short-term valid intercom permissions, which are automatically revoked after the time expires, further improving the security and management flexibility of the audio and video intercom system.

[0025] like Figure 1-2 As shown, the priority scheduling layer module is configured with a multi-level wake-up event scheduling mechanism, which is uniformly and standardized into four fixed event levels: emergency, high priority, medium priority, and low priority. For each level of event, there are preset independent trigger judgment conditions, system resource allocation ratios, audio path scheduling logic, speaker action response rules, and cloud linkage strategies. At the same time, it has a built-in multi-event concurrent preemption and queuing scheduling mechanism. High-priority events can directly preempt system resources of low-priority events, and low-priority events enter the task queue for delayed processing. Emergency-level events: The trigger condition is the recognition of a preset emergency SOS wake-up word and the matching of the voiceprint characteristics of a legitimate administrator or homeowner. The system immediately allocates 100% of the machine's computing power resources, full-bandwidth audio channels and communication resources, instantly starts the speaker at full speed and enters full-power working state, and simultaneously links local security alarm output, cloud platform alarm reporting, and related terminal message push, prioritizing rapid response to emergency scenarios; High-priority events: The triggering conditions are that the confidence level of the intercom intent predicted by the intent perception layer reaches the standard and a legitimate resident user actively initiates local voice wake-up; the system allocates 80% of the core computing power and audio resources to trigger the state pre-control layer to enter the speaker pre-control pre-warming process in advance, and simultaneously and quickly establishes the audio and video intercom link to shorten the intercom connection response latency. Medium priority event: The trigger condition is that the external remote intercom device, access control terminal, or mobile client actively initiates an intercom call request. The system allocates 50% of the normal operating resources, wakes up the speaker in the standard sequence to be ready, and normally accesses the intercom communication channel to maintain the normal intercom interaction logic without occupying additional redundant system resources. Low-priority events: The triggering condition is a suspected false wake-up signal caused by sudden environmental noise, indoor echo, or non-human voice noise, and there is no valid intercom intention to support it. The system only retains 10% of the minimum detection resources to maintain signal monitoring, does not start the speaker main drive circuit, does not establish an intercom link, directly suppresses the subsequent response of invalid events, and reduces invalid power consumption and frequent false start and stop of the speaker. When four types of events occur concurrently, the system strictly follows the order of urgency > high priority > medium priority > low priority for preemptive scheduling. Low priority events are temporarily stored in the task buffer queue and are processed in turn after the high priority events are processed, thus ensuring the timeliness of critical event response and the orderly operation of the system.

[0026] like Figure 1-2 As shown, the intent and resource dynamic matching algorithm carried by the priority scheduling layer module uses three core factors—intent confidence, scene adaptability, and user permission level coefficient—as the calculation benchmark. It employs a weighted coupling operation method to dynamically allocate core operating resources in real time, including system computing resources, audio sampling resources, speaker amplifier gain resources, and wireless communication bandwidth resources. The resource allocation coupling calculation formula is defined as follows: In the formula: This refers to the actual resource allocation percentage calculated in real time. Set a baseline resource quota for the system; The value is the confidence level of the intercom intent after normalization to 0-1; This represents the current scene adaptability coefficient. User identity and permission level coefficient; Specific parameter calibration rules: Default scene adaptation coefficient for quiet home environments. Typical office scenarios Outdoor and industrial high-noise scenarios Automatically increase resource quotas to ensure anti-interference capabilities; administrator permission level coefficients. Ordinary resident users Temporary visitors Resource allocation weights decrease progressively based on user permissions. The algorithm operates on a fixed 20ms cycle, constantly refreshing resource allocation ratios and dynamically adjusting hardware parameters such as the main processor's clock speed, audio sampling rate, speaker amplifier gain, voice wake-up detection computing power overhead, and WiFi / Bluetooth communication transmission bandwidth. This enables intelligent on-demand resource allocation for scenarios with high intent, high permissions, and complex scenarios, while streamlining resources for scenarios with low intent, low permissions, and quiet environments. This ensures the responsiveness of intercom wake-up and voice interaction while reducing system resource redundancy and waste, further lowering the average power consumption of the entire device from an algorithmic perspective.

[0027] like Figure 1-2 As shown, the three-segment speaker control architecture of the state pre-control layer module is a pre-control, follow-up, and recovery time-series closed-loop architecture. It abandons the traditional coarse switching control mode of instantaneous power-on and instantaneous power-off of speakers, and covers the entire life cycle before intercom triggering, during intercom, and after intercom. It performs refined management and control from multiple dimensions such as electrical impact, sound quality, mechanical stress, and device life. Pre-control phase: When the confidence level of the intercom intent output by the intent perception layer reaches the preset threshold, the pre-control process is started 500ms before the intercom link is officially established and the official voice is played. A weak bias current of 1 / 5 of the rated operating voltage is input to the speaker voice coil to complete the voice coil preheating, the power amplifier circuit steady state establishment, and the diaphragm mechanical attitude pre-positioning. This effectively eliminates the surge current impact generated by instantaneous full power-up, reduces the sudden change in temperature of the voice coil, and significantly shortens the speaker sound response delay. Follow-up phase: After the audio and video intercom link is established, the frequency distribution, loudness amplitude, speech rate and rhythm, and signal-to-noise ratio characteristics of the intercom speech are extracted frame by frame in real time. The speaker output power, power amplifier gain, frequency equalization and diaphragm vibration amplitude are dynamically and adaptively adjusted. The output parameters are adapted for low frequency, mid frequency and high frequency human voice segments to ensure speech clarity at different speech rates and volumes, while avoiding diaphragm overload distortion at high volumes and low signal-to-noise ratio at low volumes. Recovery Phase: After the system detects the end-of-talk voice signal and the precursor signal of link disconnection, it initiates the power smoothing recovery and fallback process 300ms in advance, and gradually reduces the speaker output power using an exponential decay law to reduce the popping noise and sudden drop in voice coil temperature caused by sudden power failure. A specific formula for calculating the power exponential attenuation of a loudspeaker is set up: In the formula: decay time The corresponding real-time output power of the speaker; This refers to the maximum rated operating power of the speaker during intercom communication. The preset power attenuation coefficient can be adaptively calibrated according to the speaker model. This refers to the duration of power attenuation. Through three-stage time-sequenced closed-loop control, the speaker can be smoothly started and stopped without current impact, the voice output parameters can be adaptively matched in real time, and the electrical and mechanical stress can be slowly released. This not only improves the audio and video intercom listening experience, but also significantly reduces the aging rate of the speaker during long-term operation and significantly extends the service life of the device.

[0028] like Figure 1-2 As shown, the dynamic power rail unit of the state pre-control layer module has a built-in 3.3V, 5V and 12V three-level switchable regulated power supply module, which consists of a high-precision voltage reference source, MOS transistor array switching circuit, LDO low voltage regulation circuit and overcurrent and overvoltage protection circuit. It has the characteristics of seamless voltage level switching, low output ripple, strong load capacity and fault self-protection. The module uses real-time ambient noise in decibels as the basis for gear switching, samples ambient noise intensity at 100ms intervals, and automatically adapts the speaker power supply voltage level using a zone threshold matching rule. In the formula: The decibel value of the ambient noise detected in real time; Match the power supply voltage level of the speaker output to the system; In quiet, low-noise home environments, with L≤40dB, the system automatically switches to the 3.3V low-voltage power supply setting to reduce the speaker amplifier's static operating current and significantly decrease static power consumption. In moderately noisy office environments, with 40dB<L≤65dB, the system switches to the 5V standard power supply setting to balance power consumption and sound volume. In high-noise outdoor and industrial environments, with L>65dB, the system switches to the 12V high-voltage power supply setting to increase speaker output power and dynamic sound range, ensuring clear and identifiable intercom voice communication even in noisy environments. Meanwhile, the module has an adaptive fine-tuning function for high and low temperature conditions. When the ambient temperature of the device is below 0℃ or above 45℃, the output voltage is slightly adjusted on the basis of the original setting to compensate for the loss of sound quality and power caused by the impedance drift of the low temperature speaker and the efficiency decay of the high temperature power amplifier. When the system enters deep sleep mode, it automatically forces a switch to the lowest voltage setting of 3.3V to maximize the reduction of standby power consumption.

[0029] like Figure 1-2 As shown, the power recovery circuit built into the state pre-control layer module consists of a freewheeling diode array, an LC filter circuit, a supercapacitor energy storage unit, a low-voltage step-down voltage regulation branch, and an anti-backflow isolation circuit. It is specifically designed to capture, rectify, filter, store, and reuse the reverse electromotive force generated under four operating conditions: speaker pre-control switching, power dynamic adjustment, power outage, and no-load standby. The loudspeaker is an inductive sound-generating device. During sudden voltage changes, power drops, or momentary power outages, the voice coil inductor generates a reverse induced electromotive force. In conventional circuits, this energy is wasted as heat. This power recovery circuit clamps the reverse voltage using a freewheeling diode, filters out high-frequency noise using an LC filter circuit, and stores the clean energy in a supercapacitor energy storage unit. The energy is then regulated to the system's low-power operating voltage via a low-voltage step-down regulator branch. The energy recovery efficiency is calculated using the following formula: In the formula: The overall energy recovery efficiency of the power recovery circuit; The total energy of the reverse electromotive force generated during the horn operating condition switching process; This is the effective electrical energy that can ultimately be stably reused by the supercapacitor energy storage unit. After being regulated, the recovered electrical energy is prioritized for use in the low-power coprocessing monitoring circuit, microphone standby acquisition circuit, and real-time clock timing circuit, replacing the external power supply. The circuit has a built-in anti-backflow isolation diode to prevent the energy storage capacitor from flowing back and damaging the front-end power module. This design can recover and reuse the idle electromagnetic energy of the speaker, effectively reducing the normal standby power consumption and average operating power consumption of the whole machine, and adapting to the low power consumption and battery life requirements of battery-powered wireless audio and video intercom devices.

[0030] like Figure 1-2 As shown, the speaker status monitoring feedback module integrates a high-precision miniature patch temperature sensor, a high-frequency AC impedance detection circuit, and a MEMS triaxial vibration sensor. The sensors are all installed close to the speaker voice coil and diaphragm. With a fixed sampling period of 50ms, it synchronously collects four core hardware status parameters in real time: speaker voice coil operating temperature T, rated AC impedance Z, diaphragm vibration amplitude A, and vibration frequency. The collected raw sensor data is processed by hardware filtering, software noise reduction, and moving average normalization to remove interference and glitch data and retain the true and valid status parameters. The module presets the speaker's safe operating threshold range: voice coil safe operating temperature 40℃~80℃, AC impedance allowable fluctuation range of ±30% of rated impedance, and diaphragm vibration amplitude not exceeding the rated full-amplitude limit. It also sets an abnormal parameter judgment formula as the basis for protection triggering. In the formula: The rated standard AC impedance of the speaker as specified by the manufacturer. The module transmits the processed real-time status parameters, threshold comparison results, and anomaly level determination results back to the priority scheduling layer module and the status pre-control layer module in a closed loop throughout the entire process. The system then dynamically completes logical self-correction based on the feedback parameters. If a high temperature or slight impedance drift is detected, the speaker amplifier output power will be automatically reduced and the no-operation sleep delay will be extended. If a high temperature overload, severe impedance abnormality, or diaphragm amplitude exceeding the second-level protection condition is triggered, the system will immediately force a 50% reduction in output power and limit the maximum volume. If the parameters continue to exceed the standard without any recovery trend, the speaker's main power supply circuit will be directly cut off, and the start / stop control logic will be locked to avoid hardware damage such as speaker burnout or amplifier breakdown. Meanwhile, the module has a historical parameter storage function, which can record the temperature, impedance and amplitude change trends of the speaker over 7 days, forming a device health status curve, providing data support for long-term aging prediction and early warning of faults.

[0031] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A smart start / stop system for audio and video intercom speakers that integrates voice wake-up, characterized in that: include: The intent perception layer module, through its built-in intercom intent predictor, two-way signal separator, scene intent matcher, and permission verification engine, and as the system's front-end multi-source information perception and pre-decision entry point, is responsible for real-time acquisition of environmental sound field signals and audio-visual intercom link signals, completing the isolation and separation of local voice wake-up signals and remote intercom signals. At the same time, it integrates voiceprint features, usage context, historical behavior, application scenarios, and user identity and permissions multi-dimensional information to predict the user's intercom behavior intent in advance and output intent confidence, scene type, and permission judgment results, providing a reliable perception data source, intent decision basis, and scene adaptation parameters for the subsequent priority scheduling layer and state pre-control layer. The priority scheduling layer module is configured with a multi-level wake-up event scheduling mechanism and a dynamic matching algorithm for intent and resources. It performs hierarchical resource allocation and action scheduling for wake-up events based on priority levels, and dynamically allocates system and speaker operating resources in combination with intent confidence, scene adaptability, and permission level coefficient. The state pre-control layer module adopts a three-segment speaker control architecture and dynamic power rail unit. After anticipating the intercom intention, it enters the pre-control stage in advance to preheat the speaker bias circuit. After the intercom is established, it dynamically adjusts the speaker output power and amplitude in real time according to the intercom voice characteristics. Before the intercom ends, it enters the recovery stage in advance to complete the power drop-off in an exponential manner. The speaker power supply voltage is dynamically adjusted through a multi-level switchable power module, and a power recovery circuit is configured to recover and utilize the speaker's idle back electromotive force. The speaker status monitoring and feedback module collects speaker voice coil temperature, impedance, and diaphragm vibration parameters in real time and transmits them back to the priority scheduling layer module and the status pre-control layer module. Combined with real-time parameters, it dynamically corrects the wake-up scheduling strategy and speaker start-stop control logic, realizing audio speaker seamless intelligent start-stop, anti-false wake-up, low-power operation, and overload self-protection, with intercom intent prediction, event hierarchical scheduling, and hardware status closed-loop perception.

2. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The intercom intent predictor in the intent perception layer module adopts a lightweight temporal attention model, builds a personalized intent model based on the user's historical intercom behavior, combines real-time sound field features to predict the user's intercom behavior in advance, and outputs the intent confidence level within the interval range.

3. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The bidirectional signal separator in the intent perception layer module adopts a phase difference and energy ratio dual-factor separation algorithm. By combining microphone array sound source phase locking with dynamic energy threshold determination, it achieves precise isolation and filtering of local wake-up signals and remote intercom signals.

4. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The scene intent matcher in the intent perception layer module has a built-in scene and intercom mode mapping library, which can automatically identify noisy application scenarios such as home, office, and outdoor. The permission verification engine adopts a voiceprint and permission dynamic binding mechanism to divide intercom access permissions into three levels: administrator, ordinary user, and visitor.

5. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The multi-level wake-up event scheduling mechanism of the priority scheduling layer module is specifically divided into four event levels: emergency, high priority, medium priority, and low priority. Differentiated system resource quotas, response times, and execution action logic are configured for different event levels.

6. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The intent and resource dynamic matching algorithm of the priority scheduling layer module uses intent confidence, scenario adaptability, and permission level coefficient as correlation factors to calculate and allocate system running resources and loudspeaker working resource quotas in a coordinated manner.

7. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The three-stage speaker control architecture of the state pre-control layer module is a closed-loop architecture of pre-control, follow-up, and recovery timing. In the pre-control stage, the speaker bias circuit is activated in advance to warm up after the intercom intention is predicted to meet the standard. In the follow-up stage, the speaker operating condition is dynamically adjusted to adapt to the intercom voice characteristics. In the recovery stage, the speaker output power is gradually reduced by exponential decay before the intercom ends.

8. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The dynamic power rail unit of the state pre-control layer module is configured with a multi-level switchable power module. The multi-level switchable power module supports adaptive switching of multiple power supply voltages to match the speaker power output requirements under different scenarios.

9. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The power recovery circuit built into the state pre-control layer module is used to collect the reverse electromotive force generated when the speaker is idle and reuse the recovered electrical energy to the system's low-power standby power supply circuit.

10. The intelligent start / stop system for audio and video intercom speakers integrating voice wake-up as described in claim 1, characterized in that: The speaker status monitoring and feedback module collects the speaker voice coil operating temperature, real-time impedance value, and diaphragm vibration amplitude parameters in real time, and transmits the collected parameters back in real time to form a closed-loop feedback, dynamically correcting the wake-up event scheduling strategy and the speaker start / stop and power control logic.