A web real-time audio monitoring intercom system, method, electronic device and storage medium
Patent Information
- Application Number
- CN202611136630.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]本发明意在提供一种Web实时音频监听对讲系统、方法、电子设备及存储介质,以解决现有Web音频监听对讲方案主线程阻塞、音频传输效率低的问题
[0008]本方案的原理及优点是:实际应用时,依托主线程与TimerWorker分离的多线程架构,将全部音频采集、编码、运算耗时操作剥离至后台线程,从根源避免页面卡顿;搭配时域分块去重、动态缓冲发包、编码转换多重数据优化手段,大幅降低音频传输带宽占用;依靠心跳保活、网络状态监听、断线自动重连机制提升复杂网络下通信稳定性;辅以设备热插拔适配、渐进式权限管理、可视化状态反馈优化人机交互,同时完整管控工作线程生命周期减少内存占用,同步解决现有技术音频卡顿、带宽消耗高、弱网易断连、设备兼容差、页面交互卡顿、内存泄漏、用户操作引导缺失等缺陷。
Smart Images

Figure CN122824722A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Web audio communication technology, specifically to a Web real-time audio monitoring and intercom system, method, electronic device, and storage medium. Background Technology
[0002] In the field of video surveillance, real-time audio monitoring and two-way communication are core functions for enhancing on-site emergency response capabilities. With the widespread application of web technology, monitoring platforms that can be accessed directly without a client or browser have become the mainstream solution, and stable and efficient real-time voice communication via the web has become an industry necessity. However, in real-world business scenarios, on-site network fluctuations are frequent, front-end audio equipment models are diverse, and the platform needs to run continuously for extended periods. Traditional web audio intercom solutions are limited by the browser's operating mechanism, making them prone to issues such as voice stuttering, disconnections, and device compatibility problems, thus failing to meet the demands for high stability.
[0003] Current mainstream web audio intercom implementation solutions mainly fall into two categories. The first category is a basic single-threaded implementation solution, where all audio sampling, data processing, and timed packet sending logic run on the main page thread, which directly blocks page rendering and causes lag in user operation response; it continuously transmits unprocessed raw audio bitstreams, resulting in high bandwidth consumption and frequent audio loss in weak network scenarios; the intercom session terminates directly after the network connection is interrupted, and the page must be manually restarted to resume; it only reads the audio device once during page initialization and requests microphone permission once, so the acquisition function fails after the device is plugged in or unplugged, and the audio module is directly paralyzed when permission is denied, without any remedial operation guidance.
[0004] The second type is an improved solution that introduces a simple working thread. However, the timing of audio acquisition and scheduling is inaccurate, and the audio output timing is unstable. The data packet sending interval is fixed, and a large amount of repeated sampling data is continuously transmitted during the silent phase, resulting in a waste of bandwidth resources. After the link is disconnected, it cannot automatically identify and reconnect to restore the intercom. Hot-plugging of audio devices cannot be adapted in real time, and microphone permissions are only requested once via a pop-up window. Background threads and audio cache resources cannot be automatically released after the intercom ends, which will cause memory leaks after long-term operation, continuously reduce page running speed, and in severe cases, cause browser crashes.
[0005] In summary, existing web audio intercom solutions have many shortcomings: audio processing consumes the main thread, causing page lag; audio data is not optimized, resulting in high bandwidth consumption and voice lag in weak network conditions; sessions cannot be automatically resumed after network interruption; changes in audio devices cannot be identified in real time, and abnormal permissions are not guided; long-term operation is prone to memory leaks; there is no real-time display of device, network, and volume status, and fault prompts and solutions are lacking, resulting in poor overall performance, stability, and user experience. Summary of the Invention
[0006] The present invention aims to provide a Web real-time audio monitoring and intercom system, method, electronic device and storage medium to solve the problems of main thread blocking and low audio transmission efficiency in existing Web audio monitoring and intercom solutions.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: A web-based real-time audio monitoring and intercom method includes: S1. The system initializes and builds a layered multi-threaded audio processing architecture, including a main thread and an independent TimerWorker timer worker thread. The main thread is used to handle lightweight UI control tasks, and the TimerWorker timer worker thread is used to execute time-consuming tasks related to audio processing. S2. Detect available audio acquisition devices through the device management module and apply for microphone permissions as needed using a progressive strategy; S3. The TimerWorker timer thread periodically collects the raw PCM audio data output by the microphone, and uses a time-domain block deduplication algorithm to remove duplicate audio data blocks. S4. Encode and convert the deduplicated PCM audio data, store it in the audio buffer, monitor the buffer data volume in real time, and execute the dynamic buffer packet sending strategy. S5. Establish a bidirectional communication link between the client and the server based on WebSocket, start a heartbeat detection mechanism, and periodically send heartbeat packets to verify the link connectivity status; S6. Monitor the browser's global network online or offline status events in real time. When the network is offline, pause audio transmission and cache the current intercom session state. After the network is restored or the link is successfully reconnected, automatically resume audio monitoring and intercom transmission. S7. The main thread receives the encoded audio data packets and status information returned by the TimerWorker timer worker thread, and calls the network communication module to complete the uplink transmission of audio data.
[0008] The principle and advantages of this solution are as follows: In practical applications, relying on a multi-threaded architecture that separates the main thread and TimerWorker, all time-consuming audio acquisition, encoding, and computation operations are offloaded to background threads, thus avoiding page lag at the source; combined with multiple data optimization methods such as time-domain block deduplication, dynamic buffering of packet sending, and encoding conversion, the bandwidth consumption of audio transmission is significantly reduced; relying on heartbeat keep-alive, network status monitoring, and automatic reconnection mechanisms after disconnection, the stability of communication under complex networks is improved; supplemented by hot-swappable device adaptation, progressive permission management, and visual status feedback to optimize human-computer interaction, while fully managing the lifecycle of worker threads to reduce memory consumption, and simultaneously solving the defects of existing technologies such as audio lag, high bandwidth consumption, easy disconnection in weak networks, poor device compatibility, page interaction lag, memory leaks, and lack of user operation guidance.
[0009] Preferably, as an improvement, in step S1, the main thread and the TimerWorker timer worker thread complete inter-thread data interaction through the postMessage and onmessage message events.
[0010] Technical effect: It realizes isolated bidirectional data communication between the main thread and background worker threads, without sharing the execution context, avoiding thread resource contention conflicts, balancing the transmission efficiency of command issuance and audio data return, and ensuring stable audio timing.
[0011] Preferably, as an improvement, in S2, the audio devices are scanned in real time by listening to the devicechange event of mediaDevices. When a device is plugged in or unplugged, a device switching notification is pushed and the current audio session is kept uninterrupted. Microphone permissions are requested only when the intercom is started, and visual guidance information is output for scenarios with abnormal permissions.
[0012] Technical benefits: Enables seamless hot-swapping of audio devices such as microphones and headphones, and eliminates the need to restart the intercom session when changing devices; progressive permission requests avoid unnecessary permission occupation, and visual prompts for permission exceptions lower the user's operating threshold, significantly improving device compatibility and smoothness of use.
[0013] Preferably, as an improvement, the time-domain block deduplication algorithm includes: using a preset number of sampling points as a detection unit, comparing adjacent audio data blocks, identifying and removing consecutive duplicate sampling data, and retaining valid speech waveforms.
[0014] Technical effect: Accurately removes duplicate sampling data in silent segments, reduces the amount of uplink data transmission without compromising the quality of human voice or causing audio dropout distortion, reduces bandwidth consumption, and alleviates audio stuttering problems in weak network environments.
[0015] Preferably, as an improvement, the dynamic buffer packet sending strategy includes: a preset buffer threshold; when the amount of data in the buffer exceeds the threshold, the packet sending interval is reduced to speed up the data sending rate; when the amount of data in the buffer is lower than the threshold, the packet sending interval is increased to reduce the frequency of network interactions.
[0016] Technical effects: The packet sending rate is dynamically adjusted according to the cache load. When the cache is piling up, the packet is pushed quickly to prevent audio delay. When the load is low, the packet sending frequency is reduced to reduce network request overhead, thus balancing the real-time performance of audio with the server and network load pressure.
[0017] Preferably, as an improvement, in step S4, the PCM audio data is converted into the G.711A encoding format.
[0018] Technical effect: The raw PCM stream is efficiently compressed, adapting to low-bandwidth transmission scenarios for monitoring and intercom. While ensuring clear voice quality, the data packet size is further reduced, improving transmission smoothness.
[0019] Preferably, as an improvement, in S5, if no heartbeat response is received for 20 consecutive seconds, the communication link is determined to be abnormal. The handling process after the link is abnormal includes: caching the current intercom session cache data and automatically initiating WebSocket reconnection; if the reconnection is successful, reading the cache data to restore the intercom session; if the reconnection fails, outputting the fault solution prompt.
[0020] Technical benefits: Heartbeat detection identifies link disconnection faults in real time, automatically caches session data and performs disconnection reconnection, and intercom can be resumed without manual restart after the network is restored. Clear fault guidance is provided for reconnection failures, improving communication reliability in weak network and network fluctuation scenarios.
[0021] A Web-based real-time audio monitoring and intercom system, using a Web-based real-time audio monitoring and intercom method, includes: a TalkAndListen main controller, a Recorder audio recorder, a TimerWorker timer worker thread module, a network communication module, and a device management module; The TalkAndListen main controller is used to manage session state; The Recorder audio recorder is used for audio acquisition, preprocessing, encoding conversion, and audio deduplication. The TimerWorker timer worker thread module is used for timed data acquisition, heartbeat packet generation, and thread data synchronization. The network communication module implements WebSocket link management, audio transmission, heartbeat detection, and disconnection reconnection. The device management module enables hardware detection, access control, hot-swappable device adaptation, and user status feedback.
[0022] An electronic device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement a web-based real-time audio monitoring and intercom method.
[0023] A computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor as a method for real-time audio monitoring and intercom via the web. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating a web-based real-time audio monitoring and intercom method. Figure 2 A schematic diagram of the multi-threaded processing flow of a Web real-time audio monitoring and intercom method; Figure 3 This is a schematic diagram of the audio data processing flow of a web-based real-time audio monitoring and intercom method. Figure 4 A schematic diagram of the network transmission process for a Web-based real-time audio monitoring and intercom method; Figure 5 A schematic diagram of the device management process for a web-based real-time audio monitoring and intercom method; Figure 6 This is a schematic diagram of the structure of a web-based real-time audio monitoring and intercom system. Detailed Implementation
[0025] The following detailed description illustrates the specific implementation method: The basic implementation examples are as follows: Figure 1 As shown: A web-based real-time audio monitoring and intercom method includes the following steps: S1, as Figure 2 As shown, during the system initialization phase, a layered, multi-threaded audio processing architecture is built, consisting of a main thread and independent TimerWorker timer worker threads. After the page loads, the main thread first instantiates basic components such as the TalkAndListen main controller, device management module, and network communication module. It then creates TimerWorker timer worker threads through the Worker interface, synchronously registering the postMessage sending interface and onmessage message listener callback for both ends, establishing a bidirectional, shared-memory-free thread communication channel. The main thread is only responsible for lightweight UI control tasks such as page rendering, user interaction, module scheduling, and session state management, while all time-consuming audio processing tasks, such as timed audio acquisition, audio encoding, and audio data processing, are independently executed by the TimerWorker timer worker threads. The two types of threads rely on postMessage and onmessage message events to complete bidirectional data interaction. The main thread sends control commands such as device selection, start / stop acquisition, and parameter configuration to TimerWorker through postMessage. TimerWorker sends back data messages such as deduplicated and encoded audio data packets, buffer status, device abnormalities, and heartbeat detection results to the main thread. The main thread receives and parses the data sent back by the worker thread through its own onmessage callback. The execution contexts of the threads are isolated from each other throughout the process, so there will be no resource contention or task blocking issues. At the same time, the main thread performs complete lifecycle management of TimerWorker, including creation, running, idle and destruction, which avoids memory leaks caused by main thread blocking and long-term resident background threads from the underlying level.
[0026] S2, as Figure 5As shown, during system operation, the device management module uniformly handles the entire process of audio device detection, microphone permission control, and hot-swappable device adaptation: After initialization, the device management module calls the mediaDevices standard interface to scan all available microphones, headphones, and other audio acquisition hardware on the local machine and generates a device list for the user to select; the permission request process adopts a progressive, on-demand triggering strategy. During the system initialization phase, microphone permissions are not actively requested. Only when the user triggers a listening or intercom operation and truly needs to use the audio acquisition hardware will the browser permission interface be called to initiate a microphone permission request; for various permission abnormal scenarios such as permission being denied by the user, device being occupied by other programs, or hardware being missing, the device management module pushes visual guidance information such as pop-ups and text prompts to the page, informing the user of the cause of the fault and providing operation guidance such as re-authorization and device replacement. Meanwhile, the device management module continuously registers and listens for the devicechange hardware change event of mediaDevices, and monitors the plugging and unplugging of audio devices in real time. Once a new or removed audio device is detected, the entire device list is automatically rescanned and refreshed, and a device switching notification pop-up is immediately pushed to the main thread for the user to select the target acquisition device. Throughout the entire process of device change detection and switching, the current audio acquisition and transmission session parameters are cached to continuously maintain the original audio session without interruption, avoiding forced disconnection of intercom and monitoring functions due to hot plugging and unplugging of devices, and ensuring the continuity of user operation.
[0027] S3, as Figure 3 As shown, when the user enables audio monitoring or intercom functions, the main thread sends a start command to the TimerWorker timer thread. This background thread continuously retrieves the raw PCM audio data output by the Recorder audio recorder at fixed intervals, relying on an internal independent high-precision timer. The raw PCM audio data is a pulse code modulation raw sampled data stream directly acquired by the microphone without compression or format conversion. After the TimerWorker timer thread acquires the segmented output raw PCM audio data, it performs time-domain block deduplication processing. Specifically, the continuous PCM sampled data stream is evenly divided into 128 sampling points as an independent detection unit. The sampling values in each adjacent detection unit are traversed and compared sequentially. Frame by frame, continuous repeated sampling data blocks with completely identical values are identified, judged as silent and invalid data, and directly removed. Only valid sampled waveform data containing speech fluctuation characteristics are retained. The size of the audio data to be transmitted is reduced without damaging the integrity of human speech or causing audio dropout distortion. The deduplicated audio data after processing continues to flow to the encoding buffer stage.
[0028] S4, after the TimerWorker thread completes the time-domain block deduplication process, it performs encoding conversion on the raw PCM audio data after removing redundant sampled data. This converts the PCM format into a G.711A audio stream suitable for low-bandwidth intercom scenarios. The converted audio data packets are then temporarily stored in a dedicated audio buffer. The thread continuously reads the accumulated data volume within the buffer in real time and compares it with a preset 320... A 4-byte buffer threshold is used for comparison and judgment to execute a dynamic buffer packet sending strategy: once the total amount of data stored in the buffer exceeds the threshold, it indicates that the buffer is piling up and there is a risk of audio delay. The system automatically shortens the data packet sending interval to 16ms to speed up the data push speed of the buffer and avoid audio backlog and stuttering. If the amount of data in the buffer is lower than the threshold, it means that the current audio generation rate is slow. The system automatically lengthens the packet sending interval to 25ms to reduce the request overhead caused by high-frequency network packet sending, balance real-time performance and network load. The audio data packets after dynamic adjustment of the packet sending interval are pushed to the main thread network communication module through the thread message channel to complete the uplink transmission.
[0029] S5, such as Figure 4 As shown, the system establishes a bidirectional data transmission and reception communication link between the client and server based on the WebSocket protocol through the network communication module. After the link is established, a heartbeat detection mechanism is started synchronously. The TimerWorker timer thread continuously sends heartbeat data packets to the server at fixed intervals to continuously verify the connectivity of the current communication link. The network communication module continuously listens for heartbeat response messages sent by the server. If no heartbeat response message is received within 20 consecutive seconds, it is determined that the current WebSocket communication link is abnormal. When the link abnormality handling process is triggered, the system first caches the configuration parameters of the current intercom session and temporarily stores the audio buffer data. Then, it automatically executes the reconnection logic of destroying and recreating the WebSocket connection to attempt to reconnect to the server. If the reconnection is successful, the system reads the previously cached session configuration and audio data and seamlessly restores the audio monitoring intercom transmission process. If multiple reconnection operations fail, the main thread outputs visual fault prompts to the front-end page and synchronously displays the corresponding fault diagnosis and solution guidance for users to handle.
[0030] S6's network communication module pre-registers continuous listener callbacks for global online and offline network status events in the browser, capturing changes in client network connectivity in real time. When an offline event is detected, the system immediately pauses the TimerWorker's push of audio data packets to the main thread, stops WebSocket uplink audio transmission, and caches the current complete intercom session parameters and any unsent buffered audio data to local storage, preserving the complete session state. When an online network recovery event is detected, or after a successful WebSocket reconnection, the system reads the locally cached session state and audio buffer data, controls the TimerWorker to restart audio processing and data push logic, and the network communication module resumes uplink audio data transmission. This eliminates the need for users to manually restart the intercom, enabling automatic service continuation after network fluctuations.
[0031] In S7, the main thread pre-registers the onmessage message listener callback and continuously receives two types of information from the TimerWorker timer worker thread via postMessage: one is the audio data packet encoded with G.711A, and the other includes various operational status information such as buffer load, device changes, heartbeat link, and real-time volume. After receiving the encoded audio data packet, the main thread directly calls the network communication module interface to transmit the audio data packet to the server via the established WebSocket link. At the same time, the main thread parses the various status information synchronously pushed by the worker thread, refreshes the front-end page visualization display area in real time, and synchronously updates the name of the currently connected audio device, the online or abnormal status of the WebSocket connection, the real-time audio volume value and volume display control, allowing users to intuitively grasp the overall operation status of the current intercom monitoring.
[0032] like Figure 6 As shown, it also includes a Web real-time audio monitoring and intercom system, which uses a Web real-time audio monitoring and intercom method, including: TalkAndListen main controller, Recorder audio recorder, TimerWorker timer worker thread module, network communication module, and device management module; The TalkAndListen main controller is used to manage session state.
[0033] The Recorder audio recorder is used for audio acquisition, preprocessing, encoding conversion, and audio deduplication.
[0034] The TimerWorker timer working thread module is used for timed data acquisition, heartbeat packet generation, and thread data synchronization.
[0035] The network communication module implements WebSocket link management, audio transmission, heartbeat detection, and disconnection reconnection.
[0036] The device management module enables hardware detection, access control, hot-swappable device adaptation, and user status feedback.
[0037] It also includes an electronic device comprising a processor and a memory, the memory storing a computer program, wherein the processor executes the computer program to implement a web-based real-time audio monitoring and intercom method.
[0038] It also includes a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor as a Web-based real-time audio monitoring and intercom method.
[0039] The above descriptions are merely embodiments of the present invention, and common knowledge such as specific technical solutions and / or characteristics are not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the technical solutions of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A web-based real-time audio monitoring and intercom method, characterized in that, include: S1. The system initializes and builds a layered multi-threaded audio processing architecture, including a main thread and an independent TimerWorker timer worker thread. The main thread is used to handle lightweight UI control tasks, and the TimerWorker timer worker thread is used to execute time-consuming tasks related to audio processing. S2. Detect available audio acquisition devices through the device management module and apply for microphone permissions as needed using a progressive strategy; S3. The TimerWorker timer thread periodically collects the raw PCM audio data output by the microphone, and uses a time-domain block deduplication algorithm to remove duplicate audio data blocks. S4. Encode and convert the deduplicated PCM audio data, store it in the audio buffer, monitor the buffer data volume in real time, and execute the dynamic buffer packet sending strategy. S5. Establish a bidirectional communication link between the client and the server based on WebSocket, start a heartbeat detection mechanism, and periodically send heartbeat packets to verify the link connectivity status; S6. Monitor the browser's global network online or offline status events in real time. When the network is offline, pause audio transmission and cache the current intercom session state. After the network is restored or the link is successfully reconnected, automatically resume audio monitoring and intercom transmission. S7. The main thread receives the encoded audio data packets and status information returned by the TimerWorker timer worker thread, and calls the network communication module to complete the uplink transmission of audio data.
2. A web-based real-time audio monitoring and intercom method according to claim 1, characterized in that: In S1, the main thread and the TimerWorker timer worker thread complete inter-thread data interaction through the postMessage and onmessage message events.
3. The Web real-time audio monitoring and intercom method according to claim 1, characterized in that: In S2, audio devices are scanned in real time by listening to the devicechange event of mediaDevices. When a device is plugged in or unplugged, a device switching notification is pushed and the current audio session is kept uninterrupted. Microphone permissions are requested only when the intercom is started, and visual guidance information is output for scenarios with abnormal permissions.
4. The Web real-time audio monitoring and intercom method according to claim 1, characterized in that: The temporal block deduplication algorithm includes: using a preset number of sampling points as a detection unit, comparing adjacent audio data blocks, identifying and removing consecutive duplicate sampling data, and retaining valid speech waveforms.
5. A Web-based real-time audio monitoring and intercom method according to claim 1, characterized in that: The dynamic buffering packet sending strategy includes: setting a preset buffer threshold; when the amount of data in the buffer exceeds the threshold, reducing the packet sending interval to speed up the data sending rate; and when the amount of data in the buffer is lower than the threshold, increasing the packet sending interval to reduce the frequency of network interactions.
6. The Web real-time audio monitoring and intercom method according to claim 1, characterized in that: In step S4, the PCM audio data is converted into the G.711A encoding format.
7. A Web-based real-time audio monitoring and intercom method according to claim 6, characterized in that: In S5, if no heartbeat response is received for 20 consecutive seconds, the communication link is determined to be abnormal. The handling process after the link is abnormal includes: caching the current intercom session cache data and automatically initiating WebSocket reconnection; if the reconnection is successful, reading the cached data to restore the intercom session; if the reconnection fails, outputting the fault solution prompt.
8. A Web-based real-time audio monitoring and intercom system, using the Web-based real-time audio monitoring and intercom method as described in any one of claims 1-7, characterized in that, include: The system includes the TalkAndListen main controller, Recorder audio recorder, TimerWorker timer worker thread module, network communication module, and device management module. The TalkAndListen main controller is used to manage session state; The Recorder audio recorder is used for audio acquisition, preprocessing, encoding conversion, and audio deduplication. The TimerWorker timer worker thread module is used for timed data acquisition, heartbeat packet generation, and thread data synchronization. The network communication module implements WebSocket link management, audio transmission, heartbeat detection, and disconnection reconnection. The device management module enables hardware detection, access control, hot-swappable device adaptation, and user status feedback.
9. An electronic device, characterized in that: It includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement a Web real-time audio monitoring and intercom method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: It stores a computer program, which is executed by a processor as described in any one of claims 1-7: a Web real-time audio monitoring and intercom method.