Multimodal real-time voice communication system, communication switching method
Patent Information
- Application Number
- CN202610985194.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-03
AI Technical Summary
[0004]本发明的目的在于提出了一种多模式实时语音通信系统、通信切换方法,旨在解决现有技术中对讲模式切换不智能、易中断的问题
[0012]采用本发明实施例,具有如下有益效果:区别于现有技术的情况,本申请在所述MESH对讲模式下,实时监测所述对讲设备的MESH信号强度、蜂窝网络信号质量;若监测到所述MESH信号强度和所述蜂窝网络信号质量满足第一设定判断条件,将所述对讲设备从所述MESH对讲模式切换至所述蜂窝网络对讲模式,以及关闭所述对讲设备的MESH音频通道;在所述蜂窝网络对讲模式下,持续监测MESH信号强度,若监测到所述MESH信号强度满足第二设定判断条件,由所述蜂窝网络对讲模式切回所述MESH对讲模式,并恢复所述对讲设备的MESH音频通道。通过上述方式,实现对讲设备之间在MESH对讲模式和蜂窝网络对讲模式的切换,通信资源自适应优化配置。
Smart Images

Figure CN122513841B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and in particular to a multi-mode real-time voice communication system and a communication switching method. Background Technology
[0002] Traditional intercom communication methods mainly include two types: mesh intercom and cellular intercom. Mesh intercom does not rely on network signals, does not consume mobile data, and has minimal latency, making it suitable for emergency scenarios without network coverage, but its communication range is limited. Cellular intercom (such as 5G / 4G intercom) relies on base station signals, consumes mobile data, and has relatively higher latency, but it has a long communication range; as long as the mobile signal is good, it can achieve ultra-long-distance communication.
[0003] In existing technologies, some dual-mode intercom devices can switch between MESH and cellular networks, but the following problems exist: the switching process requires manual intervention, resulting in a poor user experience; the original connection is disconnected during switching, causing call interruptions or stuttering; the switching decision is not intelligent and cannot automatically determine the switching timing based on signal quality; and there is a lack of refined switching strategies in group communication scenarios. Therefore, there is an urgent need for a method that can achieve intelligent, seamless, and smooth switching between MESH and cellular intercom. Summary of the Invention
[0004] The purpose of this invention is to propose a multi-mode real-time voice communication system and a communication switching method, which aims to solve the problems of unintelligent intercom mode switching and easy interruption in the prior art.
[0005] To address the aforementioned technical problems, this application provides a multi-mode real-time voice communication switching method, applied to a communication system including a group of intercom devices and user terminals, wherein the intercom devices support MESH intercom mode and cellular network intercom mode; the method includes: in the MESH intercom mode, real-time monitoring of the MESH signal strength and cellular network signal quality of the intercom devices; if the MESH signal strength and cellular network signal quality are found to meet a first preset judgment condition, switching the intercom devices from the MESH intercom mode to the cellular network intercom mode, and disabling the MESH audio channel of the intercom devices; in the cellular network intercom mode, continuously monitoring the MESH signal strength; if the MESH signal strength is found to meet a second preset judgment condition, switching back from the cellular network intercom mode to the MESH intercom mode, and restoring the MESH audio channel of the intercom devices.
[0006] In one embodiment, the first set judgment condition includes: the MESH signal strength is lower than the first predetermined threshold, and the cellular network signal quality is higher than the second predetermined threshold; the step of automatically switching the intercom device from the MESH intercom mode to the cellular network intercom mode includes: triggering the intercom device to establish a cellular network intercom connection so that the intercom device can conduct voice communication through the cellular network intercom mode; and closing the audio input channel and audio output channel corresponding to the MESH intercom module.
[0007] In one embodiment, the second set judgment condition includes: the MESH signal strength is stably higher than a third predetermined threshold for a predetermined duration.
[0008] In one embodiment, if the intercom device group includes multiple intercom devices, the method further includes: monitoring the MESH signal strength and cellular network signal between any two intercom devices; if the MESH signal strength and cellular network signal quality between two target intercom devices meet the first preset judgment condition, switching the intercom mode between the target intercom devices to the cellular network intercom mode; maintaining the intercom mode between other intercom devices besides the target intercom devices as the MESH intercom mode, and controlling the intercom mode between the target intercom device and the other intercom devices as the MESH intercom mode.
[0009] In one embodiment, the method further includes: real-time monitoring of whether a call event exists on the user terminal; if the call event exists, muting the audio input and output channels corresponding to the MESH intercom module; if the call event is detected to have ended, restoring the audio input and output channels of the intercom device to the intercom mode before switching.
[0010] In one embodiment, the real-time monitoring of whether a call event exists on the user terminal includes: if a telephone audio data stream is detected on the user terminal, identifying the audio data stream status defined by the Bluetooth hands-free specification; determining whether the telephone call is in an connected state based on the audio data stream status; if it is in the connected state, determining that a call event exists; otherwise, determining that no call event exists.
[0011] To address the aforementioned technical problems, a second aspect of this application provides a multi-mode real-time voice communication system, including a group of intercom devices and user terminals. The intercom devices and user terminals are connected in a one-to-one communication manner. Each intercom device includes a MESH intercom module, a cellular communication module, and a main control unit. The MESH intercom module and the cellular communication module are respectively used to implement MESH intercom mode and cellular network intercom mode. The main control unit is used to implement the steps of the method provided in the first aspect above.
[0012] The embodiments of this invention offer the following advantages: Unlike existing technologies, this application, in the MESH intercom mode, monitors the MESH signal strength and cellular network signal quality of the intercom device in real time. If the MESH signal strength and cellular network signal quality meet a first preset judgment condition, the intercom device is switched from the MESH intercom mode to the cellular network intercom mode, and the MESH audio channel of the intercom device is turned off. In the cellular network intercom mode, the MESH signal strength is continuously monitored. If the MESH signal strength meets a second preset judgment condition, the device is switched back from the cellular network intercom mode to the MESH intercom mode, and the MESH audio channel of the intercom device is restored. Through the above methods, the switching between MESH intercom mode and cellular network intercom mode between intercom devices is achieved, and communication resources are adaptively optimized. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] in: Figure 1 This is a schematic diagram of an embodiment of the multi-mode real-time voice communication system of this application using MESH intercom mode communication; Figure 2 This is a schematic diagram of an embodiment of a multi-mode real-time voice communication system call communication; Figure 3 This is a schematic diagram of an embodiment of the multi-mode real-time voice communication system of this application using cellular network intercom mode communication; Figure 4 This is a flowchart illustrating an embodiment of the multi-mode real-time voice communication switching method of this application; Figure 5 This is a schematic block diagram of the early switching process based on motion state prediction in the multi-mode real-time voice communication switching method of this application; Figure 6 This is a schematic block diagram of the switching scheduling process based on the group silence period in the multi-mode real-time voice communication switching method of this application. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] Please see Figures 1-3 , Figure 1 This is a schematic diagram of an embodiment of the multi-mode real-time voice communication system of this application using MESH intercom mode communication. Figure 2 This is a schematic diagram of an embodiment of a multi-mode real-time voice communication system call communication. Figure 3 This is a schematic diagram of an embodiment of the multi-mode real-time voice communication system of this application using cellular network intercom mode communication. The multi-mode real-time voice communication system of this application includes: a group of intercom devices 110 and user terminals 120, with each intercom device 110 and user terminal 120 communicating in a one-to-one correspondence. Each intercom device 110 includes a MESH intercom module, a cellular communication module, and a main control unit. The MESH intercom module and the cellular communication module are respectively used to implement MESH intercom mode and cellular network intercom mode. The main control unit is used to implement the steps of the following embodiments of the multi-mode real-time voice communication switching method; please refer to the following embodiments for details.
[0017] Please see Figure 4 , Figure 4 This is a schematic flowchart illustrating an embodiment of the multi-mode real-time voice communication switching method of this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily replace it with a similar method. Figure 1 The illustrated process sequence is limited. It includes the following steps S11~S13: S11: In the MESH intercom mode, monitor the MESH signal strength and cellular network signal quality of the intercom device in real time; S12: If the MESH signal strength and the cellular network signal quality are detected to meet the first set judgment condition, the intercom device is switched from the MESH intercom mode to the cellular network intercom mode, and the MESH audio channel of the intercom device is turned off. S13: In the cellular network intercom mode, continuously monitor the MESH signal strength. If the monitored MESH signal strength meets the second set judgment condition, switch back from the cellular network intercom mode to the MESH intercom mode and restore the MESH audio channel of the intercom device.
[0018] Among them, MESH intercom mode refers to the intercom communication mode implemented based on wireless MESH network protocols (such as BLE MESH, ZigBee MESH, or proprietary MESH protocols). Its characteristics are that it does not go through base station forwarding, and voice transmission between devices is direct or through relays. It has the advantages of minimal latency, no network dependence, and no data consumption. Cellular network intercom mode refers to the intercom mode based on 4G / 5G mobile communication network to realize voice data forwarding. It needs to communicate with remote devices through base stations. It has the advantage of long communication distance, but consumes data and has relatively large latency. MESH signal strength refers to the wireless signal strength received by the intercom device from other intercom devices in the MESH network. It is usually expressed as RSSI (Received Signal Strength Indicator) in dBm. The higher the value, the stronger the signal. Cellular network signal quality refers to the base station signal quality received by the intercom device or user terminal device. It can be characterized by parameters such as RSRP (Reference Signal Receiving Power), RSRQ (Reference Signal Receiving Quality), or CSQ.
[0019] Closing the MESH audio channel means disconnecting the microphone input channel and speaker output channel of the MESH module through the main control unit, but keeping the RF transceiver function of the MESH module active to maintain the MESH network connection and continuously monitor signal strength; restoring the MESH audio channel means reopening the previously closed microphone input channel and speaker output channel of the MESH module, so that the MESH intercom function returns to normal.
[0020] In one specific embodiment, there is a walkie-talkie device A that simultaneously supports MESH walkie-talkie mode and 5G cellular network walkie-talkie mode. Initially, walkie-talkie device A is in MESH walkie-talkie mode and communicates with walkie-talkie device B via voice.
[0021] Intercom device A performs the following monitoring operations: It reads the RSSI value between itself and intercom device B via the MESH module (e.g., currently -82dBm); it reads the current 5G network signal quality via the cellular module (e.g., the current CSQ value is 18, corresponding to approximately -65dBm). The main control unit determines that the MESH signal strength (-82dBm) and the cellular network signal quality meet the first set judgment condition, such as meeting their respective threshold conditions. The main control unit then performs a switching operation: it triggers the cellular module to initiate a 5G voice call to intercom device B to establish a connection; after confirming the successful establishment of the 5G channel, it disables the MIC bias and DAC output of the MESH module's audio codec via GPIO control signals; simultaneously, it maintains the power supply to the MESH module's RF front-end and the operation of the protocol stack. Because the MESH audio channel is disabled while the MESH RF link remains active when switching to cellular intercom mode, the MESH network connection is not interrupted. Subsequent switchbacks do not require re-networking, significantly reducing the switchback speed without any noticeable impact on the user.
[0022] Subsequently, intercom device A communicates with intercom device B via the 5G network, and the MESH module continuously monitors the RSSI between A and B in the background.
[0023] While intercom device A is in 5G intercom mode, the main control unit continuously monitors the MESH signal strength. After a period of time, the distance between A and B shortens, and the MESH RSSI rises. The main control unit determines that the signal strength meets the second preset judgment condition. The main control unit then executes a switchback operation: it sends a command to the cellular module to turn off the 5G audio channel, but maintains the cellular network registration status; simultaneously, it sends a command to the MESH module to re-enable the MIC bias and DAC output. Intercom device A seamlessly switches back to MESH intercom mode.
[0024] This embodiment simultaneously monitors MESH signal strength and cellular network signal quality in MESH intercom mode and uses both as the basis for handover decisions. This allows the handover decision to comprehensively evaluate the availability of both networks, triggering a handover only when the MESH signal is poor and the cellular network signal is good. This avoids communication interruptions caused by mistakenly switching to cellular mode in areas without cellular network coverage. In addition, because the MESH signal strength is continuously monitored in cellular intercom mode and the system automatically switches back when the second set judgment condition is met, the system can automatically return to the low-latency, data-free preferred mode as soon as the MESH signal recovers, achieving adaptive optimization of communication resources.
[0025] In one embodiment, the first set judgment condition includes: the MESH signal strength is lower than the first predetermined threshold, and the cellular network signal quality is higher than the second predetermined threshold; The step of automatically switching the intercom device from the MESH intercom mode to the cellular network intercom mode includes: Trigger the intercom device to establish a cellular network intercom connection, so that the intercom device can conduct voice communication through the cellular network intercom mode; Close the audio input and audio output channels corresponding to the MESH intercom module.
[0026] Specifically, the first predetermined threshold is a critical value for judging whether the MESH signal has deteriorated, such as -75dBm. When the RSSI is lower than this value, MESH communication may be intermittent or unable to guarantee smooth calls. The second predetermined threshold is a critical value for judging whether the cellular network is available, such as CSQ≥10 or RSRP≥-80dBm, ensuring that a stable cellular voice connection can be established after handover. Establishing a cellular network intercom connection refers to initiating a voice call to the target device through the cellular communication module (such as a 5G module) built into the intercom device, or forwarding voice data through the user terminal APP. The audio input channel refers to the microphone signal transmission path, including the microphone itself, preamplifier, ADC (analog-to-digital converter), etc.; the audio output channel refers to the speaker signal transmission path, including DAC (digital-to-analog converter), power amplifier, speaker, etc.
[0027] In one specific embodiment, the following are set: a first predetermined threshold = -75dBm, and a second predetermined threshold = -80dBm (RSRP) or CSQ ≥ 10. If intercom device A detects a MESH RSSI of -85dBm (lower than -75dBm) with intercom device B, and simultaneously detects a cellular network RSRP of -65dBm (higher than -80dBm), the first predetermined judgment condition is met. The main control unit performs the following specific operations: The first step is to trigger the establishment of a cellular network intercom connection. The main control unit controls the cellular communication module to initiate a call. Simultaneously, a request is sent to the user terminal APP, which assists in establishing a VoIP channel. After a certain period of time, the call is connected, and the cellular network intercom channel is established.
[0028] The second step is to disable the MESH audio channel. The main control unit writes 0x00 (disables MIC bias) to register 0x0A and 0x00 (disables DAC output) to register 0x0C of the audio codec chip (model NAU88L25) via the I2C bus. The RF section of the MESH module remains unchanged and continues to monitor the channel. At this time, when the user speaks into intercom device A, the sound is encoded by the 5G module and sent to the base station, then forwarded to intercom device B via the core network; the sound from intercom device B is transmitted back to the speaker of intercom device A via the 5G network. Because the MESH audio channel is disabled by directly controlling the audio codec register, the switching of the audio channel can be completed within milliseconds, achieving an extremely low latency switching response.
[0029] In this embodiment, the first setting of the judgment condition requires both poor MESH signal and good cellular network signal, which avoids accidental switching to cellular mode when the cellular network signal is poor, ensures that communication can continue normally after the switch, and realizes the reliability of the handover decision. In addition, since the MESH audio channel is turned off while the radio frequency function is kept active, the MESH module can continue to monitor the signal strength, providing real-time data support for subsequent automatic handover.
[0030] In one embodiment, the second set judgment condition includes: the MESH signal strength is stably higher than a third predetermined threshold for a predetermined duration.
[0031] The third predetermined threshold is a critical value used to determine whether the MESH signal has recovered well. This value is usually set higher than the first predetermined threshold to form a hysteresis range. For example, if the first predetermined threshold is -75dBm, the third predetermined threshold can be set to -70dBm, a difference of 5dB.
[0032] It is understandable that "stable above" means that the MESH signal strength is consistently greater than the third predetermined threshold, rather than being momentarily higher, which can be confirmed through a debouncing algorithm; the predetermined duration refers to the shortest time for the signal to be continuously higher than the third predetermined threshold, for example, 3-5 seconds. This parameter is used to eliminate misjudgments caused by instantaneous signal fluctuations.
[0033] In one specific embodiment, the third predetermined threshold is set to -70dBm, the predetermined duration is set to 5 seconds, and the data acquisition interval is set to 500ms. The MESH signal strength parameter RSSI is continuously acquired. If the MESH signal strength is detected to be greater than -70dBm for 10 consecutive times, the MESH signal is confirmed to have stabilized and the system switches back to MESH intercom mode. If the MESH signal strength parameter RSSI is detected to briefly drop to -71dBm at the third second, the counter is reset to zero, and the count is restarted. This continuous sampling and counting method for dejitter reduction makes the algorithm simple and efficient, easily implementable on resource-constrained embedded main control units, and offers the advantage of low computational overhead.
[0034] In this embodiment, the third predetermined threshold is set higher than the first predetermined threshold, so that a hysteresis interval is formed between the handover and the return switch, which effectively avoids communication instability caused by frequent handover when the signal fluctuates near the threshold. At the same time, since the signal is required to be stable above the threshold for more than a predetermined duration, a brief signal recovery will not trigger unnecessary return switches. The return switch operation is only performed after it is confirmed that the MESH signal strength has truly stabilized and recovered, thereby achieving the stability and reliability of the handover.
[0035] In one embodiment, if the intercom device group includes multiple intercom devices, the method further includes: Monitor the MESH signal strength and cellular network signal between any two of the aforementioned intercom devices; If the MESH signal strength and cellular network signal quality between two target intercom devices meet the first set judgment condition, the intercom mode between the target intercom devices is switched to the cellular network intercom mode. Maintain the intercom mode between other intercom devices besides the target intercom device as MESH intercom mode, and control the intercom mode between the target intercom device and the other intercom devices as MESH intercom mode.
[0036] The target intercom device refers to a specific pair of devices in the group whose MESH signal quality has deteriorated and need to be switched to cellular intercom mode; maintaining MESH intercom mode means that for intercom device pairs with good signal, MESH will continue to be used for voice communication without any switching.
[0037] The control mode for communication between the target intercom device and other intercom devices is set to MESH intercom mode. This means that when the target device is switched to cellular intercom mode, it still communicates with other devices in the group with good signal through MESH intercom mode, rather than through the cellular network. This means that the target device maintains the MESH radio frequency connection at the same time (only the audio is switched from MESH to cellular for communication with the other party).
[0038] In one specific embodiment, there is a MESH intercom group containing six devices: A, B, C, D, E, and F. Initially, all devices are in MESH intercom mode, and any two devices can communicate via single-hop or multi-hop relay for voice. If the signal between A and B, C, D, and E is good, but the distance between A and F is too great and obstructed by obstacles, making multi-hop relay ineffective, a cellular network intercom bridging channel is established between A and F. A's MESH module continues to maintain MESH communication with B, C, D, and E, while A and F communicate via the cellular network for voice. Specifically, when A generates intercom audio, its microphone input is simultaneously split into two streams: one stream is sent to the MESH module and transmitted to B / C / D / E, and the other stream is sent to the cellular communication module and transmitted to F. A's speaker simultaneously receives two audio streams: one from the MESH module (the audio from B / C / D / E) and the other from the cellular communication module (the audio from F). B, C, D, and E are completely unaffected; their communication with each other, as well as with A and F, continues through the low-latency MESH intercom module.
[0039] This embodiment enables cellular network bridging for specific devices in the group that have signal problems, allowing other devices in the group with good signals to continue using low-latency MESH intercom. Compared to the method of switching the entire group, this can save cellular traffic, and the larger the group size, the more obvious the saving effect. At the same time, since the target device communicates with the other party through cellular while still maintaining communication with other devices in the group through MESH, the connectivity of the entire group is not affected, achieving balanced optimization of overall communication quality.
[0040] In one embodiment, the above method further includes the following steps: Real-time monitoring of whether the user terminal has call events; If a call event occurs, mute the audio input and output channels corresponding to the MESH intercom module; If the call event is detected to have ended, the audio input and output channels of the intercom device are restored to the intercom mode before the switch.
[0041] In this context, a call event refers to a traditional circuit-switched call or a VoLTE / VoNR call conducted by a user terminal (such as a mobile phone), excluding voice communication within the intercom application. Mute refers to disabling the microphone input and speaker output of the MESH intercom module, making the MESH channel completely silent, while maintaining the MESH RF link activation and cellular network intercom connection. When a call event ends, the intercom mode is switched back to the mode state before the call was connected (which could be MESH intercom mode or cellular network intercom mode). This state needs to be recorded before mute so that it can be restored after hanging up.
[0042] In one specific embodiment, if a call event occurs, the following steps are performed in sequence: 1) User terminal A receives a call from a third party and rings.
[0043] 2) The user presses the answer button on user terminal A, and the call is connected.
[0044] 3) The main control unit of intercom device A detects that the Bluetooth HFP audio data stream has started to be transmitted and recognizes the call event.
[0045] 4) Before performing the mute operation, the main control unit of intercom device A first writes the current mode "MESH intercom mode" into a non-volatile register for storage.
[0046] 5) The main control unit sends a command to the MESH module to disable the MIC bias and DAC output via the audio codec. The MESH RF continues to operate.
[0047] 6) Users make private phone calls. During this time, the microphone of intercom device A will not pick up the call audio and send it to the MESH group, and the audio from the MESH group will not be broadcast from the speaker of intercom device A, achieving complete audio isolation.
[0048] 7) The user hangs up the phone. The main control unit detects the termination of the Bluetooth HFP audio data stream, and the call ends.
[0049] 8) The main control unit reads the previously saved mode identifier and confirms that it is MESH intercom mode.
[0050] 9) The main control unit sends a command to the MESH module to restore the MIC bias and DAC output.
[0051] 10) Intercom device A resumes MESH intercom mode, and users can continue to participate in group calls.
[0052] If the intercom device was in cellular intercom mode before switching, the main control unit will reopen the audio channel of the cellular communication module and continue to maintain cellular intercom mode when the call is disconnected and then resumed.
[0053] This embodiment assigns the highest priority to call events, ensuring that regardless of the intercom mode, the incoming call automatically obtains exclusive access to the audio channel, guaranteeing absolute privacy. Telephone audio will not interfere with the intercom channel, and intercom audio will not disrupt the call. Furthermore, since the mute operation only affects the audio channel while maintaining MESH RF link activation and cellular network registration, communication can be resumed without re-establishing the intercom connection after the call is disconnected, achieving a millisecond-level rapid recovery experience. In addition, because the current intercom mode state is recorded before mute, the system can accurately restore the original mode after disconnection, providing a completely consistent user experience with no noticeable difference to the user.
[0054] In one embodiment, the real-time monitoring of whether a call event exists on the user terminal includes the following steps: If a telephone audio data stream is detected on the user terminal, the audio data stream status defined by the Bluetooth hands-free specification is identified. Determine whether the telephone call is connected based on the status of the audio data stream. If the call is in the connected state, a call event is determined to exist; otherwise, a call event is determined not to exist.
[0055] The Bluetooth Hands-Free Profile (HFP) is a standard specification defined in the Bluetooth protocol stack, used to implement call control and audio transmission between mobile phones and hands-free devices. The audio data stream state refers to the Synchronous Connection Oriented (SCO) link state defined by the HFP. When an SCO / eSCO link is established, it indicates that a phone call is in progress; when the link is released, it indicates that the call has ended.
[0056] Specifically, determining whether a phone call is connected is based on the status of the audio data stream. In other words, the main control unit of the intercom device determines the call status by monitoring the HFP status API provided by the Bluetooth module or by parsing HFP protocol layer instructions.
[0057] This embodiment utilizes the audio data stream status of the Bluetooth HFP standard to identify call events, enabling a low-cost and highly compatible solution without the need for additional hardware circuits or proprietary protocols, and is compatible with the vast majority of smartphones on the market. At the same time, since the connection and disconnection of the call are determined by monitoring the SCO link establishment and release events, the accuracy of status recognition is high, avoiding erroneous mute or recovery failure caused by signal misjudgment.
[0058] In one embodiment, during the switch from the MESH intercom mode to the cellular intercom mode, the MESH audio channel and the cellular audio channel of the intercom device are kept open simultaneously. The two audio signals are cross-faded in and out and then the MESH audio channel is turned off. The cross-fading in and out fusion duration is dynamically adjusted according to the current audio delay of the intercom device, and during the fusion, the gain of the MESH audio channel gradually decreases while the gain of the cellular audio channel gradually increases.
[0059] As is understandable, crossfade is a smooth transition technique widely used in professional audio processing. It involves simultaneously outputting two audio signals during a transition, with one signal's amplitude gradually decreasing according to a preset curve (fade-out) and the other signal's amplitude gradually increasing (fade-in). The two signals overlap for a certain period and are then weighted and superimposed according to their respective gain coefficients before being output. Because the human ear is not sensitive to slow changes in sound amplitude but is extremely sensitive to abrupt changes in sound pressure level, crossfade effectively masks these sudden audio abrupt changes during transitions, ensuring that the user cannot hear any stutters, pops, or interruptions.
[0060] In the dual-mode switching scenario of this application, when the main control unit determines that it needs to switch from MESH intercom mode to cellular network intercom mode, the traditional approach is to immediately disconnect the MESH audio channel and then establish the cellular audio channel. During this process, the audio stream will be completely blank for tens or even hundreds of milliseconds, and the user will clearly feel the sound break or hear a harsh clicking sound.
[0061] Specifically, in this embodiment, the main control unit does not immediately shut down the MESH audio channel when the switching is triggered. Instead, it keeps both channels open simultaneously and initiates a cross-fade-in / fade-out process. The fusion duration, i.e., the total duration of the fade-in / fade-out, is not a fixed value. Instead, it is dynamically calculated based on the overall latency of the current audio processing link of the intercom device. The lower limit is set to 5 milliseconds, and the upper limit is set to 80 milliseconds. This is based on the fact that the human ear cannot perceive audio interruptions below 5 milliseconds, while users will clearly feel the delay when it exceeds 80 milliseconds. At the same time, one-quarter or one-third of the current latency is taken as the fusion duration. This ratio is based on typical empirical values of cross-fade-in / fade-out technology and can achieve a balance between smooth transition and fast switching.
[0062] The sources of audio latency are very complex, including microphone acquisition latency, which is typically 2-10 milliseconds depending on ADC sampling and buffering; MESH voice codec latency, which can be 5-20 milliseconds in low-latency mode using the Opus codec and can reach 20-40 milliseconds using AMR-WB; wireless transmission latency, which is about 5-10 milliseconds in single-hop MESH transmission and increases by about 5-10 milliseconds with each additional hop in multi-hop relay; cellular uplink and downlink latency, which is typically 20-50 milliseconds in 4G networks and can be as low as 10-20 milliseconds in 5G, but may fluctuate due to signal quality; receiver decoding latency; and speaker playback latency, which is about 10-20 milliseconds.
[0063] The main control unit can estimate the overall audio latency at the current moment by measuring the end-to-end timestamp difference of audio frames in real time or by analyzing the timestamp field in the RTP protocol. Based on this latency value, the fusion duration can be dynamically calculated, for example, taking one-quarter or one-third of the current latency, but limited to a minimum of 5 milliseconds and a maximum of 80 milliseconds. When the latency is large, such as due to increased buffering caused by MESH multi-hop relays or poor cellular network signals, the fusion duration is extended accordingly to ensure that the continuity of audio is not compromised by excessively short fusion; when the latency is small, the fusion duration is shortened to quickly complete the switching and release audio resources.
[0064] During the fusion process, the master control unit controls the gain register of the audio digital signal processor (DSP) or audio codec to gradually decrease the gain of the MESH audio channel from its maximum value to zero according to a predetermined gradient curve, while simultaneously increasing the gain of the cellular audio channel from zero to its maximum value. The actual output amplitude of each signal at any given moment is equal to the original signal amplitude of that channel multiplied by the current gain coefficient. The two signals are then superimposed in the time domain and output to the speaker. The choice of gain curve affects the auditory smoothness; a cosine curve, due to its continuous derivative, avoids the clicking sound introduced by sudden gain changes and is therefore a preferred choice. Before fusion begins, the master control unit needs to synchronously adjust the audio routing table to input both signals into the mixer; after fusion, the MESH channel is removed from the routing table, and the cellular channel takes over completely.
[0065] If a user speaks during the merging process, the voice signal will be transmitted through both channels simultaneously. However, due to the different gain weights, the signal received at the far end may contain slight timbre changes, similar to a progressively changing filter, but no voice content will be lost. Furthermore, if an anomaly occurs during the handover process, such as a cellular connection failure, the main control unit can immediately abort the merging, quickly restore the gain to the full capacity of the MESH channel, and resume the original communication, preventing call interruptions due to handover failure.
[0066] By employing the above methods, the defects of traditional hard switching solutions, such as audio interruptions, stuttering, and popping sounds, are completely eliminated, achieving truly seamless and smooth switching. Through dynamic adjustment of the fusion duration, this solution can adapt to different network environments and device states. In low-latency scenarios, the fusion duration can be shortened to five to ten milliseconds for almost instantaneous switching, while in high-latency scenarios, the fusion duration is extended to fifty to eighty milliseconds to ensure audio continuity. Furthermore, no additional hardware costs are required; it can be achieved solely through software control of the audio DSP or codec gain register, resulting in extremely low resource consumption on the embedded platform. Moreover, in group intercom scenarios, when multiple devices switch modes simultaneously or sequentially, auditory confusion caused by asynchronous switching is avoided. Because each device independently executes crossfade-in and crossfade-out, the group sound heard by the remote user is always continuous, without the phenomenon of multiple sounds overlapping and intersecting.
[0067] like Figure 5 In one embodiment, motion state information of the intercom device is acquired, the MESH signal change trend in the future time period is predicted based on the motion state information, and the connection establishment process of the cellular network intercom mode is initiated in advance based on the prediction result; the motion state information includes at least one of movement speed, movement direction and acceleration change, and when it is predicted that the MESH signal will drop below the switching threshold within a preset time window, the start time of the connection establishment process of the cellular network intercom mode is advanced accordingly.
[0068] It should be noted that motion status information refers to a series of parameters describing the physical motion state of the intercom device, including movement speed, direction of movement, and acceleration changes. Movement speed can be obtained through various sensors. Outdoors, when GPS signal is available, the GPS module can directly output the speed. In indoor environments or mines without GPS, the speed can be estimated by integrating the acceleration using an inertial measurement unit (IMU), which includes a three-axis accelerometer and a three-axis gyroscope, or indirectly by calculating the relative speed through the round-trip time change of the mesh signal. The direction of movement can be obtained from an electronic compass (magnetic meter) or calculated by differential calculation between consecutive position points. Acceleration changes reflect the degree of speed change; for example, sudden acceleration or deceleration often indicates that the device is about to enter a signal dead zone or is rapidly approaching a relay station.
[0069] Understandably, based on this motion state information and the known deployment locations of MESH relay stations (e.g., a relay station is placed at regular intervals in a mine roadway with its coordinates pre-configured in the equipment), the main control unit can predict the future trajectory of the equipment relative to the nearest relay station. The specific prediction method can employ a kinematic model, estimating the future displacement Δs = v0×t + 0.5×a×t² based on the current velocity v0, acceleration a, and time t. Then, combining this with the current distance d0 estimated from the MESH signal's round-trip time or received signal strength, the distance at time t is predicted as d(t) = |d0-Δs×cosθ|, where θ is the angle between the direction of motion and the direction towards the relay station. The strength of the MESH signal is typically inversely proportional to the distance, following a logarithmic path loss model RSSI(d) = RSSI(d0)-10×n×log10(d / d0) in a real-world environment, where n is the path loss exponent, typically 2-4.
[0070] In one feasible implementation, the future MESH signal strength RSSI(t) can be estimated based on the predicted distance d(t). When the predicted result shows that RSSI(t) will drop below the handover threshold, for example -75dBm, within a future preset time window T that can be adaptively adjusted according to the device's movement speed, the main control unit immediately initiates the connection establishment process for the cellular network intercom mode. This window T can be set to 2-5 seconds. 2 seconds is suitable for fast-moving scenarios to avoid handover lag, and 5 seconds is suitable for slow-moving scenarios to reduce unnecessary prediction calculations. The setting is based on the fact that motion prediction algorithms have high accuracy in short periods, while uncertainty increases significantly beyond 5 seconds, derived from engineering experience regarding integral drift in the field of inertial navigation. The handover threshold is set based on the lower limit of the tolerable signal strength for MESH intercom; below this value, voice quality significantly deteriorates and packet loss rate increases. This process includes waking up the dormant cellular communication module, initiating an attach request to an unattached base station, completing LTE or NR registration, establishing a PDU session or EPS bearer, obtaining an IP address, and establishing a signaling connection with the intercom server. These steps typically take hundreds of milliseconds to 1 or 2 seconds. Starting them in advance ensures that the cellular connection is already in a ready state when the MESH signal really deteriorates.
[0071] Understandably, the lead time for startup can be dynamically determined based on the predicted attenuation rate. The faster the attenuation, such as when the device moves away from the relay station at high speed or moves away from the relay station in a direction that is almost completely away from the relay station, the larger the lead time should be. For example, if it is predicted that the signal will fall below the threshold in 1 second, cell connection establishment can be started immediately. If the attenuation is slower, such as when the device moves away slowly or moves away from the relay station at a large angle, the lead time can be appropriately reduced, but it should usually be at least one full connection establishment time in advance, such as 500 milliseconds.
[0072] Furthermore, motion status information can be used to determine the necessity of a handover. If the device is rapidly returning to the relay station (i.e., its speed direction is towards the relay station and the distance is decreasing), even if the current signal is weak, the handover can be postponed until the signal recovers, avoiding unnecessary handovers and reconnections. In cellular intercom mode, the predictive mechanism also applies. When it is predicted that the MESH signal will recover well at some point in the future (i.e., the device is approaching the relay station), the MESH module's audio channel preparation can be initiated in advance, or a re-association of the MESH network can be initiated in advance to speed up the reconnection. The intercom device needs to know the approximate location information of the relay station within the mine in advance. This can be achieved through pre-setting at the factory, synchronization via the MESH network, or distribution via a server.
[0073] Ultimately, the shift from passive response to proactive prediction fundamentally shortens the actual handover interruption time. By pre-warming up the cellular connection, the handover is completed almost instantaneously, completely imperceptible to the user. Secondly, using motion state information for prediction avoids the limitations of extrapolating solely from historical signal strength. Historical signal strength can fluctuate drastically in a short time due to obstruction, interference, or rapid fading, while motion state information provides a more macroscopic and deterministic trend basis, greatly improving prediction accuracy. Thirdly, it is particularly suitable for high-speed movement scenarios, such as intercom devices on mining trucks, motorcyclists, and fast-moving inspection personnel. In these scenarios, the mesh signal can change from good to extremely poor in a very short time; pre-prediction and pre-warming up the cellular connection are key measures to ensure communication continuity. Fourthly, it can also work in conjunction with cross-fade-in / fade-out technology, initiating the cellular audio channel fade-in in advance when an impending handover is predicted, further smoothing the transition. Fifthly, it has low hardware requirements; motion sensors are widely integrated into smart intercom devices, requiring only a motion prediction module added at the software level, resulting in minimal computational overhead.
[0074] In one embodiment, the MESH signal strength is monitored with a first preset monitoring period and the cellular network signal quality is monitored with a second preset monitoring period, wherein the first preset monitoring period is shorter than the second preset monitoring period; the first preset monitoring period and the second preset monitoring period are dynamically set according to the environmental change rate of the intercom device, and the first preset monitoring period and the second preset monitoring period are shortened accordingly when the environmental change rate is higher.
[0075] In one possible implementation, the differentiated monitoring period stems from the significant differences in the physical characteristics and variation patterns of MESH signals and cellular network signals. MESH intercom relies on direct wireless links between devices, and its signal strength is affected by factors such as relative distance between devices, obstructions, multipath effects, and antenna directivity, resulting in very rapid changes. For example, in a mine tunnel, if a user carrying an intercom quickly turns a corner, the MESH signal may drop by 10-20 dB within 50-100 milliseconds; when the user passes through a metal fire door, the signal may attenuate by more than 30 dB instantaneously. Without high-frequency monitoring, the handover opportunity will be missed, leading to call interruption. Therefore, MESH signals require high-frequency monitoring, and the first preset monitoring period can be set to a value between 50-200 milliseconds, with a typical value of 100 milliseconds.
[0076] It's important to note that cellular network signals, such as 4G or 5G, originate from macro base stations or pico base stations. Their coverage radius can range from hundreds of meters to several kilometers, and signal strength changes relatively gradually. Unless a device moves rapidly away from the base station, the signal will not fluctuate drastically in a short period. Even in mobile environments, the fading rate of cellular signals is typically much lower than that of mesh signals. Therefore, cellular signals can be monitored at a lower frequency, with the second preset monitoring period set to 500-2000 milliseconds, typically 1 second. This differentiated design ensures timely monitoring while avoiding unnecessary computational and power consumption overhead, as high-frequency monitoring requires frequent processor wake-ups, signal value readings, and data processing, significantly increasing power consumption.
[0077] Furthermore, the monitoring cycle can be dynamically adjusted based on the environmental change rate. The environmental change rate is a comprehensive indicator used to quantify the dynamic nature of the environment in which the intercom equipment is located. Its calculation methods can include the intercom equipment's moving speed (provided by IMU or GPS; the higher the speed, the greater the environmental change rate); the variance or volatility of the historical sequence of MESH signal strength; the more drastic the fluctuations, the greater the environmental change rate; the rate of change of MESH network topology, such as the update frequency of the neighbor node list; the faster the update, the faster the equipment is crossing the coverage of different relay stations; and the rate of change of acceleration. Sudden acceleration or deceleration often indicates that the equipment is about to enter a signal-sensitive area.
[0078] The main control unit can periodically calculate the current value of the environmental change rate, for example, every second, and normalize it to a range of 0%-100%. Then, it adjusts the monitoring cycle according to a preset mapping relationship or fuzzy control rules. For example, when the environmental change rate is below 20%, the device is considered stationary or moving slowly, and the MESH monitoring cycle can be extended to 300 milliseconds, and the cellular monitoring cycle to 2 seconds; when the environmental change rate is between 20% and 60%, a medium cycle is used, i.e., 150 milliseconds for MESH and 1 second for cellular; when the environmental change rate is above 60%, the device is considered to be in a state of rapid movement or severe signal jitter, and the MESH monitoring cycle is shortened to 50 milliseconds, and the cellular monitoring cycle is shortened to 500 milliseconds. The aforementioned cutoff points of 20% and 60% are derived from statistical analysis of the intercom device's movement speed. Low-speed movement (speed less than 1 meter per second) typically results in an environmental change rate below 20%, medium-speed movement (speed between 1 and 3 meters per second) results in a change rate between 20% and 60%, and high-speed movement (speed greater than 3 meters per second) results in a change rate exceeding 60%. The specific value of the monitoring period is based on the test results of the balance between embedded processor power consumption and response time. 50 milliseconds can ensure fast response with limited increase in power consumption, while 300 milliseconds significantly reduces the processor wake-up frequency.
[0079] The adjustment process requires incorporating hysteresis and filtering to prevent frequent cycle changes due to instantaneous fluctuations. For example, rising and falling hysteresis can be set, and the monitoring cycle should only be changed when the environmental change rate exceeds the current interval boundary by a certain margin. Furthermore, the monitoring accuracy or the number of sampling points can be adjusted. For instance, under high change rates, multiple samplings and averaging at each monitoring point can be increased, such as taking the median of three consecutive samplings, to improve measurement reliability; under low change rates, the number of samples can be reduced to save energy. Additionally, when the device is in MESH intercom mode and actively engaged in a call, the monitoring frequency can be appropriately increased, for example, by multiplying the base cycle by 0.5, because timely switching is more critical at this time; while in standby or silent periods, the monitoring frequency can be reduced to extend battery life. This activity-adaptive strategy can be combined with environmental change rate adaptive strategies.
[0080] Ultimately, by differentiating monitoring cycles, the system's computational overhead and power consumption are significantly reduced while ensuring timely switching decisions. Real-world testing shows that compared to a uniform, fixed 100-millisecond monitoring cycle, this solution reduces power consumption while maintaining unaffected switching response speed. Secondly, the adaptive environmental change rate mechanism allows the monitoring strategy to dynamically adapt to different usage scenarios and movement states. In static or slow-moving scenarios, the monitoring frequency is reduced to extend standby time; in fast-moving or signal-fluctuating scenarios, the monitoring frequency is automatically increased to ensure timely switching decisions. Finally, it can also work in conjunction with other intelligent strategies. For example, when movement predicts an approaching signal blind zone, the monitoring frequency is temporarily increased even at low change rates to cope with sudden signal degradation. Moreover, the implementation cost of this monitoring strategy is extremely low, requiring only the addition of dynamic timer adjustment logic at the software level, without the need for additional hardware support.
[0081] In one possible implementation, it also helps to reduce the total power consumption in multi-device group scenarios, because each device can independently optimize its own monitoring frequency, avoiding excessive power consumption of the entire group due to uniform high-frequency monitoring.
[0082] In one embodiment, before the first set judgment condition is met, a registration state is pre-established with the cellular network and a background connection is maintained; when the first set judgment condition is met, the audio bearer of the background connection is activated to complete the switching; the pre-established registration state includes at least the pre-configuration of authentication parameters, bearer resources and network synchronization information, and the pre-established registration state is periodically refreshed to maintain its validity while the intercom device is in MESH intercom mode.
[0083] Specifically, pre-establishing a registration state with the cellular network refers to the proactive attachment and registration process with the cellular network by the intercom device while it is still in MESH intercom mode and the MESH signal has not deteriorated to the point where a handover is required. A complete cellular network registration typically includes: powering on the cellular communication module and searching for suitable cells; completing downlink synchronization with the base station and reading the System Information Block (SIB); initiating the Random Access Procedure (RACH); sending an attach request to the Mobility Management Entity (MME) or Access and Mobility Management Function (AMF), carrying the device identifier, capability information, etc.; the network side authenticating and encrypting the device and negotiating a security context; the network side assigning a temporary identifier such as GUTI, TMSI, and IP address to the device and establishing a default EPS (Evolved Packet System) bearer or PDU (Protocol Data Unit) session; and the device replying that the attach is complete and entering the registration state.
[0084] It should be noted that after completing these steps, the intercom device is already in a registered or idle state on the cellular network side, and the network side retains information such as user context, mobility management context, and session management context. However, at this time, the audio bearer, i.e., the dedicated bearer used for voice transmission, has not yet been activated. The intercom device will not send or receive voice data through the cellular network, but only maintains lightweight signaling such as periodic Tracking Area Update (TAU) to avoid the network releasing resources. This state is called background connection. When the first setting condition is met, i.e., the MESH signal deteriorates and the cellular signal is good, the master control unit only needs to send a service request or bearer activation command to the network to upgrade the existing default bearer to a dedicated bearer and start transmitting voice data. The signaling interaction in this process involves only a few messages such as service requests and bearer establishment responses, and the total time is usually between 50-100 milliseconds, which is much less than the 1-2 seconds of the complete registration process.
[0085] To maintain the validity of the backend connection, the intercom device needs to periodically send refresh messages to the network. If the device does not interact with the network for an extended period, the MME or AMF may assume the device has left or dropped, implicitly unattaching the device, releasing all resources, and causing subsequent activation failure. The refresh mechanism can adopt the Tracking Area Update (TAU) procedure defined in the standard 3GPP protocol, or periodic uplink NAS transmissions, such as sending empty data packets. The refresh interval can be dynamically adjusted according to the actual situation. When the MESH signal is good, the device is unlikely to need to switch, and the refresh interval can be extended, for example, to 60 seconds, to reduce unnecessary signaling overhead and power consumption. When the MESH signal is close to the handover threshold, the refresh frequency is increased, for example, to 15 seconds, to ensure the pre-registration state is valid during handover. The 60-second refresh interval is a typical empirical value based on the cellular network bearer keep-alive mechanism, achieving a reasonable balance between signaling overhead and connection validity. It avoids excessive power consumption due to overly frequent refreshes and maintains the validity of the pre-registration state in most scenarios. The 15-second refresh interval is used when the MESH signal is about to deteriorate, ensuring the pre-registration state remains valid throughout the shortest handover preparation time through more frequent refreshes.
[0086] Furthermore, after the intercom device switches back from cellular intercom mode to MESH intercom mode, the pre-registration status in this embodiment can be maintained for quick activation when needed next time. However, if it remains in MESH mode for an extended period and the signal is stable, the cellular pre-registration can be released to save power, and re-established when the signal deteriorates again. Key information that needs to be stored during pre-registration includes the assigned IP address, APN (Access Point Name), QoS (Quality of Service) parameters such as QCI value, security context such as Key Set Identifier (KSI), and encryption algorithm. This information can be stored in the intercom device's non-volatile memory for quick recovery after device restart or prolonged network outage, without needing to retrieve it from the network again.
[0087] In addition, this embodiment can also be combined with motion state prediction. When it is predicted that a switch is about to be needed, not only can a pre-registration be established in advance, but also a TAU refresh can be performed in advance to ensure that the pre-registration state is in the latest valid state.
[0088] By employing pre-registration and background connection, the actual connection establishment latency for cellular network intercom is reduced from seconds to milliseconds, shortening the voice interruption time during handover to a level imperceptible to the human ear, achieving truly seamless handover. Secondly, the periodic refresh mechanism ensures that the pre-registration status does not expire due to prolonged inactivity, avoiding the embarrassing situation of discovering the cellular connection is unavailable during emergency handover. Thirdly, this embodiment reduces the signaling impact on the cellular network because the pre-registration and activation method shifts most signaling overhead to off-peak hours, requiring only a small amount of signaling during handover, making it more network-friendly. Fourthly, this embodiment does not rely on special operator functions, only needs to comply with standard 3GPP protocols, and has good versatility and compatibility.
[0089] In one embodiment, the MESH radio frequency link of the intercom device is kept active, and the microphone signal input and speaker signal output are disabled only by controlling the input and output switches of the audio codec; while closing the MESH audio channel, the routing entry corresponding to the MESH audio channel in the audio routing table of the intercom device is marked as disabled, and the audio data stream is redirected to the cellular network audio channel.
[0090] In this embodiment, MESH radio frequency link refers to the collective term for the hardware and software stack of the MESH module in the intercom device used for transmitting and receiving wireless signals, including radio frequency front-end such as antenna, power amplifier, low noise amplifier, filter, etc., baseband processor, MAC layer protocol stack, network layer routing table, and neighbor management module for maintaining network topology.
[0091] During the handover process, although the MESH channel is no longer used to transmit audio, the MESH RF link must remain active in order to continuously monitor the MESH signal strength for a quick re-handover and to maintain MESH network connectivity with other devices in the group. For example, in a fine-grained handover strategy, the target device may still need to communicate with other unhandover devices through MESH.
[0092] It's important to note that an audio codec is typically a standalone chip or an audio interface unit integrated into a processor. It contains multiple control registers used to configure microphone bias voltage, the enable status of the analog-to-digital converter (ADC) and digital-to-analog converter (DAC), the switching of input / output channels, volume gain, and more. The main control unit can independently disable the microphone input path and speaker output path by writing control words to specific registers of the audio codec via the I2C (Inter-Integrated Circuit) or SPI (Serial Peripheral Interface) bus. For example, writing zero to the microphone bias register turns off the microphone bias voltage, thus disabling microphone signal acquisition; writing zero to the DAC output register mutes the DAC output. These operations only affect the audio signal path and do not affect the power supply, clock, or operating status of the MESH RF section.
[0093] Simultaneously, the main control unit also needs to update its internal audio routing table. The audio routing table is a logical data structure within the main control unit that records the directional mapping of audio data streams. It typically includes which module (such as the MESH module, cellular module, Bluetooth module, or recording module) the PCM data collected by the microphone should be sent to; and which playback device (such as a speaker, headphones, or Bluetooth headset) the audio streams from each module should be output to. When the MESH audio channel is closed, the main control unit marks the microphone source routing entry corresponding to the MESH audio channel in the audio routing table as disabled, meaning microphone data is no longer sent to the MESH module; it redirects the microphone data stream to the cellular network audio channel, ensuring that subsequent user speech is automatically sent to the cellular module; for downlink, it routes the audio stream from the cellular network module to the speaker, while simultaneously preventing the audio stream from the MESH module from being routed to the speaker or muting it; and it adds a new entry to the routing table indicating that the speaker's audio source is now the cellular module. The entire routing redirection process is completed in microseconds, without any perceptible delay.
[0094] Furthermore, to support rapid switchback, the main control unit can pre-reserve the configuration parameters of the MESH audio channel, such as selectable sampling rates of 8000, 16000, and 48000Hz, volume levels, equalizer settings, and noise reduction parameters. When a switchback is needed, the main control unit simply needs to re-enable the input / output switches of the audio codec and restore the entries in the audio routing table to their original state to restore the MESH audio channel. To further reduce switchback latency, it can be combined with the pre-registration concept. In cellular mode, although the MESH audio channel is turned off, the MESH RF link remains active, so the MESH network connection is always online, and there is no need to rebuild the network during a switchback, further shortening the switchback time.
[0095] By decoupling the audio channel and RF link, fine-grained control of the MESH module is achieved. This cuts off the audio path to avoid interference and resource conflicts with the cellular channel; for example, sound picked up by the microphone will not be transmitted through both MESH and cellular channels simultaneously, preventing echoes, and the speaker will not play two audio streams at the same time, causing confusion. At the same time, the RF function is retained for rapid switchback and signal monitoring. In group communication, the target device may only need to switch communication with a specific device to cellular mode, but still needs to maintain connectivity with other group members through the MESH RF. The continuous activation of the MESH RF allows the device to perform MESH and cellular communication simultaneously, achieving true dual-mode concurrency. Because the MESH RF link remains active, the device can continuously monitor the MESH signal strength, providing real-time data for switchback decisions and avoiding the delay of having to wait for signal detection again during switchback. Understandably, it can also be compatible with other audio processing strategies. For example, when the MESH audio channel is turned off, a low-volume monitoring channel can be kept open at the same time to detect whether other members are speaking in the MESH group, thereby achieving more intelligent group management.
[0096] like Figure 6 In one embodiment, voice activity statistics of the group to which the intercom device belongs are obtained, and the group is determined to be in a silent period based on the voice activity statistics. During the silent period, a switching operation from the MESH intercom mode to the cellular network intercom mode is performed. The voice activity statistics include at least one of the following: the average voice frame energy of each intercom device in the group, the voice activity density, and the continuous silence duration. When the continuous silence duration exceeds a preset duration threshold, it is determined that the group has entered a silent period.
[0097] Understandably, voice activity statistics are real-time voice transmission status data collected by intercom devices through a MESH network from all members within a group. In a MESH intercom group, voice communication is typically half-duplex, meaning only one person speaks at a time while others listen; the group is silent when no one is speaking. To effectively detect silence, each intercom device, when sending a voice frame, can include a Voice Activity Indicator (VAD) flag in the frame header or auxiliary fields, in addition to the encoded audio data itself, or directly attach a short-time voice energy value quantized between 0 and 255. When the device receives voice frames from other members, it parses these indications to determine whether anyone in the group is speaking.
[0098] To track the duration of continuous silence, each device maintains a global group silence timer. The timer resets to zero when any member's voice frame is received (i.e., the VAD flag is valid) or the energy value exceeds the background noise threshold. The timer increments when consecutive silence frames are received (i.e., the VAD flag is invalid) or when no voice frames are received for more than one frame. When the timer value exceeds a preset duration threshold configurable according to the actual scenario, such as 2 to 5 seconds, the system determines that the group has entered a silent period, meaning no one is currently speaking. This threshold is based on the average duration of natural pauses in human conversation, which is approximately 1.5 to 2 seconds. A pause exceeding 2 seconds effectively distinguishes between brief breathing pauses and true silence, preventing accidental switching due to short pauses. The upper limit of 5 seconds is suitable for noisy environments or scenarios where multiple people take turns speaking, preventing frequent switching attempts during conversation breaks.
[0099] Specifically, speech activity density refers to the percentage of frames with speech within a recent time window, such as the first 10 seconds. If the density is below 5%, it indicates that few people in the group are speaking, which can serve as an auxiliary criterion for determining the silence period. The average energy of speech frames reflects the average volume level within the group. If the average energy is below a preset background noise threshold, such as -50 dBFS, it also indicates a possible state of silence or extreme quiet. Combining multiple indicators improves robustness and avoids false triggering of the silence timer due to brief pauses in speech, such as when a user is breathing or thinking, or misjudgments caused by weak speech triggering the VAD (Voice over Action) due to low density from individual devices. By using a weighted comprehensive score, such as 60% for silence duration, 30% for activity density, and 10% for average energy, the true silence period can be determined more accurately.
[0100] Once a silent period is detected, the main control unit triggers a switchover from MESH mode to cellular mode. Since no one is speaking at this time, the switchover operation will not cause any voice interruption or missing words, and the user will be completely unaware of it. If, during the switchover process, such as when establishing a cellular connection or performing crossfade-in / fade-out, someone in the group starts speaking and a valid voice frame is received, the system can immediately stop the switchover and revert to MESH mode to maintain the original communication; or it can continue to complete the switchover but speed it up, for example, by shortening the fade-in / fade-out duration to 5 milliseconds, because at the beginning of speaking, the user may not notice the brief and slight change in sound quality; or, if speaking continues after the switchover is completed, crossfade-in / fade-out technology can be used to ensure a smooth audio transition.
[0101] Generally, the first strategy is the simplest and most reliable because the rollback operation is quick, requiring only the reset of some state variables, and does not cause voice loss. To ensure a high success rate for the handover, a handover window can be set, such as waiting a short time, like 100 milliseconds, after determining the silence period. If voice activity does not resume during this period, the handover is officially executed; once it resumes, the handover is rolled back, and the system waits for the next silence period before attempting again. Furthermore, the system can predict suitable handover times based on the group's historical voice patterns, such as using machine learning to predict peak call times, avoiding initiating handovers during periods when there is a high probability of someone speaking.
[0102] The protection of the call experience is reflected in the use of the natural voice pauses, i.e., the silence period, in group communication to perform handover, fundamentally avoiding any impact on the ongoing call. Secondly, by comprehensively judging multiple voice activity statistical indicators, such as energy, density, and silence duration, the accuracy and robustness of silence period judgment are improved, effectively distinguishing between true silence periods and brief speaking gaps. Thirdly, in group intercom scenarios, it also avoids the chaos caused by simultaneous handover of multiple devices. Because all devices independently monitor the silence period, multiple devices may trigger handover almost simultaneously when the group enters the silence period, but the handover execution time is extremely short, and after the handover is completed, the group will uniformly communicate through the cellular network, avoiding the chaotic state of some devices being in the mesh and others in the cellular.
[0103] This embodiment is particularly suitable for scenarios with extremely high requirements for call continuity, such as emergency rescue and dispatch command, fundamentally ensuring uninterrupted calls through silent period scheduling and handover. Furthermore, it can work in conjunction with audio channel control, updating the audio routing table simultaneously during handover in the silent period to ensure that subsequent calls use the cellular channel.
[0104] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0105] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0106] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A communication handover method, characterized in that, A method applicable to a communication system comprising a group of intercom devices and user terminals, wherein the group of intercom devices includes multiple intercom devices, and the intercom devices support MESH intercom mode and cellular network intercom mode; the method includes: In the MESH intercom mode, the MESH signal strength and cellular network signal quality of the intercom device are monitored in real time; and the motion status information of the intercom device is acquired. Based on the motion status information and the pre-stored deployment location of the MESH relay station, the MESH signal change trend in the future time period is predicted to obtain the prediction result. The motion status information includes at least one of movement speed, movement direction, and acceleration change. When it is predicted that the MESH signal will drop below the switching threshold within a preset time window, the start time of the connection establishment process of the cellular network intercom mode is advanced accordingly. If the MESH signal strength and the cellular network signal quality are detected to meet the first set judgment condition, the connection establishment process of the cellular network intercom mode is started in advance based on the prediction result, so as to switch the intercom device from the MESH intercom mode to the cellular network intercom mode and turn off the MESH audio channel of the intercom device. In the cellular intercom mode, the MESH signal strength is continuously monitored. If the MESH signal strength meets the second preset judgment condition, the cellular intercom mode is switched back to the MESH intercom mode, and the MESH audio channel of the intercom device is restored.
2. The method according to claim 1, characterized in that, The first set judgment condition includes: the MESH signal strength is lower than a first predetermined threshold, and the cellular network signal quality is higher than a second predetermined threshold; Switching the intercom device from the MESH intercom mode to the cellular network intercom mode includes: Before the first set judgment condition is met, a registration state is established with the cellular network and a background connection is maintained; when the first set judgment condition is met, the audio bearer of the background connection is activated to complete the switching. Keep the MESH radio frequency link of the intercom device active, disable the microphone signal input and speaker signal output by controlling the input and output switches of the audio codec, mark the routing entry corresponding to the MESH audio channel in the audio routing table of the intercom device as disabled, and redirect the audio data stream to the cellular network audio channel.
3. The method according to claim 1, characterized in that, The second set judgment conditions include: The MESH signal strength is consistently higher than a third predetermined threshold for a predetermined duration.
4. The method according to claim 1, characterized in that, If the intercom device group includes multiple intercom devices, the method further includes: Monitor the MESH signal strength and cellular network signal between any two of the aforementioned intercom devices; If the MESH signal strength and cellular network signal quality between two target intercom devices meet the first set judgment condition, the intercom mode between the target intercom devices is switched to the cellular network intercom mode. Maintain the intercom mode between the target intercom device and other intercom devices as MESH intercom mode, and control the intercom mode between the target intercom device and the other intercom devices as MESH intercom mode.
5. The method according to claim 1, characterized in that, The method further includes: Real-time monitoring of whether the user terminal has call events; If the aforementioned call event occurs, mute the audio input and output channels corresponding to the MESH intercom module; If the call event is detected to have ended, the audio input and output channels of the intercom device are restored to the intercom mode before the switch.
6. The method according to claim 5, characterized in that, The real-time monitoring of whether the user terminal has a call event includes: If a telephone audio data stream is detected on the user terminal, the audio data stream status defined by the Bluetooth hands-free specification is identified. Determine whether the telephone call is connected based on the status of the audio data stream. If the call is in the connected state, a call event is determined to exist; otherwise, a call event is determined not to exist.
7. The method according to claim 1, characterized in that, The method further includes: During the switching from the MESH intercom mode to the cellular intercom mode, the MESH audio channel and the cellular audio channel of the intercom device are kept open at the same time. The two audio signals are cross-faded in and out and then the MESH audio channel is turned off. The fusion duration of the cross-fade-in / fade-out is dynamically adjusted according to the current audio delay of the intercom device, and during the fusion, the gain of the MESH audio channel gradually decreases while the gain of the cellular network audio channel gradually increases.
8. The method according to claim 1, characterized in that, The method of real-time monitoring of the MESH signal strength and cellular network signal quality of the intercom device in the MESH intercom mode includes: The MESH signal strength is monitored for a first preset monitoring period, and the cellular network signal quality is monitored for a second preset monitoring period, wherein the first preset monitoring period is shorter than the second preset monitoring period; The first preset monitoring cycle and the second preset monitoring cycle are dynamically set according to the environmental change rate of the intercom device. The higher the environmental change rate, the shorter the first preset monitoring cycle and the second preset monitoring cycle are.
9. A multi-mode real-time voice communication system, characterized in that, The method includes a group of intercom devices and user terminals. The intercom device group contains multiple intercom devices, and each intercom device is connected to a user terminal in a one-to-one communication manner. Each intercom device includes a MESH intercom module, a cellular communication module, and a main control unit. The MESH intercom module and the cellular communication module are used to implement MESH intercom mode and cellular network intercom mode, respectively. The main control unit is used to implement the steps of the method as described in any one of claims 1-8.
Citation Information
Patent Citations
System switching method and system switching device of talkback equipment
CN117641495A
System and method for cooperative communication and predictive seamless switching of star network and cellular network
CN121151982A