Conference audio merging method supporting dual-mode coexistence and simultaneous ring

By implementing connection status monitoring, dual-mode incoming call detection, synchronous ringing control, hardware mixing processing, and audio routing optimization, the interoperability issues between VoIP devices and eLink software have been resolved, enabling seamless collaboration between VoIP and eLink devices, supporting multi-party conferencing, optimizing call quality, simplifying operation processes, and improving user experience.

CN121531072BActive Publication Date: 2026-05-15CHINA SOUTHERN POWER GRID INTERNET SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610056783.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-05-15
Estimated Expiration
2046-01-16

AI Technical Summary

Technical Problem

VoIP calling devices and eLink software cannot interact well, requiring users to frequently switch devices, affecting work efficiency and call quality. They lack audio mixing and merging functions, cannot support three-way conferencing, have poor device-software interoperability, and cannot synchronize incoming call ringing or seamlessly switch calls.

Method used

Through connection status monitoring, dual-mode incoming signal detection, synchronous ringing control, hardware mixing processing, and audio routing optimization, seamless collaboration between VoIP and eLink devices is achieved. It supports dual-mode coexistence and synchronous ringing, and uses time-domain mixing algorithms and dynamic weight adjustment, combined with echo cancellation and noise suppression technologies to ensure audio merging quality.

Benefits of technology

It improves device compatibility and collaboration, enables seamless collaborative ringing and orderly call management, supports multi-party conferences, optimizes call quality, simplifies operation processes, and improves work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531072B_ABST
    Figure CN121531072B_ABST
Patent Text Reader

Abstract

The application provides a conference audio merging method supporting dual-mode coexistence and synchronous ringing, and belongs to the technical field of audio processing. The method realizes real-time maintenance of the connection state of VoIP and eLink through a connection state monitoring unit, dual-mode incoming call signaling detection based on stable connection, synchronous ringing unified control, three-party conference through hardware mixing processing, and audio routing and optimization to improve sound quality. The method solves the problems of poor interoperability between existing VoIP devices and eLink software, inability to support three-party conference, complex operation and the like, realizes efficient conference audio merging under dual-mode coexistence, improves call efficiency and user experience, and is suitable for communication cooperation scenes in remote office.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, specifically to a method for merging conference audio that supports dual-mode coexistence and synchronized ringing. Background Technology

[0002] In modern enterprise and individual work, VoIP (Voice over IP) technology has become the mainstream communication method. However, users often face compatibility and integration issues between VoIP devices and office software systems. Especially when using cloud conferencing software such as eLink, VoIP phones and eLink software cannot interact well, forcing users to frequently switch devices, impacting work efficiency and call quality. Specific drawbacks include:

[0003] 1. Poor interoperability between the device and software, unable to synchronize incoming call ringing and seamless call switching;

[0004] 2. Lacks audio mixing and merging functions, and cannot support three-way conferencing;

[0005] 3. Switching between different devices and software is cumbersome and makes it difficult to manage calls efficiently. Summary of the Invention

[0006] The present invention aims to solve the problems mentioned in the background art by providing a conference audio merging method that supports dual-mode coexistence and synchronous ringing.

[0007] The specific technical solution is as follows:

[0008] A method for merging conference audio that supports dual-mode coexistence and synchronized ringing includes the following steps:

[0009] S1. Real-time monitoring of connection status: The connection status monitoring unit maintains the status of VoIP connection and eLink connection respectively. For VoIP connection, network probe packets are periodically sent and the connection validity is judged based on the response results. For eLink connection, the connection validity is judged by detecting specific electrical signal characteristics in the link. The status changes of the two types of connections are fed back to the central control unit in real time.

[0010] S2. Dual-mode incoming call detection: Based on the valid connection confirmed in step S1, VoIP incoming call and eLink incoming call are detected separately. VoIP incoming call detection is completed by a SIP protocol parser, which first analyzes the network data packet header to identify SIP protocol feature fields, and then deeply parses the INVITE message to extract call information. eLink incoming call detection is completed by a dedicated electrical signal detection circuit, which identifies feature signal patterns such as specific frequency pulse sequences by sensing changes in voltage and current in the link.

[0011] S3. Synchronous Ringing Unified Control: Taking the incoming call signal detected in step S2 as input, the ringing is implemented through a ringing drive circuit. The ringing drive circuit amplifies the trigger signal, generates a ringtone, and amplifies the power before driving the ringing. At the same time, the central control unit executes conflict avoidance logic. If two incoming call signals are received simultaneously, they are processed according to a preset priority. When the eLink is ringing, if a new call is a VoIP call, it is placed in the waiting queue and the ringing is triggered. At this time, the eLink session is interrupted. When the VoIP call ends, the eLink session is automatically restored. The preset priority is: the eLink incoming call signal has a higher priority than the VoIP incoming call signal. The waiting queue only supports storing unprocessed VoIP calls, storing a maximum of 3 calls, sorted by the order of incoming calls, with the first caller displayed first. For example, when A, B, and C call in sequence, the phone displays A by default, while B and C are hidden and need to be viewed using the navigation key down.

[0012] S4. Hardware mixing processing: In the call state confirmed in step S3, audio mixing is achieved through a dedicated audio processing chip. The eLink audio stream enters through the USB interface, is converted in format by the decoding circuit, and is transmitted to the mixer through the internal bus. The VoIP audio stream is conditioned and converted in format by the dedicated interface circuit and then transmitted to the mixer. The mixer uses a time-domain mixing algorithm to sample, convert, and synchronize the two audio streams, and then superimposes them according to preset weights.

[0013] S5. Audio Routing and Optimization: The audio mixed in step S4 is encoded, amplified, and then output to the speaker; at the same time, the audio picked up by the microphone is pre-amplified and enters the audio processing chip, which is then routed to the VoIP channel and the eLink channel respectively through the routing control module; the AEC module in the audio processing chip compares the speaker playback signal with the microphone pickup signal to eliminate echo, and the ANS module suppresses noise to improve sound quality.

[0014] The temporal mixing algorithm in step S4 includes dynamic weight adjustment, and the mixed audio signal y(t) is calculated as follows:

[0015] ;

[0016] in:

[0017] x v (t) and x e (t) represents the sampled values ​​of the VoIP audio signal and the eLink audio signal at time t, respectively;

[0018] E v (t) and E e (t) represents the short-time speech energy of VoIP and eLink, respectively, in dBm;

[0019] av (t) and a e (t) is the dynamic activation coefficient, defined as:

[0020] , ;

[0021] Where E thresh =−40dBm is the activation threshold;

[0022] δ is the damping constant, with a value of 10. −6 Up to 10 −3 This is used to prevent the denominator from being zero.

[0023] As a preferred embodiment of the present invention, in step S1, the periodic sending of network probe packets is 1-5 seconds, and when no response is received for 3 consecutive times, the VoIP connection is determined to be invalid.

[0024] As a preferred embodiment of the present invention, in step S1, the specific electrical signal characteristics of the eLink connection include: stable output of link voltage within the range of 3.3V±0.2V, current change rate ≤5mA / ms, and 1kHz pulse sequence with an interval of 100ms±10ms.

[0025] As a preferred embodiment of the present invention, in step S2, the SIP protocol feature fields include the “SIP / 2.0 / UDP” identifier in the Via header field and the user identification information in the From header field, and the parsing of the INVITE message includes extracting the caller ID, call timestamp and media type parameters.

[0026] As a preferred embodiment of the present invention, in step S3, the preset priority is: eLink incoming calls have a higher priority than VoIP incoming calls; the waiting queue only supports storing unprocessed VoIP calls, storing a maximum of 3 calls, sorted according to the order of incoming calls, with the first caller displayed first. For example, when A, B, and C call in sequence, the phone will display A by default, while B and C will be hidden and need to be viewed by pressing the navigation key down.

[0027] As a preferred embodiment of the present invention, in step S4, the specific processing of the time-domain mixing algorithm includes: sampling the two audio channels to 8kHz or 16kHz, achieving synchronization processing through timestamp alignment, and the superposition weight ratio of VoIP audio and eLink audio can be adjusted within the range of 3:7 to 7:3.

[0028] As a preferred embodiment of the present invention, step S4 further includes a call activation detection step: the call activation detection module monitors the call status of VoIP and eLink in real time, and when the voice energy of a certain call is detected to be ≥-40dBm, an activation signal is sent to the audio processing chip to mix the audio of that call into the current mix stream.

[0029] As a preferred embodiment of the present invention, in step S5, the audio collected by the microphone is preamplified to a gain of 20-40dB, and the audio signals routed to the VoIP channel and the eLink channel are respectively encoded using G.711 or G.729.

[0030] As a preferred embodiment of the present invention, in step S5, the echo cancellation processing of the AEC module includes: adaptively filtering the speaker playback signal to generate an echo reference signal, subtracting the microphone acquisition signal from the echo reference signal to eliminate linear echo, and updating the filtering coefficients every 20ms.

[0031] As a preferred embodiment of the present invention, in step S5, the noise suppression processing of the ANS module includes: identifying noise components in the microphone-collected signal in real time based on a pre-trained noise model, and performing inverse cancellation of the noise through spectral subtraction or Wiener filtering, wherein the noise model is generated by training on at least 100 hours of office environment noise samples.

[0032] The present invention has the following beneficial effects:

[0033] 1. Enhanced compatibility and interoperability: The dual-mode coexistence design solves the interoperability problem between VoIP devices and eLink software, enabling seamless collaboration between the two and avoiding frequent device switching for users.

[0034] 2. Achieve synchronized ringing and orderly call management: Through synchronized ringing control and conflict avoidance logic, ensure synchronized response to dual-mode calls, solve the problem of chaotic simultaneous calls, and improve call processing efficiency.

[0035] 3. Supports multi-party conferencing needs: Hardware mixing processing enables the merging of two audio streams, meeting the needs of collaborative scenarios such as enterprise three-party meetings and expanding the functionality of the device.

[0036] 4. Optimize call quality: Reduce interference and improve voice clarity through echo cancellation and noise suppression technologies, thereby enhancing the user's call experience.

[0037] 5. Simplified operation process: Fully automated processing (connection monitoring, signaling detection, audio mixing optimization, etc.) reduces user operation complexity and improves work efficiency. Attached Figure Description

[0038] Figure 1A flowchart illustrating a conference audio merging method that supports dual-mode coexistence and synchronous ringing, provided in an embodiment of the present invention. Detailed Implementation

[0039] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0040] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this application. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0041] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present application. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0042] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0043] The present invention provides a conference audio merging method that supports dual-mode coexistence and synchronous ringing, such as... Figure 1 As shown, it includes the following steps:

[0044] S1. Real-time monitoring of connection status: The connection status monitoring unit maintains the status of VoIP connection and eLink connection respectively. For VoIP connection, network probe packets are periodically sent and the connection validity is judged based on the response results. For eLink connection, the connection validity is judged by detecting specific electrical signal characteristics in the link. The status changes of the two types of connections are fed back to the central control unit in real time.

[0045] S2. Dual-mode incoming call detection: Based on the valid connection confirmed in step S1, VoIP incoming call and eLink incoming call are detected separately. VoIP incoming call detection is completed by a SIP protocol parser, which first analyzes the network data packet header to identify SIP protocol feature fields, and then deeply parses the INVITE message to extract call information. eLink incoming call detection is completed by a dedicated electrical signal detection circuit, which identifies feature signal patterns such as specific frequency pulse sequences by sensing changes in voltage and current in the link.

[0046] S3. Synchronous Ringing Unified Control: Taking the incoming call signal detected in step S2 as input, the ringing driver circuit realizes hardware ringing. The ringing driver circuit amplifies the trigger signal, generates a ringtone, and amplifies the power before driving the ringing. At the same time, the central control unit executes conflict avoidance logic. If two incoming call signals are received simultaneously, they are processed according to the preset priority. When the eLink is ringing, if a new call is a VoIP call, it is placed in the waiting queue and the ringing is triggered. At this time, the eLink session is interrupted. When the VoIP call ends, the eLink session is automatically restored. The preset priority is: the eLink incoming call signal has higher priority than the VoIP incoming call signal. The waiting queue only supports storing unprocessed VoIP calls, storing a maximum of 3 calls, sorted by the order of incoming calls, with the first caller displayed first. For example, when A, B, and C call in sequence, the phone displays A by default, while B and C are hidden and need to be viewed by pressing the navigation key down.

[0047] S4. Hardware mixing processing: In the call state confirmed in step S3, audio mixing is achieved through a dedicated audio processing chip. The eLink audio stream enters through the USB interface, is converted in format by the decoding circuit, and is transmitted to the mixer through the internal bus. The VoIP audio stream is conditioned and converted in format by the dedicated interface circuit and then transmitted to the mixer. The mixer uses a time-domain mixing algorithm to sample, convert, and synchronize the two audio streams, and then superimposes them according to preset weights.

[0048] S5. Audio Routing and Optimization: The audio mixed in step S4 is encoded, amplified, and then output to the speaker; at the same time, the audio picked up by the microphone is pre-amplified and enters the audio processing chip, which is then routed to the VoIP channel and the eLink channel respectively through the routing control module; the AEC module in the audio processing chip compares the speaker playback signal with the microphone pickup signal to eliminate echo, and the ANS module suppresses noise to improve sound quality.

[0049] The above-described technical solution supports dual-mode coexistence and synchronous ringing in conference audio merging. It ensures stable VoIP and eLink connections through real-time monitoring, providing a reliable foundation for subsequent processing; accurately identifies both modes of incoming calls through dual-mode call detection; achieves synchronized ringing for both modes and resolves conflicts caused by simultaneous calls through unified synchronous ringing control; effectively merges the two audio streams through hardware mixing, supporting multi-party conference needs; and improves call quality by eliminating echo and suppressing noise through audio routing and optimization. Overall, it achieves conference audio merging under dual-mode coexistence, improving device compatibility and user call experience, and meeting the demands for efficient communication.

[0050] The temporal mixing algorithm in step S4 includes dynamic weight adjustment, and the mixed audio signal y(t) is calculated as follows:

[0051] ;

[0052] in:

[0053] x v (t) and x e (t) represents the sampled values ​​of the VoIP audio signal and the eLink audio signal at time t, respectively. v (t) represents the decoded signal from the eLink via the USB interface; x e (t) represents the signal from VoIP after conditioning and conversion by a dedicated interface circuit, with a sampling rate of 8kHz or 16kHz.

[0054] E v (t) and E e (t) represents the short-time speech energy of VoIP and eLink, respectively, in dBm;

[0055] a v (t) and a e (t) is the dynamic activation coefficient, defined as:

[0056] , ;

[0057] Where E thresh =−40dBm is the activation threshold;

[0058] δ is the damping constant, with a value of 10. −6 Up to 10 −3 This is used to prevent the denominator from being zero.

[0059] Further explanation of the equation parameters:

[0060] x v (t) and x e(t) represents the sampled values ​​of the VoIP audio signal and the eLink audio signal at time t, respectively. These values ​​are derived from the decoded and format-converted signals in step S4, with a uniform sampling rate of 8kHz or 16kHz.

[0061] E v (t) and E e (t) represents the short-time voice energy of VoIP and eLink, respectively, in dBm. This energy is calculated in real-time by the audio processing chip using the squared average of the signal amplitude within a 20ms window, then converted to dBm (e.g., ...). ).

[0062] a v (t) and a e (t) is the dynamic activation coefficient, a binary value (0 or 1), generated by the call activation detection module. When the voice energy is ≥-40dBm, the coefficient is 1, indicating that the audio channel is activated and mixed in; otherwise, it is 0, indicating that it is not mixed in.

[0063] E thresh Fixed activation threshold, -40dBm, based on the typical range of human speech energy (speech energy in a quiet environment is approximately -30dBm to -10dBm).

[0064] δ: Damping constant, a small positive number (e.g., 10). −6 To ensure the denominator is not zero and prevent numerical overflow, the values ​​are optimized through experiments to balance numerical stability and sound quality.

[0065] Output y(t): The mixed audio signal is output to the speaker after time-domain mixing (step S5).

[0066] Example: Applying this equation in a multi-party meeting scenario within an enterprise (as shown in Example 2 below):

[0067] Scenario Description: A user simultaneously joins a VoIP conference and an eLink conference. When the VoIP participant speaks (E... v (t) = −35dBm ≥ E thresh , so a v (t)=1), eLink mute (E e (t) = −50dBm <E thresh , so a e (t)=0).

[0068] Equation calculation:

[0069] ;

[0070] The output y(t) approximates the VoIP audio, ensuring that the speaker's voice is clear.

[0071] Dynamic adjustment: If both parties speak simultaneously (E v (t) = −30 dBm, E e (t) = −25dBm, all activated), then:

[0072] The weights are automatically biased towards the higher-energy eLink audio (weight ratio ~54.5%:45.5%), which conforms to the principle of "increasing the weight of the activator" (Example 2 in the instruction manual).

[0073] Technical effect

[0074] Improve speech clarity: through dynamic activation coefficient (a v (t),a e (t) and energy ratio weights ensure that the active party's voice dominates the mixing, reducing interference from inactive audio, and significantly improving speech intelligibility in multi-party conferences.

[0075] Optimize resource utilization: Mix in only when audio is active (energy ≥ threshold) to avoid invalid audio (such as silence or background noise) occupying processing resources and effectively reduce CPU load.

[0076] Enhanced sound quality adaptability: By combining the AEC and ANS modules, this equation reduces the interference of echo and noise on the weight calculation, significantly improving the overall sound quality MOS (Mean Opinion Score) score.

[0077] Supports efficient meetings: Dynamic weighting ensures "clear distinction between primary and secondary voices" (Example 2 in the manual), solving the problem of sound overlap in three-way meetings.

[0078] Working principle and process

[0079] This equation is implemented in hardware mixing (step S4), and the workflow is as follows:

[0080] 1. Input Acquisition: The audio processing chip acquires x input from the VoIP and eLink interfaces. v (t) and x e (t) (after decoding and synchronization).

[0081] 2. Energy Calculation: Real-time calculation of short-time speech energy E v (t) and E e (t) (Updated every 20ms).

[0082] 3. Activation detection: Compare energy with threshold E thresh =−40dBm:

[0083] If E v (t)≥E thresh Let a v(t)=1 (VoIP activated).

[0084] If E e (t)≥E thresh Let a e (t)=1 (eLink activated).

[0085] 4. Dynamic mixing calculation: Substitute into the equation to calculate the mixed signal y(t):

[0086] Molecules: The energy weighted sum of the activated audio (the higher the energy, the greater the contribution).

[0087] Denominator: Normalization factor, to ensure stable output (add δ to prevent division by zero).

[0088] 5. Output and Optimization: y(t) is output to step S5 for encoding, power amplification, and audio quality optimization (such as AEC echo cancellation). Meanwhile, inactive audio (a...) v (t)=0 or a e (t)=0) is excluded, reducing the noise in the mixing stream.

[0089] This equation introduces dynamic activation coefficients (based on a patented specific threshold of -40dBm) and energy ratio normalization to form a unique mixing mechanism for dual-mode audio. For example, the general formula y(t)=w1x1(t)+w2x2(t) has no activation control; this equation adds activation coefficients and energy terms to the denominator to ensure that mixing only occurs when valid speech is present.

[0090] Addressing the issue of existing technologies' inability to support three-way conferencing:

[0091] "Selective mixing" is achieved by activating coefficients to avoid invalid audio interference (improving compatibility).

[0092] The energy ratio weighting automatically adapts to the speaking intensity, achieving "clear distinction between primary and secondary speech" (optimizing user experience).

[0093] Specifically, in this embodiment, in step S1, the periodic sending of network probe packets is 1-5 seconds, and if no response is received for three consecutive times, the VoIP connection is determined to be invalid. By adopting the above technical solution, periodically sending network probe packets and determining VoIP connection invalidation based on consecutive non-response, changes in VoIP connection status can be monitored in a timely and accurate manner, avoiding disruption to calls due to undetected connection anomalies, and ensuring the real-time performance and reliability of VoIP connection status monitoring.

[0094] Specifically, in this embodiment, in step S1, the specific electrical signal characteristics of the eLink connection include: a stable output link voltage within the range of 3.3V±0.2V, a current change rate ≤5mA / ms, and a 1kHz pulse sequence with intervals of 100ms±10ms. By employing the above technical solution, the validity of the eLink connection can be accurately determined by detecting specific voltage, current characteristics, and pulse sequences in the eLink link, ensuring the accuracy of eLink connection status monitoring and providing a stable connection status basis for the dual-mode collaborative operation of VoIP and eLink.

[0095] Specifically, in this embodiment, in step S2, the SIP protocol feature fields include the "SIP / 2.0 / UDP" identifier in the Via header field and the user identification information in the From header field. The parsing of the INVITE message includes extracting the caller ID, call timestamp, and media type parameters. By employing the above technical solution, through analyzing the SIP protocol feature fields and deeply parsing the INVITE message to extract call information, it is possible to accurately identify VoIP incoming call signals and related call details, providing precise information support for subsequent ringing control and call processing of VoIP calls, and ensuring the accuracy of VoIP incoming call signal detection.

[0096] Specifically, in this embodiment, in step S3, the preset priority is: eLink incoming calls have a higher priority than VoIP incoming calls; the waiting queue only supports storing unprocessed VoIP calls, storing a maximum of 3 calls, sorted by incoming call order, with the first caller displayed first. For example, when A, B, and C call in sequence, the phone will display A by default, while B and C are hidden and need to be viewed using the navigation key; when eLink is ringing, a new VoIP call is placed in the waiting queue and a ringing is triggered, at which point the eLink session is interrupted; when the VoIP call ends, the eLink session automatically resumes. By using the above technical solution, processing two simultaneous calls with preset priorities and managing unprocessed VoIP calls with a waiting queue, the ringing conflict problem when dual-mode calls arrive simultaneously can be effectively solved, avoiding ringing chaos and allowing users to process calls in an orderly manner, improving the convenience and orderliness of call processing.

[0097] Specifically, in this embodiment, step S4 of the time-domain mixing algorithm includes: uniformly sampling the two audio streams to 8kHz or 16kHz, achieving synchronization through timestamp alignment, and adjusting the superposition weight ratio of VoIP audio and eLink audio within the range of 3:7 to 7:3. By employing the above technical solution, the time-domain mixing algorithm performs unified sampling, synchronization processing, and weight adjustment on the two audio streams, achieving efficient merging of VoIP and eLink audio streams, ensuring the synchronization and clarity of the mixed audio, and meeting the audio merging requirements in multi-party conferences.

[0098] Specifically, in this embodiment, step S4 further includes a call activation detection step: the call activation detection module monitors the call status of VoIP and eLink in real time. When the voice energy of a certain call is detected to be ≥-40dBm, an activation signal is sent to the audio processing chip to mix that audio into the current mix stream. By adopting the above technical solution, the call activation detection module determines the call status based on voice energy and controls audio mixing, which can dynamically identify valid call audio, avoid invalid audio from being mixed into the mix stream, improve the targeting and efficiency of mixing processing, and optimize mixing quality.

[0099] Specifically, in this embodiment, in step S5, the audio collected by the microphone is pre-amplified to a gain of 20-40dB, and the audio signals routed to the VoIP channel and eLink channel are respectively encoded using G.711 or G.729. By employing the above technical solution, pre-amplifying and adapting the microphone-collected audio ensures that the collected audio signal is clear and adaptable to the transmission requirements of the VoIP and eLink channels, guaranteeing that local audio can be accurately and clearly transmitted to both channels during two-way communication, thus improving the quality of two-way communication.

[0100] Specifically, in this embodiment, step S5, the echo cancellation processing of the AEC module includes: adaptively filtering the speaker playback signal to generate an echo reference signal, subtracting the microphone acquisition signal from the echo reference signal to eliminate linear echo, and updating the filter coefficients every 20ms. Using the above technical solution, through the adaptive filtering and periodic updating of the filter coefficients by the AEC module, echo interference in calls can be effectively eliminated, avoiding the impact of echo on calls and improving the clarity of the voice signal.

[0101] Specifically, in this embodiment, step S5, the noise suppression processing of the ANS module includes: real-time identification of noise components in the microphone-collected signal based on a pre-trained noise model, and inverse cancellation of the noise through spectral subtraction or Wiener filtering. The noise model is generated through training on at least 100 hours of office environment noise samples. By employing the above technical solution, noise identification and filtering of the microphone signal based on the trained noise model can effectively suppress environmental noise, extract clean speech signals, and improve call quality and intelligibility.

[0102] The present invention also provides the following three specific embodiments:

[0103] Example 1: Dual-mode call collaboration in a remote work scenario:

[0104] In remote work scenarios, employees connect to their computers via SIP+USB dual-mode IP phones, simultaneously running a VoIP calling system and eLink collaboration software.

[0105] 1. The connection status monitoring unit periodically checks the VoIP network connection (sends probe packets) and the eLinkUSB connection (detects voltage and current characteristics) to ensure that both are in a stable state and feeds back to the central control unit;

[0106] 2. When a VoIP call comes in, the SIP protocol parser recognizes the INVITE message and extracts the call information; if a VoIP call is received while the eLink is ringing, the central control unit puts the VoIP call into the waiting queue and triggers ringing, interrupting the current eLink ringing session;

[0107] 3. After an employee answers an eLink conference call, the system stops the VoIP ringing. The dedicated audio processing chip merges the eLink conference audio (decoded via USB interface) with the subsequent VoIP call audio (conditioned and converted via dedicated interface circuit) using a time-domain mixing algorithm, and eliminates echo and noise through AEC and ANS modules. When the VoIP call ends, the eLink session automatically resumes.

[0108] 4. Employees speak through microphones, and the audio is pre-amplified and routed to eLink conferencing and VoIP call channels respectively, enabling real-time communication among the three parties.

[0109] The dual-mode call collaboration technology in this remote work scenario achieves the following effects: It ensures stable VoIP and eLink connections through connection status monitoring, preventing call interruptions due to connection drops; it resolves conflicts caused by simultaneous calls from both modes during remote work using priority processing and waiting queue mechanisms, reducing the need for users to switch devices; and it combines hardware mixing with audio quality optimization modules to achieve clear merging of cross-platform call audio, improving the fluency and efficiency of multi-party communication in remote collaboration and meeting employees' needs for multi-task call management when working from home or in different locations.

[0110] Example 2: Internal multi-party meeting scenario within an enterprise

[0111] Employees need to access two different meetings simultaneously (VoIP meeting and eLink team meeting) via dual-mode IP phones.

[0112] 1. After the connection status monitoring unit confirms that the VoIP and eLink connections are valid, the SIP protocol parser continuously monitors the VoIP conference signaling, and the eLink dedicated circuit detects the conference activation signal;

[0113] 2. When two conference calls are initiated one after the other, the system triggers a synchronized ringing in the order of receipt. After the employee answers the first conference call, the second conference call enters the waiting queue and is automatically mixed into the current call when the employee switches to the second conference call.

[0114] 3. The mixer processes the two conference audio streams synchronously and dynamically adjusts the weights based on the speaker's voice energy (increasing the weight of the active speaker) to ensure that the primary and secondary voices are clearly distinguishable.

[0115] 4. Local audio is encoded and simultaneously transmitted to two conference channels, enabling real-time collaborative discussions among multiple parties. The noise suppression module filters out background noise in the office, improving call quality.

[0116] The technical effects of this enterprise's internal multi-party conferencing scenario include: real-time connection status monitoring provides a stable foundation for parallel dual-conference operation, ensuring uninterrupted meeting processes; dynamic weighted mixing processing makes the audio of different speakers in multi-party meetings clearly distinguishable, avoiding sound overlap and confusion; noise suppression effectively filters out interference from the office environment, improving the intelligibility of meeting audio; and a waiting queue and seamless switching design allow users to flexibly manage multiple meeting accesses, simplifying the operation process of multi-party collaboration within the enterprise and improving meeting efficiency.

[0117] Example 3: Cross-platform emergency call processing

[0118] When using eLink to handle urgent matters, users need to answer important VoIP calls at the same time.

[0119] 1. The connection status monitoring unit provides real-time feedback that both VoIP and eLink are connected. When eLink is in a call, incoming VoIP calls are recognized by the SIP parser.

[0120] 2. The central control unit triggers the waiting queue mechanism, puts the VoIP call into the waiting queue and triggers ringing, interrupts the current eLink ringing session, and only indicates the eLink session status through indicator lights;

[0121] 3. After the user switches to VoIP call via the device button, the system automatically mutes the eLink call audio temporarily and keeps it in the mixing queue. The eLink audio mix is ​​restored after the VoIP call ends, and the eLink session is automatically restored after the VoIP call ends.

[0122] 4. Throughout the process, the echo cancellation module avoids audio interference during switching, ensuring a seamless connection between emergency calls and transaction processing.

[0123] The cross-platform emergency call processing technology achieves the following effects: connection status monitoring ensures the reliability of dual-mode connections in emergency situations, ensuring that no important calls are missed; the waiting queue and mute retention mechanism avoids mutual interference between emergency calls and current tasks, enabling seamless switching; the echo cancellation module prevents audio noise generated during switching, ensuring the clarity of emergency calls; the overall solution allows users to efficiently handle cross-platform calls when dealing with emergency tasks, improving communication response speed and processing convenience in emergency scenarios.

[0124] In summary, the working principle of the conference audio merging method supporting dual-mode coexistence and synchronous ringing provided in this embodiment is as follows:

[0125] 1. Connection Status Monitoring: The connection status monitoring unit maintains VoIP and eLink connections separately. For VoIP connections, network probe packets are periodically sent and the validity is determined based on the response. For eLink connections, the validity is determined by detecting specific electrical signal characteristics (such as voltage and current changes) in the link. Status changes are fed back to the central control unit in real time, providing a stable foundation for subsequent processing.

[0126] 2. Dual-mode incoming call detection: Based on a valid connection, VoIP incoming call detection uses a SIP protocol parser to identify protocol feature fields and parses INVITE messages to extract call information; eLink incoming call detection uses a dedicated circuit to detect feature signals such as specific frequency pulse sequences in the link, achieving accurate identification of the two types of incoming calls.

[0127] 3. Synchronous Ringing Control: After receiving an incoming call, the ringing drive circuit amplifies the signal, generates a ringtone, and amplifies the power to drive the hardware to ring. The central control unit processes simultaneous incoming calls through conflict avoidance logic (sorting them by priority, with unanswered calls entering the waiting queue) to ensure orderly ringing. When the eLink is ringing, if a new VoIP call is added to the waiting queue and ringing is triggered, the eLink session is interrupted. When the VoIP call ends, the eLink session automatically resumes. The waiting queue only stores unprocessed VoIP calls, up to 3 calls, displayed in order of arrival.

[0128] 4. Hardware mixing: A dedicated audio processing chip processes VoIP and eLink audio. VoIP audio is decoded via USB and then sent to the mixer, while eLink audio is conditioned and converted via the interface before being sent to the mixer. Synchronous overlay is achieved through a time-domain mixing algorithm, supporting three-way conferencing.

[0129] 5. Audio Routing and Optimization: After mixing, the audio is encoded, amplified, and output to the speaker; the audio captured by the microphone is amplified and routed to the VoIP and eLink channels respectively, while the AEC module eliminates echo and the ANS module suppresses noise to improve sound quality.

[0130] How to use

[0131] 1. Device Connection: Connect the SIP+USB dual-mode IP phone to the computer via USB to ensure simultaneous support for VoIP calling and eLink collaboration functions.

[0132] 2. Connection monitoring: The system automatically starts connection status monitoring and maintains VoIP and eLink connections in real time, without requiring manual intervention from the user.

[0133] 3. Incoming Call Handling: When a VoIP or eLink call comes in, the system synchronously detects the signaling and drives the device to ring. If a new VoIP call is received while the eLink is ringing, it is placed in the waiting queue and the ringing is triggered, at which point the eLink session is interrupted. When the VoIP call ends, the eLink session is automatically resumed. The waiting queue only stores unprocessed VoIP calls, up to 3 calls, sorted in the order of arrival (the first caller is displayed first). For example, if A, B, and C arrive in sequence, the phone will display A by default, while B and C will be hidden and can only be viewed by pressing the down navigation key.

[0134] 4. Conference audio merging: During a call, the system automatically mixes VoIP and eLink audio, supporting three-way conferencing; users can speak normally through the device, and the audio captured by the microphone is optimized and transmitted synchronously to both channels.

[0135] 5. Sound quality optimization: The system automatically eliminates echoes and suppresses noise, allowing users to obtain clear call quality without manual operation.

[0136] The above are merely preferred embodiments of the present invention and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should recognize that any equivalent substitutions and obvious changes made based on the description and illustrations of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for merging conference audio that supports dual-mode coexistence and synchronous ringing, characterized in that, Includes the following steps: S1. Real-time monitoring of connection status: The connection status monitoring unit maintains the status of VoIP connection and eLink connection respectively. For VoIP connection, network probe packets are periodically sent and the connection validity is judged based on the response results. For eLink connection, the connection validity is judged by detecting specific electrical signal characteristics in the link. The status changes of the two types of connections are fed back to the central control unit in real time. S2. Dual-mode incoming call detection: Based on the valid connection confirmed in step S1, VoIP incoming call and eLink incoming call are detected separately. VoIP incoming call detection is completed by a SIP protocol parser, which first analyzes the network data packet header to identify SIP protocol feature fields, and then deeply parses the INVITE message to extract call information. eLink incoming call detection is completed by a dedicated electrical signal detection circuit, which identifies specific frequency pulse sequence feature signal patterns by sensing voltage and current changes in the link. S3. Synchronous Ringing Unified Control: Taking the incoming call signal detected in step S2 as input, the ringing is implemented through a ringing drive circuit. The ringing drive circuit amplifies the trigger signal, generates a ringtone, and amplifies the power before driving the ringing. At the same time, the central control unit executes conflict avoidance logic. If two incoming call signals are received simultaneously, they are processed according to a preset priority. When the eLink is ringing, if a new call is a VoIP call, it is placed in the waiting queue and the ringing is triggered. At this time, the eLink session is interrupted. When the VoIP call ends, the eLink session is automatically restored. S4. Hardware mixing processing: In the call state confirmed in step S3, audio mixing is achieved through a dedicated audio processing chip. The eLink audio stream enters through the USB interface, is converted in format by the decoding circuit, and is transmitted to the mixer through the internal bus. The VoIP audio stream is conditioned and converted in format by the dedicated interface circuit and then transmitted to the mixer. The mixer uses a time-domain mixing algorithm to sample, convert, and synchronize the two audio streams, and then superimposes them according to preset weights. S5. Audio Routing and Optimization: The audio mixed in step S4 is encoded, amplified, and then output to the speaker; at the same time, the audio picked up by the microphone is pre-amplified and enters the audio processing chip, which is then routed to the VoIP channel and the eLink channel respectively through the routing control module; the AEC module in the audio processing chip compares the speaker playback signal with the microphone pickup signal to eliminate echo, and the ANS module suppresses noise to improve sound quality.

2. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, The temporal mixing algorithm in step S4 includes dynamic weight adjustment, and the mixed audio signal y(t) is calculated as follows: ; in: x v (t) and x e (t) represents the sampled values ​​of the VoIP audio signal and the eLink audio signal at time t, respectively; E v (t) and E e (t) represents the short-time speech energy of VoIP and eLink, respectively, in dBm; a v (t) and a e (t) represents the dynamic activation coefficient; δ is the damping constant, with a value of 10. −6 Up to 10 −3 .

3. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S1, the periodic sending of network probe packets is 1-5 seconds, and the VoIP connection is determined to be invalid when no response is received for 3 consecutive times. In step S1, the specific electrical signal characteristics of the eLink connection include: stable output of link voltage within the range of 3.3V±0.2V, current change rate ≤5mA / ms, and 1kHz pulse sequence with an interval of 100ms±10ms.

4. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S2, the SIP protocol feature fields include the "SIP / 2.0 / UDP" identifier in the Via header field and the user identification information in the From header field. The parsing of the INVITE message includes extracting the caller ID, call timestamp, and media type parameters.

5. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S3, the preset priority is: eLink incoming calls have a higher priority than VoIP incoming calls; the waiting queue only supports storing unprocessed VoIP calls, storing a maximum of 3 calls, sorted by the order of incoming calls, with the first caller displayed first.

6. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S4, the specific processing of the time-domain mixing algorithm includes: sampling the two audio channels to 8kHz or 16kHz, achieving synchronization processing through timestamp alignment, and adjusting the superposition weight ratio of VoIP audio and eLink audio within the range of 3:7 to 7:

3.

7. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, Step S4 also includes a call activation detection step: the call activation detection module monitors the call status of VoIP and eLink in real time. When the voice energy of a certain call is detected to be ≥-40dBm, an activation signal is sent to the audio processing chip to mix the audio of that call into the current mix stream.

8. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S5, the audio collected by the microphone is preamplified to a gain of 20-40dB, and the audio signals routed to the VoIP channel and eLink channel are respectively encoded using G.711 or G.

729.

9. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S5, the echo cancellation processing of the AEC module includes: adaptively filtering the speaker playback signal to generate an echo reference signal, subtracting the microphone acquisition signal from the echo reference signal to eliminate linear echo, and updating the filter coefficients every 20ms.

10. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S5, the noise suppression processing of the ANS module includes: identifying noise components in the microphone-collected signal in real time based on a pre-trained noise model, and performing inverse noise cancellation through spectral subtraction or Wiener filtering, wherein the noise model is generated by training on at least 100 hours of office environment noise samples.