Conference audio merging method supporting dual-mode coexistence and synchronous ringing

By monitoring the VoIP and eLink connection status in real time, identifying and processing incoming calls, and merging hardware ringing and audio, the compatibility issues between VoIP devices and office software systems are resolved, improving user experience and call efficiency.

CN121531072AActive Publication Date: 2026-02-13CHINA SOUTHERN POWER GRID INTERNET SERVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610056783.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-02-13
Estimated Expiration
2046-01-16

AI Technical Summary

Technical Problem

The compatibility and integration issues between VoIP calling devices and office software systems include poor interoperability between devices and software, inability to synchronize incoming call ringing and seamless call switching, lack of audio mixing and merging functions, cumbersome switching operations between different devices and software, and difficulty in efficiently managing calls.

Method used

Seamless collaboration between VoIP and eLink devices is achieved through connection status monitoring, dual-mode incoming call detection, synchronized ringing control, hardware audio mixing, and audio routing optimization. Specific measures include real-time monitoring of connection status, identification of incoming calls, unified ringing control, mixing of audio signals, and elimination of echo and noise.

Benefits of technology

It improves device compatibility and collaboration, enables seamless call response and audio merging, supports multi-party conferencing, optimizes call quality and operation processes, and simplifies user operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531072A_ABST
    Figure CN121531072A_ABST
Patent Text Reader

Abstract

The invention provides a conference audio merging method supporting dual-mode coexistence and synchronous ringing, which belongs to the technical field of audio processing, and is characterized in that the connection state of VoIP (Voice over Internet Protocol) and eLink is maintained in real time through a connection state monitoring unit, dual-mode incoming call signaling detection is carried out based on stable connection, and synchronous ringing unified control is realized; the three-party conference is realized through hardware sound mixing processing, and the sound quality is improved by matching with audio routing and optimization. The problems that an existing VoIP device and eLink software are poor in interoperability, cannot support a three-party conference, are complex in operation and the like are solved, efficient conference audio merging under dual-mode coexistence is achieved, the call efficiency and the user experience are improved, and the method is suitable for a communication cooperation scene in remote office.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio processing, in particular to a conference audio merging method supporting dual-mode coexistence and synchronous ringing. BACKGROUND

[0002] In modern enterprise and personal work, VoIP (Voice over IP) technology has become the mainstream communication method, but users often face compatibility and integration problems of VoIP call devices and office software systems. In particular, when using eLink and other cloud conference software, VoIP phones and eLink software cannot achieve good interaction, resulting in users needing to frequently switch devices, affecting work efficiency and call quality. Specific defects include:

[0003] 1. Poor device and software interoperability, unable to synchronize incoming call ringing and seamless call switching;

[0004] 2. Lack of audio mixing and merging functions, unable to support three-party conferences;

[0005] 3. Complicated switching operations between different devices and software, making it difficult to efficiently manage calls. SUMMARY

[0006] The present application provides a conference audio merging method supporting dual-mode coexistence and synchronous ringing to solve the problems raised in the background art.

[0007] The specific technical solution is as follows:

[0008] A conference audio merging method supporting dual-mode coexistence and synchronous ringing, comprising the following steps:

[0009] S1, Real-time connection state monitoring: The connection state monitoring unit maintains the states of VoIP connection and eLink connection respectively, wherein network probe packets are periodically sent for VoIP connection and connection validity is judged according to the response result, and for eLink connection, connection validity is judged by detecting specific electrical signal characteristics in the link, and the state changes of the two types of connections are fed back to the central control unit in real time;

[0010] S2, Dual-mode incoming call signaling detection: Based on the valid connections confirmed in step S1, VoIP incoming call signaling and eLink incoming call signaling are detected respectively, wherein VoIP incoming call signaling is completed by a SIP protocol parser, which first analyzes the network packet header to identify the SIP protocol characteristic field, and then deeply analyzes the INVITE message to extract call information; eLink incoming call signaling is completed by a special electrical signal detection circuit, which identifies characteristic signal patterns such as specific frequency pulse sequences by sensing voltage and current changes in the link;

[0011] S3, synchronous ring unified control: taking the incoming call signaling detected in step S2 as input, a ring drive circuit is used to realize hardware ringing, the ring drive circuit drives the ring through amplification, ring tone generation and power amplification; at the same time, the central control unit executes conflict avoidance logic, if two incoming call signals are received at the same time, the preset priority is processed, when eLink is ringing, the new incoming call is VoIP, and the waiting queue is put into and the ring is triggered, at this time, the eLink session is interrupted; when the VoIP call ends, the eLink session is automatically restored; the preset priority is that the eLink incoming call signaling priority is higher than the VoIP incoming call signaling; the waiting queue only supports storing VoIP unprocessed incoming calls, at most 3, sorted by incoming order, the first incoming call is displayed first, for example, A, B and C are incoming in order, the phone defaults to display A, B and C are hidden, and need to be viewed through the navigation key down key;

[0012] S4, hardware mixing processing: in the call state confirmed in step S3, audio mixing is realized through a special audio processing chip, wherein the eLink audio stream enters through the USB interface, is converted in format by the decoding circuit and is transmitted to the mixer through the internal bus; the VoIP audio stream is transmitted to the mixer after being conditioned and converted in format by the special interface circuit; the mixer uses a time domain mixing algorithm to sample, convert, synchronize and superimpose the two audio streams according to the preset weight;

[0013] S5, audio routing and optimization: the audio after mixing in step S4 is encoded and power amplified and then output to the loudspeaker; at the same time, the audio collected by the microphone is pre-amplified and then enters the audio processing chip, and is routed to the VoIP channel and the eLink channel through the routing control module; and the AEC module in the audio processing chip compares the loudspeaker playback signal and the microphone collection signal to eliminate echo, and the ANS module suppresses noise to improve sound quality.

[0014] The time domain mixing algorithm in step S4 includes dynamic weight adjustment, and the mixed audio signal y(t) is calculated as:

[0015] ;

[0016] Wherein:

[0017] x v (t) and x e (t) are the sampling values of the VoIP audio signal and the eLink audio signal at time t, respectively;

[0018] E v (t) and E e (t) are the short-time speech energies of VoIP and eLink, respectively, in dBm;

[0019] av (t) and a e (t) is the dynamic activation coefficient, defined as:

[0020] , ;

[0021] Where E thresh =−40dBm is the activation threshold;

[0022] δ is the damping constant, with a value of 10. −6 Up to 10 −3 This is used to prevent the denominator from being zero.

[0023] As a preferred embodiment of the present invention, in step S1, the periodic sending of network probe packets is 1-5 seconds, and when no response is received for 3 consecutive times, the VoIP connection is determined to be invalid.

[0024] As a preferred embodiment of the present invention, in step S1, the specific electrical signal characteristics of the eLink connection include: stable output of link voltage within the range of 3.3V±0.2V, current change rate ≤5mA / ms, and 1kHz pulse sequence with an interval of 100ms±10ms.

[0025] As a preferred embodiment of the present invention, in step S2, the SIP protocol feature fields include the “SIP / 2.0 / UDP” identifier in the Via header field and the user identification information in the From header field, and the parsing of the INVITE message includes extracting the caller ID, call timestamp and media type parameters.

[0026] As a preferred embodiment of the present invention, in step S3, the preset priority is: eLink incoming calls have a higher priority than VoIP incoming calls; the waiting queue only supports storing unprocessed VoIP calls, storing a maximum of 3 calls, sorted according to the order of incoming calls, with the first caller displayed first. For example, when A, B, and C call in sequence, the phone will display A by default, while B and C will be hidden and need to be viewed by pressing the navigation key down.

[0027] As a preferred embodiment of the present invention, in step S4, the specific processing of the time-domain mixing algorithm includes: sampling the two audio channels to 8kHz or 16kHz, achieving synchronization processing through timestamp alignment, and the superposition weight ratio of VoIP audio and eLink audio can be adjusted within the range of 3:7 to 7:3.

[0028] As a preferred scheme of the present application, in step S4, a call activation detection step is further included: the call state of VoIP and eLink is monitored in real time by a call activation detection module, and when the voice energy of a certain call is detected to be greater than or equal to -40 dBm, an activation signal is sent to the audio processing chip, and the audio of the call is mixed into the current mixed audio stream.

[0029] As a preferred scheme of the present application, in step S5, the gain of the pre-amplification of the audio collected by the microphone is 20-40 dB, and the audio signals routed to the VoIP channel and the eLink channel are respectively processed by G.711 or G.729 encoding.

[0030] As a preferred scheme of the present application, in step S5, the echo cancellation processing of the AEC module includes: adaptive filtering of the loudspeaker playback signal to generate an echo reference signal, and subtracting the microphone collected signal from the echo reference signal to cancel the linear echo, and the filtering coefficient is updated every 20 ms.

[0031] As a preferred scheme of the present application, in step S5, the noise suppression processing of the ANS module includes: real-time identification of the noise components in the microphone collected signal based on a pre-trained noise model, and inverse cancellation of the noise by spectral subtraction or Wiener filtering, wherein the noise model is generated by training at least 100 hours of office environment noise samples.

[0032] The present application has the following beneficial effects:

[0033] 1. Improved compatibility and synergy: through the design of coexistence of dual modes, the interoperability problem of VoIP devices and eLink software is solved, seamless synergy is realized, and frequent switching of devices by users is avoided.

[0034] 2. Realize synchronous ringing and orderly incoming call management: through synchronous ringing control and conflict avoidance logic, ensure synchronous response of dual-mode incoming calls, at the same time solve the confusion problem of simultaneous incoming calls, and improve the efficiency of incoming call processing.

[0035] 3. Support multi-party conference needs: through hardware mixing processing to realize two-way audio merging, meet enterprise three-party conference and other collaboration scenarios, and expand the functionality of the device.

[0036] 4. Optimize call quality: through echo cancellation and noise suppression technology, reduce interference, improve voice clarity, and improve user call experience.

[0037] 5. Simplify operation process: full-process automatic processing (connection monitoring, signaling detection, mixed audio optimization, etc.), reduce user operation complexity, and improve work efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1A flowchart of a conference audio merging method supporting dual-mode coexistence and synchronous ringing provided by an embodiment of the present application. DETAILED DESCRIPTION

[0039] The technical solutions of the present application will be further described below in combination with the drawings and through specific embodiments.

[0040] In the drawings, only for exemplary illustration, the representations are only schematic diagrams, not physical diagrams, and cannot be understood as limitations on the present application; in order to better illustrate the embodiments of the present application, some components in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings can be omitted.

[0041] In the drawings of the embodiments of the present application, the same or similar reference numerals correspond to the same or similar components; in the description of the present application, it should be understood that if the terms "upper", "lower", "left", "right", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore the terms describing the positional relationship in the drawings are only for exemplary illustration, and cannot be understood as limitations on the present application, for those skilled in the art, the specific meanings of the above terms can be understood according to the specific circumstances.

[0042] In the description of the present application, unless otherwise explicitly specified and limited, if the term "connection" and the like appear to indicate the connection relationship between components, the term should be broadly understood, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrated; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the communication or interaction relationship between two components. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0043] The conference audio merging method supporting dual-mode coexistence and synchronous ringing provided by the present application, as shown in Figure 1 includes the following steps:

[0044] S1, real-time monitoring of connection state: the connection state monitoring unit maintains the state of VoIP connection and eLink connection respectively, wherein, for VoIP connection, network probe packets are periodically sent and the connection validity is judged according to the response result, for eLink connection, the connection validity is judged by detecting the specific electrical signal characteristics in the link, and the state changes of the two types of connections are fed back to the central control unit in real time;

[0045] S2, dual-mode incoming call signaling detection: based on the valid connection confirmed in step S1, VoIP incoming call signaling and eLink incoming call signaling are detected respectively, wherein the VoIP incoming call signaling is completed through a SIP protocol parser, the SIP protocol characteristic field in the network data packet header is analyzed first, and then the INVITE message is deeply parsed to extract call information; the eLink incoming call signaling is completed through a special telecommunication signal detection circuit, and the characteristic signal mode such as a specific frequency pulse sequence is identified by sensing the voltage and current changes in the link;

[0046] S3, synchronous ring control: taking the incoming call signaling detected in step S2 as input, a ring drive circuit is used to realize hardware ring, and the ring drive circuit drives the ring after amplification, ring tone generation and power amplification of the trigger signal; at the same time, a central control unit executes a conflict avoidance logic, and if two incoming call signals are received at the same time, the pre-set priority is processed, when the eLink is ringing, the new incoming call is VoIP, and the new incoming call is put into a waiting queue and a ring is triggered, at this time, the eLink session is interrupted; when the VoIP call ends, the eLink session is automatically restored; the pre-set priority is that the eLink incoming call signaling priority is higher than the VoIP incoming call signaling; the waiting queue only supports storing VoIP unprocessed incoming calls, and at most 3 paths are stored, sorted by incoming order, and the first incoming call is displayed first, for example, A, B and C are incoming in order, the phone defaults to display A, B and C are hidden, and need to be viewed through the navigation key down key;

[0047] S4, hardware mixing processing: in the call state confirmed in step S3, a special audio processing chip is used to realize audio mixing, wherein the eLink audio stream enters through the USB interface, is converted in format by the decoding circuit and is transmitted to the mixer through the internal bus; the VoIP audio stream is transmitted to the mixer after being conditioned and converted in format by the special interface circuit; the mixer uses a time domain mixing algorithm to sample, convert and synchronize the two audio streams and then superimposes them according to the pre-set weight;

[0048] S5, audio routing and optimization: the audio after mixing in step S4 is encoded and power amplified and then output to the loudspeaker; at the same time, the audio collected by the microphone is pre-amplified and then enters the audio processing chip, and is routed to the VoIP channel and the eLink channel through the routing control module; and the AEC module in the audio processing chip compares the loudspeaker playback signal and the microphone collection signal to eliminate echo, and the ANS module suppresses noise to improve sound quality.

[0049] The conference audio merging method supporting dual-mode coexistence and synchronous ringing adopts the technical scheme, and through real-time monitoring of the connection state, ensures the stability of the VoIP and eLink connection, and provides a reliable basis for subsequent processing; through dual-mode incoming call signaling detection, accurate identification of incoming calls in two modes is realized; through synchronous ringing unified control, synchronous ringing of dual-mode incoming calls is realized and the conflict problem of simultaneous incoming calls is solved; through hardware mixing processing, effective merging of two audio signals is realized, supporting multi-party conference requirements; through audio routing and optimization, echo is eliminated and noise is suppressed, improving the call quality. The whole conference audio merging under dual-mode coexistence is realized, the device compatibility and user call experience are improved, and the efficient communication demand is met.

[0050] The time-domain mixing algorithm in step S4 includes dynamic weight adjustment, and the mixed audio signal y(t) is calculated as:

[0051]

[0052] Wherein:

[0053] x v (t) and x e (t) are the sampling values of the VoIP audio signal and the eLink audio signal at time t, respectively, x v (t) is the decoded signal from eLink via the USB interface; x e (t) is the signal converted by the special interface circuit from VoIP, and the sampling rate is unified to 8kHz or 16kHz;

[0054] E v (t) and E e (t) are the short-time speech energies of VoIP and eLink, respectively, in dBm;

[0055] a v (t) and a e (t) are dynamic activation coefficients, defined as:

[0056]

[0057] Wherein E thresh =−40dBm is the activation threshold;

[0058] δ is a damping constant, taking a value of 10 −6 to 10 −3 , used to prevent the denominator from being zero.

[0059] The equation parameters are further described as follows:

[0060] x v (t) and x e ​​​(t) are the sample values of VoIP and eLink audio signals at time t respectively, derived from the decoded and format converted signals in step S4, with a uniform sampling rate of 8 kHz or 16 kHz.

[0061] E v (t) and E e (t) are the short-time speech energies of VoIP and eLink respectively, in unit of dBm, calculated in real-time by the audio processing chip, with the calculation method being: the average of the square of the signal amplitude within a 20 ms window, converted to dBm (for example, ).

[0062] a v (t) and a e (t) is the dynamic activation coefficient, a binary value (0 or 1), generated by the call activation detection module, with the value being 1 when the speech energy ≥ -40 dBm, indicating that the audio of this path is activated and mixed in; otherwise, the value is 0, indicating that it is not mixed in.

[0063] E thresh : fixed activation threshold, -40 dBm, set based on the typical range of human speech energy (speech energy in a quiet environment is about -30 dBm to -10 dBm).

[0064] δ: damping constant, a small positive number (such as 10 −6 ), to ensure that the denominator is not zero, to prevent numerical overflow, the value is optimized through experiments, balancing numerical stability and sound quality.

[0065] Output y(t): mixed audio signal, output to the loudspeaker after time-domain mixing (step S5).

[0066] Example: apply this equation in the enterprise internal multi-party conference scenario (as described in Example Two below):

[0067] Scenario description: the user simultaneously accesses VoIP conference and eLink conference, when the VoIP party speaks (E v (t) = -35 dBm ≥ E thresh , so a v (t) = 1), the eLink party is silent (E e (t) = -50 dBm < E thresh , so a e (t) = 0).

[0068] Equation calculation:

[0069] ;

[0070] Output y(t) is approximately VoIP audio, ensuring clear speech of the speaking party.

[0071] Dynamic adjustment: if both parties speak at the same time (E v (t)=−30dBm,E e (t)=−25dBm, both activated), then:

[0072] ; the weight automatically favors the higher-energy eLink audio (weight ratio ~ 54.5%:45.5%), consistent with the "activated party weight increase" principle (Description Example II).

[0073] Technical effects

[0074] Improve speech intelligibility: through dynamic activation coefficients (a v (t),a e (t)) and energy proportion weight, ensure the activated party's voice dominates the mix, reducing non-activated audio interference, significantly improving speech intelligibility in multi-party conferences.

[0075] Optimize resource utilization: only when the audio is activated (energy ≥ threshold) is mixed in, avoiding invalid audio (such as silence or background noise) occupying processing resources, effectively reducing CPU load.

[0076] Enhance sound quality adaptability: combined with AEC and ANS modules, this equation reduces echo and noise interference in weight calculation, significantly improving overall sound quality MOS (Mean Opinion Score).

[0077] Support efficient meetings: dynamic weights achieve "primary and secondary speech clarity" (Description Example II), solving the problem of sound overlap in three-party conferences.

[0078] Working principle flow

[0079] This equation is implemented in hardware mixing processing (step S4), the working process is as follows:

[0080] 1. Input acquisition: the audio processing chip acquires x v (t) and x e (t) (decoded and synchronized) from VoIP and eLink interfaces.

[0081] 2. Energy calculation: real-time calculation of short-time speech energy E v (t) and E e (t) (updated every 20ms).

[0082] 3. Activation detection: compare energy with threshold E thresh =−40dBm:

[0083] If E v (t) ≥ E thresh , set a v(t) = 1 (VoIP activated).

[0084] If E e (t) ≥ E thresh , let a e (t) = 1 (eLink activated).

[0085] 4. Dynamic mixing calculation: substitute equation to calculate mixed signal y(t):

[0086] Numerator: energy-weighted sum of activated audio (the higher the energy, the greater the contribution).

[0087] Denominator: normalization factor to ensure output stability (add δ to prevent zero).

[0088] 5. Output and optimization: y(t) is output to step S5 for encoding, power amplification, and sound quality optimization (such as AEC echo cancellation). At the same time, non-activated audio (a v (t) = 0 or a e (t) = 0) is excluded to reduce mixed stream noise.

[0089] This equation introduces dynamic activation coefficients (based on a specific patent threshold of -40 dBm) and energy proportion normalization, forming a unique mixing mechanism for dual-mode audio. For example, the general formula y(t) = w1x1(t) + w2x2(t) has no activation control; this equation adds activation coefficients and energy terms in the denominator to ensure mixing only when valid speech exists.

[0090] Solves the problem of "unable to support three-party conference" in the prior art:

[0091] "Selective mixing" is achieved through activation coefficients to avoid invalid audio interference (improve compatibility).

[0092] Energy proportion weight automatically adapts to speaking intensity to achieve "clear and distinguishable primary and secondary speech" (optimize user experience).

[0093] Specifically, in this embodiment, in step S1, the period of periodically sending network probe packets is 1-5 seconds, and when there is no response for 3 consecutive times, it is determined that the VoIP connection is invalid. By using the above technical solution, the change of the VoIP connection state can be monitored in a timely and accurate manner by periodically sending network probe packets and determining the VoIP connection failure based on consecutive non-responses, avoiding the impact on the call due to the failure to discover the connection anomaly, and ensuring the real-time and reliability of the VoIP connection state monitoring.

[0094] Specifically, in the embodiment, in step S1, the specific electrical signal characteristics of the eLink connection include: the link voltage is stably output in the range of 3.3V±0.2V, the current change rate is ≤5mA / ms, and the 1kHz pulse sequence is separated by 100ms±10ms. By detecting the specific voltage, current characteristics and pulse sequence in the eLink link, the effectiveness of the eLink connection can be accurately judged, the accuracy of the eLink connection state monitoring is ensured, and a stable connection state basis is provided for the collaborative work of VoIP and eLink dual-mode.

[0095] Specifically, in the embodiment, in step S2, the SIP protocol characteristic field includes the "SIP / 2.0 / UDP" identifier in the Via header field, and the user identifier information in the From header field, and the analysis of the INVITE message includes extracting the calling number, call timestamp and media type parameter. By analyzing the SIP protocol characteristic field and deeply analyzing the INVITE message to extract the call information, the VoIP incoming call signaling and related call details can be accurately identified, accurate information support is provided for the subsequent ring control and call processing of VoIP incoming calls, and the accuracy of VoIP incoming call signaling detection is ensured.

[0096] Specifically, in the embodiment, in step S3, the preset priority is that the eLink incoming call signaling priority is higher than that of the VoIP incoming call signaling; the waiting queue only supports storing VoIP unprocessed incoming calls, at most 3, sorted by incoming order, and the first incoming call is displayed first, for example, when A, B and C come in order, the phone defaults to display A, and B and C are in a hidden state, which needs to be viewed through the navigation key down key; when the eLink is ringing, the new incoming call is VoIP, which is put into the waiting queue and triggers the ring, at this time the eLink session is interrupted; when the VoIP call ends, the eLink session is automatically restored. By presetting the priority to handle two simultaneous incoming calls and using the waiting queue to manage VoIP unprocessed incoming calls, the ringing conflict problem when the dual-mode simultaneously comes in can be effectively solved, the ringing confusion is avoided, the user can orderly process the incoming call, and the convenience and orderliness of incoming call processing are improved.

[0097] Specifically, in the embodiment, in step S4, the specific processing of the time domain mixing algorithm includes: uniformly sampling the two audio signals to 8kHz or 16kHz, synchronously processing through timestamp alignment, and the superposition weight ratio of VoIP audio and eLink audio can be adjusted in the range of 3:7 to 7:3. By uniformly sampling, synchronously processing and weight adjusting the two audio signals through the time domain mixing algorithm, the efficient merging of VoIP and eLink two audio signals is realized, the synchronization and clarity of the mixed audio are ensured, and the demand of audio merging in multi-party conference is met.

[0098] Specifically, in the present embodiment, in step S4, a call activation detection step is further included: the call activation detection module monitors the call state of VoIP and eLink in real time, and when it is detected that the voice energy of a certain call is greater than or equal to -40 dBm, an activation signal is sent to the audio processing chip, and the audio of the call is mixed into the current mixed audio stream. By using the above technical solution, the call state is judged based on the voice energy by the call activation detection module, and the audio is mixed in, which can dynamically identify the effective call audio, avoid mixing invalid audio into the mixed audio stream, improve the pertinence and efficiency of the mixed audio processing, and optimize the mixed audio quality.

[0099] Specifically, in the present embodiment, in step S5, the gain of the pre-amplification of the audio collected by the microphone is 20-40 dB, and the audio signals routed to the VoIP channel and the eLink channel are respectively processed by G.711 or G.729 encoding. By using the above technical solution, the pre-amplification and adaptive encoding processing of the microphone collected audio are performed to ensure that the collected audio signal is clear and adaptive to the transmission requirements of the VoIP and eLink channels, and to ensure that the local audio in the two-way call can be accurately and clearly transmitted to the two channels, thereby improving the quality of the two-way communication.

[0100] Specifically, in the present embodiment, in step S5, the echo cancellation processing of the AEC module includes: performing adaptive filtering on the loudspeaker playback signal to generate an echo reference signal, and subtracting the echo reference signal from the microphone collected signal to cancel the linear echo, and the filtering coefficient is updated every 20 ms. By using the above technical solution, the adaptive filtering and periodic updating of the filtering coefficient by the AEC module can effectively eliminate the echo interference in the call, avoid the influence of the echo on the call, and improve the intelligibility of the voice signal.

[0101] Specifically, in the present embodiment, in step S5, the noise suppression processing of the ANS module includes: identifying the noise components in the microphone collected signal in real time based on a pre-trained noise model, and performing inverse cancellation on the noise by spectral subtraction or Wiener filtering, wherein the noise model is generated by training at least 100 hours of office environment noise samples. By using the above technical solution, the noise in the microphone signal can be effectively suppressed and the pure voice signal can be extracted by identifying and filtering the noise based on the trained noise model, thereby improving the sound quality and intelligibility of the call.

[0102] The present application also provides the following three specific embodiments:

[0103] Embodiment one, dual-mode call cooperation in remote office scenario:

[0104] In the remote office scenario, the employee connects the computer through the SIP+USB dual-mode IP phone, and simultaneously runs the VoIP call system and the eLink collaboration software.

[0105] 1. Connection state monitoring unit periodically detects VoIP network connection (sends probe packets) and eLink USB connection (detects voltage, current characteristics), ensures that both are in stable state and feeds back to the central control unit;

[0106] 2. When there is a VoIP incoming call, the SIP protocol parser identifies the INVITE message to extract call information; if the eLink is ringing when the VoIP incoming call is received, the central control unit puts the VoIP incoming call into the waiting queue and triggers the ring, interrupting the current eLink ring session;

[0107] 3. After the employee answers the eLink conference, the system stops the VoIP ring, and the dedicated audio processing chip combines the eLink conference audio (decoded through the USB interface) with the subsequent VoIP call audio (converted through the dedicated interface circuit) through the time domain mixing algorithm, and eliminates echo and noise through the AEC and ANS modules; when the VoIP call ends, the eLink session is automatically restored;

[0108] 4. The employee speaks through the microphone, and the audio is routed to the eLink conference and VoIP call channels after being pre-amplified, realizing real-time three-way communication.

[0109] The dual-mode call coordination technology in this remote office scenario: through connection state monitoring to ensure stable VoIP and eLink connection, avoid call interruption due to connection interruption; use priority processing and waiting queue mechanism to solve the conflict problem of dual-mode incoming calls in remote office, reduce user operation of switching devices; combine hardware mixing and audio quality optimization module to realize clear merging of cross-platform call audio, improve the smoothness and efficiency of multi-party communication in remote collaboration, meet the needs of employees working from home or in different locations for multi-task call management.

[0110] Example Two: Enterprise Internal Multi-party Conference Scene

[0111] Enterprise employees need to access two different conferences (VoIP conference and eLink team conference) through dual-mode IP phones.

[0112] 1. After the connection state monitoring unit confirms that the VoIP and eLink connections are valid, the SIP protocol parser continuously monitors the VoIP conference signaling, and the eLink dedicated circuit detects the conference activation signal;

[0113] 2. When two conferences are initiated in sequence, the system triggers synchronous ringing according to the receiving order, and after the employee answers the first conference, the second conference enters the waiting queue, and is automatically mixed into the current call after the employee switches;

[0114] 3. The mixer synchronously processes two conference audio, dynamically adjusts the weight according to the speaker's voice energy (activates the weight of the primary speaker to increase), and ensures that the primary and secondary voices are clear and distinguishable;

[0115] 4. The local speaking audio is transmitted into two conference channels after being encoded, realizing real-time collaborative discussion of multiple parties, and the noise suppression module filters the office background noise to improve the call quality.

[0116] The technical effects of the enterprise internal multi-party conference scene: real-time connection state monitoring provides a stable foundation for parallel double conference, ensuring uninterrupted conference process; dynamic weight adjustment of the mixer makes the audio of different speakers in the multi-party conference clear and distinguishable, avoiding overlapping and confusion; the noise suppression function effectively filters the office environment interference, improving the intelligibility of conference voice; the waiting queue and seamless switching design allow users to flexibly manage multiple conference access, simplifying the operation process of enterprise internal multi-party collaboration and improving conference efficiency.

[0117] Example Three: Cross-platform emergency call processing

[0118] When the user uses eLink to handle emergency matters, they need to answer important VoIP incoming calls at the same time.

[0119] 1. The connection state monitoring unit provides real-time feedback on the connection status of VoIP and eLink. When eLink is in a call, the VoIP incoming call signaling is identified by the SIP parser;

[0120] 2. The central control unit triggers the waiting queue mechanism, puts the VoIP incoming call into the waiting queue and triggers the ring, interrupts the current eLink ring session, and only indicates the eLink session status through the indicator light;

[0121] 3. After the user switches to VoIP call through device keys, the system automatically mutes the eLink call audio temporarily and keeps it in the audio mixing queue. After the VoIP call ends, the eLink audio is restored and mixed in. After the VoIP call ends, the eLink session is automatically restored;

[0122] 4. Throughout the process, the echo cancellation module avoids audio interference during switching, ensuring seamless integration of emergency calls and transaction processing.

[0123] The technical effects of this cross-platform emergency call processing: connection state monitoring ensures the reliability of dual-mode connection in emergency situations, ensuring that important incoming calls are not missed; the waiting queue and mute retention mechanism avoid mutual interference between emergency calls and current transaction processing, enabling seamless switching; the echo cancellation module prevents audio noise during the switching process, ensuring the clarity of emergency calls; the overall solution allows users to efficiently handle cross-platform incoming calls when handling emergency matters, improving communication response speed and processing convenience in emergency scenarios.

[0124] In summary, the working principle of the conference audio merging method supporting dual-mode coexistence and synchronous ringing provided by the embodiment is as follows:

[0125] 1. Connection state monitoring: The VoIP and eLink connections are maintained by the connection state monitoring unit respectively. Network probe packets are periodically sent for VoIP connection, and the effectiveness is judged according to the response; the effectiveness of eLink connection is judged by detecting specific electrical signal characteristics (such as voltage and current changes) in the link, and the state change is fed back to the central control unit in real time, providing a stable basis for subsequent processing.

[0126] 2. Dual-mode incoming call signaling detection: Based on the effective connection, the VoIP incoming call signaling is identified by the SIP protocol parser to recognize the protocol characteristic field and parse the INVITE message to extract the call information; the eLink incoming call signaling is detected by detecting specific frequency pulse sequences and other characteristic signals in the link through a special circuit, realizing accurate identification of the two types of incoming calls.

[0127] 3. Synchronous ringing control: After the ringing drive circuit receives the incoming call signaling, it drives the hardware ringing through signal amplification, ring tone generation and power amplification; the central control unit processes the simultaneous incoming calls (sorted by priority, and the unanswered incoming calls enter the waiting queue) through the conflict avoidance logic, ensuring the order of the ringing; when eLink is ringing, the new incoming call is VoIP, which is put into the waiting queue and triggers the ring, and the eLink session is interrupted at this time; when the VoIP call ends, the eLink session is automatically restored; the waiting queue only stores VoIP incoming calls, up to 3, and displays them in order of incoming calls.

[0128] 4. Hardware mixing processing: The dedicated audio processing chip processes VoIP and eLink audio, VoIP audio is transmitted to the mixer after USB decoding, eLink audio is transmitted to the mixer after interface conditioning conversion, and synchronous superposition is realized through time domain mixing algorithm, supporting three-party conference.

[0129] 5. Audio routing and optimization: The mixed audio is encoded and amplified and output to the loudspeaker; the audio collected by the microphone is amplified and routed to the VoIP and eLink channels respectively, while the echo is eliminated through the AEC module and the noise is suppressed through the ANS module, improving the sound quality.

[0130] Method for use

[0131] 1. Device connection: Connect the SIP+USB dual-mode IP phone to the computer through USB to ensure simultaneous support of VoIP calls and eLink collaboration functions.

[0132] 2. Connection monitoring: The system automatically starts connection state monitoring, and maintains VoIP and eLink connections in real time, without manual intervention in connection detection.

[0133] 3. Incoming call processing: When VoIP or eLink has incoming call, the system synchronously detects the signaling and drives the device to ring; when eLink is ringing, new incoming call for VoIP is put into the waiting queue and triggers the ring, at this time the eLink session is interrupted; when the VoIP call ends, the eLink session is automatically resumed; the waiting queue only stores VoIP unprocessed incoming calls, at most 3 ways, sorted by incoming order (the first incoming call is displayed first), for example, A, B, C incoming calls in order, the phone defaults to display A, B and C in hidden state, which needs to be viewed through the navigation key down key.

[0134] 4. Conference audio merging: During the call, the system automatically mixes the audio of VoIP and eLink, supporting three-party conference; users can speak normally through the device, and the audio collected by the microphone is transmitted to the two channels after optimization.

[0135] 5. Audio quality optimization: The system automatically eliminates echo and suppresses noise, so users can obtain clear call effect without manual operation.

[0136] The above is only the preferred embodiment of the present application, and does not limit the implementation and protection scope of the present application. For those skilled in the art, it should be realized that any equivalent replacement and obvious changes made according to the content of the present application should be included in the protection scope of the present application.

Claims

1. A method for merging conference audio that supports dual-mode coexistence and synchronous ringing, characterized in that, Includes the following steps: S1. Real-time monitoring of connection status: The connection status monitoring unit maintains the status of VoIP connection and eLink connection respectively. For VoIP connection, network probe packets are periodically sent and the connection validity is judged based on the response results. For eLink connection, the connection validity is judged by detecting specific electrical signal characteristics in the link. The status changes of the two types of connections are fed back to the central control unit in real time. S2. Dual-mode incoming call detection: Based on the valid connection confirmed in step S1, VoIP incoming call and eLink incoming call are detected separately. VoIP incoming call detection is completed by a SIP protocol parser, which first analyzes the network data packet header to identify SIP protocol feature fields, and then deeply parses the INVITE message to extract call information. eLink incoming call detection is completed by a dedicated electrical signal detection circuit, which identifies feature signal patterns such as specific frequency pulse sequences by sensing changes in voltage and current in the link. S3. Synchronous Ringing Unified Control: Taking the incoming call signal detected in step S2 as input, the ringing is implemented through a ringing drive circuit. The ringing drive circuit amplifies the trigger signal, generates a ringtone, and amplifies the power before driving the ringing. At the same time, the central control unit executes conflict avoidance logic. If two incoming call signals are received simultaneously, they are processed according to a preset priority. When the eLink is ringing, if a new call is a VoIP call, it is placed in the waiting queue and the ringing is triggered. At this time, the eLink session is interrupted. When the VoIP call ends, the eLink session is automatically restored. S4. Hardware mixing processing: In the call state confirmed in step S3, audio mixing is achieved through a dedicated audio processing chip. The eLink audio stream enters through the USB interface, is converted in format by the decoding circuit, and is transmitted to the mixer through the internal bus. The VoIP audio stream is conditioned and converted in format by the dedicated interface circuit and then transmitted to the mixer. The mixer uses a time-domain mixing algorithm to sample, convert, and synchronize the two audio streams, and then superimposes them according to preset weights. S5. Audio Routing and Optimization: The audio mixed in step S4 is encoded, amplified, and then output to the speaker; at the same time, the audio picked up by the microphone is pre-amplified and enters the audio processing chip, which is then routed to the VoIP channel and the eLink channel respectively through the routing control module; the AEC module in the audio processing chip compares the speaker playback signal with the microphone pickup signal to eliminate echo, and the ANS module suppresses noise to improve sound quality.

2. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, The temporal mixing algorithm in step S4 includes dynamic weight adjustment, and the mixed audio signal y(t) is calculated as follows: ; in: x v (t) and x e (t) represents the sampled values ​​of the VoIP audio signal and the eLink audio signal at time t, respectively; E v (t) and E e (t) represents the short-time speech energy of VoIP and eLink, respectively, in dBm; a v (t) and a e (t) represents the dynamic activation coefficient; δ is the damping constant, with a value of 10. −6 Up to 10 −3 .

3. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S1, the periodic sending of network probe packets is 1-5 seconds, and the VoIP connection is determined to be invalid when no response is received for 3 consecutive times. In step S1, the specific electrical signal characteristics of the eLink connection include: stable output of link voltage within the range of 3.3V±0.2V, current change rate ≤5mA / ms, and 1kHz pulse sequence with an interval of 100ms±10ms.

4. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S2, the SIP protocol feature fields include the "SIP / 2.0 / UDP" identifier in the Via header field and the user identification information in the From header field. The parsing of the INVITE message includes extracting the caller ID, call timestamp, and media type parameters.

5. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S3, the preset priority is: eLink incoming calls have a higher priority than VoIP incoming calls; the waiting queue only supports storing unprocessed VoIP calls, storing a maximum of 3 calls, sorted by the order of incoming calls, with the first caller displayed first.

6. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S4, the specific processing of the time-domain mixing algorithm includes: sampling the two audio channels to 8kHz or 16kHz, achieving synchronization processing through timestamp alignment, and adjusting the superposition weight ratio of VoIP audio and eLink audio within the range of 3:7 to 7:

3.

7. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, Step S4 also includes a call activation detection step: the call activation detection module monitors the call status of VoIP and eLink in real time. When the voice energy of a certain call is detected to be ≥-40dBm, an activation signal is sent to the audio processing chip to mix the audio of that call into the current mix stream.

8. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S5, the audio collected by the microphone is preamplified to a gain of 20-40dB, and the audio signals routed to the VoIP channel and eLink channel are respectively encoded using G.711 or G.

729.

9. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S5, the echo cancellation processing of the AEC module includes: adaptively filtering the speaker playback signal to generate an echo reference signal, subtracting the microphone acquisition signal from the echo reference signal to eliminate linear echo, and updating the filter coefficients every 20ms.

10. The conference audio merging method supporting dual-mode coexistence and synchronous ringing according to claim 1, characterized in that, In step S5, the noise suppression processing of the ANS module includes: identifying noise components in the microphone-collected signal in real time based on a pre-trained noise model, and performing inverse noise cancellation through spectral subtraction or Wiener filtering, wherein the noise model is generated by training on at least 100 hours of office environment noise samples.

Citation Information

Patent Citations

  • A secure system for multi-party meetings regarding patient care

    CA2957237A1

  • Multi-seat multimedia dispatching system

    CN107592429A

  • Digital intelligent conference online fusion method and system

    CN118230110A

  • Multifunctional audio intelligent optimization conference system

    CN119135832A