Audio signal transmission method and related equipment

By preprocessing and mixing public and private network audio signals to generate Bluetooth protocol format audio signals, the problem of public and private network audio fusion and Bluetooth interoperability in multi-mode communication terminals is solved, achieving low-latency and high-compatibility audio transmission and improving user experience.

CN120980463AActive Publication Date: 2025-11-18SHENZHEN TINFULL TECH CO LTD

Patent Information

Application Number
CN202511494448.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing multimode communication terminals suffer from low device integration, high power consumption, cumbersome switching operations, and high latency when handling audio and Bluetooth interactions between public and private networks. They cannot achieve seamless integration of public and private network audio and Bluetooth interoperability, requiring additional hardware modules.

Method used

By collecting audio signals from public and private networks, preprocessing and mixing them, an audio signal in Bluetooth protocol format is generated. Then, through software-level protocol parsing and mixing, seamless integration of public and private network audio and Bluetooth collaborative output are achieved.

Benefits of technology

It achieves low-latency and high-compatibility transmission of audio over public and private networks, improves user experience and operational efficiency in complex communication scenarios, and avoids the need for additional hardware modules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980463A_ABST
    Figure CN120980463A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of communication, and relates to an audio signal transmission method and related equipment, and the method comprises the steps: collecting audio signals which comprise a public network audio signal and a private network audio signal; preprocessing the public network audio signal and the private network audio signal; mixing the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal; encoding the mixed audio signal into a Bluetooth protocol format through an application program, and generating a Bluetooth audio signal; and transmitting the Bluetooth audio signal to a Bluetooth device. According to the application, protocol analysis and mixing processing are carried out on the private network data packet through a software level, seamless audio fusion of the public network and the private network and cooperative output of the Bluetooth are realized, the problem that audio fusion of the public network and the private network and intercommunication between the Bluetooth and the private network need additional hardware modules is solved, low-delay and high-compatibility transmission is realized, and the transmission efficiency is improved. And the user experience and the operation efficiency in a complex communication scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, in particular to an audio signal transmission method and related equipment. BACKGROUND

[0002] The existing multi-mode communication terminal is limited in multi-network audio output and Bluetooth interaction, relies on additional hardware, and cannot share public network and private network Bluetooth resources, which is complicated and has high delay in switching operation.

[0003] This results in limited multi-network audio output, obstacles in Bluetooth and private network interaction, low integration and high power consumption due to reliance on additional hardware, and inability to share public network and private network Bluetooth resources.

[0004] Therefore, the present application is proposed. SUMMARY

[0005] The purpose of the embodiments of the present application is to provide an audio signal transmission method, device, computer equipment and storage medium to solve the problem of public and private network audio fusion and Bluetooth and private network interworking requiring additional hardware modules.

[0006] To solve the above technical problems, the embodiments of the present application provide an audio signal transmission method, which adopts the following technical solutions: An audio signal transmission method comprises the following steps: Collecting an audio signal, wherein the audio signal comprises a public network audio signal and a private network audio signal; Preprocessing the public network audio signal and the private network audio signal to obtain a preprocessed public network audio signal and a preprocessed private network audio signal; Mixing the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal; Encoding the mixed audio signal to generate a Bluetooth audio signal in Bluetooth protocol format; Transmitting the Bluetooth audio signal to a Bluetooth device.

[0007] Further, the preprocessing of the public network audio signal and the private network audio signal to obtain the preprocessed public network audio signal and the preprocessed private network audio signal comprises: adding time stamps to the public network audio signal and the private network audio signal to obtain the preprocessed public network audio signal and the preprocessed private network audio signal.

[0008] Further, the above-mentioned mixing the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal comprises: extracting the speech activity and signal strength of the preprocessed public network audio signal, and taking the speech activity and signal strength of the preprocessed public network audio signal as a first feature group; extracting the speech activity and signal strength of the preprocessed private network audio signal, and taking the speech activity and signal strength of the preprocessed private network audio signal as a second feature group; mixing the preprocessed public network audio signal and the preprocessed private network audio signal according to the first feature group and the second feature group to obtain a mixed audio signal.

[0009] Further, the above-mentioned mixing the preprocessed public network audio signal and the preprocessed private network audio signal according to the first feature group and the second feature group to obtain a mixed audio signal comprises: obtaining a total weight value of the first feature group and the second feature group; obtaining a public network distribution weight according to the first feature group and the total weight value; obtaining a private network distribution weight according to the second feature group and the total weight value; mixing the preprocessed public network audio signal and the preprocessed private network audio signal through a pre-set mixing model according to the public network distribution weight and the private network distribution weight to obtain the mixed audio signal.

[0010] Further, the above-mentioned obtaining a public network distribution weight according to the first feature group and the total weight value comprises: obtaining a priority of the public network audio signal; obtaining a public network distribution weight according to the priority, the first feature group and the total weight value.

[0011] Further, the above-mentioned encoding the mixed audio signal to generate a Bluetooth audio signal in a Bluetooth protocol format comprises: compressing the mixed audio signal to obtain compressed encoding data; packaging the compressed encoding data into a data packet conforming to the Bluetooth protocol format to generate the Bluetooth audio signal.

[0012] Further, after the Bluetooth audio signal is transmitted to a Bluetooth device, the method further comprises: acquire a transmission quality of the Bluetooth audio signal; adjust a parameter of the Bluetooth device according to the transmission quality and a preset threshold.

[0013] To solve the above technical problems, the embodiment of the application further provides an audio signal transmission method and device, which adopts the technical scheme as follows: An audio signal transmission device comprises: a collection module, configured to collect an audio signal, wherein the audio signal comprises a public network audio signal and a private network audio signal; a preprocessing module, configured to preprocess the public network audio signal and the private network audio signal to obtain a preprocessed public network audio signal and a preprocessed private network audio signal; a mixing module, configured to mix the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal; a conversion module, configured to encode the mixed audio signal to generate a Bluetooth audio signal in a Bluetooth protocol format; a Bluetooth output module, configured to transmit the Bluetooth audio signal to a Bluetooth device.

[0014] To solve the above technical problems, the embodiment of the application further provides a computer device, which adopts the technical scheme as follows: A computer device comprises a memory and a processor, wherein the memory stores computer readable instructions, and the processor executes the computer readable instructions to realize the steps of the audio signal transmission method.

[0015] To solve the above technical problems, the embodiment of the application further provides a computer readable storage medium, which adopts the technical scheme as follows: A computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by a processor to realize the steps of the audio signal transmission method.

[0016] Compared with the prior art, the embodiment of the application has the following beneficial effects: The application discloses an audio signal transmission method. The method comprises the following steps: collecting an audio signal, wherein the audio signal comprises a public network audio signal and a private network audio signal; pre-processing the public network audio signal and the private network audio signal; mixing the pre-processed public network audio signal and the pre-processed private network audio signal to obtain a mixed audio signal; encoding the mixed audio signal into a Bluetooth protocol format through an application program to generate a Bluetooth audio signal; and transmitting the Bluetooth audio signal to a Bluetooth device. The application realizes seamless fusion of public network audio and private network audio and Bluetooth collaborative output by performing protocol analysis and mixing processing on a private network data packet at a software level, solves the problem of additional hardware modules required for public network and private network audio fusion and interworking between private networks through Bluetooth, realizes low-delay and high-compatibility transmission, and improves user experience and operation efficiency in a complex communication scenario. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the scheme in the application, the drawings needed in the description of the embodiments of the application will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0018] Figure 1 is an exemplary system architecture diagram to which the application can be applied; Figure 2 is a flowchart of an audio signal transmission method provided by the application; Figure 3 is a structural schematic diagram of an audio signal transmission method device provided by the application; Figure 4 is a structural schematic diagram of a computer device provided by the application. DETAILED DESCRIPTION

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the application belongs; the terms used in the specification of the application are only for the purpose of describing specific embodiments and are not intended to limit the application; the specification, claims and above-described drawing of the application and the terms "include" and "have" and any variations thereof in the specification and claims of the application are intended to cover non-exclusive inclusion. The terms "first", "second" and the like in the specification and claims of the application or above-described drawings are used to distinguish different objects, not to describe a specific order.

[0020] Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in an embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is expressly understood that any of the embodiments described herein can be incorporated into any other embodiment.

[0021] In the conventional technology, a multi-mode communication terminal relies heavily on additional hardware to achieve basic functions when processing public network and private network audio. Public network and private network audio output usually needs to be switched by physical switches or implemented by simple superposition of independent audio circuits and loudspeakers, and cannot be truly integrated at the software level. At the same time, the Bluetooth module can only serve public network audio transmission. If it is to be coordinated with the private network, independent Bluetooth hardware and supporting circuits must be additionally provided for the private network module, greatly increasing the complexity and manufacturing cost of the device.

[0022] In addition, the conventional technology lacks intelligent recognition of audio content and dynamic adaptation to transmission environment. Audio signals of different sources can only be mechanically mixed, and cannot dynamically adjust the output strategy according to the importance of the voice or the quality of the signal, resulting in key communication content being easily disturbed or covered. The Bluetooth transmission process also has the problem of insufficient stability, making it difficult to guarantee the real-time and reliability required by private network communication in a complex electromagnetic environment.

[0023] In the present application, the public network is a large network infrastructure open to the public, built and operated by telecom operators. The core features of the public network include openness, sharing, large scale, easy access (usually paid by traffic or bandwidth), and standard protocol (TCP / IP).

[0024] The private network communication system is a communication technology standard designed specifically for a particular industry or scenario, and its main features are high reliability, low latency, and strong security. It includes DMR (Digital Mobile Radio), PDT (Professional Digital Trunking), TETRA (Terrestrial Trunked Radio), etc.

[0025] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings.

[0026] As Figure 1As shown, the system architecture 100 can include terminal devices (e.g., the first terminal device 101, the second terminal device 102, and the third terminal device 103), a network 104, and a server 105. The network 104 is a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0027] A user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0028] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, etc.

[0029] The server 105 can be a server providing various services, such as a background server supporting pages displayed on the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0030] It should be noted that the audio signal transmission method provided in the embodiments of the present application is generally executed by a terminal device, and accordingly, a public network and private network audio fusion and Bluetooth cooperative transmission device is generally provided in a terminal device.

[0031] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system architecture 100 is merely illustrative. Any number of terminal devices, networks, and servers can be provided according to implementation needs.

[0032] In the present embodiment, an electronic device (e.g., the first terminal device 101) on which an audio signal transmission method runs can include a processor 110, a memory 120, a communication interface 130, and a display 140. Figure 1The terminal device shown) can send or receive data through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection can include but is not limited to 3G / 4G / 5G connection, Wi-Fi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultrawideband) connection, and other now known or future developed wireless connection.

[0033] With reference to the accompanying drawings Figure 2 , a flowchart of one embodiment of an audio signal transmission method according to the present application is shown. The audio signal transmission method includes the following steps: Step S01, collecting audio signals, the audio signals including public network audio signals and private network audio signals. At least two types of network audio signals are collected, where the public network audio signal refers to voice or audio data transmitted based on a public communication network (such as 4G / 5G, broadband Internet, etc.), and the private network audio signal refers to voice or audio data transmitted based on a private communication network (such as a PDT / DMR / TETRA cluster network, an industrial private network, etc.). It can be received through a serial communication interface, which can be a UART (Universal Asynchronous Receiver / Transmitter), SPI (Serial Peripheral Interface), or an interface with serial data transmission capability including RS485 (i.e. using RS485 electrical standard to transmit UART format data), CAN (Controller Area Network), USB (pseudo serial mode), etc. as long as it can receive private network data packets.

[0034] In this embodiment, a collection device with multi-network access capability (such as a terminal device integrating a public network module and a private network module, or an external collection device connected by wired / wireless) can be used to adapt to public network protocols and private network protocols (such as PDT signaling protocols) to obtain two types of audio signals in real time or quasi-real time.

[0035] The collected audio signals include audio input signals and audio output signals, where the audio input signal refers to the audio signal received from other devices or apparatuses, and the audio output signal refers to the audio signal to be sent from the device or apparatus. The type of collection device is not limited to a specific hardware, which can include but is not limited to a microphone module, a network audio receiving module, an external audio interface, etc. to achieve the acquisition of public network and private network audio signals.

[0036] Step S02, preprocessing the public network audio signal and the private network audio signal to obtain a preprocessed public network audio signal and a preprocessed private network audio signal.

[0037] The public network and the private network have different communication protocols and transmission paths, and the audio signals may have problems such as timing misalignment (such as transmission delay difference), format incompatibility (such as different sampling rates and encoding methods), and parameter inconsistency (such as volume and gain difference). Direct fusion later will cause distortion, lag or information loss. Preprocessing is a prerequisite for ensuring the effect of subsequent fusion. The collected public network and private network audio signals are standardized in the early stage to eliminate the differences in timing, format, and parameters, so as to meet the basic requirements of subsequent fusion processing.

[0038] In this embodiment, an audio preprocessing mechanism is adopted, including but not limited to timing calibration (such as aligning signals through a unified clock reference), format standardization (such as converting to a unified sampling rate and encoding format), and parameter normalization (such as adjusting the volume to a preset range). The specific way of preprocessing is not limited to a specific algorithm, and the signal fusion and optimization are realized, for example, through a signal conversion module at the software level or an adaptive circuit at the hardware level.

[0039] Step S03, mixing the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal.

[0040] The preprocessed public network and private network audio signals are integrated into a single mixed audio signal, which needs to carry the core information of both types of signals (such as public network dispatch instructions + private network alarm voice) at the same time, and does not lose the key content of any signal.

[0041] In this embodiment, the preset mixing model can dynamically adjust the processing strategy. When only one of the public network or the private network has valid signal, the signal is directly output to avoid invalid mixing. When both signals exist, the public network audio signal and the private network audio signal are weighted and fused (such as higher priority signal has a higher proportion), time domain multiplexing (such as time slot bearing), or frequency domain separation (such as different frequency bands bearing different signals) according to the characteristic data of the audio signal (such as priority and intensity). The specific mixing algorithm is not limited to a specific type, as long as it can realize the effective integration of multiple signals.

[0042] Step S04, encoding the mixed audio signal to generate a Bluetooth audio signal in Bluetooth protocol format.

[0043] In the prior art, the terminal Bluetooth module is usually bound to the public network module, and the private network audio cannot be transmitted through Bluetooth (additional hardware support is required), which causes the user to be unable to receive the private network signal when using Bluetooth earphones and other devices, thereby limiting the flexibility of wireless interaction. Therefore, the mixed audio signal obtained in step S03 needs to be processed according to the specification of the Bluetooth transmission protocol and then transmitted to a device (such as a Bluetooth earphone, a Bluetooth speaker, etc.) with Bluetooth receiving function without adding additional hardware modules and devices.

[0044] In this embodiment, lossless encoding such as PCM (Pulse Code Modulation), A2DP (Advanced Audio Distribution Profile), etc. can be used, and low-delay lossy encoding such as LC3 (Low Complexity Communication Codec), SBC (Subband Codec), AAC (Advanced Audio Coding), etc. can also be used, which is selected according to the bandwidth requirement. After encoding, the encoded audio data is packaged into a Bluetooth data packet according to the specification of the Bluetooth audio transmission protocol (including but not limited to A2DP and LEAudio (Low Energy Audio), etc.). The data packet needs to contain a Bluetooth protocol header (including target device address, source device address, channel identifier, data type), audio metadata (such as sampling rate, encoding format, code rate) and audio data body, to ensure that the Bluetooth device can be parsed and decoded.

[0045] Step S05 transmits the Bluetooth audio signal to a Bluetooth device.

[0046] The Bluetooth module of the existing terminal is usually bound to the public network module, and only the public network audio can be transmitted. In order to realize the transmission of the private network signal, a corresponding hardware module needs to be added, and the user needs to use a wired earphone (to listen to the private network) and a Bluetooth earphone (to listen to the public network) at the same time, which is complicated and has poor experience. If multiple users need to receive mixed audio (such as a multi-person cooperative emergency team), multiple wired interfaces need to be deployed for wired output, which has high hardware cost and complex wiring. Wireless output can solve the problem of parallel reception of multiple devices. The encoded Bluetooth audio signal is transmitted from the multi-mode communication terminal to a target Bluetooth device (a device with Bluetooth audio receiving capability), realizing wireless output of mixed audio.

[0047] In this embodiment, the terminal and the Bluetooth device establish a stable Bluetooth connection through a "Bluetooth pairing process" (such as PIN code pairing, NFC quick pairing, BLE broadcast pairing), which can be one-to-one connection (such as terminal to a single Bluetooth headset) or one-to-many connection (such as terminal to multiple Bluetooth speakers). When connecting, select the transmission parameters (such as Bluetooth version, encoding format, channel bandwidth) to ensure that the capabilities match those of the Bluetooth device (such as old devices negotiating Bluetooth 4.2 + SBC, and new devices negotiating Bluetooth 5.3 + LC3).

[0048] Bluetooth audio data packets are sent through the audio channel of the Bluetooth link, such as the isochronous channel of the synchronous connection-oriented link LEAudio of the A2DP protocol, which can be a combination of continuous transmission and batch transmission. For example, real-time voice scenarios use continuous transmission (one data packet is sent every 20 ms) to ensure low latency, and non-real-time scenarios (such as audio playback) use batch transmission to improve transmission efficiency.

[0049] The application realizes seamless fusion and Bluetooth collaborative output of public network and private network audio through protocol analysis and mixing of private network data packets at the software level, solves the problem of additional hardware modules required for public and private network audio fusion and interworking between private networks, realizes low-latency, high-compatibility transmission, and improves user experience and operation efficiency in complex communication scenarios.

[0050] In some optional implementations of the embodiment, the preprocessing of the public network audio signal and the private network audio signal to obtain the preprocessed public network audio signal and the preprocessed private network audio signal includes: Adding timestamps to the public network audio signal and the private network audio signal to obtain the preprocessed public network audio signal and the preprocessed private network audio signal.

[0051] Due to differences in transmission path, protocol stack, and network load, there are transmission delay differences between public networks and private networks. Public network transmission delay is usually 20-50 ms (affected by base station load and wireless environment), and private network transmission delay is usually 10-30 ms (real-time optimized by cluster network). If mixed directly, the time difference between the two signals may exceed 10 ms, causing public network voice and private network voice to overlap and making it difficult to distinguish, which seriously affects information transmission accuracy. For example, the public network instruction "execute immediately" and the private network alarm "device failure" become "immediate device failure" due to timing misalignment.

[0052] A uniform time reference timestamp is added to the collected public network audio signal and private network audio signal, and the time difference between the public network audio signal and the private network audio signal is calculated to align the two signals on the time axis, ensuring that there is no timing misalignment (such as echo, front and back delay) when mixing. The timestamp accuracy is not less than 1 ms, and the time difference between the two signals after synchronization is ≤5 ms.

[0053] For example, based on the clock source built-in in the terminal device, which can be a real-time clock RTC, a network time service clock NTP, a GNSS synchronous clock, etc., a globally unified time reference (such as a Unix timestamp) is generated. At the moment when the public network audio signal is output from the public network module, the current time reference value is recorded as its timestamp (denoted as T1), and at the moment when the audio playback is parsed from the serial communication interface, the current time reference value is recorded as its timestamp (denoted as T2).

[0054] The time difference ΔT = |T1-T2| of the two signals is calculated, and a delay compensation is added to the signal with the earlier timestamp (for example, ΔT delay is added to the public network signal when T1 < T2), or the signal with the later timestamp is played in advance (implemented through buffer preloading), so that the time difference of the two signals is ≤5 ms. ΔT is recalculated every 100 ms to adjust the compensation amount in real time, avoiding long-term synchronization deviation, for example, when the public network load suddenly increases and the delay rises from 20 ms to 50 ms, the synchronization system can complete the compensation adjustment within 200 ms.

[0055] The application reduces the degree of confusion of the mixed audio and improves the user experience by synchronizing the time difference of the two signals and dynamically calibrating the mechanism.

[0056] In some optional implementations of the embodiment, the mixing of the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal includes: extracting the speech activity and signal strength of the preprocessed public network audio signal as a first feature group, extracting the speech activity and signal strength of the preprocessed private network audio signal as a second feature group, and mixing the preprocessed public network audio signal and the preprocessed private network audio signal according to the first feature group and the second feature group to obtain a mixed audio signal.

[0057] Fixed ratio mixing may cause the inability to adapt to dynamic changes in signals, such as when the public network changes from "effective speech" to "silence", the mixed audio volume is still reduced by 50% when mixed at a ratio of 50%. It may also cause the inability to reflect the priority difference between signals, such as mixing the private network emergency alarm and the public network ordinary speech at the same ratio, and the emergency information is diluted.

[0058] Quantitative features reflecting signal properties are extracted from the pre-processed public network and private network audio signals respectively, including signal intensity, speech activity, etc., to form a first feature group and a second feature group, which can be in the form of numerical values or labels, etc. The first feature group refers to the feature group including the quantitative features of the public network audio signal, and the second feature group refers to the feature group including the quantitative features of the private network audio signal. Finally, the public network audio signal and the private network audio signal are mixed according to the first feature group and the second feature group to obtain a mixed audio signal.

[0059] For example, speech activity can be obtained by short-time energy method (a commonly used speech signal processing method for extracting energy features of speech signals) to determine whether the signal is valid speech (non-silence, non-noise). Through preliminary screening by the short-time energy method, if the signal is determined to be valid speech (excluding obvious silence or strong noise), the audio signal is divided into continuous short segments (i.e. "audio frames") with a length of 20 ms / frame. The purpose of the frame processing is to convert the long-time varying audio signal into a short-time approximately stationary segment, which facilitates accurate calculation of features.

[0060] For each 20 ms audio frame, the "time domain square sum" (i.e. the square sum of signal amplitude) of all audio signal sampling points in the frame is calculated to obtain the short-time energy. The number of times of sign change (i.e. the number of times of signal amplitude crossing from positive to negative or vice versa) of the audio signal in the frame is counted to obtain the zero-crossing rate. The short-time energy and the zero-crossing rate of the audio frame are compared with a preset threshold value, and if the energy is >-40 dBFS and the zero-crossing rate is <50 times / frame, the active state of the frame is determined to be active (1), otherwise it is non-active (0). The active state of continuous N frames (usually 3 to 5 frames) can be obtained by proportion calculation or weighted summation. For example, the weighted proportion formula can be used to obtain: wherein, represents the 1st frame (the earliest frame) in the continuous N frames, represents the Nth frame (the latest frame) in the continuous N frames, is the active state of each frame, is the weight of each frame. N is 3 frames, and the weights of the continuous 3 frames are set to 0.5 (the latest frame), 0.3 (the middle frame), and 0.2 (the oldest frame). If all the 3 frames are active, the activity value is (1x0.5+1x0.3+1x0.2) / 1=1.0; if only the latest frame is active, the activity value is 0.5 / 1=0.5.

[0061] Signal intensity can be obtained by calculating the signal-to-noise ratio. Specifically, the noise power in the signal is estimated by using the spectral subtraction method, and then the ratio of the signal power to the noise power is calculated: The ratio is The mapping is an intensity value (e.g., 90 for SNR=30dB and 60 for SNR=20dB).

[0062] According to the priorities, signal strengths, and voice activities of the public network and the private network obtained above, a first feature group and a second feature group are respectively formed, and feature analysis and weight distribution are performed on the first feature group and the second feature group. Then, according to the weights, a signal superposition and noise reduction optimization process is performed on the preprocessed public network and private network audio signals, and the two audio signals are combined into a single mixed audio signal, which not only retains the effective information of the two signals but also highlights the core voice according to the priority.

[0063] Thus, through feature extraction and feature analysis, the information density is improved, the user can prioritize the identification of emergency information, and the user experience consistency is improved.

[0064] In some optional implementations of the embodiment, the mixing of the preprocessed public network audio signal and the preprocessed private network audio signal according to the first feature group and the second feature group to obtain a mixed audio signal includes: obtaining a total weight value of the first feature group and the second feature group, obtaining a public network distribution weight according to the first feature group and the total weight value, obtaining a private network distribution weight according to the second feature group and the total weight value, and mixing the preprocessed public network audio signal and the preprocessed private network audio signal through a preset mixing model according to the public network distribution weight and the private network distribution weight to obtain the mixed audio signal.

[0065] In the embodiment, first, the total values of the signal strengths and voice activities of the first feature group and the second feature group can be calculated to obtain a total weight value. For example, the total weight value=(private network intensity value×private network activity)+(public network intensity value×public network activity).

[0066] Then, the distribution weight of the public network audio signal is calculated through the signal strength and voice activity in the first feature group, and the distribution weight of the private network audio signal is calculated through the signal strength and voice activity in the second feature group. For example, the calculation formulas of the public network distribution weight and the private network distribution weight are as follows: public network distribution weight=(public network intensity value×public network activity) / total weight value private network distribution weight=(private network intensity value×private network activity) / total weight value Finally, the preprocessed public network and private network audio signals are superimposed through a mixing model, a public network distribution weight, and a private network distribution weight to integrate into a mixed audio signal. The signal superposition can use linear superposition or nonlinear superposition algorithms, for example, linear superposition can be used to determine the mixed audio signal=public network signal×public network weight+private network signal×private network weight.

[0067] The application can avoid invalid mixed audio, avoid mixed audio being covered by low priority signals, and improve information effectiveness by feature extraction and feature analysis.

[0068] In some optional implementations of the embodiment, the public network distribution weight is obtained according to the first feature group and the total weight value, including: obtaining the priority of the public network audio signal; obtaining the public network distribution weight according to the priority, the first feature group and the total weight value.

[0069] If only signal strength is distributed, if it is low priority, it will cover the high priority but slightly low strength signal. If only priority is distributed, if the signal strength is very low, it cannot be identified after mixing. Therefore, the two should be combined to ensure that high priority signals are given priority and to avoid blindly increasing the weight of signals with too low strength.

[0070] The priority can be obtained through signaling identification or preset label. The signaling identification (such as the priority of the scheduling resource of the 5G public network, the emergency call mark of the DMR private network) is extracted from the data of the public network and private network modules, and the priority is determined according to the signaling identification. If there is no signaling identification, the user preset label (such as public network ordinary voice for 4 and private network emergency voice for 10) is used. The priority value range can be 1 to 10 (10 is the highest).

[0071] For example, the priority (P) can be valued at 1 to 10, which can be preset by the user, such as private network emergency = 10 and public network ordinary = 6. The signal strength (S) can be valued at 0 to 100 (mapped by SNR, such as SNR = 30 dB → S = 90), if <30, then limit ≤0.3 to avoid too high weight of low definition signals. The voice activity (V) can be set to 0.8 to 1.0, and the audio signals here are all valid voice. The total weight value is adjusted according to the priority of the public network audio signal and the private network audio signal to obtain the adjusted total weight value. For example, the adjusted total weight value = The public network distribution weight The calculation formula can be: The total weight value wherein, is the public network priority and is the public network signal strength, is the private network priority and is the private network signal strength. For example, in scenario 1: private network ( = 10, = 90, V = 0.82), public network ( =6, =80,V=0.83), the public network allocation weight =(6×80×0.8) / [(10×90×0.83)+(6×80×0.8)]≈0.34. The public network allocation weight =10, =20), the private network (S =6, S=90), because the public network intensity value =20<30, the public network is forced ≤0.3, so the public network weight =0.3.

[0072] In calculating the private network allocation weight, the priority of the private network audio signal can also be obtained, and the private network allocation weight is obtained according to the priority, the second feature group and the total weight value. The specific implementation is consistent with the above-mentioned public network allocation weight, for example, in scenario 1: the private network (S =10, =90,V=0.82), the private network allocation weight =(10×90×0.83) / [(10×90×0.83)+(6×80×0.8)]≈0.66.

[0073] The public network allocation weight and the private network allocation weight are recalculated every 200ms (dynamic change with signal intensity). When the priority difference is ≥5, such as Pprivate network=10 and Ppublic network=4, even if the private network S is slightly lower (such as Sprivate network=70 and Spublic network=90), the Wprivate network≥0.55 (high priority is guaranteed) is still guaranteed.

[0074] The application determines the weight by priority and audio signal intensity, avoids the limitation of a single factor, focuses on importance and clarity, and improves the confidence of the audio signal. When the high priority signal intensity is slightly low, a reasonable weight can still be obtained, avoiding being completely covered by a low priority strong signal. The signal weight is also small when the intensity value is too small, avoiding dilution of effective information by noise or fuzzy signals.

[0075] In some optional implementations of the embodiment, the encoding of the mixed audio signal to generate a Bluetooth audio signal in a Bluetooth protocol format includes: Compressing the mixed audio signal to obtain compressed encoding data; and encapsulating the compressed encoding data into a data packet conforming to the Bluetooth protocol format to generate the Bluetooth audio signal.

[0076] If the original mixed signal is in uncompressed PCM format (16 kHz sampling rate, 16 bit depth), the code rate is 256 kbps, and the available bandwidth of Bluetooth (such as A2DP) is usually 300 to 500 kbps, if multiple signals are transmitted or other data is superimposed, it is easy to cause packet loss due to insufficient bandwidth. Therefore, the mixed audio signal needs to be compressed and encoded, and then the compressed and encoded data is added with a protocol header, metadata, and check information, so as to generate a standardized Bluetooth data packet.

[0077] For example, a built-in compression algorithm library (which can include LC3, SBC, AAC, etc.) can be used to automatically select a feasible compression algorithm according to the terminal hardware performance (such as CPU power and memory) and the Bluetooth device selection algorithm (the encoding format supported by the Bluetooth device). If the Bluetooth device supports a low-delay algorithm (such as LC3, delay ≤ 10 ms), it is preferred; if not, it is downgraded to a general algorithm (such as SBC, compatible with 99% of Bluetooth devices).

[0078] According to the Bluetooth link quality (judged by RSSI signal strength), dynamic adaptation is performed, for example, high code rate (such as 256 kbps) is used for good link (RSSI ≥ -70 dBm) to ensure sound quality, and low code rate (such as 128 kbps) is used for poor link (RSSI ≤ -85 dBm) to reduce packet loss.

[0079] Finally, it is packaged into a data packet format conforming to the Bluetooth protocol format. The protocol header (8-16 bytes) contains the source MAC address, target MAC address, protocol type, and data packet sequence number. The metadata (4-8 bytes) contains the compression algorithm identifier, sampling rate, frame length, and code rate. The data body (100-255 bytes) includes the compressed and encoded audio data. The check tail (2-4 bytes) is checked by CRC16 / CRC32 to detect data integrity.

[0080] The application enables the audio to be recognized by most Bluetooth devices (such as mobile phones, earphones, vehicles, and sound boxes) through standardized packaging, and the protocol header sequence number and check tail realize packet loss detection and packet loss discarding, thereby improving the data correctness.

[0081] In some optional implementation manners of the embodiment, after the Bluetooth audio signal is transmitted to the Bluetooth device, the following steps are included: The transmission quality of the Bluetooth audio signal is obtained from the Bluetooth device. According to the transmission quality and a preset threshold, the parameters of the next Bluetooth audio signal transmission to the Bluetooth device are adjusted.

[0082] The key quality indicators of the Bluetooth audio signal, such as transmission delay, packet loss rate, and interference strength, are collected from the Bluetooth device at a fixed time, to form quantized transmission quality data, which provides a basis for subsequent strategy adjustment.

[0083] For example, the transmission delay can be calculated by the time stamp difference method. The terminal records the sending time stamp (T_send) when sending a data packet. The Bluetooth device returns an acknowledgement frame after receiving it, containing the receiving time stamp (T_recv). The transmission delay ΔT = T_recv - T_send (excluding the Bluetooth device processing time, only counting the air transmission time). The average transmission delay is calculated every 100 ms (taking the average of 5 measurements).

[0084] The packet loss rate can be calculated by the frame number continuous check method. The terminal assigns a continuous frame number (0~65535) to each data packet. The total number of frames sent in 1 second (N_total) and the number of frames confirmed by the Bluetooth device (N_recv) are counted. Packet loss rate = (N_total - N_recv) / N_total x 100%.

[0085] The interference strength can be monitored by the channel RSSI (received signal strength) indication. The Bluetooth module scans the RSSI value of the current channel in real time (unit dBm). RSSI > -80 dBm (absolute value small) indicates strong interference, and RSSI < -90 dBm indicates weak interference. The RSSI value is recorded every 50 ms, and the average interference strength in 1 second is calculated. Real-time recording of quality data (delay, packet loss rate, interference strength).

[0086] According to the monitored transmission quality data (delay, packet loss rate, interference strength), the transmission parameters (such as channel, code rate, power) of the Bluetooth audio signal are dynamically adjusted to maintain the transmission quality within the preset threshold (delay ≤ 30 ms, packet loss rate ≤ 5%). Taking delay as an example, when delay > 30 ms or packet loss rate > 5%, reduce the compression encoding code rate (such as from 256 kbps to 192 kbps, and then to 128 kbps). When the quality recovers, such as delay ≤ 25 ms and packet loss rate ≤ 3%, in order to avoid frequent fluctuations, gradually increase the code rate.

[0087] The present application can avoid the problem from expanding when it is found by real-time detection of signal quality in the transmission process and dynamic strategy transformation. The quantitative quality data provides a clear basis for adjustment, improving the effectiveness of data.

[0088] Further referring to Figure 3 , as an implementation of the method shown in Figure 2 , the present application provides an embodiment of an audio signal transmission device, which corresponds to the method embodiment shown in Figure 2 . The device can be applied to various electronic devices.

[0089] As Figure 3As shown, the audio signal transmission device 300 described in the embodiment includes a collection module 301, a preprocessing module 302, a mixing module 303, a conversion module 304, and a Bluetooth output module 305. The collection module 301 is configured to collect audio signals, wherein the audio signals include public network audio signals and private network audio signals, and the private network audio signals are obtained by receiving data packets from a private network module through a serial communication interface. The preprocessing module 302 is configured to preprocess the public network audio signals and the private network audio signals to obtain preprocessed public network audio signals and preprocessed private network audio signals. The mixing module 303 is configured to mix the preprocessed public network audio signals and the preprocessed private network audio signals to obtain mixed audio signals. The conversion module 304 is configured to encode the mixed audio signals into a Bluetooth protocol format through an application program to generate Bluetooth audio signals. The Bluetooth output module 305 is configured to transmit the Bluetooth audio signals to a Bluetooth device.

[0090] The audio signal transmission device provided in the application realizes seamless fusion of public network audio signals and private network audio signals and Bluetooth collaborative output through protocol analysis and mixing processing of private network data packets at the software level, solves the problem of additional hardware modules required for intercommunication between public network audio signals and private network audio signals and between Bluetooth, realizes low-delay and high-compatibility transmission, and improves user experience and operation efficiency in complex communication scenarios.

[0091] In some embodiments, the audio signal transmission device provided in the application can also be applied to related hardware circuits and architectures.

[0092] Specifically, the acquisition module 301 is configured to acquire the public network audio signal and the private network audio signal, and can include a digital audio interface circuit, an analog audio acquisition circuit, a universal serial bus circuit, and the like. The preprocessing module 302 is responsible for timing calibration and format conversion of the audio signal, and the hardware structure thereof can include a DSP (Digital Signal Processor) core, a special hardware accelerator, a FPGA (Field Programmable Gate Array), a logic unit of a CPLD (Complex Programmable Logic Device), a main processor (CPU / AP), and the like. The mixing module 303 is configured to intelligently mix the two audio signals based on feature analysis, and the hardware structure thereof can include a parallel computing unit of a GPU (Graphics Processing Unit), an AI accelerator (NPU), a DSP, and a high-performance CPU. The conversion module 304 is responsible for encoding the mixed audio into a Bluetooth format, and the hardware structure thereof can include a hardware audio encoder, an encoding logic in a Bluetooth chip, and a main processor. The Bluetooth output module 305 is responsible for transmission of the Bluetooth signal, and the hardware structure thereof can include a Bluetooth baseband processor and a RF (Radio Frequency) front-end circuit (such as a power amplifier, a low-noise amplifier, and an antenna switch), and the like. The Bluetooth output module 305 modulates the digital baseband signal generated by the conversion module 304 to a 2.4 GHz frequency band, and radiates it out through an antenna.

[0093] The method and device described in the present application are not limited to pure software implementation, and can be effectively mapped into various hardware circuit architectures, including but not limited to a highly integrated SoC, a discrete component embedded system, and a slight hardware modification to an existing device, for intelligent preprocessing, feature-aware mixing, and unified Bluetooth packaging of multi-source audio.

[0094] To solve the above technical problems, the present application also provides a computer device. For details, please refer to Figure 4 , Figure 4 The basic structure block diagram of the computer device of the present embodiment is shown in FIG. 4.

[0095] The computer device 4 includes a memory 41, a processor 42, and a network interface 43, which are connected to each other through a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art can understand that the computer device herein is a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a Digital Signal Processor (DSP), an embedded device, and the like.

[0096] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server, or the like. The computer device can interact with a user through a keyboard, a mouse, a remote controller, a touchpad, a voice control device, or the like.

[0097] The memory 41 can include at least one type of readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, or the like), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, or the like. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or a memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, or the like. Of course, the memory 41 can include both an internal storage unit and an external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store an operating system and various application software installed in the computer device 4, such as computer readable instructions of the audio signal transmission method, or the like. In addition, the memory 41 can also be used to temporarily store various data that has been output or will be output.

[0098] The processor 42 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer readable instructions or process data stored in the memory 41, such as computer readable instructions of the audio signal transmission method.

[0099] The network interface 43 can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0100] The computer device provided in the application implements seamless fusion and Bluetooth collaborative output of public network and private network audio through protocol analysis and mixed processing of private network data packets at a software level, solves the problem of additional increase of hardware modules for public and private network audio fusion and interworking between Bluetooth and private network, realizes low-delay and high-compatibility transmission, and improves user experience and operation efficiency in complex communication scenarios.

[0101] The application further provides another implementation, namely providing a computer readable storage medium storing computer readable instructions, which can be executed by at least one processor to enable the at least one processor to perform the steps of the audio signal transmission method as described above.

[0102] The computer readable storage medium provided in the application implements seamless fusion and Bluetooth collaborative output of public network and private network audio through protocol analysis and mixed processing of private network data packets at a software level, solves the problem of additional increase of hardware modules for public and private network audio fusion and interworking between Bluetooth and private network, realizes low-delay and high-compatibility transmission, and improves user experience and operation efficiency in complex communication scenarios.

[0103] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and a general hardware platform, and of course, they can also be realized by hardware, but in many cases, the former is a better implementation. Based on such understanding, the technical solutions of the application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the application.

[0104] Obviously, the above-described embodiments are only some of the embodiments of the application, not all the embodiments, and the preferred embodiments of the application are given in the drawings, but do not limit the patent scope of the application. The application can be implemented in many different forms, and conversely, the purpose of providing these embodiments is to make the disclosure of the application more thorough and comprehensive. Although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the contents of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the patent protection scope of the application.

Claims

1. An audio signal transmission method, characterized by, The method comprises the following steps: Collecting an audio signal, wherein the audio signal comprises a public network audio signal and a private network audio signal; Preprocessing the public network audio signal and the private network audio signal to obtain a preprocessed public network audio signal and a preprocessed private network audio signal; Mixing the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal; Encoding the mixed audio signal to generate a Bluetooth audio signal in a Bluetooth protocol format; Transmitting the Bluetooth audio signal to a Bluetooth device.

2. The audio signal transmission method of claim 1, wherein, The preprocessing of the public network audio signal and the private network audio signal to obtain the preprocessed public network audio signal and the preprocessed private network audio signal comprises: Adding a timestamp to the public network audio signal and the private network audio signal to obtain the preprocessed public network audio signal and the preprocessed private network audio signal.

3. The audio signal transmission method of claim 1, wherein, The mixing of the preprocessed public network audio signal and the preprocessed private network audio signal to obtain the mixed audio signal comprises: Extracting the speech activity and signal strength of the preprocessed public network audio signal as a first feature group; Extracting the speech activity and signal strength of the preprocessed private network audio signal as a second feature group; Mixing the preprocessed public network audio signal and the preprocessed private network audio signal according to the first feature group and the second feature group to obtain the mixed audio signal.

4. The audio signal transmission method according to claim 3, characterized in that, The mixing of the preprocessed public network audio signal and the preprocessed private network audio signal according to the first feature group and the second feature group to obtain the mixed audio signal comprises: Obtaining a total weight value of the first feature group and the second feature group; Obtaining a public network allocation weight according to the first feature group and the total weight value; Obtaining a private network allocation weight according to the second feature group and the total weight value; Mixing the preprocessed public network audio signal and the preprocessed private network audio signal according to the public network allocation weight and the private network allocation weight through a preset mixing model to obtain the mixed audio signal.

5. The audio signal transmission method of claim 4, wherein, The obtaining of the public network allocation weight according to the first feature group and the total weight value comprises: Obtaining a priority of the public network audio signal; Obtaining the public network allocation weight according to the priority, the first feature group, and the total weight value.

6. The audio signal transmission method of claim 1, wherein, The encoding of the mixed audio signal to generate the Bluetooth audio signal in the Bluetooth protocol format comprises: Compressing the mixed audio signal to obtain compressed encoding data; Packaging the compressed encoding data into a data packet conforming to the Bluetooth protocol format to obtain the Bluetooth audio signal.

7. The audio signal transmission method of claim 1, wherein, After the transmitting of the Bluetooth audio signal to the Bluetooth device, the method further comprises: Obtaining a transmission quality of the Bluetooth audio signal; Adjusting parameters of the Bluetooth device according to the transmission quality and a preset threshold.

8. An audio signal transmission apparatus characterized by comprising: The method comprises the following steps: The collection module is used for collecting an audio signal, wherein the audio signal comprises a public network audio signal and a private network audio signal. The preprocessing module is used for preprocessing the public network audio signal and the private network audio signal to obtain a preprocessed public network audio signal and a preprocessed private network audio signal. The mixing module is used for mixing the preprocessed public network audio signal and the preprocessed private network audio signal to obtain a mixed audio signal. The conversion module is used for encoding the mixed audio signal to generate a Bluetooth audio signal in a Bluetooth protocol format. The Bluetooth output module is used for transmitting the Bluetooth audio signal to a Bluetooth device.

9. A computer device, comprising: The computer readable storage medium stores computer readable instructions, and the processor executes the computer readable instructions to realize the steps of the audio signal transmission method in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and the processor executes the computer readable instructions to realize the steps of the audio signal transmission method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Public and private network fusion system and data processing method thereof

    CN112910488A

  • Cross-platform audio data communication processing method and system based on multi-protocol support

    CN120783773A

  • Multi-input wireless digital transmitting device

    CN219268842U

  • Controller integrated audio codec for advanced audio distribution profile audio streaming applications

    US20080287063A1

Cited By

  • Real-time speech recognition method based on Bluetooth audio stream

    CN121545524A