Call recording method based on dynamic sampling rate adjustment and Bluetooth earphone

Through real-time monitoring and dynamic adjustment of the sampling rate, combining multi-microphone arrays and adaptive filtering technology for noise reduction and echo cancellation, and adopting a two-layer cache mechanism and data retransmission strategy, the problem of unstable recording quality and low transmission efficiency in wireless transmission environments is solved, achieving high-quality and efficient call recording.

CN120128979APending Publication Date: 2025-06-10VISION INTELLIGENCE CO LTD

Patent Information

Application Number
CN202510279966.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to bandwidth fluctuations in wireless transmission environments, resulting in unstable recording quality, insufficient noise resistance, low data transmission efficiency, and lack of dynamic sampling rate adjustment mechanism and multi-level cache management.

Method used

By monitoring the transmission rate of audio data packets in real time, dynamically adjusting the sampling rate, reducing it from 16kHz to 8kHz, to reduce the amount of data to adapt to bandwidth limitations, and combining multi-microphone arrays and adaptive filtering technology for noise reduction and echo cancellation, while adopting a dual-layer cache mechanism and data retransmission strategy.

Benefits of technology

It significantly improves the quality and stability of call recordings, adapts to transmission efficiency in different network environments, ensures the continuity and reliability of recordings, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128979A_ABST
    Figure CN120128979A_ABST
Patent Text Reader

Abstract

The invention discloses a call recording method based on dynamic sampling rate adjustment and a Bluetooth earphone, and aims to improve the quality and stability of call recording in a wireless transmission environment. According to the method, uplink audio data and downlink audio data are collected at the same time through the Bluetooth headset, and the transmission rate of a downlink data packet is monitored in real time. When the rate is lower than the preset threshold value, the audio sampling rate is automatically reduced from 16 kHz to 8 kHz to adapt to bandwidth limitation, and recording interruption or tone quality reduction is avoided. The adjusted audio data is cached through a double-layer cache mechanism after being coded and compressed, the first cache temporarily stores the adjusted audio data, and the second cache manages data packets to be sent and retransmitted, so that the continuity and integrity of data transmission are ensured. The method is suitable for wireless audio equipment such as a Bluetooth earphone, and the reliability, adaptability and user experience of call recording are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice communication, and in particular to a method for improving call quality by dynamically adjusting a sampling rate. Background Art

[0002] With the rapid development of mobile communications and multimedia technologies, headphones, as an important accessory for mobile terminals (such as mobile phones, tablets, and laptops), are constantly expanding their functions and application scenarios. Traditional headphones are mainly used to transmit or play audio signals to meet the user's auditory needs. However, with the diversification of user needs, modern headphones not only need to have high-quality audio playback capabilities, but also need to support call recording, noise reduction, echo cancellation and other advanced functions to enhance the user's overall experience.

[0003] In the announcement number CN1 07580102A "Headphone and Headphone Recording Method", a solution for realizing call recording on the headset side is proposed. This method sets a physical button on the headset shell, and the user can directly trigger the recording function by pressing the button without operating the mobile terminal, thereby simplifying the recording operation steps. The headset integrates a recording module and a storage unit inside, which can simultaneously record the uplink audio data transmitted from the mobile terminal (i.e., the other party's speech content) and the downlink audio data collected by the microphone (i.e., the user's speech content). After the recording is completed, the user can also directly play back the recording content through the playback module on the headset. This design effectively solves the cumbersome problem of users frequently operating the mobile terminal to start recording during a call, and improves the convenience and privacy of recording. However, this technical solution mainly focuses on the local recording function of the headset, fails to fully consider the impact of bandwidth fluctuations in the wireless transmission environment on audio data transmission and recording quality, and does not propose a corresponding sampling rate dynamic adjustment mechanism. For noise suppression and echo cancellation of multi-microphone arrays, the existing solutions also lack efficient processing methods in complex environments, resulting in the difficulty of achieving ideal recording quality in noisy or multi-echo environments, affecting user experience.

[0004] In the announcement number CN103354588A "Method, device and system for determining recording and playback sampling rate", a method for solving the voice data delay problem is proposed from the perspective of Internet phone (VoIP) application. This method polls multiple predetermined sampling rates to obtain the buffer size corresponding to each sampling rate, and selects the sampling rate corresponding to the minimum buffer duration as the recording and playback sampling rate of the current call, thereby minimizing the time required for the sound card to fill the buffer, reducing voice delay, and improving call quality. Specifically, the getMinBufferSize interface is called on the Android system to obtain the buffer size according to different sampling rates, and the sampling rate corresponding to the minimum value is selected to reduce delay. This method improves the real-time performance of calls to a certain extent, but it still has limitations in the following aspects.

[0005] First, this method mainly solves the sampling rate selection problem in a fixed Internet phone environment, and fails to take into account the wireless transmission characteristics between the Bluetooth headset and the mobile terminal and the impact of bandwidth fluctuations on audio data transmission. In a wireless transmission environment, bandwidth and channel conditions change frequently, and the selection of a fixed sampling rate may not be able to adapt to dynamically changing network conditions, resulting in the inability to effectively adjust the sound quality when the bandwidth decreases, affecting the call quality.

[0006] Secondly, this method lacks comprehensive consideration of audio data compression and dynamic adjustment. Although the delay is reduced by selecting the optimal sampling rate, in practical applications, audio data compression and transmission efficiency are equally important. Existing methods do not combine dynamic compression algorithms to adjust the audio compression ratio according to real-time bandwidth conditions to maximize transmission efficiency while ensuring sound quality.

[0007] In addition, this method does not involve multi-level cache management and data retransmission mechanism. During network transmission, data packets may be lost or transmission fails due to network congestion or signal interference. The existing technology does not provide an effective data cache and retransmission strategy, and cannot ensure the integrity and continuity of audio data. Especially in a network environment with high latency or high packet loss rate, the reliability and sound quality of call recording are still difficult to guarantee.

[0008] In summary, although the existing technologies have their own focuses in headset call recording and sampling rate adjustment, they still have many shortcomings. First, there is a lack of dynamic sampling rate adjustment mechanism for fluctuations in wireless transmission bandwidth, and it is impossible to automatically reduce the sampling rate to adapt to transmission requirements when the bandwidth is insufficient, resulting in recording interruption or deterioration of sound quality. Secondly, the audio processing technology is relatively simple, and it fails to make full use of multi-microphone arrays and advanced filtering algorithms to improve noise reduction and echo cancellation effects, making it difficult to maintain high-quality recording in complex noise environments. Thirdly, the compression and transmission efficiency of audio data have not been comprehensively optimized, and there is a lack of the ability to adjust compression parameters in real time, and it is impossible to flexibly control the amount of data under different network conditions, affecting transmission efficiency and sound quality. At the same time, the lack of multi-level cache management and data retransmission mechanism makes it easy for audio data to be lost or confused when network transmission fails or bandwidth fluctuates, affecting the continuity and reliability of recording. In addition, the existing technology also has problems in unified sampling rate and device compatibility, and cannot take into account the compatibility between different devices and the consistency of audio processing, further affecting call quality and user experience.

[0009] Therefore, how to monitor and adjust the sampling rate in real time during the headset call recording process, optimize the compression and transmission efficiency of audio data, combine multi-microphone array and adaptive filtering technology for efficient noise reduction and echo cancellation, and establish a comprehensive multi-level cache management and data retransmission mechanism has become a technical problem that needs to be solved urgently. Solving these problems can not only improve the quality and stability of headset call recording, but also significantly improve the overall user experience and meet the needs of modern users for efficient, convenient and high-quality communication tools. Summary of the invention

[0010] In view of the shortcomings of the prior art, the present invention provides a call recording method with dynamic sampling rate adjustment, which improves the quality and stability of Bluetooth headset call recording, and can also significantly improve the overall user experience, meeting the needs of modern users for efficient, convenient and high-quality communication tools.

[0011] To achieve the above object, the present invention provides the following technical solutions:

[0012] A call recording method with dynamic sampling rate adjustment,

[0013] Step 1, audio data collection: collect uplink audio data of the user's voice signal and downlink audio data of the other party's voice signal through the Bluetooth headset during the call, and pre-process the uplink audio data and the downlink audio data to reduce noise and eliminate echo;

[0014] Step 2, dynamic sampling rate adjustment: real-time monitoring of whether the transmission speed of the downlink audio data packet can reach the threshold, and comparing the monitoring result with the preset safety threshold to determine whether the current data transmission rate is lower than the safety threshold; when it is determined that the current data transmission rate is lower than the preset safety threshold, the sampling rate of the uplink audio data and the downlink audio data is automatically reduced from 16kHz to 8kHz to reduce the data volume and adapt to bandwidth limitations;

[0015] Step 3, audio data compression: compressing the uplink audio data and the downlink audio data after the sampling rate is adjusted using a coding algorithm;

[0016] Step 4, data caching: caching the uplink audio data and the downlink audio data through a double-layer caching mechanism, the first cache is used to temporarily store the uplink audio data and the downlink audio data that have been preprocessed and adjusted in sampling rate, and the second cache is used to temporarily store data packets to be sent and data packets that need to be retransmitted due to failed sending;

[0017] Step 5, audio data transmission: The compressed uplink audio data and downlink audio data are transmitted to the receiving terminal via Bluetooth for decoding and storage.

[0018] Furthermore, the uplink audio data and the downlink audio data of the step 1 are aligned by the following method: Step 2.1, uplink audio data acquisition: through the Bluetooth protocol, the microphone in the Bluetooth headset collects the user's voice signal to generate uplink audio data;

[0019] Step 2.2, downlink audio data collection: reversely collect the other party's audio signal through the built-in speaker of the Bluetooth headset to generate downlink audio data;

[0020] Step 2.3, audio data synchronization and alignment: obtaining current frame data of the uplink audio data and the downlink audio data in the memory, wherein the frame length of the current frame data is 240 bytes; compressing the frame data of the uplink audio data and the frame data of the downlink audio data respectively to obtain 40 bytes of compressed data; splicing the compressed data of the uplink audio data and the compressed data of the downlink audio data in a sequential order according to a predetermined format to form a single-frame audio of 80 bytes, and writing the single-frame audio into the audio cache queue in sequence;

[0021] Step 2.4, data cache management: the synchronized and aligned uplink audio data and downlink audio data are stored and managed respectively through a double-layer circular queue cache mechanism to ensure orderly and stable data transmission, wherein the first cache is used to store the audio data to be sent, and the second cache is used to store the command packet; the command packet is assembled by the sending thread according to the protocol requirements for the control instruction and the audio data, and the command packet is assembled and packaged by the sending thread according to the protocol requirements for the control instruction and the audio data, and is sent to the underlying serial port profile or general attribute profile service after being classified according to the command word. When the sending is successful, the sent command packet is removed from the second cache; when the sending fails due to the shortage of underlying resources, the command packet at the top of the cache is resent when the subsequent sending opportunity comes.

[0022] Furthermore, the uplink audio data and the downlink audio data in step 1 are subjected to echo cancellation and noise reduction processing by the following method:

[0023] Step 3.1, using multi-microphone arrangement and signal acquisition: at least two non-collinear miniature microphones are arranged on the Bluetooth headset body to form a microphone array, wherein each microphone synchronously collects ambient sound field signals, generates multi-channel time domain audio signals, and converts them into digital signals through an analog-to-digital converter;

[0024] Step 3.2, beamforming processing: beamforming is performed on the preprocessed multi-channel digital audio signal using minimum variance distortion-free response. According to the preset target sound source direction, each channel signal is weighted and delay compensated to achieve spatial filtering, enhance the speech signal in the target direction, and suppress interference noise from other directions.

[0025] Step 3.3, spatial filtering and noise suppression: further spatial filtering and noise suppression processing is performed on the beamformed signal, using one or more combinations of Wiener filtering, Kalman filtering, spectral subtraction or deep learning methods to calculate the power spectrum or covariance matrix of the background noise, and dynamically adjust the filter parameters according to the calculated noise characteristics to maximize the suppression of residual noise;

[0026] Step 3.3, echo cancellation: using the spatial information of the multi-microphone array and the loudspeaker output signal to establish an echo path model, using an adaptive filter based on the normalized least mean square or recursive least square algorithm to estimate and eliminate the echo signal, wherein the input of the adaptive filter includes the loudspeaker output signal and the microphone array signal, and the output is an echo estimation signal, and the echo cancellation is achieved by subtracting the echo estimation signal from the microphone signal;

[0027] Step 3.4, weighted processing of speech signals: perform post-processing on the signals after beamforming, spatial filtering and echo cancellation, and use multi-channel Wiener filtering or weighting methods based on auditory perception models to optimize the quality of the synthesized speech signal and improve speech clarity and intelligibility.

[0028] Furthermore, the method for dynamically adjusting the sampling rate in step 2 is as follows:

[0029] Step 4.1, data transmission speed monitoring step: This step is used to monitor the transmission speed of audio data in real time, so as to dynamically adjust the sampling rate according to the network status, thereby optimizing the transmission effect of audio data. The specific steps are as follows:

[0030] Step 4.1.1, data transmission rate measurement: calculate the amount of original audio data generated per second based on the audio sampling parameters; then refer to the actual audio transmission rate of the receiving terminal after compression processing, and obtain the current data transmission speed by counting the total amount of audio data received per second;

[0031] Step 4.1.2, network environment assessment: compare the calculated transmission rate with multiple rate ranges pre-defined by the system to assess the bandwidth status of the current network. If the current transmission rate is lower than the preset safety rate threshold, it is determined that the network is in a low bandwidth or high interference environment and needs to be adaptively adjusted.

[0032] Step 4.1.3, dynamic feedback mechanism: Based on the comparison result between the current transmission rate and the threshold, the system will send a feedback signal to the Bluetooth headset to control the headset to reduce the sampling rate and adjust the sampling rate strategy according to the change of the transmission rate.

[0033] Furthermore, after the transmission rate continues to be lower than the preset threshold and the sampling rate has been dynamically reduced, when the transmission rate returns to the normal range, the system will gradually restore the sampling rate according to the real-time network conditions to ensure that the audio quality is restored to the maximum extent. The specific operations of this step include: Step 4.2.1, rate recovery condition: after the sampling rate is reduced and the audio transmission speed is stable, the system determines that it can try to recover, and the system sends a pre-recovery instruction. After the Bluetooth headset receives the pre-recovery instruction, on the basis of ensuring normal audio transmission, the Bluetooth headset adds a speed verification instruction to send silent data;

[0034] Step 4.2.2, sampling rate recovery control: Once the judgment condition is met, the system will send a command to restore the sampling rate to the Bluetooth headset through a feedback signal, and the system will gradually restore the audio sampling rate to a higher sampling rate;

[0035] Step 4.2.3, sampling rate transition: After receiving the command to increase the sampling rate, the headset replies to the system with a successful command, increases the sampling rate, and simultaneously notifies the system of the new sampling rate. The system parses the audio according to the new sampling rate.

[0036] Further, the method of compression processing in step 3 to reduce the bandwidth required for transmission and improve data transmission efficiency is as follows: Step 5.1, preprocessing before data compression: Before audio data compression, the system preprocesses the uplink audio data and the downlink audio data, and this preprocessing step includes:

[0037] Step 5.1.1, echo cancellation: eliminate the echo component received by the microphone from the headphone speaker to reduce the influence of the speaker on the recording effect;

[0038] Step 5.1.2, noise reduction processing: Apply a noise reduction algorithm to eliminate the noise of the audio signal to reduce the impact of background noise on the compression effect and improve the quality of the compressed audio signal;

[0039] Step 5.1.3, audio signal enhancement: before compression, the system can enhance the audio signal according to the real-time environment to improve the audio quality;

[0040] Step 5.2, data framing and quantization during compression: Before audio data is compressed, it is divided into several frames, each frame contains an audio signal of a certain duration. In each frame, the audio signal is compressed through quantization and encoding algorithms. According to the selected encoding algorithm, the following operations are performed:

[0041] Step 5.2.1, frame processing: divide the adjusted audio data into time-continuous frames, each frame contains the characteristic data of the audio signal; frequency domain conversion and quantization: convert the audio signal to the frequency domain through methods such as fast Fourier transform or discrete cosine transform, and then perform quantization processing to reduce the amount of data in the high-frequency part, thereby achieving data compression;

[0042] Step 5.2.2, encoding and compression: Encode and compress the quantized audio data according to the selected encoding algorithm, and use a specific encoding table to map the audio signal to further reduce the amount of data;

[0043] Packaging and encapsulation of compressed data: After the audio data is compressed, the system packages the compressed audio data into a data packet format and encapsulates it, and performs packet processing according to the requirements of the Bluetooth communication protocol.

[0044] Furthermore, the echo cancellation method of step 5.1.1 includes:

[0045] Adaptive spatial filtering: The dual-microphone signals are subjected to adaptive spatial filtering and wind noise estimation, and the weighting method of the microphone array signal is adjusted based on the estimated voice direction information to form a directional beam and enhance the user's voice, while weakening the ambient noise and obtaining the wind noise distribution;

[0046] Adaptive conditioning filtering: Filters the received signal of the sensor located in the ear canal to remove external noise mixed in by the bone conduction path and different wearing states, and adaptively conditions the inner ear voice signal according to the corresponding active noise reduction state when the headset has active noise reduction function, so as to make full use of the advantages of the inner ear sensor in resisting wind noise and having a high signal-to-noise ratio, and correct possible timbre distortion;

[0047] Post-filtering: The known noise and echo information is sent to the post-filter for linear and nonlinear noise reduction processing, and the echo is further suppressed to improve the overall sound quality of the call recording.

[0048] Furthermore, the method of audio data transmission in step 4 is as follows:

[0049] Step 6.1, audio data packaging and private protocol encapsulation: Before transmission, the uplink audio data and the downlink audio data are packaged into data packets, each of which contains the following fields:

[0050] The start identifier of the data packet, used to mark the beginning of the frame;

[0051] The total length field of the data packet indicates the length of the entire data packet, including audio data and protocol control information;

[0052] A command identifier, indicating the type of data;

[0053] The payload part contains the compressed audio data;

[0054] The receiving end will only process the data packets that meet the requirements of the start identifier and command identifier fields of the data packet, and other data packets that do not meet the protocol will be ignored;

[0055] Step 6.2, data transmission process: After the data packet is encapsulated, the audio data is transmitted via the Bluetooth protocol and adapted according to the real-time network bandwidth:

[0056] Transmission of Bluetooth data packets: Data packets are transmitted on the Bluetooth connection channel and sent using the appropriate Bluetooth protocol; Dynamic bandwidth adjustment: According to the transmission rate and network conditions, the transmission rate is dynamically adjusted to adapt to the current bandwidth conditions;

[0057] Step 6.3, data verification and feedback mechanism: After receiving the audio data packet at the receiving end, the data packet is verified, including:

[0058] The start character of the received data packet is consistent with the predefined value. If it is inconsistent, the data packet is considered invalid;

[0059] Check whether the length of the data packet matches the received data. If it does not match, it is considered as packet loss or transmission error and the data packet will be discarded;

[0060] Whether the data packet is call recording data. If the value of this field does not match the predetermined command, the decoding process of the audio data will not be triggered;

[0061] If all the above checks are passed, the receiving end will continue to decode and process the audio data. If the checks fail, the receiving end will trigger the corresponding feedback mechanism according to the error type and request to resend the data packet.

[0062] Step 6.4, data decoding and audio reconstruction:

[0063] Once the receiving end passes the check and confirms the validity of the data packet, it decodes the audio data:

[0064] Audio decoding: According to the adopted encoding algorithm, the receiving end decodes the audio data and restores the compressed audio data to the original audio signal;

[0065] Step 6.5, audio data storage and subsequent processing: The decoded audio data will be stored by the receiving end and can be used later.

[0066] Furthermore, the uplink audio data and the downlink audio data are compressed using a G.722.2 encoding algorithm.

[0067] A call recording method using any of the above dynamic sampling rate adjustment methods for a Bluetooth headset, the Bluetooth headset comprising:

[0068] At least one microphone, used to collect the user's uplink audio signal;

[0069] At least one speaker, used to play the received downlink audio signal;

[0070] A Bluetooth communication module is used to transmit the collected audio data to a receiving terminal via the Bluetooth protocol;

[0071] An audio processing unit, used to perform dynamic sampling rate adjustment, audio data compression and enhancement processing, wherein the audio processing unit dynamically adjusts the sampling rate of the audio signal according to the real-time transmission rate, compresses the collected audio signal using a compression algorithm, and performs noise reduction, echo cancellation and enhancement processing on the audio signal;

[0072] The data cache module is used to cache the collected audio data and processed data packets to ensure stable data transmission.

[0073] Compared with the prior art, the present invention provides a call recording method with dynamic sampling rate adjustment, which has the following beneficial effects:

[0074] First, the present invention monitors the transmission rate of audio data packets in real time and dynamically adjusts the sampling rate. When the network bandwidth is insufficient, the sampling rate is reduced from 16kHz to 8kHz, thereby effectively reducing the amount of data and adapting to bandwidth limitations. This adjustment mechanism can ensure the continuity of call recording and the stability of sound quality when network conditions are poor, avoid recording interruptions or significant degradation of sound quality, and significantly improve the reliability and adaptability of recording.

[0075] Secondly, the present invention combines dynamic sampling rate adjustment with audio compression technology to encode and compress audio data according to real-time network conditions, effectively reducing the amount of transmitted data. When the transmission bandwidth is limited, the dual optimization strategy of compression and sampling rate adjustment can be used to improve transmission efficiency, while restoring a higher sampling rate and sound quality when network conditions improve, ensuring efficient transmission and better audio quality in different network environments.

[0076] Thirdly, the uplink and downlink audio data are efficiently processed for noise reduction and echo cancellation by using multi-microphone arrays, beamforming, adaptive filters, and post-filtering. Adaptive spatial filtering combined with wind noise estimation can enhance the user's voice signal and weaken environmental noise, while adaptive conditioning filtering uses the high signal-to-noise ratio characteristics of bone conduction signals to further improve the sound quality. These audio processing methods can still maintain the clarity and stability of recordings in complex noise environments, which is significantly better than existing technologies.

[0077] Secondly, the present invention effectively stores and manages audio data through a double-layer cache mechanism. The first cache is used to temporarily store audio data after adjusting the sampling rate, and the second cache is used to temporarily store data packets to be sent and data packets that have failed to be retransmitted. In the case of network congestion or signal interference, the cache mechanism can alleviate instantaneous transmission pressure and avoid data loss or confusion. At the same time, the integrity and continuity of the data are guaranteed through the retransmission strategy, further improving the stability of call recording.

[0078] The present invention combines the wireless transmission characteristics of Bluetooth headsets and designs a dynamic feedback and adjustment mechanism, which can maintain good audio transmission performance in low-bandwidth or high-interference environments. At the same time, the method supports a variety of encoding algorithms and Bluetooth protocols, can adapt to the communication needs of different devices, and ensure cross-device compatibility. In addition, when transmitting data packets, the present invention ensures the validity and accuracy of the data through a strict data verification mechanism, so that the overall quality of recording and user experience can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1:Flowchart of the call recording method with dynamic sampling rate adjustment in the present invention

[0080] Figure 2 : Functional block diagram of the dynamic sampling rate adjustment module in the present invention

[0081] Figure 3 : Flow chart of multi-level cache management and data transmission in the present invention DETAILED DESCRIPTION

[0082] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0083] This embodiment discloses a method for dynamically adjusting the sampling rate of a call recording for a Bluetooth headset, the purpose of which is to solve the problems of unstable recording quality, insufficient noise immunity, and low data transmission efficiency in the prior art under a wireless transmission environment. To this end, the present invention realizes efficient collection, processing, and transmission of uplink and downlink audio data under different network environments by dynamically adjusting the sampling rate, optimizing audio data compression, enhancing audio signal processing, and establishing a complete cache and retransmission mechanism, thereby significantly improving the overall quality of call recording and user experience.

[0084] Figure 1 The overall process of the method of the present invention is shown, and each step and the connection between them will be described in detail below.

[0085] First, in step S101, the system collects the user's uplink audio data (user voice signal) and downlink audio data (the other party's voice signal) through the Bluetooth headset at the same time, and pre-processes the collected data in the audio collection stage to ensure the sound quality. Specifically, the system first uses a noise reduction algorithm (such as Wiener filtering, Kalman filtering, etc.) to suppress background noise, thereby improving the clarity of the audio signal; at the same time, the system uses the spatial information of the multi-microphone array and combines the output signal of the speaker to achieve echo cancellation through an adaptive filter, effectively reducing the noise caused by the reflection of the headset speaker. In this process, the uplink audio data is collected by the built-in microphone of the Bluetooth headset, converted into a digital signal through an analog-to-digital converter (ADC), and the multi-microphone array technology is used to avoid the audio quality degradation caused by the position or angle of a single microphone; while the downlink audio data is collected by the built-in speaker of the Bluetooth headset in reverse, and the speaker transmits the audio and the microphone receives it. In a complex environment, the other party's voice signal can be accurately obtained, and the spatial filtering and beamforming processing of multiple microphone data are further reduced. Echo and interference.

[0086] In the system of this embodiment, a circular array consisting of four non-collinear MEMS microphones (model Knowles SPH0645, SNR65dB) is configured to synchronously collect the user's uplink audio data and downlink audio data. In the specific implementation:

[0087] Uplink data processing: Microphone array signals are transmitted via dual-channel DMA (source address 0x40012000, destination address 0x20001000, 240 bytes / frame cyclic mode)

[0088] The MVDR beamforming algorithm is adopted, the constraint direction angle is set to ±15° (pointing towards the user's mouth), and the diagonal loading factor is 0.1 times the covariance matrix trace. The speaker echo is eliminated through a 256-order NLMS adaptive filter (step factor μ = 0.02, regularization parameter δ = 1e 5).

[0089] Downlink data processing: The speaker signal is obtained through the reverse acquisition channel (DMA source address 0x40013000) and combined with the spatial sound field information for multi-channel Wiener filtering, dynamically adjusting the 615dB noise suppression amount. A noise classification model based on a deep neural network (input layer 128 nodes, 3-layer LSTM structure) is used to separate the ambient noise.

[0090] This solution was implemented in laboratory tests:

[0091] The signal-to-noise ratio is improved by 23.6dB (compared to the single-microphone solution)

[0092] Echo attenuation reaches 42dB (in compliance with ITU T G.168 standard)

[0093] The latency is controlled within 8.7ms (meeting the Bluetooth LE Audio synchronization requirements)

[0094] Next, in step S102, the system monitors the transmission rate of the downlink audio data packets in real time. Through comprehensive calculation of the sampling frequency, data volume and network conditions, the system can obtain the amount of audio data transmitted per second, and compare this value with the preset safety threshold to determine whether the current network conditions are suitable for high-quality audio transmission. When the monitored transmission rate is lower than the safety threshold, the system automatically enters step S103 and starts the dynamic sampling rate adjustment mechanism.

[0095] In this embodiment, the system monitors the downlink audio transmission rate in real time through a dynamic bandwidth evaluation model, and the specific implementation includes:

[0096] Safety threshold calculation:

[0097] R_safe=0.7×R_max×(1_Ploss)

[0098] R_max: Dynamically acquired based on the Bluetooth protocol version (2Mbps for BLE 5.0)

[0099] P_loss: 10-second average packet loss rate calculated using a sliding window (CRC16 checksum)

[0100] Monitoring Mechanism:

[0101] The transmission rate is sampled every 1000ms (accuracy ±3%)

[0102] When R_curr is detected three times in a row <R_safe时触发调整机制

[0103] Using the improved TCP Vegas algorithm to predict network congestion status

[0104] In step S103, in order to adapt to the bandwidth limitation, the system reduces the sampling rate of the upstream and downstream audio data from 16kHz to 8kHz to reduce the data volume and alleviate the network bandwidth burden; at the same time, the sampling rate adjustment is not a one-time operation, but a dynamic adjustment based on the real-time monitored network status. When the network environment improves, the original sampling rate can be gradually restored, thereby taking into account both transmission efficiency and sound quality requirements.

[0105] In this embodiment, the system adopts a two-level dynamic sampling rate adjustment strategy to achieve bandwidth adaptation:

[0106] Dynamic bandwidth assessment and mode switching strategy

[0107] Safe bandwidth threshold setting: Establish R_safe benchmark value (R_safe = 1.2 × average historical stable bandwidth), and dynamically update it every 5 minutes;

[0108] Emergency low bandwidth mode (8kHz / 16kbps):

[0109] Trigger condition: Activated when the current bandwidth R_curr < 0.5R_safe is detected for three consecutive sampling cycles;

[0110] Action: Enable FIR anti-aliasing filter (cutoff frequency 3.8kHz) and perform 2:1 downsampling, and switch to G.729AB codec protocol synchronously;

[0111] Normal Hi-Fi mode (16kHz / 32kbps):

[0112] Maintenance condition: When R_curr ≥ R_safe, the OPUS encoder is used to retain the full frequency band signal of 0-8kHz;

[0113] Network recovery verification and control mechanism

[0114] Stability test package design:

[0115] Generate a 3-second silent frame (filled with 0xFF), and append a CRC-16-CCITT checksum (polynomial 0x1021) to each frame.

[0116] Adopt redundant transmission strategy: send 3 test packets continuously (50ms interval), and require at least 2 packets to pass verification;

[0117] Progressive bandwidth recovery process:

[0118] Phase 1 (0-3 seconds): Maintain 8kHz sampling. If the test packet loss rate is ≤5% and the average delay is <80ms, start the transition mode.

[0119] Phase 2 (3-6 seconds): linearly increase the sampling rate to 16kHz (slope Δf = 2.67kHz / s), and monitor the jitter buffer status in real time; measured performance improvement: packet loss rate in subway environment dropped from 22.4% to 5.1%

[0120] The sampling rate switching delay is controlled within 18ms (the traditional solution requires 52ms)

[0121] CPU usage was optimized from 68% to 42% (memory usage was reduced by 29%)

[0122] Subsequently, in step S104, the audio data after the sampling rate adjustment is encoded and compressed to further reduce the bandwidth required for transmission and improve the data transmission efficiency. The compression process includes: firstly, the audio data is divided into a number of frames according to a certain time length, each frame contains the characteristic data of the audio signal; then the quantized data is compressed using a coding algorithm such as G.722.2, and after compression, the original data of 240 bytes per frame can be greatly reduced to about 40 bytes, so that the data volume is greatly reduced, which is particularly suitable for transmission requirements in low-bandwidth environments.

[0123] Perform enhanced G.722.2 encoding on the audio data after the sampling rate adjustment. The specific process includes:

[0124] Frame processing optimization: Split 240 bytes of raw data into three 80ms subframes (each subframe contains 26.6ms of audio)

[0125] Anti-packet loss design: insert 1 RS (40,32) redundant packet for every 5 data packets, which can correct up to 4 bytes of error

[0126] Key parameters (LPC coefficients, pitch period) are stored redundantly across frames.

[0127] Measured performance:

[0128] Compression latency reduced from 18.2ms to 9.7ms

[0129] The speech intelligibility MOS score remains at 4.1 (out of 5) at a 20% packet loss rate

[0130] Test environment description:

[0131] Network simulator: Uses a 20% packet loss rate + 50ms jitter channel model based on NS-3;

[0132] Voice samples: ITU-T P.501 standard test sentence library (50 sentences in Chinese and 50 sentences in English);

[0133] MOS Rating Method:

[0134] Using the ITU-T P.800 standard, 20 subjects blindly scored:

[0135] Comparison group: traditional SCB coding (non-redundant design); bandwidth saving rate reaches 83% (240B→40B / frame).

[0136] In step S105, to ensure the stability of data transmission, the system introduces a double-layer cache mechanism such as Figure 3 shown. Figure 3It shows the multi-level cache management and data transmission flow chart, and describes in detail the working logic of the double-layer cache mechanism, including the interaction process between the first cache (audio data temporary storage) and the second cache (command packet management). It also introduces how to ensure data integrity and stability through the retransmission mechanism when transmission fails.

[0137] First, the audio data is collected from the Bluetooth headset through the microphone, and then processed by noise reduction and multi-microphone (multi-mic) signal fusion to generate a clear audio signal. After the initial audio compression processing, these signals are temporarily stored in the first cache, that is, the audio data temporary storage area. At this stage, whether it is audio data from a single microphone (single mic) or multi-microphone audio data that has been noise-reduced and processed, it will first enter this cache. After the audio data is encoded and compressed, a compressed audio data group will be formed.

[0138] After the audio data is compressed, the system stores it in the first-level audio cache queue. At this stage, the key task of the system processing is to effectively manage and schedule the sending of audio data packets. These data packets will be controlled and scheduled by the second cache (command packet management area) according to the network bandwidth and the real-time status of the system. The role of the second cache is to manage the sending commands of audio data packets, ensure that the data packets can be transmitted according to priority and order, and avoid data loss or transmission confusion when the bandwidth fluctuates or the network is unstable.

[0139] When the system issues a transmission command, the second cache stores the audio data packet together with the corresponding command, ready for transmission via the Bluetooth protocol. If a data packet is lost or fails to be sent during the transmission process, the system triggers a retransmission mechanism. During this process, the system first detects whether the transmission is successful or not. If the transmission fails, the system will retry to send the data packet according to the retry mechanism. The specific retry logic includes: clearing the data packets that have not been successfully sent in the current cache, reorganizing the audio data, and adjusting the sampling rate or compression level according to the network status to ensure that the data is retransmitted smoothly.

[0140] When the system successfully sends a command and receives a response signal, the data in the second buffer will be cleared and marked as successfully sent. At this point, all audio data transmission processes are successfully completed, and the storage content in the buffer area is cleared and ready to receive the next batch of audio data.

[0141] If the transmission fails and multiple retries still fail, the system will trigger further diagnostic processes to evaluate the network status and decide whether more substantial network adaptation adjustments are needed (such as reducing the sampling rate, adjusting the compression algorithm, etc.) to ensure the quality of audio transmission. In this case, the interaction between the first cache and the second cache can effectively support the system to adjust the data stream based on real-time feedback and ensure the integrity and quality of the audio data as much as possible under different network environments.

[0142] In summary, the multi-level cache management and data transmission flow chart shows how the system can stably and efficiently transmit audio data in a complex network environment through a double-layer cache mechanism and retransmission control, ensuring that end users can experience high-quality audio effects when using Bluetooth headsets.

[0143] In this embodiment, the first cache:

[0144] 50 frame ring buffer (configurable depth)

[0145] Second cache:

[0146] Hash table storage (key value = timestamp << 16 | seq_num)

[0147] 300ms timeout elimination mechanism (with LRU algorithm)

[0148] Retransmission control logic:

[0149]

[0150] Performance improvements:

[0151] Memory usage is optimized from 256KB to 182KB

[0152] The retransmission success rate in the 1Mbps bandwidth fluctuation scenario increased from 71% to 93%

[0153] The end-to-end delay standard deviation is reduced from ±46ms to ±17ms

[0154] The solution is implemented through the dual-core architecture of the Nordic nRF5340 chip (Cortex-M33 core 1 handles cache management, core 2 performs encoding and decoding), achieving an effective transmission rate of 96.4% in outdoor multipath interference tests, significantly better than the 82.3% of traditional solutions.

[0155] In step S106, the compressed audio data is transmitted to the receiving terminal via the Bluetooth protocol. The specific transmission process includes: first, the compressed data packet is sent to the receiving device via the Bluetooth connection; second, the receiving end verifies the received data packet to ensure that the data is complete and correct, and requests to resend the data packet if the verification fails; finally, after the data packet passes the verification, the receiving end decodes it, restores the compressed audio data to the original audio signal, and stores it in the device for subsequent playback, playback or archiving.

[0156] It is worth noting that step S101 not only involves the collection of audio data, but also includes key links such as synchronization alignment, noise reduction and echo cancellation. In the process of data collection and synchronization alignment, the uplink audio data is collected by the built-in microphone, converted into a digital signal by an analog-to-digital converter, and the multi-microphone array technology is used to ensure that each microphone independently collects environmental sounds, thereby avoiding sound quality problems caused by the position or angle of a single microphone; while the downlink audio data is transmitted through the speaker and then collected by the microphone in reverse, ensuring that the other party's voice signal can still be obtained in a complex environment, and spatial filtering and beamforming are realized through the joint processing of multiple microphone data to reduce echoes and interference.

[0157] In order to achieve efficient data transmission and reduce the CPU burden, the present invention further introduces direct memory access (DMA) technology. Specifically, independent DMA channels are configured for uplink and downlink audio data respectively. During the configuration process, it is necessary to consider parameters such as the source address (i.e., the audio input device buffer address), the target address (the memory buffer address for storing audio data), the transmission length (240 bytes per frame in this solution), and the transmission mode (loop mode or automatic reload mode). When the audio input device is ready for new data, a DMA request signal is triggered. After responding, the DMA controller takes over the bus control and transfers the data directly from the input / output device buffer to the specified buffer of the memory, thereby achieving efficient transmission and synchronization of data frames.

[0158] After the DMA transfer is completed, the audio data is stored in the memory buffer. The system ensures the integrity of the data frame by extracting 240 bytes of data for each DMA transfer. Since DMA uses block transfer mode, each data block transmitted is complete data, so there is no need to worry about data loss or partial frame errors. To ensure the synchronization of audio data, the system aligns the uplink and downlink data according to the timestamp or frame index to ensure the correct timing of the audio stream. After the uplink data frame is transferred to the specified memory buffer via DMA, the system extracts 240 bytes of data and marks it as the current uplink frame; the downlink data frame is also transferred to another memory buffer via DMA, and then 240 bytes of data are extracted and marked as the current downlink frame. The two are aligned in time, providing a basis for subsequent audio compression and data splicing.

[0159] After obtaining the complete uplink and downlink audio frames, the system compresses each 240-byte frame. It is recommended to use compression algorithms such as G.722.2 to reduce the bandwidth requirements of audio data and reduce the size of data packets. After compression, each frame data becomes approximately 40 bytes, which significantly improves the transmission efficiency, especially in a wireless transmission environment with limited bandwidth. Then, the system splices the compressed uplink and downlink audio data in the order of uplink data first and downlink data last. Each compressed data block (40 bytes) is combined into an 80-byte single-frame audio data packet. This splicing method not only ensures the integrity of the audio data, but also makes the data format unified in subsequent transmission.

[0160] In terms of data caching and transmission management, the spliced ​​80-byte single-frame audio data is sent to the audio cache queue and managed through a double-layer circular queue cache mechanism. Specifically, the first cache is used to temporarily store audio data that has been preprocessed, sample rate adjusted, and compressed, and the data is queued here waiting to be sent; while the second cache is used to store data packets to be sent and failed data packets that need to be retransmitted due to transmission anomalies. Through this double-layer cache management, the system can cope with delays and packet loss problems in network transmission, ensure that data can be retransmitted when it is lost, and through the data retransmission mechanism, when a data packet transmission failure or loss is detected, these data packets are automatically marked as "pending retransmission" and resent at the next transmission, thereby ensuring data reliability and integrity.

[0161] In order to obtain high-quality recording in a complex environment, the present invention adopts a combination of multi-microphone array and advanced signal processing algorithm in noise reduction and echo cancellation. First, at least two non-collinear micro-microphones are arranged on the Bluetooth headset body to form a microphone array. Each microphone synchronously collects the environmental sound field signal and converts it into a digital signal through an analog-to-digital converter; then, the minimum variance distortion-free response (MVDR) beamforming algorithm is used to process the pre-processed multi-channel digital audio signal, and the signal of each channel is weighted and delayed according to the preset target sound source direction to achieve spatial filtering, thereby enhancing the speech signal in the target direction and suppressing interference noise from other directions. The signal after beamforming will be further spatially filtered and noise suppressed. The system can use one or more combinations based on Wiener filtering, Kalman filtering, spectral subtraction or deep learning methods to calculate the power spectrum or covariance matrix of the background noise, and dynamically adjust the filter parameters according to the calculation results to minimize the residual noise.

[0162] In terms of echo cancellation, the system uses the spatial information provided by the multi-microphone array and the speaker output signal to establish an echo path model, and uses an adaptive filter based on the normalized least mean square (NLMS) or recursive least squares (RLS) algorithm to estimate and eliminate the echo signal. The input of the adaptive filter includes the speaker output signal and the microphone array signal, and its output is the echo estimation signal. By subtracting the echo estimation signal from the microphone signal, efficient echo cancellation is achieved. Finally, the system post-processes the signal after beamforming, spatial filtering and echo cancellation, and uses multi-channel Wiener filtering or weighting methods based on auditory perception models to optimize the quality of the synthesized speech signal and further improve speech clarity and intelligibility.

[0163] In order to eliminate echo more effectively, the present invention also introduces adaptive spatial filtering, adaptive conditioning filtering and post-filtering technology. Specifically, adaptive spatial filtering uses dual microphone signals and wind noise estimation to adjust the weighting method of microphone array signals to form a directional beam, which can not only strengthen the user's voice signal, but also weaken environmental noise and wind noise; adaptive conditioning filtering filters the received signal of the sensor located in the ear canal, which not only removes the external noise introduced by the bone conduction path and different wearing states, but also when the headset has an active noise reduction function, it adaptively conditions the inner ear voice signal according to the corresponding active noise reduction state, making full use of the advantages of the inner ear sensor in wind noise resistance and high signal-to-noise ratio, while correcting possible timbre distortion; and post-filtering sends the known noise and echo information to the post-filter for linear and nonlinear noise reduction processing, and further suppresses the echo, thereby improving the overall sound quality of the call recording.

[0164] Figure 2 The functional block diagram of the dynamic sampling rate adjustment module is shown, describing the entire processing process from audio data reception to final audio signal reconstruction. The module includes several key components, namely, the audio data receiving unit D1 is composed of the Nordic nRF5340 Bluetooth chip (supporting BLE 5.2 / LE Audio), The ADS1298 multi-channel ADC (8-channel 24-bit / 125kSPS), SiTime SiT9367 clock module (±1 0ppm synchronization accuracy) and STM32H7 DMA controller work together to achieve high-precision audio acquisition and low-latency transmission. The rate monitoring unit D2 is composed of Microchip LAN9354 network coprocessor, Nordic CRYPTOCELL-312 hardware verification module and ISSIIS66WVS4M8ALLSRAM to achieve accurate bandwidth evaluation. The security threshold comparison unit D3 is composed of TI TLV3201 high-speed comparator (propagation delay <5ns), Lattice iCE40UP5K programmable logic chip and NXP 74HC595 status register to achieve dynamic threshold decision. The sampling rate adjustment control unit D4 is composed of Silicon Labs Si5341 programmable clock generator (0.001ppm jitter) and Nordic nRF5340 dual-core MCU to realize hierarchical control strategy. The sampling rate adjustment output unit D5 integrates the ADI ADSP-21489 digital filter (SHARC+ core, 450MHz) and the TI TS3A5017 audio switch matrix to achieve lossless sampling rate conversion. The data decoding unit D6 uses the Nordic nRF5340 built-in hardware decoder (supporting G.722.2 acceleration) and the ISSI dual ring buffer (each with a capacity of 60 frames) to achieve efficient data recovery. These units work together to achieve dynamic sampling rate adjustment, optimize the transmission effect of audio data, and balance the relationship between sound quality and bandwidth.

[0165] The audio data receiving unit D1 is first responsible for receiving the audio data transmitted from the Bluetooth headset. Through this unit, the system decapsulates and synchronizes the data stream to provide a stable audio signal for subsequent analysis and processing. The received data will be transmitted to the rate monitoring unit D2, which monitors the transmission rate of the audio data in real time. The rate monitoring unit D2 calculates the amount of original audio data generated per second based on the audio sampling parameters, and combines the amount of audio data packets received at the receiving end to calculate the current transmission rate by counting the amount of audio data received per second. This information is transmitted to the security threshold comparison unit D3, which compares the detected transmission rate with multiple preset rate intervals to determine the bandwidth status of the network. If the current transmission rate is lower than the preset security rate threshold, the security threshold comparison unit D3 will send a warning to the sampling rate adjustment control unit D4, indicating that the network is in a low bandwidth or high interference state and needs to be adaptively adjusted.

[0166] When the system determines that the transmission rate is lower than the safety threshold, the sampling rate adjustment control unit D4 will automatically instruct the sampling rate adjustment output unit D5 to reduce the sampling rate based on this result to reduce the pressure on the network bandwidth. For example, the sampling rate can be reduced from 16kHz to 8kHz, thereby reducing the amount of data. In order to adapt to different network conditions, the sampling rate adjustment control unit D4 can set multiple rate thresholds, each threshold corresponding to a different sampling rate, thereby achieving more refined dynamic adjustment.

[0167] Once the network transmission rate returns to the normal range, the sampling rate adjustment control unit D4 will gradually restore the sampling rate. The system continuously monitors the transmission rate through the rate monitoring unit D2 to ensure that it stabilizes and reaches a normal level. When the system confirms that the recovery conditions have been met, it will send a pre-recovery instruction to the Bluetooth headset through the sampling rate adjustment control unit D4. After receiving the instruction, the Bluetooth headset will send a speed verification instruction and confirm the network stability by transmitting silent data. Once the system confirms that the network status meets the conditions, the sampling rate adjustment control unit D4 will send an instruction to restore the sampling rate to the Bluetooth headset through a feedback signal, gradually increasing the sampling rate to a higher level to ensure a smooth recovery process.

[0168] The compression of audio data is completed by the data decoding unit D6. The audio data after the sampling rate adjustment will first go through pre-processing steps such as echo cancellation, noise reduction and audio signal enhancement. These processes provide high-quality input signals for compression by eliminating echoes, reducing noise and enhancing signals. After pre-processing, the audio data will be divided into multiple frames, each of which contains an audio signal of a certain duration. Then, the system uses technologies such as fast Fourier transform FFT or discrete cosine transform DCT to convert the time domain signal to the frequency domain, and quantize and encode and compress it. This process not only reduces the amount of data, but also retains the core information of the voice to the maximum extent. After encoding and compression, the audio data is encapsulated into a standard data packet and is ready to be transmitted through the Bluetooth protocol. During the transmission process, the system dynamically adjusts the transmission rate according to the network bandwidth monitored in real time. When the network bandwidth is sufficient, the system increases the transmission rate to achieve higher sound quality; when the bandwidth is insufficient, the transmission rate is reduced to ensure the stability of data transmission.

[0169] At the receiving end, the data decoding unit D6 will perform strict checks on the received data packets to ensure the integrity of the data. First, a start character check is performed to ensure that the start character of the data packet is consistent with the predefined value; second, a length check is performed to confirm whether the length of the data packet meets the expectation; finally, it checks whether the command identifier points to the call recording data. If the check is successful, the data packet is decoded, the original audio signal is restored, and it is used for subsequent playback or storage operations.

[0170] Through the organic combination of the above units, the system achieves efficient transmission of audio data and high-quality audio reconstruction in a low-bandwidth environment. Each unit plays a role at different stages to ensure the transmission stability of audio data and the high-quality restoration of sound quality, and through dynamic adaptation of transmission rate and sampling rate, it achieves maximum bandwidth utilization and transmission efficiency.

[0171] This embodiment also provides a Bluetooth headset, which applies the above-mentioned call recording method with dynamic sampling rate adjustment. The Bluetooth headset includes:

[0172] At least one microphone: used to collect the user's uplink audio signal. The present invention recommends using multiple microphones to form a microphone array to achieve better noise reduction and echo cancellation effects.

[0173] At least one loudspeaker: used to play received downlink audio signals.

[0174] Bluetooth communication module: used to transmit the collected audio data to the receiving terminal via the Bluetooth protocol. This module needs to support the corresponding Bluetooth protocol and audio transmission protocol.

[0175] Audio processing unit: used to perform dynamic sampling rate adjustment, audio data compression and enhancement processing. This unit dynamically adjusts the sampling rate of the audio signal according to the real-time transmission rate, compresses the collected audio signal using a compression algorithm, and performs noise reduction, echo cancellation and enhancement processing on the audio signal. This unit is the core component for implementing the present invention.

[0176] Data Cache Module: used to cache the collected audio data and processed data packets to ensure stable data transmission. This module implements a double-layer cache mechanism, including a first cache for temporarily storing audio data and a second cache for temporarily storing command packets.

[0177] This Bluetooth headset achieves high-quality and high-stability call recording function through the coordinated work of various modules.

[0178] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

Claims

1. A call recording method with dynamic sampling rate adjustment, characterized in that include: Step 1, audio data collection: collect uplink audio data of the user's voice signal and downlink audio data of the other party's voice signal through the Bluetooth headset during the call, and pre-process the uplink audio data and the downlink audio data to reduce noise and eliminate echo; Step 2, dynamic sampling rate adjustment: real-time monitoring of whether the transmission speed of the downlink audio data packet can reach the threshold, and comparing the monitoring result with the preset safety threshold to determine whether the current data transmission rate is lower than the safety threshold; when it is determined that the current data transmission rate is lower than the preset safety threshold, the sampling rate of the uplink audio data and the downlink audio data is automatically reduced from 16kHz to 8kHz to reduce the data volume and adapt to bandwidth limitations; Step 3, audio data compression: compressing the uplink audio data and the downlink audio data after the sampling rate is adjusted using a coding algorithm; Step 4, data caching: caching the uplink audio data and the downlink audio data through a double-layer caching mechanism, the first cache is used to temporarily store the uplink audio data and the downlink audio data that have been preprocessed and adjusted in sampling rate, and the second cache is used to temporarily store data packets to be sent and data packets that need to be retransmitted due to failed sending; Step 5, audio data transmission: The compressed uplink audio data and downlink audio data are transmitted to the receiving terminal via Bluetooth for decoding and storage.

2. The call recording method according to claim 1, characterized in that: The uplink audio data and the downlink audio data in step 1 are aligned by the following method: The microphone of the Bluetooth headset collects the user's voice to generate uplink audio data, and the speaker collects the other party's voice to generate downlink audio data; Compressing 240-byte current frames of uplink audio data and downlink audio data into 40-byte data packets respectively; splicing the compressed data packets of the uplink audio data and the compressed data packets of the downlink audio data into 80-byte single-frame audio in a predetermined format and storing them in a cache queue; The synchronized and aligned uplink audio data and downlink audio data are stored and managed respectively through a double-layer circular queue buffer mechanism, wherein the first buffer is used to store the audio data to be sent, and the second buffer is used to store the command packet; The command packet is assembled by the sending thread according to the protocol requirements, the control instructions and the audio data are assembled and packaged by the sending thread according to the protocol requirements, and after being classified according to the command words, it is sent to the underlying serial port configuration file or the general attribute configuration file service. When the sending is successful, the sent command packet is removed from the second cache. When the sending fails, the command packet is retained and resent later.

3. The method for adjusting call recording by dynamic sampling rate according to claim 1, characterized in that The echo cancellation and noise reduction processing of the uplink audio data and the downlink audio data in step 1 includes: Two or more non-collinear micro-microphones are arranged in the Bluetooth headset to form an array, and the ambient sound field signal is synchronously collected to generate a multi-channel time-domain audio signal, and a digital signal is obtained through analog-to-digital conversion; The minimum variance distortion-free response algorithm is used to weight and delay-compensate multi-channel digital signals, and the speech signal in the target sound source direction is enhanced and interference noise is suppressed through spatial filtering. Calculate the background noise power spectrum and covariance matrix based on one or more combinations of Wiener filtering, Kalman filtering, spectral subtraction or deep learning methods, and dynamically adjust the filter parameters to suppress residual noise; Combining the loudspeaker output signal and the spatial information of the microphone array, an adaptive filter using the normalized least mean square or recursive least squares algorithm is used to establish an echo path model, and echo cancellation is achieved by subtracting the echo estimation signal from the microphone signal. The processed signal is subjected to multi-channel Wiener filtering or auditory perception model weighted processing to optimize the clarity and intelligibility of the synthesized speech.

4. The method for adjusting call recording by dynamic sampling rate according to claim 1, characterized in that The implementation method of dynamically adjusting the sampling rate in step 2 includes: Real-time measurement of audio data transmission rate, by calculating the amount of original audio data generated and counting the amount of compressed data at the receiving end, to obtain the current actual transmission rate; The measured transmission rate is compared with the preset rate threshold interval, and when it is lower than the safety rate threshold, it is determined to be a low-bandwidth network environment; According to the network status determination result, control instructions are sent to the Bluetooth headset to dynamically reduce the audio sampling rate and establish a linkage adjustment mechanism between the transmission rate and the sampling rate.

5. The call recording method according to claim 4, characterized in that When the transmission rate returns to the normal range, the sampling rate recovery control is performed, including: After the transmission rate reaches the preset safety threshold, a verification command is sent to the Bluetooth headset and silent data is transmitted; Control the Bluetooth headset to increase the sampling rate step by step through the feedback signal; After receiving the headphone sampling rate increase confirmation command, the audio data is parsed according to the new sampling rate.

6. The call recording method according to claim 1, characterized in that: The specific implementation method of the compression process in step 3 is: Eliminate the headphone speaker echo component received by the microphone; Use noise reduction algorithm to eliminate background noise; Enhance the audio signal based on the real-time environment.

7. The call recording method according to claim 5, characterized in that: The echo cancellation method of step 5.1.1 comprises: Adaptive spatial filtering and wind noise estimation are performed on dual microphone signals. The microphone array weighting method is adjusted based on the voice direction information to form a directional beam, enhance the user's voice and suppress environmental noise, while obtaining the wind noise distribution. Filter the sensor signal in the ear canal to remove external noise introduced by the bone conduction path and wearing status, adjust the inner ear voice signal according to the active noise reduction status, take advantage of its wind noise resistance and high signal-to-noise ratio characteristics and correct the timbre distortion; The noise and echo information is input into the post-filter for linear and nonlinear noise reduction processing, further suppressing the echo to improve the call recording quality.

8. The call recording method according to claim 1, characterized in that Step 4 includes: The uplink / downlink audio data is encapsulated to form a data packet containing the following fields: The start identifier, total length field and command identifier; Transmit encapsulated data packets through the Bluetooth connection channel and dynamically adjust the transmission rate according to the real-time network bandwidth; The receiving end verifies the consistency of the start identifier of the data packet, the matching of the data packet length, and the legitimacy of the command identifier. If the verification fails, the feedback mechanism is triggered and a retransmission is requested. After verification, audio decoding is performed to restore the original audio signal; The decoded audio data is stored in the receiving end storage unit.

9. A call recording method with dynamic sampling rate adjustment according to claims 1-8, characterized in that: The uplink audio data and the downlink audio data are compressed using the G.722.2 encoding algorithm.

10. A Bluetooth headset, using any one of the dynamic sampling rate adjustment call recording methods of claims 1-8, characterized in that: This Bluetooth headset includes: At least one microphone, used to collect the user's uplink audio signal; At least one speaker, used to play the received downlink audio signal; A Bluetooth communication module is used to transmit the collected audio data to a receiving terminal via the Bluetooth protocol; An audio processing unit, used to perform dynamic sampling rate adjustment, audio data compression and enhancement processing, wherein the audio processing unit dynamically adjusts the sampling rate of the audio signal according to the real-time transmission rate, compresses the collected audio signal using a compression algorithm, and performs noise reduction, echo cancellation and enhancement processing on the audio signal; The data cache module is used to cache the collected audio data and processed data packets to ensure stable data transmission.

Citation Information

Patent Citations

  • Determination method, apparatus and system for recording and playing sampling rate

    CN103354588A

  • Earphone and earphone sound recording method

    CN107580102A

Cited By

  • Audio data processing method and device, electronic equipment and program product

    CN119229883A

  • Hearing aid dual-channel audio definition directivity enhancing method and system

    CN121013035A

  • Network call method and device, computer equipment, readable storage medium and program product

    CN121771173A