Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

268 results about "Latency (audio)" patented technology

Latency refers to a short period of delay (usually measured in milliseconds) between when an audio signal enters a system and when it emerges. Potential contributors to latency in an audio system include analog-to-digital conversion, buffering, digital signal processing, transmission time, digital-to-analog conversion and the speed of sound in the transmission medium.

Security guarantee system for cross-domain communication and data sharing of network platform software

The invention relates to the technical field of computer network security, in particular to a security guarantee system for cross-domain communication and data sharing of network platform software, which comprises a data security transmission module, an identity authentication and access control module, an encryption module, a security audit and anomaly detection module and an anti-interference disaster recovery module. According to the invention, through linkage of equipment fingerprint end-to-end encryption, threat adaptive dynamic key updating and a timestamp-random number anti-replay mechanism and AI risk pre-judgment, reinforcement learning path planning, digital twinning pre-verification and fault automatic switching technologies, collaborative guarantee of key security and transmission continuity in cross-domain audio transmission is realized; encryption intensity is dynamically upgraded along with threats, a transmission path is optimized in advance, and it is ensured that the audio stream resists various network attacks and interferences in low-delay transmission.
Owner:ICLOUDSHIELD SECURITY TECHNOLOGY CO LTD

Low-delay adaptive audio transmission method, system and device and storage medium

The invention discloses a low-delay self-adaptive audio transmission method, system and device and a storage medium, and relates to the technical field of audio, and the method comprises the steps: dynamically recognizing an optimal transmission path in wireless dual bands based on an intelligent band detection and self-adaptive channel selection mechanism; decomposing the target audio data through a hierarchical coding algorithm, and executing self-adaptive adjustment of a differential transmission strategy and a quantization parameter; the dual-band performance is monitored in real time based on a load balancing algorithm, the decomposed transmission tasks are dynamically allocated, and path switching is triggered when it is monitored that the transmission quality of the current band is abnormal; the future buffer requirement is predicted based on a predictive buffer algorithm, the size of a buffer area is dynamically adjusted according to the actual playing condition and the network state change so as to ensure the playing continuity and minimize the delay, the technical bottleneck of traditional single-frequency wireless audio transmission is broken through, an effective frequency band coordination mechanism is realized, resource waste or transmission conflicts are avoided, and the transmission efficiency is improved. And the double-frequency advantage is fully exerted.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Low-delay streaming voice interaction system with interruption processing function

The invention provides a low-delay streaming voice interaction system with an interruption processing function, and relates to the technical field of artificial intelligence. Necessary preprocessing and acoustic feature extraction are performed through a real-time acoustic processing module, and distortion and complex environmental noise introduced by an interaction channel are resisted through a robustness enhancement technology; the streaming acoustic decoding module performs acoustic modeling, language model application and decoding in real time in parallel, and outputs an ultra-low-delay text transfer result stream; the real-time acoustic processing module is combined with a signal processing technology and is responsible for detecting user voice activity in a high-precision and ultra-low-delay manner, and particularly judging the real-time voice activity state of a user through the user voice activity in an AI voice playing period; an efficient and low-delay bidirectional streaming network transmission mode is adopted between modules of the system and between the modules and a communication platform, and it is ensured that audio streams, acoustic feature streams, text streams and control signals can be transmitted and processed in real time with extremely low end-to-end delay.
Owner:GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD

Network optimization method based on online conference

The invention discloses a network optimization method based on an online conference, which relates to the technical field of real-time audio and video transmission, and comprises the following steps: in a transnational network, deploying monitoring probes at nodes participating in the online conference for monitoring network parameters such as network delay, bandwidth, packet loss rate and jitter in real time; the method comprises the following steps: processing network parameters and predicting abnormity by using a statistical analysis algorithm and a time sequence model to obtain network condition information, and then combining global network topology and using a reinforcement learning Q-Learning model and a graph neural network GNN topology prediction algorithm; according to the method, the network condition is monitored and predicted in real time, the optimal transmission path is output by using the reinforcement learning Q-Learning model and the GNN topology prediction algorithm, the bandwidth fluctuation is predicted through the LSTM, the packet loss rate is predicted based on the Bayesian network, the data transmission strategy is optimized, the network delay, the packet loss rate and the jitter are effectively reduced, and the network performance is improved. And the smoothness and the stability of the conference are improved.
Owner:申岳军

Low-delay multi-room audio playing optimization method and system, storage medium and equipment

The invention relates to the technical field of audio playing optimization, and discloses a low-delay multi-room audio playing optimization method and system, a storage medium and equipment, and the method comprises the steps: building a global time synchronization system based on a precision clock protocol in a playing equipment cluster, and enabling all playing equipment to operate based on a unified time reference; monitoring network state parameters in real time, dynamically calculating and adjusting the audio buffer depth of each playing device based on the monitored network quality, and obtaining audio data in advance according to the playing progress and the network state; generating a global playing timestamp according to the global time synchronization system, and controlling all playing devices to perform synchronous starting and progress control of audio playing based on the global playing timestamp; a transmission strategy of audio data is dynamically optimized based on real-time network state parameters, packet loss recovery and bandwidth adaptation are achieved, sub-millisecond multi-room audio synchronous playing is achieved, and meanwhile performance parameters can be automatically optimized according to different network environments and hardware configurations.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Autonomous controllable architecture software and hardware collaborative audio and video signal transmission method and system

The invention relates to the technical field of audio and video coding, decoding and processing, provides an audio and video signal transmission method and system based on software and hardware cooperation of an autonomous controllable architecture, and solves the problem of low bandwidth utilization rate. The method comprises the following steps: synchronously acquiring bandwidth fluctuation, packet loss rate and transmission delay parameters of a mobile terminal link, and generating a network state level through association analysis; performing resolution scaling on a brightness component of a video frame, performing dynamic downsampling on a chrominance component, generating a resolution reduction frame adaptive to a network state, and generating a fault-tolerant coding stream by combining a packet loss rate regulation and control quantization step size, a frame group length and an inter-frame prediction mode; performing entropy decoding and motion compensation to align brightness and chrominance components, forming a decoding reconstruction frame, performing multi-scale feature extraction and detail compensation, generating an enhanced brightness graph, and performing independent reconstruction on chrominance decoding components; and outputting an original resolution video frame through color reconstruction fusion. According to the invention, high-definition video low-delay transmission under network fluctuation is realized, and both bandwidth adaptability and image quality restoration precision are considered.
Owner:CHINA SOUTHERN TECHNOLOGY (GUANGDONG HENGQIN) CO LTD +1

Adaptive frequency optimization system in radio and television microwave transmission

The invention relates to a self-adaptive frequency optimization system in radio and television microwave transmission, and the system comprises an initial audio collection and processing module which collects an audio signal of a current frequency channel in real time, and generates a first data set; the conversion audio acquisition processing module synchronously acquires a target channel signal according to the frequency switching signal, and calculates the loudness offset between two channels; a feature extraction module extracts frequency domain energy distribution feature vectors of the first data set and the second data set; the model construction module constructs a dynamic gain compensation model and generates a compensation value of a target channel; the tone quality compensation module performs smoothing processing on the compensation value to generate an optimized third data set; and the audio optimization module updates model parameters, dynamically adjusts a frequency domain analysis strategy, and outputs low-delay and high-stability radio and television transmission signals. According to the invention, the problems of high switching delay and degraded tone quality caused by static configuration in a traditional scheme are solved, and the transmission efficiency and user experience in a complex electromagnetic environment are remarkably improved through modular dynamic calibration and compensation.
Owner:杨建梅

Wireless audio frequency band conflict avoidance method in high-density and high-strength electromagnetic environment

The invention discloses a wireless audio frequency band conflict avoidance method in a high-density and high-intensity electromagnetic environment, and the method comprises the steps: monitoring the signal intensity, the frequency band occupation condition and an interference source in the electromagnetic environment in real time, and carrying out the measurement of the signal intensity through a spectrum analyzer or a radio frequency detector; intelligent prediction and dynamic sensing are performed on the spectrum occupation condition by combining a machine learning algorithm, so that the frequency band of audio signal transmission is adjusted in real time, and the current idle frequency band with less interference is dynamically selected for audio transmission based on a spectrum sensing result, so that the conflict with the frequency bands of other equipment is avoided. Through the intelligent spectrum sensing technology, the spectrum use condition in the electromagnetic environment is monitored in real time, frequency band conflicts and interference sources can be efficiently and accurately found, smooth audio data transmission is ensured, transmission delay caused by the frequency band conflicts can be effectively reduced through dynamic frequency band selection and a rapid switching mechanism, and the low delay requirement of real-time audio transmission is ensured.
Owner:同辉佳视(北京)信息技术股份有限公司

Multi-room multi-channel audio synchronization method, apparatus, and device and storage medium

A multi-room multi-channel audio synchronization method, apparatus, and device and a storage medium are provided. In this method, an audio signal to be played in a primary audio device is determined, the audio signal is encoded according to channels by using a preset low-latency audio encoding strategy and audio encoding data is obtained. Next, the audio encoding data is optimized based on a User Datagram Protocol and a Negative Acknowledgment NACK retransmission mechanism to obtain audio coded sequence data packets. The audio encoding sequence data packets are controlled to be synchronously played in at least one secondary audio device by using a preset room synchronization control strategy and a channel synchronization control strategy.
Owner:LINKPLAY TECHNOLOGY INC

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Solution for TTS (Tone To Send) in high-concurrency scene by clone timbre

The invention belongs to the technical field of text-to-speech conversion, and particularly relates to a solution for TTS (text-to-speech) in a high-concurrency scene through tone cloning. According to the method, the text can be generated and returned to the client by segmenting and parallelizing the text and reducing the delay of the first segment, compared with a traditional serial whole segment synthesis mode, the waiting time of a user is greatly shortened, the audio is returned after being generated, the waiting time of the user can be effectively shortened under high concurrency, the network bandwidth and the delay pressure are reduced, and the user experience is improved. The real-time performance is guaranteed, reconnection disasters are avoided, underlying model services can be fully utilized, the influence of the serial processing characteristic of a CosyVoice2 model is avoided, a client can immediately start to play a first segment of audio, even if subsequent segments are still synthesized, the continuous playing experience of previous content is not influenced, and when the concurrency is increased, the continuous playing experience of the previous content is not influenced. The model gateway can dynamically disperse requests to a plurality of GPU instances, and long queue waiting after a single-card video memory is fully occupied is avoided.
Owner:BEIJING HUAYUN WORLD TECH CO LTD

Display device and screen projection display method

The embodiment of the invention discloses a display device and a screen projection display method, and the method comprises the steps: after receiving a screen projection connection instruction sent by a terminal device, sending a starting notice to a player through a screen projection service, and enabling the player to build a data pipeline based on the starting notice. And controlling the player to receive the code stream data sent by the terminal equipment after the data pipeline is established, and processing the code stream data by using the data pipeline when the player receives the code stream data to obtain a video picture and an audio. The display displays a video picture, and the audio output device plays audio. According to the embodiment of the invention, data transmission is directly carried out through the terminal equipment and the player, so that communication delay caused by inter-process communication is reduced, the received code stream data is directly processed after the player establishes the data pipeline, cache is reduced, and the delay time of screen projection display is shortened.
Owner:VIDAA (NETHERLANDS) INT HLDG LTD

Cloud desktop audio and video synchronization method and system, electronic equipment and storage medium

The invention relates to a cloud desktop audio and video synchronization method and system, electronic equipment and a storage medium. The method comprises the following steps: evaluating network quality in real time according to network state parameters to obtain a parameter evaluation result; weighting calculation is carried out on a parameter evaluation result, and grading is carried out on the network quality; for a good network, the timestamps of the audio and the video are aligned, and the audio and the video do not need to be buffered and compensated; for a mild weak network, dynamic buffer compensation, redundancy rate adjustment and audio priority processing are carried out on audio and video; and for a severe weak network, performing prediction compensation, selective frame loss and code rate adaptive processing on audio and video timestamps. According to a plurality of network state parameters such as network delay, jitter, packet loss rate and bandwidth, the network quality is evaluated in real time, and then the quality and transmission strategy of the audio and video data are automatically adjusted according to the level of the network quality, so that the audio and the video can still be efficiently and synchronously transmitted even if the network condition is poor, and the transmission efficiency is improved. And user experience is improved.
Owner:XIAN WANXIANG ELECTRONICS TECH CO LTD

Model compression and data enhancement fused lightweight deep counterfeit voice detection method

The invention provides a model compression and data enhancement fused lightweight deep forged voice detection method. The method comprises the steps of obtaining and processing a public voice data set and a large-scale self-supervision pre-training voice model; performing structured pruning and knowledge distillation to obtain a lightweight voice model; performing audio preprocessing and diversified data enhancement on the true and false voice samples to obtain an enhanced true and false voice data set; performing faking task joint fine tuning on the lightweight voice model to obtain a lightweight deep faking voice detection model; and locally deploying the model to obtain a localized counterfeit voice detection system, and carrying out real-time authenticity judgment. According to the method, calculation complexity and reasoning time delay are remarkably reduced, local deployment is carried out on a resource-limited end side platform, detection generalization and robustness are improved, and low-delay, low-power-consumption and high-robustness detection performance is achieved.
Owner:ZHEJIANG UNIV

Wireless microphone intelligent control system and method

The invention relates to the technical field of audio processing, and discloses an intelligent control system and method for a wireless microphone, and the method comprises the steps: collecting multi-point acoustic data through a distributed microphone network, and constructing a spatial acoustic topology model; based on a topological analysis result, constructing an adaptive notch filter bank to realize howling suppression; environment changes are monitored in real time through a distributed acoustic sensor network, and suppression parameters are dynamically optimized; integrating a lightweight convolutional neural network on microphone equipment to identify acoustic events; low-delay acoustic event processing and howling prevention are realized by adopting a layered edge computing architecture; according to the invention, the technical problems of unstable howling suppression, weak environmental adaptability and lack of spatial perception in a complex environment are effectively solved.
Owner:XIAMEN AKS ELECTRONICS CO LTD

System for latency-aware orchestration and performance optimization in artificial intelligence telephone communication

A system for latency-aware orchestration and performance optimization in AI-driven telephone communication, consisting of: a speech capture unit configured to capture an analog audio signal from a telephone interface and convert the analog audio signal into a digital audio signal stream; a feature extraction unit that is operationally coupled with the speech acquisition unit and is configured to generate a feature representation of the digital audio signal stream through spectral decomposition, noise reduction, and temporal segmentation; an AI inference processor communicatively connected to the feature extraction unit, configured to run one or more AI models for automatic speech recognition, natural language understanding, and emotion recognition on the feature representation to generate intermediate results for inference; a latency orchestration controller coupled to the AI ​​inference processor, wherein the latency orchestration controller is configured to monitor latency across multiple processing stages, predict cumulative delay propagation using a hybrid latency estimation model, and orchestrate the execution scheduling of the AI ​​inference processor based on the predicted latency deviation; a performance optimization unit coupled with the latency orchestration controller and configured to dynamically adjust computational accuracy, inference batch size, and feature processing resolution based on latency thresholds and quality constraints set by the latency orchestration controller; and a transmission synchronization array configured to time-align the processed output generated by the AI ​​inference processor and transmit it to a remote communication node, with the transmission synchronization array maintaining deterministic time coordination between successive packets and the orchestrated inference results.
Owner:CHEEKURI KARTHIK CHAKRAVARTHY DULUTH

Audio processing method, device and equipment for calling for help for old people and medium

The invention relates to the technical field of elderly care, solves the problems that in the prior art, due to the fact that noisy and intermittent elderly voice is lack of robust sentence bound recognition and interruption repair, misinformation and missing information of calling-for-help triggering are prone to occurring, and time delay is high, and provides an audio processing method and device for elderly calling-for-help, equipment and a medium. The method comprises the following steps: preprocessing a voice signal of the elderly to obtain a preprocessed voice signal; performing statement boundary recognition and labeling on the preprocessed voice signal to obtain a voice fragment set; carrying out interruption segment and silent interval detection on the voice segment set to obtain an initial voice segment sequence, and carrying out intelligent recombination and time sequence repair on the initial voice segment sequence to obtain a target voice segment sequence; and carrying out emergency call recognition on the target voice segment sequence to obtain an emergency call judgment result. According to the invention, the accuracy of distress call identification of the elderly is improved.
Owner:NINGBO SIMSHINE INTELLIGENT TECH CO LTD

Low-delay intelligent communication system based on end-to-end voice large model

The invention relates to the field of intelligent communication, and specifically discloses a low-delay intelligent communication system based on an end-to-end voice large model, and specifically, a sending end integrated network predictor analyzes network historical indexes in real time and predicts a future network trend. Based on the prediction results, the system dynamically generates a control signal including the target coding mode and a switching instruction. The input audio frame quantization module intelligently selects and switches a proper audio frame encoder model to quantize the audio according to the control signal, and generates a discrete token. These tokens are encapsulated and transmitted together with the control signal. And the receiving end analyzes the control signal and guides the decoding module to perform corresponding decoding. In this way, the system can pre-judge network changes, high quality is kept when the network condition is good, end-to-end delay is effectively reduced by rapidly switching to the low-delay mode when the network deteriorates, and unification of low delay, high quality and intelligence is achieved.
Owner:HANGZHOU XINGYU ARTIFICIAL INTELLIGENCE CO LTD

Wireless audio transmission protocol optimization algorithm for low-delay dynamic password

The invention discloses a wireless audio transmission protocol optimization algorithm for a low-delay dynamic password, which comprises the following steps of: generating an initial encryption key: before audio data transmission starts, generating the initial encryption key according to a current network environment, equipment computing power and audio transmission requirements; the encryption key can be generated based on a timestamp, a random number or a key exchange protocol, the randomness and the safety of the encryption key are ensured, and a proper encryption algorithm is selected: according to equipment computing resources and network conditions, a proper low-delay encryption algorithm is selected. By selecting a low-delay encryption algorithm and dynamically adjusting the encryption complexity, the influence of the encryption process on transmission delay is remarkably reduced, the real-time performance of audio transmission is optimized, the reliability of key synchronization in a wireless network is ensured through a redundant synchronous data packet and a regular key synchronization mechanism, and even if signal loss or interference occurs, the reliability of key synchronization in the wireless network is ensured. And synchronization can be recovered in time, so that decryption failure caused by asynchronous keys is avoided.
Owner:同辉佳视(北京)信息技术股份有限公司

Short-distance wireless audio multi-standard compatible communication protocol method

The invention discloses a short-distance wireless audio multi-standard compatible communication protocol method, which comprises the following steps of: intelligent protocol selection: automatically selecting a most suitable wireless communication standard to carry out audio data transmission according to information such as communication capability, network environment, bandwidth requirement, signal strength, equipment type and current network load of audio equipment; the communication standard comprises but is not limited to Bluetooth, Wi-Fi, Zigbee, NFC and the like, and multi-protocol cooperative scheduling: under the condition that a plurality of protocols coexist, the load, the bandwidth use condition, the signal quality and the like of each protocol are monitored in real time. Through protocol intelligent selection and multi-protocol cooperative scheduling, the equipment can automatically select the most suitable communication standard according to the actual situation, the problem that different protocols are incompatible is solved, the interconnection and intercommunication between the equipment are improved, the low-delay processing technology is adopted, the data packet transmission sequence and the protocol conversion process are optimized, and the communication efficiency is improved. The audio data can be transmitted and played in real time, delay is reduced, and audio experience is improved.
Owner:同辉佳视(北京)信息技术股份有限公司

Systems and methods for latency improvement for wireless speakers

The disclosure describes systems and methods for improving latency for wireless speakers. The system can establish a buffer pipe between a wireless chip of a media player and a wireless chip of a wireless speaker using an inter-IC sound protocol. The system can receive samples of audio data through the media player. The system can communicate the audio data to the wireless speaker using the buffer pipe to bypass the transport layer stack of the media player and the wireless speaker. The system can cause a transmitter of the wireless speaker to provide the audio data to a digital audio converter for output to a speaker of the wireless speaker.
Owner:AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD

WebM protocol low-delay video and audio translation and subtitle optimization method and system

The invention discloses a WebM protocol low-delay video and audio translation and subtitle optimization method and system, and belongs to the technical field of data processing, and the method comprises the steps: analyzing a WebM audio and video stream; establishing a video track-audio track association list and constructing a dynamic causal map; generating an initial translation text based on the audio frame sequence, the video frame sequence and the dynamic causal atlas, extracting lip movement and action semantic data of video frames through a reverse generation model to generate a completion translation text, and embedding SimpleBlock elements; defining a subtitle core gene, a non-core gene and a position gene according to the code rate data, the complete translation text and a user visual attention thermodynamic diagram, dynamically cutting the non-core gene based on code rate fluctuation, and adjusting the position gene in combination with the thermodynamic diagram to generate adaptive subtitle data; by synchronously playing and collecting feedback data, the edge weight and the subtitle position gene of the dynamic causal atlas are optimized. According to the invention, the robustness of audio and video translation and the self-adaptability of subtitle display are realized.
Owner:JIANGSU ZHIMENG INTELLIGENT TECH CO LTD

Ultra-low delay live broadcast method, system and device based on edge computing and medium

The invention discloses an ultra-low delay live broadcast method, system and device based on edge calculation, and a medium, and relates to the technical field of real-time audio and video interaction. The method comprises the steps that a target edge computing node receives an original media stream; monitoring an uplink network link and a downlink network link of the anchor end and the target edge computing node in real time, and determining a network link quality parameter; acquiring a historical network link quality parameter, and setting a link degradation standard and a link excellent standard according to the historical network link quality parameter; determining a target compensation strategy according to the network link quality parameter, a first preset condition and a second preset condition; and the target edge computing node processes the original media stream by adopting the target compensation strategy to generate a compensated media stream, and forwards the compensated media stream to a receiving end. By adopting the scheme, the overall performance of ultra-low delay live broadcast can be improved.
Owner:ZHEJIANG CHIGUANG DIGITAL TECHNOLOGY CO LTD

Dynamic systems and methods for media-aware transport of fragment of content in low-latency, over-the-top, and adaptive bitrate streaming

Low latency, over-the-top (OTT), and / or adaptive bitrate (ABR) content streaming is provided. Content delivery is enhanced by determining if a fragment of a content segment at a content delivery network (CDN) edge node meets a threshold for preferential encapsulation and transport. If met, preferential encapsulation and transport to the client device is provided; otherwise, it defaults to non-preferential encapsulation. The size of the fragment is quantified at a parser of the CDN edge node or an ABR segment encryption system. The ABR system may be connected between a content source and a CDN origin and may include an encryptor that sends CMAF video and audio segment's fragment byte offsets metadata. Also, the CDN edge node may include the ABR system and an encryptor that sends an encrypted CMAF segment's fragment size to a threshold calculator of an HTTP server. Related apparatuses, devices, techniques, and articles are also described.
Owner:ADEIA GUIDES INC

Telecommunications switch infrastructure for real-time audio interpretation and recording via artificial intelligence agents

A telecommunications switch-type platform positioned between a private-branch exchange and a data network dynamically hosts artificial-intelligence agents that interpret, transcribe, translate and transliterate live call traffic while the session is in progress. Machine-readable instructions stored in local memory instantiate, migrate and terminate the agents on demand across dedicated processing resources such as field-programmable gate arrays, application-specific integrated circuits or system-on-chip devices, thereby maintaining end-to-end latency below 250 milliseconds. An integrated call-audio recorder captures bidirectional media streams for secure archiving without interrupting service. The platform exposes a network interface that passes both signalling and media traffic, enabling seamless deployment inside call-centre or enterprise environments and compatibility with softswitch architectures conforming to class-4 Computational Interpretation of Communication using Artificial Intelligence standards. The same functional stack is deliverable as a computer-implemented method and as a non-transitory machine-readable medium storing the instructions executed by the processing resources.
Owner:CUNNINGHAM CHERYL

Bluetooth low-bit-rate audio encoding and decoding method and related device

The invention discloses a Bluetooth low-bit-rate audio encoding and decoding method and a related device, and the method comprises the steps: obtaining a Bluetooth real-time bandwidth, carrying out the dynamic wavelet processing of an input audio signal based on the Bluetooth real-time bandwidth, and decomposing to obtain a plurality of initial sub-band signals; quantizing the initial sub-band signal, and performing layered compression on the quantized initial sub-band signal through FSQ to obtain a compressed sub-band signal; performing differential coding processing on the compressed sub-band signal to obtain coded data; inputting the coded data into a lightweight Transform model to obtain a residual error correction item, and adding a forward error correction code to the coded data to obtain an audio data packet; packaging the residual error correction item and the audio data packet into a bit stream, and transmitting the bit stream to a decoding end through Bluetooth; and at a decoding end, performing decoding processing according to the bit stream, and recovering to obtain the target audio signal. The invention provides a high-efficiency, low-delay and low-power-consumption solution for Bluetooth audio transmission.
Owner:GUANGDONG UNIV OF TECH +1

Real-time multilingual transcription system and method

Disclosed are a method, system, and apparatus of a real-time multilingual transcription system and method. In one embodiment, a method includes continuously capturing an audio data and segment it into short segments; implementing a pre-trained enterprise-grade voice activity detection (“VAD”) system on each of the short segmental and filtering out non-speech segments to reduce computational waste, focusing resources on relevant audio data and minimizing latency. If speech is detected, a particular short segment is added to a processing queue. If speech is not detected, declining to add the particular segment to the processing queue, thereby reducing unnecessary processing.
Owner:GOVERNMENTGPT INC

Query response interface with server side generative model(s)

Various implementations include processing, at a client device, an instance of audio data capturing a user voice query using an automatic speech recognition model to generate a sequence of instances of tokenizable query text. In many implementations, one or more instances of the sequence can be transmitted to a remote computing system prior to generating the entire sequence. In a variety of implementations, each instance in the sequence can be processed using a generative model which includes a streaming multi-head attention portion. Responsive output can be transmitted from the remote computing system to the client device, where the client device renders the responsive output to the user. In many implementations, the time between the user speaking the user query and the client device rendering the responsive output is reduced, thus decreasing latency in the system.
Owner:GOOGLE LLC

Systems and methods for latency optimization for cloud applications

Described embodiments provide systems and methods for latency optimization for cloud applications. An agent of a client device comprising an audio decoder and a video decoder can monitor video and audio data paths of an application communicating audio / video (A / V) data from one or more servers to the client device. The agent can measure, using the audio decoder and the video decoder, an A / V latency and a lip-sync status of the video and audio data paths of the application. The agent can determine, based on at least one or more measurements of the A / V latency and the lip-sync status, to enable a low latency mode for at least one of the video decoder or the audio decoder. The agent can configure, responsive to the determination, the low latency mode on one of the video decoder or the audio decoder.
Owner:AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD

A method for low-latency transmission of audio and video interaction supporting multi-terminal synchronization

This invention discloses a low-latency audio and video interactive transmission method supporting multi-terminal synchronization, belonging to the field of audio and video interactive transmission technology. It includes: enabling data communication between the client and server to achieve audio and video interactive data transmission according to the WebSocket communication protocol; adaptively selecting the optimal streaming media transmission channel based on the ICE algorithm; transmitting audio and video interactive data from the client to the server according to the optimal streaming media transmission channel; monitoring the transmission status of the audio and video interactive data in real time; and dynamically adjusting and optimizing the audio and video interactive data transmission based on the monitoring results. This invention solves the problems of existing methods that cannot effectively achieve low-latency transmission and cannot achieve high-quality real-time audio and video interaction across multiple terminals. This invention can ensure synchronous low-latency transmission of audio and video interactive data, providing users with a smooth and natural interactive experience, and enabling high-quality real-time audio and video interaction across multiple terminals.
Owner:PTN ELECTRONICS LTD