Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

145 results about "Digital audio" patented technology

Digital audio is sound that has been recorded in, or converted into, digital form. In digital audio, the sound wave of the audio signal is encoded as numerical samples in continuous sequence. For example, in CD audio, samples are taken 44100 times per second each with 16 bit sample depth. Digital audio is also the name for the entire technology of sound recording and reproduction using audio signals that have been encoded in digital form. Following significant advances in digital audio technology during the 1970s, it gradually replaced analog audio technology in many areas of audio engineering and telecommunications in the 1990s and 2000s.

System and method for detecting deep fake audio

A system for analyzing audio includes a memory configured to store known digital audio representation containing known fraudulent audio streams and a processor operably coupled to the memory. The processor receives a portion of an audio stream from an external device and produces a transcript of the portion of the audio stream. The processor then determines a timing score, an emotional score, a background score, and a content score by analyzing the portion of an audio stream and the corresponding transcript and comparing them to the known digital audio representations and transcripts. The processor then determines if the audio stream is malicious by combining the timing score, emotional score, background score, and content score to produce a combined score and comparing the combined score to a threshold. The processor notifies a user that the call may be fraudulent when the combined score is greater than the threshold.
Owner:BANK OF AMERICA CORP

Digital audio processing system based on DSP

The invention discloses a digital audio processing system based on a DSP, and relates to the field of digital audio processing, and the system comprises an acquisition module which is used for capturing vehicle body vibration sound waves, environment incoming noise and in-cabin original audio signals, and generating a multi-dimensional acoustic data stream through voiceprint feature extraction so as to output an original sound signal matrix; the analysis module is used for mapping the original sound signal matrix to a phase space, extracting singular attractor features of noise evolution through a chaos theory, and generating a noise dynamic model; according to the method, various sound waves can be accurately captured and converted into multi-dimensional data, deep analysis and targeted counteracting of noise are achieved on the basis, meanwhile, an acoustic adjustment scheme matched with the environment in the cabin can be generated, the acoustic environment can be dynamically optimized in combination with the physiological state of passengers, and potential noise is pre-judged in advance for suppression.
Owner:DONGGUAN MINDONG ELECTRONIC TECH CO LTD

Intelligent sound field adaptive system of digital professional sound equipment

The invention belongs to the field of artificial intelligence, particularly relates to an intelligent sound field adaptive system of digital professional sound equipment, and aims to solve the problems of inaccurate sound field regulation and control, response lag, dependence on manual tuning and the like in a complex acoustic environment. The system comprises a sound field sensing module, an acoustic modeling and analysis module, a self-adaptive sound field regulation and control engine, a multi-channel digital audio processing unit and a feedback optimization module, and dynamic sound field modeling, real-time audio processing and environment self-adaptive regulation and control are realized through distributed sensing, hybrid modeling, multi-target optimization and closed-loop feedback. The system supports rapid re-calibration, multi-scene memory and user preference learning, ensures voice clarity and music fidelity, balances full-field hearing consistency, and significantly reduces manual intervention requirements.
Owner:深圳市多乐声电子有限公司

Spatial audio file making method and system based on Audio Vivid standard

The invention discloses a spatial audio file making method based on an Audio Vivid standard. The spatial audio file making method is characterized by comprising the following steps: loading a Panner module in a digital audio workstation (DAW) in a VST3 or AAX plug-in form; receiving at least one path of audio signal through the Panner module, and setting a spatial position, a motion track, a diffusion coefficient and a gain parameter of an audio object or a sound bed in a three-dimensional rectangular coordinate system; based on the obtained parameters, object-level metadata conforming to GY / T345-2021 and Audio Vivid specifications are generated in real time, and the Renderer module outputs monitoring signals of binaural or any loudspeaker layout in real time; and after rendering is completed, an Audio Vivid encoder is called to package the audio signal and the complete metadata into an ADM file, an AV3A file or a WAV file with a self-defined RIFF block, and the embedding rate of the metadata is 100%. According to the invention, the blank in the field of professional spatial audio production tools in China is filled, the industrial application of the Audio Vivid standard is promoted, and the technical progress and the industrial value are remarkable.
Owner:CHINA TELEVISION INFORMATION TECH BEIJINGCO

Audio channel verification method and device, equipment, storage medium and program product

The invention discloses an audio channel verification method and device, equipment, a storage medium and a program product, and relates to the technical field of communication. According to the method, under the condition that a data issuing interface receives audio algorithm data sent by an upper-layer application, the audio algorithm data is packaged into an RTAC format to obtain real-time audio calibration data; the audio link state of the digital audio processor is obtained through the real-time audio calibration module, under the condition that the audio link state is a creation state, a data issuing thread is started, and real-time audio calibration data are sent to the digital audio processor through the real-time audio calibration module. The digital audio processor is used for processing the real-time audio calibration data to obtain a return value and sending the return value to the real-time audio calibration module; and starting a result callback thread, obtaining a return value from the real-time audio calibration module through a result callback interface, and performing audio channel verification according to the real-time audio calibration data and the return value to obtain an audio channel verification result. According to the embodiment of the invention, the audio channel verification flexibility is improved.
Owner:THUNDERSOFT

Apparatus and methods for enhanced digital audio bus reliability

Apparatus and methods digital audio bus reliability are disclosed. In certain embodiments, a digital audio system includes a plurality of audio devices connected by a first digital audio chain in a clockwise direction and a second digital audio chain in a counterclockwise direction. The first digital audio chain and the second digital audio chain run concurrently, and a controller selects which audio chain to operate at a given time for audio connectivity. For example, the controller can initially select the first digital audio chain to provide audio connectivity, but transition selection from the first digital audio chain to the second digital audio chain in response to detecting a node failure in the first digital audio chain. Thus, the system is tolerant to node failures while maintaining system connectivity.
Owner:ANALOG DEVICES INT UNLTD CO

An integrated smart steering digital audio device

The application belongs to the technical field of sound equipment, and particularly relates to an integrated intelligent steering digital sound equipment. In view of the problem that the sound unit of the current sound equipment cannot follow the movement of the personnel to adjust the direction, the following scheme is provided, which comprises a main shell, a mesh panel, a sealing back plate, a front decorative panel and a rear decorative panel. A sound unit mounting groove is formed in the front side of the main shell. A steering assembly is mounted in the inside of the sound unit mounting groove. The inside of the steering assembly is mounted with a sound unit. A control panel is further arranged on the top of the main shell. The processor of the application generates data by real-time identification of the position of the human body through an algorithm, including the direction and the distance, and calibrates the generated direction and distance with the current position of the steering assembly. The four connecting rod telescopic assemblies in the steering assembly act in coordination to control the rotation of the four driving motors, so that the sound unit outputs sound aiming at the current position of the human body.
Owner:SHENZHEN ZUNTE DIGITAL CO LTD

On-line abnormal sound monitoring device and safety monitoring method for valve cooling external cold water cooling tower based on sound sensor

The invention relates to the technical field of safety protection, in particular to a valve cooling external cold water cooling tower abnormal sound online monitoring device and safety monitoring method based on a sound sensor, and the device comprises a microphone collection module which is arranged on a target valve cooling external cold water cooling tower and is used for obtaining acoustic signals generated in the operation process of the cooling tower in real time, the acoustic signal is converted into digital audio data; the MCU mainboard module is connected with the microphone acquisition module and is used for receiving the audio data of the microphone acquisition module and realizing feature extraction, model reasoning and logic control; the communication module is used for transmitting processing information of the MCU mainboard module to a mobile terminal to realize online early warning of abnormal sound of the valve cooling external cold water cooling tower, and the device realizes real-time sound monitoring of the cooling tower, judges equipment of which the sound exceeds a normal range, and sends an alarm signal at the terminal to ensure that equipment defects are found as early as possible and treated in advance; defect deterioration is avoided.
Owner:GUANGZHOU BUREAU CSG EHV POWER TRANSMISSION

Multi-channel audio input mixer

In some aspects, an audio processor (310) may provide a set of TDM clocks (330-1, 330-2) to each digital sample rate converter (325) in a time division multiplexed (TDM) data chain (320-1, 320-2), the set of TDM clocks comprising a sample rate clock input and a bit clock input. A digital sample rate converter (325) in each TDM data chain (320-1, 320-2) is connectable to respective audio ports (I2S0 to I2S7), each of which corresponds to a stereo channel. A digital sample rate converter (325) in each TDM data chain (320-1, 320-2) may receive a digital audio input via the audio ports (I2S0 to I2S7). An audio processor (310) may receive, at one or more TDM inputs (TDM (in) 0, (TDM (in) 1), a TDM audio stream from each of the one or more TDM data chains (320-1, 320-2), where the TDM audio streams mix the digital audio input based on a sample rate clock input and a bit clock input. Numerous other aspects are described.
Owner:QUALCOMM INC

Digital audio processor

The utility model discloses a digital audio processor which comprises an installation cover, a placement cavity is reserved in the installation cover, an audio processor body is arranged in the placement cavity, a water-cooling cooling mechanism is installed on one side of the installation cover, and an air-cooling cooling mechanism is installed at the top end of the installation cover. The cooling device has the beneficial effects that a water-cooling cooling mechanism consisting of a water return pump, a cooling water tank, a circulating water pipe and a water suction pump is mounted on one side of the mounting cover, and an air-cooling cooling mechanism consisting of an air blower, a semiconductor chilling plate, a cooling pipe, a concentric-square-shaped spraying cavity and an air spraying hole is mounted at the top end of the mounting cover; according to the audio processor provided by the invention, the air cooling and the water cooling are combined to efficiently dissipate heat of the audio processor body when the audio processor body is used, so that the problem that the cooling speed is slowed down due to heat absorption and temperature rise of water in the cooling process during water cooling is avoided, and the heat dissipation efficiency of the audio processor is higher when the audio processor is used.
Owner:SHANXI HUANHAO TECHNOLOGY CO LTD

Systems and methods for synchronizing and rendering audio signals via an auxiliary synchronization back channel

Methods, systems, and computer program products are presented herein for rendering synchronous audio using audio rendering devices. A networked audio rendering device may include a first wireless communications module, a second wireless communications module, a microcontroller, at least one speaker driver, and at least one digital audio amplifier. Audio data may be received by a first wireless communications module. Timecode data may be received by a second wireless communications module operating at a sub-GHz ISM band. The audio data may be received by a microcontroller. The audio data is processed by the microcontroller to generate output audio signals based on the timecode data. The output audio signals may be amplified by at least one digital audio amplifier. At least one speaker driver may be driven, by the at least one digital audio amplifier, to render an audio output.
Owner:LENBROOK IND LTD +2

An integrated smart steering digital audio device

The application discloses an integrated intelligent steering digital audio equipment, and belongs to the technical field of digital audio equipment, which comprises an equipment body, a front cover plate arranged at the front of the equipment body, two loudspeaker screen meshes clamped in the middle of the front cover plate, and supports clamped to the outer walls of the two loudspeaker screen meshes. In the application, the equipment body, the front cover plate, the loudspeaker screen meshes, the supports, the connecting cylinder, the spring, the fixing rod, the clamping rod and the clamping groove plate are arranged, the clamping rod is designed in a rectangular shape and is matched with the clamping groove plate, the purpose of limiting the connecting cylinder, the supports and the loudspeaker screen meshes is effectively achieved, the loudspeaker screen meshes are shaken to make dust and impurities fall down together, compared with the traditional external wiping mode, the loudspeaker screen meshes are changed in angle and are dislocated with the loudspeaker mouth, the purpose of cleaning both sides of the loudspeaker screen meshes is effectively achieved, some small dust is not easy to push into the loudspeaker screen meshes through the partition holes, the cleaning is convenient, and the playing sound quality of the equipment body is ensured.
Owner:WUHAN AIJIN TECH CO LTD

Error correction overwrite for audio artifact reduction

Audio communication methods, devices, and systems, are provided with error correction overwrite for audio artifact reduction. One illustrative low-latency audio streaming method includes: receiving packets of digital audio data; applying an error correction code decoder to obtain a data stream that includes error-corrected data samples; providing a correction-limited data stream by replacing any of the error-corrected data samples that are outliers; and converting the correction-limited data stream into an audio signal.
Owner:SEMICON COMPONENTS IND LLC

Ai powered digital audio workstation using quantum algorithm

PendingUS20260212848A1AlgorithmSoftware system
An AI-powered multimedia workstation integrates a digital audio workstation (DAW), visual graphics editor, and office productivity tools within a single platform. Operating on an AI / quantum-inspired software system, it performs real-time audio, visual, and data processing. A central AI / Quantum Content DNA Engine enables structural, stylistic, and semantic analysis, separation, and manipulation of multimedia content, including audio, video, images, and text. The workstation features a hybrid control surface with dynamic haptic feedback faders, high-resolution touch displays, and professional audio input / output interfaces. AI functions include automated music composition, intelligent mixing and mastering, cross-modal content generation, and context-aware workflow assistance. Quantum-inspired algorithms accelerate rendering, simulation, and large-scale data analysis. The system supports real-time global collaboration via high-speed network protocols and cross-platform compatibility with external devices and software. By combining AI-assisted and quantum-inspired processing, the workstation enhances creative production, performance, and productivity, providing an integrated environment for professional audio, visual, and office applications.
Owner:BROOKS JR ORLANDO A

Audio and video synchronization detection method and system based on deep learning

The invention discloses an audio and video synchronization detection method and system based on deep learning, and relates to the technical field of digital audio and video processing, and the method comprises the steps: carrying out the video stream and audio stream separation of a collected to-be-detected audio and video file, and generating multi-modal audio and video features through face detection and feature extraction; according to the pure data packet, a SyncNet deep double-flow network is adopted to carry out synchronism discrimination, lip movement features and voice features are extracted respectively, and a synchronization discrimination result packet is generated; and according to the detection report, verifying the audio and video synchronism by adopting a time sequence alignment algorithm, and finishing result solidification through timestamp anchoring to obtain a synchronism judgment result. According to the invention, the SyncNet deep double-flow network is adopted to carry out synchronism judgment, efficient extraction and synchronism judgment of the audio and video lip movement features and the voice features are realized in combination with deep learning, the matching degree between the target voice and the mouth shape can be accurately recognized, and the accuracy of synchronism detection is improved.
Owner:SUYUAN TECHNOLOGY (HUNAN) CO LTD

A noise monitoring digital audio tamper detection method based on power grid frequency signal

ActiveCN122173839BNoise monitoringTimestamp
The application discloses a noise monitoring digital audio tampering detection method based on a power grid frequency signal, and steps include: obtaining a noise monitoring digital audio and extracting a power grid frequency signal; signal pretreatment is performed on the extracted power grid frequency signal; the pretreated power grid frequency signal is input into a trained deep learning model to obtain the two-classification probability of each data in the power grid frequency signal being normal or tampered; the model output is normalized to calculate the tampering probability of each data in the power grid frequency signal, compared with a preset threshold, and a tampering mark is set to obtain a prediction sequence; the prediction sequence is subjected to connected domain analysis to identify continuous tampering segments, and the sample index of the tampering segments is mapped back to the timestamp of the original noise monitoring digital audio to output a corresponding detection report. The application automatically learns a tampering feature mode through a deep neural network, realizes point-by-point labeling of a tampering position, and realizes high-precision tampering detection while keeping lightweight.
Owner:HUNAN UNIV

Iot digital audio power amplifier

1. The name of the design product: Internet of Things digital audio power amplifier. 2. The use of the design product: for power amplification of audio signals. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view.
Owner:TAISI INTERNET OF THINGS TECH (GUANGZHOU) CO LTD

Generating tone-compatible synchronized neural metronomes for digital audio files

The invention relates to generating tone-compatible synchronized neural metronomes for digital audio files. Methods and systems for improved neural rhythm generation of digital audio files are provided. In one embodiment, a method is provided that includes receiving a digital audio file and a beat frequency of a neural beat. Chroma diagram features may be extracted from a digital audio file and may be used to identify dominant sound levels within the digital audio file at a plurality of timestamps. A plurality of carrier frequencies for different time periods within the digital audio file may be selected based on the dominant level. Neural metronomes may be synthesized for a digital audio file based on a metronome frequency of a plurality of carrier frequencies. The neural rhythms may be stored and / or may be combined with digital audio files to generate combined audio tracks that may be stored.
Owner:UNIVISER INT MUSIC CO LTD

An interactive system and method based on digital video and audio family tree and family memory simulation

The application discloses an interactive system and method based on digital audio and video family tree and family memory simulation, and aims to solve the problems of weak interaction of traditional family tree, rigid digital clone model and split memory personality. The system constructs an interactive family tree knowledge base containing member nodes, blood relationship and related data, and forms a structured family knowledge graph in multiple dimensions. The model graph co-evolution is proposed: multi-modal data is collected through daily interaction, incremental learning and knowledge graph updating are adopted in parallel, parameter efficient fine-tuning is used for lightweight learning of the digital clone, only a small amount of parameters are updated to overcome forgetting, and new event relationships are extracted to update the graph. Through bidirectional indexing, the personality and memory semantics are consistent, the multi-modal interaction is relationship-aware and dynamically evolved. The application synchronously updates the personality and memory, solves the problem that the digital clone cannot continuously evolve with the family, and realizes stereoscopic inheritance of family memory.
Owner:BEIJING AIHE INFORMATION TECHNOLOGY CO LTD

A communication assistance method, master device, system and storage medium

The application discloses a communication auxiliary system, a master control device and a method, and belongs to the technical field of communication. The system adopts a mediation double-agent architecture, the master control device is logically connected between a communication terminal and a headset device, and meanwhile, isolated physical connections are established by simulating hands-free (HF) device specifications and audio gateway (AG) device specifications. The master control device is internally provided with an uplink data decision module and a memory bus with a first buffer address and a backup buffer address. When a local trigger event (such as touch, action or external entity button pulse) is captured, the decision module forcibly switches the memory data reading pointer of the digital audio by using the bottom-layer hardware atomic operation in an instant, and replaces the headset pickup source with an external high-definition pickup source. The whole switching process does not trigger the bottom-layer audio routing redistribution (Audio Policy) of the communication terminal operating system, and the absolute continuity of the Bluetooth bottom-layer SCO synchronous link frame sequence is maintained. The system realizes the high-response seamless migration and injection of the external pickup source without invading the protocol stack of the host device, supports the double-track solidification of the multi-dimensional voice shunting of the communication link, and significantly improves the communication and pickup efficiency in modern mobile and complex scenarios.
Owner:金可人 +1

Scene-aware speech recognition using vision-language models

Ae system to generate a latent space model of a scene or video and apply this latent space and candidate sentences formed from digital audio to a vision-language matching model to enhance the accuracy of speech-to-text conversion. A latent space embedding of the scene is generated in which similar features are represented in the space closer to one another. An embedding for the digital audio is also generated. The vision-language matching model utilizes the latent space embedding to enhance the accuracy of transcribing / interpreting the embedding of the digital audio.
Owner:NVIDIA CORP

CALIBRATION AMPLIFICATION FOR REAL-WORLD SOUND

ActiveDE102022213018B4Application programming interfaceProgramming
Computer-readable medium containing instructions that configure a computer to: present an application programming interface (API), wherein the API includes a user-specified sound level parameter; receive, via the API, one or more values ​​for the user-specified sound level parameter, which are assigned to a digital audio asset;and incorporating code into a simulated reality application that is currently being assembled, wherein the code determines a loudness correction gain for the digital audio asset by determining a difference between i) the one or more values ​​for the user-specified sound level parameter and ii) a parameter of the playback hardware device that represents a known or predefined sound level that can be produced by a playback hardware device on which the simulated reality application is running, wherein the loudness correction gain compensates for the difference, and the loudness correction gain is then to be applied to the digital audio asset during the runtime of the simulated reality application.
Owner:APPLE INC

Audio transmission method, audio equipment and wireless audio system

The invention relates to the technical field of communication interaction, in particular to an audio transmission method, audio equipment and a wireless audio system. The audio transmission method comprises the following steps: generating and caching first digital audio data based on a collected first audio signal; when the cumulant of the first digital audio data reaches a first predetermined number, transmitting the first predetermined number of first digital audio data to a wireless transmission storage area, and determining a first starting moment for transmitting the first predetermined number of first digital audio data; according to the first starting moment, determining a first sending moment for data transmission with the second audio equipment; and sending a wireless frame to a second audio device based on the target wireless system and the first sending moment, so that the second audio device aligns the second digital audio data with the first digital audio data on the time axis based on the received first digital audio data. The time sequence jitter in data transmission can be reduced, and the certainty and consistency of cross-device audio time alignment can be improved.
Owner:HENGXUAN TECH (BEIJING) CO LTD

Robust audio watermarking method based on adaptive quantization strategy and feature classification

PendingCN121506155ASpeech analysisFeature vectorAudio watermark
The invention discloses a robust audio watermarking method based on an adaptive quantization strategy and feature classification, and belongs to the technical field of digital audio copyright protection. The method comprises the following steps: framing an audio signal, extracting a logarithmic mean feature (DWT-CLM) of discrete wavelet transform, and combining a zero-crossing rate, a variance and energy to form a frame feature vector; a Sigmoid classifier is used to discriminate frame characteristics, and a fixed or variable quantization step size is adaptively selected to embed watermark bits into approximate components; and during extraction, the watermark is accurately extracted through the same feature analysis and classifier discrimination recovery quantization mode. Experiments show that the algorithm has better inaudible property and robustness, can effectively resist attacks such as MP3 compression, resampling, low-pass filtering and re-recording, and is suitable for digital audio copyright protection scenes.
Owner:XINYANG NORMAL UNIVERSITY

Wireless transmission apparatus, display apparatus, and data transmission method and stream processing method thereof

A transmission device includes: a communication interface configured to perform communication with an external device and at least one display device; a memory configured to store analog data received from an external device; a signal processor; and a processor, where the processor is configured to control the signal processor to: decode the analog data; separating the decoded analog data into analog audio data, analog video data and analog additional data; respectively converting the analog audio data, the analog video data and the analog additional data which are separated from one another into digital audio data, digital video data and digital additional data; configuring a digital transport stream including a plurality of packets by individually packing the digital audio data, the digital video data, and the digital additional data, respectively; and transmitting the digital transport stream to the at least one display device using the communication interface.
Owner:SAMSUNG ELECTRONICS CO LTD

An audio playback device and method

The application discloses an audio playing device and method, which is applied to the technical field of electronic equipment and aims at solving the problem that the audio signal source supported by the audio playing device is relatively single in the prior art. Specifically, a first socket receives an audio signal transmitted by an external first plug; a plug detection circuit determines the plug type and generates a corresponding plug type signal and sends the plug type signal to a Bluetooth chip; a first audio transmission circuit transmits the input audio signal of an analog audio plug to the Bluetooth chip; a second audio transmission circuit transmits the input audio signal of a digital audio plug to the Bluetooth chip; and the Bluetooth chip is used for determining the plug type of the external first plug according to the plug type signal, switching to a working mode corresponding to the plug type of the external first plug and receiving the audio signal transmitted by the corresponding audio transmission circuit. In this way, the audio playing device is compatible with different types of audio signal sources, and the first socket is multiplexed, so that the design of the audio playing device is simplified.
Owner:JIANDA INTELLIGENT TECH CO LTD

Load compensation circuit

This load compensation circuit provides a way to maintain a constant power supply voltage against load fluctuations that occur in digital audio equipment and other loads, including those that cannot be handled by capacitors. [Solution] The noise detection circuit 10 detects the voltage at the midpoint of the power supply supplied to the load as a reference voltage, detects noise on the positive terminal side of the power supply based on the reference voltage, and amplifies the noise with an amplifier OP and outputs it, and the current control circuit 20 supplies current to the load in accordance with the noise amplified by the amplifier OP of the noise detection circuit 10.
Owner:高桥 功 +1

AI-based scene recognition-based intelligent control method and system for digital audio sound field parameters

PendingCN122138094ABiological modelsTransducer circuitsEnvironmental acousticsEngineering
This application discloses an AI-based method and system for intelligent control of digital speaker sound field parameters, relating to the field of digital professional audio equipment technology. The method includes: acquiring environmental acoustic parameters and user-played content, and identifying the playback scene to obtain the playback scene; performing auditory demand analysis to obtain an auditory demand solution; obtaining an initial sound field control target, and conducting an environmental impact assessment based on the environmental acoustic parameters to obtain an optimized sound field control target; iteratively optimizing the control parameters of the digital speaker in conjunction with the playback scene to obtain an optimized control scheme, and intelligently controlling the digital speaker. This solves the technical problem that existing digital speaker sound field parameter control methods cannot dynamically adapt to complex environmental changes and personalized auditory needs, resulting in poor sound quality and a poor user experience.
Owner:SHENZHEN HANKE TECH CO LTD

Method and system for producing synthesized speech digital audio content

A method for producing synthesized speech digital audio content, wherein:a feature extractor module receives an audio recording of a speaker's voice, extracts a plurality of acoustic features and converts them to an audio latent representation matrix;a phonemizing module receives as input a target text and converts the target text to a sequence of phonemes;a tokenizing module receives as input the sequence of phonemes of the target text;a linguistic encoder module receives as input the sequence of phoneme vectors and converts the sequence of phoneme vectors to a sequence of respective linguistic latent vectors;an acoustic model module produces a predicted audio latent representation matrix; anda vocoder module decodes the predicted audio latent representation matrix into the corresponding audio signal of the speech of the synthesized virtual voice.
Owner:VOISEED SRL