Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

83 results about "High fidelity" patented technology

High fidelity (often shortened to hi-fi or hifi) is a term used by listeners, audiophiles and home audio enthusiasts to refer to high-quality reproduction of sound. This is in contrast to the lower quality sound produced by inexpensive audio equipment, or the inferior quality of sound reproduction that can be heard in recordings made until the late 1940s.

System and Methods for Upsampling of Decompressed Audio Data Using a Neural Network

A computer system for upsampling decompressed audio data after lossy compression using specialized neural network techniques. The system processes compressed audio channels through an audio pre-processor that extracts spectral information, detects speech activity, segments audio, and normalizes input levels. A trained deep learning algorithm with multi-channel transformers using channel-wise and self-attention mechanisms recovers information lost during compression. The system further enhances audio quality through a time-frequency domain transformer applying Fourier transforms and Mel-scale frequency processing, while a perceptual quality assessor employing psychoacoustic models evaluates the output. This specialized audio processing approach significantly improves reconstructed audio quality by leveraging correlations between audio channels, addressing both spectral and temporal features, and optimizing for human perception characteristics, resulting in higher fidelity audio reproduction from compressed formats.
Owner:ATOMBEAM TECH INC

Loudspeaker speech enhancement method and device, electronic equipment and storage medium

The invention provides a loudspeaker speech enhancement method and device, electronic equipment and a storage medium, and belongs to the technical field of speech signal processing, and the method comprises the steps: obtaining an initial spectrum signal and an initial power spectrum signal corresponding to an initial audio signal; inputting the initial power spectrum signal into a feedback elimination network model to obtain a target frequency domain signal; and obtaining a target audio signal based on the target frequency domain signal. The method comprises the following steps: performing mask adjustment on the frequency domain amplitude of an initial frequency spectrum signal by using a real number signal mask to generate a frequency domain estimation signal, and performing mask adjustment on the frequency domain amplitude and phase of the initial frequency spectrum signal by using a complex number signal mask to obtain a frequency domain enhancement signal. And finally, a target frequency domain signal is determined based on the frequency domain estimation signal and the frequency domain enhancement signal, so that the output target frequency domain signal has higher fidelity in the aspects of amplitude and phase, environmental noise and loudspeaker feedback signals can be removed to the greatest extent, and the squeal phenomenon is effectively inhibited.
Owner:IFLYTEK CO LTD

High-fidelity, low-contact data retrieval from legacy tapes

ActiveUS12555596B2Tape carriersHeads relative to moving tapeComputer graphics (images)Data retrieval
Systems and methods are disclosed for retrieving information from legacy storage media with high fidelity and minimal contact. A read sensor with multiple read elements detects electromagnetic fields from tracks on a storage medium. Readings from different elements are combined to generate a composite interpretation, which can be used to, for example, generate digitized waveforms and determine if portions of tracks are damaged.
Owner:CHANNELSCIENCE LLC

High-fidelity AI voiceprint cloning method and system and storage medium thereof

The invention relates to the technical field of voice recognition, and particularly discloses a high-fidelity AI voiceprint cloning method and system and a storage medium thereof.The method comprises the steps that user audio data are collected, de-noising processing is conducted on the audio data, the de-noised audio data are analyzed, and a de-noising processing quality evaluation result is generated; voiceprint feature extraction is carried out on the audio data meeting the denoising success result, denoising processing is carried out on the audio data not meeting the denoising success result again, parameters in the denoising processing process are adjusted, a cloned voice signal is generated, a cloned evaluation result of the cloned voice signal is collected, and the cloned voice signal is obtained. And storing the cloned voice signals meeting the voiceprint cloning success result, carrying out voiceprint feature extraction on the cloned voice signals not meeting the voiceprint cloning success result again, optimizing voiceprint feature extraction process parameters, and finally obtaining the high-fidelity AI cloned voice signals. And the problem of sound quality degradation caused by separation of de-noising processing and feature extraction links is solved.
Owner:HUNAN BOJI LIFE TECHNOLOGY CO LTD

Directional pickup method and device, computer equipment and medium

The invention relates to the technical field of directional pickup, and discloses a directional pickup method and device, computer equipment and a medium, and the method comprises the steps: obtaining time domain voice signals collected by at least two microphones, and carrying out the time-frequency conversion; calculating an observation phase difference of each frequency point, and determining a target theoretical phase difference according to a target pickup direction specified by a user and the geometric parameters of the microphone; encoding the target theoretical phase difference into a query vector, encoding the observation phase difference into a key vector, and generating a space confidence mask through a geometric cross attention mechanism; and filtering the frequency domain signal by using the mask, and outputting enhanced voice in a target direction through inverse time-frequency transformation. According to the invention, through an attention mechanism guided by physical prior, the problem of phase winding is effectively solved, and high-robustness and high-fidelity speech pickup in any direction is realized in a strong reverberation and low signal-to-noise ratio environment.
Owner:深圳市友杰智新科技有限公司

Digital human image generation method combining cold start driving and active learning mechanism

The invention discloses a digital human image generation method combining cold start driving and an active learning mechanism, and the method comprises the steps: processing input data through a preset few-sample emotional speech generation model which is derived from a cold start process and has high individuation and dynamic adaptability; according to the method, efficient personalized and multi-emotion voice batch generation can be realized on large-scale unlabeled input data, so that a target audio file with high fidelity and consistent emotion and personality expression is output; the preset few-sample emotional speech generation model is obtained by training a first qualified sample screened by a first cold start quality evaluator on the basis of candidate training samples, so that the model can be started under the condition of extremely little data, and automatic generation of a high-quality personalized digital human image under the condition of small samples is realized; besides, by inputting a target emotion embedding vector and a target character feature vector, synchronous expression generation under voice driving is realized, and the natural interaction capability and style consistency of the digital human are enhanced.
Owner:SHAANXI JIEQI NETWORK TECHNOLOGY CO LTD

electrostatic speaker

To provide an electrostatic speaker capable of maintaining high fidelity sound quality output even in a wide range of sound. [Solution] The electrostatic speaker includes an electrostatic sound generating diaphragm unit consisting of two parallel, spaced-apart electrode plates (1) and a diaphragm (2) between the electrode plates. The surface of the electrode plate facing the diaphragm is coated with a conductive layer, and both sides of the diaphragm are also coated with conductive layers. A voltage is applied to the conductive layers of the electrode plate and the diaphragm. Hollow sound-emitting holes are uniformly formed in the electrode plate, and an insulating buffer layer (3) is filled between the electrode plate and the diaphragm. Even if the applied high-voltage electric field voltage causes the diaphragm to distort due to excessive amplitude, the insulating buffer layer limits the diaphragm's amplitude, preventing distortion. The insulating buffer layer also prevents contact between the diaphragm and the electrode plate and maintains an appropriate minimum distance, preventing destructive discharges caused by too close a distance between the diaphragm and the electrode plate. This prevents the appearance of high-voltage discharge noise and prevents the diaphragm from being damaged by high-voltage discharge.
Owner:ABLE AUDIO TECHNOLOGY CO LTD

Underground space optical cable high-fidelity sound restoration method based on DAS system

The application discloses a kind of underground space optical cable high fidelity sound restoration method based on DAS system, comprising: the real-time signal collected to optical cable is carried out data demodulation, identify whether it contains energy impact signal generated by the optical cable after being padded and being hit, if it contains, the energy abnormal distribution position in energy impact signal is positioned, obtains positioning point;Continuous analysis time domain signal of positioning point, after filtering low-frequency interference signal, phase anomaly signal is rejected;The sound signals of a plurality of equivalent sensing nodes adjacent to the positioning point are superimposed and averaged to suppress additive noise, and then the sound signal is restored.The application can realize rapid positioning according to the spatial distribution mapping relationship of optical cable, and after filtering background noise from the collected signal, the highly correlated signals of multiple spatial sensing nodes near the positioning position are linearly superimposed to suppress additive random noise, and high-fidelity sound signal restoration is realized.
Owner:NANJING UNIV

USB high-fidelity decoder with vibration unit

The utility model discloses a USB high-fidelity decoder with a vibration unit, which comprises an integrated board. According to the utility model, the USB assembly, the system control assembly, the Bluetooth assembly, the first motor assembly, the first vibration driving assembly, the audio signal processor assembly, the loudspeaker assembly, the second motor assembly, the second vibration driving assembly, the earphone assembly and the high-fidelity digital-to-analog conversion assembly are integrated in the integrated board, so that the integrated board forms an integral module; the module can be attached to the surface of an object such as a mobile phone and can also be embedded into a mobile phone shell, the design that sound is experienced through the touch sense of the human body is achieved, a user can feel the sound in multiple dimensions on the basis of traditional auditory sense, the overall experience of entertainment equipment is enhanced, and meanwhile brand new understanding of music / sound is achieved.
Owner:SHENZHEN ZHIHENG TIMES TECHNOLOGY CO LTD

A method and apparatus for audio processing based on semantic driving

The application discloses a kind of based on semantic drive's audio processing method and device;The method comprises: spatial decomposition is carried out to multi-channel audio signal, and forward target signal and backward environmental signal are extracted;Respectively carry out statistical noise reduction, obtain forward noise reduction gain and backward noise reduction gain;Respectively input acoustic semantic perception network, obtain forward and backward semantic probability vector;Based on semantic probability vector and semantic complementary weight matrix, the dynamic transmission weight of backward environmental signal is calculated;Weighted fusion is carried out using forward noise reduction gain, backward noise reduction gain and dynamic transmission weight, and enhanced audio signal is output;The present application transmits side rear important acoustic event on demand through semantic gating mechanism, retains environmental situation awareness capability while suppressing environmental noise, adopts statistical noise reduction and lightweight network heterogeneous deployment, meets the requirement of low delay and high fidelity of edge device, significantly improves the user experience of auditory enhancement device.
Owner:TSINGHUA UNIVERSITY +1

High Fidelity Player (H41)

1. The name of the design product: high fidelity player (H41). 2. The use of the design product: the design product is used for high fidelity player. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view 1.
Owner:SHENZHEN LE CROSS COUNTRY TECHNOLOGY CO LTD

A dual-channel audio control circuit

This invention relates to a dual-channel audio control circuit, comprising a dual potentiometer circuit, a left channel driver circuit, a left channel output filter circuit, a right channel driver circuit, a right channel output filter circuit, a left channel speaker, and a right channel speaker. The dual potentiometer circuit receives and adjusts the dual-channel audio signals. The input terminal of the left channel driver circuit is connected to the left output terminal of the dual potentiometer circuit, and the input terminal of the right channel driver circuit is connected to the right output terminal of the dual potentiometer circuit. The left channel output filter circuit is connected between the left channel driver circuit and the left channel speaker, and the right channel output filter circuit is connected between the right channel driver circuit and the right channel speaker. This invention offers significant advantages in elevator voice systems, providing a more natural and immersive listening experience while meeting the requirements for high fidelity, flexible control, and system integration.
Owner:TIANJIN XINBAOLONG ELEVATOR GRP

A replaceable cord type high fidelity noise reduction earphone based on a Loth iron moving unit

The application belongs to the technical field of electro-acoustic conversion equipment, and particularly relates to a replaceable-line high-fidelity noise-reducing in-ear earphone based on a floor iron moving unit, which comprises a 5-unit 4-frequency division hybrid driving module formed by the floor iron moving unit and a micro moving coil unit, and the two achieve signal cooperation through a phase synchronization circuit; full-frequency sound quality optimization and hybrid noise reduction are achieved through the phase synchronization circuit and a DSP+ANC chip, a sound cavity is matched with a shape memory alloy damping adjusting sheet, magnetic attraction and sealing of a 2-pin interface improve durability and sealing performance, a double-mode switching audio line and a three-in-one replaceable plug adapt to multiple devices, and touch control and a selectable Bluetooth kit realize convenient interaction. The application breaks through the sound quality limitation of a single unit through a double-unit cooperative driving structure, realizes dynamic adaptation of noise reduction and sound quality through an adaptive acoustic adjusting system, improves durability and personalized upgrade space through a reinforced replaceable line design, enriches use scenarios through an intelligent interaction module, and meets the multiple needs of high-end users for high-fidelity sound quality, flexible use and intelligent experience.
Owner:山东峰启文化创意产业有限公司

Hearing aid device, method and device with self-adaptive hearing compensation function

The invention provides a hearing aid device, method and device with a self-adaptive hearing compensation function, the hearing aid device comprises a wireless communication module, a main control chip and an audio output module, and a pre-training hearing compensation model is built in the main control chip. A full-band continuous personalized frequency response gain curve with gain amplitude positively correlated with hearing loss is output through model reasoning, an audio output module performs full-band frequency-point-by-frequency balanced linear gain compensation on an original digital audio signal played by the device based on the curve, the audio dynamic range is not compressed in the whole process, the original spectrum structure is not changed, and the gain amplitude of the original digital audio signal is not changed. And the high-fidelity tone quality is kept. Different from a traditional hearing aid and a conventional audio playing device, only for audio processing inside the device, pickup and noise reduction are not needed, hearing features of a hearing impaired user can be accurately adapted, clear and balanced listening can be achieved, tone quality distortion can be avoided, and compensation precision, use comfort and scene universality are considered.
Owner:SHENZHEN POROS TECH CO LTD

Correlating scene-based audio data for psychoacoustic audio coding

To improve the encoding of scene-based audio data, after generating background components, foreground audio signals, and corresponding spatial components from scene-based audio data (such as high-order high-fidelity stereo HOA coefficients), correlation analysis is performed to determine the ordering pairs of correlation components from the foreground and background audio signals that will undergo stereo psychoacoustic audio encoding. On the decoding side, before reconstructing the scene-based audio data using the reordered correlation components, the encoded correlation coefficient pairs are decoded using stereo decoding and reordered using reordering information included in the bitstream.
Owner:QUALCOMM INC

Training method of music generation model, music generation method, equipment and medium

The invention provides a training method of a music generation model, a music generation method, equipment and a medium, and relates to the technical field of artificial intelligence. Performing first up-sampling training on the training audio coding information based on a first generator in a vocoder in a music generation model to obtain intermediate training audio data of a first sampling rate, the intermediate training audio data is subjected to first up-sampling training based on a first generator in the vocoder, second up-sampling training is carried out on the intermediate training audio data based on a second generator in the vocoder to obtain training audio data of a second sampling rate, the second sampling rate is larger than the first sampling rate, and generative adversarial training is carried out on the vocoder based on the training audio data to obtain a trained music generation model. By adopting the method and the device, the dependence of model training on high-sampling-rate audio data can be reduced while the music generation model is ensured to output high-fidelity songs, so that the model training cost is controlled.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Front cover for loudspeaker, loudspeaker module and mobile terminal

The utility model belongs to the technical field of mobile communication equipment, and particularly relates to a front cover for a loudspeaker, a loudspeaker module and a mobile terminal. The front cover for the loudspeaker comprises a metal sheet, the metal sheet is provided with a via hole, and the surface of the metal sheet is provided with a silica gel ring. According to the utility model, a better sound transmission effect of the loudspeaker can be realized, the sound transmission obstruction is small, the sound can clearly and naturally pass through the loudspeaker, the sound output of audio equipment can maintain high fidelity, and the loudspeaker can have stable performance under different frequencies; moreover, good damping characteristics are achieved, appropriate damping adjustment can be carried out on sound, reflection and scattering of the sound are reduced, distortion is reduced, and the sound quality is improved; the loudspeaker module also has excellent protection performance, realizes effective water and dust prevention, and can effectively prevent water, dust and other small particles from entering the loudspeaker module; and the anti-impact performance is higher, good protection is provided for internal parts, and damage caused by external force collision is reduced.
Owner:DONGGUAN FUYING ELECTRONIC MATERIALS CO LTD

Active cancellation of a height-channel soundbar array's forward sound radiation

A multi-driver multi-channel single enclosure Height-Channel (e.g., ATMOS™ or DTS-X®) enabled soundbar loudspeaker system 260 uses a novel signal processing system, driver mounting configuration (310L, 310R) and method to provide a high fidelity home theater listening experience, in a manner which relies on a new method for cancellation of unwanted direct (not ceiling-bounced) radiation of the Height-Channel (or virtual height envelopment) channel's sound 213DS.
Owner:POLK AUDIO LLC

Partition purification power supply management time sequencer for high-fidelity sound system

The invention discloses a partition purification power supply management time sequencer for a high-fidelity sound system, and the time sequencer comprises a relay board, the relay board is provided with a first power supply purification circuit and a second power supply purification circuit, two differential mode capacitors in the second power supply purification circuit are in short circuit, and the second power supply purification circuit is used for supplying power to a post-stage power amplifier. The two differential mode capacitors in the first power supply purification circuit are electrically connected through the common mode inductor, and the first power supply purification circuit is used for supplying power to the access turntable, the multicast, the decoding or the pre-stage power amplifier. By optimizing the filter circuit, the requirements of other sound equipment except the post-stage power amplifier on the current purity can be met, and the phenomena of'soft ', 'weak' and sound lag of bass in the use process of the post-stage power amplifier can be avoided; therefore, the sound heard by a user of the high-fidelity sound system is pure in sound quality and vigorous and powerful in bass, and is rich in impact force and infectivity.
Owner:河南广播电视台

Hybrid rendering

An apparatus includes a memory configured to store first audio data and second audio data. The device also includes one or more processors coupled to the memory and configured to determine priorities of audio sources of the audio scene. The one or more processors are further configured to render the first audio data using the object renderer to generate a first audio signal. The first audio data represents a first audio source associated with a first priority. The one or more processors are further configured to render the second audio data using the first Ambisonics renderer to generate a second audio signal. The second audio data represents a second audio source associated with a second priority.
Owner:QUALCOMM INC

Wi-Fi streaming media high-fidelity audio real-time transmission system and method adopting dynamic code rate adaptation

The invention relates to a Wi-Fi streaming media high-fidelity audio real-time transmission system and a Wi-Fi streaming media high-fidelity audio real-time transmission method adopting dynamic code rate adaptation. The system comprises a transmitting end and a receiving end, wherein the transmitting end realizes layered encoding and packaging of audio signals through an audio acquisition module, a layered encoder, a forward complexity analyzer and a data packaging module; and a receiving end completes data analysis and low-power-consumption decoding through a data analyzer, a heterogeneous decoding processor and a power consumption scheduling manager. A dynamic code rate control mechanism based on a network state and a buffer area state is adopted, and a power consumption pre-scheduling strategy of forward complexity analysis is combined, so that collaborative optimization of code rate self-adaptive adjustment of a sending end and dynamic configuration of processing resources of a receiving end is realized. The technical problem that sound quality, delay and power consumption are difficult to consider in traditional wireless audio transmission is effectively solved, the power consumption of the receiving end is remarkably reduced while high-fidelity sound quality and low transmission delay are guaranteed, the method is particularly suitable for portable equipment such as earphones, and the endurance time of the portable equipment can be greatly prolonged.
Owner:HEAD DIRECT (KUNSHAN) CO LTD

Game earphone receiver

The utility model discloses a game earphone receiver, and relates to the technical field of receivers, the game earphone receiver comprises a support, the support is provided with six outer ring holes, the six outer ring holes are uniformly distributed on the support in a ring shape, the inner sides of the six outer ring holes are provided with ten inner ring holes, and the inner ring holes are communicated with the inner ring holes. The ten inner ring holes are evenly distributed in the support in an annular mode, a tuning plate is arranged at the bottoms of the ten inner ring holes and detachably connected with the support, and an installation opening is formed in the center of the support. According to the unique tuning design, double-layer tuning holes of the support are matched with York middle holes and tuning mesh cloth, airflow is dispersed and circulated, fine sound is restored, game sound details are presented, the requirements of game players are met, the noise reduction function is achieved for development of the game players, and the players concentrate on games, wear comfortably and enjoy the high-fidelity sound effect. And novel environment-friendly UV glue is used, so that the efficiency is improved, and the production period is shortened.
Owner:音品电子(深圳)有限公司

Vehicle-mounted multi-channel wireless sound system and vehicle

The utility model belongs to the audio signal transmission technical field, concretely relates to a kind of vehicle-mounted multi-channel wireless sound system and vehicle, it include: multiple loudspeaker, and: control module, control module includes main control control module and controlled control module;UWB component, including main control UWB component and controlled UWB component, main control UWB component is electrically connected with main control control module, controlled UWB component is electrically connected with controlled control module, each controlled control module is electrically connected with one or more loudspeaker.The utility model solves the problem that vehicle internal high fidelity audio signal transmission is limited by wiring harness cost, utilizes UWB transmission technology while high fidelity transmission audio signal, avoids complex audio line wiring, and single control loudspeaker can be formed multiple sound fields in vehicle.
Owner:重庆云辉新能源科技有限公司

Industrial heritage digital display system

The invention, which relates to the technical field of industrial heritage display, discloses an industrial heritage digital display system comprising a cross-modal time sequence observation module, a causal coherent decomposition module, an anti-factual shadow restoration module and a double-mirror-image time scale control module. And the cross-modal time sequence observation module is used for constructing a cross-modal time sequence integrity observation layer based on the dynamic refreshing tracks of the three-dimensional model rendering data and the acoustic playback data, and continuously acquiring a rendering delay drift curve of the three-dimensional model rendering data and an acoustic phase deviation curve of the acoustic playback data. Through cross-modal time sequence observation and causal decomposition, accurate synchronization of three-dimensional rendering data and acoustic playback data is realized, sound picture dislocation caused by dynamic phase drift is eliminated, and the authenticity and immersion of display are improved; and meanwhile, through anti-fact shadow repair and double-mirror-image time scale control, self-adaptive time sequence repair and dynamic closed-loop adjustment are realized, long-term stable synchronization of sound and pictures is ensured, and high fidelity and continuity of industrial heritage display are ensured.
Owner:KUNMING UNIV OF SCI & TECH

Photoelectric artificial cochlea

The invention discloses a photoelectric artificial cochlea, which comprises an extracorporeal machine and an implant, and is characterized in that the extracorporeal machine is configured to receive an original sound signal, process the original sound signal into a first analog signal or a conversion signal of the first analog signal, and transmit the first analog signal or the conversion signal to the implant; the implant comprises a stimulator, the stimulator comprises a photoelectric hybrid chip, the photoelectric hybrid chip is configured to receive an input signal and modulate a carrier optical signal by using the input signal to output a modulated carrier optical signal, the modulated carrier optical signal is an analog signal, and the input signal is a first analog signal or a conversion signal of the first analog signal; the implant comprises an optical fiber electrode, and the optical fiber electrode is implanted from the round window of the human ear, is arranged in the drum step, extends from the cochlea bottom to the cochlea top and directly acts on the auditory neuron, so that the high-frequency component of the modulated carrier optical signal is preferentially demodulated at the cochlea bottom, and the low-frequency component of the modulated carrier optical signal is preferentially demodulated at the cochlea top. Therefore, the photoelectric artificial cochlea system which is low in power consumption, high in fidelity and accurate in directional stimulation is realized.
Owner:SHANGHAI MISTAR MEDICAL TECH CO LTD

High fidelity speech synthesis using adversarial networks

The invention relates to high fidelity speech synthesis with adversarial networks. Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating output audio examples using a generative neural network. One of the methods includes obtaining a training adjustment text input; processing a training generation input comprising the training adjustment text input using a feedforward generation neural network to generate a training audio output; processing the training audio output using each of a plurality of discriminators, where the plurality of discriminators includes one or more conditional discriminators and one or more unconditional discriminators; determining a first combined prediction by combining the respective predictions of the plurality of discriminators; and determining an update to a current value of a plurality of generation parameters of the feed-forward generation neural network to increase a first error in the first combined prediction.
Owner:DEEPMIND TECH LTD

Hi-fi micro system

ActiveCN224418906UGreat restorationgood hi-fiHigh fidelitySound box
The utility model discloses a HI -FI micro system sound, including face shell, top shell, wood top bottom plate and bottom shell, bottom shell is fixed on wood top bottom plate, top shell covers the top of wood top bottom plate, and the both sides of top shell are locked on bottom shell, and the both sides of bottom shell all are equipped with tray support, and the fixed mainboard support of tray support has the inside slide connection of core, and the front end of top shell is equipped with the screen PCB, and face shell covers the outside of screen PCB, and the both sides of top shell are provided with wood sound box, and the front end of wood sound box is equipped with wood box front plate, and the wood box front plate is covered with the sound box cloth net frame. This HI -FI micro system sound is used to play the similar sound height of original, and receives CD and enters the core, releases the sound through the wood sound box of both sides, greatly restores the sound, and the high fidelity effect is good.
Owner:DONGGUAN JINWEIJU TECH CO LTD

Lightweight voice band extension method and device for edge device, terminal and medium

The edge device-oriented lightweight voice band expansion method, device, terminal and medium provided by the application belong to the technical field of voice signal processing, and the method comprises the following steps: obtaining a logarithmic domain narrowband audio signal amplitude spectrum and a logarithmic domain mixed amplitude spectrum; constructing a white noise amplitude spectrum; inputting the logarithmic domain narrowband audio signal amplitude spectrum, the logarithmic domain mixed amplitude spectrum and the white noise amplitude spectrum into a trained voice bandwidth expansion model to generate a first high-frequency component corresponding to a consonant in the narrowband audio signal and a second high-frequency component corresponding to a vowel in the narrowband audio signal, and then obtaining a predicted amplitude spectrum; expanding phase information of the narrowband audio signal, generating a wideband audio signal according to the predicted amplitude spectrum and the expanded phase information of the narrowband audio signal, and outputting the wideband audio signal. The voice bandwidth expansion model is used to reconstruct the consonant, so that the reconstructed consonant component has higher fidelity, and the intelligibility of the reconstructed voice is ensured.
Owner:ELEVOC TECH CO LTD

Digital human image generation method combining cold start driving and active learning mechanism

The application discloses a digital human image generation method combining cold start driving and active learning mechanism. The preset few-sample emotional speech generation model derived from the cold start process and having high personalization and dynamic adaptability is used to process input data, efficient personalized and multi-emotional speech batch generation can be realized on large-scale unlabeled input data, and a target audio file with high fidelity and consistent emotional and personal expression can be output. The preset few-sample emotional speech generation model is trained based on the first qualified sample obtained by screening the candidate training sample by using the first cold start quality evaluator, so that the model can be started under the condition of very few data, and high-quality personalized digital human image automatic generation is realized under the condition of small samples. In addition, by inputting a target emotion embedding vector and a target personality feature vector, expression synchronous generation under speech driving is realized, and the natural interaction ability and style consistency of the digital human are enhanced.
Owner:SHAANXI JIEQI NETWORK TECHNOLOGY CO LTD