Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41 results about "Quality voice" patented technology

Voice quality is that component of speech which gives the primary distinction to a given speaker's voice when pitch and loudness are excluded. It involves both phonatory and resonatory characteristics. Some of the descriptions of voice quality are harshness, breathiness and nasality.

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Service compliance dynamic evaluation system based on real-time voice recognition

The invention relates to a service compliance dynamic evaluation system based on real-time voice recognition, and belongs to the cross technical field of artificial intelligence and service compliance management. The system comprises a voice signal enhancement acquisition unit, a voice analysis and risk initial judgment unit, a dynamic evaluation iteration unit and a compliance risk disposal unit. The voice signal enhancement acquisition unit processes low-quality voice and generates a voice data set; the voice analysis and risk initial judgment unit is used for transferring voice in real time, extracting a dialogue intention and identifying violation contents and interactive emotions; the dynamic evaluation iteration unit outputs compliance scores and violation details and is in butt joint with a manual labeling platform; and the compliance risk disposal unit generates a risk label, carries out linkage training and appeal review, and forms a compliance risk disposal scheme. According to the invention, full quality inspection of service scenes in multiple industries such as finance, operators and the like is realized, the voice transfer accuracy and compliance rectification rate are improved, the verbal skill violation complaint rate is reduced, and dynamic monitoring and closed-loop disposal of service compliance are achieved.
Owner:SHANGHAI RONGDA DIGITAL TECH CO LTD

Voice noise reduction method and device for multi-person scene, electronic equipment and storage medium

The invention discloses a voice noise reduction method and device for a multi-person scene, electronic equipment and a storage medium, and relates to the technical field of voice signal processing. The method comprises the following steps: acquiring an audio signal and a video image of a target space scene; determining audio perception information matched with the audio signal and a visual voice activity detection result corresponding to the video image; fusing the audio perception information and the visual voice activity detection result, and determining a target voice sounding object when the voice activity fusion result indicates that the voice activity exists; and positioning a human face corresponding to the target voice sounding object to obtain position change information, updating a pickup direction, controlling beam forming processing, and outputting a denoised target voice signal. According to the scheme provided by the invention, the voice of the current effective speaker can be accurately separated and enhanced from the aliasing audio signals, meanwhile, the interference of other speakers and environmental noise is inhibited, and help is provided for realizing high-quality voice interaction and recognition.
Owner:ANHUI JISHEN YINSAI TECHNOLOGY CO LTD

Controllable voice generation method and device based on multi-agent dynamic scheduling

The invention discloses a controllable voice generation method and device based on multi-agent dynamic scheduling, and belongs to the technical field of voice synthesis, and the method comprises the steps: constructing a multi-agent voice generation frame comprising a central scheduling module, an identity agent, an emotion agent and an environment agent; a central scheduling module analyzes a user instruction and outputs a structured task plan to drive each agent to generate primary voice output of identity, emotion and environment dimensions, and a cooperation cost matrix between the agents is constructed according to the primary voice output; the optimal execution path of the multiple agents is solved through an optimization algorithm based on the cost matrix, the agents are executed in a cascading mode based on the optimal execution path, and finally sound mixing is completed and high-quality voice is output. According to the method, the naturalness, the semantic consistency and the overall quality of the synthesized voice in a complex scene can be remarkably improved, and the method is suitable for various man-machine interaction application scenes such as intelligent voice assistants, virtual digital humans, immersive entertainment, barrier-free voice services, personalized content creation and the like.
Owner:ZHEJIANG UNIV OF TECH

Extremely low rate voice communication method and related device

The invention provides an extremely low rate voice communication method and related equipment. The method comprises the following steps: extracting a first feature and a second feature from acquired voice information; the first feature reflects semantic information of the voice information; the second feature reflects the acoustic feature of the voice information; calculating an entropy value of variational posteriori distribution of the first feature, and adjusting the first feature according to the entropy value to obtain a third feature; performing semantic quantization on the third feature to obtain a first index; performing acoustic quantization on the second feature to obtain a second index; and gradually restoring details of the voice information based on the first index and the second index so as to complete transmission of the voice information. According to the embodiment of the invention, through semantic and acoustic feature fusion and entropy modulation and quantization, the compression efficiency is improved, the bit demand is reduced, and high-quality voice transmission and restoration are realized.
Owner:BEIJING UNIV OF POSTS & TELECOMM +1

Driving emotion early warning method and system based on voice analysis

The invention provides a driving emotion early warning method and system based on voice analysis, relates to the technical field of voice emotion recognition, and effectively solves the interference problem of vehicle-mounted dynamic strong noise by deploying a microphone array and a noise sensor in a cockpit and combining multi-channel noise reduction algorithms such as beam forming and linear beam minimum variance. A high-quality voice instruction is extracted from a source; on the basis, the driving intention of the clear voice signal is accurately recognized by using the deep neural network, so that the accuracy is remarkably improved; more importantly, according to the method, the instability Ls and the high interactive friction probability Pf are recognized innovatively through calculation instructions, real-time quantitative evaluation on the man-machine interaction fluency is achieved, when the interactive friction probability Pf exceeds a preset threshold value, the system can automatically trigger self-adaptive strategy adjustment, and therefore the situation that a driver is distracted due to recognition errors is avoided, and user experience is improved. And the interaction experience and the driving safety are greatly enhanced.
Owner:FUJIAN LUOYUAN COUNTY SENIOR VOCATIONAL HIGH SCHOOL

Sound processing system and device

The utility model discloses a sound processing system, comprising a pickup module used for picking up the sound of a user; the playback module is used for converting the electric signal into a sound signal and playing the sound signal; the information processing module is used for receiving and processing the signal of the pickup module and outputting the enhanced voice signal to the playback module; and the power management module is used for supplying power to the sound processing system and comprises a direct current-direct current charging management unit and an overvoltage and overcurrent protection device. By means of the mode, on the premise that normal oxygen uptake of a patient is guaranteed, barrier-free communication between the patient and the outside world is achieved, voice enhancement is achieved by combining noise elimination, echo suppression, howling suppression, gain control and the like, and high-quality voice transmission is achieved.
Owner:SUZHOU WENLEXING TECHNOLOGY CO LTD

Voice data loss recovery processing method, device and equipment and storage medium

The application discloses a voice data loss recovery processing method and device, equipment and a storage medium, and relates to the technical field of voice transmission. The method comprises the following steps: performing frame processing on received voice message data to obtain frame voice data; when signal loss is detected, performing voice reconstruction on the frame voice data based on compressed sensing to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data. The voice message data is first subjected to frame processing, the voice is reconstructed based on compressed sensing by using the sparsity of the voice signal when the signal is lost, and then dynamic filtering is performed through the preset filtering network, so that high-quality voice data is finally output. The method solves the problems of bandwidth waste caused by redundant transmission and poor effect of traditional interpolation filtering in the prior art, realizes the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality, and meets the demand for high-quality voice communication.
Owner:SHENZHEN DINSTAR TECH

Voice command-based processing method, device, and system

The embodiments of the present disclosure disclose a processing method, device and system based on voice instructions. A specific implementation of the method includes: recognizing voice data, and synchronously sending the above-mentioned voice data to the terminal and the cloud; in response to receiving the user intention information sent by the above-mentioned terminal or the above-mentioned cloud or identifying the user intention information, executing the preset operation corresponding to the above-mentioned user intention information, wherein the user intention information sent by the above-mentioned terminal is obtained by the above-mentioned terminal recognizing the received voice data, and the user intention information sent by the above-mentioned cloud is obtained by the above-mentioned cloud recognizing the received voice data; in response to recognizing the user intention information, sending a cancellation signal to the above-mentioned terminal and the above-mentioned cloud, wherein the above-mentioned cancellation signal is used to interrupt the task corresponding to the above-mentioned voice data in the above-mentioned terminal and the above-mentioned cloud. This implementation method achieves higher-quality voice instruction processing through multi-terminal recognition and collaboration, thereby improving user experience and system practicality.
Owner:HANGZHOU LINGBAN TECH CO LTD

Voice cloning method and system

The invention provides a voice cloning system and method, and the system comprises a data preprocessing module, a feature extraction module, a model training module, and a reasoning generation module, and the data preprocessing module is used for carrying out the human voice accompaniment separation, audio cutting, and automatic marking of an original audio; the feature extraction module is used for extracting a self-encoding feature and a Mel spectrum feature from the audio and pairing the self-encoding feature and the Mel spectrum feature with the text length; the model training module is used for performing fine tuning training on the basis of a small amount of reference audio data by combining a GPT-like model and a VITS and Valle model architecture; and the reasoning generation module is used for generating target audio according to the reference audio and the text provided by the user. According to the method and the device, high-quality voice cloning and text-to-voice conversion can be realized only by a small amount of data sets, so that the user experience is improved, and the development cost is also reduced.
Owner:SHANGHAI MANJU NETWORK TECHNOLOGY CO LTD

An earphone box, a wireless communication device and a method of using the same

This application provides an earphone case, a wireless communication device, and a method for using the same. The earphone case includes a case body, a cover assembly, and a magnetic component. The cover assembly includes a lid, a microphone module, and a playback module. The lid can be closed and detached from the case body. The playback module is used to play sound. At least a portion of the lid is metal. The magnetic component is detachably mounted to the case body and can fix the lid when the cover assembly is separated from the case body, allowing the cover assembly to function as an independent module for both microphone and playback. In this embodiment, the earphone case not only stores wireless earphones but also allows the cover assembly to function independently as an earphone. During two-way communication, one person wears the wireless earphones while the other wears the cover assembly with the help of the magnetic component, reducing the distance limitation for communication and enabling high-quality voice interaction.
Owner:1MORE ACOUSTIC TECH CO LTD

Ultra-low delay signal sequence conversion method and system

The invention provides an ultra-low time delay signal sequence conversion method and system, which can be widely applied to various time sequence conversion scenes, including hearing enhancement systems and equipment, such as hearing aid, denoising, simultaneous translation, sound conversion signal sequence, intelligent glasses, brain-computer systems and the like. According to the invention, the rapid high-fidelity noise elimination method is adopted, so that the limitation and difficulty in the traditional technology can be overcome. And for the problems of noise uncertainty and voice distortion in the sound related field, the high-quality voice can be recovered by recognizing the voice content as the intermediate language representation with high generalization, and meanwhile, the voice of the same speaker is synthesized by using a pre-trained artificial intelligence module. And the ultra-low time delay is realized by using a variable multi-partition block and a multi-head prediction mechanism in the process. Therefore, in the specific hearing enhancement field, the voice of the target speaker can be focused by using the voice separation technology, and the uncertainty of background noise is avoided.
Owner:SHANGHAI PEDAWISE INTELLIGENT TECH CO LTD

A microphone for classroom use

This utility model relates to the field of audio acquisition equipment technology and discloses a microphone for classrooms, including a microphone body comprising a housing. This microphone for classrooms constructs a highly efficient dual noise reduction system by incorporating a noise-reducing cover and composite sound-absorbing cotton. The noise-reducing cover tightly wraps around the housing, utilizing its material properties to block most external environmental noise, such as noise from outside the window and footsteps outside the classroom. The composite sound-absorbing cotton, filled between the housing and the noise-reducing cover, further absorbs residual noise penetrating the cover while reducing noise generated by vibrations of internal components. This design allows the microphone to effectively filter out non-target sounds in complex classroom environments, significantly improving the clarity of target sounds such as teachers lecturing and students answering questions. It avoids the audio blurring problem caused by noise interference in traditional microphones, providing higher-quality sound materials for classroom recording, remote teaching, and other scenarios.
Owner:YUNNAN ZHILAN CLOUD PIGEON INFORMATION TECH CO LTD

High-quality voice signal processing device and method through removal of ambient noise based on multi-sensor signal fusion

A high-quality voice signal processing device through removal of ambient noise based on multi-sensor signal fusion, includes: a voice microphone sensor that senses and outputs a speaker's voice signal; an accelerometer sensor that senses vibration of the speaker's vocal cords and outputs a signal; a noise reduction processing MCU that extracts a voice section according to vocal cord vibration using the output signal of the accelerometer sensor, synthesizes a low-frequency component of the accelerometer sensor and a low-frequency component of the voice microphone sensor at different synthesis ratios based on a level of noise extracted from the output signal of the voice microphone sensor using voice section information, and restores and outputs a voice signal by adding the synthesized low-frequency components and a high-frequency component of the voice microphone sensor; and a wireless communication module that externally outputs the restored voice signal.
Owner:INTUS CO LTD

Signal generation processing device

The present application realizes a signal generation processing device that realizes a voice synthesis processing or an image signal generation processing that can maintain the speed of the voice synthesis processing or the image signal generation and obtain a high-quality voice signal or an image signal. In the signal generation processing device, first to Nth sub-model sections respectively perform learning processing of learning models included in the first to Nth sub-model sections using noise levels included in different noise level ranges, thereby acquiring learned models. That is, in the signal generation processing device, processing can be performed in parallel for each sub-model section, and as a result, the learning processing can be performed at high speed. In addition, in the signal generation processing device, at the time of prediction processing, a sub-model section to be used for processing can be appropriately selected, and thus a high-precision voice synthesis processing or an image generation processing can be performed.
Owner:NAT INST OF INFORMATION & COMM TECH

Distributable ai voice upscaling

A method for distributable upscaling of audio signals includes receiving, over a communication channel by an electronic device of a first user, a low quality voice communication from a second user. The method includes accessing an artificial intelligence (“AI”) voice upscaling model of the second user. The AI voice upscaling model is trained on a voice of the second user. The method includes using the AI voice upscaling model to improve the quality of the low quality voice communication to create a higher quality voice communication of the second user. The method includes transmitting the higher quality voice communication to a speaker connected to the electronic device of the first user.
Owner:LENOVO ENTERPRISE SOLUTIONS (SINGAPORE) PTE LTD

An embedded audio processing system and method supporting AI language enhancement

The application relates to the technical field of audio processing technology, in particular to an embedded audio processing system and method supporting AI language enhancement, which comprises a data acquisition module, an audio analysis module, an AI enhancement module, an audio processing module and an interactive output module. The modules of the system cooperate, the data acquisition module collects audio, the analysis module accurately classifies and calibrates, the enhancement module improves the quality of fuzzy audio, the processing module generates corresponding content, the output module outputs, accurate, efficient and high-quality voice interaction is realized, the speed of vehicle-mounted AI voice answering voice questions is improved, and the accuracy of vehicle-mounted AI voice answering voice questions is improved.
Owner:BEIJING HANGYU XINGZHOU TECHNOLOGY CO LTD

Locomotive CIR host supporting 5G broadband voice AMR-WB dispatch communication and dispatch communication method

This invention relates to a locomotive CIR host and scheduling communication method supporting 5G broadband voice AMR-WB scheduling communication. The CIR host is implemented by replacing the main control unit on an existing CIR host. The CIR host includes a backplane, and the main control unit includes a base plate and a core board. The core board is mounted on the base plate, which is connected to the backplane. The base plate has an Ethernet port, and the CIR host connects to a 5G communication module via the Ethernet interface to achieve 5G broadband voice AMR-WB scheduling communication. Compared with existing technologies, this invention offers advantages such as providing higher quality voice communication for locomotives, low modification costs, convenience, and strong versatility.
Owner:CRSC COMM & INFORMATION

Desktop operation and maintenance integrated machine

ActiveCN309547457SMachineSpeech sound
1. The name of the design product: desktop operation and maintenance all-in-one machine. 2. The use of the design product: for remote maintenance assistance, realizing high-quality voice connection with remote service center. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view. 5. The bottom view is not a common view, and the bottom view is omitted.
Owner:AKSU ZHENGHENG INFORMATION TECHNOLOGY CO LTD

Voice coding and decoding method based on principal component analysis and multi-scale depth attention

The invention discloses a voice coding and decoding method based on principal component analysis and multi-scale depth attention, and relates to the technical field of voice signal processing, and the method comprises the steps: carrying out the multi-scale depth attention convolution operation of a voice signal for many times, and obtaining a first feature; wherein after feature extraction and feature fusion are respectively carried out through a plurality of parallel depth separable convolutions with different depth convolution kernels, channel-by-channel multiplication is carried out on fused features according to channel attention weights; performing projection processing on the first feature according to the projection matrix and the feature mean value to obtain a second feature; wherein principal component analysis is carried out on first features of a training set, a feature mean value is determined, and a previous feature vector with the highest cumulative variance contribution rate is selected to form a projection matrix; performing multi-stage vector quantization on the second feature in sequence to obtain a reconstructed vector; and carrying out inverse projection processing on the reconstructed vector, and then decoding and outputting reconstructed voice. And ultra-low-bit-rate and high-quality voice communication is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

A method, device, computer equipment and readable storage medium for implementing cloud-rendered voice interactive digital human

The present invention discloses a method, apparatus, computer device, and readable storage medium for implementing a cloud-rendered voice-interactive digital human. The method includes: first, constructing a voice-interactive digital human, performing audio lip synchronization and anthropomorphic configuration, and then placing the human on a front-end page via real-time cloud rendering. After obtaining the user's front-end voice interaction input, the method generates the digital human's interactive feedback based on the previously configured input, and then renders this feedback to the front-end page in real-time via the cloud. This method not only meets the demand for high-quality voice-interactive digital humans across various industries, but also leverages cloud rendering technology to reduce user device requirements and enhance the user experience.
Owner:DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD

Vehicle-mounted voice endpoint detection method and device, equipment and storage medium

The invention provides a vehicle-mounted voice endpoint detection method, device and equipment and a storage medium, and belongs to the technical field of vehicle-mounted voice interaction, and the method comprises the steps: collecting a digital voice signal; performing framing processing and adaptive noise elimination on the digital voice signal, and marking a voice frame and a non-voice frame in a current frame; marking a voice starting endpoint and a voice ending endpoint in the voice frame; dividing the continuous voice frames into one or more voice segments, and adjusting the dynamic mute tolerance time and the voice ending endpoint; updating a voice starting endpoint and a voice ending endpoint in the voice frame based on the non-voice frame; and extracting a corresponding voice segment according to the voice starting endpoint and the voice ending endpoint in the updated voice frame. Through the technical scheme in the embodiment of the invention, the starting endpoint and the ending endpoint of the voice can be accurately and stably detected, and high-quality voice segment input is provided for vehicle-mounted voice recognition.
Owner:DONGFENG MOTOR GRP

Voice data loss recovery processing method and device, equipment and storage medium

The invention discloses a voice data loss recovery processing method, device and equipment and a storage medium, and relates to the technical field of voice transmission, and the method comprises the steps: carrying out the framing processing of received voice message data, and obtaining the framed voice data; when detecting that the signal is lost, performing voice reconstruction based on compressed sensing on the framed voice data to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data. According to the method, the voice message data is framed, the voice is reconstructed based on compressed sensing by using the sparsity of the voice signal when the signal is lost, then dynamic filtering is performed through the preset filtering network, and finally the high-quality voice data is output. The problems of bandwidth waste caused by redundant transmission and poor traditional interpolation filtering effect in the prior art are solved, the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality are achieved, and the high-quality voice communication requirement is met.
Owner:SHENZHEN DINSTAR TECH

A dialect speech recognition method and system in a vehicle-mounted scenario

The present invention discloses a dialect speech recognition method and system in an in-vehicle scenario, belonging to the field of speech recognition technology. In response to the challenges of noise, dialect diversity, and speaker overlap in the complex environment of the vehicle, a systematic solution from data acquisition, data preprocessing to model training and decoding is proposed. In the data acquisition stage, by arranging far-field microphones, wearing near-field microphones, and recording the noise in the vehicle, a variety of high-quality voice data is obtained, and text and speaker timestamps are annotated. The present invention significantly improves the performance of the speech recognition system in complex in-vehicle scenarios and is suitable for practical application scenarios such as in-vehicle navigation and voice assistants in the case of dialects.
Owner:TIANJIN UNIV

Device and Method for Processing High Quality Voice signal Using Removing Ambient Noise based on Multi Sensor Signal Fusion

ActiveKR102993224B1Noise levelNoise
The present invention relates to a high-quality voice signal processing device and method through ambient noise removal based on multi-sensor signal fusion, which enables voice signal processing robust to external noise environments through multi-sensor signal fusion using an accelerometer sensor (ACC) and a voice microphone sensor (MIC). The device comprises: a voice microphone sensor (MIC) that senses and outputs a voice signal of a speaker; an accelerometer sensor (ACC) that detects vocal cord vibration of a speaker and outputs a signal; a noise reduction processing MCU that extracts a vocalization segment based on vocal cord vibration using the output signal of the accelerometer sensor (ACC), synthesizes the low-frequency component of the accelerometer sensor (ACC) and the low-frequency component of the voice microphone sensor (MIC) by varying the synthesis ratio based on the noise level extracted from the output signal of the voice microphone sensor (MIC) using the vocalization segment information, and restores and outputs a voice signal by adding the synthesized low-frequency component and the high-frequency component of the voice microphone sensor (MIC); and a wireless communication module that outputs the restored voice signal externally.
Owner:INTUS CO LTD

Earphone box and wireless communication equipment

The embodiment of the utility model provides an earphone box and wireless communication equipment, the earphone box comprises a box body, a cover body assembly and a magnetic part, the cover body assembly comprises a box cover, a sound receiving module and a sound playing module, and the box cover can cover and separate from the box body; the playback module is used for playing sound; at least part of the box cover is a metal part, the magnetic part and the box body are detachably arranged, and the magnetic part can fix the box cover when the cover body assembly is separated from the box body, so that the cover body assembly serves as an independent module to achieve sound receiving and playing. According to the earphone box provided by the embodiment of the invention, while the earphone box is used for accommodating the wireless earphone, the cover body assembly can be used as the earphone to work independently. When communication between two parties is carried out, one person wears the wireless earphone, and the other person wears the cover body assembly through cooperation of the magnetic part, so that limitation of the communication distance between the two parties is reduced, and high-quality voice interaction is realized.
Owner:1MORE ACOUSTIC TECH CO LTD

Text-to-voice method and device, computer equipment and storage medium

The invention relates to the field of artificial intelligence, is applied to financial and medical scenes, and discloses a text-to-voice method and device, computer equipment and a storage medium, and the method comprises the steps: receiving a target text, a voice prompt and an emotion label; performing semantic coding on the target text to extract text semantic features; performing acoustic coding on the voice prompt to extract voice acoustic features; based on the emotion label and the text semantic feature, generating a dynamic emotion feature through a time sequence emotion model; performing cross-modal fusion on the voice acoustic features and the dynamic emotion features to obtain cross-modal alignment features, and inputting the cross-modal alignment features and the text semantic features into a pre-training language model for feature fusion and voice token prediction to obtain an initial voice token sequence; and performing semantic alignment and spectrum conversion on the initial voice token sequence to generate a target voice waveform. According to the invention, high-quality voice with natural emotional circulation, controllable timbre and high conformity with professional contexts can be synthesized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Mobile device and self-noise elimination method thereof

The invention provides a mobile device and a self-noise elimination method thereof, and relates to the technical field of denoising. The method comprises the following steps: collecting motion state data of the mobile device in real time; inputting the motion state data into a noise feature pre-judgment model to obtain noise feature parameters output by the noise feature pre-judgment model; the noise characteristic parameters comprise a target amplitude, a target frequency and a target phase of self-noise generated by the mobile device; and based on the noise characteristic parameter, generating a reverse counteracting signal, and controlling a reverse counteracting component of the mobile device to output the reverse counteracting signal. According to the invention, the self-noise is actively offset from the source through hardware, the interference of complex multi-source self-noise is reduced in a real-time and preposed manner, and a solid foundation is laid for high-quality voice interaction of a mobile device in a severe operation environment, that is, high-efficiency and high-quality self-noise elimination is realized, and high-efficiency and high-quality voice interaction is ensured.
Owner:IFLYTEK CO LTD

Method for automatically generating lecturer video based on AI speech synthesis and animation driving

The application discloses a kind of based on AI speech synthesis and animation driving's lecturer video automatic generation method, comprising: by improved inner point method and incremental shortest path algorithm to PPT or text script is structured and analyzed;Application full dynamic parallel single link clustering algorithm carries out semantic grouping and generates enhanced script with expressive mark;High-quality voice data stream is generated by using CosyVoice technology combined with low-rank approximation method;The mapping relationship between content and action expression is established by semantic analysis, and a complete set of action expression instructions is generated;The museTalk technology is used to drive the digital human model, and the final lecturer teaching video is generated by using parallel rendering algorithm.The application realizes the efficient and automatic generation of educational videos, significantly improves the content production efficiency, reduces the production cost, and ensures the professionalism and expressiveness of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Embedded audio processing system and method supporting AI language enhancement

The invention relates to the technical field of audio processing, in particular to an embedded audio processing system and method supporting AI language enhancement, and the system comprises a data collection module, an audio analysis module, an AI enhancement module, and an audio processing module. And an interactive output module. The modules of the system cooperate, the data acquisition module collects audios, the analysis module accurately classifies and calibrates the audios, the enhancement module improves the fuzzy audio quality, the processing module generates corresponding contents, and the output module outputs the contents, so that accurate, efficient and high-quality voice interaction is realized, and the speed of answering voice questions by vehicle-mounted AI voice is improved. And the accuracy of answering the voice question by the vehicle-mounted AI voice is improved.
Owner:BEIJING HANGYU XINGZHOU TECHNOLOGY CO LTD