Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97 results about "Speech quality" patented technology

Measuring Speech Quality. In telecommunication, speech quality is an important contributing factor to the success of a product and to the success of the communication itself. High speech quality guarantees that the effort the users have to put forward in order to correctly perceive the communication is low.

Far-field single-channel speech enhancement method

The invention relates to the technical field of speech enhancement, in particular to a far-field single-channel speech enhancement method based on an MFSE (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error) (Maximum Free Square Error)). The method comprises the following steps: step 1, processing a far-field voice signal to obtain a complex spectrogram of a noise voice signal; 2, inputting the compressed complex spectrogram into a feature encoder, processing the output of the complex spectrogram by N MamAttention blocks, and then sending the processed complex spectrogram into an amplitude mask decoder and a phase decoder to respectively predict a clean compressed amplitude mask and a phase spectrum; step 3, preheating and training the MamAttention model, and performing supervised confrontation training by taking the MamAttention model as a generator and the multi-resolution discriminator as a discriminator; and step 4, inputting test voice into the trained model to realize far-field single-channel voice enhancement. According to the method, the supervised adversarial training strategy and the MamAttention model are combined, so that the problems of signal attenuation, noise and reverberation interference in far-field voice are effectively solved, and the voice quality is remarkably improved in a scene that the distance of a loudspeaker exceeds 5 meters.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Zero sample speech synthesis method and device, computer equipment and storage medium

The invention relates to a zero-sample speech synthesis method and device, computer equipment, a storage medium and a program product, and the method comprises the steps: obtaining a target coding feature according to a reference speech and a target text; inputting the target coding features into a stream matching model to obtain a conditional velocity field and an unconditional velocity field; inputting the target coding feature into a prior model to obtain a prior speech feature; obtaining a prior generation flow field according to the prior voice features and the standard Gaussian noise; calculating a KL divergence value between the priori generated flow field and a preset real generated flow field, and taking a moment when the KL divergence value is smaller than or equal to a preset KL divergence threshold as an initial moment of the priori generated flow field; fusing the conditional velocity field and the unconditional velocity field to obtain a fused velocity field; and inputting the fusion velocity field from the initial moment to the target moment and the priori generated flow field into an ordinary differential equation solver to obtain the target speech features, thereby improving the speech quality of the synthesized speech.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Speech synthesis style transfer method and device, computer device and storage medium

The present application relates to the field of speech synthesis, and particularly relates to a speech synthesis style migration method and device, computer equipment and a storage medium. The method comprises the following steps: inputting the text data to be converted into a speech synthesis encoder to obtain text encoding; obtaining multi-level speech style representation of the speech to be migrated; inputting the text data to be converted and the multi-level speech style representation into a style predictor to obtain a predicted speech style; performing multi-style layer regularization on the text encoding and the predicted speech style to obtain style-regularized text encoding; inputting the style-regularized text encoding into a pitch predictor to obtain rhythm change data; inputting the rhythm change data and the style-regularized text encoding into a speech synthesis decoder to obtain a mel-frequency spectrum diagram and generate target speech. The present application considers rhythm information and the style of words and phonemes, and the overall rhythm style of the speaker, thereby improving the speech quality of the target speech.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio processing method and device, server and storage medium

The invention provides an audio processing method and device, a server and a storage medium. The method comprises the following steps: receiving a to-be-adjusted original audio and a corresponding line; converting the lines into a corresponding text sequence; performing alignment operation on the original audio and the text sequence to obtain an audio-text corresponding relation; receiving a target text; determining a to-be-replaced first text from the text sequence according to the target text and frame information of the first text in the audio in the corresponding relation; replacing the first text with the target text by using the frame information to obtain an updated text sequence; and generating a new audio by using the updated text sequence. Therefore, the problems that the edited voice is stiff, the complex audio effect is poor, and the quality is possibly poor if the voice is modified too much can be solved, and the quality of the generated voice can be better controlled by controlling the durations of different granularities.
Owner:CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD

Voice processing method and device, equipment, storage medium and program product

The embodiment of the invention provides a voice processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a voice signal; performing signal processing on the voice signal to obtain a first time frequency signal; performing adaptive normalization processing on the first time-frequency signal to obtain a first feature map; performing two-dimensional modeling processing on the first feature map to obtain a second feature map; and performing synthesis processing on the second feature map to obtain an enhanced voice signal. The method can improve the voice quality.
Owner:RDA CHONGQING MICROELECTRONICS TECH CO LTD

System

An object of a system according to an embodiment is to analyze biological data of participants and automatically organize an effective breakout room.SOLUTION: A system includes a biological data collection section, an analysis section, a classification section, and an organization section. The biological data collection unit collects biological data such as facial expressions and voice quality from a camera or a microphone of a terminal of a participant. The analysis section analyzes the biological data collected by the biological data collection section. The classification component classifies the participants by character on the basis of the data analyzed by the analysis component. The organization unit automatically organizes a breakout room based on the characters classified by the classification unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

radio

PendingJP2026043102ATransmissionNoiseSpeech sound
To provide a radio capable of performing squelch control with little erroneous determination regardless of the voice quality of a speaker. [Solution] The radio 20 in this example comprises an SSB demodulation unit 21 that performs SSB demodulation processing on the received signal, a noise cancellation unit 22 that performs noise cancellation processing on the audio signal resulting from demodulation by the SSB demodulation unit 21, a pitch period calculation unit 23 that calculates the pitch period of the audio from the audio signal after noise cancellation processing, and a squelch control unit 24 that controls the opening and closing of the squelch of a squelch circuit 25 based on the pitch period calculated by the pitch period calculation unit 23.
Owner:KOKUSAI DENKI ELECTRIC INC

Voice quality evaluation method and device, equipment and storage medium

The invention discloses a voice quality evaluation method and device, equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps of obtaining to-be-evaluated voice generated through text-to-voice conversion and style information for the to-be-evaluated voice, and extracting basic features of the to-be-evaluated voice; extracting text-to-speech features from the basic features; the text-to-speech features comprise any one or more of rhythm features, tone features and text matching features; calculating a style matching degree between an actual style embedding vector corresponding to the text-to-speech feature and a style embedding vector template corresponding to the style information; and predicting a speech degradation score based on the text-to-speech features, and determining speech quality based on the speech degradation score and the style matching degree. By extracting rhythm features, timbre features and text matching features, identifying specific quality defects in text-to-speech conversion; and in combination with the style matching degree, the problem of confusion of style differences and quality defects is solved, and the accuracy of voice quality scoring is improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Voice quality testing including codec rate change detection

A disclosed method may include (i) detecting a codec rate change during a voice quality test of a testing UE that emulates a subscriber UE connecting to a base station in a mobile network such that a test segment during which the codec rate change occurred is identified and (ii) reducing, based on detecting that the codec rate change occurred during the test segment, a weight of the test segment within a voice quality assessment.
Owner:BOOST SUBSCRIBERCO LLC

Voice generation method and device based on preference alignment, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a speech generation method and device based on preference alignment, equipment and a medium, and the method comprises the steps: obtaining a pre-training speech generation model and a preference training sample pair, constructing a preference alignment model and a non-preference alignment model, and taking the pre-training model as a reference; determining a loss function based on the preference samples and the non-preference samples, and updating model parameters to obtain a trained model; and receiving a target condition item and generating an agent prompt in a reasoning stage, fusing speed prediction results of the two types of models, and generating a target voice based on a fusion result in a stream matching process. According to the method, human preference signals are fused through a preference alignment mechanism, so that the naturalness and semantic consistency are considered in the speech generation process, and the speech quality and the personalized expression ability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Hearing aid control method and system for eliminating ear blockage effect

PendingCN121568021AHearing aids signal processingOcclusion effectSpeech sound
A hearing aid control method for eliminating an ear blockage effect comprises the following steps: acquiring an audio decibel recorded by a first microphone as a first decibel; the decibel interval where the first decibel falls is judged, the access mode of the microphones is selected based on the decibel interval, and the first access mode is that the first microphone and the second microphone are both electrically connected with the audio access interface, and the valve of the air hole of the hearing aid is in a closed state; the second access mode is that the first microphone is connected with the hearing aid control chip; and when the second access mode is executed, the valve opening degree of the hearing aid air hole is controlled based on the first decibel, and whether the second microphone is electrically connected with the audio access interface is controlled. According to the invention, the automatic opening and closing of the air hole are realized, while the ear plugging effect and the ear canal damp-heat problem are alleviated, the sound quality is ensured and the howling is effectively inhibited through an acoustic management mechanism of software and hardware collaboration, so that the hearing aid can give consideration to both comfort and definition in different use scenes.
Owner:FOSHAN VOHOM TECHNOLOGY CO LTD +1

Call voice real-time noise reduction method and system based on dynamic noise perception

The application discloses a call voice real-time noise reduction method and system based on dynamic noise perception, acquires a noisy original voice signal, carries out frame division and windowing pretreatment on the digital voice signal with noise, extracts the spectral characteristics of the noise, divides the noise into steady-state noise, non-steady-state noise and burst noise according to the spectral characteristics of the noise, constructs a lightweight neural network structure through model pruning and quantization, adjusts the parameters and strategies of the types of the steady-state noise, the non-steady-state noise and the burst noise to realize optimized reduction, carries out post-processing such as window removal and overlap addition on the voice signal after noise reduction, and outputs the voice signal after noise reduction; the application effectively removes the background noise in a complex noise environment through dynamic noise perception and an adaptive noise reduction strategy, significantly improves the voice quality, meets the real-time requirement, realizes real-time noise reduction processing under low delay conditions through lightweight model design, and meets the real-time requirement of voice communication and voice control.
Owner:JIANGXI RUI TECH CO LTD

Voice quality evaluation system and method

The invention relates to the technical field of voice signal processing, and provides a voice quality evaluation system and method. The system comprises an acoustic feature extraction module used for extracting acoustic information from an input voice signal and outputting a first time sequence feature; the semantic feature extraction module is used for extracting linguistic content and semantic information from the voice signal and outputting a second time sequence feature; the bidirectional cross-attention fusion module is used for carrying out interactive fusion on the first time sequence feature and the second time sequence feature to generate a cross-enhanced feature vector; and the score prediction module is used for predicting a voice quality score according to the cross enhancement feature vector. According to the method and the device, through the acoustic branch, the semantic branch and a bidirectional cross-attention fusion mechanism, deep interactive modeling of acoustic fidelity and semantic content is realized, systematic deviation in the prior art is effectively reduced, and the accuracy and generalization ability of a voice quality evaluation model in diversified and unknown scenes are remarkably improved.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Dual-group human voice synchronous acquisition frequency domain noise reduction and restoration method and system

The application relates to the technical field of voice communication, and discloses a frequency domain noise reduction and restoration method and system for double-group human voice synchronous acquisition, wherein the method acquires a synchronous signal through double-group acquisition points, constructs a double-source frequency domain coordinate system and maps the double-source frequency domain coordinate system into energy distribution points; a spatial relationship is calculated to identify an emotional distortion state, four quadrants are divided to generate distortion type marks; the frequency domain signals are segmented according to the marks, targeted reconstruction and correction are performed, and finally, the energy of a frequency band is enhanced to significantly reduce a part; the application can accurately process voice composite distortion under extreme emotions, and provides reliable voice quality guarantee for accurate communication of key information in an emergency communication scene.
Owner:GUANGZHOU CMX AUDIO CO LTD

Robust speech enhancement method based on adaptive beam forming and sparse spectrum constraint

The invention discloses a robust speech enhancement method based on adaptive beam forming and sparse spectrum constraint. The robust speech enhancement method comprises the following steps: receiving a multichannel observation signal through a microphone array and constructing a signal model of a generalized sidelobe canceller structure; based on the signal model, a beam forming optimization model is constructed in combination with the least square criterion and the speech spectrum sparse constraint; performing iterative solution on the beam forming optimization model by adopting a Lagrange multiplier alternating direction method to obtain an adaptive filter weight vector in a generalized sidelobe canceller; and calculating and outputting an enhanced voice signal based on the weight vector. According to the method, the voice quality and intelligibility can be effectively improved in a strong interference and reverberation environment, interference and reverberation are remarkably suppressed, and the robustness of an algorithm is improved.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Voice packet loss hiding method and device, electronic equipment and storage medium

PendingCN121789695ASpeech analysisTelecommunicationsLinear prediction coding
The invention provides a voice packet loss hiding method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the framing processing of a received voice signal, and obtaining a current frame signal and a historical reference frame signal; based on the current frame signal, calculating an LPC coefficient through linear predictive coding LPC analysis; performing echo detection on the current frame signal and the historical reference frame signal, and judging whether an echo exists in the current frame signal; if it is judged that the echo exists, echo cancellation processing is started, and an LPC coefficient is updated based on an echo cancellation result to obtain an updated LPC coefficient; the updated LPC coefficient is used for frame reconstruction when subsequent voice packet loss occurs. According to the invention, through real-time echo detection and elimination of echo components in the reconstructed voice, the core problem of echo generation due to limitation of an LPC model in a transient phase is solved in a targeted manner, so that the reconstructed voice is purer and more natural, and the voice quality is remarkably improved.
Owner:TELINK SEMICON SHANGHAI

Voice quality detection method and device, medium and equipment

The embodiment of the invention discloses a voice quality detection method, and the method comprises the steps: determining a voice classification result of each frame through a preset voice activity detection algorithm, dividing audio data into a voice segment and a non-voice segment, taking each frame of audio of which the classification result is the voice data in the non-voice segment as an interference frame, and carrying out the elimination. And calculating a signal-to-noise ratio of the voice segment to determine a quality detection result. And the interference frame which is misjudged as voice is eliminated from the non-voice segment, and purer noise estimation is obtained. Based on the calculated signal-to-noise ratio, the voice signal quality can be reflected more truly, and the problems of noise power overestimation and signal-to-noise ratio underestimation caused by VAD false detection are effectively solved, so that the accuracy and reliability of voice quality detection are improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Voice service channel switching method, terminal, system and storage medium

PendingCN122395552ASignal qualitySpeech sound
The application provides a voice service channel switching method, a terminal, a system and a storage medium, and relates to the technical field of communication. Based on a first terminal and a second terminal connected by wire, when voice service is initiated at any time, a path with the best network signal quality is taken as a target voice service path according to a comparison result of a voice quality level on the first terminal side and a voice quality level on the second terminal side, and voice service is performed by the terminal corresponding to the path. The method can not only switch to the best voice service channel for voice service in time and accurately, but also effectively improve the voice service quality of converged services under different network modes, especially the voice service quality of wide / narrowband converged services, so as to provide the best call quality for users and avoid the situation that the user cannot make a call in an extreme case.
Owner:HYTERA COMM CORP

Speech enhancement method, electronic equipment, vehicle, medium and product

The invention discloses a speech enhancement method, electronic equipment, a vehicle, a computer readable storage medium and a computer program product. The method comprises the steps that first position information of a vehicle microphone is determined according to state information of a vehicle, and the position of the vehicle microphone is adjustable; and according to the first position information and an audio signal received by the vehicle microphone, adjusting the position of the vehicle microphone so as to reduce noise received by the vehicle microphone, and performing noise reduction processing on the audio signal. Thus, according to the state information of the vehicle, the first position information of the vehicle microphone can be determined to adapt to different working conditions of the vehicle, and an accurate data basis is provided for microphone position adjustment. According to the first position information and the audio signal received by the vehicle microphone, the position of the vehicle microphone can be adjusted in real time, the noise can be quickly and accurately suppressed, and the voice quality can be improved to a certain extent.
Owner:BYD CO LTD

Voice generation method based on two-channel semantic token and block condition flow matching

The invention relates to the technical field of speech synthesis, and particularly discloses a speech generation method based on two-channel semantic token and block condition flow matching. According to the method, the first semantic token sequence and the second semantic token sequence are generated through double-channel semantic parallel processing, the two channels can mutually perceive interaction dynamic information, interaction information between speakers can be accurately obtained, the problem that interaction dynamic information cannot be extracted through a single channel is solved, and naturalness and interactivity of synthesized voice are improved; besides, the long token sequence is divided into token blocks with fixed lengths, and each token block is processed by adopting a conditional stream matching model, so that high-quality synthesis is realized, and the naturalness of the synthesized speech is further improved. When the method is applied to business scenes of medical knowledge popularization, financial product interpretation and the like, natural and smooth dialogue audio with high interactivity can be generated, and the popularization, introduction, interpretation and promotion video voice quality is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech-based linear prediction codec post-processing method, device, equipment and medium

PendingCN121938385ASpeech analysisBiological modelsSpeech soundSpeech spectrum
The invention discloses a speech-based linear prediction codec post-processing method and device, equipment and a medium, and relates to the technical field of computers, and the method comprises the steps: obtaining original decoding parameters of all speech frames obtained by processing an original speech signal through a standard linear prediction codec; performing parameter enhancement on the original decoding parameters by using a preset parameter enhancement model corresponding to each original decoding parameter to obtain enhanced parameters; obtaining an original voice spectrum based on the original decoding parameter, obtaining an enhanced voice spectrum based on the enhanced parameter, and splicing the original voice spectrum and the enhanced voice spectrum to obtain a spliced voice spectrum; processing the spliced speech spectrum by using a preset spectrum enhancement network to obtain a depth filter coefficient; and performing depth filtering on the enhanced speech spectrum based on the depth filter coefficient to obtain a target speech spectrum, and obtaining a target speech signal based on the target speech spectrum. The voice quality can be improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Microphone design method with voiceprint privacy protection function

The invention discloses a microphone design method with a voiceprint privacy protection function, and belongs to the technical field of voice privacy protection, and the method comprises the steps: S1, training a semantic feature extraction model; s2, according to a semantic feature extraction model, determining spectral distribution of the semantic features of the specified speaker, and obtaining a frequency band corresponding to the semantic features; s3, designing a microphone device according to the frequency band corresponding to the semantic feature, and obtaining the microphone device without the voiceprint feature; and S4, recording an audio according to the microphone device from which the voiceprint features are removed, and complementing the frequency spectrum by using the generative model to obtain a complete audible audio. The voice privacy protection method solves the problems that an existing voice privacy protection method cannot achieve real-time protection and is high in computing power requirement. According to the method, the specific spectrum distribution of the semantic features of the speaker is disclosed from the signal level, the positions of the semantic features in the voice in the spectrum can be effectively positioned, the influence on the voice quality is small while the voiceprint privacy is protected, and the method is high in real-time performance and high in safety.
Owner:ZHEJIANG UNIV

Ultra-long distance speech enhancement system based on PDMA

PendingCN121905204Aquality improvementimprove intelligibilitySpeech analysisIntelligibility (communication)Testing Methods
The invention provides a PDMA-based ultra-long distance speech enhancement system, which comprises a hardware acquisition end and a signal processing end, and is characterized in that the hardware acquisition end comprises a parabolic reflector and a differential microphone array, the hardware acquisition end is used for acquiring an original mixed signal, the original signal passes through the parabolic reflector to obtain a pre-enhanced signal, and the pre-enhanced signal is used for processing the pre-enhanced signal; an original signal passes through the differential microphone array to obtain a differential signal, the obtained pre-enhanced signal and the differential signal are input to the signal processing end, and the signal processing end carries out fusion processing on the received pre-enhanced signal and the differential signal to realize secondary enhancement of a target and output a target voice signal. According to the invention, the voice signal can be effectively enhanced in an ultra-long distance scene of more than 20 meters, the voice quality and intelligibility can be obviously improved, the requirements of daily communication transmission are met, the complexity of the system is reduced, the declaration cost is controlled, and the performance of the system is improved.
Owner:TIANJIN UNIV

Radio station voice silencing method

ActiveCN121506164ASpeech analysisStationary noiseNoise (radio)
The invention discloses a radio station voice squelch method, relates to the technical field of radio station squelch, and solves the technical problem that in the prior art, voice and noise are difficult to accurately distinguish under the condition of a low signal-to-noise ratio, and misjudgment is easily caused. The method comprises the following steps: pre-processing an input digital voice signal, performing voice characteristic parameter calculation on the pre-processed digital voice signal, and processing an obtained calculation result in two paths: in one path, directly mapping the calculation result into a voice quality grade through a binary search method to obtain a real-time voice quality grade; the other path performs smooth filtering on the calculation result and then looks up a table through a binary search method to map the calculation result into a voice quality level, so that an average voice quality level is obtained; finally, the real-time voice quality grade and the average voice quality grade are sent to a squelch judgment module for squelch judgment, and squelch switch output is obtained; the method is based on voice characteristic parameter calculation and analysis, and has good adaptability to stable noise.
Owner:CHENGDUSCEON TECH

Voice quality early warning method and device, electronic equipment and storage medium

ActiveCN117240959BSimulationSpeech sound
The application discloses a voice quality early warning method and device, electronic equipment and storage medium, and belongs to the field of wireless communication, to improve the early warning efficiency of voice quality, the method comprises the following steps: acquiring at least one first quality difference call data of a preset area in a preset time period; according to the first user position information corresponding to each first quality difference call data, each first quality difference call data is converged to the corresponding element point in the preset quality difference grid, wherein the quality difference grid is obtained by gridding the preset area; the element points in the quality difference grid are clustered to obtain all quality difference problem clusters, and all quality difference problem clusters are displayed in a map tool including the preset area, wherein the position of any quality difference problem cluster in the map tool corresponds to the position of the quality difference problem cluster in the quality difference grid; if it is detected in the map tool that the quality difference problem cluster corresponding to any position point meets the preset early warning condition, voice quality early warning is performed for any position point.
Owner:XINYANG BRANCH HENAN CO LTD OF CHINA MOBILE COMM CORP +1

A model training method and device, a chip, and a module equipment

The application discloses a model training method and device, a chip and a module equipment. The method comprises the following steps: based on the model parameters of a first evaluation model obtained through i-th iteration training and the model parameters of a second evaluation model obtained through (i-1)-th iteration training, adjusting the model parameters of the second evaluation model in the i-th iteration training; inputting sample audio data into the first evaluation model and the second evaluation model respectively in (i+1)-th iteration training, obtaining a first speech quality evaluation result output by the first evaluation model and a second speech quality evaluation result output by the second evaluation model; and based on the first speech quality evaluation result, the second speech quality evaluation result and a speech quality evaluation result label, adjusting the model parameters of the first evaluation model in (i+1)-th iteration training. According to the method described in the application, the model obtained through training can more objectively and accurately evaluate the speech quality.
Owner:UNISOC CHONGQING TECH CO LTD

End-to-end voice encryption method compatible with broadband and narrowband wireless communication systems

The invention discloses an end-to-end voice encryption method compatible with broadband and narrowband wireless communication systems, a sender in a broadband system sends voice data of a basic layer and voice data of an enhancement layer at the same time, a receiver in a narrowband system only processes the voice data of the basic layer, high-bandwidth resources can be dynamically allocated and used in a broadband and narrowband mixed group calling scene, and the user experience is improved. And the problem that the voice quality is reduced due to forced and unified adoption of a low-bit-rate vocoder is avoided. According to the difference between the voice sending period of the broadband system and the voice sending period of the narrowband system, the unpacking / packing operation is carried out among different systems through the interconnection equipment, and extra time delay caused by cross-system group calling is eliminated. According to the invention, after the encrypted synchronization information sent intermittently is subjected to redundant coding, the encrypted synchronization information is split and embedded into the continuous RTP voice packets, so that the length of all the RTP packets is ensured to be constant, an SPS scheduling mechanism is enabled to accurately reserve resources according to actual service requirements, and wireless resource waste is effectively eliminated.
Owner:THE FIRST RES INST OF MIN OF PUBLIC SECURITY

A pickup method of a microphone array, an electronic device, and a storage medium

The application discloses a microphone array sound pickup method, electronic equipment and computer readable storage medium. The sound pickup method comprises the following steps: fixed beam forming is performed on a voice signal received by a microphone array, and a beam forming direction of the microphone array is pointed to an estimated expected direction of arrival; blocking processing is performed on the processed voice signal, so that the voice signal from the expected direction of arrival is blocked, and only the voice signal from a non-expected direction of arrival is reserved; the processed signal is used as a reference signal, a first filter is used to filter out the signal from the non-expected direction of arrival and reserve the signal from the expected direction of arrival; and an update factor of the first filter is calculated according to the following formula (I), wherein the update factor of the first filter of the mth microphone channel is SNR f,d (ω, l) is Y f,d (ω, l) is a signal-to-noise ratio of Y f,d (ω, l) is a delay signal obtained by delaying the signal processed in the step S1, SNR m (ω, l) is a signal U m (ω, l) is a signal-to-noise ratio of Y. The application further improves voice quality.
Owner:COLSONIC SUZHOU ELECTRONICS CO LTD

Systems and methods for enhancing speech audio signals

A method and device for enhancing speech audio signals of an individual in a noisy environment based on a user's gaze and a captured image of the user's environment. A direction of a user's gaze is determined using image sensors configured to capture an orientation of a user's eyes and an image of the user environment is captures. Spatial audio is captured and analyzed along with the direction of gaze and image of the user environment to enhance audio of an active speaker.
Owner:ADEIA GUIDES INC