Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
23 results about "Speech perception" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
Speech perception is the process by which the sounds of language are heard, interpreted and understood. The study of speech perception is closely linked to the fields of phonology and phonetics in linguistics and cognitive psychology and perception in psychology. Research in speech perception seeks to understand how human listeners recognize speech sounds and use this information to understand spoken language. Speech perception research has applications in building computer systems that can recognize speech, in improving speech recognition for hearing- and language-impaired listeners, and in foreign-language teaching.
The invention relates to the technical field of medical health and artificial intelligence, in particular to a system and method for improving the cognitive function of old people based on emotion voice perception training. Comprises: a voice stimulation presentation module for controlling playing of target voice stimulation and masking voice; the training program control module is used for controlling and adjusting the voice stimulation presentation module; the reaction collection module collects the reaction of the patient and outputs the reaction to the feedback and enhancement module; the feedback and strengthening module is used for providing training feedback according to the input of the reaction acquisition module; the data processing and evaluation module is used for processing the input of the feedback and enhancement module; and the user interface and visualization module is used for providing operation and result display for patients and researchers. According to the method, emotional voice perception training and voice recognition training in a noise environment are combined for the first time, a comprehensive cognitive intervention mode based on acoustic signalprocessing and emotional rhythm recognition is established, and the limitation that a traditional cognitive training visual channel is poor in dependence and migration effect is broken through.
The invention provides an ultra-low bit rate voice coding and decoding system based on text semantic information fidelity, and relates to the technical field of voice coding and decoding. The system comprises the following steps: performing voice feature extraction and text feature extraction on original voice through a multi-modal text-voice combined encoder to obtain voice features and text features, embedding the text features into the voice features to obtain text-voice features of the original voice, and performing voice feature extraction and text feature extraction on the original voice through a multi-modal text-voice combined encoder; sending the text-speech features of the original speech to a receiving end in a semantic communication system; and the receiving end inputs the received text-speech features into a text semantic fidelity decoder based on an attention mechanism to obtain reconstructed speech features and reconstructed text features, decodes the reconstructed speech features layer by layer, and performs attention calculation on the reconstructed speech features and the reconstructed text features to obtain decoded speech. By means of the voice coding and decoding technology, good voice perception quality can still be kept under the ultra-low code rate, and voice intelligibility is guaranteed.
Presented herein are techniques for improving speech perception for hearing device recipients through the use of phonemic information. For example, in certain embodiments, automatic speech recognition (ASR) is used to enhance the perception of phonemic information by a recipient. Such techniques can, for example, provide improved speech intelligibility in noise for hearing device recipients.
The invention discloses a visual lip auxiliary evaluation method based on voice perception in a noise environment of electroencephalogram. The visual lip auxiliary evaluation method comprises the following steps: step 1, implementing a stimulation experiment to collect electroencephalogram data; 2, preprocessing the electroencephalogram data collected in the step 1, and grouping according to noise types, signal-to-noise ratio levels and visual conditions to obtain an electroencephalogram data set; 3, electroencephalogram features are extracted based on the electroencephalogram data set obtained in the step 2, wherein the electroencephalogram features comprise PLV and COH; and 4, analyzing an auxiliary function of visual lip information on voice perception by combining the correlation between the electroencephalogram characteristics obtained in the step 3 and the voice recognition accuracy obtained in the step 1. According to the method, the limitation of traditional behavioral detection is broken through, and evaluation upgrading from perception result description to neural mechanism analysis is realized.
The application discloses a speech text error correction method, a speech text error correction device, a speech text error correction equipment and a storage medium. The method comprises the following steps: obtaining speech information, obtaining a Chinese character sequence in the speech information and a pinyin sequence corresponding to the Chinese character sequence, generating a speech perception sequence according to the combination of the pinyin sequence and the Chinese character sequence, introducing an encoder with a separation mask, forming a coding sequence according to the speech perception sequence and the segmentation mask, establishing a mapping model according to the dependency relationship between an error text and a correct text, and obtaining a correct Chinese text corresponding to the speech information according to the coding sequence and the mapping model. The method can more accurately locate the part that may have errors, can better handle the complexity of the language structure, and thus improves the accuracy of error correction.
The application discloses a post-processing method and device for improving voice perception, equipment and medium. The pre-configured contrast stretching maximum value can limit the degree of contrast stretching, preventing signaldistortion caused by excessive stretching. When determining the target contrast stretching parameter of any frequency, the pre-set contrast stretching preset parameter corresponding to the frequency is divided by the pre-determined contrast stretching maximum value, which can effectively reduce the energy of the non-sensitive frequency of the human ear and maintain the energy of the sensitive frequency of the human ear. Subsequently, the amplitude spectrum of the current voice frame is stretched by applying the target contrast stretching parameter corresponding to each frequency to obtain the stretched amplitude spectrum. The stretched amplitude spectrum maintains the energy of the sensitive frequency of the human ear while reducing the energy of the non-sensitive frequency, thereby avoiding excessive stretching of the amplitude of the sensitive frequency of the human ear and the problem of numerical overflow.
The application relates to a multi-modal fusion emotion recognition method and system for a social robot, and the method comprises the following steps: obtaining real-time sensing information of target personnel, wherein the real-time sensing information comprises image information, tactile information and sound information; pre-processing the real-time sensing information to obtain pre-processed sensing information; and obtaining an emotion recognition result according to the pre-processed sensing information. Compared with the prior art, the application fuses visual, voice and tactile sensing data for emotion recognition, uses a microphone and an electromyography sensor to acquire voice sensing data, and uses multiple key points to capture human parameters for visual sensing data, so that the recognition accuracy is improved and the method is suitable for more environments.
A team sports vision training system based on extended reality, voice interaction and action recognition is configured to train vision and an action of a user. A head-mounted display device includes a task scenario player and a speech sensing module. An action capture device generates an action message. A computing server stores a scenario setting parameter group and includes a task scenario generating module, a speech recognition module and an action recognition module. The task scenario generating module generates a virtual task scenario image and a task parameter group according to the scenario setting parameter group. The speech recognition module generates a speech recognition result and a vision training result. Then action recognition module generates an action recognition result and a sport training result. The vision training result and the sport training result are configured to judge whether the user meets a training requirement.
This invention discloses an environment-aware, controllable background removal and preservation speech synthesissystem, relating to the field of speech. The system proposes a speech synthesissystem capable of perceiving the acoustic environment based on noisy cues, thereby performing controllable background removal and preservation. It takes text, cues, and task-related control signals as input and includes a duration predictor, an acoustic model, and a dual-cue speech encoder. In terms of training strategy, based on a stream matching algorithm, a controllable masked speech prediction training strategy is further proposed, achieving controllable background removal and preservation by providing noisy cues. This invention improves the robustness and controllability of the system in handling noisy, reverberant, and speaker-interference cues, effectively controlling the removal and preservation of background contained in the cues during speech generation, achieving higher generated speech quality and a more similar acoustic background.
The application discloses a visual lip auxiliary evaluation method for speech perception in a noise environment based on electroencephalogram (EEG), and comprises the following steps: step 1, collecting EEG data by implementing a stimulation experiment; step 2, pre-processing the EEG data collected in step 1, and grouping the EEG data according to noise types, signal-to-noise ratio (SNR) levels and visual conditions to obtain an EEG data set; step 3, extracting EEG features based on the EEG data set obtained in step 2, wherein the EEG features comprise phase locking value (PLV) and cross-frequency coupling (COH); and step 4, analyzing the auxiliary function of visual lip information for speech perception in combination with the correlation between the EEG features obtained in step 3 and the speech recognition accuracy obtained in step 1. The application breaks through the limitation of traditional behavior detection, and realizes evaluation upgrading from 'perception result description' to 'neural mechanism analysis'.
The application provides a kind of based on text semantic information fidelity super low code ratespeech codingsystem, it is related to speech coding technical field.The system includes: through multimodal text-speech joint encoder, speech feature extraction and text feature extraction are carried out to original speech, speech feature and text feature are obtained, and text feature is embedded into speech feature, text-speech feature of original speech is obtained, and text-speech feature of original speech is sent to receiving end in semantic communication system;The text-speech feature received by receiving end is input into text semantic fidelity decoder based on attention mechanism, to obtain reconstructed speech feature and reconstructed text feature, and the reconstructed speech feature is decoded layer by layer and attention calculation is carried out with the reconstructed text feature, to obtain decoded speech.Through the speech coding technology of the application, good speech perceptual quality can still be maintained under super low code rate, and the speech intelligibility is guaranteed.
Methods, device and system for determining a target voice parameters. A location within a 2D search space is assigned to parameterized voices, perceptually similar voices being proximate. Candidate-voices are inserted into a candidate list when a resemblance threshold is reached; A choice between two unmixed voices is received. The plurality of underlying parameters of the unmixed voices are mixed into a mixed voice towards the target-voice. The plurality of underlying parameters from the candidate list are identified. The unadjusted voice is adjusted into an adjusted voice by altering values of the plurality of underlying parameters towards the target-voice. A user interface module receives a choice of a candidate-voice from the 2D search space. An audio playback device plays back at least a portion of the candidate-voice.
This invention belongs to the field of artificial intelligence education technology and discloses a humanoid robot multimodal control system for artificial intelligence education. The privacy perception module integrates visual perception, voice perception, tactile perception, and environmental perception units to accurately capture students' learning behaviors and complete knowledge point association filtering, teaching semantic noise reduction, and graded perception of minors' contact ability, filtering invalid data. The semantic fusion module completes spatiotemporal alignment based on the teaching timeline, generates semantic constraint rules based on the subject cognitive map, and uses an improved Transformer model to achieve deep multimodal fusion, dynamically outputting teaching-specific semantic labels, and transforming multimodal data into cognitive states that can directly support teaching decisions. The closed-loop decision module, based on the K12 full-subject cognitive map and dynamic student cognitive profile, generates customized teaching strategies through deep reinforcement learning and automatically converts them into full-dimensional instructions such as voice and actions.
The present application relates to the technical field of speech emotion recognition, in particular to a human-computer interaction speech perception method and system based on gradient intelligent tonal net pool. The method comprises: obtaining an emotion dataset; constructing a human-computer interaction speech perception model based on gradient intelligent tonal net pool, which comprises an acoustic cue perception purification module, a hierarchical acoustic essence coding module, a gradient harmonization sub-net pool module, a task-specific feature extraction module, a focus and confidence joint calibration module, an adaptive optimization strategy module, and a real-time reasoning and decision fusion module; using the constructed human-computer interaction speech perception model to make an emotion decision; and outputting the decision result. The present application fundamentally solves the problem of emotion information distortion and identity feature confusion caused by real environment noise through the acoustic cue perception purification module and the hierarchical acoustic essence coding.
The invention discloses a lightweight speech enhancement method based on a grouped dual-path LSTM (Long Short Term Memory). The lightweight speech enhancement method comprises the following steps: downloading and preprocessing a VoiceBank + DEMAND data set required by a model; performing short-time Fourier transform on the noisy voice to convert the noisy voice into a frequency domain; compressing the high-frequency spectrum to an equivalent rectangular bandwidth (ERB) sub-band space by using a frequency band compression module; the compressed spectrum features are input into an encoder to extract high-order time-frequency features, in-depth modeling is carried out on contextual information in time and frequency dimensions through a grouping double-path long and short term memory module, and then the features are restored through a decoder; reconstructing an original spectrum resolution by means of a frequency bandrecovery operation; converting a result from a frequency domain to a time domain through short-time inverse Fourier transform, and reconstructing a voice waveform; constructing a joint loss function; and the result is converted from the frequency domain to the time domain through short-time inverse Fourier transform (iSTFT), the voice waveform is reconstructed, and the performance of the proposed model is evaluated. Through spectrum compression and a grouping parallel modeling strategy, while the voice perception quality, the voice definition and the background noise suppression capability are improved, the model parameter quantity and the calculation complexity are remarkably reduced.
A voice perception method based on RFID, through a conditional denoising autoencoder network (CDAE), the collected RF signals are preprocessed to eliminate the interference of body movement, and then the preprocessed RF signals are extracted through a recurrent neural network (RNN), a convolutional recurrent neural network (CRNN) and a deep residual shrinkage network (DRSN) to obtain facial movement, skeletal vibration and air vibration features, and after fusion, the comparative learning network is input to remove the user-specificity in the facial voice dynamic feature, obtain the facial voice dynamic feature related to the sound content, and then construct a facial voice dynamic model according to the facial voice dynamic feature, and realize voice perception. The application can perceive the facial voice dynamics when people speak by pasting a RFID tag on ordinary glasses without continuous power supply and low cost, and can realize voice perception in a natural and robust manner, and has wide application scenarios.
This invention discloses an emotion monitoring system and method based on voice perception of smart safety helmets, belonging to the technical field of speech processing. The system acquires video streams from a construction site and extracts audio and visual features. These features are then fused to obtain multimodal features, which are input into an improved model to obtain a binary spectral mask. The binary spectral mask is combined with the audio features and subjected to time-frequency domain transformation to obtain clean speech features. A text sequence is generated based on these clean speech features. The clean speech features, text sequence, and video stream are input into a target model to obtain emotion tags. Work logs are retrieved, and emotion tags are added to the work logs to obtain digital archives. By fusing audio and visual features from the construction video stream, the improved model achieves speech separation under complex noise conditions, extracts work content, generates text, and accurately identifies the emotions of mask-wearing workers by combining multi-dimensional information, generating digital archives with emotion tags.