Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
16 results about "Speech perception" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
Speech perception is the process by which the sounds of language are heard, interpreted and understood. The study of speech perception is closely linked to the fields of phonology and phonetics in linguistics and cognitive psychology and perception in psychology. Research in speech perception seeks to understand how human listeners recognize speech sounds and use this information to understand spoken language. Speech perception research has applications in building computer systems that can recognize speech, in improving speech recognition for hearing- and language-impaired listeners, and in foreign-language teaching.
The invention relates to the technical field of medical health and artificial intelligence, in particular to a system and method for improving the cognitive function of old people based on emotion voice perception training. Comprises: a voice stimulation presentation module for controlling playing of target voice stimulation and masking voice; the training program control module is used for controlling and adjusting the voice stimulation presentation module; the reaction collection module collects the reaction of the patient and outputs the reaction to the feedback and enhancement module; the feedback and strengthening module is used for providing training feedback according to the input of the reaction acquisition module; the data processing and evaluation module is used for processing the input of the feedback and enhancement module; and the user interface and visualization module is used for providing operation and result display for patients and researchers. According to the method, emotional voice perception training and voice recognition training in a noise environment are combined for the first time, a comprehensive cognitive intervention mode based on acoustic signalprocessing and emotional rhythm recognition is established, and the limitation that a traditional cognitive training visual channel is poor in dependence and migration effect is broken through.
The invention provides an ultra-low bit rate voice coding and decoding system based on text semantic information fidelity, and relates to the technical field of voice coding and decoding. The system comprises the following steps: performing voice feature extraction and text feature extraction on original voice through a multi-modal text-voice combined encoder to obtain voice features and text features, embedding the text features into the voice features to obtain text-voice features of the original voice, and performing voice feature extraction and text feature extraction on the original voice through a multi-modal text-voice combined encoder; sending the text-speech features of the original speech to a receiving end in a semantic communication system; and the receiving end inputs the received text-speech features into a text semantic fidelity decoder based on an attention mechanism to obtain reconstructed speech features and reconstructed text features, decodes the reconstructed speech features layer by layer, and performs attention calculation on the reconstructed speech features and the reconstructed text features to obtain decoded speech. By means of the voice coding and decoding technology, good voice perception quality can still be kept under the ultra-low code rate, and voice intelligibility is guaranteed.
The invention discloses a visual lip auxiliary evaluation method based on voice perception in a noise environment of electroencephalogram. The visual lip auxiliary evaluation method comprises the following steps: step 1, implementing a stimulation experiment to collect electroencephalogram data; 2, preprocessing the electroencephalogram data collected in the step 1, and grouping according to noise types, signal-to-noise ratio levels and visual conditions to obtain an electroencephalogram data set; 3, electroencephalogram features are extracted based on the electroencephalogram data set obtained in the step 2, wherein the electroencephalogram features comprise PLV and COH; and 4, analyzing an auxiliary function of visual lip information on voice perception by combining the correlation between the electroencephalogram characteristics obtained in the step 3 and the voice recognition accuracy obtained in the step 1. According to the method, the limitation of traditional behavioral detection is broken through, and evaluation upgrading from perception result description to neural mechanism analysis is realized.
A team sports vision training system based on extended reality, voice interaction and action recognition is configured to train vision and an action of a user. A head-mounted display device includes a task scenario player and a speech sensing module. An action capture device generates an action message. A computing server stores a scenario setting parameter group and includes a task scenario generating module, a speech recognition module and an action recognition module. The task scenario generating module generates a virtual task scenario image and a task parameter group according to the scenario setting parameter group. The speech recognition module generates a speech recognition result and a vision training result. Then action recognition module generates an action recognition result and a sport training result. The vision training result and the sport training result are configured to judge whether the user meets a training requirement.
The application discloses a visual lip auxiliary evaluation method for speech perception in a noise environment based on electroencephalogram (EEG), and comprises the following steps: step 1, collecting EEG data by implementing a stimulation experiment; step 2, pre-processing the EEG data collected in step 1, and grouping the EEG data according to noise types, signal-to-noise ratio (SNR) levels and visual conditions to obtain an EEG data set; step 3, extracting EEG features based on the EEG data set obtained in step 2, wherein the EEG features comprise phase locking value (PLV) and cross-frequency coupling (COH); and step 4, analyzing the auxiliary function of visual lip information for speech perception in combination with the correlation between the EEG features obtained in step 3 and the speech recognition accuracy obtained in step 1. The application breaks through the limitation of traditional behavior detection, and realizes evaluation upgrading from 'perception result description' to 'neural mechanism analysis'.
The application provides a kind of based on text semantic information fidelity super low code ratespeech codingsystem, it is related to speech coding technical field.The system includes: through multimodal text-speech joint encoder, speech feature extraction and text feature extraction are carried out to original speech, speech feature and text feature are obtained, and text feature is embedded into speech feature, text-speech feature of original speech is obtained, and text-speech feature of original speech is sent to receiving end in semantic communication system;The text-speech feature received by receiving end is input into text semantic fidelity decoder based on attention mechanism, to obtain reconstructed speech feature and reconstructed text feature, and the reconstructed speech feature is decoded layer by layer and attention calculation is carried out with the reconstructed text feature, to obtain decoded speech.Through the speech coding technology of the application, good speech perceptual quality can still be maintained under super low code rate, and the speech intelligibility is guaranteed.
Methods, device and system for determining a target voice parameters. A location within a 2D search space is assigned to parameterized voices, perceptually similar voices being proximate. Candidate-voices are inserted into a candidate list when a resemblance threshold is reached; A choice between two unmixed voices is received. The plurality of underlying parameters of the unmixed voices are mixed into a mixed voice towards the target-voice. The plurality of underlying parameters from the candidate list are identified. The unadjusted voice is adjusted into an adjusted voice by altering values of the plurality of underlying parameters towards the target-voice. A user interface module receives a choice of a candidate-voice from the 2D search space. An audio playback device plays back at least a portion of the candidate-voice.
This invention belongs to the field of artificial intelligence education technology and discloses a humanoid robot multimodal control system for artificial intelligence education. The privacy perception module integrates visual perception, voice perception, tactile perception, and environmental perception units to accurately capture students' learning behaviors and complete knowledge point association filtering, teaching semantic noise reduction, and graded perception of minors' contact ability, filtering invalid data. The semantic fusion module completes spatiotemporal alignment based on the teaching timeline, generates semantic constraint rules based on the subject cognitive map, and uses an improved Transformer model to achieve deep multimodal fusion, dynamically outputting teaching-specific semantic labels, and transforming multimodal data into cognitive states that can directly support teaching decisions. The closed-loop decision module, based on the K12 full-subject cognitive map and dynamic student cognitive profile, generates customized teaching strategies through deep reinforcement learning and automatically converts them into full-dimensional instructions such as voice and actions.
The present application relates to the technical field of speech emotion recognition, in particular to a human-computer interaction speech perception method and system based on gradient intelligent tonal net pool. The method comprises: obtaining an emotion dataset; constructing a human-computer interaction speech perception model based on gradient intelligent tonal net pool, which comprises an acoustic cue perception purification module, a hierarchical acoustic essence coding module, a gradient harmonization sub-net pool module, a task-specific feature extraction module, a focus and confidence joint calibration module, an adaptive optimization strategy module, and a real-time reasoning and decision fusion module; using the constructed human-computer interaction speech perception model to make an emotion decision; and outputting the decision result. The present application fundamentally solves the problem of emotion information distortion and identity feature confusion caused by real environment noise through the acoustic cue perception purification module and the hierarchical acoustic essence coding.
The invention discloses a lightweight speech enhancement method based on a grouped dual-path LSTM (Long Short Term Memory). The lightweight speech enhancement method comprises the following steps: downloading and preprocessing a VoiceBank + DEMAND data set required by a model; performing short-time Fourier transform on the noisy voice to convert the noisy voice into a frequency domain; compressing the high-frequency spectrum to an equivalent rectangular bandwidth (ERB) sub-band space by using a frequency band compression module; the compressed spectrum features are input into an encoder to extract high-order time-frequency features, in-depth modeling is carried out on contextual information in time and frequency dimensions through a grouping double-path long and short term memory module, and then the features are restored through a decoder; reconstructing an original spectrum resolution by means of a frequency bandrecovery operation; converting a result from a frequency domain to a time domain through short-time inverse Fourier transform, and reconstructing a voice waveform; constructing a joint loss function; and the result is converted from the frequency domain to the time domain through short-time inverse Fourier transform (iSTFT), the voice waveform is reconstructed, and the performance of the proposed model is evaluated. Through spectrum compression and a grouping parallel modeling strategy, while the voice perception quality, the voice definition and the background noise suppression capability are improved, the model parameter quantity and the calculation complexity are remarkably reduced.
This invention discloses an emotion monitoring system and method based on voice perception of smart safety helmets, belonging to the technical field of speech processing. The system acquires video streams from a construction site and extracts audio and visual features. These features are then fused to obtain multimodal features, which are input into an improved model to obtain a binary spectral mask. The binary spectral mask is combined with the audio features and subjected to time-frequency domain transformation to obtain clean speech features. A text sequence is generated based on these clean speech features. The clean speech features, text sequence, and video stream are input into a target model to obtain emotion tags. Work logs are retrieved, and emotion tags are added to the work logs to obtain digital archives. By fusing audio and visual features from the construction video stream, the improved model achieves speech separation under complex noise conditions, extracts work content, generates text, and accurately identifies the emotions of mask-wearing workers by combining multi-dimensional information, generating digital archives with emotion tags.