Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

74 results about "Speech sounds" patented technology

Speech Sounds. sounds formed for the purpose of verbal communication by the human vocal apparatus (the lungs; larynx and vocal cords; pharynx; oral cavity with the tongue; lips; uvula; and the nasal cavity). There are three aspects of speech sounds: the articulatory, the acoustical, and the linguistic, or social.

Vocal music training method and system based on artificial intelligence

The invention relates to the technical field of voice analysis, and particularly discloses a vocal music training method and system based on artificial intelligence, and the method comprises the steps: collecting and analyzing the body posture parameters of a target vocal music training person, carrying out the preprocessing, detecting the posture abnormality, forming a first abnormal training set, capturing and analyzing the pronunciation feature parameters, and recognizing the pronunciation abnormality, thereby obtaining a vocal music training result; and integrating the first abnormal training set and the second abnormal training set to generate a total abnormal parameter training set, matching a correction set, and visually displaying the body posture abnormal degree and the abnormal reason and the pronunciation abnormal degree and the reason, so that the target vocal music training personnel can visually see own posture and pronunciation problems, and self-adjustment and correction are facilitated.
Owner:SHANGLUO UNIV

Estimation method, recording medium, and estimation device

An estimation method includes: obtaining a first voice feature group of a plurality of persons who speak a first language; obtaining a second voice feature group of a plurality of persons who speak a second language; obtaining a voice feature of a subject; correcting the voice feature of the subject according to a relationship between the first voice feature group and the second voice feature group; estimating, from the voice feature of the subject that has been corrected, an oral function or a cognitive function of the subject by using an estimation process for an oral function or a cognitive function based on the second language; and outputting a result of estimation of the oral function or the cognitive function of the subject.
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

AI (artificial intelligence) English learning system capable of correcting pronunciation

The invention relates to the technical field of English pronunciation learning, and discloses an AI (artificial intelligence) English learning system capable of correcting pronunciation. The system comprises a multi-source voice acquisition layer, an acoustic feature analysis layer, a pronunciation quality evaluation layer, a correction decision generation layer and a biological feedback execution layer. The multi-source voice acquisition layer captures a multi-channel sound wave signal and a maxillofacial electromyographic signal in real time when a user pronounces, and generates an original pronunciation data stream; the acoustic feature analysis layer performs time-frequency domain conversion on original data, separates multi-scale acoustic features and constructs a three-dimensional feature tensor; the pronunciation quality evaluation layer compares the standard template library to output scores and phoneme offset vectors; the correction decision generation layer constructs a pronunciation error topological graph and generates a correction instruction set containing tongue position, airflow and vocal cord parameters; and the biological feedback execution layer synchronously monitors the physiological response and updates the weight of the template library. According to the system, accurate acquisition, analysis, evaluation, correction and dynamic optimization of pronunciation learning are realized, and the English pronunciation learning effect is improved.
Owner:JIANGXI TELLHOW ANIMATION VOCATIONAL COLLEGE

Intelligent exoskeleton bone auxiliary conduction hearing aid system capable of being remotely controlled

The invention discloses an intelligent exoskeleton bone auxiliary conduction hearing aid system capable of being remotely controlled and having an electroencephalogram control function, and relates to the technical field of intelligent rehabilitation robots and hearing aid devices. The wearable mechanical part comprises a flexible actuator, an elastic fabric base material, an inertial measurement unit IMU, a plantar pressure sensor array, a surface myoelectricity EMG sensor and an electroencephalogram EEG sensor. And the information processing communication part comprises a central controller, a bone conduction hearing aid module, a multi-mode sensor, an embedded SoC, a DSP / AI accelerator, a 4G / 5G / Wi-Fi module and a power management unit. The flexible exoskeleton walking aid system and the bone conduction hearing aid module are integrated, the electroencephalogram EEG sensor is additionally arranged to achieve the electroencephalogram control function, and cooperation of the three functions of movement assistance, auditory perception and neural intention recognition is achieved. A user can directly control the exoskeleton action mode through electroencephalogram signals, meanwhile, intelligent assistance of lower limb joints is obtained in the walking process, and environment sounds and voice prompts are clearly received in a bone conduction mode. According to the system, the perception ability, the control flexibility and the social participation degree of old people or rehabilitation patients who are inconvenient to move and accompanied with hearing impairment in a complex environment are remarkably improved, and more natural and convenient man-machine interaction experience is provided for users.
Owner:SUZHOU ZHILINGDA INTELLIGENT TECHNOLOGY CO LTD

Synthesizing speech from facial skin movements

A method for generating speech includes uploading a reference set of features that were extracted from sensed movements of one or more target regions of skin on faces of one or more reference human subjects in response to words articulated by the subjects and without contacting the one or more target regions. A test set of features is extracted a from the sensed movements of at least one of the target regions of skin on a face of a test subject in response to words articulated silently by the test subject and without contacting the one or more target regions. The extracted test set of features is compared to the reference set of features, and, based on the comparison, a speech output is generated, that includes the articulated words of the test subject.
Owner:APPLE INC

Apparatus for the self-administration of therapies for the treatment of speech disorders

PCT designated stageWO2026028021A1Psychotechnic devicesSensorsAuditory stimuliSound sources
The present invention relates to an apparatus (10) for the self-administration of therapies for the treatment of speech disorders comprising a stimulus delivery device (11,12) with a horizontal development substantially curved along an angular portion so as to outline a portion of a circle, comprising at least one panel (11) for the delivery of visual stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); at least one acoustic source (12) for the delivery of acoustic stimuli spatially distributed in an angular range substantially equal to the angular portion of horizontal development of the stimulus delivery device (11,12); and a local electronic processing unit configured to control the generation of visual and acoustic stimuli by means of the at least one panel (11) and the at least one acoustic source (12), and characterized in that it comprises at least one first acoustic detector (15) for detecting a reproduction of a verbal element produced by the patient, wherein the at least one first acoustic detector (15) is carried by the at least one panel (11) and positioned so as to detect acoustic signals generated in a space inside the portion of circle outlined by the stimulus delivery device (11,12).
Owner:LINARI MEDICAL SRL

Voice interaction massage device

Voice-interaction massage device for massaging at least one penis, comprising the following: a housing comprising a first massage chamber and a second massage chamber, each of the first massage chamber and the second massage chamber being configured to accommodate at least part of the penis; a massage module incorporated in the housing, wherein the massage module is configured to massage at least a part of the at least one penis that is inserted into at least one of the two massage chambers; a loudspeaker; and a first sensor configured to detect movements of the penis in the first massage chamber; a control unit that is electrically connected to the speaker and the first sensor; the control unit, as soon as the first sensor detects movements of the penis in the first massage chamber, activates the speaker to output an initial voice message.
Owner:SHENZHEN SHECLONE TECHNOLOGY CO LTD

Model training method and bipolar disorder personalized treatment dialogue generation method and device

The application provides a model training method and a bipolar disorder personalized treatment dialogue generation method and device, belonging to the technical field of deep learning, aiming at the problems that the current diagnosis accuracy of bipolar mood disorder is not high and the matching of treatment dialogue with patients is poor, the model training can be performed on a multi-modal language model according to a voice signal with significant characteristics of bipolar mood disorder, and a target classification model for bipolar mood disorder classification is obtained. Since the voice signal contains characteristics such as semantics, tone, pronunciation rhythm and other characteristics directly related to bipolar mood disorder expressed by the patient, the target classification model can be used for accurate classification of bipolar mood disorder, and when the target classification model is used for classification of bipolar mood disorder, the excessive dependence on the subjective diagnosis of doctors can be avoided, and the accuracy of the classification prediction of bipolar mood disorder is effectively improved.
Owner:THE SECOND HOSPITAL AFFILIATED TO WENZHOU MEDICAL COLLEGE

A data processing method and system for multimodal facial motion point data and vocal cord motion data

The present invention discloses a data processing method and system for multimodal facial motion point data and vocal cord motion data. The method includes providing text, collecting continuous facial images or videos and laryngeal vibration data of a normal person speaking, preprocessing the data, extracting temporal and spatial features, establishing a facial and neck motion model for Chinese pronunciation, and having a deaf-mute person imitate the pronunciation according to the model and obtain feedback. The system includes a depth camera, a laryngeal vibration sensor, and a microphone. By comprehensively utilizing multimodal data, it provides instant feedback to the deaf-mute, lowers the learning threshold, and improves communication efficiency. It is applicable to deaf-mute groups worldwide. This invention promotes speech vocalization training and has broad application prospects and social significance.
Owner:ZHEJIANG UNIV

Speech disorder detection method, device, equipment and readable storage medium

The present disclosure relates to a speech disorder detection method, device, equipment and readable storage medium. By acquiring standard audio-visual materials, in response to the pronunciation operation of the to-be-detected object to the standard audio-visual materials, multi-modal pronunciation data is collected, audio acoustic features are extracted based on the pronunciation audio, video visual features are extracted based on the video of the face and oral cavity activity, the audio acoustic features, the video visual features and the demographic information coding data are fused to obtain a fusion feature vector, and based on the fusion feature vector and a pre-trained prediction model, a speech disorder detection result of the to-be-detected object is obtained. Compared with the prior art, the embodiment of the present disclosure can improve the accuracy and comprehensiveness of speech disorder detection, improve the diagnosis efficiency, reduce the dependence on professionals, reduce the burden of medical resources, and clearly determine the specific type of pronunciation problem, thereby providing a scientific basis for subsequent individualized intervention treatment.
Owner:AFFILIATED CHILDRENS HOSPITAL OF CAPITAL INST OF PEDIATRICS

Method and system for realizing global nonlinear dynamic vocal cord model

The invention discloses a global nonlinear dynamics vocal cord model implementation method and system, and relates to the technical field of nonlinear dynamics and biomechanics, and the method comprises the steps: introducing a vocal cord tension coefficient, and constructing a left-right asymmetric coefficient; introducing a Brown random jitter factor to correct the left-right asymmetric coefficient so as to construct a global nonlinear dynamic vocal cord model; and introducing a river horse optimization module to realize dynamic updating of vocal cord model parameters. According to the method, the left and right asymmetric coefficients and the vocal cord tension coefficient are introduced, so that the dynamic change of the vocal cord structure can be represented more meticulously, particularly, the difference of the mass, the elastic coefficient and the damping coefficient of the left and right sides of the vocal cord is accurately described, and the dynamic tracking of the change of the vocal cord structure is realized; through the construction of a global nonlinear model, the local change of the vocal cord is combined with the overall dynamic behavior, the dynamic characteristics of the vocal cord under different sounding states or pathological conditions are reflected, and a finer modeling basis is provided for vocal cord disease diagnosis and voice research.
Owner:SUZHOU UNIV

Digital human generation system and method based on language model

PendingCN121811908ASpeech analysisTongue tipSpeech sounds
The invention discloses a digital human generation system and method based on a language model, and relates to the technical field of artificial intelligence. A complete digital human pronunciation modeling path from voice signal acquisition, formant frequency extraction, deviation analysis and pronunciation stability evaluation to motion compensation control and three-dimensional animation fusion is realized. Compared with the existing mode of driving the digital human pattern only based on energy envelope or mouth shape classification, the method has the advantages that the multi-frame formant frequency of the user voice is extracted and processed, so that the digital human can generate dynamic lingual surface and tongue tip actions according to the sound channel change in the real pronunciation process of the user; therefore, the physical correspondence and linguistic consistency of the pronunciation actions of the digital human are improved. According to the method, the direct mapping between the pronunciation characteristics and the three-dimensional tongue actions is realized, and the authenticity of the pronunciation animation of the digital human and the guidance in a teaching scene are remarkably improved.
Owner:SUZHOU LVHUA TECH CO LTD

System

An object of the system according to the embodiment is to analyze a cry or an action of a pet and accurately convert an intention or an emotion into a human language.SOLUTION: A system includes a collection unit, an analysis unit, an estimation unit, and a conversion unit. The collection unit collects a cry or an action. The analysis unit analyzes the data collected by the collection unit. The estimation unit estimates an intention or a feeling on the basis of an analysis result obtained by the analysis unit. The conversion unit converts the intention or the feeling estimated by the estimation unit into human language.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

A teaching model for nursing care of voice tremor

ActiveCN224287680UImprove the level of clinical skills trainingincrease awarenessCosmonautic condition simulationsEducational modelsPneumothoraxBronchial tube
This utility model discloses a voice tremor nursing teaching model in the field of nursing teaching aids. It includes a three-dimensional humanoid model with a left lung model and a right lung model positioned at the chest area. The left lung model is covered with a silicone membrane, forming a pleural cavity model. A catheter is located at the top of the left lung model, with one end extending into the pleural cavity model. Air or liquid is injected into the pleural cavity through the catheter to simulate pneumothorax or pleural effusion. The right lung model includes a detachable and replaceable lung cavity model, lung tumor model, and bronchial obstruction model. A bronchial model is located between the left and right lung models, with a microphone connected to the top of the bronchial model. The two forked ends of the bronchial model at the bottom are connected to the left and right lung models, respectively. This device can replace different models to simulate pathological states, allowing students to intuitively experience the differences in voice tremor under different pathological conditions, improving teaching effectiveness and students' clinical skills training level.
Owner:YANGZHOU UNIV

A humanized voice intervention method for inhalation error of an oral-nasal aerosol dispenser

PendingCN122337174AGuaranteed therapeutic effectRealize humanized voice interventionAerosol drug deliveryPatient suctioning
This invention discloses a humanized voice intervention method for inhalation errors in nasal and oral aerosol delivery devices, aiming to address the lack of humanization in existing voice intervention methods, which leads to patient misunderstanding and resistance. The method includes: establishing a patient-doctor corpus by combining expert experience, literature, and guidelines; fine-tuning a language parsing model to obtain an error-sensitive parsing model; using the error-sensitive parsing model to obtain a total intervention feature library and an urgency level sub-library; establishing a two-level query mechanism for patient inhalation errors to query intervention text in the urgency level sub-library; and fine-tuning the voice generation model based on the aforementioned patient-doctor corpus, incorporating new tokens, and evolving through prompting learning into a patient-friendly intervention voice generation model capable of generating intervention voices that are appropriate for the urgency level. The intervention voices generated by this invention can form a humanized voice intervention that approximates the tone of clinical medical staff, and can adapt the tone according to the urgency level of the error, effectively improving the patient's understanding and acceptance of the intervention suggestions.
Owner:CHILDRENS HOSPITAL OF CHONGQING MEDICAL UNIV

Ear-nose-throat specialized clinical teaching effect evaluation system based on artificial intelligence

The invention provides an artificial intelligence-based ear-nose-throat specialized clinical teaching effect evaluation system. The system comprises a natural language processing module, a computer vision analysis module, a virtual patient evaluation module, an otoscope image acquisition device, an image recognition module, an intelligent scoring and feedback module, a literature retrieval scoring module, a data management module and a visual presentation module. According to the system, multi-modal data of students in the processes of voice answering, skill operation, virtual inquiry, otoscope image recognition and literature retrieval are collected, automatic analysis and scoring are carried out, an ability evaluation report and personalized feedback are generated, standardization and intelligence of teaching evaluation are achieved, and the teaching efficiency is improved. The comprehensive ability evaluation device is suitable for comprehensive ability evaluation in the ear-nose-throat specialized medical teaching process. The system has the advantages of objective scoring, efficient processing, visual feedback and the like.
Owner:ZHANGJIAGANG FIRST PEOPLES HOSPITAL

A device for collecting exhaled gas in stages

The present invention relates to the technical field of exhaled gas sampling, and discloses an exhaled gas phased collection device, including a host system and a power supply module; the host system includes a gas drying module, an information interaction module, a gas phased collection module, a main control module and a data processing module. The information interaction module is used to select the exhaled gas collection mode and the interaction of voice and display information. The main control module controls the gas phased collection module according to the selected exhaled gas collection mode to simultaneously collect single or multiple selected exhaled gases in the ambient gas, dead space gas, and alveolar gas in phases, detect the gas collection status, and the installation status of the disposable mouthpiece and air bag during the exhaled gas collection process, and the inflation status of the air bag at the gas collection port. The above-mentioned exhaled gas phased collection device is intended to divide the gas exhaled by the human body into different stages for collection, provide technical means for collecting the corresponding stage gas for exhaled gas detection of different diseases, and improve and promote the development of disease exhaled gas detection technology.
Owner:CHONGQING UNIV

Adaptive wavelet transform-based easily-confused emotional feature extraction method

PendingCN121789725ARealize multi-resolution time-frequency feature extractionSpeech analysisFeature extractionAlgorithm
The invention provides an easily-confused emotion feature extraction method based on adaptive wavelet transform, and aims to improve the recognition performance of easily-confused emotions. According to the method, learnable parameters beta, gamma and a are introduced into a wavelet basis function, the time-frequency form of the wavelet basis function is dynamically adjusted, multi-scale time-frequency decomposition is carried out on input voice signals, two-dimensional tensor feature representation is obtained, and extracted features reserve time local changes of the voice signals and differences of different frequency bands in emotion distinguishing. Meanwhile, a projection gradient descent method is adopted to carry out constraint optimization on parameters, the stability and energy convergence of a primary function are ensured, and a gradient return updating mechanism is utilized to enable wavelet parameters to be adaptively optimized in the training process. According to the method, emotional characteristics with higher discrimination and robustness can be extracted, and the accuracy and generalization ability of emotion recognition are improved.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY

Cerebral stroke rehabilitation evaluation method and system based on voice multi-task learning

The invention relates to the technical field of rehabilitation assessment, in particular to a cerebral apoplexy rehabilitation assessment method and system based on voice multi-task learning, and the method comprises the following steps: carrying out the psychological emotion recognition of assessment voice data, and obtaining a cerebral apoplexy rehabilitation subjective assessment task used for perceiving the subjective rehabilitation feeling of an assessment object; performing physiological function recognition on the evaluation voice data to obtain a cerebral apoplexy rehabilitation objective evaluation task for perceiving the objective rehabilitation state of the evaluation object; and performing weight adaptive combination on the stroke rehabilitation subjective evaluation task and the stroke rehabilitation objective evaluation task by using a multi-task learning mechanism to obtain a stroke rehabilitation evaluation model. According to the method, objective function states and subjective emotional feelings fed back from the voices of the patient can be taken into consideration during rehabilitation evaluation, pain and discomfort of the stroke patient can be perceived during functional evaluation, the individual accuracy of stroke rehabilitation evaluation is improved, and the user experience is enhanced.
Owner:THE FIRST AFFILIATED HOSPITAL OF WANNAN MEDICAL COLLEGE (YIJISHAN HOSPITAL OF WANNAN MEDICAL COLLEGE)

Speech anomaly detection method for screening Parkinson's disease in proactive period

PendingCN121459853ASpeech analysisPatient complianceVoice abnormality
The invention discloses a voice anomaly detection method for pre-driving period screening of Parkinson's disease, which is mainly used for performing non-invasive pre-driving period screening of the Parkinson's disease on voice information of a subject, analyzing by using daily voice, and discovering early signs of the Parkinson's disease without invasive examination such as blood drawing, imaging and the like. The method provides a new way for large-scale population screening, reduces the detection cost, is good in patient compliance, is high in result interpretation, is combined with a knowledge graph or feature importance analysis, provides a basis for an anomaly detection result, enables a clinician to understand the correlation between voice anomaly and Parkinson pathology, and improves the detection accuracy. And the credibility of the AI result is improved.
Owner:GYENNO TECH

Filter element for respiratory gas conditioning with speech function

The invention, which relates to a filter element (1) for respiratory gas conditioning with a speech function, is based on the objective of providing a filter element (1) for use with a tracheotomized or laryngectomized person, wherein the filter element (1) is robust, easy to operate, and inexpensive to manufacture. This objective is achieved by the housing (2) comprising a first housing part (5) and a second housing part (6) connected to the first housing part (5), the first opening (3) being arranged in the second housing part (6), the first housing part (5) being a single piece comprising an outer ring (7), a central upper part (8), and elastically deformable webs (9) arranged between the outer ring (7) and the upper part (8).
Owner:PRIMED HALBERSTADT MEDIZINTECHN

Multifunctional swallowing training and sound production auxiliary exerciser

The invention discloses a multifunctional swallowing training and sound production auxiliary exercising device, and belongs to the technical field of sound production training devices. Comprising an arc-shaped plate fixed to the lower jaw of a user, a supporting plate fixed to the chest of the user, a fixing column connecting the arc-shaped plate and the supporting plate and a rear neck adjusting belt, and a pickup, a displayer, a display lamp and a control terminal are integrated. The sound pick-up collects voice signals, the displayer displays volume, and the display lamp feeds back sound production quality through the control terminal. A piezoelectric sensor is embedded in the supporting plate to monitor the thoracic expansion amplitude; when chest breathing is detected, the clavicle heating module is linked to heat up, phrenic nerve reflex is activated through thermal stimulation to guide abdominal breathing, vocal cord closure is improved, and the risk of aspiration by mistake is reduced. The equipment is integrated with a rotating assembly to realize multi-angle adjustment and automatic switching of the display; a multi-frequency vibration module and a temperature control heating module are arranged in the arc-shaped plate, and targeted muscle stimulation is provided. Through mechanical linkage and multi-mode feedback, the breathing-sounding-swallowing cooperative training efficiency is optimized, and the user suitability and the rehabilitation effect are improved.
Owner:SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)

Morning check device

ActiveCN309673867SHuman bodySpeech sounds
1. The name of the design product: morning inspection device. 2. The use of the design product: health detection in schools and other places, with functions such as face recognition, oral cavity photographing and temperature measurement. 3. The design points of the design product: the combination of shape and pattern. 4. The picture or photo that best indicates the design points: perspective view. 5. Other circumstances that need to be explained: application example: the morning inspection device is fixed at the entrance of each class. After sensing the human body information of the students, the instruction triggers the voice module to play the oral temperature measurement prompt. The face image of the student is recognized, and the oral temperature of the student is sensed and the oral temperature information is obtained. If the oral temperature is abnormal, an alarm prompt is played.
Owner:陈宇煊

Electric motor control methods and apparatuses, and device and storage medium

Disclosed in the embodiments of the present application are electric motor control methods and apparatuses, and a device and a storage medium. A method comprises: driving an electric motor to vibrate, wherein the vibration is used for generating a sound while implementing a cleaning operation. When driven, the electric motor can output a sound while implementing an oral cleaning operation, so that a user of an oral care device can listen to music while cleaning the oral cavity, or can hear a prompt speech while cleaning the oral cavity.
Owner:GUANGZHOU STARS PULSE CO LTD

Selection of speech features for building a model to detect medical conditions

Providing a selection of speech features for building a model to detect medical conditions. [Solution] A mathematical model can be trained to diagnose a person's medical condition by processing the acoustic features 1021 and linguistic features 1022 of a person's voice. The performance of the mathematical model can be improved by appropriately selecting the features 1021 and 1022 used with the mathematical model. Features 1021 and 1022 can be selected by calculating a feature selection score 1031 for each acoustic feature 1021 and each linguistic feature 1022, and then using the score 1031 to select the feature 1021 or 1022 that gave the highest score 1031.
Owner:CANARY SPEECH LLC

A post-processing method, device and equipment for improving speech perception and a medium

ActiveCN119446163BSpeech analysisSpeech soundsSpeech perception
The application discloses a post-processing method and device for improving voice perception, equipment and medium. The pre-configured contrast stretching maximum value can limit the degree of contrast stretching, preventing signal distortion caused by excessive stretching. When determining the target contrast stretching parameter of any frequency, the pre-set contrast stretching preset parameter corresponding to the frequency is divided by the pre-determined contrast stretching maximum value, which can effectively reduce the energy of the non-sensitive frequency of the human ear and maintain the energy of the sensitive frequency of the human ear. Subsequently, the amplitude spectrum of the current voice frame is stretched by applying the target contrast stretching parameter corresponding to each frequency to obtain the stretched amplitude spectrum. The stretched amplitude spectrum maintains the energy of the sensitive frequency of the human ear while reducing the energy of the non-sensitive frequency, thereby avoiding excessive stretching of the amplitude of the sensitive frequency of the human ear and the problem of numerical overflow.
Owner:BEIJING UNISOUND INFORMATION TECH CO LTD

A bionic oral cavity sound production device, a robot and a dynamic bionic oral cavity sound production control method

The application provides a bionic oral cavity sound production device, a robot and a dynamic bionic oral cavity sound production control method. The bionic oral cavity sound production device comprises a sound production element, the sound production element is arranged in a sound production box, the sound production box comprises a movable component, and the movable component is used for changing the form of the internal space of the sound production box. The dynamic bionic oral cavity sound production control method comprises the following steps: obtaining semantic information of a target pronunciation; determining a target phoneme sequence according to the semantic information; querying a preset phoneme-cavity form mapping database according to the phoneme sequence to obtain a movement parameter sequence of the movable component; and adjusting the spatial form of the sound production box according to the movement parameter sequence. The application drives the sound production box to change the form through the movable component, realizes dynamic adjustment matched with the physiological process of human sound production, and has the advantages of improving the naturalness of speech and the accuracy of pronunciation by controlling the cavity form in real time.
Owner:SHANGHAI TODAY XINDONG TECHNOLOGY CO LTD

A speech-based emotion recognition method

PendingCN122347964APhonic TicLinear predictive coding
The application relates to the technical field of speech recognition, in particular to a speech-based emotion recognition method, which comprises the following steps: according to an input speech signal, frame division and windowing are carried out. The application can obtain more profound insights into emotional expression by decomposing the input speech signal into two physical sources of glottal excitation and vocal tract response for independent modeling and analysis, obtaining a vocal tract transfer function set via linear predictive coding operation, and applying inverse filtering to the original speech signal to reconstruct an approximate glottal pulse sequence, effectively stripping the influence of vocal tract resonance on the signal, so that perturbation parameters, open quotient and closed quotient and the like representing vocal cord vibration patterns can be directly calculated, at the same time, the vocal tract transfer function set is used to identify and track the dynamic trajectory of the formant, and the phase difference cosine mean between adjacent frames is combined to quantify the sound production stability, and subtle dynamic adjustment and control stability of sound production organs such as the oral cavity and tongue position caused by emotional changes are captured.
Owner:SHANGHAI LIXIN UNIV OF ACCOUNTING & FINANCE