Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Speech Acoustics" patented technology

The acoustic aspects of speech in terms of frequency, intensity, and time.

Computer-aided senile language erosion assessment method and assessment system

The invention discloses a computer-aided old-age language erosion assessment method and assessment system, and relates to the technical field of old-age health assessment, and the method comprises the following steps: S1, collecting the voice data, text input data and interactive behavior data of an old-age user; s2, preprocessing the data, and extracting voice acoustic features, language structure features and cognitive behavior features; s3, inputting the extracted features into a pre-trained language erosion evaluation model, and outputting a language ability score and an erosion type classification result; and S4, generating a visual evaluation report, wherein the visual evaluation report comprises the language ability degradation degree, key obstacle points and intervention suggestions. According to the method, the score and classification result is automatically output through the multi-modal data acquisition and pre-trained deep learning model, so that the evaluation time is greatly shortened, the subjective deviation is eliminated, the result objectivity is ensured, the problems of low manual evaluation efficiency and high subjectivity are solved, and the effect of automatic evaluation is realized.
Owner:BEIJING FOREIGN STUDIES UNIVERSITY

WIFI-based indoor voice positioning method and device

The application discloses a WIFI-based indoor voice positioning method and device. The method utilizes existing terminal equipment and is based on voice acoustic principles. The corresponding terminal equipment can be controlled to emit and receive voice information in a WIFI communication mode, so that the position of the terminal equipment is determined, and the positioning of a sound source is completed. The indoor positioning operation is simplified, the cost of system deployment is reduced, the accuracy of indoor positioning results is increased, and the convenience of indoor positioning in application is improved.
Owner:FOSHAN VIOMI ELECTRICAL TECH

Anomaly processing method and device based on multi-dimensional classifier chain, equipment and medium

This invention relates to the field of intelligent decision-making technology and can be applied to business scenarios such as fintech and healthcare. It discloses an anomaly handling method, apparatus, device, and medium based on a multi-dimensional classifier chain, comprising: acquiring a voice data stream and generating a text data sequence with role identifiers, while simultaneously extracting voice acoustic features; performing semantic parsing on the text data sequence and determining dialogue structure features; fusing three types of features to form a unified input feature vector; performing chain-like prediction based on label dependencies using a multi-dimensional classifier chain model and outputting multi-dimensional classification prediction results; determining the anomaly level based on the prediction results and matching the target strategy to execute the corresponding action. This invention improves the accuracy of multi-dimensional prediction by fusing multi-source features and utilizing a classifier chain model to model label dependencies, thereby more reliably identifying anomaly levels and triggering matching strategies, achieving more accurate customer service analysis and recovery processing.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Hearing-impaired voice conversion and generation method based on multi-language self-adaption

PendingCN121483225ABiological modelsSpeech recognitionIntelligibility (communication)Speech Acoustics
The invention relates to a hearing-impaired voice conversion and generation method based on multi-language self-adaption, and belongs to the field of voice processing and artificial intelligence. According to the invention, hearing-impaired voices in multiple languages are converted into standard voices, so that normal communication with hearing-impaired people is facilitated. According to the method, acoustic features of hearing-impaired speech are extracted and aligned with target language phonemes, the acoustic features are input into a speech adaptation model, standard speech features are generated, multi-language training is carried out based on the standard speech features and target language corpus, and the training result is used for generating target language natural speech. According to the invention, the intelligibility of hearing-impaired voice can be improved, so that the voice is closer to natural pronunciation. According to the invention, rapid adaptation of multiple languages can be realized, and the cost of cross-language training is reduced. The system can be applied to education, medical rehabilitation, cross-border communication and other scenes, and provides effective voice communication assistance for hearing-impaired people.
Owner:BEIJING INST OF COMP TECH & APPL

Smart home user behavior prediction method based on semantic analysis

The invention discloses a smart home user behavior prediction method based on semantic analysis, and relates to the technical field of natural language processing. The method comprises the following specific steps: data acquisition: respectively acquiring user interaction data, equipment operation historical data and user interaction feedback data through a voice acquisition device, a user input interface, a smart home gateway and a feedback interface, and uniformly storing the data after formatting processing; the speech acoustic features and the text features are extracted on the emotional semantic recognition level, the pre-training model is used for outputting the user emotional state, the defect that in the prior art, emotional perception is lacked is overcome, in the semantic ambiguity resolution aspect, a learning mechanism with user interaction feedback data as an optimization signal is adopted, and the user experience is improved. When the same number of times of correction of the same fuzzy instruction by the user reaches a preset threshold value, the semantic mapping rule is automatically generated and written into the rule base, and the semantic analysis priority can be adjusted according to the family structure change.
Owner:ZHEJIANG UNIV OF FINANCE & ECONOMICS

Model training method, speech recognition method, device, medium, and program product

ActiveCN121565156BSpeech recognitionSpeech AcousticsAutomatic speech
This invention discloses a model training method, a speech recognition method, a device, a medium, and a program product. The method includes: decoupling sample speech information by an acoustic feature encoder to determine sample speech content features and sample speech acoustic style features; performing acoustic feature discrimination on the sample speech content features using an acoustic feature discriminator to determine adversarial loss; fusing the sample speech content features and sample speech acoustic style features using a feature fusion module to determine fused sample speech features; extracting text information from the fused sample speech features using an automatic speech recognition decoder, and determining the automatic speech recognition loss based on the extracted text and speech recognition training samples; and substituting the adversarial loss and the automatic speech recognition loss into a pre-constructed total loss function to complete the training of the speech recognition model. This enables the speech recognition model to recognize speech data acquired with different accents and in different environments.
Owner:SUZHOU KEDA TECH +1

A speech speaker conversion point detection method and a model training method thereof

PendingCN122369471APattern recognitionNoise
This invention discloses a method for detecting speaker transition points and its model training method. The detection method includes: extracting acoustic features of the speech, inputting them into a pre-trained multi-task coupled learning model to obtain a probability sequence of transition points, and determining the location of the transition points through post-processing. The innovation of this model lies in: employing a shared speaker embedding extraction network and introducing a frame-by-frame feature gating unit in its speaker transition point regression branch to suppress noise interference using speech activity probability; simultaneously, optimizing the model training using a regression loss based on bell-shaped distribution soft labels to achieve refined boundary localization. The corresponding training method clarifies the model construction, soft label generation, and multi-task joint training process. This invention solves the problems of low accuracy in speaker transition point detection, sensitivity to silence and noise, and inaccurate boundary localization in existing technologies, effectively improving the accuracy and robustness of detection.
Owner:JIANGSU UNIV OF SCI & TECH

A system and method for predicting the risk of alzheimer's disease

PendingCN122436204ABlood biomarkersData acquisition
The application discloses an Alzheimer's disease risk prediction system and method, and belongs to the field of biomedical detection and artificial intelligence technology, and the system comprises: a multi-modal data acquisition module for acquiring biological samples, images, voices and demographic information; a biomarker detection module for detecting the expression levels of p-tau217 and GFAP; an image analysis module for extracting fat metabolism parameters in PET / CT images; a voice feature extraction module for extracting acoustic features; and a risk prediction module with an integrated prediction model combining multiple machine learning models, which takes the above multi-modal data as input variables, comprehensively analyzes and outputs risk probabilities. The application combines blood biomarkers, image metabolism parameters, voice acoustic features and demographic information, and through multi-dimensional data complementation and correction, the accuracy, convenience and reliability of early risk prediction of Alzheimer's disease are significantly improved.
Owner:GUANGZHOU NANFANG COLLEGE

Homophone error correction method and device and storage medium

The invention provides a homophone error correction method and device and a storage medium, and belongs to the technical field of character error correction, and the method comprises the steps: importing to-be-corrected voice data, original voice data and real text data; performing voice recognition on the to-be-corrected voice data and the original voice data to obtain to-be-corrected text features and original text features; constructing a training model, and performing model analysis on the training model according to the original text features and the real text data to obtain a homophone error correction model; and performing error correction analysis on the to-be-corrected text features through the homophone error correction model to obtain a homophone error correction result. The method improves the context association capability, does not need to depend on shallow text matching, improves the adaptability of a low-resource scene and a dynamic scene, fully considers the association of voice acoustic features, improves the error-tolerant rate of a rare homophone combination, and also can remarkably improve the homophone error correction accuracy.
Owner:SI-TECH INFORMATION TECH CO LTD +1

Mental electrophysiology assessment method and system based on video follow-up visit system

The invention relates to the technical field of mental health assessment, in particular to a mental electrophysiology assessment method based on a video follow-up visit system, which comprises the following steps of: acquiring doctor-patient double-channel audio and video data through the video follow-up visit system; preprocessing the collected data and extracting multi-modal features including facial behavior features, voice acoustic features and motion dynamics features; and constructing a spatial model of the mental pathological state of the patient based on the extracted multi-modal features, wherein the spatial model describes a dynamic evolution process of a multi-dimensional state variable through a stochastic differential equation. According to the method, a multi-modal feature extraction system of doctor-patient dual-channel audio and video data is constructed, and then a psychiatric pathology state space model based on a stochastic differential equation is established, so that the problems that subjective scale and static observation are mostly adopted in a traditional mental assessment method, and due to lack of objective quantitative indexes and a dynamic tracking mechanism, the accuracy of mental assessment is poor are solved. Therefore, the evaluation result is greatly influenced by subjective experience of doctors, and the dynamic change of symptoms cannot be captured.
Owner:河南医药大学第二附属医院(河南省精神病医院)

Text-to-voice method and device, computer equipment and storage medium

The invention relates to the field of artificial intelligence, is applied to financial and medical scenes, and discloses a text-to-voice method and device, computer equipment and a storage medium, and the method comprises the steps: receiving a target text, a voice prompt and an emotion label; performing semantic coding on the target text to extract text semantic features; performing acoustic coding on the voice prompt to extract voice acoustic features; based on the emotion label and the text semantic feature, generating a dynamic emotion feature through a time sequence emotion model; performing cross-modal fusion on the voice acoustic features and the dynamic emotion features to obtain cross-modal alignment features, and inputting the cross-modal alignment features and the text semantic features into a pre-training language model for feature fusion and voice token prediction to obtain an initial voice token sequence; and performing semantic alignment and spectrum conversion on the initial voice token sequence to generate a target voice waveform. According to the invention, high-quality voice with natural emotional circulation, controllable timbre and high conformity with professional contexts can be synthesized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent multi-mode voice conference interaction method and system

The invention relates to the field of voice interaction, in particular to an intelligent multi-mode voice conference interaction method and system. The method comprises the following steps: collecting historical voice signal data, carrying out feature extraction and analysis on original voice signal data, generating a voice acoustic vector, analyzing the voice acoustic vector, and constructing a voice baseline vector; collecting conference voice data, analyzing the conference voice data, and constructing a conference semantic state diagram; acquiring real-time voice signal data, and analyzing the real-time voice signal data based on the voice baseline vector to obtain a voice deviation parameter; and on the basis of the voice deviation parameter and the voice baseline vector, adjusting voice recognition adaptation, generating an adjustment strategy, stopping recording the conference semantic state diagram before conference interruption when the conference interruption is detected, and performing reconnection processing on the conference semantic state diagram after conference reconnection. According to the invention, the voice recognition suitability and the semantic understanding continuity can be improved.
Owner:JIAN XIANGE ACOUSTIC ELECTRONIC CO LTD

Music generation method, electronic device, storage medium and computer program product

The embodiment of the invention discloses a music generation method, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: performing feature extraction on voice information of a target object to obtain voice semantic features and voice acoustic features of the target object; generating a human voice audio based on the voice semantic features and the voice acoustic features; based on the voice audio, generating a composing and music-making audio; and generating target music based on the voice audio and the arrangement music audio.
Owner:MIGU CO LTD +1

Speech synthesis method and electronic equipment

PendingCN121938342ASpeech synthesisSynthesis methodsSpeech Acoustics
The invention provides a speech synthesis method and electronic equipment, and the method comprises the steps: obtaining a first feature inputted by a speech synthesis model, the first feature being obtained based on multi-modal data processing; obtaining a second feature generated by the speech synthesis model based on the first feature reasoning, wherein the second feature is a speech semantic feature associated with the speech acoustic attribute; based on the first feature and the second feature, performing iterative prediction through a reasoning prediction model to obtain at least one candidate acoustic unit of respective corresponding output positions of continuous target iteration prediction; based on the plurality of candidate acoustic units predicted by continuous target iterations, screening a plurality of target acoustic units with continuously adjacent output positions through a speech synthesis model to synthesize target speech; wherein the model parameters of the reasoning prediction model are smaller than the model parameters of the speech synthesis model.
Owner:LENOVO (BEIJING) LTD

An operator-oriented multi-modal AI large model voice call abnormal risk real-time quality inspection method

PendingCN122340215ACommunications securityAbnormal voice
This invention belongs to the field of communication security and artificial intelligence quality inspection technology, and relates to a real-time quality inspection method and system for abnormal voice calls using a multimodal AI large-scale model for telecom operators. The method receives call information, industry information, and compliance script information; aligns and verifies detailed call records with audio recordings; and performs speech recognition and feature extraction. It combines text semantic features, speech acoustic features, call behavior features, voiceprint biometric features, and deviation features of the actual call content from the reported information to construct multimodal evidence units and evidence sequences. It uses an AI large-scale model to identify candidate risk segments, and uses recording integrity identifiers and reported deviation features as pre-gating to determine the risk level of candidate risk segments and perform review and triage, outputting a real-time quality inspection evidence package. This invention can improve the evidence completeness, risk identification accuracy, and traceability of handling in the quality inspection of abnormal voice calls for telecom operators.
Owner:HANGZHOU AITA TECH CO LTD

Speech synthesis method and device for long text data, equipment and medium

PendingCN121811847ASpeech synthesisSynthesis methodsSpeech Acoustics
The invention relates to the technical field of speech semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a speech synthesis method, device and equipment for long text data and a medium, and the method comprises the steps: obtaining the long text data and corresponding historical speech data, and extracting global text semantic features and speech acoustic features; re-characterizing the global text semantic feature and the voice acoustic feature to obtain a text potential representation and a voice potential representation; performing cross-modal fusion on the text potential representation and the voice potential representation to obtain fusion features; retrieving a target context feature from the historical context information, and fusing the target context feature with the long-term memory feature to obtain a target long-term memory feature; performing attention interaction processing on the target long-term memory feature and the local text semantic feature to obtain a context enhancement feature; and generating a target synthetic speech according to the context enhancement feature. According to the invention, the accuracy of long text data speech synthesis can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech reconstruction method and device based on intracranial neural electrical signals and electronic equipment

PendingCN122392482ASpeech reconstructionEngineering
The present application relates to a kind of speech reconstruction methods, device and electronic equipment based on intracranial nerve electric signal, the method comprises: the neural feature time sequence stream corresponding to the intracranial nerve electric signal continuously collected by the language-related brain area of subject is acquired;Neural feature time sequence stream is input to a flow causal decoding model, to generate corresponding speech acoustic feature time sequence stream in real time;Wherein, flow causal decoding model is the causal deep neural network model trained based on sample data, and sample data includes sample neural signal feature sequence and its corresponding sample speech acoustic feature sequence;Speech acoustic feature time sequence stream is integrated into speech waveform stream and output.The present application realizes real-time speech synthesis of millisecond level delay by flow causal decoding model and dynamic block strategy, solves the high delay problem of existing scheme, simultaneously avoids auditory feedback interference by causal mask mechanism, can adapt to the application demand in silent scene.
Owner:AFFILIATED HUSN HOSPITAL OF FUDAN UNIV

Emotional dialogue speech synthesis method and system based on heterogeneous subgraph comparative learning

PendingCN121600906ASpeech synthesisSynthesis methodsConversational speech
The invention belongs to the technical field of voice data processing, and particularly discloses an emotional dialogue voice synthesis method and system based on heterogeneous subgraph comparative learning. The method comprises the following steps: obtaining a dialogue voice data set and carrying out cleaning preprocessing, carrying out characterization processing on dialogue voice data, constructing a heterogeneous dialogue graph, carrying out context emotion modeling based on heterogeneous subgraph comparison learning, and carrying out voice synthesis based on emotion perception. According to the method, through heterogeneous subgraph comparative learning, context emotion clues can be focused, and the problem that emotion and context are disjointed in a traditional method is solved; continuous emotion intensity and dimension changes are supported through spherical coordinate coding, and the limitation of discrete emotion labels is broken through; the heterogeneous graph can dynamically integrate multiple rounds of dialogue information and can adapt to emotional evolution of a long dialogue scene; the emotion parameters and the voice acoustic features are directly mapped, so that mismatching of the emotion and the voice features is effectively avoided.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

An office worker emotion recognition method and system based on visible light and voice signals

PendingCN122451571AFacial movementVoice source
The application discloses an office staff emotion recognition method and system based on visible light and voice signals, relates to the technical field of computer vision and voice signal processing, and the method is characterized in that: face images and voice signals are synchronously collected, first, it is judged whether a staff is in a silent state or a speaking state, if the staff is in the silent state, emotion is recognized by analyzing facial motion unit features of the whole face, if the staff is in the speaking state, the voice source is confirmed by lip reading verification, and then emotion is recognized by fusing facial motion features based on only eyebrow and eye regions and voice acoustic features. The application effectively solves the interference problem of mouth movement on expression recognition when speaking, and improves the accuracy and robustness of emotion state monitoring.
Owner:HUNAN INST OF TECH

Model training method, speech recognition method, device, medium and program product

ActiveCN121565156ASpeech recognitionSpeech AcousticsAutomatic speech
The embodiment of the invention discloses a model training method, a speech recognition method, equipment, a medium and a program product. Comprising the steps that an acoustic feature encoder decouples sample voice information, and sample voice content features and sample voice acoustic style features are determined; performing acoustic feature discrimination on the sample voice content features through an acoustic feature discriminator, and determining adversarial loss; fusing the sample voice content features and the sample voice acoustic style features through a feature fusion module to determine fused sample voice features; performing text information extraction on the fused sample speech features through an automatic speech recognition decoder, and determining automatic speech recognition loss according to a sample text extraction result and a speech recognition training sample; and substituting the confrontation loss and the automatic speech recognition loss into a pre-constructed total loss function to complete training of a speech recognition model. Therefore, the voice recognition model can recognize the voice data which have different accent and are acquired in different environments.
Owner:SUZHOU KEDA TECH +1

A multi-modal speech interaction large model training method and system based on speech acoustic feature regulation, a terminal device, and a medium

ActiveCN120954388BSpeech recognitionSpeech synthesisSpeech comprehensionModal voice
The application discloses a kind of based on multi-modal speech interaction big model training method, system, terminal equipment and medium of voice acoustic feature regulation, it is related to multi-modal speech interaction technical field, the method includes: obtaining the text token of text training sample and constructs corresponding voice token, obtains the pre-training data for converting text token into voice token;Combining multi-modal input sample and pre-training data, construct fine-tuning training data for speech understanding and dialogue generation;Using pre-training data constructs and pre-trains basic model;Based on pre-training basic model, build multi-modal speech interaction big model, with fine-tuning data training, so that it can be based on multi-modal input regulation voice acoustic feature and output voice.The application is trained by the alignment and stage of text token and voice token, realize the fine regulation of voice acoustic feature, improve long speech coherence and interactive naturalness, efficiently give model controllable timbre, the speech interaction ability of emotion.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)