Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1295 results about "Speech characteristics" patented technology

Speech characteristics are features of speech that in varying may affect intelligibility. They include: Articulation. Pronunciation. Speech disfluency. Speech pauses. Speech pitch. Speech rate.

Intelligent multi-mode virtual digital human interaction system based on AI language large model, interaction method and application

The invention discloses an intelligent multi-modal virtual digital human interaction system based on an AI language large model. The system comprises a high-authenticity face generation module; the high-authenticity face generation module uses an AdaAN network, based on adaptive feature fusion and voice driving and time sequence modeling of voice features, feature information related to voice is extracted, the extracted voice features are processed through a deep neural network, it is ensured that the voice and facial expressions are highly aligned in time and space, and the face recognition accuracy is improved. Collecting a bio-electricity signal, mapping the signal to facial muscle movement, generating a final facial expression, and interacting with a user; the system further comprises an intelligent interaction module, a training optimization and efficient generation module, an efficient integration module, a multi-modal data acquisition module, an AI large model core processing module, a digital human image generation and driving module, an interaction scene adaptation module and a feedback optimization module. The invention further discloses a multi-mode digital human interaction method which has wide application value.
Owner:EAST CHINA NORMAL UNIV

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Speech feature processing method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice feature processing method, device, equipment and medium. Performing time resolution analysis based on the fused Mel band energy to generate a multi-scale Mel spectrum amplitude value, and performing nonlinear transformation on the multi-scale Mel spectrum amplitude value according to the noise intensity parameter to generate a noise suppression Mel component; and generating a perception weighting coefficient according to an auditory perception model, and executing frequency domain energy adjustment on the noise suppression Mel component to generate Mel spectrum representation. On the basis of frequency resolution self-adaption, time resolution dynamic adjustment and auditory perception modeling, nonlinear transformation and perception weighting processing are applied to the multi-scale Mel spectrum amplitude value, the influence of noise interference on voice features can be effectively reduced, and the key information retention capacity of voice signals is enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Precise international communication digital human real-time dialogue method fused with multi-modal technology

The invention discloses an accurate international communication digital human real-time dialogue method fused with a multi-modal technology. The method comprises the following steps: S1, constructing a digital human image and tone; s2, propagation content generation and problem guidance; s3, semantic analysis and intention clarification based on the real-time voice dialogue; s4, geographic preference modeling and path planning; s5, cross-context propagation content generated based on retrieval enhancement is generated; s6, visually displaying the propagation content; and S7, carrying out digital human-driven multi-language propagation content real-time output and feedback closed-loop optimization. Through accurate utterance expression analysis, accurate international propagation problem recommendation is provided, cross-context propagation content generation based on semantic understanding is realized, digital people with voice features and visual images are constructed, real-time dialogue interaction of users is realized, and user experience is improved. The system can carry out geographic modeling according to the region where the accurate problem is located, language preference and propagation object culture characteristics, and differential propagation path planning is achieved.
Owner:HUNAN NORMAL UNIVERSITY

Aircraft cabin environment personalized adjustment method based on sentiment analysis

The invention discloses an aircraft cabin environment personalized adjustment method based on sentiment analysis, and the method comprises the steps: collecting the facial expression, voice waveform and environment parameters of a passenger in real time through a cabin multi-source sensor, and generating a standardized physiological signal matrix and anonymized voice features through noise reduction and feature extraction; inputting the physiological signal matrix and the voice features into a pre-trained deep learning model, outputting an emotion index and a classification label, and updating model parameters through a federal learning framework; dynamically generating temperature, humidity and oxygen concentration adjusting instructions and dynamic weights by adopting a fuzzy reasoning system in combination with the emotion indexes and passenger preset preferences; environment adjustment is executed through a closed-loop control system, and environment parameter errors are fed back; synchronously updating a fuzzy inference system rule base and deep learning model parameters by utilizing reinforcement learning in combination with environmental parameter errors and emotion index changes, and completing optimization of a closed loop; according to the invention, real-time dynamic regulation and control and continuous adaptation optimization of the personalized cabin environment can be realized.
Owner:WENZHOU DOVER AVIATION IND GROUP CO LTD

Voice interaction task execution method and device based on large model, equipment and medium

The invention discloses a voice interaction task execution method and device based on a large model, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the voice recognition of voice features, obtaining a text character sequence, and carrying out the optimization of the text character sequence through a preset large language model, and obtaining a target text character sequence; performing entity recognition on the target text character sequence by using a preset large language model to obtain an entity recognition result, determining a relationship type among entities in the target text character sequence according to the entity recognition result, and constructing a knowledge graph according to the relationship type among the entities; generating an initial triple based on the knowledge graph, and optimizing the initial triple by using a preset large language model to obtain a target triple; and fusing the target triple with the initial knowledge graph, and executing a voice interaction task in the target voice interaction scene based on the updated knowledge graph. According to the method and the device, the efficiency and the accuracy of extracting the structured knowledge from the Chinese speech are improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Multi-modal driven virtual digital human face animation generation method and system

The invention discloses a multi-mode driven virtual digital human face animation generation method and system, and relates to the field of computer graphics, and the method comprises the steps: obtaining voice input and text input, and extracting voice features and text features; dynamically fusing the two features through an attention fusion model to generate facial expression and head posture control parameters, and dynamically adjusting contribution weights of a voice mode and a text mode to the control parameters by adopting a driving strategy of differentiation of facial upper expression and facial lower expression; and performing local deformation on the virtual digital human face image based on the control parameter to generate an initial animation frame, and performing refining processing on the initial animation by using a generative adversarial network to obtain a refined face animation. By means of the technical scheme, the natural and vivid virtual human face animation matched with the voice content and the text semantics can be generated.
Owner:LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

An ear-worn device is provided that operates to isolate and individually treat the received speech of a target speaker or multiple target speakers from an audio input signal detected in a multi-speaker environment. The ear-worn device uses a machine learning model that receives a voice signature of each of one or more target speakers as input signals, to identify and isolate the component of the audio input signal attributable to the target speaker(s). Once isolated, the target speaker's speech may be enhanced, de-emphasized, or otherwise processed in a manner desired by the wearer of the ear-worn device. The wearer may use an external electronic device, e.g., a phone, to select one or more target speakers in a conversation and / or configure various settings associated with processing the speech on the ear-worn device.
Owner:FORTELL RESEARCH INC

Vector voice interaction risk control method and system based on voiceprint features and lip synchronization

The invention relates to a borrower voice interaction risk control method and system based on voiceprint features and lip synchronization. The method comprises the following steps: synchronously acquiring a voice signal and face video stream data of a user through acquisition equipment; performing voice feature extraction on the voice signal to generate a voiceprint feature vector and a voice time sequence feature vector; lip motion feature analysis is carried out on the face video stream data, and lip motion track feature vectors representing lip dynamic changes are extracted; performing identity matching on the voiceprint feature vector and a pre-stored reference voiceprint database to generate a first matching degree score; performing cross-modal time sequence alignment analysis on the voice time sequence feature vector and the lip movement trajectory feature vector to generate a second matching degree score; and generating a comprehensive risk control scoring result according to the first matching degree score and the second matching degree score. According to the scheme provided by the invention, a multi-modal cross validation mechanism can be formed, and the limitation of a traditional single-modal detection scheme is effectively overcome.
Owner:FEIHU INTERACTIVE TECH BEIJING CO LTD

Artificial intelligence speech recognition system

The invention discloses an artificial intelligence speech recognition system, and the system comprises a multi-modal feature extraction module which employs an improved Conformer architecture to synchronously extract the time-frequency features and text embedding vectors of speech signals; the joint training module is used for performing joint optimization on ASR and NMT loss functions through an adversarial training strategy, learning voice recognition and machine translation tasks at the same time through joint training, and completing direct mapping from voice features to a target language; the context perception translation engine is used for integrating an attention mechanism of a pre-training language model, carrying out deep coding on the extracted speech features and generating cross-language semantic representation; the self-adaptive post-processing module is used for dynamically optimizing an output result by adopting a reinforcement learning framework, dynamically adjusting the output result according to a reward function, and optimizing translation quality and a speech synthesis effect; the dynamic language recognition module is a real-time language classifier based on a Wave2Vec 2.0 framework and is used for recognizing the language of the input voice in real time; and the incremental field adaptation module is used for quickly updating a field term library by using a LoRA fine tuning technology.
Owner:ANKANG UNIV

Automated generation of targeted feedback using speech characteristics extracted from audio samples to address speech defects

Provided herein are systems and methods for providing instructions for speech based on speech classifications of verbal communications from users. Ae computing system can generate speech characteristics for a first verbal communication using a first audio sample from a user. The computing system can determine, from a plurality of speech classifications, a first speech classification for the first verbal communication based on the speech characteristics. The computing system can select, from a plurality of actions, an action comprising modifying one or more of the speech characteristics to define an utterance for the user based on the first speech classification. The computing system can provide an instruction presenting a message to prompt the user to perform the utterance defined by the action selected from the plurality of actions. The efficacy of the medication that the user is taking to address the condition may be increased.
Owner:CLICK THERAPEUTICS INC

Multi-speaker dialogue voice analysis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a multi-speaker dialogue voice analysis method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-analyzed multi-speaker dialogue voice, determining a naturalness score based on an acoustic feature, and obtaining a multi-speaker dialogue voice analysis result; determining a semantic consistency score based on voice embedding and semantic embedding corresponding to a preset text, determining a speaker consistency score based on embedding of a plurality of speakers of the same speaker, determining an interaction rationality score based on voice alternate overlapping duration, determining a diversity score based on a variance of voice features, and fusing the scores, the comprehensive mass fraction is obtained. According to the invention, through quantitative evaluation of five dimensions of naturalness, semantic consistency, speaker consistency, interaction rationality and diversity, a comprehensive quality scoring system is established, so that the evaluation result simultaneously reflects voice fluency, content matching degree, identity stability, interaction rhythm rationality and feature richness.
Owner:PING AN TECH (SHENZHEN) CO LTD

Classroom emotion recognition method based on multi-view neural network

The invention discloses a classroom emotion recognition method based on a multi-view graph neural network, and relates to the technical field of deep learning and emotion recognition, and the method comprises the steps: respectively constructing a visual modal graph, a voice modal graph and a text modal graph based on visual features, voice features and text features; video segments are used as nodes in each modal diagram, and visual features, voice features and text features corresponding to each video segment are respectively used as node features of the corresponding modal diagram; respectively carrying out GAT coding on each modal diagram by utilizing a diagram attention network, and respectively updating node features in each modal diagram to obtain each updated modal diagram; performing multi-modal adaptive fusion on the basis of the updated modal diagrams to obtain multi-modal adaptive fusion features; and predicting the classroom emotion of the student by using the multi-modal adaptive fusion features. According to the method, the multi-view features are constructed by using the graph neural network, and the classroom emotions of the students are identified more comprehensively and accurately in combination with a multi-view feature fusion technology.
Owner:DATA SPACE RES INST

Digital human real-time interaction system and AI voice conversation chip circuit system thereof

The invention discloses a digital human real-time interaction system and an AI voice dialogue chip circuit system thereof, and the system employs a modular design, and comprises the following modules: a BaseReal module which is responsible for audio frame management, TTS service initialization, video recording and customized audio and video cycle management; the media stream management module has an audio stream processing function and is responsible for real-time transmission of audio and video streams; the speech recognition module is responsible for speech feature extraction and text conversion; the text-to-voice module is used for converting the collected audio data into a text; the lip shape synchronization module is used for realizing real-time lip shape synchronization animation generation through lip shape synchronization algorithm optimization and inheritance of a basic class; and the large model module comprises an LLM large model based on mass data fine tuning, an api interface or a workflow agent. The method has the effects of voice recognition, text-to-voice conversion, lip shape synchronization, real-time audio and video interaction, real-time translation and the like.
Owner:王蕊

Emotional voice generation method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses an emotional voice generation method, device, equipment and medium, and the method comprises the steps: obtaining a target text of a to-be-generated voice and a target emotional prompt text based on natural language description; inputting the target emotion prompt text into a pre-trained emotion encoder to obtain an emotion embedding vector; inputting the target text into a pre-trained text encoder to obtain a semantic embedding vector; inputting the emotion embedding vector and the semantic embedding vector into a pre-trained joint speech generation model, performing fused speech feature prediction on the emotion embedding vector and the semantic embedding vector, and generating a speech feature with an expected emotion; and decoding the voice features to generate target emotional voice corresponding to the target text. The speech emotion is controlled through the emotion prompt text described by the natural language, and the flexibility and synthesis effect of emotion speech generation interaction are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-mode emotion recognition method, system, electronic device and storage medium

Disclosed are a multi-mode emotion recognition method, a system, an electronic device, and a storage medium. The method includes obtaining a spectrogram of a voice to be recognized and a corresponding text and inputting the spectrogram and the text into a multi-mode emotion recognition model to obtain an emotion recognition result output by the multi-mode emotion recognition model. The multi-mode emotion recognition model is trained based on a sample spectrogram, and a corresponding sample text, and a sample emotion recognition result, and is configured to extract a feature from the spectrogram and the text by a self-attention mechanism to obtain the voice features and the text feature, fuse the text feature and voice feature to obtain a multi-mode fusion feature, and make an emotion classification decision to obtain an emotion recognition result based on the text feature, the voice feature, and the multi-mode fusion feature.
Owner:HUAZHONG NORMAL UNIV

Interactive multi-round dialogue digital human modeling system and method

The invention relates to an interactive multi-round dialogue digital human modeling method, which comprises the following steps of: extracting double-person multi-modal characteristics in an interactive multi-round dialogue scene, including voice characteristics and expression characteristics of a speaker in a current dialogue round and voice characteristics of a listener in a previous dialogue round in the current dialogue round; performing time sequence alignment and reinforcement on the extracted double-person multi-modal features based on a time dimension to obtain a combined feature sequence; according to the joint feature sequence, generating a voice text of the listener in the current dialogue round and synchronous expression parameters based on a codec fused with an attention mechanism; and generating a corresponding 3D facial animation frame sequence according to the expression parameters of the listener in the current dialogue round.
Owner:RENMIN UNIVERSITY OF CHINA

Generating speaker video and audio in multiple languages for videoconferencing

Systems and methods for generating speaker video and audio in multiple languages for videoconferencing are provided. For example, a computing device can access a speaker speech audio signal that includes a speaker speech in a first language, a video of the speaker and a translated speech audio signal of the speaker speech in a second language. The computing device generates, based on the translated speech audio signal, a converted translated speech audio signal that includes a speech in the second language having voice characteristics in the speaker speech. The computing device further generates a lip-synched speaker video based on the video of the speaker and the converted translated speech audio signal. Lip movements in the lip-synched speaker video correspond to the converted translated speech audio signal. The converted translated speech audio signal and the lip-synched speaker video are transmitted to a video conference provider configured to host the video conference.
Owner:ZOOM VIDEO COMM INC

Intelligent conference memo generation method based on robot

The invention provides an intelligent conference memo generation method based on a robot, and the method comprises the steps: collecting a multi-channel voice signal through a microphone array, and enhancing the voice of a target speaker through a beam forming technology; background noise is separated by adopting a self-adaptive filtering algorithm, and the voice signal quality is improved; speech features are extracted, language model parameters are adjusted, and a transliteration text is generated; analyzing the text structure, and extracting conference themes, participants and decision contents to form structured information; calculating semantic similarity among the knowledge graph nodes, and if the semantic similarity is higher than a preset threshold value, associating historical records to generate extension information; constructing a structured memorandum based on the extended information, organizing a conference theme, participants, decision contents and associated historical records, and generating an initial memorandum; and monitoring a memorandum editing operation, updating a knowledge graph node relationship, and generating a final memorandum document. According to the method, the accuracy, integrity and availability of conference records are effectively improved, and the conference efficiency is remarkably improved.
Owner:HUNAN HEXIN ANHUA BLOCKCHAIN TECH CO LTD

Laying hen voice recognition method and system fusing acoustic features and deep learning features

The invention provides a laying hen voice recognition method and system fusing acoustic features and deep learning features. The method comprises the steps of obtaining a to-be-recognized original audio signal and a voice recognition model; wherein the voice recognition model comprises a feature extraction network, a feature fusion network and a classification recognition network; performing feature extraction on the original audio signal by using the feature extraction network to obtain a spectrogram feature, a Mel-frequency cepstrum coefficient feature and a deep speech feature; the feature fusion network performs feature fusion on the spectrogram features, the Mel-frequency cepstrum coefficient features and the deep speech features by using a collaborative attention mechanism or a multi-head attention mechanism to obtain fused features; and inputting the fused features into a classification recognition network to obtain a voice recognition result. According to the method, the advantages of various characteristics can be fully utilized, and the sound signals are described and analyzed from multiple angles, so that the voiceprint of the laying hen is more accurately recognized, and the voiceprint recognition accuracy of the laying hen is remarkably improved.
Owner:BEIJING RES CENT FOR INFORMATION TECH & AGRI

Cross-domain voice classification method and device based on feature decoupling and multi-task learning

PendingCN120452429ASpeech recognitionPhonetic environmentData set
The invention relates to a cross-domain voice classification method and device based on feature decoupling and multi-task learning. The method comprises the following steps: firstly, acquiring a multi-data-domain voice file, and preprocessing the multi-data-domain voice file to obtain a cross-domain voice classification data set; then, constructing a cross-domain voice classification model which comprises a voice feature encoder module, a data domain classification module, a supervised comparative learning module and a multi-task classification module; then, a joint optimization loss function is constructed based on data field classification loss, supervised contrast learning loss and task classification loss, and a gradient descent algorithm is adopted to train and optimize the cross-domain voice classification model based on the cross-domain voice classification data set and the joint optimization loss function; and finally, inputting to-be-classified voice into the trained cross-domain voice classification model to obtain a voice corresponding category. The discrimination ability and generalization ability of the model in a cross-domain scene are significantly improved, so that the model still maintains high classification precision in a complex multi-source voice environment.
Owner:SICHUAN UNIV

Anger emotion recognition method based on video stream

The invention discloses an angry emotion recognition method based on a video stream, and belongs to the technical field of emotion recognition. The method comprises the following steps: preprocessing a video stream of a visitor to obtain a time-space aligned standardized facial image sequence, a 3D skeleton sequence and a voice segment; performing feature extraction on the standardized facial image sequence by using a 3D-ResNet34 network and a micro-expression optical flow enhancement algorithm to obtain optical flow facial features; performing feature extraction on the 3D skeleton sequence by using a space-time diagram convolutional network to obtain limb action features; performing feature extraction on the voice segments to obtain voice features; and performing fusion processing on the facial features, the body movement features and the voice features by using a multi-head space-time cross attention mechanism and an expansion time sequence convolution algorithm, and outputting an angry emotion recognition result through dynamic weight distribution classification. According to the method, the angry emotion recognition accuracy of the system is greatly improved on the whole.
Owner:SICHUAN CANCER HOSPITAL

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

Voice compression method and system based on multi-scale back projection feature fusion

The invention discloses a voice compression method and system based on multi-scale back projection feature fusion, and belongs to the technical field of voice synthesis, and the method comprises the steps: calling a plurality of multi-scale back projection feature fusion layers in an encoder to encode a to-be-synthesized voice signal, and obtaining voice features; calling a plurality of multi-scale back projection feature fusion layers in a decoder to decode the voice features to obtain synthetic voice; wherein the process of encoding or decoding the input features by the multi-scale back projection feature fusion layer comprises the following steps: carrying out cross learning on the input features by using convolution kernels of different scales to obtain multi-scale features, carrying out back projection on the multi-scale features respectively to obtain back projection features, the back projection features and the multi-scale features are fused and then fused with features input into the multi-scale back projection feature fusion layer, and output features of the multi-scale back projection feature fusion layer are obtained. The speech synthesis quality is improved, and the technical problem that the speech synthesis quality of a current method is limited is solved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Situation simulation method and system based on vehicle-mounted large language model

The invention discloses a situation simulation method and system based on a vehicle-mounted large language model, and belongs to the technical field of vehicle-mounted artificial intelligence, and the method comprises the steps: carrying out simulation situation construction and role setting based on a simulation situation type and a role selected by a user, and generating a virtual dialogue scene; voice features and emotional states of the user are analyzed, and simulation situation construction and role setting are dynamically adjusted; according to the voice features and the emotional state, selecting a proper voice synthesis role from a vehicle-mounted voice database to represent the voice of the role of the opposite side; based on the language and the expression content of the user, using a large language model to generate words and logics of the role of the opposite side; and real-time situational simulation feedback of the opposite role is provided according to the synthesized role and the generated words and logics of the opposite role in combination with real-time analysis of the voice feature and the emotional state of the user. The method helps the user to improve the expression ability and emotion management ability in a specific scene. A user is allowed to select an interested simulation situation from a preset situation library.
Owner:CHERY AUTOMOBILE CO LTD

Knowledge distillation method, system and equipment of multi-modal large model and storage medium

The invention provides a knowledge distillation method, system and equipment for a multi-modal large model and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: inputting multi-modal data sent by a user side into a visual modal analysis unit, a language modal analysis unit and a voice modal analysis unit, correspondingly generating a visual feature vector, a language feature vector and a voice feature vector; carrying out feature fusion to generate a fused feature matrix; on the basis of the fusion feature matrix, migrating knowledge of the teacher model to the student model; obtaining a prediction result output by the student model, feeding back the prediction result to the user side, and receiving feedback information sent by the user side; and the feedback information is stored in a dynamic memory bank, and incremental learning is performed on the student model based on the gradient direction vector stored in the dynamic memory bank. The method has the technical effects that dynamic optimization is carried out according to the actual use condition of the user, so that the model performance meets the actual application requirement.
Owner:北京思普艾斯科技有限公司

Knowledge distillation-based Hainan dialect speech recognition optimization system

The invention discloses a knowledge distillation-based Hainan dialect speech recognition optimization system, which comprises a data preprocessing module for processing Hainan dialect speech data and Hainan dialect text data, and outputting to obtain an MFCC speech feature sequence, a labeled text label and Hainan dialect text data; the teacher model module comprises an RNN (Recurrent Neural Network) language model, a CNN (Convolutional Neural Network) language model and a Transform language model, dialect text data are respectively input into the RNN language model, the CNN language model and the Transform language model, and a soft label and a middle layer feature are obtained through dynamic temperature regulation; the student model training module is used for performing knowledge distillation and parameter adjustment and optimization on a student model according to the MFCC voice feature sequence, the marked text tag, the soft tag and the middle layer feature; the output module is used for performing voice recognition on the Hainan dialect by using the student model after knowledge distillation and parameter optimization to obtain a recognition result; moreover, the trained student model is small in size and low in calculation complexity, and an accurate recognition result can be obtained finally.
Owner:海南经贸职业技术学院

LED display screen interaction method and system supporting voice interaction

The invention provides an LED display screen interaction method and system supporting voice interaction, and the method comprises the steps: carrying out the voice feature extraction processing through obtaining a continuous voice data stream sent by a user, and obtaining an acoustic feature sequence and a semantic association feature set of a voice segment; and calling a pre-trained voice semantic understanding model to perform joint semantic analysis processing on the acoustic feature sequence and the semantic association feature set, generating a semantic understanding result of the voice segment, and determining a user interaction intention type corresponding to the continuous voice data stream and key information positioning features of the user interaction intention in the voice content based on the semantic understanding result. And generating an LED display control instruction containing an interaction content identifier based on the user interaction intention type and the key information positioning feature, and sending the LED display control instruction to the target LED display screen to execute an interaction display operation. The real intention and key information in the continuous voice of the user can be accurately understood, the accuracy and naturalness of voice interaction between the LED display screen and the user are remarkably improved, and the interaction experience of the user is effectively improved.
Owner:SHANXI LAMPSON TECHNOLOGY CO LTD

Real-time pronunciation correction equipment for English learner

A real-time pronunciation correction device for English learners belongs to the field of voice processing and comprises an acquisition and preprocessing module, a voice feature extraction module, a standard pronunciation model construction module, a data customization management unit, a pronunciation correction suggestion module, a correction and progress tracking module and an adjustment and optimization module. By combining the deep neural network and the convolutional neural network deep learning technology, the pronunciation characteristics of the learner can be efficiently analyzed and processed. By comparing the difference between the pronunciation of the learner and the standard pronunciation in real time, pronunciation errors are accurately detected, so that high-accuracy pronunciation correction is realized. And in combination with machine learning and data analysis, the equipment can dynamically adjust a pronunciation training strategy according to individual requirements of the learner, and provide customized pronunciation guidance and suggestions. The personalized feedback helps the learner to improve the pronunciation error of the learner in a targeted manner, and the learning efficiency is improved.
Owner:NANJING COLLEGE OF CHEM TECH

Speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis

The invention relates to a speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis, and relates to the field of speech recognizing.The speech recognition and transcription method comprises the steps that a target speech signal and auxiliary modal information of synchronous visual information and text context information are obtained firstly, and the speech signal is segmented and recognized to obtain speech feature vectors; the method comprises the following steps: extracting text context information to obtain a text auxiliary feature vector, carrying out multi-modal fusion on the text auxiliary feature vector and the text auxiliary feature vector to generate fusion feature representation so as to carry out voice transcription to obtain an initial transcription text, and carrying out sentiment analysis according to visual information and the initial transcription text to generate a sentiment feature tag; and finally, optimizing and correcting the initial transliteration text based on the label to obtain a target transliteration text, thereby solving the technical problems that the speech recognition transliteration is difficult to adapt to dialect diversity and the recognition accuracy and robustness are insufficient due to neglect of emotion information, and improving the recognition accuracy and robustness through fusion of multi-modal information and emotion analysis. The voice content can be recognized more accurately, the transcription text can be optimized, and the accuracy and quality of voice recognition transcription are improved.
Owner:山西益通电网保护自动化有限责任公司