Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

154 results about "Speech rate" patented technology

Call center dialogue sentiment analysis method and system fusing voice and text

The invention provides a call center dialogue sentiment analysis method and system fusing voice and text, and relates to the technical field of voice signal processing, and the method comprises the steps: obtaining a voice signal of a call center and a corresponding transliteration text, extracting intonation, speed, sound intensity and pause sentiment features from the voice signal, meanwhile, deep language analysis is carried out on the transliterated text, and semantic emotion features related to context are extracted; and performing cross-modal correlation analysis on the voice emotion features and the semantic emotion features, and generating a time sequence correction coefficient for feature alignment by constructing a corresponding relation analysis framework between feature sequences. According to the method, language emotion information in voice and text is integrated, more accurate and comprehensive recognition of conversation emotion of the call center is realized, customer satisfaction is improved, and reliable emotion analysis basis is provided for efficiently processing customer appeals.
Owner:SHENZHEN ROADTEL DIGITAL TECH CO LTD

Voice interaction system and method for customer service based on artificial intelligence

The invention relates to the technical field of voice recognition, in particular to a voice interaction system and method for customer service based on artificial intelligence, and the system comprises a voice input processing module, an intention classification and routing module, a context dynamic adjustment module, a user behavior learning module, a multi-level intention fusion module and a final result module. According to the method, a multi-dimensional feature system is constructed by extracting tone intensity, speech speed frequency and emotional fluctuation amplitude, intention categories, priority weights and confidence scores are generated to realize accurate acquisition of appeals, and dialogue history, context and emotional change dynamic reconstruction path nodes, switching rules and response time sequences are tracked during interaction. Historical behavior mining preference features, habit fusion intention relevance, emergency calculation of an optimal strategy, construction of service steps, resource allocation schemes and execution timelines, adjustment of an interactive interface, a service process and a feedback mechanism according to multi-dimensional analysis, guarantee of differentiated service experience, and improvement of response accuracy and user satisfaction.
Owner:NANJING XIUGUO INTELLIGENT TECH CO LTD

AI agent personality switching method and device based on NFC identification

The invention relates to the technical field of artificial intelligence and Internet of Things equipment, and discloses an AI (artificial intelligence) agent personality switching method based on NFC (near field communication) identification, which realizes dynamic switching of AI agent personality in a physical mode, and comprises the following steps: S1, an equipment end obtains and decrypts an NFC label signal of an agent; s2, the equipment end obtains an AI virtual role ID corresponding to the intelligent agent according to the decryption information, and sends a request parameter to a cloud server; s3, the cloud server obtains personality configuration of the intelligent agent from a role database of the cloud server according to the request parameters and returns the personality configuration to the equipment end, wherein the personality configuration comprises an image, voice tone, speed, tone, role background, story index information and a knowledge template; and S4, the device side updates and switches the personality state corresponding to the intelligent agent according to the personality configuration, and performs voice interaction with the user according to the current personality state parameter. The invention further discloses an AI agent personality switching device for implementing the AI agent personality switching method.
Owner:SHENZHEN GIEC DIGITAL CO LTD

AI-based emotional text voice conversion method and device

The invention discloses an AI-based emotional text speech conversion method and device. The method comprises the following steps: acquiring speech segments and text records from historical data of a user; performing noise reduction processing and feature extraction according to the voice segments and the text records to obtain voice features; inputting the voice features into a pre-constructed emotional tendency model, and outputting emotional tendency and emotional intensity; according to the emotional tendency and the emotional intensity, adjusting a tone weight, a speech speed and a volume to obtain a speech parameter; extracting new voice features according to the voice parameters to perform scene emotion label matching, and determining voice adjustment parameters through a linear regression model; according to the emotional tendency and the new voice features, generating an emotional type through a pre-established emotional intention classification model, and calculating a voice parameter weight in combination with a pre-established emotional mapping table; and according to the voice adjustment parameter and the voice parameter weight, performing language synthesis to generate personalized voice. According to the method, personalized expression can be accurately generated according to the scene.
Owner:FUJIAN YUANZHI UNIVERSE CULTURE COMMUNICATION CO LTD

Sound source localization and identification system for electric power intelligent service and operation method of sound source localization and identification system

The invention discloses a sound source positioning identification system for electric power intelligent service and an operation method, and the system comprises an input module which is used for receiving an interaction demand of a user, and uploading the interaction demand of the user to an identification module; the identification module receives a user interaction demand, identifies and positions a user sound position based on the user interaction demand, and uploads an identification and positioning result to the processing module; one end of the processing module is connected to the recognition module, the other end of the processing module is connected to the output module, the processing module analyzes and processes the user question according to the received recognition and positioning result, and the output module answers the user question based on the analysis and processing result. According to the method, after the noise doped in the utterance spoken by the user is removed, the timbre, the speech speed, the audio frequency and the like in the statement of the user are recorded, and meanwhile, the user is captured, so that the virtual digital human can more accurately recognize and locate the user through a sound source, the interaction error is reduced, and the naturalness during interaction is improved.
Owner:GUANGXI POWER GRID CORP

AI call content optimization method and device based on voiceprint recognition

The invention relates to the technical field of intelligent voice processing and communication, and discloses an AI call content optimization method and device based on voiceprint recognition. The method comprises the following steps: acquiring an original voice data stream from a user call, and separating voice speed, tone and frequency characteristics to obtain a dynamic voice characteristic set; determining a microphone frequency response deviation and a speech speed change rate according to the microphone frequency response deviation and the speech speed change rate to form a speech feature parameter set; if the parameter exceeds the threshold value, redistributing a speech speed weight to generate an adjusted speech data stream; noise reduction is carried out to obtain a pure data stream, and features are fused to generate a personalized sound effect adjustment curve; adapting the equipment difference to obtain an adaptation curve; compressing the voice data according to the adaptive curve and optimizing the transmission priority to obtain an optimized transmission data stream; and combining the transmission data stream with the adaptive curve to generate final call voice output. According to the method, conversation content dynamic adaptation and whole-process optimization are realized, and conversation quality stability and cross-scene applicability are improved.
Owner:QUANZHOU YUANZHISHI ELECTRONIC COMMERCE CO LTD

Interactive robot intelligent man-machine interaction method

The invention discloses an intelligent man-machine interaction method for an interactive robot. According to the method, the interaction accuracy and adaptability are remarkably improved through multi-dimensional technology fusion. On the intention understanding level, deep analysis of user requirements is achieved through dynamic feature weight distribution, and a strategy is generated by combining cross-domain knowledge graph embedding and emotional expression, so that response content can accurately match user core appeals and can also be naturally fused into professional knowledge and emotional temperature, and information transmission deviation is effectively reduced. According to the real-time feedback mechanism, a dynamic preference model is constructed by capturing facial expressions, voice intonation and behavior data, so that the system can adjust a content generation strategy according to instant emotional fluctuation and attention change of a user, personalized experience in an interaction process is enhanced, mechanical feeling brought by a fixed response mode is avoided, and user experience is improved. Cooperation of content generation and rhythm control is realized, so that the speed, pause and emotion expression of voice response form a natural rhythm, and the comfort level of information receiving is improved.
Owner:BEIJING HAIBAICHUAN TECH CO LTD

A voice conversion method, device, equipment and readable storage medium

The application provides a speech conversion method, device and equipment and a readable storage medium. The method comprises the following steps: obtaining speech information to be processed; based on a three-head encoder, encoding and modeling speech content, environmental noise and fundamental frequency information in the speech information to be processed respectively to obtain encoded and modeled speech information; changing the time sequence of the encoded and modeled speech information to adjust the speech speed of the encoded speech information; inputting the speech information with adjusted speech speed into a previously trained timbre conversion model corresponding to a target user to obtain target acoustic features, wherein the timbre of the target acoustic features is the same as the timbre of the target user. Thus, the timbre conversion of the speech information to be processed can be performed as required, and the conversion method is more efficient and accurate. Multi-dimensional encoding of the speech information can improve the robustness of the speech in a noisy environment. Speech speed control can make the speech more in line with user requirements.
Owner:MIGU CO LTD +1

System

To provide a system capable of effectively training a presentation in an environment close to a real world.SOLUTION: The specification processing unit 290 of the data processing device 12 in the system receives the voice, the facial expression, the gesture, and the slide material input by the user, converts the received voice data into text, evaluates the speaking speed, the volume, and the pause, analyzes the facial expression, the line of sight, and the gesture of the user from the received video data, evaluates the degree of calmness and the degree of confidence of the presentation, analyzes the slide material, evaluates the consistency of the content and the visual effect, generates feedback to the user based on these evaluation results, and transmits the feedback to the terminal of the user.SELECTED DRAWING: Figure 2
Owner:SOFTBANK GROUP CORP

Vehicle-mounted voice interaction method and device, computer readable medium and electronic equipment

The application discloses a vehicle-mounted voice interaction method and device, a computer readable medium and an electronic device. The method comprises the following steps: receiving a voice input of a target user, wherein the target user is any one of all people in a current vehicle; identifying the voice input to determine a user age, a conversation speed and a conversation habit of the target user, wherein the conversation habit is used to represent a content detail degree of the target user when conversing; performing semantic recognition on text information corresponding to the voice input to determine a conversation intention of the target user; determining corresponding interaction reply content and a playing speed thereof according to the conversation intention, the user age, the conversation speed and the conversation habit; and performing voice playing on the interaction reply content according to the playing speed. The technical scheme provided by the application can adapt to the conversation habits of different users and ensure user experience.
Owner:VOYAH AUTOMOBILE TECH CO LTD

Debt conciliation verification method based on multi-modal characteristics and electronic equipment

The embodiment of the invention relates to a debt mediation verification method based on multi-modal features and electronic equipment, and belongs to the technical field of financial credit risk assessment or fraud detection.The method comprises the steps that first and second voice subjects in voice information are recognized, debtor declaration income of second voice content is extracted, and the debtor declaration income of the second voice content is obtained; and comparing the actual income of the debtor with the declared income of the debtor to obtain a flow matching deviation, analyzing an acoustic index, carrying out matching degree detection on the current voiceprint, and obtaining a credibility score of the debtor according to a comprehensive evaluation formula. According to the embodiment of the invention, the voice interaction content of the first voice main body and the second voice main body is split into multi-modal information to obtain the corresponding voiceprint and acoustic index, and the credibility score is calculated by using the comprehensive evaluation formula in combination with the multi-modal characteristics such as the voice content credibility, the flow matching deviation, the voiceprint matching degree, the environmental noise index and the voice speed fluctuation coefficient. And finally, the real repayment capability of the debtor is restored.
Owner:SUYUAN TECHNOLOGY (HUNAN) CO LTD

Multi-language adaptive identification method based on AI

The invention discloses an AI-based multi-language adaptive recognition method, and relates to the technical field of language information processing, and the method comprises the following steps: collecting rhythm features and pause nodes in a continuous voice stream, extracting a speed change track, and generating a rhythm basic draft for representing the rhythm change trend of a voice signal in a time dimension; and performing speed change decomposition on the voice signal based on the rhythm basic draft, determining a speed increasing area and a speed slowing area, generating a time adjustment table, and recording the duration and rhythm span of each speed change section. According to the invention, through double-layer time control of a rhythm basic draft and a time adjustment table, dynamic time mapping and feature extraction continuity under speech speed change are realized; and through a dynamic rhythm adjusting mechanism, executing time extension in a speech speed increasing area, executing rhythm forward movement in a speech speed reducing area, and correcting semantic dislocation and fracture in real time, so that speech recognition keeps semantic integrity and recognition stability in a multi-language and variable-speed scene.
Owner:FUJIAN SANQINGNIAO TECH CO LTD

Visual language recognition method based on spatio-temporal local harmonic neural network and application

The application discloses a visual language recognition method based on a space-time local harmonic neural network and application, and steps of the method comprise the following steps: 1, data preprocessing; 2, constructing a visual language recognition model based on a space-time local harmonic neural network; 3, training of the network model.The method can solve the problems of the existing visual language recognition methods, such as insufficient extraction ability of time characteristics and space characteristics, single method, and not paying attention to the difference of different information, so that the method can accurately recognize the word content in the scene where the speaker posture and the speech speed frequently change, and further provides a new solution for visual language recognition.
Owner:HEFEI UNIV OF TECH +3

Speech synthesis method and device

The embodiment of the invention provides a voice synthesis method and device. A server receives a voice request from a user terminal; personal information and a service scene of the target object are obtained, the personal information comprises identity information and voice information, the identity information comprises age, region and identity feature information, and the voice information comprises speed and tone; inputting the personal information and the service scene into a pre-trained voice synthesis model to obtain prompt voice data; and sending the prompt voice data to the user terminal to play the prompt voice. Visibly, according to the personal information and the service scene information of the target object, the personalized voice meeting the user requirements can be generated according to the characteristics and the specific service scenes of different users, the naturalness of the synthesized voice is improved, machinery is reduced, and the experience of the user in the interaction process with the voice synthesis system is improved.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD

Digital population voice synchronization generation method and system based on time sequence decoupling

The embodiment of the application provides a kind of based on timing decoupling's digital population voice synchronous generation method and system, belong to digital human technical field;The method includes face detection and cutting to original video sequence, obtain standard face image sequence, the feature extraction of target audio signal is carried out, and deep audio feature sequence is obtained;Mask processing is carried out to each frame, and the image of mouth area to be driven is obtained;It is constructed as multi-channel input tensor;The mouth image that pre-trained mouth generation network outputs is synchronized with target audio signal;After replacing the mouth image to the corresponding position of original video sequence, it is combined with target audio signal, and the digital human video of mouth voice synchronization is output.The deep audio feature extraction and audio-video accurate alignment of the application improve the synchronization accuracy of mouth and voice, the timing dependence is modeled when multi-channel input, the transition of mouth sequence is smooth, and the adaptive feature fusion mechanism can automatically adapt to different phonemes and speech rate.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Method for simultaneous call interpretation using headset

The present application relates to the field of call translation. Disclosed in the present application is a method for simultaneous call interpretation using a headset, which solves the problem of severe information distortion and loss of connotation often being caused during simultaneous call interpretation if the meaning of an original speech is merely rigidly conveyed in a target language, without fully capturing and expressing the connotation and context of an original text. By simulating the tone of a speaker, the present invention can significantly improve the quality of simultaneous call interpretation, enhance the expression of emotions, increase the accuracy of information, strengthen communication effects, and improve the user experience; in terms of technical implementation, speech synthesis and emotion analysis techniques can provide strong support, such that tone simulation is more natural and real; and speech data of different accents, different speakers and different background noises and speaking speeds is collected, such that a more robust and accurate speech recognition model can be trained, thereby adapting to diverse speech inputs, and improving the generalization capability of the model in different scenarios.
Owner:VISION INTELLIGENCE CO LTD

A psychological counseling real-time speech recognition method based on multi-modal data

The present application relates to the technical field of speech recognition, and discloses a psychological counseling real-time speech recognition method based on multi-modal data, comprising: constructing a gender-specific speech psychological feature mapping baseline, analyzing and judging the emotional stability tendency degree, and combining the amplitude peak value proportion to judge the expression tendency degree; dynamically correcting the fundamental frequency mean amplitude to solve the psychological feature misjudgment caused by individual pronunciation difference; detecting the stress position, key words, speech speed and pause features through the fundamental frequency mutation and amplitude mutation, constructing a multi-dimensional emotion analysis model, and optimizing the emotion focus positioning, emotion type judgment and change trend tracking problems; designing a comprehensive stability value calculation method, simultaneously reflecting the emotional stability and the influence degree of key words, and providing a psychological health evaluation quantitative index; constructing a three-level processing mechanism, and through the fusion of the exclusive baseline and multi-modal evidence, performing feature adaptation, cross-validation and decision correction.
Owner:MEDICAL HEALTHCARE DIGITAL TECH (SHENZHEN) CO LTD

A high-precision speech recognition and semantic understanding method

PendingCN122369448ASpeech rateSound sources
The application belongs to the technical field of speech recognition and semantic understanding, and discloses a high-precision speech recognition and semantic understanding method. The method comprises the following steps: collecting a far-field multi-sound-source original speech signal to generate a discrete speech sampling sequence; performing frequency domain transformation on the original speech signal to extract an amplitude-frequency distortion quantization parameter; segmenting the original speech signal to generate a speech frame sequence and calculating a non-steady-state speech speed quantization parameter; extracting an original mel-frequency cepstrum feature and performing distortion correction to obtain a weighted acoustic feature through dynamic weighting; constructing a preset semantic feature library and mapping to generate an initial semantic matching vector; and recursively iterating to calibrate the semantic confidence and screening an optimal term to output a result. The application solves the problem of low far-field speech recognition accuracy in the prior art, realizes high-precision speech recognition and semantic understanding through multi-link collaborative optimization, and improves recognition stability.
Owner:SHANGHAI MAIJUN TECHNOLOGY CO LTD

A motorcycle off-line voice interaction system and method thereof

This invention discloses an offline voice interaction system and method for motorcycles, belonging to the field of voice interaction technology. The invention aims to solve the technical problems of low recognition rate and lack of driving safety control in existing motorcycle voice systems under high-speed wind noise environments. Its technical solution includes: synchronously acquiring microphone array audio signals and CAN bus vehicle driving status data; dynamically loading a high-noise robust simplified instruction library or a full-function instruction library based on the real-time environmental signal-to-noise ratio to adapt to different riding scenarios; using a safety interlock decision module to match voice operation intentions with physical states such as vehicle speed and tilt angle in a permission mapping table, intercepting high-risk instructions or triggering secondary confirmation; and adaptively adjusting the gain and speech rate of the feedback voice based on the environmental signal-to-noise ratio. This invention can achieve highly reliable voice interaction across the entire speed range while ensuring riding safety.
Owner:ZHEJIANG QIANJIANG MOTORCYCLE

Estimating keyword length refinement based on speech rate classification

Systems and techniques for processing one or more audio samples are provided. For example, a process may include: detecting spoken keywords within audio samples in the one or more audio samples using a first keyword detection model; determining an estimated keyword index corresponding to the detection of spoken keywords within the audio samples, the estimated keyword index including an estimated keyword start index and an estimated keyword end index; using a speech rate classification machine learning network to determine speech rate information corresponding to the audio samples; obtaining an average spoken length value corresponding to the spoken keywords and speech rate information; and generating a refined keyword index based on the estimated keyword index and the average spoken length value, wherein the refined keyword index includes a refined keyword start index offset to a time earlier than the estimated keyword start index.
Owner:QUALCOMM INC

Systems and methods for intelligent playback

Systems and methods for intelligent playback of media content may include an intelligent media playback system that, in response to determining the speech tempo in audio content by measuring syllable density of speech in the audio content, automatically adjusts a playback speed of the audio content as the audio content is being played based on the determined speech tempo. In some embodiments, the system may automatically and dynamically adjust the playback speed to result in a desired target speech tempo. In addition, the system may determine whether to automatically adjust playback speed of the audio content, as the media is being played, based on the detected speech tempo of the speech in the audio content and the determined type of content of media. Such automatic adjustments in playback speed result in more efficient playback of the audio content.
Owner:DISH NETWORK TECHNOLOGIES INDIA PTE LTD +1

Emotion perception and empathetic interaction adjustment method and system for embodied agent

The application discloses a method and system for emotion perception and empathetic interaction adjustment of embodied agents. Non-contact acquisition of user physiological signals: millimeter wave radar measures respiratory rate and heart rate, a camera measures heart rate and / or heart rate variability through remote photoplethysmography rPPG, and wearable device signals can be optionally fused; real-time evaluation of signal quality of each mode and adaptive fusion, estimation of emotional / stress state containing physiological arousal based on individualized baseline; adjustment of robot behavior strategy based on the estimation, output of embodied behavior parameters such as speech speed, action amplitude, close distance and dialogue content; the behavior strategy adopts reinforcement learning, and physiological arousal is included in the reward function, so that the robot learns embodied behaviors for reducing user discomfort and improving comfort, and an empathetic interaction closed loop of perception-estimation-decision-execution-re-perception is formed. The application does not need long-term wearing, can more truly reflect the internal state, is robust, safe and privacy-friendly, and is suitable for companion and education robots.
Owner:ANHUI HEARTVOICE MEDICAL TECH CO LTD

Digital NPC role emotion expression and interaction system

The invention discloses a digital NPC role emotion expression and interaction system, relates to the technical field of digital interaction, and aims to improve the accuracy and dimension coverage of user emotion state recognition, comprehensively capture dominant emotion expression and implicit emotion tendency of a user, form more accurate and fine emotion state features and improve the user emotion state recognition efficiency. According to the method and the system, the NPC role response adaptation degree and the user interaction experience satisfaction degree can be improved, the user experience can be improved, the user experience can be improved, and the user experience can be further improved by searching and matching each allowable response emotion type and comprehensively evaluating to obtain a target response emotion type. And meanwhile, the voice output and the body movement of the NPC role present a natural and smooth rhythm sensation, the emotion change of the user can be continuously monitored in the interaction process, and the speech speed, the volume and the movement amplitude parameters of the next interaction segment are dynamically adjusted by calculating the emotion change trend deviation value, so that the real-time adaptive optimization of the NPC role interaction expression is realized.
Owner:YUANZHIUNIVERSE (FUJIAN) TECH GRP CO LTD +1

System

A system is provided.SOLUTION: A system comprising: means for recording content of a meeting; means for extracting audio from the recorded data; means for converting the extracted audio to text; means for analyzing the text to evaluate linguistics, syntax, and structure of discussion; means for analyzing audio data to evaluate voice tone and speaking speed; means for generating feedback based on the analysis; and means for presenting the generated feedback to a user.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Multi-emotion speech synthesis method based on pre-training embedding and label control

The invention provides a multi-emotion speech synthesis method based on pre-training embedding and label control, and relates to the technical field of speech synthesis. According to the method, effective alignment of text and acoustic features is realized through joint modeling of emotion embedding and text features, high-quality spectrum features are generated by using a bidirectional normalized stream and a de-noising latent space diffusion model, and finally voice waveforms are synthesized through HIFIGAN decoding. The method has the beneficial effects that (1) the dependence on a predefined emotion label is eliminated, and the emotion information can be directly captured and utilized from the original voice; (2) realizing multi-dimensional flexible control of synthetic speech, including adjusting an emotion feature vector to control emotion intensity; changing discrete emotion labels to convert emotion categories; adjusting a duration predictor to realize speech speed control; and (3) the synthesized speech emotion expression is natural and rich, and effective cloning of unknown new emotions is supported.
Owner:CHONGQING UNIV OF TECH

Dynamically adapting media playback rate

A method provides techniques for dynamically adapting a media playback rate. A speaker analysis (SA) module operating on an electronic device is configured to cause the electronic device to obtain one or more playback preferences for a user. Audio data of recorded speech including spoken words is obtained and analyzed. The analysis can include determining a speech rate, accent, subject matter, genre, and / or other parameters pertaining to the spoken words. Based on the analysis, and user preferences, a recommended playback rate is computed, based at least in part on the playback rate preferences of a user and the one or more speech parameters. The playback rate can be automatically set to the recommended playback rate, and the media asset is rendered at the recommended playback rate.
Owner:MOTOROLA MOBILITY LLC

Digital human navigation type marketing short video generation method and system

The invention relates to the technical field of video generation, discloses a digital human guide type marketing short video generation method and system, and aims to solve the problems of fixed speech speed, lack of dynamic adjustment and low matching degree of content and users in the prior art. The method comprises the following core steps: evaluating user interaction strength based on a user sliding speed and a reviewing rate; evaluating the content complexity based on the keyword density and the numerical information density; integrating the user interaction strength, the content complexity, the platform recommendation score and the real-time conversion rate to calculate the adaptation degree; based on the split-screen state, the continuous watching time and the environment noise level, evaluating the attention distraction degree of the user; and finally, dynamically regulating and controlling the explanation speed of the digital human to a target value according to the adaptation degree, the attention dispersion coefficient and the historical average playing speed multiplying power of the user. The corresponding system module is used for executing the process. According to the invention, real-time personalized adaptation of the explanation speed is realized, and the information transmission efficiency and the user watching experience are improved.
Owner:TAIDOU TECH GRP CO LTD

Selective visual display

According to one aspect of the invention, a multi-modal reading system is provided, comprising: a display screen; an eye movement tracking device; a processor; one or more computer storage devices; and a speaker, wherein the processor is configured to: measure, by the eye movement tracking device, eye movement of the user to determine a particular word gazed by the user; modifying the display of the text content, and visually highlighting the word or phrase of the current fixation point of the user; calculating the number of words read by the user per minute based on continuous measurements of the specific words watched by the user; according to the calculated reading rate, delay is set between continuous text element presentation so as to adapt to user cognition processing time; and playing the text content to the user at a speech speed matched with the reading speed of the user through a loudspeaker.
Owner:德查姆斯·理查德·克里斯托弗

Psychological counseling real-time speech recognition method based on multi-modal data

The invention relates to the technical field of speech recognition, and discloses a psychological counseling real-time speech recognition method based on multi-modal data, which comprises the following steps: analyzing and judging an emotion stability tendency degree by constructing a gender-exclusive speech psychological feature mapping baseline, and judging an expression tendency degree in combination with an amplitude peak value proportion; the fundamental frequency mean amplitude is dynamically corrected, and psychological feature misjudgment caused by individual pronunciation difference is solved; detecting accent positions, keywords, speech speed and pause features through fundamental frequency abrupt change and amplitude abrupt change, constructing a multi-dimensional emotion analysis model, and optimizing emotion focus positioning, emotion type judgment and change trend tracking problems; designing a comprehensive stable value calculation method, reflecting the emotional stability and the keyword influence degree at the same time, and providing quantitative indexes for mental health assessment; and constructing a three-level processing mechanism, and performing feature adaptation, cross validation and decision correction through exclusive baseline and multi-modal evidence fusion.
Owner:MEDICAL HEALTHCARE DIGITAL TECH (SHENZHEN) CO LTD