Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

275 results about "Speech rate" patented technology

Robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation

The invention discloses a robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation. The method comprises the following steps: S1, dynamically fusing multi-modal emotions; the method comprises the following steps: S1, synchronously acquiring voice, visual and text signals through a multi-source heterogeneous sensor, capturing a user voice stream by a high-fidelity microphone array, and extracting acoustic characteristics such as intonation and speed, S2, performing cross-modal reasoning; s3, synchronously generating contents; step S4: style migration; step S5, anthropomorphic voice and expression generation; according to the method, man-machine interaction emotion is analyzed and generated by utilizing a large language model and multi-modal information fusion, the singleness of interaction emotion and the deficiency of emotional sharing ability are avoided, a strong emotion interaction characteristic is achieved, the image of the robot is obtained through a generative technology and can be migrated to any image, the limitation that a specific image is independently made is broken through, and the interaction effect of the robot is improved. The advantage that one robot can be suitable for different scenes is achieved.
Owner:JIANGSU YUNMU ZHIZAO TECH CO LTD

AI-driven personalized voice training and pronunciation correction system

The invention provides an AI-driven personalized voice training and pronunciation correction system. The system firstly collects original voice data of a user in a voice training process, and extracts multi-dimensional voice feature vectors including pitch, speed, intonation, formant parameters and corresponding texts; based on the feature vector, recognizing a context tag where the current voice is located, and constructing a pronunciation feature vector of the user; the system further calls a standard pronunciation database according to the context tag, generates a target pronunciation feature vector under the corresponding context, and performs multi-dimensional comparison on the user pronunciation and the target pronunciation to obtain a difference parameter set containing pronunciation parts, speed and emotional differences; finally, the system generates correction suggestions including pronunciation action guidance, intonation adjustment prompts and semantic emotion enhancement instructions, and receives user feedback information to improve the personalized training effect; according to the invention, accurate pronunciation correction under context perception can be realized, and the personalized and intelligent level of voice training is enhanced.
Owner:GUANGZHOU SENJI SOFTWARE TECH CO LTD

Transform speech recognition method, system and device based on multi-modal audio-visual fusion and medium

The invention discloses a transformer speech recognition method, system and device based on multi-modal audio-visual fusion and a medium, and the method comprises the steps: obtaining audio and video data to be processed, the audio and video data comprising paired audio data and video data; extracting audio data features to obtain audio features; extracting video data features to obtain video features; inputting the extracted audio features and the extracted video features into a Transform model, and outputting predicted text information; and the Transform model comprises an encoder, a decoder and a hybrid CTC / attention (China Table Control Code / Attention). According to the method, after an original signal is converted into a feature vector which can be processed by a Transform model, information of audio and video modalities is integrated, and dynamic weight distribution is applied to balance information contribution among different modalities; according to the method, the conversion from voice to text is realized by utilizing an encoder and decoder structure, and meanwhile, the dependency relationship among positions in an input sequence is captured by virtue of a multi-head self-attention mechanism, so that the problem that the expression of voice recognition in a complex environment is limited by noise, accent and speech speed is solved.
Owner:XIJING UNIV

Voice emotion recognition method, device and equipment and readable storage medium

The invention relates to the technical field of voice signal processing, and discloses a voice emotion recognition method, device and equipment and a readable storage medium, and the voice emotion recognition method comprises the steps: obtaining a target voice signal, and executing sampling processing to obtain voice sampling data; extracting a plurality of acoustic features based on the voice sampling data, and constructing initial feature representation; in combination with environment signal-to-noise ratio information, performing channel weighted fusion processing to generate fusion feature representation; inputting the fusion feature representation into a time sequence modeling network, and extracting context information to obtain time sequence abstract features; and executing emotion recognition processing based on the time sequence abstract features to generate a corresponding emotion recognition result. According to the method, the emotion recognition accuracy in a multi-noise environment is improved, the expression ability of fine-grained emotion features such as speech speed changes and rhythm fluctuations is enhanced, the dynamic adaptability to environment changes in the feature fusion process is achieved, and higher recognition stability and environment adaptability are achieved.
Owner:ULTIMATE IOT (HENAN) TECHNOLOGY LTD +1

Speech cloning system and method fusing rhythm characteristics

The invention discloses a voice cloning system and method fusing rhythm characteristics, belongs to the technical field of voice synthesis and natural language, and is applied to the aspect of fine-grained rhythm control in zero-sample voice synthesis. The implementation method comprises the following steps of: 1, extracting rhythm features and audio features of an audio file, and further respectively acquiring pause features, speed features and tone features in the rhythm features by sequentially adopting transcriptional text inverse coding, syllable-level speed registration quantization and pitch sequence feature splicing modes; 2, fusing the features of the audio files in a manner of discarding feature screening without guidance of a classifier; 3, generating a target Mel spectrogram based on conditional flow matching; generating a target audio file from a to-be-cloned audio file through the trained voice cloning model controlled by the fusion rhythm; compared with the prior art, fine-grained rhythm control of tone and rhythm feature decoupling is realized in zero-sample speech synthesis, so that intonation accuracy based on a context scene is improved.
Owner:BEIJING INST OF TECH

Spoken language pronunciation training correction system based on intelligent equipment

The invention belongs to the technical field of intelligent voice processing, and particularly relates to a spoken language pronunciation training and correcting system based on intelligent equipment, which acquires a user rhythm feature set including pitch change rate, accent intensity, pause duration and intonation contour by acquiring spoken language audio data sent by a user for a target text, and corrects the spoken language pronunciation training and correcting system. The method comprises the following steps: acquiring spoken language audio data, converting the spoken language audio data into a phoneme sequence aligned with target text time, generating a rhythm deviation degree report according to comparison evaluation of a user rhythm feature set and the phoneme sequence, an execution intonation mode, accent distribution, speech stream sound change and speech speed rhythm, and determining the rhythm deviation degree according to a rhythm problem type in the report in combination with user historical learning data. And forming and outputting a correction scheme including a text prompt, a targeted minimum contrast training unit and a listen-and-read simulation task, solving the problem of weak capability of correcting hyper-phoneme in a second language spoken language of an adult, and improving the authentic and fluency of the spoken language, thereby realizing efficient personalized learning.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Call center dialogue sentiment analysis method and system fusing voice and text

The invention provides a call center dialogue sentiment analysis method and system fusing voice and text, and relates to the technical field of voice signal processing, and the method comprises the steps: obtaining a voice signal of a call center and a corresponding transliteration text, extracting intonation, speed, sound intensity and pause sentiment features from the voice signal, meanwhile, deep language analysis is carried out on the transliterated text, and semantic emotion features related to context are extracted; and performing cross-modal correlation analysis on the voice emotion features and the semantic emotion features, and generating a time sequence correction coefficient for feature alignment by constructing a corresponding relation analysis framework between feature sequences. According to the method, language emotion information in voice and text is integrated, more accurate and comprehensive recognition of conversation emotion of the call center is realized, customer satisfaction is improved, and reliable emotion analysis basis is provided for efficiently processing customer appeals.
Owner:SHENZHEN ROADTEL DIGITAL TECH CO LTD

Speech translation system based on Bluetooth earphone

The invention discloses a voice translation system based on a Bluetooth headset, which relates to the technical field of voice translation of Bluetooth headsets and comprises a voice acquisition module, a voice speed detection module, a rhythm evaluation module, a voice regulation and control module and a text verification module. The rhythm evaluation module is used for extracting rhythm characteristic information in a user voice signal collected by the Bluetooth earphone in the voice speed fluctuation time period, analyzing the rhythm characteristic information and evaluating the intensity of statement rhythm mutation in user expression in the voice speed fluctuation time period; and the voice regulation and control module is used for selecting and executing a corresponding voice processing strategy according to the evaluation result, and carrying out dynamic regulation and control in the execution process. According to the method, the problem that statement rhythm mutation caused by speech speed fluctuation cannot be dynamically regulated and controlled is solved, intelligent adaptation and whole-process response of a speech processing strategy are realized, structural optimization is performed on a translation result, and translation accuracy and semantic coherence are remarkably improved.
Owner:VISION INTELLIGENCE CO LTD

Voice interaction system and method for customer service based on artificial intelligence

The invention relates to the technical field of voice recognition, in particular to a voice interaction system and method for customer service based on artificial intelligence, and the system comprises a voice input processing module, an intention classification and routing module, a context dynamic adjustment module, a user behavior learning module, a multi-level intention fusion module and a final result module. According to the method, a multi-dimensional feature system is constructed by extracting tone intensity, speech speed frequency and emotional fluctuation amplitude, intention categories, priority weights and confidence scores are generated to realize accurate acquisition of appeals, and dialogue history, context and emotional change dynamic reconstruction path nodes, switching rules and response time sequences are tracked during interaction. Historical behavior mining preference features, habit fusion intention relevance, emergency calculation of an optimal strategy, construction of service steps, resource allocation schemes and execution timelines, adjustment of an interactive interface, a service process and a feedback mechanism according to multi-dimensional analysis, guarantee of differentiated service experience, and improvement of response accuracy and user satisfaction.
Owner:NANJING XIUGUO INTELLIGENT TECH CO LTD

AI-driven video content subtitle synchronous translation method and system

The invention discloses an AI-driven video content subtitle synchronous translation method and system, and relates to the technical field of video subtitle synchronization, the system is combined with a video frame acquisition module and a face and lip action recognition module, and the system can accurately obtain the lip opening and closing vertical distance and opening and closing times of each role. The data are used for calculating the actual speaking speed, and the actual speaking speed is compared with a traditional speed index to obtain a first calibration difference coefficient. According to the method, the timestamps of the subtitles are effectively adjusted, the time deviation caused by the speech speed difference is reduced, the subtitles and the actual speech are more synchronous, and therefore the accuracy of the subtitles and the film watching experience of audiences are improved. The multi-person talking overlapping recognition module can accurately detect and mark the voice overlapping condition. And if the overlapped voice influence factor D exceeds the abnormal threshold F, the system triggers the second correction instruction to further calibrate the subtitle timestamp, so that the synchronization problem caused by voice overlapping is avoided.
Owner:CHONGQING MALYA MEDIA CO LTD

Enhanced wireless communication handover management system

To effectively resolve the issue of inappropriate handovers during communication sessions, methods, systems, and machine-readable mediums which utilize speech features captured by microphones in the original wireless peripheral device and / or the wireless peripheral device to which the communication session is to be handed over to determine if a handover should proceed or should be reversed. Speech features refer to the composite attributes of spoken language that encompass both acoustic and linguistic features. Acoustic features characterize the sound properties of speech and include, but are not limited to, timbre, pitch, intonation, speaking rate, articulation, prosody, melody, spectral features, formant frequencies, and the like. Linguistic features pertain to the actual content conveyed, comprising words, phrases, syntax, and semantics. This includes the analysis of words, phrases, syntax, and semantics to understand the context and continuity of the conversation.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Bluetooth earphone AI voice control method and system

The invention relates to the technical field of voice recognition, in particular to a Bluetooth headset AI voice control method and system, and the method comprises the following steps: collecting a voice sample, carrying out the feature extraction of a voice signal through employing a Mel-frequency cepstrum coefficient, and generating voice feature data; and inputting the voice feature data into an acoustic model, and improving the recognition rate of the key instruction words by the acoustic model through learning features to obtain an optimized acoustic model. According to the method, acoustic feature capture is realized through voice signal feature extraction, the recognition accuracy of the key instruction is improved in cooperation with deep learning of acoustic features, the influence of different user speech speeds on the recognition effect is overcome by applying the time alignment technology of voice input, and the stability of the control instruction in different use scenes and speech speed changes is ensured. In addition, the recognition capability of a specific command is further mined and optimized by means of statistical characteristic analysis of voice, and efficient and automatic recognition and extraction of key control commands in continuous speech streams are completed.
Owner:JIANGXI CHANGRONG TECHNOLOGY CO LTD

AI agent personality switching method and device based on NFC identification

The invention relates to the technical field of artificial intelligence and Internet of Things equipment, and discloses an AI (artificial intelligence) agent personality switching method based on NFC (near field communication) identification, which realizes dynamic switching of AI agent personality in a physical mode, and comprises the following steps: S1, an equipment end obtains and decrypts an NFC label signal of an agent; s2, the equipment end obtains an AI virtual role ID corresponding to the intelligent agent according to the decryption information, and sends a request parameter to a cloud server; s3, the cloud server obtains personality configuration of the intelligent agent from a role database of the cloud server according to the request parameters and returns the personality configuration to the equipment end, wherein the personality configuration comprises an image, voice tone, speed, tone, role background, story index information and a knowledge template; and S4, the device side updates and switches the personality state corresponding to the intelligent agent according to the personality configuration, and performs voice interaction with the user according to the current personality state parameter. The invention further discloses an AI agent personality switching device for implementing the AI agent personality switching method.
Owner:SHENZHEN GIEC DIGITAL CO LTD

Teaching quality evaluation method based on teaching robot

The invention discloses a teaching quality evaluation method based on a teaching robot, and relates to the field of artificial intelligence education application, and the method comprises the steps: collecting classroom teaching video data and audio data through a sensor of the teaching robot; performing face recognition according to the classroom teaching video data to obtain student expression data, and performing acoustic feature extraction on the audio data to obtain teacher voice data; sending the student expression data into a pre-trained emotion recognition model to obtain a classroom student attention score, and sending the teacher voice data into a voice evaluation model to obtain a teacher speed score; and establishing a weighted calculation formula based on the classroom student attention score and the teacher speech speed score, and calculating to obtain a classroom teaching quality score. According to the invention, the classroom video and audio data are collected and processed in real time, the student attention score and the teacher speed score are calculated, and weighted fusion is carried out, so that the objective quantitative evaluation of the classroom teaching quality is realized.
Owner:SHANDONG HONGRU SMART EDUCATION TECHNOLOGY CO LTD

Anti-fraud outbound call identification method for cross-number-segment voiceprint tracking

The invention relates to an anti-fraud outbound recognition method for cross-number-segment voiceprint tracking, and the method comprises the steps: introducing a time domain voiceprint adversarial extraction mechanism, and adding an adversarial voice discriminator, an active inhibition device, a channel, and the interference of non-identity factors of language emotion in voiceprint modeling; constructing a speaker constant feature residual channel: taking stable syllable fragments including vowels and gutto vowels in call content as a reference extraction window, and filtering emotion or speech speed driving components; outputting a steady-state voiceprint contour flow as a unique basis for subsequent cross-number identity fusion; designing a dynamic number fusion graph based on the voiceprint steady-state flow; each time of voiceprint appearance is regarded as a single time point node, and all similar historical voiceprints form a serial number merging path; a number time sequence edge weight function is set; the occurrence time interval, the use frequency and the regional jump of the new number and the old number are synthesized, and whether the number belongs to the same user or not is evaluated; and establishing atlas connection between the voiceprint main body node and all number nodes thereof.
Owner:SHIJIAZHUANG LINGYUE TECHNOLOGY CO LTD

Cabin voice interaction method and device and storage medium

The invention discloses a cockpit voice interaction method and device and a storage medium, and belongs to the technical field of vehicle control. The method comprises the steps of collecting sound information of a driver; acquiring a detection result and a control instruction of the elderly driver based on the sound information of the driver, wherein the detection result of the elderly driver is used for indicating whether the driver is an elderly driver; in response to the obtained detection result that the driver is the elderly driver, the vehicle is controlled to start a working mode suitable for the elderly; controlling the vehicle to execute the control instruction; collecting comprehensive noise values inside and outside the vehicle; calculating a broadcast volume and a broadcast speed based on the comprehensive noise value; and broadcasting the execution condition of the control instruction according to the broadcasting volume and the broadcasting speed. The accuracy of instruction recognition and voice broadcast functions of the cabin is improved, voice instructions of old people can be recognized more easily, voice broadcast content can be heard more easily, and therefore driving experience and driving safety are improved.
Owner:CHERY AUTOMOBILE CO LTD

Cross-platform calling method and system for corpus training library

The invention relates to the technical field of speech synthesis, in particular to a corpus training library cross-platform calling method and system, and the method comprises the steps: obtaining the text content of a to-be-synthesized audio, disassembling the text content to obtain a pause position sequence, selecting a corresponding corpus training library according to a preset language identifier, extracting the pronunciation duration of character phonemes, and calculating the duration of basic phonemes. Text emotion features are recognized through an emotion classification model, a three-dimensional emotion feature vector containing emotion density, emotion variance and distribution entropy is constructed, and a first phoneme adjustment coefficient is determined according to the three-dimensional emotion feature vector to correct basic phoneme duration. And when a difference value exists between the corrected phoneme duration and the target audio duration, carrying out compression processing on the phonemes when the difference value is a positive number, and generating an optimal second phoneme adjustment coefficient sequence and a pause duration sequence by using a genetic algorithm when the difference value is a negative number. According to the invention, the problem of speech speed and pause duration optimization under the preset audio length is solved.
Owner:SICHUAN NORMAL UNIV

Information processing method and equipment for man-machine practice, and medium

The invention relates to an information processing method and device for man-machine practice and a medium, and belongs to the technical field of man-machine interaction, and the method comprises the following steps: obtaining interaction voice of a user for question reply in real time, and converting the interaction voice into text information with a timestamp; performing preprocessing and key semantic information extraction on the text information, and performing emotion recognition, mute recognition and speech speed recognition; evaluating the practice process in real time according to the quality inspection rule, and judging whether interaction needs to be interrupted or not according to an evaluation result; performing intention recognition based on the extracted key semantic information, mapping the recognized intention to a corresponding node of a knowledge graph, iteratively updating an associated question-chasing path according to the knowledge graph, generating a response question in combination with a scene label and a process rule, and sending the response question to a user; and obtaining the interactive voice of the user for the response question, and carrying out the next round of man-machine practice interaction. Compared with the prior art, the method has the advantages of being capable of achieving personalized practice and the like.
Owner:中国太平洋人寿保险股份有限公司

Tourism guidance system based on big data AI

The invention discloses a tourism guidance system based on big data AI, and belongs to the technical field of tourism guidance, the tourism guidance system comprises an interaction engine design module and a multi-source module information fusion positioning module, the interaction engine design module comprises: a big language model hybrid architecture; a context sensing and semantic tracking processing mechanism; according to the method, interaction naturalness is improved, context association and language naturalness of dialogues are remarkably improved through a multi-round semantic understanding mechanism based on a large language model, the large language model is used for supporting multi-round semantic interaction, slot position guiding and context maintaining, a system can understand continuous intentions of users, and the interaction naturalness is improved through a dialogue state machine and a semantic memory bank. The system can recognize the emotion of the user through tone, speed and context, adjust the explanation tone and content according to the emotional state, have the emotional common feeling ability and improve the user satisfaction degree.
Owner:HEFEI TINGTING ARTIFICIAL INTELLIGENCE APPLICATION TECHNOLOGY SERVICE CO LTD

Digital human contradictory dispute mediation method based on intelligent perception and emotion regulation

The invention discloses a digital human contradictory dispute mediation method based on intelligent perception and emotion regulation. The method comprises the following steps: firstly, acquiring voice information and visual information of two contradictory parties; calculating a voice emotion score by using the voice information, calculating a visual emotion score by using the visual information, and carrying out weighted calculation to obtain a comprehensive emotion score; then generating an initial speech speed, an initial expression and an initial gesture of the digital human; matching a proper response strategy from the knowledge base and transmitting the response strategy to the digital person; and finally, mapping voice information in the voice emotion score to a voice synthesis parameter, generating emotional voice signals consistent with both parties of a contradictory dispute in language by combining a response strategy, and outputting the emotional voice signals by a digital person for contradictory dispute mediation. In the mediation process, the digital person monitors the comprehensive emotion score and the contradictory scene in real time for dynamic adjustment. The real emotional state of the user is accurately captured based on intelligent perception, emotion mediation is carried out based on emotion perception, and the reality sense of digital human interaction and the user trust degree are improved.
Owner:GUANGDONG UNIV OF EDUCATION

English emotion intonation reading device

The invention discloses an English emotion intonation reading device, which relates to the technical field of English emotion intonation reading and comprises a text input module, a semantic analysis module, an emotion recognition module, an intonation adjusting module, a speech synthesis module, an audio output module and a self-learning module. The text input module is used for acquiring an English text by typing, uploading a document and adopting an OCR (Optical Character Recognition) image recognition mode to obtain original text data; the semantic analysis module is used for analyzing original text data by adopting a method for analyzing a syntactic structure, keywords and context information to obtain semantic information of a text; the emotion recognition module is used for processing semantic information of the text in combination with a rule base and a deep neural network, recognizing emotion categories in the text and obtaining emotion tags; and the intonation adjusting module is used for adjusting voice expression parameters in the emotion label by adopting a strategy of controlling the speed, pitch, rhythm, pause and accent to obtain a voice regulation and control scheme.
Owner:SICHUAN HEALTH REHABILITATION VOCATIONAL COLLEGE

Voice interactive chart dynamic generation method and system based on large language model

InactiveCN120496560ASpeech recognitionFrequency spectrumNoise power spectrum
The invention discloses a voice interactive chart dynamic generation method and system based on a large language model, and relates to the technical field of chart generation, and the method comprises the steps: collecting a user voice signal, carrying out the preliminary framing, calculating the short-time energy, and analyzing a mute segment set signal; frequency domain signals are extracted from the set signals, gain is calculated through a filter, noise reduction is carried out on the frequency domain signals, the zero-crossing rate and noise reduction short-time energy are calculated, effective voice frames are screened, peak detection is carried out, the average voice speed is calculated, voice speed normalization is carried out, and signal values of sampling points in the frames are extracted. According to the method, the noise power spectrum density is extracted through FFT on the mute section, a frequency domain model of background noise is effectively established, a follow-up filter can accurately act on an actually existing frequency band interference area, weakening of the voice main signal spectrum is avoided, the Mel filter bank is accessed after the power spectrum is calculated through the FFT after frame windowing, and the noise power spectrum density is calculated through the FFT after frame windowing. Energy can be redistributed on the logarithmic Mel scale according to human ear perception characteristics.
Owner:DONGQU INTELLIGENT TRANSPORTATION INFRASTRUCTURE TECH (JIANGSU) CO LTD

Voice interaction method, server and computer readable storage medium

The invention discloses a voice interaction method, a server and a computer readable storage medium. The method comprises: determining an emotional state identifier according to acoustic features of a received voice request, the acoustic features including at least one of a tone feature, a speech speed feature and an energy feature; and according to the emotional state identifier and the voice request, generating an emotional adaptive pad call so as to complete the voice interaction. Therefore, by analyzing the acoustic features, identifying the emotional state identifier, accurately generating the emotional adaptive pad call corresponding to the voice request, and establishing emotional connection with the user, the user experience is enhanced. Moreover, by generating the emotion-adaptive pad call, a waiting blank period from the time after the user sends an instruction to the time before the system returns a formal result can be filled up, the experience of'instant response 'is created, the user is prevented from anxiety caused by too long waiting time, and the satisfaction degree and the credibility of the system are further improved.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

AI tourism personalized explanation system and service system

The invention relates to the technical field of intelligent tourism services, and particularly discloses an AI tourism personalized explanation system and service system.The personalized explanation system comprises a positioning module used for obtaining user position information; the audio acquisition module is used for acquiring environmental noise data; the attribute acquisition module is used for acquiring user attribute data; the intelligent explanation module is used for determining explanation content based on the user position information; determining a reference speed and a reference volume based on the user attribute data; respectively compensating the reference speed and the reference volume based on the environmental noise data to generate an explanation speed and an explanation volume; and explaining the explaining content based on the explaining speed and the explaining volume. According to the method, the real-time position information of the user, the environment noise data and the attribute characteristics of the user are integrated, the interpretation content is dynamically matched, the speech speed and the volume are adaptively adjusted, and the interpretation effect and the user experience are improved.
Owner:SICHUAN XINGANYAN CULTURE TECHNOLOGY CO LTD

AI speech emotion analysis method and platform of deep learning architecture

The invention relates to the technical field of speech sentiment analysis, and discloses an AI speech sentiment analysis method and platform of a deep learning architecture, and the platform comprises a signal collection unit, a noise reduction processing unit, a speech separation unit, a feature extraction unit, and a sentiment classification unit. In the signal processing link, analog-to-digital conversion is carried out according to the Nyquist sampling theorem, and short-time Fourier transform is utilized to respectively process noisy voice and noise, so that background interference is effectively eliminated; in the speaker separation stage, voiceprint features are extracted by means of an LPCC method, and the Euclidean distance clustering technology is combined, so that the voice of a known speaker can be accurately distinguished, and a new sounding individual can be recognized; the emotion feature extraction layer is used for mining emotion information in the speech from multiple dimensions by calculating the speech speed and analyzing intonation changes; in a sentiment classification link, the multi-layer perceptron model carries out deep processing and probability calculation on sentiment features through collaborative operation of an input layer, a hiding layer and an output layer, and it is ensured that a sentiment prediction result is scientific and reliable.
Owner:SHANDONG QIANGBI INFORMATION TECHNOLOGY CO LTD

System and method for dubbing by automatically aligning time axis

The invention provides a system and a method for automatically aligning time axis dubbing, and the system comprises a silent video extraction module which is used for removing the audio of an original video; the OCR subtitle extraction module is used for generating a subtitle file in an SRT format; the AI translation module translates the subtitles into a target language; the subtitle erasing module is used for positioning and erasing original subtitles in the video; the TTS speech synthesis module is used for generating a dubbing file; the time axis alignment module is used for dynamically adjusting the speech speed, the video speed and the subtitle time; and the video synthesis module is used for integrating the silent video, dubbing and subtitles. According to the method, OCR subtitle extraction, AI subtitle translation and generative dubbing technologies are combined, and the voice speed of the voice is dynamically adjusted to keep the voice aligned with the time axis of the video and the time axis of the subtitle, so that the quality of the translation play can be improved, the cost can be greatly reduced, the product marketing is accelerated, and great convenience and competitive advantages are brought to content creators and issuers.
Owner:SHENZHEN MAPLE LEAF INTERACTIVE TECHNOLOGY CO LTD

Automatic intelligent customer service system and method thereof

The invention relates to the field of intelligent customer service, in particular to an automatic intelligent customer service system and a method thereof, and the system comprises an image analysis unit which is used for determining whether to carry out image auxiliary analysis or not according to an image shielding reference value and an image defect reference value of a target area; the auxiliary analysis unit is used for executing image auxiliary analysis; the audio analysis unit is used for determining the category of each audio paragraph as a first-class audio paragraph or a second-class audio paragraph according to the speech speed fluctuation coefficient and the fundamental frequency mutation coefficient; the difficulty evaluation unit is used for determining difficulty coefficients of the audio paragraphs according to the number and the intensity of the first-class audio paragraphs in the target audio; the distribution processing unit is used for determining a task execution condition according to the difficulty coefficient balance degree and correspondingly determining that a task distribution mode is task distribution according to the task difficulty coefficient or the task repetition frequency; according to the invention, the voice recognition task distribution efficiency is improved.
Owner:TIANJIN UNIVERSITY OF TECHNOLOGY

Intelligent call dynamic response method integrating ASR and emotion recognition

The invention discloses an intelligent call dynamic response method and system fusing ASR (automatic voice recognition) and emotion recognition, and is applied to an interaction scene of a voice robot and a call center. The robot response strategy is dynamically adjusted by synchronously analyzing the text content (ASR) and emotional characteristics (such as intonation, speech speed and energy) of the user voice in real time. When the negative emotion is detected, a manual seat is automatically triggered to switch over or switch the pacifying verbal skill; and for the positive emotion of the high-value customer, a precision marketing module is started. The method solves the problems that an existing call center is single in response mode and cannot sense the emotion of a user, and the customer satisfaction and the service conversion rate are remarkably improved.
Owner:SHANGHAI ZHAOKUN INFORMATION TECHNOLOGY CO LTD

Voice data extraction method and system adopting artificial intelligence

The invention relates to a voice data extraction method and system adopting artificial intelligence, and relates to the technical field of voice intelligent extraction, and the method comprises the steps: monitoring and collecting video and audio data generated by video communication, and obtaining a corresponding data sequence; carrying out speech speed and volume identification on the audio sequence, obtaining characteristic parameters, and respectively configuring mouth shape and audio weight to obtain a first group of weight and a second group of weight; performing mouth shape image proportion analysis on the video sequence, and obtaining a third group of weights and a fourth group of weights in combination with vocabulary analysis of historical voice data of the user; and fusing the four groups of mouth shapes and audio weights, and performing voice content recognition extraction on the audio and video sequence to obtain final voice data. The problems that in the prior art, dependence on voice data extraction is single, the voice data extraction quality is poor, and personalized adaptation is lacked are solved.
Owner:HANGZHOU ZHILIAO INFORMATION TECH CO LTD

AI-based emotional text voice conversion method and device

The invention discloses an AI-based emotional text speech conversion method and device. The method comprises the following steps: acquiring speech segments and text records from historical data of a user; performing noise reduction processing and feature extraction according to the voice segments and the text records to obtain voice features; inputting the voice features into a pre-constructed emotional tendency model, and outputting emotional tendency and emotional intensity; according to the emotional tendency and the emotional intensity, adjusting a tone weight, a speech speed and a volume to obtain a speech parameter; extracting new voice features according to the voice parameters to perform scene emotion label matching, and determining voice adjustment parameters through a linear regression model; according to the emotional tendency and the new voice features, generating an emotional type through a pre-established emotional intention classification model, and calculating a voice parameter weight in combination with a pre-established emotional mapping table; and according to the voice adjustment parameter and the voice parameter weight, performing language synthesis to generate personalized voice. According to the method, personalized expression can be accurately generated according to the scene.
Owner:FUJIAN YUANZHI UNIVERSE CULTURE COMMUNICATION CO LTD