Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

209 results about "Speech rate" patented technology

Robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation

The invention discloses a robot anthropomorphic interaction method based on multi-modal emotion recognition and customized portrait generation. The method comprises the following steps: S1, dynamically fusing multi-modal emotions; the method comprises the following steps: S1, synchronously acquiring voice, visual and text signals through a multi-source heterogeneous sensor, capturing a user voice stream by a high-fidelity microphone array, and extracting acoustic characteristics such as intonation and speed, S2, performing cross-modal reasoning; s3, synchronously generating contents; step S4: style migration; step S5, anthropomorphic voice and expression generation; according to the method, man-machine interaction emotion is analyzed and generated by utilizing a large language model and multi-modal information fusion, the singleness of interaction emotion and the deficiency of emotional sharing ability are avoided, a strong emotion interaction characteristic is achieved, the image of the robot is obtained through a generative technology and can be migrated to any image, the limitation that a specific image is independently made is broken through, and the interaction effect of the robot is improved. The advantage that one robot can be suitable for different scenes is achieved.
Owner:JIANGSU YUNMU ZHIZAO TECH CO LTD

Spoken language pronunciation training correction system based on intelligent equipment

The invention belongs to the technical field of intelligent voice processing, and particularly relates to a spoken language pronunciation training and correcting system based on intelligent equipment, which acquires a user rhythm feature set including pitch change rate, accent intensity, pause duration and intonation contour by acquiring spoken language audio data sent by a user for a target text, and corrects the spoken language pronunciation training and correcting system. The method comprises the following steps: acquiring spoken language audio data, converting the spoken language audio data into a phoneme sequence aligned with target text time, generating a rhythm deviation degree report according to comparison evaluation of a user rhythm feature set and the phoneme sequence, an execution intonation mode, accent distribution, speech stream sound change and speech speed rhythm, and determining the rhythm deviation degree according to a rhythm problem type in the report in combination with user historical learning data. And forming and outputting a correction scheme including a text prompt, a targeted minimum contrast training unit and a listen-and-read simulation task, solving the problem of weak capability of correcting hyper-phoneme in a second language spoken language of an adult, and improving the authentic and fluency of the spoken language, thereby realizing efficient personalized learning.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Call center dialogue sentiment analysis method and system fusing voice and text

The invention provides a call center dialogue sentiment analysis method and system fusing voice and text, and relates to the technical field of voice signal processing, and the method comprises the steps: obtaining a voice signal of a call center and a corresponding transliteration text, extracting intonation, speed, sound intensity and pause sentiment features from the voice signal, meanwhile, deep language analysis is carried out on the transliterated text, and semantic emotion features related to context are extracted; and performing cross-modal correlation analysis on the voice emotion features and the semantic emotion features, and generating a time sequence correction coefficient for feature alignment by constructing a corresponding relation analysis framework between feature sequences. According to the method, language emotion information in voice and text is integrated, more accurate and comprehensive recognition of conversation emotion of the call center is realized, customer satisfaction is improved, and reliable emotion analysis basis is provided for efficiently processing customer appeals.
Owner:SHENZHEN ROADTEL DIGITAL TECH CO LTD

Voice interaction system and method for customer service based on artificial intelligence

The invention relates to the technical field of voice recognition, in particular to a voice interaction system and method for customer service based on artificial intelligence, and the system comprises a voice input processing module, an intention classification and routing module, a context dynamic adjustment module, a user behavior learning module, a multi-level intention fusion module and a final result module. According to the method, a multi-dimensional feature system is constructed by extracting tone intensity, speech speed frequency and emotional fluctuation amplitude, intention categories, priority weights and confidence scores are generated to realize accurate acquisition of appeals, and dialogue history, context and emotional change dynamic reconstruction path nodes, switching rules and response time sequences are tracked during interaction. Historical behavior mining preference features, habit fusion intention relevance, emergency calculation of an optimal strategy, construction of service steps, resource allocation schemes and execution timelines, adjustment of an interactive interface, a service process and a feedback mechanism according to multi-dimensional analysis, guarantee of differentiated service experience, and improvement of response accuracy and user satisfaction.
Owner:NANJING XIUGUO INTELLIGENT TECH CO LTD

AI agent personality switching method and device based on NFC identification

The invention relates to the technical field of artificial intelligence and Internet of Things equipment, and discloses an AI (artificial intelligence) agent personality switching method based on NFC (near field communication) identification, which realizes dynamic switching of AI agent personality in a physical mode, and comprises the following steps: S1, an equipment end obtains and decrypts an NFC label signal of an agent; s2, the equipment end obtains an AI virtual role ID corresponding to the intelligent agent according to the decryption information, and sends a request parameter to a cloud server; s3, the cloud server obtains personality configuration of the intelligent agent from a role database of the cloud server according to the request parameters and returns the personality configuration to the equipment end, wherein the personality configuration comprises an image, voice tone, speed, tone, role background, story index information and a knowledge template; and S4, the device side updates and switches the personality state corresponding to the intelligent agent according to the personality configuration, and performs voice interaction with the user according to the current personality state parameter. The invention further discloses an AI agent personality switching device for implementing the AI agent personality switching method.
Owner:SHENZHEN GIEC DIGITAL CO LTD

Anti-fraud outbound call identification method for cross-number-segment voiceprint tracking

The invention relates to an anti-fraud outbound recognition method for cross-number-segment voiceprint tracking, and the method comprises the steps: introducing a time domain voiceprint adversarial extraction mechanism, and adding an adversarial voice discriminator, an active inhibition device, a channel, and the interference of non-identity factors of language emotion in voiceprint modeling; constructing a speaker constant feature residual channel: taking stable syllable fragments including vowels and gutto vowels in call content as a reference extraction window, and filtering emotion or speech speed driving components; outputting a steady-state voiceprint contour flow as a unique basis for subsequent cross-number identity fusion; designing a dynamic number fusion graph based on the voiceprint steady-state flow; each time of voiceprint appearance is regarded as a single time point node, and all similar historical voiceprints form a serial number merging path; a number time sequence edge weight function is set; the occurrence time interval, the use frequency and the regional jump of the new number and the old number are synthesized, and whether the number belongs to the same user or not is evaluated; and establishing atlas connection between the voiceprint main body node and all number nodes thereof.
Owner:SHIJIAZHUANG LINGYUE TECHNOLOGY CO LTD

Cross-platform calling method and system for corpus training library

The invention relates to the technical field of speech synthesis, in particular to a corpus training library cross-platform calling method and system, and the method comprises the steps: obtaining the text content of a to-be-synthesized audio, disassembling the text content to obtain a pause position sequence, selecting a corresponding corpus training library according to a preset language identifier, extracting the pronunciation duration of character phonemes, and calculating the duration of basic phonemes. Text emotion features are recognized through an emotion classification model, a three-dimensional emotion feature vector containing emotion density, emotion variance and distribution entropy is constructed, and a first phoneme adjustment coefficient is determined according to the three-dimensional emotion feature vector to correct basic phoneme duration. And when a difference value exists between the corrected phoneme duration and the target audio duration, carrying out compression processing on the phonemes when the difference value is a positive number, and generating an optimal second phoneme adjustment coefficient sequence and a pause duration sequence by using a genetic algorithm when the difference value is a negative number. According to the invention, the problem of speech speed and pause duration optimization under the preset audio length is solved.
Owner:SICHUAN NORMAL UNIV

Tourism guidance system based on big data AI

The invention discloses a tourism guidance system based on big data AI, and belongs to the technical field of tourism guidance, the tourism guidance system comprises an interaction engine design module and a multi-source module information fusion positioning module, the interaction engine design module comprises: a big language model hybrid architecture; a context sensing and semantic tracking processing mechanism; according to the method, interaction naturalness is improved, context association and language naturalness of dialogues are remarkably improved through a multi-round semantic understanding mechanism based on a large language model, the large language model is used for supporting multi-round semantic interaction, slot position guiding and context maintaining, a system can understand continuous intentions of users, and the interaction naturalness is improved through a dialogue state machine and a semantic memory bank. The system can recognize the emotion of the user through tone, speed and context, adjust the explanation tone and content according to the emotional state, have the emotional common feeling ability and improve the user satisfaction degree.
Owner:HEFEI TINGTING ARTIFICIAL INTELLIGENCE APPLICATION TECHNOLOGY SERVICE CO LTD

Digital human contradictory dispute mediation method based on intelligent perception and emotion regulation

The invention discloses a digital human contradictory dispute mediation method based on intelligent perception and emotion regulation. The method comprises the following steps: firstly, acquiring voice information and visual information of two contradictory parties; calculating a voice emotion score by using the voice information, calculating a visual emotion score by using the visual information, and carrying out weighted calculation to obtain a comprehensive emotion score; then generating an initial speech speed, an initial expression and an initial gesture of the digital human; matching a proper response strategy from the knowledge base and transmitting the response strategy to the digital person; and finally, mapping voice information in the voice emotion score to a voice synthesis parameter, generating emotional voice signals consistent with both parties of a contradictory dispute in language by combining a response strategy, and outputting the emotional voice signals by a digital person for contradictory dispute mediation. In the mediation process, the digital person monitors the comprehensive emotion score and the contradictory scene in real time for dynamic adjustment. The real emotional state of the user is accurately captured based on intelligent perception, emotion mediation is carried out based on emotion perception, and the reality sense of digital human interaction and the user trust degree are improved.
Owner:GUANGDONG UNIV OF EDUCATION

Voice interaction method, server and computer readable storage medium

The invention discloses a voice interaction method, a server and a computer readable storage medium. The method comprises: determining an emotional state identifier according to acoustic features of a received voice request, the acoustic features including at least one of a tone feature, a speech speed feature and an energy feature; and according to the emotional state identifier and the voice request, generating an emotional adaptive pad call so as to complete the voice interaction. Therefore, by analyzing the acoustic features, identifying the emotional state identifier, accurately generating the emotional adaptive pad call corresponding to the voice request, and establishing emotional connection with the user, the user experience is enhanced. Moreover, by generating the emotion-adaptive pad call, a waiting blank period from the time after the user sends an instruction to the time before the system returns a formal result can be filled up, the experience of'instant response 'is created, the user is prevented from anxiety caused by too long waiting time, and the satisfaction degree and the credibility of the system are further improved.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

AI tourism personalized explanation system and service system

The invention relates to the technical field of intelligent tourism services, and particularly discloses an AI tourism personalized explanation system and service system.The personalized explanation system comprises a positioning module used for obtaining user position information; the audio acquisition module is used for acquiring environmental noise data; the attribute acquisition module is used for acquiring user attribute data; the intelligent explanation module is used for determining explanation content based on the user position information; determining a reference speed and a reference volume based on the user attribute data; respectively compensating the reference speed and the reference volume based on the environmental noise data to generate an explanation speed and an explanation volume; and explaining the explaining content based on the explaining speed and the explaining volume. According to the method, the real-time position information of the user, the environment noise data and the attribute characteristics of the user are integrated, the interpretation content is dynamically matched, the speech speed and the volume are adaptively adjusted, and the interpretation effect and the user experience are improved.
Owner:SICHUAN XINGANYAN CULTURE TECHNOLOGY CO LTD

Intelligent call dynamic response method integrating ASR and emotion recognition

The invention discloses an intelligent call dynamic response method and system fusing ASR (automatic voice recognition) and emotion recognition, and is applied to an interaction scene of a voice robot and a call center. The robot response strategy is dynamically adjusted by synchronously analyzing the text content (ASR) and emotional characteristics (such as intonation, speech speed and energy) of the user voice in real time. When the negative emotion is detected, a manual seat is automatically triggered to switch over or switch the pacifying verbal skill; and for the positive emotion of the high-value customer, a precision marketing module is started. The method solves the problems that an existing call center is single in response mode and cannot sense the emotion of a user, and the customer satisfaction and the service conversion rate are remarkably improved.
Owner:SHANGHAI ZHAOKUN INFORMATION TECHNOLOGY CO LTD

AI-based emotional text voice conversion method and device

The invention discloses an AI-based emotional text speech conversion method and device. The method comprises the following steps: acquiring speech segments and text records from historical data of a user; performing noise reduction processing and feature extraction according to the voice segments and the text records to obtain voice features; inputting the voice features into a pre-constructed emotional tendency model, and outputting emotional tendency and emotional intensity; according to the emotional tendency and the emotional intensity, adjusting a tone weight, a speech speed and a volume to obtain a speech parameter; extracting new voice features according to the voice parameters to perform scene emotion label matching, and determining voice adjustment parameters through a linear regression model; according to the emotional tendency and the new voice features, generating an emotional type through a pre-established emotional intention classification model, and calculating a voice parameter weight in combination with a pre-established emotional mapping table; and according to the voice adjustment parameter and the voice parameter weight, performing language synthesis to generate personalized voice. According to the method, personalized expression can be accurately generated according to the scene.
Owner:FUJIAN YUANZHI UNIVERSE CULTURE COMMUNICATION CO LTD

Sound source localization and identification system for electric power intelligent service and operation method of sound source localization and identification system

The invention discloses a sound source positioning identification system for electric power intelligent service and an operation method, and the system comprises an input module which is used for receiving an interaction demand of a user, and uploading the interaction demand of the user to an identification module; the identification module receives a user interaction demand, identifies and positions a user sound position based on the user interaction demand, and uploads an identification and positioning result to the processing module; one end of the processing module is connected to the recognition module, the other end of the processing module is connected to the output module, the processing module analyzes and processes the user question according to the received recognition and positioning result, and the output module answers the user question based on the analysis and processing result. According to the method, after the noise doped in the utterance spoken by the user is removed, the timbre, the speech speed, the audio frequency and the like in the statement of the user are recorded, and meanwhile, the user is captured, so that the virtual digital human can more accurately recognize and locate the user through a sound source, the interaction error is reduced, and the naturalness during interaction is improved.
Owner:GUANGXI POWER GRID CORP

Intelligent microphone starting method based on voice recognition

InactiveCN121053971ASpeech recognitionSyllableRAPID SPEECH
The invention discloses an intelligent microphone starting method based on voice recognition, and relates to the technical field of voice recognition. The method comprises the following steps: acquiring a user voice signal and surrounding environment audio data to obtain a mixed audio signal stream; identifying an abnormal waveform and determining a frequency feature vector; determining a waveform fluctuation amplitude and obtaining an enhanced waveform stable representation; recognizing syllable interval duration and analyzing interval shortening degree; analyzing the urgent speech speed characteristics to determine a time domain stretching compensation value, and adjusting the syllable interval duration to obtain a syllable sequence; identifying spectrum features of the environment background audio and determining candidate wake-up words; performing similarity matching with an emergency wake-up word template to determine an intention recognition result; and analyzing the confidence score to obtain an emergency response instruction execution signal. Through multi-level signal processing and intelligent feature extraction, high-precision voice intention recognition in a complex environment is realized, the robustness and reliability of an emergency response system are remarkably improved, and rapid help seeking of a user in a crisis scene is guaranteed.
Owner:GUANGZHOU AOYUAN ELECTRONICS CO LTD

Agricultural information handling method and system based on virtual digital human and AI

PendingCN120581002AAnimationSpeech recognitionRAPID SPEECHEngineering
The invention discloses an agricultural information handling method and system based on a virtual digital human and an AI, and relates to the technical field of agricultural information handling, and the method comprises the following steps: firstly, capturing user voice input data in real time, collecting user voice parameters, and providing original data support for subsequent analysis and processing; and summarizing the obtained voice data parameters into an analysis set, and preprocessing the collected voice data. According to the method, the voice of the user is captured in real time, the change of the voice speed is dynamically analyzed, the voice speed is intelligently evaluated by using the deep learning model, adaptive adjustment is performed for urgent voice, the sensitivity is reduced, the processing time is prolonged, and the accuracy of voice recognition is ensured. The method is especially suitable for agricultural information handling, helps farmers to obtain accurate agricultural production suggestions, reduces misrecognition, decision errors, resource waste and production loss, and improves agricultural benefits and service quality.
Owner:XIAN XINGCHEN CLOUD DATA TECH CO LTD

AI call content optimization method and device based on voiceprint recognition

The invention relates to the technical field of intelligent voice processing and communication, and discloses an AI call content optimization method and device based on voiceprint recognition. The method comprises the following steps: acquiring an original voice data stream from a user call, and separating voice speed, tone and frequency characteristics to obtain a dynamic voice characteristic set; determining a microphone frequency response deviation and a speech speed change rate according to the microphone frequency response deviation and the speech speed change rate to form a speech feature parameter set; if the parameter exceeds the threshold value, redistributing a speech speed weight to generate an adjusted speech data stream; noise reduction is carried out to obtain a pure data stream, and features are fused to generate a personalized sound effect adjustment curve; adapting the equipment difference to obtain an adaptation curve; compressing the voice data according to the adaptive curve and optimizing the transmission priority to obtain an optimized transmission data stream; and combining the transmission data stream with the adaptive curve to generate final call voice output. According to the method, conversation content dynamic adaptation and whole-process optimization are realized, and conversation quality stability and cross-scene applicability are improved.
Owner:QUANZHOU YUANZHISHI ELECTRONIC COMMERCE CO LTD

Insurance intention identification method and system

The invention relates to the technical field of intention recognition, in particular to an insurance intention recognition method and system.The insurance intention recognition method is provided with a data acquisition module, a feature pre-analysis module, a feature recognition module, a feature analysis module and an intention tendency judgment module; the feature pre-analysis module is used for marking feature customers according to the speech speed uniform tendency coefficient, the feature recognition module is used for determining the polling tendency category of the feature customers based on the comparison condition of the active polling amount of the feature customers and the total double-end polling amount, and the feature analysis module is used for determining the acquisition mode of intention tendency parameters based on the polling tendency category. And marking a high-insurance intention tendency customer through an intention tendency judgment module. According to the invention, rapid customer preliminary screening is realized according to the voice features of the customers, the feature acquisition mode is adaptively adjusted according to the intention expression mode difference of the customers, and the efficiency and reliability of the insurance intention recognition system are improved.
Owner:BAOZHONGLIAN TECHNOLOGY CO LTD

Interactive robot intelligent man-machine interaction method

The invention discloses an intelligent man-machine interaction method for an interactive robot. According to the method, the interaction accuracy and adaptability are remarkably improved through multi-dimensional technology fusion. On the intention understanding level, deep analysis of user requirements is achieved through dynamic feature weight distribution, and a strategy is generated by combining cross-domain knowledge graph embedding and emotional expression, so that response content can accurately match user core appeals and can also be naturally fused into professional knowledge and emotional temperature, and information transmission deviation is effectively reduced. According to the real-time feedback mechanism, a dynamic preference model is constructed by capturing facial expressions, voice intonation and behavior data, so that the system can adjust a content generation strategy according to instant emotional fluctuation and attention change of a user, personalized experience in an interaction process is enhanced, mechanical feeling brought by a fixed response mode is avoided, and user experience is improved. Cooperation of content generation and rhythm control is realized, so that the speed, pause and emotion expression of voice response form a natural rhythm, and the comfort level of information receiving is improved.
Owner:BEIJING HAIBAICHUAN TECH CO LTD

A voice conversion method, device, equipment and readable storage medium

The application provides a speech conversion method, device and equipment and a readable storage medium. The method comprises the following steps: obtaining speech information to be processed; based on a three-head encoder, encoding and modeling speech content, environmental noise and fundamental frequency information in the speech information to be processed respectively to obtain encoded and modeled speech information; changing the time sequence of the encoded and modeled speech information to adjust the speech speed of the encoded speech information; inputting the speech information with adjusted speech speed into a previously trained timbre conversion model corresponding to a target user to obtain target acoustic features, wherein the timbre of the target acoustic features is the same as the timbre of the target user. Thus, the timbre conversion of the speech information to be processed can be performed as required, and the conversion method is more efficient and accurate. Multi-dimensional encoding of the speech information can improve the robustness of the speech in a noisy environment. Speech speed control can make the speech more in line with user requirements.
Owner:MIGU CO LTD +1

AI-based telephone answering system

The application discloses an AI-based telephone answering system, relates to the technical field of telephone answering, and comprises a real-time voice acquisition and preprocessing module, a speech speed detection and prediction module, a self-adaptive ASR dynamic regulation module, a refined voice slicing and decoding module, a keyword detection and voice recognition module, a semantic understanding and intention analysis module and an emergency response and dispatching module; the real-time voice acquisition and preprocessing module acquires user voice data of an emergency help telephone in real time, establishes a real-time audio stream transmission channel, and rapidly preprocesses the acquired audio signal. The application rapidly and accurately captures the voice features of a user under high speech speed, dynamically adjusts the voice segmentation length of an ASR engine, makes each voice segment more clear and accurate, avoids the splitting and recognition errors of cross-segment words caused by too fast speech speed, and effectively avoids the omission of user emergency information and the risk of misjudgment in an emergency help scene.
Owner:SHANDONG ZHIQUN INFORMATION TECH CO LTD

System

To provide a system capable of effectively training a presentation in an environment close to a real world.SOLUTION: The specification processing unit 290 of the data processing device 12 in the system receives the voice, the facial expression, the gesture, and the slide material input by the user, converts the received voice data into text, evaluates the speaking speed, the volume, and the pause, analyzes the facial expression, the line of sight, and the gesture of the user from the received video data, evaluates the degree of calmness and the degree of confidence of the presentation, analyzes the slide material, evaluates the consistency of the content and the visual effect, generates feedback to the user based on these evaluation results, and transmits the feedback to the terminal of the user.SELECTED DRAWING: Figure 2
Owner:SOFTBANK GROUP CORP

Vehicle-mounted voice interaction method and device, computer readable medium and electronic equipment

The application discloses a vehicle-mounted voice interaction method and device, a computer readable medium and an electronic device. The method comprises the following steps: receiving a voice input of a target user, wherein the target user is any one of all people in a current vehicle; identifying the voice input to determine a user age, a conversation speed and a conversation habit of the target user, wherein the conversation habit is used to represent a content detail degree of the target user when conversing; performing semantic recognition on text information corresponding to the voice input to determine a conversation intention of the target user; determining corresponding interaction reply content and a playing speed thereof according to the conversation intention, the user age, the conversation speed and the conversation habit; and performing voice playing on the interaction reply content according to the playing speed. The technical scheme provided by the application can adapt to the conversation habits of different users and ensure user experience.
Owner:VOYAH AUTOMOBILE TECH CO LTD

Debt conciliation verification method based on multi-modal characteristics and electronic equipment

The embodiment of the invention relates to a debt mediation verification method based on multi-modal features and electronic equipment, and belongs to the technical field of financial credit risk assessment or fraud detection.The method comprises the steps that first and second voice subjects in voice information are recognized, debtor declaration income of second voice content is extracted, and the debtor declaration income of the second voice content is obtained; and comparing the actual income of the debtor with the declared income of the debtor to obtain a flow matching deviation, analyzing an acoustic index, carrying out matching degree detection on the current voiceprint, and obtaining a credibility score of the debtor according to a comprehensive evaluation formula. According to the embodiment of the invention, the voice interaction content of the first voice main body and the second voice main body is split into multi-modal information to obtain the corresponding voiceprint and acoustic index, and the credibility score is calculated by using the comprehensive evaluation formula in combination with the multi-modal characteristics such as the voice content credibility, the flow matching deviation, the voiceprint matching degree, the environmental noise index and the voice speed fluctuation coefficient. And finally, the real repayment capability of the debtor is restored.
Owner:SUYUAN TECHNOLOGY (HUNAN) CO LTD

Multi-language adaptive identification method based on AI

The invention discloses an AI-based multi-language adaptive recognition method, and relates to the technical field of language information processing, and the method comprises the following steps: collecting rhythm features and pause nodes in a continuous voice stream, extracting a speed change track, and generating a rhythm basic draft for representing the rhythm change trend of a voice signal in a time dimension; and performing speed change decomposition on the voice signal based on the rhythm basic draft, determining a speed increasing area and a speed slowing area, generating a time adjustment table, and recording the duration and rhythm span of each speed change section. According to the invention, through double-layer time control of a rhythm basic draft and a time adjustment table, dynamic time mapping and feature extraction continuity under speech speed change are realized; and through a dynamic rhythm adjusting mechanism, executing time extension in a speech speed increasing area, executing rhythm forward movement in a speech speed reducing area, and correcting semantic dislocation and fracture in real time, so that speech recognition keeps semantic integrity and recognition stability in a multi-language and variable-speed scene.
Owner:FUJIAN SANQINGNIAO TECH CO LTD

Visual language recognition method based on spatio-temporal local harmonic neural network and application

The application discloses a visual language recognition method based on a space-time local harmonic neural network and application, and steps of the method comprise the following steps: 1, data preprocessing; 2, constructing a visual language recognition model based on a space-time local harmonic neural network; 3, training of the network model.The method can solve the problems of the existing visual language recognition methods, such as insufficient extraction ability of time characteristics and space characteristics, single method, and not paying attention to the difference of different information, so that the method can accurately recognize the word content in the scene where the speaker posture and the speech speed frequently change, and further provides a new solution for visual language recognition.
Owner:HEFEI UNIV OF TECH +3

Speech synthesis method and device

The embodiment of the invention provides a voice synthesis method and device. A server receives a voice request from a user terminal; personal information and a service scene of the target object are obtained, the personal information comprises identity information and voice information, the identity information comprises age, region and identity feature information, and the voice information comprises speed and tone; inputting the personal information and the service scene into a pre-trained voice synthesis model to obtain prompt voice data; and sending the prompt voice data to the user terminal to play the prompt voice. Visibly, according to the personal information and the service scene information of the target object, the personalized voice meeting the user requirements can be generated according to the characteristics and the specific service scenes of different users, the naturalness of the synthesized voice is improved, machinery is reduced, and the experience of the user in the interaction process with the voice synthesis system is improved.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD

Digital population voice synchronization generation method and system based on time sequence decoupling

The embodiment of the application provides a kind of based on timing decoupling's digital population voice synchronous generation method and system, belong to digital human technical field;The method includes face detection and cutting to original video sequence, obtain standard face image sequence, the feature extraction of target audio signal is carried out, and deep audio feature sequence is obtained;Mask processing is carried out to each frame, and the image of mouth area to be driven is obtained;It is constructed as multi-channel input tensor;The mouth image that pre-trained mouth generation network outputs is synchronized with target audio signal;After replacing the mouth image to the corresponding position of original video sequence, it is combined with target audio signal, and the digital human video of mouth voice synchronization is output.The deep audio feature extraction and audio-video accurate alignment of the application improve the synchronization accuracy of mouth and voice, the timing dependence is modeled when multi-channel input, the transition of mouth sequence is smooth, and the adaptive feature fusion mechanism can automatically adapt to different phonemes and speech rate.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Method for simultaneous call interpretation using headset

The present application relates to the field of call translation. Disclosed in the present application is a method for simultaneous call interpretation using a headset, which solves the problem of severe information distortion and loss of connotation often being caused during simultaneous call interpretation if the meaning of an original speech is merely rigidly conveyed in a target language, without fully capturing and expressing the connotation and context of an original text. By simulating the tone of a speaker, the present invention can significantly improve the quality of simultaneous call interpretation, enhance the expression of emotions, increase the accuracy of information, strengthen communication effects, and improve the user experience; in terms of technical implementation, speech synthesis and emotion analysis techniques can provide strong support, such that tone simulation is more natural and real; and speech data of different accents, different speakers and different background noises and speaking speeds is collected, such that a more robust and accurate speech recognition model can be trained, thereby adapting to diverse speech inputs, and improving the generalization capability of the model in different scenarios.
Owner:VISION INTELLIGENCE CO LTD